Assembly-based QUality Assessment for Microbial Isolate Sequencing
2.5K
[2023-05-10] We switched the deployment strategy to docker-compose (see below).
AQUAMIS is a pipeline for routine assembly and quality assessment of microbial isolate sequencing experiments. It is based on snakemake and includes the following tools:
It will read untrimmed fastq data from your Illumina sequencing experiments as paired .fastq.gz-files. These are then trimmed, assembled and polished. Besides generating ready-to-use contigs, AQUAMIS will select the closest reference genome from NCBI RefSeq and produce an intuitive, detailed report on your data and assemblies to evaluate its reliability for further analyses. It relies on reference-based and reference-free measures such as coverage depth, gene content, genome completeness and contamination, assembly length and many more. Based on the experience from thousands of sequencing experiments, threshold sets for different species have been defined to detect potentially poor results.
The AQUAMIS project website is https://gitlab.com/bfr_bioinformatics/AQUAMIS. There, you can find the latest version, source code and documentation. An example AQUAMIS report can be viewed here.
To process data and write results, Docker needs a volume mapping from a host directory containing your sequence data (<path_to_data>) to the Docker container (/AQUAMIS/analysis).
Your sample list (samples.tsv) needs to be located within <path_to_data> and contain relative paths to your NGS reads in the same or another child directory.
Which Docker Tags? Gitlab-based Docker image only include the version in their tags. The images version and latest are Gitlab-based.
Advantage: Custom entrypoint script with LOCAL_USER_ID option so that results will be written with a custom user ownership. Fully functional.
(1) For easy installation please review and run:
cd <path_to_parent_directory>
bash <(curl -s https://gitlab.com/bfr_bioinformatics/AQUAMIS/-/raw/master/docker/get_aquamis_docker.sh)
You may generate a Docker-compatible sample list in your host directory (<path_to_data>/samples.tsv) by executing the create_sampleSheet.sh from the container with the following terminal commands:
host:<path_to_data>$ ls fastq/
sample1_R1.fastq sample1_R2.fastq sample2_R1.fastq sample2_R2.fastq
(2) Start-up the containers:
export LOCAL_USER_ID=$(id -u $USER)
export AQUAMIS_HOST_PATH=<path_to_data>
docker compose --file <path_to_aquamis>/docker/docker-compose.yaml up --detach
(3) Generate a sample table:
docker exec --user $LOCAL_USER_ID aquamis_app \
/AQUAMIS/scripts/create_sampleSheet.sh \
--mode ncbi \
--fastxDir /AQUAMIS/analysis/fastq \
--outDir /AQUAMIS/analysis
(4) With the following command, AQUAMIS is started within the Docker container and will process any options appended:
docker exec --user $LOCAL_USER_ID aquamis_app \
micromamba run -n aquamis \
/AQUAMIS/aquamis.py \
--docker <path_to_data> \
--working_directory /AQUAMIS/analysis \
--sample_list /AQUAMIS/analysis/samples.tsv \
--<any_other_AQUAMIS_options>
Notes on Docker images: The Docker image latest represents the latest build from the Gitlab repository. It requires additional reference databases (also provided via the setup script) as well as a set of test data files (fastq) for validation purposes via Docker volume binds. These volume binds are provided by the docker-compose setup.
Notes on Docker usage: The container path /AQUAMIS/analysis is fixed and may not be altered. Any subdirectories of <path_to_data> will be available as subdirectories under /AQUAMIS/analysis/. Our container is able to write results with the Linux user and group ID of your choice (UID and GID, respectively) to blend into your host file permission setup. With the environment variable LOCAL_USER_ID=$(id -u $USER) the UID of the currently executing user is inherited, change it according to your needs. The absolute host path mapped to the container has to be provided as the aquamis.py argument --docker <path_to_data>, too. It is used for correcting file paths in the result JSON files of each sample to match the host perspective.
For reference, please consult the /AQUAMIS/scripts/aquamis_wrapper.sh module run_docker_compose(). This module may need to be uncommmented/activated for execution in the function run_modules().
Advanced users may copy the wrapper script to your host directory (prerequisite: docker compose up, see described above) for a more convenient execution.
Adapt aquamis_wrapper.sh and the run sheet aquamis_runs.tsv according to your needs:
docker cp aquamis_app:/AQUAMIS/scripts/aquamis_wrapper.sh ./
docker cp aquamis_app:/AQUAMIS/scripts/helper_functions.sh ./
docker cp aquamis_app:/AQUAMIS/scripts/aquamis_runs.tsv ./
Which Docker Tags? All Bioconda-based Docker images have tags ending in "_bioconda". The test dataset and databases are included.
Advantage: Smaller Docker imgages.
Disadvantage: Results will be written with file ownership ROOT.
Hotfixes: The tzdata package is missing in the Busybox Build from Bioconda. Until we update the AQUAMIS Bioconda Recipe, a copy of /usr/share/zoneinfo from our build system is included in the Docker hub Bioconda images.
Nofix: The QUAST module Augustus has a read/write problem, hence the module BUSCO will not find any orthologues. This applies only to the Bioconda Docker build.
With the following command, AQUAMIS is started within the Docker container and will process any options appended:
docker run --rm \
-v <path_to_data>:/AQUAMIS/analysis \
$DOCKER_IMAGE \
aquamis \
--docker <path_to_data> \
--working_directory /AQUAMIS/analysis \
--sample_list /AQUAMIS/test_data/samples.tsv \
--threads $CORES \
--threads_sample $CORES_SAMPLE
--<any_other_AQUAMIS_options>
Content type
Image
Digest
sha256:a053f189d…
Size
1.3 GB
Last updated
about 3 years ago
docker pull bfrbioinformatics/aquamis