Docker container for the IDBA assembler, including the IDBA-UD metagenomic assembler
604
:warning: WORK IN PROGRESS :warning:
Shotgun metagenomics offers relatively unbiased pathogen detection and characterization in a single methodological step, with the drawbacks of increased data volume and complexity, requiring expert handling and appropriate infrastructure.
The most common strategy for analysing metagenomic data is through de novo assembly followed by a combination of different tools for characterization. This approach creates longer sequences, which are more informative and can provide better insights into the structure of microbial communities. Several dedicated metagenomic assembly tools are available. These tools apply strategies that promise better handling of intragenomic and intergenomic repeats and uneven sequencing coverage when compared to traditional genome assemblers. However, there are no comprehensive benchmarks focusing on the performance of these tools.
To assess the performance and limitations of currently available de novo assembly algorithms, we propose a benchmarking of currently available and recently mantained de novo assembly tools, both traditional and dedicated for metagenomic data.
We've compiled a list of 24 different de novo assembly tools, including 5 implementing Overlap, Layout and Consensus
(OLC) assembly algorithm, 18 De Bruijn graph assemblers, of which 10 are single k-mer value assemblers and 8 implement
a multiple k-mer value approach, and a hybrid assembler using both OLC and single k-mer De Bruijn algorythms. Of these,
12 were developed explicitly to handle metagenomic data. The complete list is available as a google sheet.
To assess the performance and limitations of currently available de novo assembly algorithms, we benchmarked 11 assemblers, including traditional and dedicated metagenomic assemblers, both using OLC, and De Bruijn with single and multiple k-mer values. The creteria was based on date of last update.
The benchmarked assembly tools include 7 traditional - bcalm2, Minia, Unicycler, SPAdes, Skesa, PANDAseq and VelvetOptimizer - and 4 dedicated for metagenomic data - GATB-Minia Pipeline, MEGAHIT, MetaSPAdes, and IDBA-UD.
Docker containers for all assemblers were created, and the corresponding docker files can be found in the docker/
repository.
An assembly pipeline to perform concorrent assemblies with all tools in the benchmark was developed, in Nextflow,
and it's available in the nextflow/ directory of this repository. All resulting assemblies are filtered for a minimum
contig length of 1000bp.
The assembly performance is evaluated through mapping of the obtained assemblies to the reference sequences, with
minimap2. The assessment of the quality of the assembly is done after computation of
the percentage of mapped contigs, contiguity, average identity, and breadth of coverage. Assembly statistics, such as
number of basepairs, number of contigs and N50 were also determined for each assembler. The scripts developed for this
are available in the scripts/ directory, including the mapping commands.
Performance metrics such as average run time, max memory usage, and average data read and written, are obtained through
the Nextflow pipeline.
This assembler, publiched by Chikhi et al, 2016 in Bioinformatics, is a fast and low memory algorithm for graph compaction, consisting of three stages: careful distribution of input k-mers into buckets, parallel compaction of the buckets, and a parallel reunification step to glue together the compacted strings into unitigs. It's a traditional single k-mer value De Bruijn assembler.
cimendes/bcalm:2.2.1-1GATB-Minia is an assembly pipeline, still unpublished, that consists Bloocoo for error correction, Minia 3 for contigs assembly, which is based on the BCALM2 assembler, and the BESST for scaffolding. It was developed to extend Minia assembler to use multiple k-mer values. It was developed to extend the Minia assembler to use De Bruijn algorithm with multiple k-mer values. It's explicit for metagenomic data.
cimendes/gatb-minia-pipeline:17.09.2019-1This tool, published by Chikhi & Rizk, 2013 in *Algorithms for Molecular Biology, performs the assembly on a data structure based on unitigs produced by the BCALM software and using graph simplifications that are heavily inspired by the SPAdes assembler.
cimendes/gatb-minia-pipeline:17.09.2019-1An assembly pipeline for bacterial genomes that can do long-read assembly, hybrid assembly and short-read assembly. When asseblying Illumina-only read sets where it functions as a SPAdes-optimiser, using a De Bruijn algorithm with multiple k-mer values. It was published by Wick et al. 2017.
cimendes/unicycler:0.4.8-1MEGAHIT, published by Li et al. 2015, de novo assembler for assembling large and complex metagenomics data in a time- and cost-efficient manner. It makes use of succinct de Bruijn graph, with a a multiple k-mer size strategy. In each iteration, MEGAHIT cleans potentially erroneous edges by removing tips, merging bubbles and removing low local coverage edges,specially useful for metagenomics which suffers from non-uniform sequencing depths.
cimendes/megahit-assembler:12.08.19-1SPAdes started out as a tool aiming to resolve uneven coverage in single cell genome data, but later metaSPAdes was released, building specific metagenomic pipeline on top of SPAdes. It was published by Nurk et al. 2017, and like SPAdes, it uses multiple k-mer sizes of de Bruijn graph, starting with lowest kmer size and adding hypothetical kmers to connect graph. It's available at http://cab.spbu.ru/software/spades/ and https://github.com/ablab/spades
cimendes/metaspades:11.10.2018-1A tool aiming to resolve uneven coverage in single cell genome data through multiple k-mer sizes of De Brujin graphs. It starts with the smallest k-mer size and and adds hypotetical k-mers to connect graph.
cimendes/metaspades:11.10.2018-1This de novo sequence read assembler is based on De Bruijn graphs and uses conservative heuristics and is designed to create breaks at repeat regions in the genome, creating shorter assemblies but with greater sequence quality. It tries to obtain good contiguity by using k-mers longer than mate length and up to insert size. It was recently published by Souvorov et al. 2018 and it's available at https://github.com/ncbi/SKESA.
flowcraft/skesa:2.3.0-1This assembler, published by [Masella et al. 2012] in BMC Bioinformatics, implements an OLC algorithm to assemble genomic data. It align Illumina reads, optionally with PCR primers embedded in the sequence, and reconstruct an overlapping sequence.
cimendes/pandaseq:2.11-1This optimizing pipeline, developed by Torsten Seeman, is still unpublished but extends the original Velvet assembler by performing several assemblies with variable k-mer sizes. It searches a supplied hash value range for the optimum, estimates the expected coverage and then searches for the optimum coverage cutoff. It uses Velvet's internal mechanism for estimating insert lengths for paired end libraries. It can optimise the assemblies by either the default optimisation condition or by a user supplied one. It outputs the results to a subdirectory and records all its operations in a logfile.
cimendes/velvetoptimiser:2.2.6-1Published by Peng et al. 2012, it's a De Brujin graph assembler for assembling reads from single-cell sequencing or metagenomic sequencing technologies with uneven sequencing depths. It employs multiple depth relative thresholds to remove erroneous k-mers in both low-depth and high-depth regions. The technique of local assembly with paired-end information is used to solve the branch problem of low-depth short repeat regions. To speed up the process, an error correction step is conducted to correct reads of high-depth regions that can be aligned to high confidence contigs.
cimendes/idba:31.12.2016-3Two commercially available mock communities containing 10 microbial species (ZymoBIOMICS Microbial Community Standards) were sequences by Nicholls et al. 2019. Shotgun sequencing of the Even and Log communities was performed with the same protocol, with the exception that the Log community was sequenced individually on 2 flowcell lanes and the Even community was instead sequenced on an Illumina MiSeq using 2×151 bp (paired-end) sequencing. They are available under accession numbers:
With the M3S3 tool, a simulation sample was obtained made up of the Zymobiomics microbial community standasrd species of bacteria (n=8). The sample has the following composition (obtained with Kraken2 with the minimkaken2_V1 database):
:warning:TODO:warning:
Rick et al. 2019 proposed the use of a triple reference to assess chromosome contiguity while benchmaking long-read genomic assemblers. This measure is the longest single alignment between the assembly and the reference. Therefore, in terms of the reference length:
More information is available on Rick's Assembly Benchmark GitHub page
For each reference in the Zymos community standard, we've generated a chromosomal triple reference, available at
data/refereces/Zymos_Genomes_triple_chromosomes. Each assembly was mapped against the triple chromosome references
with minimap2, published by Li, 2018.
using the container cimendes/minimap2:2.17-1.
Because each metagenomic sample is composed by more than one genome, the contings in the assemblies are mapped against
all references and only the best result is kept with --secondary=no flag. The resulting mapping file, in
PAF format describes the approximate mapping positions between each
mapped contig and the reference.
To obtain the mapping statistics, including contiguity, assembly_stats.py script is used available in the scripts/
directory.
:warning:TODO:warning:
The results are available in the results/ directory of this repository.
:warning:TODO:warning:
Content type
Image
Digest
Size
193.7 MB
Last updated
over 6 years ago
docker pull cimendes/idba:1.1.3-1