docker container for megahit metagenomic assembler
573
While searching for a benchmark with relevant metagenomic assembly software, I gave up and decided to make my own. WORK IN PROGRESS
TODO
TODO
With the M3S3 tool, a simulation sample was obtained made up of the Zymobiomics microbial community standasrd species of bacteria and yeast. The sample has the following composition (obtained with Kraken2 with the minimkaken2_V1 database):
Two commercially available mock communities containing 10 microbial species (ZymoBIOMICS Microbial Community Standards) were sequences by Nicholls et al. 2019. Shotgun sequencing of the Even and Log communities was performed with the same protocol, with the exception that the Log community was sequenced individually on 2 flowcell lanes and the Even community was instead sequenced on an Illumina MiSeq using 2×151 bp (paired-end) sequencing. They are available under accession numbers ERR2984773 (even) and ERR2935805.
As a real metagenomic dataset, the three CAMI synthetic metagenome challenge sets was used (Low (L), Medium with differential abundance (M), and High complexity with time series (H)). These datasets are publicly available in GigaDB.
I've followed the following metagenomic assembly tools table, published by Ayling et al. 2019, as base for selecting metagenomic software to bee tested. In total 17 tools are presented. IVA and SAVAGE were removed as they were aimed at viruses, as well as Genovo and MAP due to inavilability of software, VICUNA due to requiring registeration, Omega as it was an assembly pipeline, abd BBAP, MetaVelvet, MEtaVelvet-SL, PRICE and Ray Meta due to no update sinde 2016. The following tools will be tested:
Published by Peng et al. 2012, it's a De Brujin graph assembler for assembling reads from single-cell sequencing or metagenomic sequencing technologies with uneven sequencing depths. It employs multiple depthrelative thresholds to remove erroneous k-mers in both low-depth and high-depth regions. The technique of local assembly with paired-end information is used to solve the branch problem of low-depth short repeat regions. To speed up the process, an error correction step is conducted to correct reads of high-depth regions that can be aligned to highconfident contigs. The latest version is available at https://github.com/loneknightpy/idba, and an official docker image at https://hub.docker.com/r/loneknightpy/idba Last update: 31/12/2016 (GitHub)
Published by Le et al. 2017, MegaGTA is a gene-targeted assembler that utilizes iterative de Bruijn graphs. It tries to improve on Xander assembler, implementing the same method of using the trained Hidden Markov Model (HMM) to guide the traversal of de Bruijn graph, but using mutiple k-mer sizes to take full advantage of multiple k-mer sizes to make the best of both sensitivity and accuracy. The latest version is available at https://github.com/HKU-BAL/MegaGTA. Last update: 16/06/2016 (GitHub)
MEGAHIT, published by Li et al. 2015, de novo assembler for assembling large and complex metagenomics data in a time- and cost-efficient manner. It makes use of succinct de Bruijn graph, with a a multiple k-mer size strategy. In each iteration, MEGAHIT cleans potentially erroneous edges by removing tips, merging bubbles and removing low local coverage edges,specially useful for metagenomics which suffers from non-uniform sequencing depths. The latest version is available at https://github.com/voutcn/megahit Last update: 12/08/2019 (GitHub)
Published by Gregot et al. 2016, Snowball is a novel strain aware gene assembler for shotgun metagenomic data that does not require closely related reference genomes to be available. Like MegaGTA and Xander, it uses profile hidden Markov models (HMMs) of gene domains of interest to guide the assembly. The latest version is available at https://github.com/hzi-bifo/snowball Last update: 24/09/2017 (GitHub)
SPAdes started out as a tool aiming to resolve uneven coverage in single cell genome data, but later metaSPAdes was released, building specific metagenomic pipeline on top of SPAdes. IT was published by Nurk et al. 2017, it uses multiple k-mer sizes of de Bruijn graph, starting with lowest kmer size and adding hypothetical kmers to connect graph. It's available at http://cab.spbu.ru/software/spades/ and https://github.com/ablab/spades Last update: 11/10/2018 (release) and 24/04/2019 (GitHub)
Like Snowball and MegaGTA, Xander employs HMM profile model to perform a guided assebly targeting specific genes. These are used to create a novel combined weighted assembly graph. Xander performs both assembly and annotation concomitantly using information incorporated in this graph. It was published by Wang et al. 2015 and it's available at https://github.com/rdpstaff/Xander_assembler. Last update: 27/10/2017 (GitHub)
Content type
Image
Digest
Size
174.9 MB
Last updated
about 6 years ago
docker pull cimendes/megahit-assembler:1.2.9-1