Sign inSign up

cimendes/skesa

By cimendes

•Updated about 5 years ago

docker container for skesa

Image
0

1.4K

cimendes/skesa repository overview

⁠Benchmarking of de novo metagenomic assembly software

:warning: WORK IN PROGRESS :warning:

⁠Table of Contents

⁠Introduction

Shotgun metagenomics offers relatively unbiased pathogen detection and characterization in a single methodological step, with the drawbacks of increased data volume and complexity, requiring expert handling and appropriate infrastructure.

The most common strategy for analysing metagenomic data is through de novo assembly followed by a combination of different tools for characterization. This approach creates longer sequences, which are more informative and can provide better insights into the structure of microbial communities. Several dedicated metagenomic assembly tools are available. These tools apply strategies that promise better handling of intragenomic and intergenomic repeats and uneven sequencing coverage when compared to traditional genome assemblers. However, there are no comprehensive benchmarks focusing on the performance of these tools.

To assess the performance and limitations of currently available de novo assembly algorithms, we propose a benchmarking of currently available and recently mantained de novo assembly tools, both traditional and dedicated for metagenomic data.

We've compiled a list of 24 different de novo assembly tools, including 5 implementing Overlap, Layout and Consensus
(OLC) assembly algorithm, 18 De Bruijn graph assemblers, of which 10 are single k-mer value assemblers and 8 implement a multiple k-mer value approach, and a hybrid assembler using both OLC and single k-mer De Bruijn algorythms. Of these, 12 were developed explicitly to handle metagenomic data. The complete list is available as a google sheet⁠.

⁠Methods

⁠de novo Assembly tools

To assess the performance and limitations of currently available de novo assembly algorithms, we benchmarked 11 assemblers, including traditional and dedicated metagenomic assemblers, both using OLC, and De Bruijn with single and multiple k-mer values. The creteria was based on date of last update.

The benchmarked assembly tools include 7 traditional - bcalm2, Minia, Unicycler, SPAdes, Skesa, PANDAseq and VelvetOptimizer - and 4 dedicated for metagenomic data - GATB-Minia Pipeline, MEGAHIT, MetaSPAdes, and IDBA-UD.

Docker containers for all assemblers were created, and the corresponding docker files can be found in the docker/ repository.

An assembly pipeline to perform concorrent assemblies with all tools in the benchmark was developed, in Nextflow⁠, and it's available in the nextflow/ directory of this repository. All resulting assemblies are filtered for a minimum contig length of 1000bp.

The assembly performance is evaluated through mapping of the obtained assemblies to the reference sequences, with minimap2⁠. The assessment of the quality of the assembly is done after computation of the percentage of mapped contigs, contiguity, average identity, and breadth of coverage. Assembly statistics, such as number of basepairs, number of contigs and N50 were also determined for each assembler. The scripts developed for this are available in the scripts/ directory, including the mapping commands. Performance metrics such as average run time, max memory usage, and average data read and written, are obtained through the Nextflow pipeline.

⁠bcalm2

This assembler, publiched by Chikhi et al, 2016⁠ in Bioinformatics, is a fast and low memory algorithm for graph compaction, consisting of three stages: careful distribution of input k-mers into buckets, parallel compaction of the buckets, and a parallel reunification step to glue together the compacted strings into unitigs. It's a traditional single k-mer value De Bruijn assembler.

⁠GATB-Minia Pipeline

GATB-Minia is an assembly pipeline, still unpublished, that consists Bloocoo for error correction, Minia 3 for contigs assembly, which is based on the BCALM2 assembler, and the BESST for scaffolding. It was developed to extend Minia assembler to use multiple k-mer values. It was developed to extend the Minia⁠ assembler to use De Bruijn algorithm with multiple k-mer values. It's explicit for metagenomic data.

⁠Minia

This tool, published by Chikhi & Rizk, 2013⁠ in *Algorithms for Molecular Biology, performs the assembly on a data structure based on unitigs produced by the BCALM software and using graph simplifications that are heavily inspired by the SPAdes⁠ assembler.

⁠Unicycler

An assembly pipeline for bacterial genomes that can do long-read assembly, hybrid assembly and short-read assembly. When asseblying Illumina-only read sets where it functions as a SPAdes-optimiser, using a De Bruijn algorithm with multiple k-mer values. It was published by Wick et al. 2017⁠.

⁠MEGAHIT

MEGAHIT, published by Li et al. 2015⁠, de novo assembler for assembling large and complex metagenomics data in a time- and cost-efficient manner. It makes use of succinct de Bruijn graph, with a a multiple k-mer size strategy. In each iteration, MEGAHIT cleans potentially erroneous edges by removing tips, merging bubbles and removing low local coverage edges,specially useful for metagenomics which suffers from non-uniform sequencing depths.

⁠MetaSPAdes

SPAdes⁠ started out as a tool aiming to resolve uneven coverage in single cell genome data, but later metaSPAdes was released, building specific metagenomic pipeline on top of SPAdes⁠. It was published by Nurk et al. 2017⁠, and like SPAdes, it uses multiple k-mer sizes of de Bruijn graph, starting with lowest kmer size and adding hypothetical kmers to connect graph. It's available at http://cab.spbu.ru/software/spades/⁠ and https://github.com/ablab/spades⁠

⁠SPAdes

A tool aiming to resolve uneven coverage in single cell genome data through multiple k-mer sizes of De Brujin graphs. It starts with the smallest k-mer size and and adds hypotetical k-mers to connect graph.

⁠SKESA

This de novo sequence read assembler is based on De Bruijn graphs and uses conservative heuristics and is designed to create breaks at repeat regions in the genome, creating shorter assemblies but with greater sequence quality. It tries to obtain good contiguity by using k-mers longer than mate length and up to insert size. It was recently published by Souvorov et al. 2018⁠ and it's available at https://github.com/ncbi/SKESA⁠.

⁠PANDAseq

This assembler, published by [Masella et al. 2012] in BMC Bioinformatics, implements an OLC algorithm to assemble genomic data. It align Illumina reads, optionally with PCR primers embedded in the sequence, and reconstruct an overlapping sequence.

⁠VelvetOptimier

This optimizing pipeline, developed by Torsten Seeman, is still unpublished but extends the original Velvet⁠ assembler by performing several assemblies with variable k-mer sizes. It searches a supplied hash value range for the optimum, estimates the expected coverage and then searches for the optimum coverage cutoff. It uses Velvet's internal mechanism for estimating insert lengths for paired end libraries. It can optimise the assemblies by either the default optimisation condition or by a user supplied one. It outputs the results to a subdirectory and records all its operations in a logfile.

⁠IDBA-UD

Published by Peng et al. 2012⁠, it's a De Brujin graph assembler for assembling reads from single-cell sequencing or metagenomic sequencing technologies with uneven sequencing depths. It employs multiple depth relative thresholds to remove erroneous k-mers in both low-depth and high-depth regions. The technique of local assembly with paired-end information is used to solve the branch problem of low-depth short repeat regions. To speed up the process, an error correction step is conducted to correct reads of high-depth regions that can be aligned to high confidence contigs.

⁠Metagenomic datasets
⁠Zymobiomics Community Standard

Two commercially available mock communities containing 10 microbial species (ZymoBIOMICS Microbial Community Standards) were sequences by Nicholls et al. 2019⁠. Shotgun sequencing of the Even and Log communities was performed with the same protocol, with the exception that the Log community was sequenced individually on 2 flowcell lanes and the Even community was instead sequenced on an Illumina MiSeq using 2×151 bp (paired-end) sequencing. They are available under accession numbers:

With the M3S3 tool⁠, a simulation sample was obtained made up of the Zymobiomics microbial community standasrd species of bacteria (n=8). The sample has the following composition (obtained with Kraken2 with the minimkaken2_V1 database):

zymos_zimulated_kraken2

⁠BMock12

:warning:TODO:warning:

⁠Assessing Metagenomic Assembly Success
⁠Assembly Continuity

Rick et al. 2019⁠ proposed the use of a triple reference to assess chromosome contiguity while benchmaking long-read genomic assemblers. This measure is the longest single alignment between the assembly and the reference. Therefore, in terms of the reference length:

  • Contiguity is 100% if the assembly went perfectly.
  • Contiguity slightly less than 100% (e.g. 99.99%) indicates that the assembly was complete, but some bases were lost at the start/end of the circular chromosome.
  • Contiguity more than 100% (e.g. 102%) indicates that the assembly contains duplicated sequence via a start-end overlap.
  • Much lower contiguity (e.g. 70%) indicate that the assembly was not complete, either due to fragmentation or missassembly.

More information is available on Rick's Assembly Benchmark GitHub page⁠

For each reference in the Zymos community standard, we've generated a chromosomal triple reference, available at data/refereces/Zymos_Genomes_triple_chromosomes. Each assembly was mapped against the triple chromosome references with minimap2⁠, published by Li, 2018⁠. using the container cimendes/minimap2:2.17-1. Because each metagenomic sample is composed by more than one genome, the contings in the assemblies are mapped against all references and only the best result is kept with --secondary=no flag. The resulting mapping file, in PAF format⁠ describes the approximate mapping positions between each mapped contig and the reference.

To obtain the mapping statistics, including contiguity, assembly_stats.py script is used available in the scripts/ directory.

⁠Chimera Assessment

:warning:TODO:warning:

⁠Results

⁠Zymobiomics Community Standard

The results are available in the results/ directory of this repository. :warning:TODO:warning:

⁠Mock Community
⁠Even distributed
⁠Log distributed

⁠Authors

  • Inês Mendes
    • Instituto de Microbiologia, Instituto de Medicina Molecular, Faculdade de Medicina, Universidade de Lisboa, Lisboa, Portugal;
    • University of Groningen, University Medical Center Groningen, Department of Medical Microbiology and Infection Prevention, Groningen, The Netherlands
  • Rafael Mamede
    • Instituto de Microbiologia, Instituto de Medicina Molecular, Faculdade de Medicina, Universidade de Lisboa, Lisboa, Portugal;
  • Yair Motro
    • Department of Health System Management, School of Public Health, Faculty of Health Sciences, Ben-Gurion University of the Negev, Beer-Sheva, Israel.

Tag summary

Content type

Image

Digest

Size

168.8 MB

Last updated

about 5 years ago

docker pull cimendes/skesa:2.5.0-1