Analyzing gene flow in the deep biosphere via "multi-omics" integrated analysis.
config/ : configuration files
data/ : raw data goes here
resources/ : databases, references etc.
workflow/ : main directory for the snakemake workflow
data/testdata/ :
This folder contains two small test datasets:
This dataset is comprised of 25k paired-end reads from the SCAPP test data + 25k paired-end reads each from three plasmids:
CEX4 plasmid pCEX4 (LC556220.1, Enterobacter cloacae)
unnamed plasmid (NC_012780.1, [Eubacterium] eligens ATCC 27750)
plasmid pBPSE01 (NZ_KF418775.1, Burkholderia pseudomallei strain MSHR1950)
generated using randomreads.sh from bbmap.
mock
comprised of 50k reads from the minced testdata of the Aquifex aeolicus VF5
genome (generated with randomreads.sh) + 50k reads subsampled from a
synthetic mock metagenome (using seqtk).
Virsorter uses curated protein databases of viral genes which are queried with
hmmsearch or blastp using genes predicted on contigs as queries. Metrics
are then computed using sliding windows to look for enrichment of 'viral'
signals, followed by calculation of a significance score. It also includes a
step to identify circular sequences as an initial step. The author notes that:
for fragmented genomes, category 3 predictions help recover more viral sequences, but do so at the cost of increased false-positives.
MetaviralSPADES (Antipov et al 2020)
MetaviralSPADES is essentially a modified version of the MetaSPADES assembler that looks for viral contigs in the MetaSPADES assembly graph using differences in coverage. Further steps include viralVerify and viralComplete to verify contigs and check viral genome completeness respectively.
MARVEL (Amgarten et al 2018)
Marvel uses a random forest classifier to classify metagenomic bins as phage/bacteria based on features such as 1) gene density, 2) strand shifts and 3) fraction of hits to pVOG database. The classifier was trained on 1,247 phage and 1,029 bacterial genomes. The authors note that:
MARVEL has high F1 scores and accuracy for all bin lengths analyzed, but especially for bins composed of contigs 4 kbp long and longer.
VIBRANT uses machine learning (neural networks) based on annotation metrics derived from HMM searches against KEGG, PFAM AND VOG.
Content type
Image
Digest
Size
386.5 MB
Last updated
almost 6 years ago
docker pull nbisweden/deepbio_gf