Sign inSign up

nbisweden/deepbio_gf

By nbisweden

•Updated almost 6 years ago

Image
0

1.7K

nbisweden/deepbio_gf repository overview

CI

⁠Deep Biosphere Geneflow project

Analyzing gene flow in the deep biosphere via "multi-omics" integrated analysis.

⁠Repository organization

config/ : configuration files

data/ : raw data goes here

resources/ : databases, references etc.

workflow/ : main directory for the snakemake workflow

⁠Data locations

data/testdata/ :

This folder contains two small test datasets:

  • test

This dataset is comprised of 25k paired-end reads from the SCAPP⁠ test data + 25k paired-end reads each from three plasmids:

  • CEX4 plasmid pCEX4 (LC556220.1, Enterobacter cloacae)

  • unnamed plasmid (NC_012780.1, [Eubacterium] eligens ATCC 27750)

  • plasmid pBPSE01 (NZ_KF418775.1, Burkholderia pseudomallei strain MSHR1950) generated using randomreads.sh from bbmap.

  • mock

comprised of 50k reads from the minced testdata of the Aquifex aeolicus VF5 genome (generated with randomreads.sh) + 50k reads subsampled from a synthetic mock⁠ metagenome (using seqtk).

⁠Tools and outline

⁠Plasmids
⁠Phages
  • Virsorter⁠ (Roux et al 2015⁠)

    Virsorter uses curated protein databases of viral genes which are queried with hmmsearch or blastp using genes predicted on contigs as queries. Metrics are then computed using sliding windows to look for enrichment of 'viral' signals, followed by calculation of a significance score. It also includes a step to identify circular sequences as an initial step. The author notes that:

    for fragmented genomes, category 3 predictions help recover more viral sequences, but do so at the cost of increased false-positives.

  • MetaviralSPADES⁠ (Antipov et al 2020⁠)

    MetaviralSPADES is essentially a modified version of the MetaSPADES assembler that looks for viral contigs in the MetaSPADES assembly graph using differences in coverage. Further steps include viralVerify and viralComplete to verify contigs and check viral genome completeness respectively.

  • MARVEL⁠ (Amgarten et al 2018⁠)

    Marvel uses a random forest classifier to classify metagenomic bins as phage/bacteria based on features such as 1) gene density, 2) strand shifts and 3) fraction of hits to pVOG database⁠. The classifier was trained on 1,247 phage and 1,029 bacterial genomes. The authors note that:

    MARVEL has high F1 scores and accuracy for all bin lengths analyzed, but especially for bins composed of contigs 4 kbp long and longer.

  • VIBRANT⁠ (Kieft et al 2020⁠)

    VIBRANT uses machine learning (neural networks) based on annotation metrics derived from HMM searches against KEGG, PFAM AND VOG⁠.

⁠CRISPR
⁠Genome Islands

Tag summary

Content type

Image

Digest

Size

386.5 MB

Last updated

almost 6 years ago

docker pull nbisweden/deepbio_gf