Sign inSign up

etheleon/pass

By etheleon

•Updated over 8 years ago

Reference protein guided assembly short reads in gene family bins

Image
0

1.1K

etheleon/pass repository overview

⁠PADI - Protein-guided Assembly and Diversity Indexing

Build Status DockerImage Image description

⁠DOI:

DOI

⁠Table of Contents

⁠Build

⁠Description

PADI short for Protein-guided Assembly and Diversity Indexing is the core module / tool in a pipeline for accessing diversity in complex microbial communities. It predicts number of genes within each KEGG (proteinaceous) gene family. PADI dynamically identifies regions of high diversity from contigs aligned to a guide consisting of reference protein sequences.

⁠Usage

maxDiversity    --megan </path/to/MEGAN> \
                --meganLicense <path/to/MEGAN5-academic-license.txt> \
                -f --outputDIR </path/to/output/dir> \
                --contigs ./example/data/contigs/ \
                --refseqKO ./example/refSeqProtDB/ \
                --threads 20

--contigs Folder containing binned contigs according to their KEGG family / orthology (NEWBLER 2.6 (20110517_1502))
--refseqKO Reference sequences grouped by their gene families 

Use pipeline⁠ to get from raw reads to end of procedure.

⁠Procedure
  1. BlastX of nucleotide contigs against reference prot sequences
  2. Assign NTcontigs to best matched prot REFSEQ2.
  3. MUSCLE generates MSA of reference prot sequences
  4. Align NTcontigs to global REFSEQ MSA
  5. Search for MAX diversity region
⁠Filters
  • GAPS (ie. more gaps than are bases)

    • When the the gap ratio (NT:GAP) exceeds 1:10
  • KOs where there are too few contigs are not considered

    • this or may not be IDEAL cause we’re looking for cases where there’s a large amt of RNA reads mapped to a low number o contigs

⁠Installation

Install⁠ Docker

⁠Example: SINGLE COPY GENES
cp SingleCopyGene SCG
mkdir SCG/out SCG/misc
#Place MEGAN5 license file inside misc
cp MEGAN5-academic-license.txt SCG/misc/

docker run --rm \
    -v `pwd`/SCG/data/konr:/data/refSeqProtDB \
    -v `pwd`/SCG/data/newbler:/data/contigs \
    -v `pwd`/SCG/out:/data/out \
    -v `pwd`/SCG/misc/MEGAN5-academic-license.txt:/data/misc/MEGAN5-academic-license.txt \
    etheleon/pass:0.1.2
    
⁠Installation from source
⁠1. Dependencies/Pre-requisites
  • Tools

  • Perl

    • v5.2X.X is required:
      • Use tokuhirom/plenv⁠ to install a local version of the latest perl (>=5.21.0)
      • Install CPANMINUS, using plenv install-cpanm
  • R

    • ggplot2
    • dplyr
    • Biostrings (from Bioconductor)
⁠2. Install
cpanm https://github.com/etheleon/pAss.git

⁠Description of pipeline

The pAss pipeline requires one to provide contigs grouped by the their Ortholog groups. The scripts should be run in series from 00 to XX.

Script NameDescription
pAss.00Given contigs assembled using reads binned into relavant KEGG orthologs we blast them against known reference sequences in the same categories
pAss.01Calls MEGAN to output the blastx alignment of contigs same KO Refseq sequences
pAss.03Calls the PASS::Alignment package; Details below
pAss.04Generates diagnostic plots
pAss.10Scans for maxdiversity region
pAss.11Outputs MAX diversity sequence
pAss.12Outputs as fasta MAX DIVERSITY region for each contig

⁠Publication

In preparation

⁠Future

Include a sister software which uses protein HMMs on top of sequence similarity as a method to search for distantly related sequences.

Tag summary

Content type

Image

Digest

Size

1 GB

Last updated

about 9 years ago

docker pull etheleon/pass