Reference protein guided assembly short reads in gene family bins
1.1K
PADI short for Protein-guided Assembly and Diversity Indexing is the core module / tool in a pipeline for accessing diversity in complex microbial communities. It predicts number of genes within each KEGG (proteinaceous) gene family. PADI dynamically identifies regions of high diversity from contigs aligned to a guide consisting of reference protein sequences.
maxDiversity --megan </path/to/MEGAN> \
--meganLicense <path/to/MEGAN5-academic-license.txt> \
-f --outputDIR </path/to/output/dir> \
--contigs ./example/data/contigs/ \
--refseqKO ./example/refSeqProtDB/ \
--threads 20
--contigs Folder containing binned contigs according to their KEGG family / orthology (NEWBLER 2.6 (20110517_1502))
--refseqKO Reference sequences grouped by their gene families
Use pipeline to get from raw reads to end of procedure.
GAPS (ie. more gaps than are bases)
KOs where there are too few contigs are not considered
Install Docker
cp SingleCopyGene SCG
mkdir SCG/out SCG/misc
#Place MEGAN5 license file inside misc
cp MEGAN5-academic-license.txt SCG/misc/
docker run --rm \
-v `pwd`/SCG/data/konr:/data/refSeqProtDB \
-v `pwd`/SCG/data/newbler:/data/contigs \
-v `pwd`/SCG/out:/data/out \
-v `pwd`/SCG/misc/MEGAN5-academic-license.txt:/data/misc/MEGAN5-academic-license.txt \
etheleon/pass:0.1.2
Tools
Perl
plenv install-cpanmR
cpanm https://github.com/etheleon/pAss.git
The pAss pipeline requires one to provide contigs grouped by the their Ortholog groups. The scripts should be run in series from 00 to XX.
| Script Name | Description |
|---|---|
| pAss.00 | Given contigs assembled using reads binned into relavant KEGG orthologs we blast them against known reference sequences in the same categories |
| pAss.01 | Calls MEGAN to output the blastx alignment of contigs same KO Refseq sequences |
| pAss.03 | Calls the PASS::Alignment package; Details below |
| pAss.04 | Generates diagnostic plots |
| pAss.10 | Scans for maxdiversity region |
| pAss.11 | Outputs MAX diversity sequence |
| pAss.12 | Outputs as fasta MAX DIVERSITY region for each contig |
In preparation
Include a sister software which uses protein HMMs on top of sequence similarity as a method to search for distantly related sequences.
Content type
Image
Digest
Size
1 GB
Last updated
about 9 years ago
docker pull etheleon/pass