identify species and inter-species hybrids and chromosome copy number variants from short-read data
1.8K
sppIDer is a pipeline for looking at genome composition in hybrid genomes and checking for chromosomal copy variants in single species strains.
sppIDer.py is the main wrapper that calls established bioinformatic tools and custom scripts. This pipeline needs a combination reference genome and one or more short read (fastq) files.
The sppIDer docker image is a self-contained platform capable of executing its pipeline without requiring cumbersome managment and installation of prerequisite tools.
Changes to this source repo are automatically built into an updated docker image, available from docker hub at glbrc/sppider.
Additional detailed usage information is available in the sppIDer manual.
docker run --rm -it glbrc/sppider [pipeline_script] --help
pipeline scripts:
sppIDer.py
mitoSppIDer.py
combineRefGenomes.py
docker run --rm -it glbrc/sppider sppIDer.py -h
usage: sppIDer.py [-h] --out OUT --ref REF --r1 R1 [--r2 R2] [--byBP]
[--byGroup]
Run full sppIDer
optional arguments:
-h, --help show this help message and exit
--out OUT Output prefix, required
--ref REF Reference Genome, required
--r1 R1 Read1, required
--r2 R2 Read2, optional
--byBP Calculate coverage by basepair, optional, DEFAULT, can't be used
with -byGroup
--byGroup Calculate coverage by chunks of same coverage, optional, can't
be used with -byBP
Workflow:
Notes:
docker run \
--rm -it \
--mount type=bind,src=$(pwd),target=/tmp/sppIDer/working \
--user "$UID:$(id -g $USERNAME)" \
glbrc/sppider \
combineRefGenomes.py
--out REF.fasta \
--key KEY.txt
An optional --trim can be used to trim short uninformative contigs for reference genomes with many short contigs. All contigs shorter than the supplied interger will be ignored. The KEY.txt file must be tab delimited and the reference genome unique name cannot contain hyphens. See example.
docker run \
--rm -it \
--mount type=bind,src=$(pwd),target=/tmp/sppIDer/working \
--user "$UID:$(id -g $USERNAME)" \
glbrc/sppider \
sppIDer.py \
--out OUT \
--ref REF.fasta \
--r1 R1.fastq \
--r2 R2.fastq
An optional --byGroup flag can be used for very large combination genomes. This produce a bedfile that doesn't have coverage information for each basepair but by groups. Which speeds up the run.
docker run \
--rm -it \
--mount type=bind,src=$(pwd),target=/tmp/sppIDer/working \
--user "$UID:$(id -g $USERNAME)" \
glbrc/sppider \
combineGFF.py
--out REF.gff \
--key GFF_KEY.txt
docker run \
--rm -it \
--mount type=bind,src=$(pwd),target=/tmp/sppIDer/working \
--user "$UID:$(id -g $USERNAME)" \
glbrc/sppider \
mitoSppIDer.py \
--out OUT \
--ref MITO_REF.fasta \
--r1 R1.fastq \
--r2 R2.fastq
An optional --gff can be used if you are providing a combined gff of the regions that should be marked on the final plots.
This pipeline has been tested CentOS 7.5 (1804) running Docker Community Edition (CE) Stable.
Content type
Image
Digest
Size
653.3 MB
Last updated
about 8 years ago
docker pull glbrc/sppider