Sign inSign up

pseudogene/diet_12s

By pseudogene

•Updated over 8 years ago
Archived

For internal use only (beta version)

Image
1

216

pseudogene/diet_12s repository overview

⁠diet_12s

Diet_12S is a pipeline for analysing mitochondrial 12S & species identification - Now in a Docker

⁠Running pipeline

Download and prepare a fresh mitochondrial database (once only) in the current directory:

docker run --rm -v $(pwd):/data pseudogene/diet_12s make_obidb.sh

Run the full pipeline in the current directory:

  • --forward: Forward reads file e.g. R1.fastq.
  • --reverse: Reverse reads file e.g. R2.fastq.
  • --prefix (optional): Output files prefix [default is diet]
  • --threads (optional): Number of threads/CPU to be used. [default is 1]
  • --ncbi: Version of the mitochondrial database (e.g. 88 if the file generated by make_obidb.sh is db_88.fasta)
  • --ngsfilter: File with the barcode used. See requirement of the ngsfilter⁠ tool. An example can be found in test/ngsfilter.txt:
test    sample1      acacaca  TTAGATACCCCACTATGC    TAGAACAGGCTCCTCTAG     F       @
test    sample2      acacgac  TTAGATACCCCACTATGC    TAGAACAGGCTCCTCTAG     F       @
test    sample3      tagtgca  TTAGATACCCCACTATGC    TAGAACAGGCTCCTCTAG     F       @
test    sample4      agatctc  TTAGATACCCCACTATGC    TAGAACAGGCTCCTCTAG     F       @
docker run --rm -v $(pwd):/data pseudogene/diet_12s run_pipeline.sh --forward R1.fastq --reverse R2.fastq --ncbi 88 --ngsfilter test.ngsfilter.txt

⁠Super short example

Assuming that Docker⁠ is installed on your system AND that you have forward (R1.fastq), reverse (R2.fastq) and filter (ngsfilter.txt) files in the current directory:

docker run --rm -v $(pwd):/data pseudogene/diet_12s make_obidb.sh
docker run --rm -v $(pwd):/data pseudogene/diet_12s run_pipeline.sh --forward R1.fastq --reverse R2.fastq --ncbi $(ls db_*.fasta | cut -d _ -f 2 | cut -d . -f 1) --ngsfilter ngsfilter.txt

⁠Further analysis

Few files are generated by the pipeline. Including a file called diet.tsv or <prefix>.tsv if you provided your own prefix to run_pipeline.sh. A simplify visual report can be generated with R (threshold of 1000 reads):

library(reshape2)
library(ggplot2)

tag = read.csv("diet.tsv", header=TRUE, sep="\t")
tag_melted<-melt(tag)
t<-apply(tag_melted, 1, function(row) any(as.numeric(row[3]) > 1000 ))

ggplot(tag_melted[t,], aes(variable, Taxon)) + geom_point(aes(size=value, colour=value, fill=value)) + scale_size_area() + theme_bw() + theme(axis.text.x = element_text(angle=90, vjust=0.4, hjust=1), panel.border = element_blank(), panel.grid.major = element_blank(), panel.grid.minor = element_blank(), plot.background  = element_blank()) + labs(x="Samples", y="Taxon")

Tag summary

Content type

Image

Digest

Size

207.8 MB

Last updated

over 8 years ago

docker pull pseudogene/diet_12s