A Snakemake pipeline for the analysis of messenger RNA-seq data. It processes mRNA-seq fastq files and delivers both raw and normalised/scaled count tables. This pipeline also outputs a QC report per fastq file and a .bam mapping file to use with a genome browser for instance.
This pipeline can process single or paired-end data and is mostly suited for Illumina sequencing data.
This pipeline analyses the raw RNA-seq data and produces two files containing the raw and normalized counts.
fastp.STAR.subread featureCounts.DESeq2 median of ratios method to generate the scaled ("normalized") counts.config/samples.tsv file. Specify a sample name (e.g. "Sample_A") in the sample column and the paths to the forward read (fq1) and to the reverse read (fq2). If you have single-end reads, leave the fq2 column empty.gffread my.gff3 -T -o my.gtf. :warning: for featureCounts to work, the feature in the GTF file should be exon while the meta-feature has to be transcript_id.Below is an example of a GTF file format. :warning: a real GTF file does not have column names (seqname, source, etc.). Remove all non-data rows.
| seqname | source | feature | start | end | score | strand | frame | attributes |
|---|---|---|---|---|---|---|---|---|
| SL4.0ch01 | maker_ITAG | CDS | 279 | 743 | . | + | 0 | transcript_id "Solyc01g004000.1.1"; gene_id "gene:Solyc01g004000.1"; gene_name "Solyc01g004000.1"; |
| SL4.0ch01 | maker_ITAG | exon | 1173 | 1616 | . | + | . | transcript_id "Solyc01g004002.1.1"; gene_id "gene:Solyc01g004002.1"; gene_name "Solyc01g004002.1"; |
| SL4.0ch01 | maker_ITAG | exon | 3793 | 3971 | . | + | . | transcript_id "Solyc01g004002.1.1"; gene_id "gene:Solyc01g004002.1"; gene_name "Solyc01g004002.1"; |
raw_counts.txt: this table can be used to perform a differential gene expression analysis with DESeq2.scaled_counts.tsv: this table can be used to perform an Exploratory Data Analysis with a PCA, heatmaps, sample clustering, etc.wget and scp commands).Snakefile: a master file that contains the desired outputs and the rules to generate them from the input files.config/samples.tsv: a file containing sample names and the paths to the forward and eventually reverse reads (if paired-end). This file has to be adapted to your sample names before running the pipeline.config/config.yaml: the configuration files making the Snakefile adaptable to any input files, genome and parameter for the rules.config/refs/: a folder containing
S_lycopersicum_chromosomes.4.00.chrom1.fa is placed for testing purposes.ITAG4.0_gene_models.sub.gtf for testing purposes..fastq/: a (hidden) folder containing subsetted paired-end fastq files used to test locally the pipeline. Generated using Seqtk:
seqtk sample -s100 <inputfile> 250000 > <output file>
This folder should contain the fastq of the paired-end RNA-seq data, you want to run.envs/: a folder containing the environments needed for the pipeline:
environment.yaml is used by the conda package manager to create a working environment (see below).Dockerfile is a Docker file used to build the docker image by refering to the environment.yaml (see below).You will need a local copy of the GitHub snakemake_rnaseq repository on your machine. You can either:
git clone [email protected]:BleekerLab/snakemake_rnaseq.git.download.snakemake_rnaseq folder using Shell commands.You'll need to change a few things to accomodate this pipeline to your needs. Make sure you have changed the parameters in the config/config.yaml file that specifies where to find the sample data file, the genomic and transcriptomic reference fasta files to use and the parameters for certains rules etc.
This file is used so the Snakefile does not need to be changed when locations or parameters need to be changed.
Using the conda package manager, you need to create an environment where core softwares such as Snakemake will be installed.
rnaseq using the envs/environment.yaml file with the following command: conda env create --name rnaseq --file envs/environment.yamlsource activate rnaseq.While a conda environment will in most cases work just fine, Docker is the recommended solution as it increases pipeline execution reproducibility.
:round_pushpin: Option 2: using a Docker container
docker pull bleekerlab/snakemake_rnaseq:4.7.12 to retrieve a Docker image that includes the pipeline required softwares (Snakemake and conda and many others).docker run --rm -v $PWD:/home/snakemake/ bleekerlab/snakemake_rnaseq:4.7.12 and add any options for snakemake (-n, --cores 10) etc.
The image was built using a Dockerfile based on the 4.7.12 Miniconda3 official Docker image.singularity run docker://bleekerlab/snakemake_rnaseq:4.7.12 to retrieve a Docker image that includes the pipeline required software (Snakemake and conda and many others).singularity run snakemake_rnaseq_4.7.12.sif and add any options for snakemake (-n, --cores 10) etc.
The directory where the sif file is stored will automatically be mapped to /home/snakemake. Results will be written to a folder named $PWD/results/ (you can change results to something you like in the result_dir parameter of the config.yaml).snakemake -np to perform a dry run that prints out the rules and commands.docker run With conda: snakemake --cores 10
You will need a local copy of the GitHub snakemake_rnaseq repository on your machine. On a HPC system, you will have to clone it using the Shell command-line: git clone [email protected]:BleekerLab/snakemake_rnaseq.git.
download.snakemake_rnaseq folder using Shell commands.See the detailed protocol here.

Johannes Köster; creator of Snakemake.
Content type
Image
Digest
Size
790.4 MB
Last updated
over 5 years ago
docker pull bleekerlab/snakemake_rnaseq