Sign inSign up

nfcore/cageseq

By nfcore

•Updated over 5 years ago

CAGE-sequencing analysis pipeline with trimming, alignment and counting of CAGE tags.

Image
Machine learning & AI
Data science
0

6.0K

nfcore/cageseq repository overview

nf-core/cageseq

GitHub Actions CI Status GitHub Actions Linting Status Nextflow

install with bioconda Docker Get help on Slack

⁠Introduction

nf-core/cageseq is a bioinformatics analysis pipeline used for CAGE-seq sequencing data.

The pipeline takes raw demultiplexed fastq-files as input and includes steps for linker and artefact trimming (cutadapt⁠), rRNA removal (SortMeRNA⁠, alignment to a reference genome (STAR⁠ or bowtie1⁠) and CAGE tag counting and clustering (paraclu⁠). Additionally, several quality control steps (FastQC⁠, RSeQC⁠, MultiQC⁠) are included to allow for easy verification of the results after a run.

The pipeline is built using Nextflow⁠, a workflow tool to run tasks across multiple compute infrastructures in a very portable manner. It comes with docker containers making installation trivial and results highly reproducible.

⁠Quick Start

  1. Install nextflow⁠

  2. Install either Docker⁠ or Singularity⁠ for full pipeline reproducibility (please only use Conda⁠ as a last resort; see docs⁠)

  3. Download the pipeline and test it on a minimal dataset with a single command:

    nextflow run nf-core/cageseq -profile test,<docker/singularity/conda/institute>
    

    Please check nf-core/configs⁠ to see if a custom config file to run nf-core pipelines already exists for your Institute. If so, you can simply use -profile <institute> in your command. This will enable either docker or singularity and set the appropriate execution settings for your local compute environment.

  4. Start running your own analysis!

nextflow run nf-core/cageseq -profile <docker/singularity/conda/institute> --input '*_R1.fastq.gz' --aligner <'star'/'bowtie1'> --genome GRCh38

See usage docs⁠ for all of the available options when running the pipeline.

⁠Documentation

The nf-core/cageseq pipeline comes with documentation about the pipeline which you can read at https://nf-co.re/cageseq/usage⁠ and https://nf-co.re/cageseq/output⁠ or find in the docs/ directory⁠.

⁠Credits

nf-core/cageseq was originally written by Kevin Menden (@KevinMenden⁠) and Tristan Kast (@TrisKast⁠) and updated by Matthias Hörtenhuber (@mashehu⁠).

⁠Contributions and Support

If you would like to contribute to this pipeline, please see the contributing guidelines⁠.

For further information or help, don't hesitate to get in touch on the Slack #cageseq channel⁠ (you can join with this invite⁠).

⁠Citation

You can cite the nf-core publication as follows:

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x⁠. ReadCube: Full Access Link⁠

Tag summary

Content type

Image

Digest

Size

789 MB

Last updated

over 5 years ago

docker pull nfcore/cageseq