Sign inSign up

gynecoloji/chipseq_pipeline

By gynecoloji

•Updated 9 months ago

A comprehensive, production-ready Snakemake workflow for ChIP-seq data analysis

Image
Data science
0

146

gynecoloji/chipseq_pipeline repository overview

⁠ChIP-seq Analysis Pipeline

A comprehensive, production-ready Snakemake workflow for ChIP-seq data analysis, fully containerized and optimized for both local Docker execution and HPC deployment via Singularity.

⁠🚀 Quick Start

Pull the image:

docker pull gynecoloji/chipseq_pipeline:v1.0

Run the pipeline:

docker run --rm \
    -v $(pwd)/data:/pipeline/data:ro \
    -v $(pwd)/ref:/pipeline/ref:ro \
    -v $(pwd)/results:/pipeline/results \
    -v $(pwd)/logs:/pipeline/logs \
    gynecoloji/chipseq_pipeline:v1.0 \
    --cores 8 --use-conda -s /pipeline/snakefile_tolerant_ChIPseq -p

For HPC (Singularity):

singularity pull chipseq_pipeline.sif docker://gynecoloji/chipseq_pipeline:v1.0
singularity run --bind $(pwd)/data:/pipeline/data chipseq_pipeline.sif --cores 8 --use-conda -s /pipeline/snakefile_tolerant_ChIPseq -p

⁠✨ Features

  • Complete ChIP-seq workflow: FastQC → Fastp → HISAT2 → Filtering → MACS2 → BigWig
  • Flexible peak calling: Supports both narrow (TFs) and broad (histone marks) modes
  • Input control support: Optional normalization for accurate peak calling
  • Quality control: Comprehensive QC at multiple pipeline steps
  • Blacklist filtering: Removes reads from ENCODE blacklist regions
  • PCR duplicate removal: Picard-based deduplication
  • Visualization ready: Generates BigWig tracks for genome browsers
  • HPC compatible: Fully tested with Singularity on SLURM clusters

⁠🔧 Included Tools

ToolVersionPurpose
FastQC0.12.1Quality control
Fastp0.24.1Read trimming
HISAT22.2.1Alignment
Samtools1.21BAM processing
Picard3.1.1Duplicate removal
BEDTools2.31.1Blacklist filtering
MACS22.2.7.1Peak calling
deepTools3.5.6BigWig generation
Snakemake9.3.2Workflow management

⁠📋 Requirements

Input files:

  • Paired-end FASTQ files (.fastq.gz)
  • Sample metadata CSV (samples.csv)
  • HISAT2 genome index
  • ENCODE blacklist BED file
  • Configuration file (config.yaml)

System requirements:

  • 16GB+ RAM recommended
  • 8+ CPU cores for optimal performance
  • 50GB+ storage for intermediate files

⁠📂 Directory Structure

Mount these directories when running:

/pipeline/data     - Input FASTQ files
/pipeline/ref      - Reference files (index, blacklist, config)
/pipeline/results  - Output directory
/pipeline/logs     - Log files

⁠🎯 Use Cases

  1. Transcription Factor ChIP-seq - Narrow peak calling mode
  2. Histone Modification ChIP-seq - Broad peak calling mode
  3. With/without input controls - Flexible normalization
  4. Multi-sample comparisons - Batch processing support

⁠📖 Documentation

Full documentation: https://github.com/gynecoloji/Snakemake_ChIP_seq_containerization⁠

Includes:

  • Detailed installation instructions
  • Sample configuration examples
  • SLURM job submission scripts
  • Troubleshooting guide
  • Advanced usage options

⁠💡 Example Workflow

  1. Prepare your data and reference files
  2. Create samples.csv with sample metadata
  3. Configure ref/config.yaml with paths
  4. Run pipeline with desired number of cores
  5. Find results in organized output directories

⁠🔬 Output

The pipeline generates:

  • Quality control reports (FastQC, Fastp)
  • Alignment statistics
  • Deduplicated BAM files
  • Peak calls (narrowPeak/broadPeak formats)
  • RPGC-normalized BigWig tracks
  • Input-normalized BigWig files (when controls available)
  • Comprehensive QC summary

⁠🏆 Highlights

  • Production-ready: Extensively tested on real ChIP-seq datasets
  • Reproducible: Containerized for consistency across systems
  • HPC-optimized: Designed for high-performance computing environments
  • Well-documented: Comprehensive README and inline documentation
  • Actively maintained: Regular updates and bug fixes

⁠📞 Support

⁠📜 License

MIT License - Free for academic and commercial use

⁠🙏 Acknowledgments

Built with tools from the Bioconda community and ENCODE Consortium resources.


Tags: chipseq bioinformatics snakemake peak-calling macs2 hisat2 genomics ngs hpc singularity

Image Size: ~4 GB (includes all bioinformatics tools and conda environments)

Last Updated: January 2026

Tag summary

Content type

Image

Digest

sha256:b59a8922c…

Size

3.8 GB

Last updated

9 months ago

docker pull gynecoloji/chipseq_pipeline:v1.0