Sign inSign up

engbio/pannagram

By engbio

•Updated about 1 year ago

Image
0

70

engbio/pannagram repository overview

Pannagram is a command-line toolkit for the construction of unbiased pangenome alignments and the detection of the Mobilome from complete genome assemblies. The software suite consists of

Bash pipelines for the core analysis and an R library for in-depth analysis and visualization.

This repository provides a self-contained, automated environment using Docker to run the complete Pannagram analysis pipeline. It takes raw genome FASTA files and produces a comprehensive pangenome alignment, from which it extracts synteny blocks, SNPs, and structural variants (SVs).

Official Pannagram Page: https://github.com/iganna/pannagram⁠

šŸš€ Quick Use with Docker

This guide provides the steps to run the complete pangenome analysis using our automated pipeline.

Prerequisites:

#Docker must be installed on your system.

#You must have a folder containing your genome assembly files (.fasta or .fa).

Step 1: Pull the Pannagram Image

Download the pre-built image from Docker Hub. This contains Pannagram, all its dependencies, and the automated pipeline script, fully configured.

docker pull engbio/pannagram:v1.0

Step 2: Verify the Image

Check if the image was successfully downloaded to your machine.

docker images

You should see engbio/pannagram with the tag v1.0 in the list.

Step 3: Prepare Your Data Folders The container is designed to work with a specific folder structure.

Create a main project directory (e.g., my_pannagram_project).

Inside it, create a subdirectory for your genomes: genomes.

Also create an empty subdirectory for the results: results.

Your structure should look like this:

my_pannagram_project/
ā”œā”€ā”€ genomes/
│   ā”œā”€ā”€ genome1.fasta
│   ā”œā”€ā”€ genome2.fa
│   └── ...
└── results/

Step 4: Run the Automated Pangenome Analysis

Navigate your terminal to your main project directory (my_pannagram_project) and run the container. Bash

    docker run --rm \
        -v "$(pwd)/genomes:/data/input" \
        -v "$(pwd)/results:/data/output" \
        -e NUM_CORES=8 \
        engbio/pannagram:v1.0

This single command will:

Start the container.

Mount your local genomes folder for reading and your results folder for writing.

Set the number of CPU cores to use (e.g., -e NUM_CORES=8).

Automatically execute the internal pipeline script, which performs the full analysis:

Runs pannagram to build the pangenome alignment in reference-free mode.

Runs features to extract synteny blocks, consensus sequences, SNPs, and structural variants from the alignment.

Step 5: Fix Output File Permissions The results generated in your local results folder will be owned by root. To reclaim ownership, run this command after the analysis is complete: Bash

    sudo chown -R $(id -u):$(id -g) results/

āš™ļø For Advanced Users: Configuration Files These are the files used to build the engbio/pannagram:v1.0 image and automate the pipeline.

Dockerfile

This Dockerfile creates the complete environment, installs Pannagram, and sets the internal pipeline script as the main command.

# Dockerfile (Final Version for Automated Pipeline)

FROM mambaorg/micromamba:1.5.8

ARG GIT_REPO="https://github.com/iganna/pannagram.git"
ARG GIT_BRANCH="main"

# Use ROOT to ensure full permissions during build
USER root

# 1. Install system dependencies for compilation
RUN apt-get update && apt-get install -y git build-essential && apt-get clean && rm -rf /var/lib/apt/lists/*

# 2. Clone the repository
RUN git clone --branch ${GIT_BRANCH} ${GIT_REPO} /opt/pannagram

WORKDIR /opt/pannagram

# 3. Create the conda environment with all dependencies
RUN micromamba env create -f pannagram.yml --yes && \
    micromamba clean --all --yes

# 4. Set the SHELL to run subsequent commands inside the conda environment
SHELL ["micromamba", "run", "-n", "pannagram", "/bin/bash", "-c"]

# 5. The crucial step: Build and install the Pannagram R package and scripts
RUN bash build.sh

# 6. Revert to the default shell
SHELL ["/bin/bash", "-c"]

# 7. Add the environment's bin directory to the PATH for the final image
ENV PATH /opt/conda/envs/pannagram/bin:$PATH

# 8. Copy the automation script and set it as the main command
COPY internal_pipeline.sh /usr/local/bin/internal_pipeline.sh
RUN chmod +x /usr/local/bin/internal_pipeline.sh
ENTRYPOINT ["/usr/local/bin/internal_pipeline.sh"]

Script to Automate Pipeline (internal_pipeline.sh)

This is the Bash script included in the image that orchestrates the entire two-stage analysis.

#!/bin/bash
set -e

# Configure variables from environment variables or use default values
PATH_IN="${INPUT_DIR:-/data/input}"
PATH_OUT="${OUTPUT_DIR:-/data/output}"
CORES="${NUM_CORES:-8}"

# --- Print Execution Status ---
echo "--- Starting Pannagram Automated Pipeline (Reference-Free Mode) ---"
echo "Input Directory: ${PATH_IN}"
echo "Output Directory: ${PATH_OUT}"
echo "Cores: ${CORES}"
echo "-----------------------------------------------------------"

# --- Step 1: Run Pangenome Alignment ---
echo "STEP 1: Running pangenome alignment with 'pannagram'..."
pannagram \
    -path_in "${PATH_IN}" \
    -path_out "${PATH_OUT}/alignment_results" \
    -cores "${CORES}"

echo "Alignment complete."
echo "-----------------------------------------------------------"

# --- Step 2: Run Feature Extraction ---
echo "STEP 2: Running feature extraction with 'features'..."
features \
    -path_in "${PATH_OUT}/alignment_results" \
    -blocks \
    -seq \
    -snp \
    -sv_call \
    -cores "${CORES}"

echo "Feature extraction complete."
echo "--- Pannagram Pipeline Finished Successfully! ---"

Tag summary

Content type

Image

Digest

sha256:7217dc53d…

Size

1 GB

Last updated

about 1 year ago

docker pull engbio/pannagram:v1.0