Pannagram is a command-line toolkit for the construction of unbiased pangenome alignments and the detection of the Mobilome from complete genome assemblies. The software suite consists of
Bash pipelines for the core analysis and an R library for in-depth analysis and visualization.
This repository provides a self-contained, automated environment using Docker to run the complete Pannagram analysis pipeline. It takes raw genome FASTA files and produces a comprehensive pangenome alignment, from which it extracts synteny blocks, SNPs, and structural variants (SVs).
Official Pannagram Page: https://github.com/iganna/pannagramā
š Quick Use with Docker
This guide provides the steps to run the complete pangenome analysis using our automated pipeline.
Prerequisites:
#Docker must be installed on your system.
#You must have a folder containing your genome assembly files (.fasta or .fa).
Step 1: Pull the Pannagram Image
Download the pre-built image from Docker Hub. This contains Pannagram, all its dependencies, and the automated pipeline script, fully configured.
docker pull engbio/pannagram:v1.0
Step 2: Verify the Image
Check if the image was successfully downloaded to your machine.
docker images
You should see engbio/pannagram with the tag v1.0 in the list.
Step 3: Prepare Your Data Folders The container is designed to work with a specific folder structure.
Create a main project directory (e.g., my_pannagram_project).
Inside it, create a subdirectory for your genomes: genomes.
Also create an empty subdirectory for the results: results.
Your structure should look like this:
my_pannagram_project/
āāā genomes/
ā āāā genome1.fasta
ā āāā genome2.fa
ā āāā ...
āāā results/
Step 4: Run the Automated Pangenome Analysis
Navigate your terminal to your main project directory (my_pannagram_project) and run the container. Bash
docker run --rm \
-v "$(pwd)/genomes:/data/input" \
-v "$(pwd)/results:/data/output" \
-e NUM_CORES=8 \
engbio/pannagram:v1.0
This single command will:
Start the container.
Mount your local genomes folder for reading and your results folder for writing.
Set the number of CPU cores to use (e.g., -e NUM_CORES=8).
Automatically execute the internal pipeline script, which performs the full analysis:
Runs pannagram to build the pangenome alignment in reference-free mode.
Runs features to extract synteny blocks, consensus sequences, SNPs, and structural variants from the alignment.
Step 5: Fix Output File Permissions The results generated in your local results folder will be owned by root. To reclaim ownership, run this command after the analysis is complete: Bash
sudo chown -R $(id -u):$(id -g) results/
āļø For Advanced Users: Configuration Files These are the files used to build the engbio/pannagram:v1.0 image and automate the pipeline.
Dockerfile
This Dockerfile creates the complete environment, installs Pannagram, and sets the internal pipeline script as the main command.
# Dockerfile (Final Version for Automated Pipeline)
FROM mambaorg/micromamba:1.5.8
ARG GIT_REPO="https://github.com/iganna/pannagram.git"
ARG GIT_BRANCH="main"
# Use ROOT to ensure full permissions during build
USER root
# 1. Install system dependencies for compilation
RUN apt-get update && apt-get install -y git build-essential && apt-get clean && rm -rf /var/lib/apt/lists/*
# 2. Clone the repository
RUN git clone --branch ${GIT_BRANCH} ${GIT_REPO} /opt/pannagram
WORKDIR /opt/pannagram
# 3. Create the conda environment with all dependencies
RUN micromamba env create -f pannagram.yml --yes && \
micromamba clean --all --yes
# 4. Set the SHELL to run subsequent commands inside the conda environment
SHELL ["micromamba", "run", "-n", "pannagram", "/bin/bash", "-c"]
# 5. The crucial step: Build and install the Pannagram R package and scripts
RUN bash build.sh
# 6. Revert to the default shell
SHELL ["/bin/bash", "-c"]
# 7. Add the environment's bin directory to the PATH for the final image
ENV PATH /opt/conda/envs/pannagram/bin:$PATH
# 8. Copy the automation script and set it as the main command
COPY internal_pipeline.sh /usr/local/bin/internal_pipeline.sh
RUN chmod +x /usr/local/bin/internal_pipeline.sh
ENTRYPOINT ["/usr/local/bin/internal_pipeline.sh"]
Script to Automate Pipeline (internal_pipeline.sh)
This is the Bash script included in the image that orchestrates the entire two-stage analysis.
#!/bin/bash
set -e
# Configure variables from environment variables or use default values
PATH_IN="${INPUT_DIR:-/data/input}"
PATH_OUT="${OUTPUT_DIR:-/data/output}"
CORES="${NUM_CORES:-8}"
# --- Print Execution Status ---
echo "--- Starting Pannagram Automated Pipeline (Reference-Free Mode) ---"
echo "Input Directory: ${PATH_IN}"
echo "Output Directory: ${PATH_OUT}"
echo "Cores: ${CORES}"
echo "-----------------------------------------------------------"
# --- Step 1: Run Pangenome Alignment ---
echo "STEP 1: Running pangenome alignment with 'pannagram'..."
pannagram \
-path_in "${PATH_IN}" \
-path_out "${PATH_OUT}/alignment_results" \
-cores "${CORES}"
echo "Alignment complete."
echo "-----------------------------------------------------------"
# --- Step 2: Run Feature Extraction ---
echo "STEP 2: Running feature extraction with 'features'..."
features \
-path_in "${PATH_OUT}/alignment_results" \
-blocks \
-seq \
-snp \
-sv_call \
-cores "${CORES}"
echo "Feature extraction complete."
echo "--- Pannagram Pipeline Finished Successfully! ---"
Content type
Image
Digest
sha256:7217dc53dā¦
Size
1 GB
Last updated
about 1 year ago
docker pull engbio/pannagram:v1.0