Sign inSign up

pinellolab/crisprme

By pinellolab

•Updated 4 days ago

Variant and haplotype aware enumeration of potential CRISPR genome editing off-targets

Image
0

10K+

pinellolab/crisprme repository overview

install with bioconda GitHub release (latest by date) Conda license

crisprme-logo.png

⁠CRISPRme

CRISPRme is a comprehensive tool designed for thorough off-target assessment in CRISPR-Cas systems. Available as a web application (http://crisprme.di.univr.it/⁠), offline tool, and command-line interface, it integrates human genetic variant datasets with orthogonal genomic annotations to predict and prioritize potential off-target sites at scale. CRISPRme accounts for single-nucleotide variants (SNVs) and indels, considers bona fide haplotypes, and allows for spacer:protospacer mismatches and bulges, making it well-suited for both population-wide and personal genome analyses. CRISPRme automates the entire workflow, from data download to executing the search, and delivers detailed reports complete with tables and figures through an interactive web-based interface.

⁠Table Of Contents

0 System Requirements⁠
1 Installation⁠
  1.1 Install CRISPRme via Conda/Mamba⁠
    1.1.1 Installing Conda or Mamba⁠
    1.1.2 Installing CRISPRme⁠
    1.1.3 Updating CRISPRme⁠
  1.2 Install CRISPRme via Docker⁠
    1.2.1 Installing Docker⁠
    1.2.2 Building and Pulling CRISPRme Docker Image⁠
2 Usage⁠
  2.1 Directory Structure⁠
  2.2 CRISPRme Functions⁠
     2.2.1 Complete Search⁠
     2.2.2 Complete Test⁠
     2.2.3 Off-target sites validation Test⁠
     2.2.4 Targets Integration⁠
     2.2.5 GNOMAD Converter⁠
     2.2.6 Generate Personal Card⁠
     2.2.7 Web Interface⁠
3 Test⁠
  3.1 Quick Test⁠
  3.2 Detailed Test⁠
     3.2.1 Single Chromosome Test⁠
     3.2.2 Full Genome Test⁠
4 Citation⁠
5 Contacts⁠
6 License⁠

⁠0 System Requirements

To ensure optimal performance, CRISPRme requires the following:

  • Minimum Memory (RAM): 32 GB
    Suitable for typical use cases and smaller datasets.

  • Recommended Memory for Large Analyses: 64 GB or more
    Necessary for intensive operations such as whole-genome searches and processing large variant datasets.

For best results, confirm that your system meets or exceeds these specifications before running CRISPRme.

⁠1 Installation

This section outlines the steps to install CRISPRme, tailored to suit different operating systems. Select the method that best matches your setup:

Each method ensures a streamlined and efficient installation, enabling you to use CRISPRme with minimal effort. Follow the detailed instructions provided in the respective sections below.

⁠1.1 Install CRISPRme via Conda/Mamba

This section is organized into three subsections to guide you through the installation and maintenance of CRISPRme:

  • Installing Conda or Mamba⁠:
    This subsection provides step-by-step instructions to install either Conda or Mamba. Begin here if you do not have these package managers installed on your machine.

  • Installing CRISPRme⁠:
    Once you have Conda or Mamba set up, proceed to this subsection for detailed instructions on creating the CRISPRme environment and installing the necessary dependencies.

  • Updating CRISPRme⁠:
    Learn how to update an existing CRISPRme installation to the latest version, ensuring access to new features and bug fixes.

⁠1.1.1 Installing Conda or Mamba

Before installing CRISPRme, ensure you have either Conda or Mamba installed on your machine. Based on recommendations from the Bioconda community, we highly recommend using Mamba over Conda. Mamba is a faster, more efficient drop-in replacement for Conda, leveraging a high-performance dependency solver and components optimized in C++.

Step1: Install Conda or Mamba

Step 2: Configure Bioconda Channels

Once Mamba is installed, configure it to use Bioconda and related channels by running the following one-time setup commands:

mamba config --add channels bioconda
mamba config --add channels defaults
mamba config --add channels conda-forge
mamba config --set channel_priority strict

Note: If you prefer to use Conda, replace mamba with conda in the commands above

By completing these steps, your system will be fully prepared for installing CRISPRme.

⁠1.1.2 Installing CRISPRme

We strongly recommend using Mamba to create CRISPRme's conda environment due to its superior speed and reliability in dependency management. However, if you prefer Conda, you can replace mamba with conda in all the commands below.

Step 1: Create CRISPRme's Environment

Open a terminal and execute the following command:

mamba create -n crisprme python=3.9 crisprme -y  # Install CRISPRme and its dependencies

This command sets up a dedicated conda environment named crisprme, installing CRISPRme along with all required dependencies.

Step 2: Activate the Environment

To activate the newly created CRISPRme environment, type:

mamba activate crisprme  # Enable the CRISPRme environment

Step 3: Test the Installation

To verify that CRISPRme is correctly installed, run the following commands in your terminal:

crisprme.py --version  # Display the installed CRISPRme version
crisprme.py            # List CRISPRme functionalities
  • The first command will output the version of CRISPRme (e.g., 2.1.6).
  • The second command should display CRISPRme's functionalities.

If both commands execute successfully, your installation is complete, and CRISPRme is ready to use.

⁠1.1.3 Updating CRISPRme

To update an existing CRISPRme installation using Mamba or Conda, follow the steps below:

Step 1: Check the Latest Version

Visit the CRISPRme README to identify the latest version of the tool.

Step 2: Update CRISPRme

Run the following command in your terminal, replacing <latest_version> with the desired version number:

mamba install crisprme=<latest_version>  # Update CRISPRme to the specified version

For example, to update CRISPRme to version 2.1.6, execute:

mamba install crisprme=2.1.6

If you're using Conda, replace mamba with conda in the commands above.

Step 3: Verify the Update

After the update completes, ensure the installation was successful by checking the version:

crisprme.py --version  # Confirm the installed version

If the displayed version matches the one you installed, the update was successful.

⁠1.2 Install CRISPRme via Docker

This section is organized into two subsections to guide you through the setup of CRISPRme using Docker:

  • Installing Docker⁠:
    Provides step-by-step instructions for installing Docker on your system, ensuring compatibility with all operating systems, including Linux, macOS, and Windows.

  • Building and Pulling CRISPRme Docker Image⁠:
    Explains how to create or download the CRISPRme Docker image to set up a containerized environment for seamless execution.

Follow the subsections in order if Docker is not yet installed on your machine. If Docker is already installed, skip to the second subsection.

⁠1.2.1 Installing Docker

MacOS and Windows users are encouraged to install Docker⁠ to use CRISPRme. Linux users may also choose Docker for convenience and compatibility.

Docker provides tailored distributions for different operating systems. Follow the official Docker installation guide specific to your OS:

Linux-Specific Post-Installation Steps

If you're using Linux, additional configuration steps are required:

  1. Create the Docker Group:
sudo groupadd docker
  1. Add Your User to the Docker Group:
sudo usermod -aG docker $USER

Repeat this command for any additional users you want to include in the Docker Group.

  1. Restart Your Machine
    Log out and log back in, or restart your machine to apply the changes.

Testing Docker Installation

Once Docker is installed, verify the setup by opening a terminal window and typing:

docker run hello-world

If Docker is installed correctly, you should see output like this:

Hello from Docker!
This message shows that your installation appears to be working correctly.

To generate this message, Docker took the following steps:
 1. The Docker client contacted the Docker daemon.
 2. The Docker daemon pulled the "hello-world" image from the Docker Hub.
 3. The Docker daemon created a new container from that image, which runs the executable that produces this output.
 4. The Docker daemon streamed this output to the Docker client, which displayed it on your terminal.

For more examples and ideas, visit:
 https://docs.docker.com/get-started/
⁠1.2.2 Building and Pulling CRISPRme Docker Image

After installing Docker, you can download and build the CRISPRme Docker image by running the following command in a terminal:

docker pull pinellolab/crisprme

This command retrieves the latest pre-built CRISPRme image from Docker Hub and sets it up on your system, ensuring all required dependencies and configurations are included.

Once the download is complete, the CRISPRme Docker image will be ready for use. To confirm the image is successfully installed, you can list all available Docker images by typing:

docker images

Look for an entry similar to the following:

REPOSITORY          TAG       IMAGE ID       CREATED        SIZE
pinellolab/crisprme latest    <image_id>     <timestamp>    <size>

You are now ready to run CRISPRme using Docker.

⁠2 Usage

CRISPRme is a tool designed for variant- and haplotype-aware CRISPR off-target analysis. It integrates robust functionalities for off-target detection, variant-aware search, and result analysis. The tool also includes a user-friendly graphical interface, which can be deployed locally to streamline its usage.

⁠2.1 Directory Structure

CRISPRme operates within a specific directory structure to manage input data and outputs efficiently. To ensure proper functionality, your working directory must include the following main subdirectories:

  • Genomes

    • Purpose: Stores reference genomes.
    • Structure: Each reference genome resides in its own subdirectory.
    • Requirements: The genome must be split into separate files, each representing a single chromosome.
  • VCFs

    • Purpose: Contains variant data in VCF format.
    • Structure: Similar to the Genomes directory, each dataset has a dedicated subdirectory with VCF files split by chromosome.
    • Requirements: Files must be compressed using bgzip (with a .gz extension).
  • sampleIDs

    • Purpose: Lists the sample identifiers corresponding to the VCF datasets.
    • Structure: Tab-separated files, one for each VCF dataset, specifying the sample IDs.
  • Annotations

    • Purpose: Provides genome annotation data.
    • Format: Annotation files must be in BED format.
  • PAMs

    • Purpose: Specifies the Protospacer Adjacent Motif (PAM) sequences for off-target search.
    • Format: Text files containing PAM sequences.

The directory organization required by CRISPRme is illustrated below:

crisprme_dirtree.png

⁠2.2 CRISPRme Functions

This section provides a comprehensive overview of CRISPRme's core functions, detailing each feature, the required input data and formats, and the resulting outputs. The following is a summary of CRISPRme's key features:

  • Complete Search⁠ (complete-search)
    Executes a genome-wide off-targets search across both reference and variant datasets (if specified), conducts Cutting Frequency Determination (CFD) and CRISTA analyses (if applicable), and identifies candidate targets.

  • Complete Test⁠ (complete-test)
    Tests CRISPRme pipeline on a small input dataset or the full genome, enabling users to validate the tool's functionality before performing large-scale analyses.

  • Off-target sites validation Test⁠ (validate-test)
    Validates off-target sites generated by the Complete Test workflow by comparing CRISPRme predictions against brute-force ground-truth alignments derived from 1000 Genomes variant data.

  • Targets Integration⁠ (targets-integration)
    Combines in silico predicted targets with experimental data to create a finalized target panel.

  • GNOMAD Converter⁠ (gnomAD-converter)
    Transforms GNOMAD VCFs (vcf.bgz format) into a format compatible with CRISPRme. The function supports VCFs from GNOMAD v3.1, v4.0, and v4.1, including joint VCFs.

  • Generate Personal Card⁠ (generate-personal-card)
    Generates a personalized summary for a specific sample, identifying all private off-targets unique to that individual.

  • Web Interface⁠ (web-interface)
    Launches CRISPRme's interactive web interface, allowing users to manage and execute tasks directly via a local browser.


The Complete Search function performs an exhaustive variant- and haplotype-aware off-target analysis, leveraging the provided reference genome and variant datasets to deliver comprehensive results. This feature integrates all critical stages of the CRISPRme pipeline, encompassing off-target identification, functional annotation, and detailed reporting.

Key highlights of the Complete Search functionality include:

  • Variant- and Haplotype-Awareness
    Accurately incorporates genetic variation, including population- and sample-specific variants, and haplotypes data, to identify off-targets that reflect real-world genomic diversity.

  • Comprehensive Off-Target Discovery
    Searches both the reference genome and user-specified variant datasets for potential off-targets, including those encompassing mismatches and bulges.

  • Functional Annotation
    Annotates off-targets with relevant genomic features, such as coding/non-coding regions, regulatory elements, and gene proximity.

  • Detailed Reporting
    Generates population-specific and sample-specific off-target summaries, highlighting variations that may impact specificity or introduce novel PAM sites. Provides CFD (Cutting Frequency Determination) and CRISTA scores, and mismatches and bulges counts to rank off-targets based on their potential impact. Includes graphical representations of findings to facilitate result interpretation.

  • Output Formats
    Produces user-friendly output files, including text-based tables and visualization-ready graphical summaries.

Usage Example for the Complete Search function:

  • Via Conda/Mamba

    crisprme.py complete-search \
      --genome Genomes/hg38 \  # reference genome directory
      --vcf vcf_config.1000G.HGDP.txt \  # config file declaring usage of 1000G and HGDP variant datasets
      --guide sg1617.txt \  # guide 
      --pam PAMs/20bp-NGG-spCas9.txt \  # NGG PAM file
      --annotation Annotations/dhs+gencode+encode.hg38.bed \  # annotation BED
      --gene_annotation Annotations/gencode.protein_coding.bed \  # gene proximity annotation BED
      --samplesID samplesIDs.1000G.HGDP.txt \  # config file declaring usage of 1000G and HGDP samples
      --be-window 4,8 \  # base editing window start and stop positions within off-targets
      --be-base A,G \  # nucleotide to test base editing potential (A>G)
      --mm 6 \  # number of max mismatches
      --bDNA 2 \  # number of max DNA bulges
      --bRNA 2 \  # number of max RNA bulges
      --merge 3 \  # merge off-targets mapped within 3 bp in clusters
      --sorting-criteria-scoring mm+bulges \  # prioritize within each cluster off-targets with highest score and lowest mm+bulges (CFD and CRISTA reports only)
      --sorting-criteria mm,bulges \  # prioritize within each cluster off-targets with lowest mm and bulges counts
      --output sg1617-NGG-1000G-HGDP \  # output directory name
      --thread 8  # number of threads 
    
  • Via Docker

    docker run -v ${PWD}:/DATA -w /DATA -i pinellolab/crisprme \
      crisprme.py complete-search \
      --genome Genomes/hg38 \  # reference genome directory
      --vcf vcf_config.1000G.HGDP.txt \  # config file declaring usage of 1000G and HGDP variant datasets
      --guide sg1617.txt \  # guide 
      --pam PAMs/20bp-NGG-spCas9.txt \  # NGG PAM file
      --annotation Annotations/dhs+gencode+encode.hg38.bed \  # annotation BED
      --gene_annotation Annotations/gencode.protein_coding.bed \  # gene proximity annotation BED
      --samplesID samplesIDs.1000G.HGDP.txt \  # config file declaring usage of 1000G and HGDP samples
      --be-window 4,8 \  # base editing window start and stop positions within off-targets
      --be-base A,G \  # nucleotide to test base editing potential (A>G)
      --mm 6 \  # number of max mismatches
      --bDNA 2 \  # number of max DNA bulges
      --bRNA 2 \  # number of max RNA bulges
      --merge 3 \  # merge off-targets mapped within 3 bp in clusters
      --sorting-criteria-scoring mm+bulges \  # prioritize within each cluster off-targets with highest score and lowest mm+bulges (CFD and CRISTA reports only)
      --sorting-criteria mm,bulges \  # prioritize within each cluster off-targets with lowest mm and bulges counts
      --output sg1617-NGG-1000G-HGDP \  # output directory name
      --thread 8  # number of threads 
    
⁠Input Arguments

Below is a detailed list of the input arguments required or optionally used by the Complete Search function. Each parameter is explained to ensure clarity in its purpose and usage:

General Parameters

  • --help
    Displays the help message with usage details and exits. Useful for quickly referencing all available options.

  • --output (Required)
    Specifies the name of the output directory where all results from the analysis will be saved. This directory will be created within the Results directory.

  • --thread (Optional - Default: 4)
    Defines the number of CPU threads to use for parallel computation. Increasing the number of threads can speed up analysis on systems with multiple cores.

  • --debug (Optional)
    Runs the tool in debug mode.

Input Data Parameters

  • --genome (Required)
    Path to the directory containing the reference genome in FASTA format. Each chromosome must be in a separate file (e.g., chr1.fa, chr2.fa, etc.).

  • --vcf (Optional)
    Path to text config file listing the directories containing VCF files to be integrated into the analysis. When provided, CRISPRme conducts variant- and haplotype-aware searches. If not specified, the tool searches only on the reference genome.

  • --guide
    Path to a text file containing one or more guide RNA sequences (one per line) to search for in the input genome and variants. This argument cannot be used together with --sequence.

  • --sequence
    Path to a FASTA file listing guide RNA sequences. This argument is an alternative to --guide and cannot be used simultaneously.

  • --pam (Required)
    Path to a text file specifying the PAM sequence(s) required for the search. The file should define the PAM format (e.g., NGG for SpCas9).

Annotation Parameters (Optional)

  • --annotation
    Path to a BED file containing genomic annotations, such as regulatory regions (e.g., DNase hypersensitive sites, enhancers, promoters). These annotations provide functional context for identified off-targets.

  • --gene_annotation
    Path to a BED file containing gene information, such as Gencode protein-coding annotations. This is used to calculate the proximity of off-targets to genes for downstream analyses.

Base Editing Parameters (Optional)

  • --be-window
    Specifies the editing window for base editors, defined as start and stop positions relative to the guide RNA (1-based, comma-separated). This defines the region of interest for base-editing analysis.

  • --be-base
    Defines the target nucleotide(s) for base editing. This is only used when base editing functionality is needed.

Sample-Specific Parameters (Optional)

  • --samplesID
    Path to a text config file listing sample identifiers (one per line) corresponding to VCF datasets. This enables sample-specific off-target analyses. Mandatory if --vcf is specified.

Search and Merging Parameters

  • --mm (Required)
    Maximum number of mismatches allowed during off-target identification.

  • --bDNA (Required)
    Maximum allowable DNA bulge size.

  • --bRNA (Required)
    Maximum allowable RNA bulge size.

  • --merge (Optional - Default: 3)
    Defines the window size (in base pairs) used to merge closely spaced off-targets. Pivot targets are selected based on the highest score (e.g., CFD, CRISTA) or criteria defined by the --sorting-criteria.

  • --sorting-criteria-scoring (Optional - Default: mm+bulges)
    Specifies sorting criteria for merging when using CFD/CRISTA scores. Options include:

    • mm: Number of mismatches.
    • bulges: Total bulge size.
    • mm+bulges: Combined mismatches and bulges.
  • --sorting-criteria (Optional - Default: mm+bulges,mm)
    Sorting criteria used when CFD/CRISTA scores are unavailable. Options are similar to --sorting-criteria-scoring but tailored for simpler analyses.

Note 1: Ensure compatibility between input files and genome builds (e.g., hg38 or hg19) to avoid alignment issues.

Note 2: Optional arguments can be omitted when not applicable, but required arguments must always be specified.

⁠Output Data Overview

The Complete Search function generates a comprehensive suite of reports, detailing the identified and prioritized targets, along with statistical and graphical summaries. These outputs are essential for interpreting the search results and understanding the impact of genetic diversity on CRISPR off-target.

Output Off-targets Files Description

  1. *.integrated_results.tsv

    • Contents: A detailed file containing the top targets (`*.bestMerge.tx

Tag summary

Content type

Image

Digest

sha256:c8b908d99…

Size

4.5 GB

Last updated

4 days ago

docker pull pinellolab/crisprme