Variant and haplotype aware enumeration of potential CRISPR genome editing off-targets
10K+
CRISPRme is a comprehensive tool designed for thorough off-target assessment in CRISPR-Cas systems. Available as a web application (http://crisprme.di.univr.it/), offline tool, and command-line interface, it integrates human genetic variant datasets with orthogonal genomic annotations to predict and prioritize potential off-target sites at scale. CRISPRme accounts for single-nucleotide variants (SNVs) and indels, considers bona fide haplotypes, and allows for spacer:protospacer mismatches and bulges, making it well-suited for both population-wide and personal genome analyses. CRISPRme automates the entire workflow, from data download to executing the search, and delivers detailed reports complete with tables and figures through an interactive web-based interface.
0 System Requirements
1 Installation
1.1 Install CRISPRme via Conda/Mamba
1.1.1 Installing Conda or Mamba
1.1.2 Installing CRISPRme
1.1.3 Updating CRISPRme
1.2 Install CRISPRme via Docker
1.2.1 Installing Docker
1.2.2 Building and Pulling CRISPRme Docker Image
2 Usage
2.1 Directory Structure
2.2 CRISPRme Functions
2.2.1 Complete Search
2.2.2 Complete Test
2.2.3 Off-target sites validation Test
2.2.4 Targets Integration
2.2.5 GNOMAD Converter
2.2.6 Generate Personal Card
2.2.7 Web Interface
3 Test
3.1 Quick Test
3.2 Detailed Test
3.2.1 Single Chromosome Test
3.2.2 Full Genome Test
4 Citation
5 Contacts
6 License
To ensure optimal performance, CRISPRme requires the following:
Minimum Memory (RAM): 32 GB
Suitable for typical use cases and smaller datasets.
Recommended Memory for Large Analyses: 64 GB or more
Necessary for intensive operations such as whole-genome searches and
processing large variant datasets.
For best results, confirm that your system meets or exceeds these specifications before running CRISPRme.
This section outlines the steps to install CRISPRme, tailored to suit different operating systems. Select the method that best matches your setup:
Each method ensures a streamlined and efficient installation, enabling you to use CRISPRme with minimal effort. Follow the detailed instructions provided in the respective sections below.
This section is organized into three subsections to guide you through the installation and maintenance of CRISPRme:
Installing Conda or Mamba:
This subsection provides step-by-step instructions to install either
Conda or Mamba. Begin here if you do not have these package managers installed
on your machine.
Installing CRISPRme:
Once you have Conda or Mamba set up, proceed to this subsection for detailed
instructions on creating the CRISPRme environment and installing the necessary
dependencies.
Updating CRISPRme:
Learn how to update an existing CRISPRme installation to the latest version,
ensuring access to new features and bug fixes.
Before installing CRISPRme, ensure you have either Conda or Mamba installed on your machine. Based on recommendations from the Bioconda community, we highly recommend using Mamba over Conda. Mamba is a faster, more efficient drop-in replacement for Conda, leveraging a high-performance dependency solver and components optimized in C++.
Step1: Install Conda or Mamba
To install Conda, refer to the official installation guide:
Conda Installation Guide
To install Mamba, refer to the official installation guide:
Mamba Installation Guide
Step 2: Configure Bioconda Channels
Once Mamba is installed, configure it to use Bioconda and related channels by
running the following one-time setup commands:
mamba config --add channels bioconda
mamba config --add channels defaults
mamba config --add channels conda-forge
mamba config --set channel_priority strict
Note: If you prefer to use
Conda, replacemambawithcondain the commands above
By completing these steps, your system will be fully prepared for installing CRISPRme.
We strongly recommend using Mamba to create CRISPRme's conda environment
due to its superior speed and reliability in dependency management. However, if
you prefer Conda, you can replace mamba with conda in all the commands below.
Step 1: Create CRISPRme's Environment
Open a terminal and execute the following command:
mamba create -n crisprme python=3.9 crisprme -y # Install CRISPRme and its dependencies
This command sets up a dedicated conda environment named crisprme, installing
CRISPRme along with all required dependencies.
Step 2: Activate the Environment
To activate the newly created CRISPRme environment, type:
mamba activate crisprme # Enable the CRISPRme environment
Step 3: Test the Installation
To verify that CRISPRme is correctly installed, run the following commands in your terminal:
crisprme.py --version # Display the installed CRISPRme version
crisprme.py # List CRISPRme functionalities
2.1.6).If both commands execute successfully, your installation is complete, and CRISPRme is ready to use.
To update an existing CRISPRme installation using Mamba or Conda, follow the
steps below:
Step 1: Check the Latest Version
Visit the CRISPRme README to identify the latest version of the tool.
Step 2: Update CRISPRme
Run the following command in your terminal, replacing <latest_version> with the desired version number:
mamba install crisprme=<latest_version> # Update CRISPRme to the specified version
For example, to update CRISPRme to version 2.1.6, execute:
mamba install crisprme=2.1.6
If you're using Conda, replace mamba with conda in the commands above.
Step 3: Verify the Update
After the update completes, ensure the installation was successful by checking the version:
crisprme.py --version # Confirm the installed version
If the displayed version matches the one you installed, the update was successful.
This section is organized into two subsections to guide you through the setup of CRISPRme using Docker:
Installing Docker:
Provides step-by-step instructions for installing Docker on your system,
ensuring compatibility with all operating systems, including Linux, macOS, and Windows.
Building and Pulling CRISPRme Docker Image:
Explains how to create or download the CRISPRme Docker image to set up a
containerized environment for seamless execution.
Follow the subsections in order if Docker is not yet installed on your machine. If Docker is already installed, skip to the second subsection.
MacOS and Windows users are encouraged to install Docker to use CRISPRme. Linux users may also choose Docker for convenience and compatibility.
Docker provides tailored distributions for different operating systems. Follow the official Docker installation guide specific to your OS:
Linux-Specific Post-Installation Steps
If you're using Linux, additional configuration steps are required:
sudo groupadd docker
sudo usermod -aG docker $USER
Repeat this command for any additional users you want to include in the Docker Group.
Testing Docker Installation
Once Docker is installed, verify the setup by opening a terminal window and typing:
docker run hello-world
If Docker is installed correctly, you should see output like this:
Hello from Docker!
This message shows that your installation appears to be working correctly.
To generate this message, Docker took the following steps:
1. The Docker client contacted the Docker daemon.
2. The Docker daemon pulled the "hello-world" image from the Docker Hub.
3. The Docker daemon created a new container from that image, which runs the executable that produces this output.
4. The Docker daemon streamed this output to the Docker client, which displayed it on your terminal.
For more examples and ideas, visit:
https://docs.docker.com/get-started/
After installing Docker, you can download and build the CRISPRme Docker image by running the following command in a terminal:
docker pull pinellolab/crisprme
This command retrieves the latest pre-built CRISPRme image from Docker Hub and sets it up on your system, ensuring all required dependencies and configurations are included.
Once the download is complete, the CRISPRme Docker image will be ready for use. To confirm the image is successfully installed, you can list all available Docker images by typing:
docker images
Look for an entry similar to the following:
REPOSITORY TAG IMAGE ID CREATED SIZE
pinellolab/crisprme latest <image_id> <timestamp> <size>
You are now ready to run CRISPRme using Docker.
CRISPRme is a tool designed for variant- and haplotype-aware CRISPR off-target analysis. It integrates robust functionalities for off-target detection, variant-aware search, and result analysis. The tool also includes a user-friendly graphical interface, which can be deployed locally to streamline its usage.
CRISPRme operates within a specific directory structure to manage input data and outputs efficiently. To ensure proper functionality, your working directory must include the following main subdirectories:
Genomes
VCFs
bgzip (with a .gz
extension).sampleIDs
Annotations
PAMs
The directory organization required by CRISPRme is illustrated below:
This section provides a comprehensive overview of CRISPRme's core functions, detailing each feature, the required input data and formats, and the resulting outputs. The following is a summary of CRISPRme's key features:
Complete Search (complete-search)
Executes a genome-wide off-targets
search across both reference and variant datasets (if specified), conducts
Cutting Frequency Determination (CFD) and CRISTA analyses (if applicable), and
identifies candidate targets.
Complete Test (complete-test)
Tests CRISPRme pipeline on a small input dataset or the full genome,
enabling users to validate the tool's functionality before performing
large-scale analyses.
Off-target sites validation Test (validate-test)
Validates off-target sites generated by the Complete Test workflow by
comparing CRISPRme predictions against brute-force ground-truth alignments
derived from 1000 Genomes variant data.
Targets Integration (targets-integration)
Combines in silico predicted targets with experimental data to create a
finalized target panel.
GNOMAD Converter (gnomAD-converter)
Transforms GNOMAD VCFs (vcf.bgz format) into a format compatible with
CRISPRme. The function supports VCFs from GNOMAD v3.1, v4.0, and v4.1,
including joint VCFs.
Generate Personal Card (generate-personal-card)
Generates a personalized summary for a specific sample, identifying all
private off-targets unique to that individual.
Web Interface (web-interface)
Launches CRISPRme's interactive web interface, allowing users to manage
and execute tasks directly via a local browser.
The Complete Search function performs an exhaustive variant- and haplotype-aware off-target analysis, leveraging the provided reference genome and variant datasets to deliver comprehensive results. This feature integrates all critical stages of the CRISPRme pipeline, encompassing off-target identification, functional annotation, and detailed reporting.
Key highlights of the Complete Search functionality include:
Variant- and Haplotype-Awareness
Accurately incorporates genetic variation, including population- and
sample-specific variants, and haplotypes data, to identify off-targets that
reflect real-world genomic diversity.
Comprehensive Off-Target Discovery
Searches both the reference genome and user-specified variant datasets
for potential off-targets, including those encompassing mismatches and bulges.
Functional Annotation
Annotates off-targets with relevant genomic features, such as
coding/non-coding regions, regulatory elements, and gene proximity.
Detailed Reporting
Generates population-specific and sample-specific off-target summaries,
highlighting variations that may impact specificity or introduce novel PAM
sites. Provides CFD (Cutting Frequency Determination) and CRISTA scores, and
mismatches and bulges counts to rank off-targets based on their potential
impact. Includes graphical representations of findings to facilitate result
interpretation.
Output Formats
Produces user-friendly output files, including text-based tables and
visualization-ready graphical summaries.
Usage Example for the Complete Search function:
Via Conda/Mamba
crisprme.py complete-search \
--genome Genomes/hg38 \ # reference genome directory
--vcf vcf_config.1000G.HGDP.txt \ # config file declaring usage of 1000G and HGDP variant datasets
--guide sg1617.txt \ # guide
--pam PAMs/20bp-NGG-spCas9.txt \ # NGG PAM file
--annotation Annotations/dhs+gencode+encode.hg38.bed \ # annotation BED
--gene_annotation Annotations/gencode.protein_coding.bed \ # gene proximity annotation BED
--samplesID samplesIDs.1000G.HGDP.txt \ # config file declaring usage of 1000G and HGDP samples
--be-window 4,8 \ # base editing window start and stop positions within off-targets
--be-base A,G \ # nucleotide to test base editing potential (A>G)
--mm 6 \ # number of max mismatches
--bDNA 2 \ # number of max DNA bulges
--bRNA 2 \ # number of max RNA bulges
--merge 3 \ # merge off-targets mapped within 3 bp in clusters
--sorting-criteria-scoring mm+bulges \ # prioritize within each cluster off-targets with highest score and lowest mm+bulges (CFD and CRISTA reports only)
--sorting-criteria mm,bulges \ # prioritize within each cluster off-targets with lowest mm and bulges counts
--output sg1617-NGG-1000G-HGDP \ # output directory name
--thread 8 # number of threads
Via Docker
docker run -v ${PWD}:/DATA -w /DATA -i pinellolab/crisprme \
crisprme.py complete-search \
--genome Genomes/hg38 \ # reference genome directory
--vcf vcf_config.1000G.HGDP.txt \ # config file declaring usage of 1000G and HGDP variant datasets
--guide sg1617.txt \ # guide
--pam PAMs/20bp-NGG-spCas9.txt \ # NGG PAM file
--annotation Annotations/dhs+gencode+encode.hg38.bed \ # annotation BED
--gene_annotation Annotations/gencode.protein_coding.bed \ # gene proximity annotation BED
--samplesID samplesIDs.1000G.HGDP.txt \ # config file declaring usage of 1000G and HGDP samples
--be-window 4,8 \ # base editing window start and stop positions within off-targets
--be-base A,G \ # nucleotide to test base editing potential (A>G)
--mm 6 \ # number of max mismatches
--bDNA 2 \ # number of max DNA bulges
--bRNA 2 \ # number of max RNA bulges
--merge 3 \ # merge off-targets mapped within 3 bp in clusters
--sorting-criteria-scoring mm+bulges \ # prioritize within each cluster off-targets with highest score and lowest mm+bulges (CFD and CRISTA reports only)
--sorting-criteria mm,bulges \ # prioritize within each cluster off-targets with lowest mm and bulges counts
--output sg1617-NGG-1000G-HGDP \ # output directory name
--thread 8 # number of threads
Below is a detailed list of the input arguments required or optionally used by the Complete Search function. Each parameter is explained to ensure clarity in its purpose and usage:
General Parameters
--help
Displays the help message with usage details and exits. Useful for quickly
referencing all available options.
--output (Required)
Specifies the name of the output directory where all results from the
analysis will be saved. This directory will be created within the Results
directory.
--thread (Optional - Default: 4)
Defines the number of CPU threads to use for parallel computation.
Increasing the number of threads can speed up analysis on systems with
multiple cores.
--debug (Optional)
Runs the tool in debug mode.
Input Data Parameters
--genome (Required)
Path to the directory containing the reference genome in FASTA format.
Each chromosome must be in a separate file (e.g., chr1.fa, chr2.fa, etc.).
--vcf (Optional)
Path to text config file listing the directories containing VCF files to
be integrated into the analysis. When provided, CRISPRme conducts variant-
and haplotype-aware searches. If not specified, the tool searches only on the
reference genome.
--guide
Path to a text file containing one or more guide RNA sequences (one per
line) to search for in the input genome and variants. This argument cannot be
used together with --sequence.
--sequence
Path to a FASTA file listing guide RNA sequences. This argument is an
alternative to --guide and cannot be used simultaneously.
--pam (Required)
Path to a text file specifying the PAM sequence(s) required for the
search. The file should define the PAM format (e.g., NGG for SpCas9).
Annotation Parameters (Optional)
--annotation
Path to a BED file containing genomic annotations, such as regulatory
regions (e.g., DNase hypersensitive sites, enhancers, promoters). These
annotations provide functional context for identified off-targets.
--gene_annotation
Path to a BED file containing gene information, such as Gencode
protein-coding annotations. This is used to calculate the proximity of
off-targets to genes for downstream analyses.
Base Editing Parameters (Optional)
--be-window
Specifies the editing window for base editors, defined as start and stop
positions relative to the guide RNA (1-based, comma-separated). This defines
the region of interest for base-editing analysis.
--be-base
Defines the target nucleotide(s) for base editing. This is only used when
base editing functionality is needed.
Sample-Specific Parameters (Optional)
--samplesID
--vcf is specified.Search and Merging Parameters
--mm (Required)
Maximum number of mismatches allowed during off-target identification.
--bDNA (Required)
Maximum allowable DNA bulge size.
--bRNA (Required)
Maximum allowable RNA bulge size.
--merge (Optional - Default: 3)
Defines the window size (in base pairs) used to merge closely spaced
off-targets. Pivot targets are selected based on the highest score (e.g.,
CFD, CRISTA) or criteria defined by the --sorting-criteria.
--sorting-criteria-scoring (Optional - Default: mm+bulges)
Specifies sorting criteria for merging when using CFD/CRISTA scores.
Options include:
--sorting-criteria (Optional - Default: mm+bulges,mm)
Sorting criteria used when CFD/CRISTA scores are unavailable. Options are
similar to --sorting-criteria-scoring but tailored for simpler analyses.
Note 1: Ensure compatibility between input files and genome builds (e.g., hg38 or hg19) to avoid alignment issues.
Note 2: Optional arguments can be omitted when not applicable, but required arguments must always be specified.
The Complete Search function generates a comprehensive suite of reports, detailing the identified and prioritized targets, along with statistical and graphical summaries. These outputs are essential for interpreting the search results and understanding the impact of genetic diversity on CRISPR off-target.
Output Off-targets Files Description
*.integrated_results.tsv
Content type
Image
Digest
sha256:c8b908d99…
Size
4.5 GB
Last updated
4 days ago
docker pull pinellolab/crisprme