Web-App for continuous tracking and filtering of SARS-CoV-2 mutations
10K+
Continuously tracking and filtering SARS-CoV-2 mutations
You can run the snakemake pipeline on a Linux system with anaconda installed or use the docker container. For a quick example you can also look into our Google Colab Notebook.
This pipeline requires a Linux system and conda to manage all of its dependencies.
Clone this repository:
git clone https://gitlab.com/dacs-hpi/covradar.git covradar
Create and activate conda environment:
cd covradar/backend
conda env create -n covradar_pipeline --file envs/covradar_pipeline.yml
conda activate covradar_pipeline
Note: Generating the PDF output of the pipeline further required libffi6. If it is not already installed on your system. you can download and install it via:
curl -LO http://archive.ubuntu.com/ubuntu/pool/main/libf/libffi/libffi6_3.2.1-8_amd64.deb
sudo dpkg -i libffi6_3.2.1-8_amd64.deb
Add data to backend/data. Example data:
mkdir data/charite
wget -O data/charite/charite-SARS-CoV-2.fasta.gz https://civnb.info/public/charite-SARS-CoV-2.fasta.gz
wget -O data/charite/charite-SARS-CoV-2.tsv.gz https://civnb.info/public/charite-SARS-CoV-2.tsv.gz
Create config.yaml for the pipeline configurations (see here for more details):
python create_config.py
# OR
nano config.yaml
With the example Charité data and python create_config.py, "samples" in config.yaml will be updated to:
samples:
charite:
genomes: 'data/charite/charite-SARS-CoV-2.fasta.gz'
metadata: 'data/charite/charite-SARS-CoV-2.tsv.gz'
Run the pipeline (needs activate conda environment):
snakemake --use-conda --cores 28
Clone this repository and switch to backend
git clone [email protected]:dacs-hpi/covradar.git covradar
cd backend/
Copy your data in an own directory covradar/backend/data/example_dir (see here)
Example data:
cp -r tests/testdata/mink/ data/
Pull and run docker:
docker pull dacshpi/covradar-pipeline
docker run --rm -v $(pwd):/covradar/backend dacshpi/covradar-pipeline /bin/bash -c "python create_config.py --modus pdf; snakemake --use-conda --cores 4"
When finished, you can access the results in: covradar/backend/results
The pipeline needs sequence data in (compressed) FASTA format and metadata in (compressed) TSV format in an extra subfolder in backend/data. It is possible to combine several data sources, for example in-house and data from the COVID-19 Data Portal.
.
├── backend
│ ├── data
│ ├── INHOUSE
│ ├── sequences.fasta
│ ├── meta.tsv
│ ├── EBI
│ ├── sequences.fasta
│ ├── meta.tsv
CovRadar accepts as input compressed FASTA files containing the genomes and TSV files with their respective metadata. The file extraction is automated by a shell script capable of recognizing and extracting multiple formats, namely: 7z (.7z), bzip2 (.bz2), gzip (.gz), RAR (.rar), TAR (.tar), TBZ2 (.tar.bz2 or .tbz2), TGZ (.tar.gz or .tgz), Z (.Z) and zip (.zip).
cp -r tests/testdata/mink/ data/
Metadata and sequence data are merged over the TSV strain column / FASTA headers. They need to be the same and unique.
Minimum required metadata fields in TSV:
Additional sequence filters in the web application can be used if also available:
covradar/backend/config.yaml contains the configurations for the snakemake pipeline and can be created by the provided Python script or by modifying the existing file.
nano config.yaml
# OR
python create_config.py
Note: It needs an activated conda environment.
conda env create -n covradar_pipeline --file envs/covradar_pipeline.yml
conda activate covradar_pipeline
The most important parameters to modify are:
pwd, e.g., /home/fabio/Desktop/covradar/backend;The file backend/data/voc.tsv containes all mutations to be used for the app's global map.
We are providing all characterstic spike mutations of Variants of Concern (≥75% frequency) taken from outbreak.info.
It is required for the pipeline's website mode (see here). You can add the path to your mutation table in config.yaml:
voc: 'data/voc.tsv'
backend/data/voc.tsv needs to be a TSV file with the following header:
nucleotide amino acid
Please be aware of our mutation nomenclature:
| Mutation | Nucleotide | Amino Acid |
|---|---|---|
| SNP | C21614T | L18F |
| Insertion | ins:22861:3:TAC ins:22873:6:TACTAC | Y433-434ins YY437-438ins |
| Deletion | del:21767:6 del:21992:3 | HV69-70del Y144-144del |
Combination of mutations are also possible:
| Combination | Nucleotide | Amino Acid |
|---|---|---|
| Two SNPs on different positions leading to one amino acid change | A22461G+G22462C | K300S |
| Two SNPs on the same position leading to one amino acid change | G22132C/G22132T | R190S |
| VOC is more than one amino acid mutation | G22161C,C22311A | Y200C, T250N |
After running the analytical pipeline, you can access the results in: covradar/backend/results.
results/benchmarks: run metrics for single rulesresults/extracted_spikes: extracted spike sequencesresults/filter_and_merge: discarded sequencesresults/logs: log metrics for single rulesresults/multiple_sequence_alignment: MSA and aligned index sequence (NC_045512.2)results/numbering: table with MSA position index positionresults/consensus: consensus sequences for each calendar week and countryresults/website: files important for the web applicationresults/website/bcftools: VCF statisticsresults/website/variant_calling: VCF generated from MSAresults/website/voc: used mutation file for global mapresults/website/voc_of_seqs: boolean table for each sequence mutation combinationYou can install the app manually on a Linux system with Anaconda installed or via Docker.
CovRadar is available for latest sequences from the COVID19-Data Portal at:
You can run the app without docker on a Linux system.
Clone Covradar
git clone https://gitlab.com/dacs-hpi/covradar.git covradar
Install and activate the environment.
conda env create -f frontend/environment.yml -n app_env
conda activate app_env
Install Redis for caching.
sudo apt install redis-server
sudo service redis-server {start | stop | status}
Test successful Redis installation
redis-cli
Type 'ping': 127.0.0.1:6379> ping
output will be PONG
leave Redis: 127.0.0.1:6379> exit
Add REDIS_CACHE_URL to .bashrc file:
export REDIS_CACHE_URL=redis://localhost:6379/0
If not done, add data into the database add data into the database.
Start the web server:
python server.py
This will start the app under http://localhost:8888/report.
We also provide a containerized environment to run the app. You only need Docker and docker-compose installed.
You can use the app as a Docker container together with a MySQL database. If you have a SQL dump, you can use docker-compose.
Add the data to frontend/db.
Example SQL dump (~130 MB)
mkdir frontend/db
wget https://osf.io/2h98n/download -O frontend/db/Spike.sql
Or add your own SQL dump to frontend/db, see add data into the database and creating an SQL dump.
# build the app and database container
docker-compose build
# run the app
docker-compose up
The build and up process will take a couple of minutes. After the successful launch, you can explore your data at http://0.0.0.0:8888/report/ on any browser to start.
Pull the docker container:
docker pull dacshpi/covradar:latest
Run the docker container, providing the MySQL credentials and database name stored in environment variables (see here):
docker run -e MYSQL_USER="$MYSQL_USER" -e MYSQL_PW="$MYSQL_PW" -e MYSQL_HOST="$MYSQL_HOST" \
-e MYSQL_DBglobal="$MYSQL_DBglobal" -p 127.0.0.1:8888:8888 dacshpi/covradar:latest
The credentials and database name come from the system environment variables. Add them into .bashrc with nano ~/.bashrc.
export MYSQL_USER=user_name
export MYSQL_PW=password
export MYSQL_HOST=localhost
export MYSQL_DBglobal=global
export MYSQL_DBlatest=latest
CovRadar's app is actually two dash apps integrated into one Flask app. On the landing page you have the option to select one of the two apps. The layout and functionality are the same, the difference is that the apps access their own database. These are defined in the environments: MYSQL_DBglobal and MYSQL_DBlates. You don't have to use two databases and you can define only one database.
To use an example dataset (results from Covid-19 Data Portal from March 2022), download:
# DOI 10.17605/OSF.IO/7XAJG
wget https://osf.io/2h98n/download -O Spike.sql
Create database and import SQL dump:
mysql -u USER -p -e "CREATE DATABASE Spike;"
mysql -u USER -p Spike < Spike.sql
# if you get the error: ERROR 1273 (HY000) at line 25: Unknown collation: 'utf8mb4_0900_ai_ci'
sed -i Spike.sql -e 's/utf8mb4_0900_ai_ci/utf8mb4_unicode_ci/g'
The app gets the database name over the enviroment variables:
export MYSQL_DBglobal=Spike
See also System environment variables.
After a successful pipeline (in website mode, see here run, we can create MySQL tables for the frontend.
Create and activate conda environment
conda env create -f backend/sql/environment.yml -n database
conda activate database
Create empty database
mysql -u $MYSQL_USER -p $MYSQL_PW -h $MYSQL_HOST -e "create database DB_NAME;"
upload backend/results to database DB_NAME
cd backend/sql
python create_spike_database.py -db DB_NAME
For info about options regarding database creation, use python create_spike_database -h.
Note that the mysql server needs to permit loading data from files on the database client. This can be done by adding
[mysqld]
local-infile
[mysql]
local-infile
to /etc/mysql/my.cnf
The app gets the database name over the enviroment variables:
export MYSQL_DBglobal=DB_NAME
See also System environment variables.
CovRadar's global map in the app shows the data on country level by default but it is also possible to show higher resolved locations like postal codes.
To visualize the mutations per location in the map and in location specific plots, the following must be done:
By assigning the sequences to different locations, the spatial resolution can be adjusted as desired.
Database schema table location_coordinates:
Fields (type): location_ID (int NOT NULL, primary key), name (varchar), lat (decimal(15,10)), lon (decimal(15,10))
This will create an SQL dump from the table DBNAME.
DBNAME=Spike
mysqldump --no-tablespaces -u $MYSQL_USER --password=$MYSQL_PW $DBNAME -h $MYSQL_HOST > $DBNAME.sql
This is necessary when merging different database in a single SQL dump, e.g., EBI + German RKI sequences. You can use this to use two databases with docker-compose.
DBNAME1=latest
DBNAME2=global
mysqldump -u $MYSQL_USER --password=$MYSQL_PW -h $MYSQL_HOST \
--skip-add-drop-table --databases $DBNAME1 $DBNAME2 > Spike.sql
CovRadar further offers small functionalities for nucleotide and amino acid translations and position conversions that can be executed without the need to run the pipeline beforehand. The scripts for these tasks are located in the covradar/shared/residue_converter/ folder. All the subtools are explained in detail in the README.md.
For the demand to visualize and investigate nucleotide and amino acid frequencies for specific positions and regions generated by the pipeline without using the app, CovRadar provides scripts in the covradar/analysis/ folder. How to use these scripts is explained in the README.md.
To perform a pipeline-independent extraction of gene sequences from a whole-genome Multi-FASTA file, the script gene_extraction.py within the folder shared/gene_extraction/ can be used. The detailed workflow and usage is explained in the README.md.
Content type
Image
Digest
sha256:971cc3ad3…
Size
1.3 GB
Last updated
over 3 years ago
docker pull dacshpi/covradar