Container environment to launch Caulifinder pipelines that detect all Caulimoviridae in a genome
1.6K
The Dockerfile contains the instruction to create the necessary environment for Caulifinder from event_caulifinder master (https://forgemia.inra.fr/urgi-anagen/event_caulifinder). It will also install REPET with its dependencies. This docker builds on the giovtorres/docker-centos7-slurm:17.02.9 image.
/!\ Caulifinder image will not work on debian
To pull the image from the docker repository, run:
>docker pull urgi/docker_caulifinder
Create 'myproject' directory, copy in it your genome 'mygenome.fa'
>mkdir myproject
>chmod oug+w myproject
>cd myproject
>cp PATHTO/mygenome.fa .
From 'myproject' folder, execute an interactive bash shell on the container with mount volume on 'myproject' folder.
docker run -ti -h ernie --volume $PWD:/home/centos --name caulifinder_docker urgi/docker_caulifinder:latest
In the container, your are in /home/centos directory containing different databanks and pre-configured files used by caulifinder.
The directory also contains the script to retrieve the CDD database, run:
/home/centos> . forCDD_database.sh
To launch caulifinder Branch A, run:
/home/centos> Caulifinder_branch_A.py -g mygenome.fa -l Caulimoviridae_ref_genomes.fa -b baits.fa -c baits.fa -db Cdd_LE_db/Cdd 1>& Caulifinder_branch_A.txt &
To launch caulifinder Branch B, run:
/home/centos> Caulifinder_branch_B.sh mygenome.fa 1>& Caulifinder_branch_B.txt &
From the 'caulifinder_docker' container, you quit and still run it with 'ctrl+p ctrl +q' command. So you are returned in your 'myproject' folder with all results (files and folders) :
/home/centos> ctrl+p ctrl+q
From the 'caulifinder_docker' container, you quit and stop it with 'exit' command. So you are returned in your 'myproject' folder with all results (files and folders) :
/home/centos> exit
To come back to 'caulifinder_docker' container, from your 'myproject' folder:
docker container rm -f caulifinder_docker
docker run -ti -h ernie --volume $PWD:/home/centos --name caulifinder_docker urgi/docker_caulifinder:latest
Caulifinder is a tool that allows discovering Caulimoviridae sequences in plant genomes. It is composed of two branches with complementary purposes: Branch A builds consensus sequences of endogenous Caulimoviridae Branch B produces a phylogenetic analysis of marker genes.
Please read the following for more details regarding the process and usage of each branch.
Branch A - Caulifinder_branch_A.py - aims at constructing a library of consensus sequences that are representative of repetitive elements with significant homology to Caulimoviridae found in a plant genome. This library typically contains complete Caulimoviridae genomes, almost complete genomes and genomic fragments. Branch A also builds groups of consensus sequences sharing high similarity in order to establish tentative biological links between consensus. This shall allow for instance grouping the different parts of multipartite genomes and connecting short fragments with longer sequences.
The main steps of branch A are as follows:
The following files are provided to run branch A:
Before launching: The working directory must also contain the input genome provided by user (headers should NOT contain space characters or symbols such as "=", ";", ":", "|"...).
Mandatory parameters:
Additional options
To launch, in background, Caulifinder_branch_A.py on myGenome.fa, using default names for outputs, collect the standard output/standard error in Caulifinder_branch_A.txt, run:
Caulifinder_branch_A.py -g mygenome.fa -l Caulimoviridae_ref_genomes.fa -b baits.fa -c baits.fa -db Cdd_LE_db/Cdd 1>& Caulifinder_branch_A.txt &
To launch, in background, Caulifinder_branch_A.py using -o option and an extension set to 3000 bp, run:
Caulifinder_branch_A.py -g mygenome.fa -l Caulimoviridae_ref_genomes.fa -b baits.fa -c baits.fa -db Cdd_LE_db/Cdd -X 3000 -o results_3000.csv 1>& Caulifinder_branch_A_3000.txt &
Main output files: in summary_branchA/
Branch B - Caulifinder_branch_B.sh - aims at collecting amino acid RT sequences from endogenous Caulimoviridae and to use these, in combination to reference sequences, to produce a phylogenetic tree. The classification suggested by the pylogenetic tree can be more accurate than that proposed by best hit in Branch A, especially with sequences that do not fall within any reference clade. In addition, branch B also allows detecting low copy RT loci by contrast to branch A which inherently only collects repetitive elements.
The main steps of branch B are as follows:
The following files are provided to run branch B:
In addition, the input genome file in fasta format should be in the running directory (headers in genome should NOT contain space characters or symbols such as "=", ";", ":", "|"...).
To launch branch B, run:
Caulifinder_branch_B.sh mygenome.fa 1>& launch.txt &
Main output files: in summary_branchB/
All intermediate files created during the workflow are kept. For instance, cluster sizes can be checked in UCLUST output file "results.clstr"
Content type
Image
Digest
Size
2.3 GB
Last updated
about 4 years ago
docker pull urgi/docker_caulifinder