Tool for the generation of pseudo-long-reads
815
PLR-GEN is a tool for the generation of pseudo-long-reads (PLRs) by using short-reads of a metagenomic sample and microbial reference genome sequences as input.
You can see PLR-GEN Github
Input single-end or paired-end NGS short reads
-s (incompatible with -1 and -2)-1 and -2 (incompatible with -s)A list of reference genome sequences
-r or -ref option (incompatible with -tama)with a list of all the file paths of reference genomes to set up the references for running PLR-GEN(Additional) If TAMA is installed, you can predict putative bacterial species in your dataset using TAMA, running PLR-GEN using -tama option (incompatible with -r).
Running PLR-GEN
PLR-GEN.pl [options] -1 <pe1> -2 <pe2> (or -s <se>) -r <ref_list> -o <out_dir>
Input options
-s : an input sequence file with unpaired reads (single-end reads); a fastq or fastq.gz file-1 and -2 : input sequence files with paired-end reads; fastq or fastq.gz files-r or -ref : a list with all file paths of reference genome sequences; a text file-tama : if this option is used, TAMA program will be used for reference preparation instead of -r (default: not used)-sampling : if this option is used, reference genomes will be randomly selected and used (default: not used)Output options
-o : a path of output directory (default: ./PR_out; final output: ./PR_out/pseudo-long-read.fa)-t : if this option is used, all intermediated output files are saved (default: not used)Running options
-p : the number of threads-l or -min_length : minimum length cutoff of generated pseudo-long reads (PLRs) (default: 100)-q or -mapq : cutoff of mapping quality (used in samtools) (default: 20)-c or -min_count : cutoff of mapping depth for generating normal nodes and bubbles; after piling-up of read mapping, alignment column less than the cutoff value will be discarded (default: 1)-d or -min_depth : cutoff of mapping depth of bubbles for filtering bubble nodes; a bubble with less than the cutoff % of mapping depth will be converted to a normal nodeHelp page of PLR-GEN
Usage: PLR-GEN.pl [options] -1 <pe1> -2 <pe2> (or -s <se>) -r <ref_list> -o <out_dir>
== MANDATORY
-s <se> File with unpaired reads [incompatible with -1 and -2]
-1 <pe1> File with #1 mates (paired 1) [incompatible with -s]
-2 <pe2> File with #2 mates (paired 2) [incompatible with -s]
-r|-ref <ref_list> The list of reference genome sequence files
-tama Reference preparation using TAMA [incompatible with -r|-ref]
-sampling <proportion> proportion to random sampling for references (default: off, range: 0-1)
-o <out_dir> Output directory (default: ./PR.out)
==Running and filtering options
-p|-core <integer> The number of threads (default: 1)
-q|-mapq <integer> Minimum mapping quality (default: 20)
-l|-min_length <integer> Cutoff of minimum length of pseudo-long reads (default: 100bp)
-c|-min_count <integer> Cutoff of minimum mapping depth for each node (default: 1)
-d|-min_depth <integer> Cutoff of mapping depth of bubbles (default: 1, range: 0-100)
0: all bubbles are used.
1: bubbles with less than 1% mapping depth from mapping depth distribution of bubbles are converted to normal nodes.
100: all bubbles are converted to normal nodes
==Other options
-t|-temp If you use -t option, all intermediate files are left.
Please careful to use this option because it has to be needed very large space.
-h|-help Print help page.
-tama option for preparation of reference genomes, you can install TAMA using below commands. It will automatically download and install TAMA into the given path, and set ready-made species-level TAMA databases. These TAMA databases require total 300GB disk space, so please carefully set up the path for installation of TAMA. If you do not specify a path for TAMA, the TAMA package and TAMA databases will be set inside "bin" directory in the PLR-GEN.Installation of TAMA with Docker
docker run --rm -v /PATH/TO/TAMA_DIR:/tama_dir -t jkimlab/plrgen:latest /work_dir/src/TAMA_install.pl /tama_dir
Manual installation
-tama option for PLR-GEN, you need to install TAMA and ready-made species-level databases into /LOCAL_DISK/PLR-GEN/bin/ directory or make a link of path of installed TAMA directory to /LOCAL_DISK/PLR-GEN/bin/ directory.For more information of TAMA, see TAMA Github
Suppose (1) you pulled the PLR-GEN docker image (jkimlab/plrgen), (2) read_1.fq, read_2.fq are in /LOCAL_DISK/DATA directory, and (3) all reference sequence files are in /LOCAL_DISK/DATA/fasta/ directory. In this situation, you can run PLR-GEN with docker image as followed command,
Make a reference list file in the /LOCAL_DISK/DATA directory with the file system for the docker image, for an example,
file: /DATA/reference_list.txt
/data_dir/fasta/REF_1.fa
/data_dir/fasta/REF_2.fa
/data_dir/fasta/REF_3.fa
Run the docker image, mounting /DATA directory to /data_dir directory of docker image, for an example,
docker run -v /LOCAL_DISK/DATA:/data_dir -t jkimlab/plrgen PLR-GEN.pl -1 /data_dir/read_1.fq -2 /data_dir/read_2.fq -r /data_dir/reference_list.txt -o /data_dir/output
If you want to use TAMA installed at /LOCAL_DISK/PROGRAMS/TAMA, you need to mount the TAMA directory to /work_dir/bin/ directory of docker image, for an example,
docker run -v /LOCAL_DISK/PROGRAMS/TAMA:/work_dir/bin/TAMA -v /LOCAL_DISK/DATA:/data_dir -t jkimlab/plrgen PLR-GEN.pl -1 /data_dir/read_1.fq -2 /data_dir/read_2.fq -r /data_dir/reference_list.txt -o /data_dir/output
Content type
Image
Digest
Size
1 GB
Last updated
over 4 years ago
docker pull jkimlab/plrgen