Sign inSign up

bioedge/edge_ubuntu18_bsve

By bioedge

•Updated over 7 years ago

Image
0

384

bioedge/edge_ubuntu18_bsve repository overview

⁠EDGE BSVE

⁠How to use this image

⁠Install Docker

See Docker at https://www.docker.com/⁠

⁠Obtain the docker image

docker pull bioedge/edge_ubuntu18_bsve

⁠Download database
# Pipeline database is ~18 Gb and contains the other databases needed for EDGE
  wget -c https://edge-dl.lanl.gov/EDGE/dev/edge_dev_pipeline_databases.tgz

# HOST genomes BWA index is ~48Gb for Host removal, including human, bacteria, phiX, viruses, invertebrate vectors of human pathogens
 wget -c https://edge-dl.lanl.gov/EDGE/dev/edge_dev_HostIndex.tgz

# NCBI Genomes is ~21Gb and contain the full genomes for prokaryotes and some viruses
 wget -c https://edge-dl.lanl.gov/EDGE/dev/edge_dev_NCBI_genomes.tgz

# GOTTCHA2 databases are 37Gb for the GOTTCHA2 taxonomic identification pipeline
 wget -c https://edge-dl.lanl.gov/EDGE/DB/edge_GOTTCHA2_db_20181115.tgz

# Kraken2 database is 26Gb contains the databases used for the Kraken2 taxonomic identification pipeline 
 wget -c https://edge-dl.lanl.gov/EDGE/DB/edge_Kraken2_db_20190104.tgz

# Centrifuge databases are 6.2Gb for the Centrifuge taxonomic identification pipeline
 wget -c https://edge-dl.lanl.gov/EDGE/DB/edge_Centrifuge_db_20181220.tgz

# PanGIA databases

(Optional)
# Other Host bwa index ~18Gb for host removal, including pig, sheep, cow, monkey, hamster. and goat.
 wget -c https://edge-dl.lanl.gov/EDGE/DB/edge_dev_otherHostIndex.tgz

Decompressed database files for later use.

⁠Start EDGE bioinformatics instance
    $ docker pull bioedge/edge_ubuntu_mysql
    $ docker create --name mysql_data --volume /var/lib/mysql  bioedge/edge_ubuntu_mysql
    $ docker run -d --privileged=true --security-opt "seccomp:unconfined" \
    --cap-add=SYS_ADMIN --cap-add=SYS_PTRACE  \
    --volumes-from mysql_data  \
    -v /path/to/database:/home/edge/database \
    -v /path/to/EDGE_output:/home/edge/EDGE_output \
    -v /path/to/EDGE_input:/home/edge/EDGE_input \
    -v /path/to/EDGE_report:/home/edge/EDGE_report \
    -p 80:80 -p 8080:8080 --name edge_bsve bioedge/edge_ubuntu18_bsve

Wait for a minute or so for the docker image to start EDGE service and Open http://localhost/⁠ on the browser to start experience EDGE GUI. See below sections for command line usage.

⁠Test Run
$ docker exec -it edge_bsve bash -c "/home/edge/edge/testData/runReadsTaxonomyTest/runTest.sh"
⁠Usage
$ docker exec edge_bsve bash -c "/home/edge/edge/runPipeline -h"

     Usage: perl /home/edge/edge/runPipeline [options] -c config.txt -p reads1.fastq reads2.fastq -o out_directory
     Version 2.4.0
     Input File:
            -u            Unpaired reads, Single end reads in fastq
            
            -p            Paired reads in two fastq files and separate by space

            -contigs      Contig Fasta File.

            -c            Config File
     Output:
            -o            Output directory.
  
     Options:
            -ref          Reference genome file in fasta        
                          It will find the genbank file (same prefix) in the same location if any.
            
            -cpu          number of CPUs (default: 8)

            -data_cleanup remove .sam .bam .gz .fastq .fq. tgz files after run finished.
 
            -version      print verison

⁠Example Config file
$ docker exec edge_bsve bash -c "cat /home/edge/edge/testData/runReadsTaxonomyTest/config.txt"

### There is a section for taxonomy classification. Four tools are enabled. ###
[Reads Taxonomy Classification]
## boolean, 1=yes, 0=no
DoReadsTaxonomy=1
## If reference genome exists, only use unmapped reads to do Taxonomy Classification. 
## Turn on AllReads=1 will use all reads instead.
AllReads=1
enabledTools=gottcha2-speDB-b,pangia,centrifuge,kraken2
###
⁠SRA download
$ docker exec edge_bsve bash -c "/home/edge/edge/scripts/sra2fastq.pl"

[DESCRIPTION]
    A script retrieves sequence project in FASTQ files from 
NCBI-SRA/EBI-ENA/DDBJ database using `curl` or `wget`. Input accession number
supports studies (SRP*/ERP*/DRP*), experiments (SRX*/ERX*/DRX*), 
samples (SRS*/ERS*/DRS*), runs (SRR*/ERR*/DRR*), or submissions 
(SRA*/ERA*/DRA*).

[USAGE]
     /home/edge/edge/scripts/sra2fastq.pl [OPTIONS] <Accession#> (<Accession# 2> <Accession# 3>...)

[OPTIONS]
    --outdir|d             Output directory
    --clean                clean up temp directory
    --platform-restrict    Only allow a specific platform
    --filesize-restrict    (in MB) Only allow to download less than a specific
	                       total size of files.
    --run-restrict         Only allow download less than a specific number
	                       of runs.
    --download-interface   curl or wget [default: curl]
    --help/h/?             display this help
⁠Default credentials for GUI
  • EDGE user: [email protected]⁠/admin

  • For security, you may need update the credentials if the server will be used by others or public.

⁠Note
  • The image size is around 12.4GB: If you have issue pulling this image, you may increase the basesize when launch docker daemon or using different Storage Driver⁠. see issue⁠.
  • This image is built on top of official Ubuntu 18.04, and is officially supported on Docker version 2.0.0.0-mac81 (29211)
  • The user management can be accessed by http://localhost:8080/userManagement⁠ if host port is 8080
⁠Contact Info

Chien-Chi Lo: [email protected]⁠

Tag summary

Content type

Image

Digest

Size

4.1 GB

Last updated

over 7 years ago

docker pull bioedge/edge_ubuntu18_bsve