Original Quality Functional Equivalent Pipeline
1M+
oqfe image do?This app processes whole genome sequencing (WGS) or whole exome sequencing (WES) data - it aligns raw sequencing data to the human reference genome and additionally performs duplicate marking. It generates an OQFE CRAM file from a FASTQ or mapped CRAM file, following the Original Quality Functionally Equivalent (OQFE) protocol, which is a revision of the Functionally Equivalent (FE) protocol.
The app follows the FE methodology by: the adoption of a standard GRCh38 reference genome with alternate loci, read alignment with BWA-MEM, inclusion of supplementary alignments, duplicate marking that includes the supplementary alignments, CRAM compression and restricted tag usage.
In accordance with the OQFE protocol, the app retains the original quality scores of the reads, which allows for recovery of original FASTQs from the generated OQFE CRAM, and implements updated versions of the used programs.
It can be used to process WGS or WES data according to an agreed, widely adopted standard, which results in less variability in the downstream analyses of the data.
This application requires input files in either gzipped FASTQ (*.fastq.gz or *.fq.gz) or CRAM (*.cram) format.
For FASTQ inputs, a single array of files is needed for unpaired reads, and two file arrays (left and right read mates) are required for paired reads.
For CRAM input, a single mapped and merged CRAM file should be provided, as well as a reference genome file needed for CRAM decompression (the same that was used to create the CRAM file). Additionally, the flag reuse_cram_header can be set to specify if the Read Group (RG) header should be obtained directly from the input CRAM header or derived from read names (default behavior). The input CRAM RG must meet OQFE requirements (see below).
Sample name is required regardless of the input format. It is added to the RG header (RG SM) and used to create output file names.
A reference genome sequence for reads alignment is not expected since it is locked in the image and cannot be overridden, according to the FE protocol. It is set to GRCh38 reference genome with alternate loci from the 1000 Genome Project.
Finally, optical duplicate pixel distance for "Picard MarkDuplicates" is a parameter that is set in the app to 2500 and can be overridden with the --optical-duplicate-pixel-distance flag.
This app outputs the mappings as a coordinate-sorted CRAM file (*.oqfe.cram) and a corresponding mappings index (*.oqfe.crai). It also returns a duplication metrics file (*.oqfe.markdup_stats.txt).
The following is the summary of steps performed by the OQFE pipeline run in the Docker container:
If a CRAM file is passed on the input, it is converted to FASTQ files by first splitting the CRAM file into BAM by read names (with "samtools split"), then sorting the BAM (with "samtools sort -n"), and finally converting each BAM into FASTQ files.
The FASTQ files, provided directly to the app or extracted from an input CRAM, are aligned using BWA-MEM with parameters standardized by the FE protocol. Each FASTQ pair is mapped separately. MC and MQ tags are added using samblaster with the "-a --addMateTags" parameters.
According to the FE protocol, the header line for the RG should contain minimally the tags: ID, PL, PU, SM, and LB. Only if all of these tags are present in the input CRAM can they be used for the output OQFE CRAM header, otherwise the app will return an error and exit.
OQFE CRAM constructs the RG header line as follows:
ID: ID will be [SampleName].[LaneID]
PL: Platform will default to ILLUMINA unless otherwise specified
PU: Platform unit is obtained from the first line in the FASTQ (FlowcellID.LaneID)
SM: Sample name is app input parameter
LB: LB tag will be populated with Flowcell ID and is obtained from the first line in the FASTQ
Each BAM file is name-sorted with "sambamba sort -n".
The resulting BAM files from all lanes are merged with "sambamba merge".
Duplicates are marked by "Picard MarkDuplicates" with the parameters "ASSUME_SORT_ORDER = queryname" and OPTICAL_DUPLICATE_PIXEL_DISTANCE = 2500 (can be overridden), and the results are coordinate-sorted using "sambamba sort".
The duplicate-marked BAM is converted to CRAM and indexed using "samtools".
To run OQFE with a CRAM file as an input, execute:
docker run -v $(pwd):/home -w /home/ \
dnanexus/oqfe \
-1 /home/input.cram \
--sample SampleId \
--cram-reference-fasta /home/ref_genome_hs38DH.fa
Mounting the current directory to /home in the running Docker container with -v will make the local files (in this case "input.cram" and "ref_genome_hs38DH.fa") available to the oqfe application in the container. The -w option will ensure the outputs will be in your local directory when the Docker container finished running and exits. Note that /home is an example and can be replaced with a different directory in the container.
Optionally, you can pass the --optical-duplicate-pixel-distance <num>flag, which is used in the "Picard MarkDuplicates" execution (by default it is set to 2500)
In order to see full help message of the application run:
docker run dnanexus/oqfe -h
Content type
Image
Digest
Size
4.3 GB
Last updated
over 6 years ago
docker pull dnanexus/oqfe