Sign inSign up

dnanexus/oqfe

By dnanexus

•Updated over 6 years ago

Original Quality Functional Equivalent Pipeline

Image
1

1M+

dnanexus/oqfe repository overview

⁠Original Quality Functional Equivalent (OQFE) Pipeline - Docker Image

⁠What does the app in the oqfe image do?

This app processes whole genome sequencing (WGS) or whole exome sequencing (WES) data - it aligns raw sequencing data to the human reference genome and additionally performs duplicate marking. It generates an OQFE CRAM file from a FASTQ or mapped CRAM file, following the Original Quality Functionally Equivalent (OQFE) protocol, which is a revision of the Functionally Equivalent (FE) protocol⁠.

The app follows the FE methodology by: the adoption of a standard GRCh38 reference genome with alternate loci, read alignment with BWA-MEM, inclusion of supplementary alignments, duplicate marking that includes the supplementary alignments, CRAM compression and restricted tag usage.

In accordance with the OQFE protocol, the app retains the original quality scores of the reads, which allows for recovery of original FASTQs from the generated OQFE CRAM, and implements updated versions of the used programs.

⁠What are typical use cases for this image?

It can be used to process WGS or WES data according to an agreed, widely adopted standard, which results in less variability in the downstream analyses of the data.

⁠What data are required to run it?

This application requires input files in either gzipped FASTQ (*.fastq.gz or *.fq.gz) or CRAM (*.cram) format.

For FASTQ inputs, a single array of files is needed for unpaired reads, and two file arrays (left and right read mates) are required for paired reads.

For CRAM input, a single mapped and merged CRAM file should be provided, as well as a reference genome file needed for CRAM decompression (the same that was used to create the CRAM file). Additionally, the flag reuse_cram_header can be set to specify if the Read Group (RG) header should be obtained directly from the input CRAM header or derived from read names (default behavior). The input CRAM RG must meet OQFE requirements (see below).

Sample name is required regardless of the input format. It is added to the RG header (RG SM) and used to create output file names.

A reference genome sequence for reads alignment is not expected since it is locked in the image and cannot be overridden, according to the FE protocol. It is set to GRCh38 reference genome with alternate loci⁠ from the 1000 Genome Project.

Finally, optical duplicate pixel distance for "Picard MarkDuplicates" is a parameter that is set in the app to 2500 and can be overridden with the --optical-duplicate-pixel-distance flag.

⁠What does this app output?

This app outputs the mappings as a coordinate-sorted CRAM file (*.oqfe.cram) and a corresponding mappings index (*.oqfe.crai). It also returns a duplication metrics file (*.oqfe.markdup_stats.txt).

⁠How does the OQFE application work?

The following is the summary of steps performed by the OQFE pipeline run in the Docker container:

⁠CRAM to FASTQ conversion

If a CRAM file is passed on the input, it is converted to FASTQ files by first splitting the CRAM file into BAM by read names (with "samtools split"), then sorting the BAM (with "samtools sort -n"), and finally converting each BAM into FASTQ files.

⁠Mapping

The FASTQ files, provided directly to the app or extracted from an input CRAM, are aligned using BWA-MEM with parameters standardized by the FE protocol. Each FASTQ pair is mapped separately. MC and MQ tags are added using samblaster with the "-a --addMateTags" parameters.

⁠ReadGroup (RG) header line specification

According to the FE protocol, the header line for the RG should contain minimally the tags: ID, PL, PU, SM, and LB. Only if all of these tags are present in the input CRAM can they be used for the output OQFE CRAM header, otherwise the app will return an error and exit.

OQFE CRAM constructs the RG header line as follows:

  • ID: ID will be [SampleName].[LaneID]

    • SampleName is the app input parameter
    • Lane is obtained from the first line in the FASTQ
    • Example:
    • @A00001:122:HWXYZDSXX:1:1101:1000:11083 RG:Z:Coriell_NA12878_NA12878-09102019.1
    • Lane: 1 (fourth field split by colon)
    • ID: [SampleName].1
  • PL: Platform will default to ILLUMINA unless otherwise specified

  • PU: Platform unit is obtained from the first line in the FASTQ (FlowcellID.LaneID)

    • Example:
    • @A00001:122:HWXYZDSXX:1:1101:1000:11083 RG:Z:Coriell_NA12878_NA12878-09102019.1
    • PU: HWXYZDSXX.1 (third and fourth fields split by colon)
  • SM: Sample name is app input parameter

  • LB: LB tag will be populated with Flowcell ID and is obtained from the first line in the FASTQ

    • Example:
    • @A00001:122:HWXYZDSXX:1:1101:1000:11083 RG:Z:Coriell_NA12878_NA12878-09102019.1
    • LB: HWXYZDSXX (third field split by colon)
⁠Sorting

Each BAM file is name-sorted with "sambamba sort -n".

⁠Merging

The resulting BAM files from all lanes are merged with "sambamba merge".

⁠Duplicate marking

Duplicates are marked by "Picard MarkDuplicates" with the parameters "ASSUME_SORT_ORDER = queryname" and OPTICAL_DUPLICATE_PIXEL_DISTANCE = 2500 (can be overridden), and the results are coordinate-sorted using "sambamba sort".

⁠CRAM generation and indexing

The duplicate-marked BAM is converted to CRAM and indexed using "samtools".

⁠Example

To run OQFE with a CRAM file as an input, execute:

docker run -v $(pwd):/home -w /home/ \
    dnanexus/oqfe \
    -1 /home/input.cram \
    --sample SampleId \ 
    --cram-reference-fasta /home/ref_genome_hs38DH.fa

Mounting the current directory to /home in the running Docker container with -v will make the local files (in this case "input.cram" and "ref_genome_hs38DH.fa") available to the oqfe application in the container. The -w option will ensure the outputs will be in your local directory when the Docker container finished running and exits. Note that /home is an example and can be replaced with a different directory in the container.

Optionally, you can pass the --optical-duplicate-pixel-distance <num>flag, which is used in the "Picard MarkDuplicates" execution (by default it is set to 2500)

In order to see full help message of the application run:

docker run dnanexus/oqfe -h
⁠Last updated: June 12th 2020

Tag summary

Content type

Image

Digest

Size

4.3 GB

Last updated

over 6 years ago

docker pull dnanexus/oqfe