Estimate NGS Dataset qualitity in terms of its ability to detect mutations of predefined spectrum
972
EphaGen - a package to estimate NGS Dataset qualitity in terms of its ability to detect mutations of predefined spectrum #REQUIREMENTS
To run EphaGen you need:
#Run EphaGen To view run options execute:
docker run m4merg/ephagen perl /EphaGen/src/ephagen.pl
This will produce Help message describing input/output files description and parameters. To run EphaGen execute:
docker run -v /path/to/data/:/data/:rw m4merg/ephagen perl /EphaGen/src/ephagen.pl --bam /data/demo_data.bam --ref /data/brac.fasta --vcf /data/BRAC.vcf --out /data/demo_out.tsv --out_vcf /data/demo_out.vcf
#Notes: /path/to/data - folder must contain all required input files (including bam file "--bam"; reference file "--ref"; vcf file "--vcf"). Output data will be stored here as well. Note that bam index files also should be stored in /path/to/data folder before running EphaGen.
--ref FILE Path to reference genome fasta (required)
--bam FILE Path to sample BAM file (required). BAM file should be aligned to the reference specified with "--ref" option, sorted and indexed.
--vcf FILE Path to VCF file with target mutations. Note that mutation positions in VCF file should be in concordance with input reference genome file. Required unless --vcf_ref used.
--out FILE Path to output file containing result dataset sensitivity analysis results. This is tab-separated file where first raw contains information on mean coverage and sensitivity for the whole dataset. Further lines contain results of downsample analysis based on random sampling of reads from the whole datasets. For each read fraction random sampling is performed several times. Mean sensitivity is written for each read fraction. Mean coverage calculation is carried only across defined mutation sites.
--out_vcf FILE Path to output VCF file containing sensitivity analysis per each mutation site
#OUTPUT FILE DESCRIPTION
EphaGen will produce two output files: general sensitivity analysis results in tsv format defined by "--out" option and VCF with probabilities for each mutation from input VCF file defined by "--out_vcf" option.
General sensitivity analysis results provide sensitivity calculation results for each downsample iteration line by line. If no downsample was carried out, file contains only one line for 100% of reads. In addition to sensitivity mean coverage is provided.
Output VCF file copies the input VCF files provided with "--vcf" option with INFO section is expanded. Field "SSS" is added into INFO section providing information for Single Site Sensitivity for each mutation:
##INFO=<ID=SSS,Number=R,Type=Float,Description="Single Site Sensitivity">
Variants in the output VCF file are sorted based on the false negative rate in descending order.
#Limitations Of Usage Please make sure that:
-EphaGen handle only Single Nucleotide Variation, Multiple Nucleotide Variations (up to 50b.p. long), short insertions and deletions (up to 50b.p. long). Large genomic rearrangements, CNV, exon deletions and insertions are not supported.
-Reference VCF files, stored in the /reference directory (for BRCA and CFTR pathogenic mutation analysis) are made based on databases BREAST CANCER INFORMATION CORE and CFTR2. Allele counts taken from these databases refer to general population and may be inappropriate for some populations, especially for minor populations. Moreover, allele frequency spectrum based on these allele counts can possess bias towards over-represented variations.
Content type
Image
Digest
Size
392.2 MB
Last updated
over 7 years ago
docker pull m4merg/ephagen