Single Variant Association pipeline for whole genome sequence analysis
5.1K
This workflow implements a single variant association analysis of genotype data with a single trait using mixed models. The primary code is written in R using the GENESIS package for model fitting and association testing. The workflow can either generate a null model from traits and covariates and a genetic related matrix or use a pre-generated null model obtained from GENESIS.
Also included is a workflow for regenerating the final summary results of the single variant association pipeline. This summary workflow takes a CSV file of association results (with the same format of the final summary output) and regenerates plots and output files.
This workflow is produced and maintained by the Manning Lab. Contributing authors include:
All code and necessary packages are available through a Docker image as well as through the Github repository.
This script preprocesses the phenotype file for a conditional analysis on some snps, extracting the genotype information and adding to the phenotypes for each sample.
Inputs:
Outputs :
This script generates a null model to be used in association testing in Genesis
Inputs:
Outputs:
This script performs an association test to generate p-values for each variant included.
Inputs:
Outputs:
Generate a summary of association results including quantile-quantile and manhattan plots for variants subseted by minor allele frequency (all variants, maf < 5%, maf >= 5%). Also generates CSV files of all variants and variants with P < pval_threshold.
Inputs:
Outputs:
The main outputs of the single variant association workflow come from summary.R. Two CSV files are produced:
Each of these files have a single row for each variant included (multi-allelics are split with a single row per alternate allele). Each file has a minimum set of columns with a few extra columns that depend on the type of association test performed. The columns that are always present in both output CSV files are:
1. MarkerName: a unique identifier of the variant with the format: `chromosome-position-reference_allele-alternate_allele`, ex: `chr1-123456-A-C` (string)
2. chr: chromosome of the variant (string)
3. pos: position of the variant (int)
4. ref: reference allele of the variant as encoded in the input GDS file (string)
5. alt: alternate allele of the variant as encoded in the input GDS file (string)
6. minor.allele: allele with lower maf between ref and alt
7. maf: frequency of `minor.allele` in sample tested for association (float)
8. pvalue: pvalue generated in association analysis (float)
9. n: number of samples used in generated association statistics (int)
10. mac: minor allele count
If the trait of interest is dichotomous:
11. homref.case: number of homozygous reference cases (int)
12. homref.control: number of homozygous reference controls (int)
13. het.case: number of heterozygous reference cases (int)
14. het.control: number of heterozygous reference controls (int)
15. homalt.case: number of homozygous alternate cases (int)
16. homalt.control: number of homozygous alternate controls (int)
If the statistical test is the Score test:
16. Score.stat: score statistic of the variant (float)
If the statistical test is the Wald test:
9. Wald.Stat: Wald statistic of the variant (float)
Below is an example output file using 1000 genomes data and simulated phenotypes. Note that case/control genotype counts is only calculated for variants with minor allele frequency < 5% and pvalue < 0.01. (maf, pvalue, and Score.Stat have been truncated for display only)
| MarkerName | chr | pos | ref | alt | minor.allele | maf | pvalue | n | Score.Stat | homref.case | homref.control | het.case | het.control | homalt.case | homalt.control |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| chr22-16052271-G-A | 22 | 16052271 | G | A | alt | 0.0075878 | 0.0598677 | 2504 | 3.541045 | ||||||
| chr22-16052986-C-A | 22 | 16052986 | C | A | alt | 0.0740814 | 0.0013632 | 2504 | 10.25490 | ||||||
| chr22-16911246-C-T | 22 | 16911246 | C | T | alt | 0.0229632 | 0.0049379 | 2504 | 7.902014 | 485 | 1909 | 11 | 94 | 0 | 5 |
| chr22-17008128-C-T | 22 | 17008128 | C | T | alt | 0.0021964 | 0.5074979 | 2504 | 0.439222 | ||||||
| chr22-17008980-C-T | 22 | 17008980 | C | T | alt | 0.0658945 | 0.7537240 | 2504 | 0.098428 | ||||||
| chr22-17009923-G-C | 22 | 17009923 | G | C | alt | 0.0119808 | 0.0024070 | 2504 | 9.209963 | 475 | 1969 | 21 | 39 | 0 | 0 |
summary.R also produces a PNG file with Manhattan and quantile-quantile plots. Three plots are shown in this image (from top-left):

Content type
Image
Digest
Size
3.4 GB
Last updated
about 7 years ago
docker pull manninglab/singlevariantassociation