Sign inSign up

kbchoi/gbrs

By kbchoi

•Updated over 6 years ago

Image
1

1.2K

kbchoi/gbrs repository overview

GBRS is a suite of tools for reconstructing individualized diploid genomes using RNA-Seq data from multiparent population and quantifying allele specific expression. Although we tested it with mouse models only, GBRS should work for any multiparent populations. If you are interested in using it for other multiparent population model, please contact⁠ Kwangbom “KB” Choi at The Jackson Laboratory. For the Diversity Outbred (DO⁠), Collaborative Cross (CC⁠) mice, or F1 hybrids of CC’s (CCRIX), required data files are available at ftp://churchill-lab.jax.org/software/GBRS/⁠.

  • Free software: MIT license

⁠⁠Usage

First of all, please pull and run the image.

$ docker pull kbchoi/gbrs
$ docker run -i -t kbchoi/gbrs
(base) root@a6b69ed57991:/#

The environment variable, $GBRS_DATA is set and all the required files are available there already.

(base) root@a6b69ed57991:/# echo $GBRS_DATA
/usr/share/gbrs/R84-REL1505
(base) root@a6b69ed57991:/# ls -lh $GBRS_DATA
total 3.0G
-rw-r--r-- 1 root root  6.0M May 16  2016 avecs.npz
-rw-r--r-- 1 root root    80 May 16  2016 founder.hexcolor.info
-rw-r--r-- 1 root root   23M May 16  2016 gbrs.hybridized.targets.info
-rw-r--r-- 1 root root   627 May 16  2016 ref.fa.fai
-rw-r--r-- 1 root root  3.0M May 16  2016 ref.gene2transcripts.tsv
-rw-r--r-- 1 root root  155K May 16  2016 ref.gene_ids.ordered.npz
-rw-r--r-- 1 root root  380K May 16  2016 ref.gene_pos.ordered.npz
-rw-r--r-- 1 root root  1.5M Jul 12  2016 ref.genome_grid.64k.noYnoMT.txt
-rw-r--r-- 1 root root  1.6M May 16  2016 ref.genome_grid.64k.txt
-rw-r--r-- 1 root root  1.9M Mar 18  2019 ref.genome_grid.69k.noYnoMT.txt
-rw-r--r-- 1 root root  2.0M Mar 18  2019 ref.genome_grid.69k.txt
-rw-r--r-- 1 root root  2.6M May 16  2016 ref.transcripts.info
-rw-r--r-- 1 root root   16M May 16  2016 tranprob.DO.G0.F.npz
-rw-r--r-- 1 root root   16M May 16  2016 tranprob.DO.G0.M.npz
...
-rw-r--r-- 1 5156 10000 462M May 11  2016 transcripts.1.ebwt
-rw-r--r-- 1 5156 10000 187M May 11  2016 transcripts.2.ebwt
-rw-r--r-- 1 5156 10000 8.0M May 11  2016 transcripts.3.ebwt
-rw-r--r-- 1 5156 10000 373M May 11  2016 transcripts.4.ebwt
-rw-r--r-- 1 5156 10000 462M May 11  2016 transcripts.rev.1.ebwt
-rw-r--r-- 1 5156 10000 187M May 11  2016 transcripts.rev.2.ebwt

GBRS is installed in its own conda virtual environment. First of all, you have to “activate” the virtual environment by doing the following:

(base) root@a6b69ed57991:/# conda activate gbrs

The first step is to align our RNA-Seq reads against pooled transcriptome of all founder strains:

(gbrs) root@a6b69ed57991:/# bowtie -q -a --best --strata --sam -v 3 ${GBRS_DATA}/transcripts ${FASTQ} \
    | samtools view -bS - > ${BAM_FILE}

Before quantifying multiway allele specificity, bam file should be converted into emase format:

(gbrs) root@a6b69ed57991:/# gbrs bam2emase -i ${BAM_FILE} \
                 -m ${GBRS_DATA}/ref.transcripts.info \
                 -s A,B,C,D,E,F,G,H \
                 -o ${EMASE_FILE}

We can compress EMASE format alignment incidence matrix:

(gbrs) root@a6b69ed57991:/# gbrs compress -i ${EMASE_FILE} \
                -o ${COMPRESSED_EMASE_FILE}

If storage space is tight, you may want to delete ${BAM_FILE} or ${EMASE_FILE} at this point since ${COMPRESSED_EMASE_FILE} has all the information the following steps would need. If you want to merge emase format files in order to, for example, pool technical replicates, you run ‘compress’ once more listing files you want to merge with commas:

(gbrs) root@a6b69ed57991:/# gbrs compress -i ${COMPRESSED_EMASE_FILE1},${COMPRESSED_EMASE_FILE2},... \
                -o ${MERGED_COMPRESSED_EMASE_FILE}

and use ${MERGED_COMPRESSED_EMASE_FILE} in the following steps. Now we are ready to quantify multiway allele specificity:

(gbrs) root@a6b69ed57991:/# gbrs quantify -i ${COMPRESSED_EMASE_FILE} \
                -g ${GBRS_DATA}/ref.gene2transcripts.tsv \
                -L ${GBRS_DATA}/gbrs.hybridized.targets.info \
                -M 4 --report-alignment-counts

Then, we reconstruct the genome based upon gene-level TPM quantities (assuming the sample is a female from the 20th generation Diversity Outbred mice population)

(gbrs) root@a6b69ed57991:/# gbrs reconstruct -e gbrs.quantified.multiway.genes.tpm \
                   -t ${GBRS_DATA}/tranprob.DO.G20.F.npz \
                   -x ${GBRS_DATA}/avecs.npz \
                   -g ${GBRS_DATA}/ref.gene_pos.ordered.npz

We can now quantify allele-specific expressions on diploid transcriptome:

(gbrs) root@a6b69ed57991:/# gbrs quantify -i ${COMPRESSED_EMASE_FILE} \
                -G gbrs.reconstructed.genotypes.tsv \
                -g ${GBRS_DATA}/ref.gene2transcripts.tsv \
                -L ${GBRS_DATA}/gbrs.hybridized.targets.info \
                -M 4 --report-alignment-counts

Genotype probabilities are on a grid of genes. For eQTL mapping or plotting genome reconstruction, we may want to interpolate probability on a decently-spaced grid of the reference genome.:

(gbrs) root@a6b69ed57991:/# gbrs interpolate -i gbrs.reconstructed.genoprobs.npz \
                   -g ${GBRS_DATA}/ref.genome_grid.69k.noYnoMT.txt \
                   -p ${GBRS_DATA}/ref.gene_pos.ordered.npz \
                   -o gbrs.interpolated.genoprobs.npz

To plot a reconstructed genome:

(gbrs) root@a6b69ed57991:/# gbrs plot -i gbrs.interpolated.genoprobs.npz \
            -o gbrs.plotted.genome.pdf \
            -n ${SAMPLE_ID}

Tag summary

Content type

Image

Digest

Size

2.8 GB

Last updated

over 6 years ago

docker pull kbchoi/gbrs