Sign inSign up

knoweng/gene_prioritization_pipeline

By knoweng

•Updated over 8 years ago

Gene_Prioritization_Pipeline auto build

Image
0

109

knoweng/gene_prioritization_pipeline repository overview

⁠KnowEnG's Gene Prioritization Pipeline

This is the Knowledge Engine for Genomics (KnowEnG), an NIH, BD2K Center of Excellence, Gene Prioritization Pipeline.

This pipeline ranks the rows of a given spreadsheet, where spreadsheet's rows correspond to gene-labels and columns correspond to sample-labels. The ranking is based on correlating gene expression data (network smoothed) against pheno-type data.

There are four prioritization methods, using either pearson or t-test as the measure of correlation:

OptionsMethodParameters
Simple Correlationsimple correlationcorrelation
Bootstrap Correlationbootstrap sampling correlationbootstrap_correlation
Correlation with network regularizationnetwork-based correlationnet_correlation
Bootstrap Correlation with network regularizationbootstrapping w network correlationbootstrap_net_correlation

Note: all of the correlation methods mentioned above use the Pearson or t-test correlation measure method.


⁠How to run this pipeline with Our data


⁠1. Clone the Gene_Prioritization_Pipeline Repo
 git clone https://github.com/KnowEnG/Gene_Prioritization_Pipeline.git
⁠2. Install the following (Ubuntu or Linux)
apt-get install -y python3-pip
apt-get install -y libblas-dev liblapack-dev libatlas-base-dev gfortran
pip3 install numpy==1.11.1
pip3 install pandas==0.18.1
pip3 install scipy==0.19.1
pip3 install scikit-learn==0.17.1
apt-get install -y libfreetype6-dev libxft-dev
pip3 install matplotlib==1.4.2
pip3 install pyyaml
pip3 install knpackage
⁠3. Change directory to Gene_Prioritization_Pipeline
cd Gene_Prioritization_Pipeline
⁠4. Change directory to test
cd test
⁠5. Create a local directory "run_dir" and place all the run files in it
make env_setup
⁠6. Use one of the following "make" commands to select and run a clustering option:
CommandOption
make run_pearsonpearson correlation
make run_bootstrap_pearsonbootstrap sampling with pearson correlation
make run_net_pearsonpearson correlation with network regularization
make run_bootstrap_net_pearsonbootstrap pearson correlation with network regularization
make run_t_testt-test correlation
make run_bootstrap_t_testbootstrap sampling with t-test correlation
make run_net_t_testt-test correlation with network regularization
make run_bootstrap_net_t_testbootstrap t-test correlation with network regularization

⁠How to run this pipeline with Your data


Follow steps 1-3 above then do the following:

⁠* Create your run directory
mkdir run_directory
⁠* Change directory to the run_directory
cd run_directory
⁠* Create your results directory
mkdir results_directory
⁠* Create run_paramters file (YAML Format)
Look for examples of run_parameters in ./Gene_Prioritization_Pipeline/data/run_files/zTEMPLATE_GP_BENCHMARKS.yml
⁠* Modify run_paramters file (YAML Format)
set the spreadsheet, network and phenotype data file names to point to your data
⁠* Run the Gene Prioritization Pipeline:
  • Update PYTHONPATH enviroment variable
export PYTHONPATH='../src':$PYTHONPATH    
  • Run (in test directory with env_setup as described above)
python3 ../src/gene_prioritization.py -run_directory ./run_dir -run_file zTEMPLATE_GP_BENCHMARKS.yml

⁠Description of "run_parameters" file


KeyValueComments
methodcorrelation or net_correlation or bootstrap_correlation or bootstrap_net_correlationChoose gene prioritization method
correlation_measurepearson or t_testChoose correlation measure method
gg_network_name_full_pathdirectory+gg_network_namePath and file name of the 4 col network file
spreadsheet_name_full_pathdirectory+spreadsheet_namePath and file name of user supplied gene sets
phenotype_name_full_pathdirectory+phenotype_responsePath and file name of user supplied phenotype response file
results_directorydirectoryDirectory to save the output files
number_of_bootstraps5Number of random samplings
cols_sampling_fraction0.9Select 90% of spreadsheet columns
rwr_max_iterations100Maximum number of iterations without convergence in random walk with restart
rwr_convergence_tolerence1.0e-2Frobenius norm tolerence of spreadsheet vector in random walk
rwr_restart_probability0.5alpha in V_(n+1) = alpha * N * Vn + (1-alpha) * Vo
top_beta_of_sort100Number of top genes selected
top_gamma_of_sort50Number of top genes reported

gg_network_name = STRING_experimental_gene_gene.edge
spreadsheet_name = CCLE_Expression_ensembl.df
phenotype_name = CCLE_drug_ec50_cleaned_NAremoved_pearson.txt


⁠Description of Output files saved in results directory


  • Any method saves separate files per phenotype with name {phenotype}_{method}_{correlation_measure}_{timestamp}_viz.tsv. Genes are sorted in descending order based on visualization_score.
ResponseGene_ENSEMBL_IDquantitative_sorting_scorevisualization_scorebaseline_score
phenotype 1gene 1floatfloatfloat
...............
phenotype 1gene nfloatfloatfloat
  • Any method saves sorted genes for each phenotype with name ranked_genes_per_phenotype_{method}_{correlation_measure}_{timestamp}_download.tsv.
Rankingphenotype 1phenotype 2...phenotype n
1gene
(most significant)
gene
(most significant)
...gene
(most significant)
...............
ngene
(least significant)
gene
(least significant)
...gene
(least significant)
  • Any method saves spreadsheet with top ranked genes per phenotype with name top_genes_per_phenotype_{method}_{correlation_measure}_{timestamp}_download.tsv.
Genesphenotype 1...phenotype n
gene 11/0...1/0
............
gene n1/0...1/0

Tag summary

Content type

Image

Digest

Size

388.3 MB

Last updated

over 8 years ago

docker pull knoweng/gene_prioritization_pipeline:02_19_2018