Sign inSign up

knoweng/general_clustering_pipeline

By knoweng

•Updated almost 7 years ago

General_Clustering_Pipeline auto build.

Image
0

97

knoweng/general_clustering_pipeline repository overview

⁠KnowEnG's General Clustering Pipeline

This is the Knowledge Engine for Genomics (KnowEnG), an NIH BD2K Center of Excellence, General Clustering Pipeline.

This pipeline clusters a spreadsheet's columns, with various methods:

OptionsMethodParameters
K-meansK Meanskmeans
hierarchical clusteringhierarchical clusteringhclust
Linked hierarchical clusteringhierarchical clustering constraintlink_hclust
Bootstrapped hierarchical clusteringconsensus hierarchical clusteringcc_ hclust
Bootstrapped K-meansconsensus K Meanscc_kmeans
Bootstrapped Linked hierarchical clusteringconsensus linked hierarchical clusteringcc_link_hclust

⁠How to run this pipeline with Our data


⁠1. Clone the General_Clustering_Pipeline Repo
 git clone https://github.com/KnowEnG-Research/General_Clustering_Pipeline.git
⁠2. Install the following, for Linux
 apt-get install -y python3-pip libfreetype6-dev libxft-dev libblas-dev liblapack-dev libatlas-base-dev gfortran
 pip3 install pyyaml knpackage scipy==0.19.1 numpy==1.11.1 pandas==0.18.1 matplotlib==1.4.2 scikit-learn==0.17.1 
⁠3. Change directory to General_Clustering_Pipeline
cd General_Clustering_Pipeline
⁠4. Change directory to test
cd test
⁠5. Create a local directory "run_dir" and place all the run files in it
make env_setup
⁠6. Use one of the following "make" commands to select and run a clustering option:
CommandOption
make run_kmeans_binaryClustering with k-means
make run_kmeans_continuous
make run_hclust_binaryHierarchical Clustering
make run_hclust_continuous
make run_link_hclust_binaryHierarchical linkage Clustering
make run_link_hclust_continuous
make run_cc_kmeans_binaryConsensus Clustering with k-means
make run_cc_kmeans_continuous
make run_cc_hclust_binaryConsensus Hierarchical Clustering
make run_cc_hclust_continuous
make run_cc_link_hclust_binaryConsensus Hierarchical linkage Clustering

⁠How to run this pipeline with Your data


Follow steps 1-5 above then do the following:

⁠* Create your run directory
mkdir run_dir
⁠* Change directory to the run directory
cd run_dir
⁠* Create your results directory
mkdir results
⁠* Create run_paramters file (YAML Format)
Look for examples of run_parameters in the General_Clustering_Pipeline/data/run_files zTEMPLATE_cc_hclust.yml
⁠* Modify run_paramters file (YAML Format)

Change processing_method to one of: serial, parallel depending on your machine.

processing_method: serial

set the data file targets to the files you want to run, and the parameters as appropriate for your data.

⁠* Run the General Clustering Pipeline:
  • Update PYTHONPATH enviroment variable
export PYTHONPATH='../':$PYTHONPATH    
  • Run
python3 -m kngeneralclustering.general_clustering -run_directory ./run_dir -run_file zTEMPLATE_cc_net_nmf.yml

⁠Description of "run_parameters" file


KeyValueComments
methodkmeans,hclust,link_hclust,cc_kmeans, cc_hclust, cc_link_hclustChoose clustering method
affinity_metriceuclidean, manhattan, jaccardChoose clustering affinity
linkage_criterionward, complete, averageChoose clustering affinity
spreadsheet_name_full_pathdirectory+spreadsheet_namePath and file name of user supplied gene sets
results_directorydirectoryDirectory to save the output files
tmp_directory./run_dir/tmpDirectory to save the temporary files
number_of_clusters3Estimated number of clusters
number_of_bootstraps4Number of bootstraps for cc_kmeans, cc_hclust and cc_link_hclust
rows_sampling_fraction0.8Select 80% of spreadsheet rows
cols_sampling_fraction0.8Select 80% of spreadsheet columns
top_number_of_rows10Top number of features to analyze
processing_methodserial or parallel or distributeChoose processing method
parallelismnumber of coresSet number of cores for speed or memory
threshold10Threshold to define categorical data and continuous data in evaluation toolbox
nearest_neighbors10Number of Nearest Neighbors in cc_link_hclust method

spreadsheet_name = EXPR_GSE_METABRIC_lymphN_binary.tsv.gz


⁠Description of Output files saved in results directory


  • Output files of all methods save row by col heatmap variances per row with name row_variance_{method}_{timestamp}_viz.tsv.
variance
row 1float
......
row mfloat
  • Output files of all the methods save row by col heatmap with name row_by_col_heatmp_{method}_{timestamp}_viz.tsv.
col 1...col n
row 1float...float
............
row mfloat...float
  • Output files of all methods save col to cluster map with name col_labeled_by_cluster_{method}_{timestamp}_viz.tsv.
cluster
col 1int
......
col nint
  • Output files of all methods save row scores by cluster with name row_averages_by_cluster_{method}_{timestamp}_viz.tsv.
cluster 1...cluster k
row 1float...float
............
row mfloat...float
  • Output files of all methods save spreadsheet with top ranked rows per column with name top_row_by_cluster_{method}_{timestamp}_download.tsv.
cluster 1...cluster k
row 11/0...1/0
............
row m1/0...1/0
  • All methods save three silhouette scores: silhouette overall score, silhouette per cluster score and silhouette per sample with name silhouette_{method}_{timestamp}_viz.tsv.
    1. silhouette overall score file: | number of clusters | silhouette score |

    2. silhouette per cluster score file: | ith clusters | corresponding silhouette score |

    3. silhouette per sample score file: | ith sample | corresponding silhouette score|

Tag summary

Content type

Image

Digest

Size

470 MB

Last updated

almost 7 years ago

docker pull knoweng/general_clustering_pipeline:mjberry_create_package