
This pipeline is designed for image patch and whole slide level nuclei detection, segmentation, and TME feature extraction. The pipeline uses HD-Yolo algorithm by default for real-time nuclei segmentation across the whole slide and utilizes the results for TME feature analysis. There are three main functions/components in this repo. A docker image here contains all the code, model, and test data.
Rong R, Sheng H, Jin KW, Wu F, Luo D, Wen Z, Tang C, Yang DM, Jia L, Amgad M, Cooper LAD, Xie Y, Zhan X, Wang S, Xiao G. A Deep Learning Approach for Histology-Based Nucleus Segmentation and Tumor Microenvironment Characterization. Mod Pathol. 2023 Apr 24:100196. doi: 10.1016/j.modpat.2023.100196. Epub ahead of print. PMID: 37100227. [Pubmed]
If you have any questions or suggestions, please contact the following:
Developer: Ruichen Rong ([email protected])
Maintainer: Hudanyun Sheng ([email protected])
Corresponding: Shidan Wang ([email protected])
Corresponding: Xiaowei Zhan ([email protected])
Corresponding: Guanghua Xiao ([email protected])
The first component (run_patch_inference.py) runs a pre-trained object detection/segmentation model on image patches and outputs nuclei locations, types, and masks. This process can be done in real time and is integrated into our real-time deepzoom server.
| Sample input patch | Nuclei segmentation |
|---|---|
![]() | ![]() |
The second component (run_wsi_inference.py) applies a pre-trained object detection/segmentation model on whole slide images and outputs nuclei locations, types and masks. Results can be viewed through deep zoom server.
| Sample input slide | Nuclei segmentation |
|---|---|
![]() | ![]() |
The third component (summarize_tme_features.py) extracts nuclei morphological features and further utilizes a density-based feature extraction module to summarize slide level and ROI level TME features. All nuclei detected above will be allocated into a 2d point cloud and further smoothed into density maps that reflect nuclei distribution. Then multiple Delaunay graph-based and density-based TME features will be calculated.
| Nuclei scatter plot | Nuclei density plot |
|---|---|
![]() | ![]() |
Whole slide pathology image (WSI) analysis requires the necessary computer hardware settings due to the big size of the pathology image. The type 20X image size might range between 100MB and 600MB. Moreover, the A Nvidia GPU device is recommended to run the WSI analysis on a large dataset (our docker doesn’t support non-Nvidia GPU, such as intel integrated GPU or ATI GPU). Lastly, Function (a) and Function (b/c) has different hardware requirements due to the input file size). By default, this docker image is running CPU-only. If the Nvidia card is applied, the calculated speed for WSI data analysis may increase 2-10 times.
run_patch_inference.pyrun_wsi_inference.pysummarize_tme_features.pyCheck docker version
docker version
It is better to upgrade your version >= 23.*.*
docker pull utsw1qbrc/hd_wsi
docker tag utsw1qbrc/hd_wsi:latest hd_wsi:latest
Linux / Mac Os
# create the folder for the files to analyze
cd ~
mkdir hd_wsi_slides
# initiate the container
docker run --ipc=host -dit -v ~/hd_wsi_slides:/usr/src/hd_wsi/slides_folder --name hd_wsi hd_wsi:latest
docker exec -it hd_wsi /bin/bash; conda activate hd_env
Windows (Use Windows PowerShell to run the commands
# create the folder for the files to analyze
# Go to C disk
# Create a folder 'hd_wsi_slides'
# initiate the container
docker run --ipc=host -dit -v c:\hd_wsi_slides:/usr/src/hd_wsi/slides_folder --name hd_wsi hd_wsi:latest
docker exec -it hd_wsi /bin/bash
conda activate hd_env
The script run_patch_inference.py analyzes a single image patch or a folder of image patches with given model and outputs nuclei detection and segmentation results. Image patches are analyzed one-by-one without parallel, so image patches with different sizes are allowed. By default, model uses breast cancer HD-Yolo model to segment the following nuclei: tumor, stromal, immune, blood, macrophage, necrosis and others.
Input files: small image patch files; each path file size is <= 10MB
python run_patch_inference.py --data_path test_data/patch_example.png --output_dir /usr/src/hd_wsi/slides_folder/patch_result
# on your computer / server
# put one of your image patch file (such as patch1.png) under ~/hd_wsi_slides
python run_patch_inference.py --data_path /usr/src/hd_wsi/slides_folder/patch1.png --output_dir /usr/src/hd_wsi/slides_folder/patch_result
# on your computer / server
# put all your image patch files under ~/hd_wsi_slides
python run_patch_inference.py --data_path /usr/src/hd_wsi/slides_folder --output_dir /usr/src/hd_wsi/slides_folder/patch_result
You can read all the results under ~/hd_wsi_slides/patch_result
[Hardware Requirement]
Whole slide image (WSI) data processing requires the necessary CPU and memory size, please refer to 2.1.2
The default options of our pipeline is CPU-only. However, the pipeline supports NVidia GPU card. By adding GPU options, the calculate speed will increase 2-10 times due to GPU performance.
python -u run_wsi_inference.py --data_path test_data/wsi_sample.svs --output_dir /usr/src/hd_wsi/slides_folder/wsi_result_default
You can read all the results under ~/hd_wsi_slides/wsi_result_default
python -u run_wsi_inference.py --data_path sample.svs --model_path lung --output_dir /usr/src/hd_wsi/slides_folder/wsi_result_box_only --box_only --save_csv
You can read all the results under ~/hd_wsi_slides/wsi_result_box_only
# clean the folder ./slides on your computer/server
# put multiple WSI files under ./slides
# data process one wsi by one
python -u run_wsi_inference.py \
--data_path /usr/src/hd_wsi/slides_folder \
--meta_info text_color_info.yaml \
--model_path /path/to/model/ \
--output_dir /path/to/output_wsi \
--batch_size 4 --roi tissue
--save_img --save_csv \
--export_mask --export_text
The script summarize_tme_features.py takes the outputs from a) to extract a variety of nuclei and TME features. Currently, the following nuclei features and TME features are calculated:
By default, Delaunay graph based features are summarized from a maximum of 10 random selected patches (2048*2048) in tumor region. Density based features are calculated with kernel smoothed density map that is 1/32 of the original image scale. Parameters can be changed based on needs, run python summarize_tme_features.py -h for detailed information about all options.
Example 1. Analyze results from sample.svs and save plots with default settings.
python -u summarize_tme_features.py --model_res_path test_wsi --output_dir ./test_features --n_classes 3 --save_images
Example 2. Customize Delaunay graph based and density based features: randomly select 100 512x512 patches with more than 10 tumor and calculate density under 1/16
python -u summarize_tme_features.py \
--model_res_path /path/to/output_wsi \
--output_dir /path/to/output_tme \
--n_patches 100 --patch_size 512 \
--score_thresh 10 --scale_factor 16 \
--save_images \
slide_id_pred.pt: The torch file stores the outputs and inference time.slide_id_pred.png: The image file displays the segmentation/detection results.slide_id_pred.csv: The csv file contains boxes, scores, labels and masks info if --box_only is not triggered.slide_id.pt: The compressed object contains slide information, nuclei locations, scores and types, as well as inference time.slide_id.masks.pt: If model outputs masks and --box_only is not enabled, all nuclei masks shrinked into 28x28 pixel are stored in this file.slide_id.tiff: If --save_img is enabled, script will plot a large pyramid tiff image with the same size as input slide (nuclei color are provided through --meta_info with default transparency = 0.3). This file can be viewed through openslide and other tiff viewers. Don't enable this option for large image as it will take extremely long time to plot and save.slide_id.csv: If --save_csv is enabled, script will export boxes, scores, labels into this csv file. If --export_text is enabled, script will replace numeric labels with text labels defined in --meta_info. If --export_mask is enabled, an extra column contains masks in polygon format will be added to the csv file. Note that TME feature extraction pipeline takes slide_id.pt as input, the csv file is not necessary for downstream analysis. Export csv with text and masks will cost extra time and take more space.slide_id/feature_summary.csv: The csv file of slide_id x tme_features. Currently the following TME features are calculated.slide_id/feature_summary.pkl: The pickle file contains all the raw features without normalization/standardization (count, norm, dotproduct, etc.). This is useful when merging multiple slides under same patient.slide_id.slide_img.png: If --save_images is enabled, script will generate this thumbnail image.slide_id.scatter_img.png: If --save_images is enabled, script will generate a scatter plot for detection result.slide_id.density_img.png: If --save_images is enabled, script will export density plot for core nuclei type (by default, green: tumor, red: stromal, blue: immune).slide_id.roi_mask.png: If --save_images is enabled, script will export the roi region for feature extraction. (default is tumor region)Content type
Image
Digest
sha256:7f6a2bcf7…
Size
4.5 GB
Last updated
over 3 years ago
docker pull utsw1qbrc/hd_wsi