Sign inSign up

mertcaglar/segm

By mertcaglar

โ€ขUpdated over 1 year ago

Histopathology image segmentation with segmentation models pytorch as the backbone.

Image
Machine learning & AI
Data science
1

183

mertcaglar/segm repository overview

โ ๐Ÿ”„ WSI Downscaling

To downscale WSIs:

  1. Run the dsconvert.py script.
  2. Update folder paths in the script to match your local directory structure.
  3. WSIs are split into patches (default: 2048x2048), and each patch is individually downscaled based on a specified scale factor.
  4. Segmentation masks are generated from GeoJSON annotations provided with the dataset.
  5. Downscaled patches are reassembled into .PNG images and masks using lossless compression.

โ ๐Ÿ“ Folder Structure

The expected folder structure is as follows:

ICIP2025/
โ”œโ”€โ”€ geojson/
โ”œโ”€โ”€ images/
โ”œโ”€โ”€ dataset/
โ”‚   โ”œโ”€โ”€ images/
โ”‚   โ”‚   โ”œโ”€โ”€ train/
โ”‚   โ”‚   โ”œโ”€โ”€ test/
โ”‚   โ”‚   โ””โ”€โ”€ validation/
โ”‚   โ””โ”€โ”€ annotations/
โ”‚       โ”œโ”€โ”€ train/
โ”‚       โ””โ”€โ”€ validation/

โ ๐Ÿณ Docker Support

A prebuilt Docker image is available:

docker pull mertcaglar/segm

To run the container with GPU support and mount the current working directory:

docker run --gpus all -it -v $(pwd):/host -w /host mertcaglar/segm

This will launch the container and run the adaptive_trainer.py script.


โ ๐Ÿงช Experimentation & Training

โ ๐Ÿงฌ Adaptive Data Augmentation

All training experiments use adaptive data augmentation, optimized via the OpenAI API. To use this feature:

  • Provide your API key in augmentation_strategy.py.
  • Adjust the prompt to define your augmentation policy.
  • adaptive_trainer.py will then apply the generated augmentations using the Albumentationsโ  library (Buslaev et al., 2020โ ).

โ ๐Ÿ› ๏ธ Training Configurations
DownscaleArchitecturePatch SizeStrideEncoderLoss Function
60xDPT22456maxvit_large_tf_224โ Dice
40xDPT51264maxvit_large_tf_512โ Jaccard
20xDPT512256maxvit_large_tf_512โ Dice
20xDPT512256maxvit_xlarge_tf_512โ Tversky
20xDPT512256maxvit_large_tf_512โ Lovasz

Backbone: MaxViT โ€“ Tu et al., "MaxViT: Multi-Axis Vision Transformer" (arXiv:2204.01697โ ) Hugging Face models via timmโ 


โ ๐Ÿ” Probability Matrix Inference

Once models are trained using the configurations above, generate probability maps for test images using:

infer_probs.py

Make sure to modify the script configuration to match your model and dataset settings.


โ ๐Ÿงฎ Top-N Soft Biased Voting

Use Top-N Soft Biased Voting to ensemble predictions across the best-performing models. To generate final masks from the probability matrices:

softvote_topN.py

This method improves accuracy by aggregating predictions from multiple models using soft-weighted voting.


โ ๐Ÿง  DPT Architecture Overview

DPT (Dense Prediction Transformer) is a vision transformer tailored for dense prediction tasks like semantic segmentation.

  • Replaces conventional convolutional backbones with a transformer-based encoder, enabling global context at every layer.
  • Intermediate transformer tokens are reconstructed into spatial feature maps and progressively decoded via a convolutional decoder.
  • Offers fine-grained segmentation and superior global consistency compared to CNN-based approaches.

DPT: Ranftl et al., "Vision Transformers for Dense Prediction" DPT Paper (arXiv:2103.13413)โ  Hugging Face Model Cardโ 


Tag summary

Content type

Image

Digest

sha256:6e123d02bโ€ฆ

Size

9.7 GB

Last updated

over 1 year ago

docker pull mertcaglar/segm