Histopathology image segmentation with segmentation models pytorch as the backbone.
183
To downscale WSIs:
dsconvert.py script.2048x2048), and each patch is individually downscaled based on a specified scale factor.GeoJSON annotations provided with the dataset..PNG images and masks using lossless compression.The expected folder structure is as follows:
ICIP2025/
โโโ geojson/
โโโ images/
โโโ dataset/
โ โโโ images/
โ โ โโโ train/
โ โ โโโ test/
โ โ โโโ validation/
โ โโโ annotations/
โ โโโ train/
โ โโโ validation/
A prebuilt Docker image is available:
docker pull mertcaglar/segm
To run the container with GPU support and mount the current working directory:
docker run --gpus all -it -v $(pwd):/host -w /host mertcaglar/segm
This will launch the container and run the adaptive_trainer.py script.
All training experiments use adaptive data augmentation, optimized via the OpenAI API. To use this feature:
augmentation_strategy.py.adaptive_trainer.py will then apply the generated augmentations using the Albumentationsโ library (Buslaev et al., 2020โ ).| Downscale | Architecture | Patch Size | Stride | Encoder | Loss Function |
|---|---|---|---|---|---|
| 60x | DPT | 224 | 56 | maxvit_large_tf_224โ | Dice |
| 40x | DPT | 512 | 64 | maxvit_large_tf_512โ | Jaccard |
| 20x | DPT | 512 | 256 | maxvit_large_tf_512โ | Dice |
| 20x | DPT | 512 | 256 | maxvit_xlarge_tf_512โ | Tversky |
| 20x | DPT | 512 | 256 | maxvit_large_tf_512โ | Lovasz |
Backbone: MaxViT โ Tu et al., "MaxViT: Multi-Axis Vision Transformer" (arXiv:2204.01697โ ) Hugging Face models via timmโ
Once models are trained using the configurations above, generate probability maps for test images using:
infer_probs.py
Make sure to modify the script configuration to match your model and dataset settings.
Use Top-N Soft Biased Voting to ensemble predictions across the best-performing models. To generate final masks from the probability matrices:
softvote_topN.py
This method improves accuracy by aggregating predictions from multiple models using soft-weighted voting.
DPT (Dense Prediction Transformer) is a vision transformer tailored for dense prediction tasks like semantic segmentation.
DPT: Ranftl et al., "Vision Transformers for Dense Prediction" DPT Paper (arXiv:2103.13413)โ Hugging Face Model Cardโ
Content type
Image
Digest
sha256:6e123d02bโฆ
Size
9.7 GB
Last updated
over 1 year ago
docker pull mertcaglar/segm