An image for serving GLM 5.2 (lowbitcoffee/GLM-5.2-W4A16) on NVIDIA SM80 architecture via vLLM.
277
An image for serving GLM 5.2 (lowbitcoffee/GLM-5.2-W4A16) on NVIDIA SM80 architecture. Tested on 8xA100-80GB.
Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/glm52-a100
Served with
--tensor-parallel-size=8
--max-num-seqs=64
--tool-call-parser=glm47
--enable-auto-tool-choice
--reasoning-parser=glm45
--max-model-len=262144
--gpu-memory-utilization=0.95
--enable-chunked-prefill
--max-num-batched-tokens=4096
--dtype=bfloat16
--safetensors-load-strategy=lazy
--load-format=instanttensor
-cc.pass_config.fuse_allreduce_rms=False
Content type
Image
Digest
sha256:9e755a13f…
Size
8.5 GB
Last updated
23 days ago
docker pull malavp/vllm-glm52:sm80