An image for serving zai.org/GLM-5.3-Flash on SM80. Single request decode 93 tok/s on 8xA100-80GB.
2.2K
An image for serving zai.org/GLM-5.3-Flash on NVIDIA SM80 architecture. Tested on 8xA100-80GB with single stream decode speed of ~93 tok/s.
Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/glm53flash-a100
Served with
--tensor-parallel-size=8
--max-num-seqs=32
--gpu-memory-utilization=0.95
--max-model-len=262144
--tool-call-parser=glm47
--reasoning-parser=glm45
--enable-auto-tool-choice
--safetensors-load-strategy=lazy
--load-format=instanttensor
--no-enable-flashinfer-autotune
--attention-backend=TRITON_MLA_SPARSE
--kv-cache-dtype=bfloat16
Content type
Image
Digest
sha256:0a5d4d85f…
Size
8.1 GB
Last updated
3 days ago
docker pull malavp/vllm-glm53flash:sm80