Sign inSign up

malavp/vllm-glm53flash

By malavp

•Updated 3 days ago

An image for serving zai.org/GLM-5.3-Flash on SM80. Single request decode 93 tok/s on 8xA100-80GB.

Image
0

2.2K

malavp/vllm-glm53flash repository overview

An image for serving zai.org/GLM-5.3-Flash on NVIDIA SM80 architecture. Tested on 8xA100-80GB with single stream decode speed of ~93 tok/s.

⁠Image Details

Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/glm53flash-a100⁠

⁠Serving Recipe

Served with

  --tensor-parallel-size=8
  --max-num-seqs=32
  --gpu-memory-utilization=0.95
  --max-model-len=262144
  --tool-call-parser=glm47
  --reasoning-parser=glm45
  --enable-auto-tool-choice
  --safetensors-load-strategy=lazy
  --load-format=instanttensor
  --no-enable-flashinfer-autotune
  --attention-backend=TRITON_MLA_SPARSE
  --kv-cache-dtype=bfloat16

Tag summary

Content type

Image

Digest

sha256:0a5d4d85f…

Size

8.1 GB

Last updated

3 days ago

docker pull malavp/vllm-glm53flash:sm80