Sign inSign up

malavp/vllm-glm52

By malavp

•Updated 23 days ago

An image for serving GLM 5.2 (lowbitcoffee/GLM-5.2-W4A16) on NVIDIA SM80 architecture via vLLM.

Image
0

277

malavp/vllm-glm52 repository overview

An image for serving GLM 5.2 (lowbitcoffee/GLM-5.2-W4A16) on NVIDIA SM80 architecture. Tested on 8xA100-80GB.

⁠Image Details

Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/glm52-a100⁠

⁠Serving Recipe

Served with

  --tensor-parallel-size=8
  --max-num-seqs=64
  --tool-call-parser=glm47
  --enable-auto-tool-choice
  --reasoning-parser=glm45
  --max-model-len=262144
  --gpu-memory-utilization=0.95
  --enable-chunked-prefill
  --max-num-batched-tokens=4096
  --dtype=bfloat16
  --safetensors-load-strategy=lazy
  --load-format=instanttensor
  -cc.pass_config.fuse_allreduce_rms=False

Tag summary

Content type

Image

Digest

sha256:9e755a13f…

Size

8.5 GB

Last updated

23 days ago

docker pull malavp/vllm-glm52:sm80