Sign inSign up

madeby561/vllm

By madeby561

•Updated 16 days ago

GLM-5.2 NVFP4 REAP vLLM: voipmonitor dark-devotion 2026-06-21 + MTP-draft DCP topk_scores_buffer pat

Image
0

2.9K

madeby561/vllm repository overview

⁠GLM-5.2 NVFP4 REAP - vLLM serving image (dcpglobaltopk b12x + MTP-draft DCP fix)

Base: voipmonitor/vllm:dark-devotion-df8ad3b-b12x5af873a-dcpglobaltopk-cu132-20260621 Added: a one-file patch to vllm/model_executor/models/deepseek_mtp.py that makes VLLM_DCP_SHARD_DRAFT=1 work on GLM-5.2 + b12x.

⁠What VLLM_DCP_SHARD_DRAFT does, and what we fixed

VLLM_DCP_SHARD_DRAFT=1 is the official flag that runs the MTP speculative draft DCP-parallel - the draft is sharded across the decode-context-parallel (DCP) ranks like the main model, instead of being replicated whole on every GPU. (It is the clean replacement for the old manual create_draft_parallel_config one-liner.)

But once the draft runs DCP-parallel (dcp_world_size > 1), its sparse attention must do top-k the same way the main model does under DCP: each rank holds part of the KV, computes its local candidate top-k, then the ranks all-gather their candidate scores and merge to the true GLOBAL top-k (global-topk). That cross-rank merge needs a scratch buffer - topk_scores_buffer - to hold each token`s per-rank scores during the gather.

The stock image allocates that buffer for the main model (DeepseekV2Model / GLM-5.2s GlmMoeDsaForCausalLM, both in deepseek_v2.py) but **NOT for the draft**: DeepSeekMTPModelindeepseek_mtp.pyonly allocatedtopk_indices_buffer. So the instant you set VLLM_DCP_SHARD_DRAFT=1` with GLM-5.2 + the b12x sparse indexer, the sharded draft hits the b12x DCP global-topk path, finds the scores buffer missing, and crashes during speculator cudagraph capture:

RuntimeError: B12X sparse indexer DCP requires topk_scores_buffer.
at vllm/model_executor/layers/sparse_attn_indexer.py:1596
(speculator.py _prefill -> deepseek_mtp.py forward -> sparse_attn_indexer)

⁠The fix (this image)

In deepseek_mtp.py, allocate topk_scores_buffer (shape [max_num_batched_tokens, index_topk], fp32) under the exact same gate the main model uses - decode_context_parallel_size > 1 and use_b12x_sparse_indexer() - and thread it through the drafts DeepseekV2DecoderLayer-> MLA attention ->Indexer->SparseAttnIndexer (that chain already accepts the arg; only the MTP wrapper omitted it). That is the whole patch: the sharded draft now has the buffer it needs for the cross-rank top-k merge, so **VLLM_DCP_SHARD_DRAFT=1+VLLM_DCP_GLOBAL_TOPK=1` + b12x run cleanly on GLM-5.2**.

⁠Measured (4x RTX PRO 6000, PCIe, TP4 / DCP4 / MTP5)

  • 70-94 t/s codegen single-stream
  • ~50 t/s @ 256k context (vs ~8 t/s unpatched -> decode stays flat at depth)
  • 256k cold prefill: no wedge
  • ~200 t/s aggregate @ concurrency 4 on the long-context estonia profile
  • KV pool 474k tokens @ util 0.95, MAX_MODEL_LEN=300000

⁠Run flags

VLLM_DCP_GLOBAL_TOPK=1, VLLM_DCP_SHARD_DRAFT=1, MTP=1, NUM_SPECULATIVE_TOKENS=5, DCP_SIZE=4, GPU_MEMORY_UTILIZATION=0.95, MAX_MODEL_LEN=300000, MAX_NUM_SEQS=8, MAX_NUM_BATCHED_TOKENS=8192, and a SPEC_CONFIG with use_local_argmax_reduction:true.

Weights not included - pair with a GLM-5.2-NVFP4-REAP checkpoint (e.g. madeby561/GLM-5.2-NVFP4-REAP-504B-term).

Upstream note: the gap is in the stock image too - deepseek_mtp.py should allocate topk_scores_buffer the same way DeepseekV2Model does. This image is just that fix baked on top.

Tag summary

Content type

Image

Digest

sha256:5ea2dad96…

Size

12 GB

Last updated

16 days ago

docker pull madeby561/vllm:mimo-v26-flash-b12x-20260923-rc4