Sign inSign up

stilldeadcode/vllm-radiance

By stilldeadcode

•Updated about 1 month ago

Image
18

10K+

stilldeadcode/vllm-radiance repository overview

⁠vllm-radiance

vLLM inference server for the AMD Radeon AI PRO R9700 (gfx1201 / RDNA4). Bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack with the RDNA4 patches and custom kernels needed to run vLLM on this card, so you don't have to build the stack yourself.

⁠Tested so far

ModelsQwen3.8-27B-FP8 / Qwen3.6-27B-FP8 (gated-delta-net hybrids, architecturally identical), Qwen3.6-35B-A3B-FP8 (fine-grained MoE, 256 experts top-8), Gemma-4-31B-it-FP8 (dense, sliding + global attention, vision)
KV cachefp8, bf16 or auto
GPUs2x R9700, tensor parallel (TP=2)

Untested: any other model, non-FP8 weights, single GPU, more than two GPUs, non-R9700 hardware. The defaults below are a starting point for these setups, not a general recommendation.

Qwen3.8-27B-FP8 / Qwen3.6-27B-FP8 are what everything here is tuned around: 64 layers, 48 of them linear attention (GDN) and 16 full attention, hidden 5120, head_dim 256 with 6 query heads per KV head. That is exactly the geometry the hand-written kernels are compiled for, so the whole R4D library engages on them. The MTP head is inside the checkpoint, so speculative decoding needs no separate drafter. They are vision-language checkpoints -- pass --language-model-only to skip the vision tower when serving text.

Qwen3.6-35B-A3B-FP8. RDNA4-tuned fused-MoE configs and the skinny MoE-gate GEMM activate automatically. One requirement: with --mamba-cache-mode=align its attention block size is 2240 and align asserts block_size <= max_num_batched_tokens, so pass --max-num-batched-tokens >= 2240.

Gemma-4-31B-it-FP8 (e.g. RedHatAI/gemma-4-31B-it-FP8-block). Quantization is auto-detected; tuned block-FP8 GEMM configs load for its shapes. Text and vision both work. Not a GDN hybrid, so drop --mamba-cache-mode, and it uses its own chat template and parsers. Long-context prefill is tuned for its head-512 global layers (up to -38% TTFT at 64K, -46% at 120K; inert on other head sizes). It carries a lot of KV (60 layers), so at a given --gpu-memory-utilization it wants a smaller --max-model-len than the Qwen models. For a large decode speedup pair it with Google's drafter google/gemma-4-31B-it-assistant, whose head-512 layer needs "attention_backend":"ROCM_AITER_UNIFIED_ATTN" in the speculative config.

⁠Why this exists

vLLM's ROCm builds target datacenter cards (MI300 / CDNA). RDNA4 workstation cards like the R9700 (gfx1201) don't work out of the box: AITER isn't enabled for the arch, GPU enumeration fails, several kernels need patching, and the vendor attention and GEMM paths aren't tuned for RDNA4. This image pins a working combination, builds AITER from source for gfx1201, applies the fixes, and adds tuned kernels.

⁠Stack

Everything below is compiled from source for gfx1201 in the image build (0.5.0 onward); nothing is pulled from a prebuilt wheel index.

ComponentVersion
vLLM0.27.1
PyTorch2.11.0
Triton3.6.0
torchvision0.24.1
AITER0.1.17
transformers5.14.1 (pinned)
ROCm userspace7.14, bundled
Baserocm/dev-ubuntu-24.04:7.14.0-full (Ubuntu 24.04, Python 3.12)

The PyTorch / Triton / torchvision versions are the ones upstream builds vLLM against on ROCm, not a newer combination chosen for this image. Read the ROCm numbers, not pyproject.toml: 0.27.1's build-system asks for torch == 2.13.0, which is the CUDA build, while upstream's own ROCm image builds torch release/2.11 with torchvision 0.24.1 and pins no torch in requirements/rocm.txt. That distinction is deliberate: see the tensor-parallel hang note below.

transformers is pinned because vLLM does not pin it -- requirements/common.txt asks only for transformers >= 5.5.3, so an unpinned rebuild picks up whatever is newest and the stack moves underneath the build. transformers 5.15.0 made Gemma-4's head_dim a per-layer attribute and turned the global read into an exception that no released vLLM handles, so a Gemma-4 checkpoint fails during argument parsing, before a model or an attention backend exists. 5.14.1 is the last release before that change and loads every architecture this image serves. If you build your own image, keep the pin.

There is no flash-attention package: the vendor flash kernels have no gfx1201 device code. Attention runs on the AITER unified path, and the vision tower on the image's own Triton flash kernel.

⁠What it patches (to make vLLM run on gfx1201)

  • GPU enumeration (amdsmi init order). Without it the platform is undetected and device count reads 0.
  • AITER enablement for gfx12x (upstream gates it to MI3xx).
  • Triton driver activation for the GPU-less model-inspection subprocess.
  • Native sampler fallback (AITER's top-k/top-p kernel doesn't build on RDNA4).
  • Tool-parser streaming vs non-streaming consistency.
  • from_json Jinja filter for tool-calling chat templates.
  • MTP drafter unpadding, so --speculative-config's disable_padded_drafter_batch:true works (the single-stream MTP speed path).
  • MTP drafter multimodal mask alignment, so speculative decoding works with image inputs (otherwise the vision-placeholder mask outlives the compacted draft batch and the engine crashes).
  • torch.compile telemetry JSON encoding, which otherwise raises TypeError: Object of type function is not JSON serializable at startup on this torch version (harmless but alarming: the serve came up anyway).

⁠Custom kernels and tuning (on by default, env-gated)

The hand-written kernels are a separate library, libr4d⁠, written for gfx1201 rather than for any one model. Each entry point is named for the geometry it is compiled for and refuses anything else, so import r4d; r4d.kernels() inside the image lists exactly what it covers. The build clones a pinned tag and compiles it with its own hipcc.

Env varDefaultWhat it does
RADIANCE_USE_R4D1Master switch for the hand-written gfx1201 kernels: the paged attention behind --attention-backend R4D; the whole gated-delta-net layer for hybrid linear-attention models (2.80x on the fused prefill scan in isolation, +1.8-2.2% prefill end to end); the head_dim-72 vision-encoder kernel; the TP=2 all-reduce; and the skinny bf16 GEMM. Each is compiled for a specific geometry and declines anything else, so all of it is inert on a model it does not fit. The startup log prints which kernel each part of the model resolved to. Set 0 and every path reverts to stock -- the quickest way to tell whether a problem is ours or upstream's.
RADIANCE_PRESHUFFLE1preshuffled AITER FP8 blockscale GEMM
RADIANCE_SKINNY_GEMM1Route bf16 projections too small for rocBLAS to fill the machine to the R4D split-K kernel, for M in [6,64]. 1 covers the shapes that are a clear win alone (the MoE gate on fine-grained MoE models: 9.6 -> 3.2 us). all adds shapes that differ from rocBLAS at a bf16 ULP on a few elements in ten thousand -- notably the gated-delta-net in_proj_ba, 480 KiB run 48 times per step, 28.5 us against 3.6 us. Under speculative decoding a ULP-level change can move drafting acceptance, which is why those are not in the default set; measured on a DFlash2 drafter over four paired compiles, all was worth -3.9% on the decode step with no acceptance cost.
RADIANCE_GDN_META1build the gated-delta-net attention metadata with numpy on the host instead of the stock tensor path. Byte-identical output, and it removes host work from every step of a hybrid linear-attention model.
RADIANCE_FAST_DRAFT0The tuned drafter stack, as one switch -- no sub-knobs; each was a sweep and the answer is baked in. Each half engages only where it applies, so it is safe to leave on across models. The draft head goes to 2 bits with an exact rerank (any drafter): 0.167 GiB/rank instead of 1.18, the coarse pass emits the best 8 candidates of each 64-wide block for free and the top 32 are rescored exactly against the bf16 weight -- +16.6% tokens/s single-stream, +12.5% at 8 concurrent, acceptance unchanged, and exact (it matches the bf16 argmax on all 8192 real inputs tested). A dflash drafter's weights go to 4 bits (inert under mtp): signed symmetric int4, one f16 scale per 128 input channels, no zero point, 4.25 bits/weight, on two gfx1201 kernels -- f16 matrix-core at or below 16 rows, int8 above, because this chip's f16 matrix instruction is half the rate of its int8 one. Codes are derived at load from the weight: no calibration data, no offline step, nothing on disk. On Qwen3.8-27B + DFlash2 the draft pass falls 9.1% at a drafter batch of 64; with RADIANCE_SKINNY_GEMM=all the decode step falls 5.1% for +3.5% tokens/s. It cannot change what the model emits -- a drafter only chooses what is proposed, and the target verifies every token with its own weights. Pair with RADIANCE_DRAFT_TAU=0.28 under mtp.
(always on)Shard-local draft confidence. The draft controller reads the drafted token id and its top-1 probability from each rank's own vocab shard rather than gathering the full logits, which is exact and cuts the per-step all-reduce by 42%.
RADIANCE_USE_R4D_AR1custom PCIe peer-to-peer all-reduce for TP=2, byte-identical to RCCL, falls back to RCCL if P2P is unavailable
RADIANCE_USE_R4D_AR_QUANT1Compress the all-reduce payload: each group of 64 is rotated by a Walsh-Hadamard, scaled by its own amplitude and stored in 6 uniform bits, so a message costs 6.25/16 of its bf16 size. +7.2% prefill at 16K context, +3.5% at 32K; decode untouched. NOT bit-identical to RCCL, though the two TP ranks stay bit-identical to each other. Set 0 for the exact bf16 all-reduce.
RADIANCE_FUSE_RMS_QUANT1folds group-FP8 quant into the RMSNorm epilogue
RADIANCE_DYNAMIC_DRAFT1Dynamic MTP draft depth: per request, a per-slot confidence gate decides how deep to draft (up to num_speculative_tokens) and whether to take a verbatim n-gram continuation -- deep on high-acceptance content like code and JSON, shallow on prose. Lossless. mtp only, by mechanism: it works by stopping a serial loop of draft forwards early, and a dflash drafter emits every position in one forward pass at a depth fixed when its CUDA graph is captured, so this does nothing there.
RADIANCE_DRAFT_SCHEDULE1:8,2:7,4:6,8:5,16:4bs:max_depth pairs (carry-forward): caps how many serial MTP forwards run at each batch size, so drafting stays deep single-stream and shallower at concurrency. The free n-gram tail is unaffected.
RADIANCE_DRAFT_TAU0.35Confidence-product stop threshold: keep drafting while the running product of the drafter's top-1 confidences stays >= TAU. Lower = deeper. The baked 0.35 suits the default bf16 head; with RADIANCE_FAST_DRAFT=1 use 0.28 (+5.3% over keeping 0.35), since a cheaper draft step lowers the acceptance a position must clear.
--attention-backend R4Doff (opt-in CLI flag, not an env var)R4D attention: purpose-built gfx1201 attention kernels, in place of the tuned AITER unified attention. Prefill and decode are hand-written HIP built around a transposed score matrix, S^T = K.Q^T, so a wave32 matrix-core fragment gives each lane exactly one query row and the softmax stays inside the lane. In the serve the prefill kernel is 1.65x the AITER one: +14.6% prefill throughput at 64K context (attention is 34% of prefill GPU time there), +4.1% at 16K (11.8%), decode unchanged within noise. More accurate, not less: 1.69e-03 relative to an fp32 oracle against 2.28e-03. Requires head_dim 256, paged block 16, 6 query heads per KV head, causal decoder attention and a bf16 or fp8_e4m3 KV cache; any other shape is refused at startup with the reason. Give the drafter the same backend with "attention_backend": "R4D" inside --speculative-config.
RADIANCE_RUN_BWTEST1run the GPU topology + bandwidth sweep at startup (rocm-bandwidth-test, compiled into the image): device list, P2P access matrix, NUMA distances, and peak uni/bidirectional copy bandwidth per agent pair. Backgrounded and takes about a second, so it never delays the serve; the report lands in the log a few seconds in. Set 0 to skip it.
RADIANCE_NUMA_BINDunset (off)NUMA pinning for multi-node hosts; see below. Same as --numa-bind, which wins if both are given
RADIANCE_BANNER_PLAIN0set 1 for a startup banner without ANSI colour (log scrapers, CI). NO_COLOR does the same

R4D attention (--attention-backend R4D, opt-in). Purpose-built attention kernels for this GPU instead of the tuned AITER path. The core is a transposed score matrix, S^T = K.Q^T: a wave32 matrix-core fragment splits a 16x16 tile column-wise, so with the score matrix transposed each lane owns exactly one query row and the softmax never leaves the lane. Measured in the serve against ROCM_AITER_UNIFIED_ATTN, same image and flags: +14.6% prefill throughput at 64K context (65.6K-token prompt, 21.25 s -> 18.54 s to first token), +4.1% at 16K, decode unchanged within noise. The gain scales with context because attention's share of prefill does -- 34% of prefill GPU time at 64K, 11.8% at 16K -- so this is a TTFT feature, not a tokens/s one. It is also more accurate than what it replaces: 1.69e-03 relative error against an fp32 oracle, against 2.28e-03.

Shape support is narrow on purpose: head_dim 256, paged block 16, 6 query heads per KV head, causal decoder attention, bf16 query, bf16 or fp8_e4m3 KV. Anything else is refused at startup with the reason and the backends that would work instead. With speculative decoding, give the drafter the same backend: "attention_backend":"R4D" in the speculative config.

All of these are baked ON in the image. Set RADIANCE_DYNAMIC_DRAFT=0 to turn draft control off (RADIANCE_DRAFT_SCHEDULE and RADIANCE_DRAFT_TAU are values, not toggles). RADIANCE_DYNAMIC_DRAFT only does anything when speculative decoding is enabled; it is lossless (it changes only how many tokens are drafted and whether they come from MTP or a verbatim copy of earlier text, never what the model verifies).

NUMA pinning (RADIANCE_NUMA_BIND / --numa-bind, off by default). On a multi-NUMA-node host, pin the server and its TP workers to the node(s) local to the GPUs. auto detects from the visible GPUs; SPEC may also be explicit nodes, bind=, interleave, preferred= or none. No-op on single-node hosts; needs --cap-add SYS_NICE under Docker's default seccomp.

⁠Requirements

  • AMD Radeon AI PRO R9700 (gfx1201). Compiled for gfx1201 only, won't run on other GPUs. Two GPUs (TP=2) is the only configuration tested so far.
  • Linux host with the amdgpu kernel driver and /dev/kfd + /dev/dri. ROCm userspace is inside the image.
  • Docker with device passthrough.

⁠Run

On start the image prints a banner and a short preamble (GPU count, gfx1201 check, P2P, enabled optimizations, component versions), then hands off to vllm serve. First argument is the model path, the rest are vllm serve flags. The RADIANCE_* vars below are already baked ON in the image; they are spelled out here only so they are visible and easy to flip off.

docker run --rm -it \
  --device /dev/kfd --device /dev/dri \
  --group-add "$(getent group render | cut -d: -f3)" \
  --group-add "$(getent group video  | cut -d: -f3)" \
  --shm-size 4g --cap-add SYS_PTRACE --security-opt seccomp=unconfined \
  -v /path/to/models:/models:ro \
  -v "$PWD/vllm-cache:/cache" \
  -p 8000:8000 \
  -e HIP_VISIBLE_DEVICES=0,1 \
  -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \
  -e VLLM_ROCM_USE_AITER_MHA=0 -e VLLM_ROCM_USE_AITER_MLA=0 -e VLLM_ROCM_USE_AITER_MOE=0 \
  -e VLLM_ROCM_USE_AITER_LINEAR=0 -e VLLM_ROCM_USE_AITER_FP8BMM=0 \
  -e VLLM_ROCM_USE_AITER_FP4BMM=0 -e VLLM_ROCM_USE_AITER_RMSNORM=0 \
  -e NCCL_PROTO=Simple \
  -e RADIANCE_PRESHUFFLE=1 \
  -e RADIANCE_USE_R4D_AR=1 \
  -e RADIANCE_USE_R4D_AR_QUANT=1 \
  -e RADIANCE_FUSE_RMS_QUANT=1 \
  -e RADIANCE_DYNAMIC_DRAFT=1 \
  -e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
  -e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter \
  -e TRITON_CACHE_AUTOTUNING=1 \
  stilldeadcode/vllm-radiance:0.9.3 \
    /models/YourOrg/Your-Model-FP8 \
    --served-model-name my-model \
    --quantization fp8 --kv-cache-dtype fp8 \
    --tensor-parallel-size 2 \
    --gpu-memory-utilization 0.92 \
    --attention-backend ROCM_AITER_UNIFIED_ATTN \
    --enable-prefix-caching --mamba-cache-mode align \
    --speculative-config '{"method":"mtp","num_speculative_tokens":8,"attention_backend":"ROCM_AITER_UNIFIED_ATTN","disable_padded_drafter_batch":true}' \
    --no-async-scheduling \
    --host 0.0.0.0 --port 8000

Test:

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"my-model","messages":[{"role":"user","content":"Hello!"}]}'
⁠compose

A ready-to-edit docker-compose.yml with the same flags, every device / group / volume already wired, and each knob commented lives in the source repo: https://codeberg.org/StillDeadcode/vllm-radiance⁠.

⁠First run is slower

With an empty cache the first start spends a few extra minutes compiling Triton / inductor kernels before the engine comes up; it looks idle but it is compiling. (Older builds spent 15 to 20 minutes here, dominated by the gated-delta-net fp32 autotune; that path is gone on this stack.) Mount a persistent cache so restarts stay fast:

  -v /path/to/vllm-cache:/cache \
  -e VLLM_CACHE_ROOT=/cache/vllm \
  -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
  -e TRITON_CACHE_DIR=/cache/triton \
  -e AITER_ROOT_DIR=/cache/aiter \
  -e TRITON_CACHE_AUTOTUNING=1

⁠Flags

FlagSuggestedNotes
--tensor-parallel-size2one rank per R9700
--quantizationfp8tuned for FP8 weights
--kv-cache-dtypefp8, bf16, or autofp8 = 1 byte/elem (most KV capacity); bf16 / auto keep full precision
--attention-backendROCM_AITER_UNIFIED_ATTNrequired for the tuned attention path
--max-model-lenmodel dependentcontext length per request
--max-num-seqsworkload dependentmax concurrent sequences
--gpu-memory-utilization0.90 to 0.97VRAM fraction for weights + KV
--enable-prefix-cachingon for shared prefixesenables automatic prefix caching; required: hybrid (GDN/mamba) models leave it off by default even though the engine default looks on
--mamba-cache-modealign (hybrid models)makes the linear-attention (GDN) layers prefix-cacheable; pair with --enable-prefix-caching on this hybrid. none disables mamba-layer caching; all is unsupported by this model
--numa-bindomit (off)multi-NUMA-node hosts only: pin the fleet to the GPU-local node(s). auto / <nodes> / interleave / preferred=<n> / none. Same as RADIANCE_NUMA_BIND; needs --cap-add SYS_NICE. See NUMA pinning above.

Speculative decoding (MTP). Two forms depending on where the MTP head lives:

# Qwen3.8-27B / Qwen3.6-27B / 35B: the MTP head is in the target checkpoint, so no separate drafter
--speculative-config '{"method":"mtp","num_speculative_tokens":8,"attention_backend":"ROCM_AITER_UNIFIED_ATTN","disable_padded_drafter_batch":true}'

# ...and on the 27B hybrids, give the drafter the same R4D backend if the target uses it
--attention-backend R4D --speculative-config '{"method":"mtp","num_speculative_tokens":8,"attention_backend":"R4D","disable_padded_drafter_batch":true}'

# Gemma-4-31B: the drafter is a separate model, so add "model" (and --trust-remote-code --no-async-scheduling)
--speculative-config '{"method":"mtp","model":"/models/google/gemma-4-31B-it-assistant","num_speculative_tokens":8,"attention_backend":"ROCM_AITER_UNIFIED_ATTN","disable_padded_drafter_batch":true}'

What num_speculative_tokens means here. Under mtp with RADIANCE_DYNAMIC_DRAFT=1 (baked on) it is a ceiling, not a fixed cost: the controller drafts up to that many, stops early on low-acceptance content, and may take a verbatim n-gram continuation. So 8 is a good default -- it gives the drafter room on code and JSON without adding fixed overhead on prose. With RADIANCE_DYNAMIC_DRAFT=0, or under dflash, it is a fixed width again and a smaller value is typical.

DFlash drafters. dflash is a different speculative shape from mtp: instead of running one draft forward per speculative position, it emits all of them in a single pass, so the draft is one CUDA-graph replay whose depth is fixed at capture. That has three consequences worth knowing.

--speculative-config '{"method":"dflash","model":"/models/<drafter>","num_speculative_tokens":7,"attention_backend":"TRITON_ATTN"}'
  • RADIANCE_DYNAMIC_DRAFT does nothing here (see the table above) -- num_speculative_tokens is a fixed width again, so it is worth tuning: on Qwen3.8-27B the per-position acceptance rate falls to ~0.10 by the seventh, and each extra position widens both the draft pass and the target's verify.
  • The drafter's attention backend must be one that supports full CUDA graphs, or vLLM logs that it is "running the draft eagerly" and you lose the graph. TRITON_ATTN does; check the startup log for Capturing dflash CUDA graphs (FULL).
  • Prefix caching with --mamba-cache-mode align has only been verified against mtp on hybrid models; it is a different speculative shape and has not been re-verified here.

disable_padded_drafter_batch:true is the key single-stream lever (~+50% on the 27B hybrids): it drops the drafter's batch padding, and the image bakes the vLLM unpad patch this relies on. Leave it on. Note it is incompatible with async scheduling: pass --no-async-scheduling to disable it explicitly (otherwise vLLM auto-enables async scheduling and then disables it with a runtime warning; --async-scheduling would hard-error).

Prefix caching (shared system prompts, RAG, agentic context):

--enable-prefix-caching --mamba-cache-mode align

Automatic prefix caching reuses a shared prompt prefix across requests so only the new suffix is prefilled -- a large TTFT drop when requests share a system prompt or document. On a GDN hybrid you must pass both flags: hybrid models default their prefix-caching support flag off, so vLLM silently disables it without --enable-prefix-caching, and align is what makes the linear-attention layers cacheable by snapshotting their conv + recurrent state at block boundaries. That restore is verified bit-identical to a full recompute, so outputs are unchanged and the win is purely latency (~3.6x faster TTFT on shared prefixes). Trade-offs: align raises the attention block size to 1664 tokens and adds one state block per linear-attention layer, so prefix hits land on 1664-token boundaries and max concurrency at full context drops slightly. Do not use --mamba-cache-mode all (unsupported, raises at startup) or set VLLM_SSM_CONV_STATE_LAYOUT=DS (asserts under MTP + align).

Tool-calling and reasoning:

--enable-auto-tool-choice --tool-call-parser <parser> --reasoning-parser <parser>

Pass a template with --chat-template file.jinja if the model needs one. The image ships the from_json filter those templates often rely on.

Tag summary

Content type

Image

Digest

sha256:456942091…

Size

3.7 GB

Last updated

about 1 month ago

docker pull stilldeadcode/vllm-radiance:0.9.3