tcclaviger/vllm - ROCm/RDNA4 vLLM with native HIP kernels, baked-in quantizers + tuner
vLLM 0.30 tree on ROCm 10.0, torch 2.11, Python 3.14, built for gfx1201 (Radeon R9700). With no tool flag the image runs vllm serve "$@", so it drops into an existing vLLM deployment unchanged.
Docs, examples and serving recipes: https://blog.robai.net/vllmdocs/
Contact: [email protected] . Bug reports and feature requests welcome; handled as my schedule allows.
Tags
latest - stable, use it for serving. dev - newer, less settled. exp - experimental, may change or break. Numbered tags (e.g. 29.04.1) pin a build.
All tags ship libr4d , the RDNA4 HIP kernel library: https://codeberg.org/StillDeadcode/libr4d . Unlisted features may be present as drafts.
Credits
StillDeadcode - libr4d
Davetha - int4 draft lm_head, LRU expert cache behind expert offload
Pat Carter - RDNA4 GDN kernels, rdna4-vllm
Dylhun - KVA and bug hunts
Code 000 - refinements
ggz14 - kernels
Everyone on Launch80 who contributed ideas, suggestions and testing
Change log
:latest (29.07.1)
added simple kv offload to system ram or disk with eager or lazy store policy
validated with PLE in RAM, PLE LRU cache and PLE NVMe direct on TP2 and TP4
kv offload disk tier now raises on a disk io error and takes the engine down with the cause named instead of hanging
fixed kv offload startup assert on the qsa key ring group for flash next
fixed kv offload store completion assert on mamba align boundary blocks
offload scheduler log line now prints the tier size in tokens
:dev (29.07.2)
initial adjustment to improve TP2 stability
:exp (29.07.16)
expert offloading is completely reworked as a single copy plan that sizes itself from the model, users opt in with these two environment variables
CLAV_OFFLOAD_V2=1
R4D_LRU_UVA_TOKENS=32
DeepSeek V4 Flash experimental kernels updated for the speculator, the DSpark draft step now runs through a fused libr4d preamble
DeepSeek V4 Flash now boots under the reworked offload with its kv cache sized exactly from the model config
fixed the mxfp4 expert shape check that stopped DeepSeek V4 Flash loading under the reworked offload
draft model expert layers for DSpark and MTP stay resident on the gpu under expert offload