Sign inSign up

tcclaviger/vllm

By tcclaviger

•Updated 1 day ago

Image
12

10K+

tcclaviger/vllm repository overview

⁠tcclaviger/vllm - ROCm/RDNA4 vLLM with native HIP kernels, baked-in quantizers + tuner

vLLM 0.30 tree on ROCm 10.0, torch 2.11, Python 3.14, built for gfx1201 (Radeon R9700). With no tool flag the image runs vllm serve "$@", so it drops into an existing vLLM deployment unchanged.

Docs, examples and serving recipes: https://blog.robai.net/vllmdocs/⁠ Contact: [email protected]⁠. Bug reports and feature requests welcome; handled as my schedule allows.

⁠Tags

  • latest - stable, use it for serving. dev - newer, less settled. exp - experimental, may change or break. Numbered tags (e.g. 29.04.1) pin a build.

All tags ship libr4d, the RDNA4 HIP kernel library: https://codeberg.org/StillDeadcode/libr4d⁠. Unlisted features may be present as drafts.

⁠Credits

  • StillDeadcode - libr4d
  • Davetha - int4 draft lm_head, LRU expert cache behind expert offload
  • Pat Carter - RDNA4 GDN kernels, rdna4-vllm⁠
  • Dylhun - KVA and bug hunts
  • Code 000 - refinements
  • ggz14 - kernels
  • Everyone on Launch80 who contributed ideas, suggestions and testing

⁠Change log

⁠:latest (29.07.1)
  • added simple kv offload to system ram or disk with eager or lazy store policy
    • validated with PLE in RAM, PLE LRU cache and PLE NVMe direct on TP2 and TP4
  • kv offload disk tier now raises on a disk io error and takes the engine down with the cause named instead of hanging
  • fixed kv offload startup assert on the qsa key ring group for flash next
  • fixed kv offload store completion assert on mamba align boundary blocks
  • offload scheduler log line now prints the tier size in tokens
⁠:dev (29.07.2)
  • initial adjustment to improve TP2 stability
⁠:exp (29.07.16)
  • expert offloading is completely reworked as a single copy plan that sizes itself from the model, users opt in with these two environment variables
    • CLAV_OFFLOAD_V2=1
    • R4D_LRU_UVA_TOKENS=32
  • DeepSeek V4 Flash experimental kernels updated for the speculator, the DSpark draft step now runs through a fused libr4d preamble
  • DeepSeek V4 Flash now boots under the reworked offload with its kv cache sized exactly from the model config
  • fixed the mxfp4 expert shape check that stopped DeepSeek V4 Flash loading under the reworked offload
  • draft model expert layers for DSpark and MTP stay resident on the gpu under expert offload

Tag summary

Content type

Image

Digest

sha256:a7b08ce32…

Size

3.9 GB

Last updated

3 days ago

docker pull tcclaviger/vllm