vLLM for AMD gfx1030 (RDNA2) — rdna2_extras @47d92b6cf, ROCm 7.14
158
vLLM for AMD gfx1030 (RDNA2 — Radeon PRO V620 / RX 6800-class), ROCm 7.14.
Built from BlivionIaG/vllm branch rdna2_extras
@ 47d92b6cf, including:
head_size=256 on gfx10x/gfx1x (prefill + decode)q_gemm_rdna2) with v_dot2_f32_f16 packed mathamdgpu_waves_per_eu, unpinned wvSplitK)Tags: 47d92b6cf (branch HEAD), latest.
C0 checkpoint matrix (medians of 3, fp8 KV cache, fresh container per leg). TP=2 numbers are the best configurations ever measured on this hardware:
| Model | Config | Single-stream TPS | TTFT @16k | KV pool |
|---|---|---|---|---|
| Qwen3.5-4B-AWQ-vd | TP=1 | 93.1 | 16.4 s | 933k tok |
| Qwen3.5-4B-AWQ-vd | TP=2 | 113.9 | 8.8 s | 2.18M tok |
| Qwen3.8-27B-AWQ-vd | TP=1 | 17.2 | 109.6 s | 96k tok |
| Qwen3.8-27B-AWQ-vd | TP=2 | 28.5 | 48.6 s | 750k tok (44×16k) |
Kernel-level (gfx1030 microbench, medians):
VLLM_AWQ_FP16_MATMUL_MIN_M=128 (see our PR to the fork)v_dot2_f32_f16 issue rate = 0.96× scalar (viable); v_pk_fma_f16 is
half-rate (avoid); v_pk_mul_f16 full-rate. First measured RDNA2 packed-ISA table.Notes: decode throughput scales ~1/context (RDNA2 paged-attention decode path);
VLLM_USE_BREAKABLE_CUDAGRAPH=1 is required (torch.compile freezes gfx1030);
ROCM_ATTN backend (TRITON_ATTN exceeds the 64KB LDS/WG limit); KV fp8 costs
~10% TPS for 2× capacity.
services:
qwen27b:
image: exaroot/gfx1030-vllm-rdna2-extras:latest
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- /path/to/Qwen3.8-27B-AWQ-vd-lmhead-int4:/model:ro
shm_size: "8g"
environment:
- HSA_OVERRIDE_GFX_VERSION=10.3.0
- HIP_FORCE_DEV_KERNARG=1
- OMP_NUM_THREADS=8
- TOKENIZERS_PARALLELISM=false
- VLLM_USE_TRITON_AWQ=1
- VLLM_USE_DEEP_GEMM=0
- VLLM_USE_FLASHINFER_SAMPLER=0
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
- VLLM_USE_BREAKABLE_CUDAGRAPH=1
- VLLM_AWQ_FP16_MATMUL_MIN_M=128
devices:
- /dev/dri
- /dev/kfd
group_add: ["render", "video"]
cap_add: [SYS_PTRACE]
security_opt: [seccomp=unconfined]
command: >
serve /model --quantization awq --tensor-parallel-size 2
--kv-cache-dtype fp8 --max-model-len 17000 --max-num-seqs 8
--max-num-batched-tokens 8192 --trust-remote-code --dtype float16
--attention-backend ROCM_ATTN --gpu-memory-utilization 0.90
--generation-config vllm --enable-chunked-prefill
--limit-mm-per-prompt '{"image": 1}'
--mm-processor-kwargs '{"max_pixels": 1003520}'
--host 0.0.0.0 --port 8000
For the 4B: change the model mount and keep everything else (113.9 TPS, 2.18M pool).
Raise --max-model-len for longer context (32k works; budget KV from the pool table).
rdna2_extrasContent type
Image
Digest
sha256:334cca941…
Size
12.7 GB
Last updated
about 1 month ago
docker pull exaroot/gfx1030-vllm-rdna2-extras