Sign inSign up

exaroot/gfx1030-vllm-rdna2-extras

By exaroot

•Updated about 1 month ago

vLLM for AMD gfx1030 (RDNA2) — rdna2_extras @47d92b6cf, ROCm 7.14

Image
0

158

exaroot/gfx1030-vllm-rdna2-extras repository overview

⁠gfx1030-vllm-rdna2-extras

vLLM for AMD gfx1030 (RDNA2 — Radeon PRO V620 / RX 6800-class), ROCm 7.14.

Built from BlivionIaG/vllm⁠ branch rdna2_extras @ 47d92b6cf, including:

  • FA-RDNA2 attention for head_size=256 on gfx10x/gfx1x (prefill + decode)
  • W4A16 AWQ kernel (q_gemm_rdna2) with v_dot2_f32_f16 packed math
  • GDN packed decode hand-ported to HIP RDNA2 (register-resident fp32 state)
  • Decode-kernel occupancy fixes (amdgpu_waves_per_eu, unpinned wvSplitK)

Tags: 47d92b6cf (branch HEAD), latest.

⁠Measured performance (2× Radeon PRO V620, our AWQ-INT4 checkpoints)

C0 checkpoint matrix (medians of 3, fp8 KV cache, fresh container per leg). TP=2 numbers are the best configurations ever measured on this hardware:

ModelConfigSingle-stream TPSTTFT @16kKV pool
Qwen3.5-4B-AWQ-vdTP=193.116.4 s933k tok
Qwen3.5-4B-AWQ-vdTP=2113.98.8 s2.18M tok
Qwen3.8-27B-AWQ-vdTP=117.2109.6 s96k tok
Qwen3.8-27B-AWQ-vdTP=228.548.6 s750k tok (44×16k)

Kernel-level (gfx1030 microbench, medians):

  • AWQ GEMM at M=128: dequant + rocBLAS-fp16 = 16.75 TF vs fused Triton AWQ 7.64 TF (2.2×) — use VLLM_AWQ_FP16_MATMUL_MIN_M=128 (see our PR to the fork)
  • v_dot2_f32_f16 issue rate = 0.96× scalar (viable); v_pk_fma_f16 is half-rate (avoid); v_pk_mul_f16 full-rate. First measured RDNA2 packed-ISA table.

Notes: decode throughput scales ~1/context (RDNA2 paged-attention decode path); VLLM_USE_BREAKABLE_CUDAGRAPH=1 is required (torch.compile freezes gfx1030); ROCM_ATTN backend (TRITON_ATTN exceeds the 64KB LDS/WG limit); KV fp8 costs ~10% TPS for 2× capacity.

⁠Best docker-compose.yml (27B, TP=2 — the 28.5 TPS config)

services:
  qwen27b:
    image: exaroot/gfx1030-vllm-rdna2-extras:latest
    restart: unless-stopped
    ports:
      - "8000:8000"
    volumes:
      - /path/to/Qwen3.8-27B-AWQ-vd-lmhead-int4:/model:ro
    shm_size: "8g"
    environment:
      - HSA_OVERRIDE_GFX_VERSION=10.3.0
      - HIP_FORCE_DEV_KERNARG=1
      - OMP_NUM_THREADS=8
      - TOKENIZERS_PARALLELISM=false
      - VLLM_USE_TRITON_AWQ=1
      - VLLM_USE_DEEP_GEMM=0
      - VLLM_USE_FLASHINFER_SAMPLER=0
      - PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
      - VLLM_USE_BREAKABLE_CUDAGRAPH=1
      - VLLM_AWQ_FP16_MATMUL_MIN_M=128
    devices:
      - /dev/dri
      - /dev/kfd
    group_add: ["render", "video"]
    cap_add: [SYS_PTRACE]
    security_opt: [seccomp=unconfined]
    command: >
      serve /model --quantization awq --tensor-parallel-size 2
      --kv-cache-dtype fp8 --max-model-len 17000 --max-num-seqs 8
      --max-num-batched-tokens 8192 --trust-remote-code --dtype float16
      --attention-backend ROCM_ATTN --gpu-memory-utilization 0.90
      --generation-config vllm --enable-chunked-prefill
      --limit-mm-per-prompt '{"image": 1}'
      --mm-processor-kwargs '{"max_pixels": 1003520}'
      --host 0.0.0.0 --port 8000

For the 4B: change the model mount and keep everything else (113.9 TPS, 2.18M pool). Raise --max-model-len for longer context (32k works; budget KV from the pool table).

⁠Credits

  • Base kernels and fork: BlivionIaG⁠ rdna2_extras
  • AWQ-INT4 checkpoints (full-INT4 incl. attention/GDN/lm_head): ikantkode⁠ — quant recipe at awq-quant-recipe⁠
  • Tuning, benchmark matrix, and dispatch work: the gfx1030-vllm campaign (per-shape AWQ tables, LLMM1 arch-gate unlock, fused RMSNorm, G0 ISA-rate table, C0 matrix)

Tag summary

Content type

Image

Digest

sha256:334cca941…

Size

12.7 GB

Last updated

about 1 month ago

docker pull exaroot/gfx1030-vllm-rdna2-extras