Sign inSign up

christopherlin/vllm-gfx906

By christopherlin

•Updated 11 months ago

The docker image of https://github.com/nlzy/vllm-gfx906 with PR#46

Image
0

2.6K

christopherlin/vllm-gfx906 repository overview

This image packages a gfx906-optimized build of vLLM for AMD Instinct MI50/MI60 and Radeon VII class GPUs, running on ROCm PyTorch. This PR contributes a reproducible ROCm 6.3 vLLM Docker image for gfx906 (MI50-class) and wires up model routes so that Qwen3-Next, Qwen3-VL and Qwen3-MoE run out-of-the-box via vLLM’s AMD backend. It aligns with upstream vLLM ROCm guidance and leverages vLLM’s PagedAttention/continuous batching for high-throughput serving on ROCm.

Dependency stack

  • Ubuntu 24.04 base
  • ROCm 6.3.3
  • PyTorch 2.8.0
  • transformers 4.57.1
  • triton 3.3.0 (gfx906)

What this PR does

  • Adds a Docker build targeting Ubuntu 24.04 + ROCm 6.3 that preinstalls the AMD ROCm runtime and PyTorch ROCm wheels recommended by upstream, ensuring compatibility with vLLM on MI50/gfx906.
  • Bundles vLLM + Triton (ROCm) wheels, avoiding CUDA-only artifacts and keeping the image coherent for AMD deployments. (General vLLM AMD install flow referenced.)
  • Unblocks Qwen3-Next and Qwen3-VL routes: tested that the image can serve current Qwen3 variants via vLLM (text + vision). References for model families included below.

Performance & serving notes

  • The image is tuned for vLLM’s PagedAttention + continuous batching, which are the core enablers of high-throughput, long-context serving; these apply on ROCm as well as CUDA.
  • On multi-GPU ROCm, PyTorch uses AMD collectives (RCCL) and math libraries (e.g., hipBLASLt/Tensile) that the ROCm toolchain exposes; the Docker baseline here is compatible with those components.

Confirmed model pathways unlocked

  • Qwen3-Next (text) — modern Qwen3 family; available from major providers and used widely for LLM inference.
  • Qwen3-VL (vision-language) — latest multimodal Qwen3-VL series; supported in cloud/runtime ecosystems and compatible with vLLM’s Hugging Face pathways.
  • Remain the performance of Qwen3-MoE in pervious vllm+gfx906 version.

p.s. When running the model in the Docker image, you will see the current version as 0.9.3, which is an uncorrected version number. It is actually running version 0.11.0 of vllm, which will be fixed in a future code update (this is not a big problem).

other performance details in https://github.com/nlzy/vllm-gfx906/pull/46⁠

Tag summary

Content type

Image

Digest

sha256:4192bd9b2…

Size

6.9 GB

Last updated

11 months ago

docker pull christopherlin/vllm-gfx906