This image packages a gfx906-optimized build of vLLM for AMD Instinct MI50/MI60 and Radeon VII class GPUs, running on ROCm PyTorch. This PR contributes a reproducible ROCm 6.3 vLLM Docker image for gfx906 (MI50-class) and wires up model routes so that Qwen3-Next, Qwen3-VL and Qwen3-MoE run out-of-the-box via vLLM’s AMD backend. It aligns with upstream vLLM ROCm guidance and leverages vLLM’s PagedAttention/continuous batching for high-throughput serving on ROCm.
Dependency stack
Ubuntu 24.04 base
ROCm 6.3.3
PyTorch 2.8.0
transformers 4.57.1
triton 3.3.0 (gfx906)
What this PR does
Adds a Docker build targeting Ubuntu 24.04 + ROCm 6.3 that preinstalls the AMD ROCm runtime and PyTorch ROCm wheels recommended by upstream, ensuring compatibility with vLLM on MI50/gfx906.
Bundles vLLM + Triton (ROCm) wheels, avoiding CUDA-only artifacts and keeping the image coherent for AMD deployments. (General vLLM AMD install flow referenced.)
Unblocks Qwen3-Next and Qwen3-VL routes: tested that the image can serve current Qwen3 variants via vLLM (text + vision). References for model families included below.
Performance & serving notes
The image is tuned for vLLM’s PagedAttention + continuous batching, which are the core enablers of high-throughput, long-context serving; these apply on ROCm as well as CUDA.
On multi-GPU ROCm, PyTorch uses AMD collectives (RCCL) and math libraries (e.g., hipBLASLt/Tensile) that the ROCm toolchain exposes; the Docker baseline here is compatible with those components.
Confirmed model pathways unlocked
Qwen3-Next (text) — modern Qwen3 family; available from major providers and used widely for LLM inference.
Qwen3-VL (vision-language) — latest multimodal Qwen3-VL series; supported in cloud/runtime ecosystems and compatible with vLLM’s Hugging Face pathways.
Remain the performance of Qwen3-MoE in pervious vllm+gfx906 version.
p.s. When running the model in the Docker image, you will see the current version as 0.9.3, which is an uncorrected version number. It is actually running version 0.11.0 of vllm, which will be fixed in a future code update (this is not a big problem).