vLLM inference stack for RTX PRO 6000 Blackwell (SM120a) with CUDA 13.2.
nvidia-cutlass-dsl[cu13] ships two sub-packages that both write _cutlass_ir.so:
libs-base (mandatory) — NVVM/ptxas 12.9 (can't compile _mma.block_scale for SM120)libs-cu13 (optional extra) — NVVM/ptxas 13.1 (works)pip installs libs-base after libs-cu13, overwriting the CUDA 13 binary.
Fix: pip install --force-reinstall --no-deps nvidia-cutlass-dsl-libs-cu13 as last step.
See full Dockerfile in the source repository.
Content type
Image
Digest
sha256:679526c83…
Size
11.9 GB
Last updated
11 days ago
docker pull voipmonitor/vllm:kimi-k3-kk-cu134-tp9-dspark-k5-20260922-r2