vLLM 0.28.1 for DGX Spark GB10: DeepSeek-V4-Flash vision + DSpark + PCP; NCCL 2.30.7; sm_121a
154
中文版
面向 NVIDIA DGX Spark / ASUS Ascent GX10(GB10,SM_121a,ARM64) 预构建的 vLLM 运行时镜像,针对 DeepSeek-V4-Flash(含 Vision-Exp 视觉) 双机 TP=2 部署优化。
| 组件 | 版本 |
|---|---|
| vLLM | 0.28.1rc1.dev602+gd365b35c4(基于 upstream main) |
| CUDA | 13.0.3 |
| PyTorch | 2.13.0+cu130 |
| FlashInfer | 0.6.18 |
| b12x | 1.3.0(MXFP4 MoE) |
| Triton | 3.7.1 |
| CUTLASS DSL | 4.6.2 |
| NCCL | 2.30.7(宿主机源码编译版,已打包进镜像 /opt/nccl/lib) |
| Python | 3.12.3 |
| 架构 | linux/arm64 (aarch64),sm_121a |
Prebuilt vLLM runtime image for NVIDIA DGX Spark / ASUS Ascent GX10 (GB10, SM_121a, ARM64), optimized for 2-node TP=2 deployment of DeepSeek-V4-Flash (incl. Vision-Exp).
| Component | Version |
|---|---|
| vLLM | 0.28.1rc1.dev602+gd365b35c4 (upstream main based) |
| CUDA | 13.0.3 |
| PyTorch | 2.13.0+cu130 |
| FlashInfer | 0.6.18 |
| b12x | 1.3.0 (MXFP4 MoE) |
| Triton | 3.7.1 |
| CUTLASS DSL | 4.6.2 |
| NCCL | 2.30.7 (host source build, bundled at /opt/nccl/lib) |
| Python | 3.12.3 |
| Architecture | linux/arm64 (aarch64), sm_121a |
Content type
Image
Digest
sha256:dac36a09b…
Size
8.8 GB
Last updated
about 1 month ago
docker pull luojiecong/dgx-spark-vllm:v0.28.1