Sign inSign up

luojiecong/dgx-spark-vllm

By luojiecong

•Updated about 1 month ago

vLLM 0.28.1 for DGX Spark GB10: DeepSeek-V4-Flash vision + DSpark + PCP; NCCL 2.30.7; sm_121a

Image
0

154

luojiecong/dgx-spark-vllm repository overview

中文版

⁠DGX Spark (GB10) · DeepSeek-V4-Flash vLLM 运行时

面向 NVIDIA DGX Spark / ASUS Ascent GX10(GB10,SM_121a,ARM64) 预构建的 vLLM 运行时镜像,针对 DeepSeek-V4-Flash(含 Vision-Exp 视觉) 双机 TP=2 部署优化。

⁠组件版本

组件版本
vLLM0.28.1rc1.dev602+gd365b35c4(基于 upstream main)
CUDA13.0.3
PyTorch2.13.0+cu130
FlashInfer0.6.18
b12x1.3.0(MXFP4 MoE)
Triton3.7.1
CUTLASS DSL4.6.2
NCCL2.30.7(宿主机源码编译版,已打包进镜像 /opt/nccl/lib)
Python3.12.3
架构linux/arm64 (aarch64),sm_121a

⁠内置特性

  • DeepSeek-V4-Flash / Vision-Exp 原生支持(MLA + Lightning Indexer + C4A/C128A 压缩)
  • DSpark 投机解码(spec decode,num_speculative_tokens 可配)
  • PCP(Prefill Context Parallel),含 DSpark 组合支持(upstream PR #53427)
  • 多模态 × PCP(视觉模型 inputs_embeds 按 rank 切分)
  • FP8 DS-MLA KV cache(fp8_ds_mla,584 B/token)
  • MoE 走 B12X MXFP4×MXFP8 原生 tensor-core 路径
  • NCCL 2.30.7 内置(LD_LIBRARY_PATH=/opt/nccl/lib 优先于 site-packages)

⁠适用硬件

  • 2× NVIDIA DGX Spark (GB10) 或 ASUS Ascent GX10
  • ConnectX-7 RoCE 互联(100G+)

⁠说明

  • 镜像不含模型权重,需自行挂载
  • PCP 的设备数 = TP × PCP(例如 TP=2 + PCP=2 需要 4 台节点) English version

⁠DGX Spark (GB10) · DeepSeek-V4-Flash vLLM Runtime

Prebuilt vLLM runtime image for NVIDIA DGX Spark / ASUS Ascent GX10 (GB10, SM_121a, ARM64), optimized for 2-node TP=2 deployment of DeepSeek-V4-Flash (incl. Vision-Exp).

⁠Component versions

ComponentVersion
vLLM0.28.1rc1.dev602+gd365b35c4 (upstream main based)
CUDA13.0.3
PyTorch2.13.0+cu130
FlashInfer0.6.18
b12x1.3.0 (MXFP4 MoE)
Triton3.7.1
CUTLASS DSL4.6.2
NCCL2.30.7 (host source build, bundled at /opt/nccl/lib)
Python3.12.3
Architecturelinux/arm64 (aarch64), sm_121a

⁠Built-in features

  • Native DeepSeek-V4-Flash / Vision-Exp support (MLA + Lightning Indexer + C4A/C128A compression)
  • DSpark speculative decoding
  • PCP (Prefill Context Parallel), including DSpark combination (upstream PR #53427)
  • Multimodal × PCP (vision inputs_embeds sliced per rank)
  • FP8 DS-MLA KV cache (fp8_ds_mla, 584 B/token)
  • MoE via B12X MXFP4×MXFP8 native tensor-core path
  • NCCL 2.30.7 bundled (takes precedence via LD_LIBRARY_PATH=/opt/nccl/lib)

⁠Target hardware

  • 2× NVIDIA DGX Spark (GB10) or ASUS Ascent GX10
  • ConnectX-7 RoCE interconnect (100G+)

⁠Notes

  • Model weights are NOT included; mount your own
  • PCP requires world_size = TP × PCP devices (e.g. TP=2 + PCP=2 needs 4 nodes)

Tag summary

Content type

Image

Digest

sha256:dac36a09b…

Size

8.8 GB

Last updated

about 1 month ago

docker pull luojiecong/dgx-spark-vllm:v0.28.1