Sign inSign up

lovedheart/sglang

By lovedheart

•Updated 22 days ago

Qwen3.8-Flash-Next on sglang (#37500 merged, FP4/FP8 QSA indexer, SM120 b12x)

Image
0

311

lovedheart/sglang repository overview

⁠Qwen3.8-Flash-Next (sglang)

Serving image for Qwen3.8-Flash-Next (qwen4_exp / QSA sparse attention / PLE n-gram), built from branch qwen-flash-next-dev-pre-rollback⁠ (commit 40d55d7da, upstream main 52fecfdf0 merged in - includes the official PR #37500 "support qwen 3.8 flash next") on top of lmsysorg/sglang:dev-cu13 (CUDA 13.0.3, torch 2.13+cu130). Source tree is overlaid; all compiled deps inherited.

flashinfer 0.6.18 (python + jit-cache cu130 + cubin), aligned with the branch pin.

⁠Launch command (SM120, NVFP4 checkpoint, single GPU)

Standard NVIDIA entrypoint - flags go straight after the image name:

docker run --gpus all --ipc=host --network host \
  -v /path/to/models:/models:ro \
  -v /mnt/hicache:/mnt/hicache \
  -e SGLANG_QSA_USE_FP8_INDEXER=1 \
  lovedheart/qwen38-flash-next:latest \
  python3 -m sglang.launch_server \
    --model-path /models/Qwen3.8-Flash-Next-NVFP4-w4a16-4o6-attnFP8b128-full/ \
    --served-model-name Qwen3.8-Flash-Next \
    --host 0.0.0.0 --port 8070 --trust-remote-code \
    --tensor-parallel-size 1 --max-running-requests 3 \
    --chunked-prefill-size 4096 --max-prefill-tokens 12288 \
    --mem-fraction-static 0.96 \
    --kv-cache-dtype nvfp4 \
    --mamba-ssm-dtype bfloat16 --mamba-radix-cache-strategy extra_buffer_lazy \
    --max-mamba-cache-size 16 \
    --schedule-policy lpm --allow-auto-truncate --sleep-on-idle \
    --enable-cache-report --enable-metrics --enable-session-radix-cache \
    --attention-backend triton --decode-attention-backend trtllm_mha \
    --linear-attn-decode-backend flashinfer --linear-attn-prefill-backend flashinfer \
    --sampling-backend flashinfer \
    --moe-runner-backend flashinfer_cutlass \
    --fp4-gemm-backend flashinfer_b12x \
    --ple-offload-embedding \
    --speculative-algo NEXTN --speculative-num-steps 3 \
    --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
    --speculative-draft-kv-cache-dtype bfloat16 \
    --enable-linear-replayssm-spec \
    --enable-hierarchical-cache --hicache-ratio 0.5 \
    --hicache-write-policy write_through --hicache-io-backend direct \
    --hicache-mem-layout page_first_direct \
    --mm-feature-transport cpu --mm-attention-backend flashinfer_cudnn \
    --reasoning-parser auto --tool-call-parser auto

Notes:

  • Model is not baked in - mount it read-only at /models.
  • Mount the hicache storage dir read-write if hierarchical cache is enabled.
  • JIT caches live under /root/.cache (flashinfer JIT: /root/.cache/flashinfer) - mount a host dir to persist compiled kernels across containers.

⁠QSA indexer backends (SM120/SM121)

-e SGLANG_QSA_USE_FP8_INDEXER=1   # FP8 prefill scoring via DeepGEMM fp8_mqa_logits
-e SGLANG_QSA_USE_FP4_INDEXER=1   # FP4 packed prefill scoring + FP4 index-K pages (takes precedence)

FP8/FP4 scoring targets long-context prefill (measured at 512K); decode stays on the BF16 paged path. Use --max-running-requests 1 at 512K context. Legacy env SGLANG_QWEN_DSA_USE_FP8_INDEXER still works.

⁠What's new vs the previous latest (upstream merge @ 40d55d7da)

  • Upstream main merged (incl. official #37500 Qwen 3.8 Flash Next support); flashinfer 0.6.17 -> 0.6.18.
  • FP4 QSA indexer: packed prefill scoring + FP4 index-K pages (tokenwise & compressed).
  • FP4 KV pools (--kv-cache-dtype nvfp4) in QSA backends with dequant-gather reads; speculative draft KV pinned to BF16 under an FP4 target cache.
  • SM120/SM121 FP4 GEMM: flashinfer_b12x backend, auto-selected on SM12x.
  • Deferred clearing of recycled mamba ping-pong track slots (fixes ABAB parity corruption across cold runs).
  • PLE side-state scatter fix on the ReplaySSM fold/circular speculative commit paths.
  • Spec-verify pending-ring stride widened to 2x compress ratio (stops draft clobber).
  • Deterministic radix fast_topk tie selection; optional topk sort (SGLANG_QSA_SORT_TOPK=1).
  • PPTRACE forensics hook (SGLANG_PPTRACE=1, diagnostics only).

⁠Tags

tagcontents
latestthis branch merged with upstream main (= sglang:qwen-flash-next-latest)
dev / dev-prev / modelshistorical builds (pre-merge; flashinfer 0.6.17 era)

Mirror repo with the same image: lovedheart/sglang:qwen-flash-next-latest.

Tag summary

Content type

Image

Digest

sha256:e58bf01ad…

Size

21.3 GB

Last updated

22 days ago

docker pull lovedheart/sglang:qwen-flash-next-latest