Qwen3.8-Flash-Next on sglang (#37500 merged, FP4/FP8 QSA indexer, SM120 b12x)
311
Serving image for Qwen3.8-Flash-Next (qwen4_exp / QSA sparse attention / PLE n-gram),
built from branch qwen-flash-next-dev-pre-rollback
(commit 40d55d7da, upstream main 52fecfdf0 merged in - includes the official PR #37500
"support qwen 3.8 flash next") on top of lmsysorg/sglang:dev-cu13
(CUDA 13.0.3, torch 2.13+cu130). Source tree is overlaid; all compiled deps inherited.
flashinfer 0.6.18 (python + jit-cache cu130 + cubin), aligned with the branch pin.
Standard NVIDIA entrypoint - flags go straight after the image name:
docker run --gpus all --ipc=host --network host \
-v /path/to/models:/models:ro \
-v /mnt/hicache:/mnt/hicache \
-e SGLANG_QSA_USE_FP8_INDEXER=1 \
lovedheart/qwen38-flash-next:latest \
python3 -m sglang.launch_server \
--model-path /models/Qwen3.8-Flash-Next-NVFP4-w4a16-4o6-attnFP8b128-full/ \
--served-model-name Qwen3.8-Flash-Next \
--host 0.0.0.0 --port 8070 --trust-remote-code \
--tensor-parallel-size 1 --max-running-requests 3 \
--chunked-prefill-size 4096 --max-prefill-tokens 12288 \
--mem-fraction-static 0.96 \
--kv-cache-dtype nvfp4 \
--mamba-ssm-dtype bfloat16 --mamba-radix-cache-strategy extra_buffer_lazy \
--max-mamba-cache-size 16 \
--schedule-policy lpm --allow-auto-truncate --sleep-on-idle \
--enable-cache-report --enable-metrics --enable-session-radix-cache \
--attention-backend triton --decode-attention-backend trtllm_mha \
--linear-attn-decode-backend flashinfer --linear-attn-prefill-backend flashinfer \
--sampling-backend flashinfer \
--moe-runner-backend flashinfer_cutlass \
--fp4-gemm-backend flashinfer_b12x \
--ple-offload-embedding \
--speculative-algo NEXTN --speculative-num-steps 3 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
--speculative-draft-kv-cache-dtype bfloat16 \
--enable-linear-replayssm-spec \
--enable-hierarchical-cache --hicache-ratio 0.5 \
--hicache-write-policy write_through --hicache-io-backend direct \
--hicache-mem-layout page_first_direct \
--mm-feature-transport cpu --mm-attention-backend flashinfer_cudnn \
--reasoning-parser auto --tool-call-parser auto
Notes:
/models./root/.cache (flashinfer JIT: /root/.cache/flashinfer) - mount a host dir to persist compiled kernels across containers.-e SGLANG_QSA_USE_FP8_INDEXER=1 # FP8 prefill scoring via DeepGEMM fp8_mqa_logits
-e SGLANG_QSA_USE_FP4_INDEXER=1 # FP4 packed prefill scoring + FP4 index-K pages (takes precedence)
FP8/FP4 scoring targets long-context prefill (measured at 512K); decode stays on the BF16 paged path.
Use --max-running-requests 1 at 512K context. Legacy env SGLANG_QWEN_DSA_USE_FP8_INDEXER still works.
latest (upstream merge @ 40d55d7da)--kv-cache-dtype nvfp4) in QSA backends with dequant-gather reads; speculative draft KV pinned to BF16 under an FP4 target cache.flashinfer_b12x backend, auto-selected on SM12x.SGLANG_QSA_SORT_TOPK=1).SGLANG_PPTRACE=1, diagnostics only).| tag | contents |
|---|---|
latest | this branch merged with upstream main (= sglang:qwen-flash-next-latest) |
dev / dev-prev / models | historical builds (pre-merge; flashinfer 0.6.17 era) |
Mirror repo with the same image: lovedheart/sglang:qwen-flash-next-latest.
Content type
Image
Digest
sha256:e58bf01ad…
Size
21.3 GB
Last updated
22 days ago
docker pull lovedheart/sglang:qwen-flash-next-latest