vLLM image for XiaomiMiMo/MiMo-V2.6-Flash-RL on SM80. Single request decode 100 tok/s on 8xA100-40GB
436
An image for serving XiaomiMiMo/MiMo-V2.6-Flash-RL on NVIDIA SM80 architecture. Tested on 8xA100-40GB with single stream decode speed of ~100 tok/s.
Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/mimo26flash-a100
Served with
--tensor-parallel-size=8
--trust-remote-code
--load-format=instanttensor
--gpu-memory-utilization=0.90
--max-model-len=131072
--kv-cache-dtype=auto
--reasoning-parser=mimo
--tool-call-parser=mimo
--enable-auto-tool-choice
--generation-config=vllm
Content type
Image
Digest
sha256:d0e321ff8…
Size
8.2 GB
Last updated
7 days ago
docker pull malavp/vllm-mimo26flash:sm80