Sign inSign up

malavp/vllm-mimo26flash

By malavp

•Updated 7 days ago

vLLM image for XiaomiMiMo/MiMo-V2.6-Flash-RL on SM80. Single request decode 100 tok/s on 8xA100-40GB

Image
0

436

malavp/vllm-mimo26flash repository overview

An image for serving XiaomiMiMo/MiMo-V2.6-Flash-RL on NVIDIA SM80 architecture. Tested on 8xA100-40GB with single stream decode speed of ~100 tok/s.

⁠Image Details

Built for linux/amd64. This image was built from a fork of vllm available at https://github.com/Malav-P/vllm/tree/mimo26flash-a100⁠

⁠Serving Recipe

Served with

  --tensor-parallel-size=8
  --trust-remote-code
  --load-format=instanttensor
  --gpu-memory-utilization=0.90
  --max-model-len=131072  
  --kv-cache-dtype=auto
  --reasoning-parser=mimo
  --tool-call-parser=mimo
  --enable-auto-tool-choice
  --generation-config=vllm

Tag summary

Content type

Image

Digest

sha256:d0e321ff8…

Size

8.2 GB

Last updated

7 days ago

docker pull malavp/vllm-mimo26flash:sm80