Sign inSign up

mitakad/vllm

By mitakad

•Updated 7 days ago

vLLM Distributed Inference on Jetson Orin AGX-built by https://github.com/dusty-nv/jetson-containers

Image
Machine learning & AI
3

4.1K

mitakad/vllm repository overview

!!! Note for mitakad/vllm:0.27.0.dev0-r39.2.tegra-aarch64-cp312-cu132-24.04-commit.96add73-pruned

vLLM removed the automatic disable_custom_all_reduce = True for
multi-node setup. Add --disable-custom-all-reduce to your serve command like:

  • On Head Node (in case of 2 nodes cluster with Jetson Orin AGX 64GB RAM, set $HEAD_IP):
vllm serve Kwaipilot/KAT-Coder-V2.5-Dev --port 8000 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --mm-encoder-tp-mode data \
  --max-model-len auto \
  --gpu-memory-utilization 0.85 \
  --max-num-seqs 8 \
  --kv-cache-memory-bytes 20g \
  --max-num-batched-tokens 8192 \
  --language-model-only \
  --disable-custom-all-reduce \
  --enable-prefix-caching \
  --tensor-parallel-size 2 \
  --nnodes 2 \
  --master-addr $HEAD_IP 
  • On Worker Node:
vllm serve Kwaipilot/KAT-Coder-V2.5-Dev --port 8000 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --mm-encoder-tp-mode data \
  --max-model-len auto \
  --gpu-memory-utilization 0.85 \
  --max-num-seqs 8 \
  --kv-cache-memory-bytes 20g \
  --max-num-batched-tokens 8192 \
  --language-model-only \
  --disable-custom-all-reduce \
  --enable-prefix-caching \
  --tensor-parallel-size 2 \
  --nnodes 2 \
  --master-addr $HEAD_IP \
  --headless \
  --node-rank 1

!!! Note for 0.24.0.pr.45544-r39.2.tegra-aarch64-cp312-cu132-24.04

Includes changes from https://github.com/vllm-project/vllm/pull/45544⁠ rebased on v0.24.0 release.

This way you can run modelopt quantized LLM's (from this PR: https://github.com/vllm-project/vllm/pull/45306⁠):

vllm serve nvidia/Gemma-4-26B-A4B-NVFP4 --port 8000 \
  --quantization modelopt \
  --kv-cache-dtype bfloat16 \
  --tensor-parallel-size 1 \
  --trust-remote-code \
  --tool-call-parser gemma4 \
  --reasoning-parser gemma4 \
  --enable-auto-tool-choice \
  --max-model-len 4096 \
  --max-num-seqs 1 \
  --gpu-memory-utilization 0.85

!!! Note for vllm images tagged with TurboQuant:

Example running on Jetson Orin AGX with 64GB of RAM:

VLLM_MARLIN_USE_ATOMIC_ADD=1 \
vllm serve Qwen/Qwen3.6-35B-A3B-FP8 --port 8000 \
  --enable-auto-tool-choice \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --max-model-len auto \
  --gpu-memory-utilization 0.85 \
  --max-num-seqs 10 \
  --kv-cache-memory-bytes 10g \
  --language-model-only \
  --kv-cache-dtype turboquant_4bit_nc \
  --compilation-config '{"cudagraph_mode": "PIECEWISE"}' \
  --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'

Shows :

  • RAM at 51.5 GB
  • Add 3 padding layers, may waste at most 10.00% KV cache memory
  • Auto-fit max_model_len: full model context length 262144 fits in available GPU memory
  • GPU KV cache size: 432,640 tokens
  • Maximum concurrency for 262,144 tokens per request: 5.71x

More info about TurboQuant in PR: [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity #38479⁠


!!! Note: for vLLM:v0.15.0, if you experience degraded performance, check the vllm serve logs if using ...TRITON_MLA attention backend..., this could degrade performance with larger inputs. To disble MLA, set VLLM_MLA_DISABLE=1, example: VLLM_MLA_DISABLE=1 vllm serve ...

Built on Jetson Orin AGX (SM 87)

vLLM docs: https://docs.vllm.ai/en/stable/⁠


Distributed LLM inferencing with vLLM, as described in the Multi-node case here https://docs.vllm.ai/en/latest/serving/parallelism_scaling.html#multi-node-⁠

How to use: https://github.com/dusty-nv/jetson-containers/blob/213ecc5b8db55ad10b875f5afb61569470f3541d/docs/distributed-inference.md⁠

Tested with 2x Jetson AGX Orin dev kits.

Also try EAGLE-3 speculative decoding for faster inference (as explained here⁠) :

vllm serve RedHatAI/Qwen3-8B-quantized.w4a16 \
  --port 8000 \
  --gpu-memory-utilization 0.6 \
  --kv-cache-memory-bytes 5G \
  --max-model-len 2048 \
  --speculative-config '{
    "model": "RedHatAI/Qwen3-8B-speculator.eagle3",
    "num_speculative_tokens": 3,
    "method": "eagle3"
  }'

Tag summary

Content type

Image

Digest

sha256:b95e0b9c1…

Size

18.5 GB

Last updated

7 days ago

docker pull mitakad/vllm:0.30.0-r39.2.tegra-aarch64-cp312-cu132-24.04