vLLM Distributed Inference on Jetson Orin AGX-built by https://github.com/dusty-nv/jetson-containers
4.1K
!!! Note for mitakad/vllm:0.27.0.dev0-r39.2.tegra-aarch64-cp312-cu132-24.04-commit.96add73-pruned
vLLM removed the automatic disable_custom_all_reduce = True for
multi-node setup.
Add --disable-custom-all-reduce to your serve command like:
Head Node (in case of 2 nodes cluster with Jetson Orin AGX 64GB RAM, set $HEAD_IP):vllm serve Kwaipilot/KAT-Coder-V2.5-Dev --port 8000 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--mm-encoder-tp-mode data \
--max-model-len auto \
--gpu-memory-utilization 0.85 \
--max-num-seqs 8 \
--kv-cache-memory-bytes 20g \
--max-num-batched-tokens 8192 \
--language-model-only \
--disable-custom-all-reduce \
--enable-prefix-caching \
--tensor-parallel-size 2 \
--nnodes 2 \
--master-addr $HEAD_IP
Worker Node:vllm serve Kwaipilot/KAT-Coder-V2.5-Dev --port 8000 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--mm-encoder-tp-mode data \
--max-model-len auto \
--gpu-memory-utilization 0.85 \
--max-num-seqs 8 \
--kv-cache-memory-bytes 20g \
--max-num-batched-tokens 8192 \
--language-model-only \
--disable-custom-all-reduce \
--enable-prefix-caching \
--tensor-parallel-size 2 \
--nnodes 2 \
--master-addr $HEAD_IP \
--headless \
--node-rank 1
!!! Note for 0.24.0.pr.45544-r39.2.tegra-aarch64-cp312-cu132-24.04
Includes changes from https://github.com/vllm-project/vllm/pull/45544 rebased on v0.24.0 release.
This way you can run modelopt quantized LLM's (from this PR: https://github.com/vllm-project/vllm/pull/45306):
vllm serve nvidia/Gemma-4-26B-A4B-NVFP4 --port 8000 \
--quantization modelopt \
--kv-cache-dtype bfloat16 \
--tensor-parallel-size 1 \
--trust-remote-code \
--tool-call-parser gemma4 \
--reasoning-parser gemma4 \
--enable-auto-tool-choice \
--max-model-len 4096 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.85
!!! Note for vllm images tagged with TurboQuant:
turboquant_k8v4 will fail on Jetson Orin (sm_87): not supportedturboquant_4bit_ncExample running on Jetson Orin AGX with 64GB of RAM:
VLLM_MARLIN_USE_ATOMIC_ADD=1 \
vllm serve Qwen/Qwen3.6-35B-A3B-FP8 --port 8000 \
--enable-auto-tool-choice \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--max-model-len auto \
--gpu-memory-utilization 0.85 \
--max-num-seqs 10 \
--kv-cache-memory-bytes 10g \
--language-model-only \
--kv-cache-dtype turboquant_4bit_nc \
--compilation-config '{"cudagraph_mode": "PIECEWISE"}' \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
Shows :
51.5 GB262144 fits in available GPU memory432,640 tokens262,144 tokens per request: 5.71xMore info about TurboQuant in PR: [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity #38479
!!! Note: for vLLM:v0.15.0, if you experience degraded performance, check the vllm serve logs if using ...TRITON_MLA attention backend..., this could degrade performance with larger inputs. To disble MLA, set VLLM_MLA_DISABLE=1, example: VLLM_MLA_DISABLE=1 vllm serve ...
Built on Jetson Orin AGX (SM 87)
vLLM docs: https://docs.vllm.ai/en/stable/
Distributed LLM inferencing with vLLM, as described in the Multi-node case here https://docs.vllm.ai/en/latest/serving/parallelism_scaling.html#multi-node-
Tested with 2x Jetson AGX Orin dev kits.
Also try EAGLE-3 speculative decoding for faster inference (as explained here) :
vllm serve RedHatAI/Qwen3-8B-quantized.w4a16 \
--port 8000 \
--gpu-memory-utilization 0.6 \
--kv-cache-memory-bytes 5G \
--max-model-len 2048 \
--speculative-config '{
"model": "RedHatAI/Qwen3-8B-speculator.eagle3",
"num_speculative_tokens": 3,
"method": "eagle3"
}'
Content type
Image
Digest
sha256:b95e0b9c1…
Size
18.5 GB
Last updated
7 days ago
docker pull mitakad/vllm:0.30.0-r39.2.tegra-aarch64-cp312-cu132-24.04