SGLang ( +parallelism ) on Jetson Orin AGX: built by https://github.com/dusty-nv/jetson-containers
1.1K
Built on Jetson Orin AGX (SM 87)
SGLang docs for NVIDIA Jetson Orin: https://docs.sglang.io/platforms/nvidia_jetson.html
Note: PyTorch is build with experimental NCCL support for Jetson, meaning multi-node (parallel) inference should be possible (more info https://github.com/sgl-project/sglang/issues/8164)
Example usage with GPTQ quantized models (tested with SGL v0.6.0):
SGLANG_ENABLE_SPEC_V2=1 SGLANG_DISABLE_CUDNN_CHECK=1 \
sglang serve --host 0.0.0.0 --port 8000 \
--model-path Qwen/Qwen3.5-35B-A3B-GPTQ-Int4 \
--tp-size 1 \
--mem-fraction-static 0.9 \
--context-length 2048 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-algo NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--quantization moe_wna16 \
--mamba-scheduler-strategy extra_buffer
python -m sglang.launch_server --host 0.0.0.0 --port 8000 \
--model-path Qwen/Qwen3-0.6B-GPTQ-Int8 \
--tp-size 1 \
--mem-fraction-static 0.8 \
--context-length 2048 \
--reasoning-parser qwen3 \
--attention-backend flashinfer \
--quantization gptq
Example usage with FP8 quantized models (tested with SGL v0.6.0):
!!! Note: the Qwen/Qwen3.5-35B-A3B-FP8 timeouts on inference.
python -m sglang.launch_server --host 0.0.0.0 --port 8000 \
--model-path Qwen/Qwen3-0.6B-FP8 \
--tp-size 1 \
--mem-fraction-static 0.8 \
--context-length 2048 \
--reasoning-parser qwen3 \
--attention-backend flashinfer \
--quantization fp8
Example usage:
Run with jetson-containers run IMAGE_NAME
To serve for example Qwen/Qwen3-4B-Instruct-2507:
python3 -m sglang.launch_server --host 0.0.0.0 --port 8000 \
--model-path Qwen/Qwen3-4B-Instruct-2507 \
--mem-fraction-static 0.5 \
--context-length 8192
curl --location 'http://localhost:8000/v1/chat/completions' \
--header 'Content-Type: application/json' \
--data '{
"model": "Qwen/Qwen3-4B-Instruct-2507",
"messages": [
{
"role": "user",
"content": "Why is the sky blue?"
}
]
}'
mitakad/sglang:0.5.4-r36.4.tegra-aarch64-cp310-cu126-22.04-truncatedBuilt with: ENABLE_DISTRIBUTED_JETSON_NCCL=1 PYTORCH_FORCE_BUILD=on CUDA_VERSION=12.6 PYTHON_VERSION=3.10 LSB_RELEASE=22.04 PYTORCH_VERSION=2.9 jetson-containers build sglang:0.5.4-builder
testing SGLang...
✅ Memory cleared
Python: 3.12.12 (main, Oct 14 2025, 21:26:46) [Clang 20.1.4 ]
CUDA available: True
GPU 0: Orin
GPU 0 Compute Capability: 8.7
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 12.9, V12.9.86
CUDA Driver Version: 540.4.0
PyTorch: 2.9.0
sglang: 0.5.3.post3
sgl_kernel: 0.3.16.post3
flashinfer_python: 0.4.1
triton: 3.4.0
transformers: 4.57.1
torchao: 0.9.0
numpy: 2.3.4
aiohttp: 3.13.1
fastapi: 0.119.1
hf_transfer: 0.1.9
huggingface_hub: 0.35.3
interegular: 0.3.3
modelscope: 1.31.0
orjson: 3.11.3
outlines: 1.2.7
packaging: 25.0
psutil: 7.1.1
pydantic: 2.12.3
python-multipart: 0.0.20
pyzmq: 27.1.0
uvicorn: 0.38.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.1.25
openai: 2.6.0
tiktoken: 0.12.0
anthropic: 0.71.0
litellm: Module Not Found
decord: Module Not Found
ulimit soft: 1048576
SGLang OK
mitakad/sglang:0.5.4-r36.4.tegra-aarch64-cp312-cu129-24.04-truncatedBuilt with: ENABLE_DISTRIBUTED_JETSON_NCCL=1 PYTORCH_FORCE_BUILD=on CUDA_VERSION=12.9 PYTHON_VERSION=3.12 LSB_RELEASE=24.04 PYTORCH_VERSION=2.9 jetson-containers build sglang:0.5.4-builder
testing SGLang...
✅ Memory cleared
Python: 3.12.12 (main, Oct 14 2025, 21:26:46) [Clang 20.1.4 ]
CUDA available: True
GPU 0: Orin
GPU 0 Compute Capability: 8.7
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 12.9, V12.9.86
CUDA Driver Version: 540.4.0
PyTorch: 2.9.0
sglang: 0.5.4
sgl_kernel: 0.3.16.post3
flashinfer_python: 0.4.1
triton: 3.4.0
transformers: 4.57.1
torchao: 0.9.0
numpy: 2.3.4
aiohttp: 3.13.1
fastapi: 0.120.0
hf_transfer: 0.1.9
huggingface_hub: 0.36.0
interegular: 0.3.3
modelscope: 1.31.0
orjson: 3.11.4
outlines: 1.2.7
packaging: 25.0
psutil: 7.1.1
pydantic: 2.12.3
python-multipart: 0.0.20
pyzmq: 27.1.0
uvicorn: 0.38.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.1.25
openai: 2.6.1
tiktoken: 0.12.0
anthropic: 0.71.0
litellm: 1.79.0
decord2: 2.0.0
ulimit soft: 1048576
SGLang OK
Content type
Image
Digest
sha256:eb4fe0b69…
Size
9.2 GB
Last updated
7 months ago
docker pull mitakad/sglang:0.6.0-r36.5.tegra-aarch64-cp312-cu129-24.04-at-commit-d093e70