Sign inSign up

mitakad/sglang

By mitakad

•Updated 7 months ago

SGLang ( +parallelism ) on Jetson Orin AGX: built by https://github.com/dusty-nv/jetson-containers

Image
Machine learning & AI
2

1.1K

mitakad/sglang repository overview

Built on Jetson Orin AGX (SM 87)

SGLang docs for NVIDIA Jetson Orin: https://docs.sglang.io/platforms/nvidia_jetson.html⁠

Note: PyTorch is build with experimental NCCL support for Jetson, meaning multi-node (parallel) inference should be possible (more info https://github.com/sgl-project/sglang/issues/8164⁠)

Example usage with GPTQ quantized models (tested with SGL v0.6.0):

SGLANG_ENABLE_SPEC_V2=1 SGLANG_DISABLE_CUDNN_CHECK=1 \
  sglang serve --host 0.0.0.0 --port 8000 \
    --model-path Qwen/Qwen3.5-35B-A3B-GPTQ-Int4 \
    --tp-size 1 \
    --mem-fraction-static 0.9 \
    --context-length 2048 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --speculative-algo NEXTN \
    --speculative-num-steps 3 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 4 \
    --quantization moe_wna16 \
    --mamba-scheduler-strategy extra_buffer
python -m sglang.launch_server --host 0.0.0.0 --port 8000 \
    --model-path Qwen/Qwen3-0.6B-GPTQ-Int8 \
    --tp-size 1 \
    --mem-fraction-static 0.8 \
    --context-length 2048 \
    --reasoning-parser qwen3 \
    --attention-backend flashinfer \
    --quantization gptq

Example usage with FP8 quantized models (tested with SGL v0.6.0):

!!! Note: the Qwen/Qwen3.5-35B-A3B-FP8 timeouts on inference.

  python -m sglang.launch_server --host 0.0.0.0 --port 8000 \
    --model-path Qwen/Qwen3-0.6B-FP8 \
    --tp-size 1 \
    --mem-fraction-static 0.8 \
    --context-length 2048 \
    --reasoning-parser qwen3 \
    --attention-backend flashinfer \
    --quantization fp8

Example usage:

  • Run with jetson-containers run IMAGE_NAME

  • To serve for example Qwen/Qwen3-4B-Instruct-2507:

python3 -m sglang.launch_server --host 0.0.0.0   --port 8000 \
  --model-path Qwen/Qwen3-4B-Instruct-2507 \
  --mem-fraction-static 0.5 \
  --context-length 8192
  • Request:
curl --location 'http://localhost:8000/v1/chat/completions' \
--header 'Content-Type: application/json' \
--data '{
    "model": "Qwen/Qwen3-4B-Instruct-2507",
    "messages": [
        {
            "role": "user",
            "content": "Why is the sky blue?"
        }
    ]
}'
  • Image contents/ check_env() output for mitakad/sglang:0.5.4-r36.4.tegra-aarch64-cp310-cu126-22.04-truncated
Built with: ENABLE_DISTRIBUTED_JETSON_NCCL=1 PYTORCH_FORCE_BUILD=on CUDA_VERSION=12.6 PYTHON_VERSION=3.10 LSB_RELEASE=22.04 PYTORCH_VERSION=2.9 jetson-containers build sglang:0.5.4-builder

testing SGLang...
✅ Memory cleared
Python: 3.12.12 (main, Oct 14 2025, 21:26:46) [Clang 20.1.4 ]
CUDA available: True
GPU 0: Orin
GPU 0 Compute Capability: 8.7
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 12.9, V12.9.86
CUDA Driver Version: 540.4.0
PyTorch: 2.9.0
sglang: 0.5.3.post3
sgl_kernel: 0.3.16.post3
flashinfer_python: 0.4.1
triton: 3.4.0
transformers: 4.57.1
torchao: 0.9.0
numpy: 2.3.4
aiohttp: 3.13.1
fastapi: 0.119.1
hf_transfer: 0.1.9
huggingface_hub: 0.35.3
interegular: 0.3.3
modelscope: 1.31.0
orjson: 3.11.3
outlines: 1.2.7
packaging: 25.0
psutil: 7.1.1
pydantic: 2.12.3
python-multipart: 0.0.20
pyzmq: 27.1.0
uvicorn: 0.38.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.1.25
openai: 2.6.0
tiktoken: 0.12.0
anthropic: 0.71.0
litellm: Module Not Found
decord: Module Not Found
ulimit soft: 1048576
SGLang OK
  • Image contents/ check_env() output for mitakad/sglang:0.5.4-r36.4.tegra-aarch64-cp312-cu129-24.04-truncated
Built with: ENABLE_DISTRIBUTED_JETSON_NCCL=1 PYTORCH_FORCE_BUILD=on CUDA_VERSION=12.9 PYTHON_VERSION=3.12 LSB_RELEASE=24.04 PYTORCH_VERSION=2.9 jetson-containers build sglang:0.5.4-builder

testing SGLang...
✅ Memory cleared
Python: 3.12.12 (main, Oct 14 2025, 21:26:46) [Clang 20.1.4 ]
CUDA available: True
GPU 0: Orin
GPU 0 Compute Capability: 8.7
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 12.9, V12.9.86
CUDA Driver Version: 540.4.0
PyTorch: 2.9.0
sglang: 0.5.4
sgl_kernel: 0.3.16.post3
flashinfer_python: 0.4.1
triton: 3.4.0
transformers: 4.57.1
torchao: 0.9.0
numpy: 2.3.4
aiohttp: 3.13.1
fastapi: 0.120.0
hf_transfer: 0.1.9
huggingface_hub: 0.36.0
interegular: 0.3.3
modelscope: 1.31.0
orjson: 3.11.4
outlines: 1.2.7
packaging: 25.0
psutil: 7.1.1
pydantic: 2.12.3
python-multipart: 0.0.20
pyzmq: 27.1.0
uvicorn: 0.38.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.1.25
openai: 2.6.1
tiktoken: 0.12.0
anthropic: 0.71.0
litellm: 1.79.0
decord2: 2.0.0
ulimit soft: 1048576
SGLang OK

Tag summary

Content type

Image

Digest

sha256:eb4fe0b69…

Size

9.2 GB

Last updated

7 months ago

docker pull mitakad/sglang:0.6.0-r36.5.tegra-aarch64-cp312-cu129-24.04-at-commit-d093e70