SGLang + Codec binary transport (PRs #24483, #24557) + codec-supervisor control plane (BUSL-1.1)
1.6K
SGLang inference server with Codec token-native binary transport plus the codec-supervisor control plane. One container, GPU-accelerated, OpenAI-compatible.
License: BUSL-1.1 for the codec-supervisor wrapper. The bundled sglang remains under Apache-2.0. Production use under BUSL-1.1 is permitted up to USD $5M annual gross revenue; above that, contact [email protected].
docker run -d --gpus all \
-p 8080:8080 \
-v codec-models:/models \
-v codec-hf-cache:/root/.cache/huggingface \
-e CODEC_INITIAL_MODEL=Qwen/Qwen2.5-0.5B-Instruct \
--shm-size 8g \
wdunn001/codec-sglang:latest
Then:
# OpenAI-compatible
curl http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-d '{"model":"x","prompt":"Hello","max_tokens":20}'
# Codec wire format (msgpack frames of token IDs — ~200x smaller)
curl http://localhost:8080/v1/completions \
-d '{"model":"x","prompt":"Hello","max_tokens":20,"stream":true,"stream_format":"msgpack"}'
Three ways:
1. Override CODEC_INITIAL_MODEL — any HF repo id:
docker run --gpus all -p 8080:8080 \
-e CODEC_INITIAL_MODEL=meta-llama/Llama-3.1-8B-Instruct \
-e HF_TOKEN=hf_xxx \
wdunn001/codec-sglang:latest
2. Bind-mount a local checkpoint — fine-tunes you don't want on HF:
docker run --gpus all -p 8080:8080 \
-e CODEC_INITIAL_MODEL=/models/my-finetune \
-v /path/to/my-finetune:/models/my-finetune:ro \
wdunn001/codec-sglang:latest
3. Hot-swap via admin API — supervisor stays up, model swaps:
# Pull from HF
curl -X POST http://localhost:8080/admin/models/pull \
-H "Content-Type: application/json" \
-d '{"repo_id": "Qwen/Qwen2.5-7B-Instruct"}'
# Upload a tarball
tar -cf my-finetune.tar -C ./checkpoints/my-finetune .
curl -X POST "http://localhost:8080/admin/models/upload?name=my-finetune" \
-F "[email protected]"
# Hot-swap
curl -X POST http://localhost:8080/admin/load \
-H "Content-Type: application/json" \
-d '{"name":"my-finetune"}'
Pass CODEC_BACKEND_ARGS to tune sglang (--tp 2 --quantization fp8 ...).
lmsysorg/sglang:latest:8080 for model upload, HF pull, hot-swap, and an OpenAI-compatible reverse proxy.| Method | Path | What |
|---|---|---|
GET | /health | supervisor liveness |
GET | /admin/status | current model + uptime |
GET | /admin/models | list models in /models |
POST | /admin/models/pull | snapshot_download from HF — body {"repo_id":"..."} |
POST | /admin/models/upload | multipart tarball — name is a query param |
DELETE | /admin/models/{name} | remove from /models |
POST | /admin/load | hot-swap — body {"name":"...","allow_remote":true} for direct HF id |
POST | /admin/stop | stop backend (supervisor stays up) |
* | /v1/* | proxied to the backend |
Content type
Image
Digest
sha256:d3fe54495…
Size
12.4 GB
Last updated
4 months ago
docker pull wdunn001/codec-sglang