Sign inSign up

wdunn001/codec-sglang

By wdunn001

Updated 4 months ago

SGLang + Codec binary transport (PRs #24483, #24557) + codec-supervisor control plane (BUSL-1.1)

Image
0

1.6K

wdunn001/codec-sglang repository overview

codec-sglang

SGLang inference server with Codec token-native binary transport plus the codec-supervisor control plane. One container, GPU-accelerated, OpenAI-compatible.

License: BUSL-1.1 for the codec-supervisor wrapper. The bundled sglang remains under Apache-2.0. Production use under BUSL-1.1 is permitted up to USD $5M annual gross revenue; above that, contact [email protected].

Quick start

docker run -d --gpus all \
  -p 8080:8080 \
  -v codec-models:/models \
  -v codec-hf-cache:/root/.cache/huggingface \
  -e CODEC_INITIAL_MODEL=Qwen/Qwen2.5-0.5B-Instruct \
  --shm-size 8g \
  wdunn001/codec-sglang:latest

Then:

# OpenAI-compatible
curl http://localhost:8080/v1/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"x","prompt":"Hello","max_tokens":20}'

# Codec wire format (msgpack frames of token IDs — ~200x smaller)
curl http://localhost:8080/v1/completions \
  -d '{"model":"x","prompt":"Hello","max_tokens":20,"stream":true,"stream_format":"msgpack"}'

Run your own model

Three ways:

1. Override CODEC_INITIAL_MODEL — any HF repo id:

docker run --gpus all -p 8080:8080 \
  -e CODEC_INITIAL_MODEL=meta-llama/Llama-3.1-8B-Instruct \
  -e HF_TOKEN=hf_xxx \
  wdunn001/codec-sglang:latest

2. Bind-mount a local checkpoint — fine-tunes you don't want on HF:

docker run --gpus all -p 8080:8080 \
  -e CODEC_INITIAL_MODEL=/models/my-finetune \
  -v /path/to/my-finetune:/models/my-finetune:ro \
  wdunn001/codec-sglang:latest

3. Hot-swap via admin API — supervisor stays up, model swaps:

# Pull from HF
curl -X POST http://localhost:8080/admin/models/pull \
  -H "Content-Type: application/json" \
  -d '{"repo_id": "Qwen/Qwen2.5-7B-Instruct"}'

# Upload a tarball
tar -cf my-finetune.tar -C ./checkpoints/my-finetune .
curl -X POST "http://localhost:8080/admin/models/upload?name=my-finetune" \
  -F "[email protected]"

# Hot-swap
curl -X POST http://localhost:8080/admin/load \
  -H "Content-Type: application/json" \
  -d '{"name":"my-finetune"}'

Pass CODEC_BACKEND_ARGS to tune sglang (--tp 2 --quantization fp8 ...).

What's inside

  • Base: lmsysorg/sglang:latest
  • Codec patches: sgl-project/sglang#24483 (token-native binary transport) and #24557 (server-side ToolWatcher), applied as an editable overlay so all upstream kernels (flash-attn, sgl_kernel, triton) stay intact.
  • Control plane: codec-supervisor — FastAPI admin on :8080 for model upload, HF pull, hot-swap, and an OpenAI-compatible reverse proxy.

Admin endpoints

MethodPathWhat
GET/healthsupervisor liveness
GET/admin/statuscurrent model + uptime
GET/admin/modelslist models in /models
POST/admin/models/pullsnapshot_download from HF — body {"repo_id":"..."}
POST/admin/models/uploadmultipart tarball — name is a query param
DELETE/admin/models/{name}remove from /models
POST/admin/loadhot-swap — body {"name":"...","allow_remote":true} for direct HF id
POST/admin/stopstop backend (supervisor stays up)
*/v1/*proxied to the backend

Source & docs

Tag summary

Content type

Image

Digest

sha256:d3fe54495

Size

12.4 GB

Last updated

4 months ago

docker pull wdunn001/codec-sglang