an ultra-light LLM API gateway in a single binary with embedded SQLite, ready out of the box.
335
Vanex Lite is an ultra-lightweight LLM API gateway: a single binary with embedded SQLite, ready to use out of the box. Clients connect via a unified OpenAI-compatible protocol, while the gateway converts and forwards requests to upstream providers such as OpenAI, Anthropic, Google Gemini, Alibaba Cloud DashScope, and Volcano Ark — with priority routing, failover, cooldown, and auto-disable built in.
This guide is aimed at end users and covers: deployment & startup, configuration, core concepts, client integration, the admin API, routing & fault-tolerance behavior, data backup, and day-to-day operations.
The Lite edition has a three-layer data model:
| Concept | Description |
|---|---|
| Provider | One upstream account: name + endpoint + protocol + API key + whether to use the proxy |
| Model Pool | The "model name" exposed to clients. The pool name is exactly the value of the model field in client requests (case-insensitive) |
| Route | The link between a pool and a provider, with a priority and an optional upstream model mapping (upstream_model) |
A typical relationship:
Client request model="vanex"
│
▼
Pool vanex ──route1(priority=10)──▶ Provider A (openai) upstream_model=gpt-4o
└─route2(priority=100)──▶ Provider B (gemini) upstream_model=gemini-2.5-pro
vanex is created automatically. It cannot be renamed or deleted.upstream_model, the request's model value is replaced with it when forwarding; otherwise the original model is passed through unchanged.| Protocol | Upstream | Supported Inbound Endpoints |
|---|---|---|
openai | OpenAI and any OpenAI-compatible service | All OpenAI-compatible endpoints |
anthropic | Anthropic official API | v1/chat/completions, v1/completions (auto-converted to the messages protocol) |
gemini | Google Gemini | v1/chat/completions, v1/completions (auto-converted to the generateContent protocol) |
dashscope | Alibaba Cloud DashScope (Wanxiang image/video) | v1/images/generations, v1/videos/generations |
ark | Volcano Ark (Seedream image / Seedance video) | v1/images/generations, v1/videos/generations |
cd lite
# Build the image (multi-stage build, producing an alpine runtime image)
docker build -t vanex-lite:latest .
# Start the container
docker run -d \
--name vanex-lite \
-p 6661:6661 \
-v $(pwd)/vanexlite-data:/app/data \
vanex-lite:latest
# Verify
curl http://localhost:6661/admin/status
Notes:
/app/data/vanexlite.db (set by the built-in VANEX_DB_PATH environment variable). Mount this directory for persistence.appuser. Make sure the host mount directory is writable by it (uid 1000): chown -R 1000:1000 vanexlite-data./admin/status every 30s), usable directly for Docker/K8s health checks.# Run in any directory (the static/ directory provides the admin UI; without it there is no UI)
mkdir -p /opt/vanexlite && cd /opt/vanexlite
cp /path/to/vanexlite .
cp -r /path/to/static .
./vanexlite
cd lite
RUST_LOG=info cargo build --release
./target/release/vanexlite
After a successful start, the service listens on [::]:6661 (change with VANEX_PORT). Visit http://localhost:6661 to open the admin panel.
# 1. Create a provider (using an OpenAI-compatible upstream as an example)
curl -X POST http://localhost:6661/admin/providers \
-H 'Content-Type: application/json' \
-d '{
"name": "my-openai",
"endpoint": "https://api.openai.com/v1",
"protocol": "openai",
"use_proxy": false,
"api_key": "sk-xxxx"
}'
# 2. Attach the provider to the default pool vanex (assuming the returned provider id is 1)
curl -X POST http://localhost:6661/admin/pools/1/routes \
-H 'Content-Type: application/json' \
-d '{ "provider_id": 1, "upstream_model": "gpt-4o", "priority": 10 }'
# 3. Call it
curl http://localhost:6661/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vanex",
"messages": [{"role": "user", "content": "Hello"}]
}'
Configuration precedence: environment variables > config.toml > built-in defaults.
| Variable | Default | Description |
|---|---|---|
VANEX_PORT | 6661 | HTTP listen port |
VANEX_DB_PATH | vanexlite.db | SQLite database file path |
VANEX_PROXY | none | Global HTTP proxy URL (used by providers with use_proxy=true) |
VANEX_COOLDOWN_SECS | 60 | Default cooldown duration (seconds) after a route is rate-limited (429) |
VANEX_COOLDOWN_MAX_SECS | 300 | Upper bound for cooldown duration (automatically kept ≥ VANEX_COOLDOWN_SECS) |
VANEX_STREAM_IDLE_TIMEOUT_SECS | 30 | Idle timeout for streaming responses (seconds) |
VANEX_FAILURE_THRESHOLD | 3 | Number of consecutive failures before a route is auto-disabled (minimum 1) |
VANEX_CONFIG | none | Explicit path to config.toml |
RUST_LOG | info | Log level (standard tracing filter, e.g. RUST_LOG=debug) |
Config file lookup order: path from VANEX_CONFIG > ~/.config/vanexlite/config.toml > ./config.toml in the current directory.
port = 6661
[database]
path = "./vanexlite.db"
[proxy]
url = "http://127.0.0.1:7890"
[limits]
cooldown_secs = 60
cooldown_max_secs = 300
stream_idle_timeout_secs = 30
failure_threshold = 3
The gateway itself performs no authentication (designed for intranet/trusted environments). Clients do not need a gateway token; the Authorization header in requests is ignored, and actual upstream authentication uses the API key configured on the provider.
| Endpoint | Description |
|---|---|
POST /v1/chat/completions | Chat completions (supports stream: true) |
POST /v1/completions | Text completions |
POST /v1/embeddings | Embeddings |
POST /v1/moderations | Content moderation |
POST /v1/files / GET /v1/files | File upload / list |
POST /v1/images/generations | Image generation |
POST /v1/videos/generations | Video generation |
POST /v1/audio/speech | Text-to-speech |
POST /v1/audio/transcriptions | Speech transcription (multipart) |
POST /v1/audio/translations | Speech translation (multipart) |
GET /v1/models | List all model pool names |
Request body limit is 64MB.
# Non-streaming
curl http://localhost:6661/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "vanex", "messages": [{"role": "user", "content": "Hello"}]}'
# Streaming (SSE)
curl -N http://localhost:6661/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "vanex", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
from openai import OpenAI
client = OpenAI(base_url="http://localhost:6661/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="vanex", # use the model pool name
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:6661/v1", apiKey: "not-needed" });
const resp = await client.chat.completions.create({
model: "vanex",
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
The
modelfield takes a model pool name, case-insensitive. Whenmodelis omitted, the global default route (the overall highest-priority one) is used.
POST /anthropic/v1/messages
curl http://localhost:6661/anthropic/v1/messages \
-H 'Content-Type: application/json' \
-d '{
"model": "vanex",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hi"}]
}'
POST /gemini/v1beta/models/{model}:generateContent
POST /gemini/v1beta/models/{model}:streamGenerateContent?alt=sse
curl http://localhost:6661/gemini/v1beta/models/vanex:generateContent \
-H 'Content-Type: application/json' \
-d '{"contents": [{"parts": [{"text": "Write a poem"}]}]}'
Successful responses carry an x-vanex-route: <pool>,<route_id>,<attempt_index> header, useful for identifying which upstream route actually served the request.
Visit http://localhost:6661/ in a browser to open the built-in static admin page (no login required), where you can:
All UI operations call the REST APIs in the next section, so scripted operations can simply use curl.
The admin API has no authentication — do not expose it directly to the public internet (see Security Recommendations).
GET /admin/status
Returns every pool's candidate routes (sorted by priority) with real-time status: active (available) / cooldown (cooling down, with remaining seconds) / disabled (disabled, with reason).
GET /admin/providers # list
POST /admin/providers # create
PATCH /admin/providers/:id # update (all fields)
DELETE /admin/providers/:id # delete (cascades to its routes)
GET /admin/providers/:id/models # fetch the provider's upstream model list
Create/update request body:
| Field | Type | Required | Description |
|---|---|---|---|
name | string | yes | Provider name (unique) |
endpoint | string | yes | Upstream base URL, e.g. https://api.openai.com/v1 |
protocol | string | no | openai (default) / anthropic / gemini / dashscope / ark |
use_proxy | bool | no | Whether to use the VANEX_PROXY proxy, default true |
api_key | string | no | Upstream API key, defaults to empty |
curl -X POST http://localhost:6661/admin/providers \
-H 'Content-Type: application/json' \
-d '{"name":"my-gemini","endpoint":"https://generativelanguage.googleapis.com","protocol":"gemini","use_proxy":false,"api_key":"AIza..."}'
# => 201 {"id": 2}
GET /admin/pools # list (including route details per pool)
POST /admin/pools # create
PATCH /admin/pools/:id # rename (the default pool vanex is forbidden)
DELETE /admin/pools/:id # delete (the default pool vanex is forbidden; cascades to routes)
curl -X POST http://localhost:6661/admin/pools \
-H 'Content-Type: application/json' \
-d '{"name": "my-claude"}'
POST /admin/pools/:pool_id/routes # attach a route
PATCH /admin/routes/:id # change priority
DELETE /admin/routes/:id # delete a route
Attach request body:
| Field | Type | Required | Description |
|---|---|---|---|
provider_id | int | yes | Provider ID |
upstream_model | string | no | Real upstream model name; if omitted, the client's model is passed through |
priority | int | no | Priority, default 100, lower value = higher priority |
curl -X POST http://localhost:6661/admin/pools/1/routes \
-H 'Content-Type: application/json' \
-d '{"provider_id": 2, "upstream_model": "gemini-2.5-pro", "priority": 20}'
GET /admin/settings/:key # read a setting
POST /admin/settings # write a setting (upsert)
The special setting active_route controls the global routing mode:
# auto: priority-based scheduling + failover (default)
curl -X POST http://localhost:6661/admin/settings \
-H 'Content-Type: application/json' \
-d '{"key": "active_route", "value": "auto"}'
# Pin to a specific route (by route ID); only that route is used
curl -X POST http://localhost:6661/admin/settings \
-H 'Content-Type: application/json' \
-d '{"key": "active_route", "value": "3"}'
This setting is persisted in the settings table of the database, survives restarts, and is synced to memory immediately upon write.
model against pool names (case-insensitive) and take all of the pool's routes, sorted by ascending priority;active=false), in cooldown, or have reached the consecutive-failure threshold;active_route ≠ auto), keep only that route;| Upstream Response | Handling |
|---|---|
| 2xx | Return success and reset the route's failure counter |
| 429 | Enter cooldown: seconds taken from the upstream Retry-After header or VANEX_COOLDOWN_SECS, capped at VANEX_COOLDOWN_MAX_SECS; then fail over |
| 5xx / 403 | Increment failure counter and fail over to the next route |
| Connection error | Cooldown + increment failure counter, then fail over |
Consecutive failures ≥ VANEX_FAILURE_THRESHOLD | Auto-disable the route and persist it (recording disabled_reason); manual re-enable via the admin panel/API is required |
| No route available | Return 429 no route available, retry later |
Note: the gateway's own errors are returned in a format normalized to the inbound protocol (OpenAI endpoints return OpenAI-style errors; Anthropic/Gemini native endpoints return their respective formats), so clients need no special handling.
All data lives in a single SQLite file (vanexlite.db by default), with 4 tables:
| Table | Content |
|---|---|
providers | Providers (name/endpoint/protocol/use_proxy/api_key) |
pools | Model pools |
model_api_keys | Routes (pool_id/provider_id/upstream_model/priority/active/disabled_reason) |
settings | Key-value settings (e.g. active_route) |
# Backup (you can copy the file while running, but with SQLite WAL a sqlite3 backup or a stopped service is safer)
cp vanexlite.db vanexlite.db.bak
# Safer online backup
sqlite3 vanexlite.db ".backup /backup/vanexlite-$(date +%F).db"
# Restore: stop the service → overwrite the file → restart
sqlite3 vanexlite.db ".restore /backup/vanexlite-2026-08-23.db"
SQLite uses a single-writer model. Ensure only one vanexlite process uses a given database file.
Logging goes through tracing (stderr), defaulting to the info level:
RUST_LOG=debug ./vanexlite # debugging
RUST_LOG=info ./vanexlite 2>&1 | tee vanexlite.log
Key log events: request received, forwarding to upstream, rate limited, route in cooldown, failover to next route, route auto-disabled after repeated failures.
[Unit]
Description=Vanex Lite LLM Gateway
After=network.target
[Service]
Type=simple
WorkingDirectory=/opt/vanexlite
ExecStart=/opt/vanexlite/vanexlite
Restart=always
RestartSec=5
Environment="VANEX_PORT=6661"
Environment="VANEX_DB_PATH=/opt/vanexlite/vanexlite.db"
Environment="RUST_LOG=info"
[Install]
WantedBy=multi-user.target
The service supports graceful shutdown on Ctrl+C / SIGINT (waiting up to 30 seconds for in-flight requests).
curl -f http://localhost:6661/admin/status
The Lite edition's admin API has no authentication and is intended for intranet use only. If public exposure is required, choose one of:
ufw allow from 192.168.1.0/24 to any port 6661);ssh -L 6661:localhost:6661 user@server).Q: Port already in use (Failed to bind)?
Change the port: VANEX_PORT=6662 ./vanexlite.
Q: The request returns Model not found?
The model field must match a model pool name (case-insensitive). Use GET /v1/models or GET /admin/pools to see existing pools.
Q: The request returns no route available, retry later?
All routes in the pool are in cooldown or disabled. Check GET /admin/status for the cause; auto-disabled routes must be re-enabled manually (enable them in the admin panel, or recreate the route).
Q: Database is locked error?
The same database file is being opened by multiple processes. Ensure only one vanexlite instance is running.
Q: I changed a provider/route but it didn't take effect? Admin API writes automatically hot-reload the in-memory route table, so no restart is needed; if you modified the database directly with sqlite3, a process restart is required.
Q: The container can't write to the database?
This is a mount directory permission issue — the container runs as uid 1000: chown -R 1000:1000 <mount-directory>.
Q: How do I change the port while keeping the image's health check working?
The health check uses ${VANEX_PORT:-6661}, so changing the VANEX_PORT environment variable adapts automatically; remember to update the container port mapping accordingly.
Content type
Image
Digest
sha256:210182819…
Size
10 MB
Last updated
about 1 month ago
docker pull vanexllm/vanex-lite