Sign inSign up

vanexllm/vanex-lite

By vanexllm

•Updated about 1 month ago

an ultra-light LLM API gateway in a single binary with embedded SQLite, ready out of the box.

Image
Networking
API management
Machine learning & AI
1

335

vanexllm/vanex-lite repository overview

⁠Vanex Lite User Guide

Vanex Lite is an ultra-lightweight LLM API gateway: a single binary with embedded SQLite, ready to use out of the box. Clients connect via a unified OpenAI-compatible protocol, while the gateway converts and forwards requests to upstream providers such as OpenAI, Anthropic, Google Gemini, Alibaba Cloud DashScope, and Volcano Ark — with priority routing, failover, cooldown, and auto-disable built in.

This guide is aimed at end users and covers: deployment & startup, configuration, core concepts, client integration, the admin API, routing & fault-tolerance behavior, data backup, and day-to-day operations.


⁠Table of Contents

  1. Core Concepts⁠
  2. Quick Start⁠
  3. Configuration Reference⁠
  4. Client Integration⁠
  5. Admin Panel⁠
  6. Admin API Reference⁠
  7. Routing & Fault Tolerance⁠
  8. Data Storage & Backup⁠
  9. Operations Guide⁠
  10. FAQ⁠

⁠1. Core Concepts

The Lite edition has a three-layer data model:

ConceptDescription
ProviderOne upstream account: name + endpoint + protocol + API key + whether to use the proxy
Model PoolThe "model name" exposed to clients. The pool name is exactly the value of the model field in client requests (case-insensitive)
RouteThe link between a pool and a provider, with a priority and an optional upstream model mapping (upstream_model)

A typical relationship:

Client request model="vanex"
      │
      ▼
Pool vanex ──route1(priority=10)──▶ Provider A (openai)   upstream_model=gpt-4o
           └─route2(priority=100)──▶ Provider B (gemini)  upstream_model=gemini-2.5-pro
  • On first startup, a default pool named vanex is created automatically. It cannot be renamed or deleted.
  • A pool can hold multiple routes, tried in ascending order of priority (lower value = higher priority, default 100).
  • If a route sets upstream_model, the request's model value is replaced with it when forwarding; otherwise the original model is passed through unchanged.
⁠Supported Upstream Protocols
ProtocolUpstreamSupported Inbound Endpoints
openaiOpenAI and any OpenAI-compatible serviceAll OpenAI-compatible endpoints
anthropicAnthropic official APIv1/chat/completions, v1/completions (auto-converted to the messages protocol)
geminiGoogle Geminiv1/chat/completions, v1/completions (auto-converted to the generateContent protocol)
dashscopeAlibaba Cloud DashScope (Wanxiang image/video)v1/images/generations, v1/videos/generations
arkVolcano Ark (Seedream image / Seedance video)v1/images/generations, v1/videos/generations

⁠2. Quick Start

cd lite

# Build the image (multi-stage build, producing an alpine runtime image)
docker build -t vanex-lite:latest .

# Start the container
docker run -d \
    --name vanex-lite \
    -p 6661:6661 \
    -v $(pwd)/vanexlite-data:/app/data \
    vanex-lite:latest

# Verify
curl http://localhost:6661/admin/status

Notes:

  • Inside the container the database is fixed at /app/data/vanexlite.db (set by the built-in VANEX_DB_PATH environment variable). Mount this directory for persistence.
  • The image runs as the unprivileged user appuser. Make sure the host mount directory is writable by it (uid 1000): chown -R 1000:1000 vanexlite-data.
  • A HEALTHCHECK is built into the image (probing /admin/status every 30s), usable directly for Docker/K8s health checks.
⁠Option B: Running the Binary Directly
# Run in any directory (the static/ directory provides the admin UI; without it there is no UI)
mkdir -p /opt/vanexlite && cd /opt/vanexlite
cp /path/to/vanexlite .
cp -r /path/to/static .

./vanexlite
⁠Option C: Building from Source
cd lite
RUST_LOG=info cargo build --release
./target/release/vanexlite

After a successful start, the service listens on [::]:6661 (change with VANEX_PORT). Visit http://localhost:6661 to open the admin panel.

⁠Three Steps to First Configuration
# 1. Create a provider (using an OpenAI-compatible upstream as an example)
curl -X POST http://localhost:6661/admin/providers \
  -H 'Content-Type: application/json' \
  -d '{
    "name": "my-openai",
    "endpoint": "https://api.openai.com/v1",
    "protocol": "openai",
    "use_proxy": false,
    "api_key": "sk-xxxx"
  }'

# 2. Attach the provider to the default pool vanex (assuming the returned provider id is 1)
curl -X POST http://localhost:6661/admin/pools/1/routes \
  -H 'Content-Type: application/json' \
  -d '{ "provider_id": 1, "upstream_model": "gpt-4o", "priority": 10 }'

# 3. Call it
curl http://localhost:6661/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "vanex",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

⁠3. Configuration Reference

Configuration precedence: environment variables > config.toml > built-in defaults.

⁠Environment Variables
VariableDefaultDescription
VANEX_PORT6661HTTP listen port
VANEX_DB_PATHvanexlite.dbSQLite database file path
VANEX_PROXYnoneGlobal HTTP proxy URL (used by providers with use_proxy=true)
VANEX_COOLDOWN_SECS60Default cooldown duration (seconds) after a route is rate-limited (429)
VANEX_COOLDOWN_MAX_SECS300Upper bound for cooldown duration (automatically kept ≥ VANEX_COOLDOWN_SECS)
VANEX_STREAM_IDLE_TIMEOUT_SECS30Idle timeout for streaming responses (seconds)
VANEX_FAILURE_THRESHOLD3Number of consecutive failures before a route is auto-disabled (minimum 1)
VANEX_CONFIGnoneExplicit path to config.toml
RUST_LOGinfoLog level (standard tracing filter, e.g. RUST_LOG=debug)
⁠config.toml

Config file lookup order: path from VANEX_CONFIG > ~/.config/vanexlite/config.toml > ./config.toml in the current directory.

port = 6661

[database]
path = "./vanexlite.db"

[proxy]
url = "http://127.0.0.1:7890"

[limits]
cooldown_secs = 60
cooldown_max_secs = 300
stream_idle_timeout_secs = 30
failure_threshold = 3

⁠4. Client Integration

The gateway itself performs no authentication (designed for intranet/trusted environments). Clients do not need a gateway token; the Authorization header in requests is ignored, and actual upstream authentication uses the API key configured on the provider.

⁠4.1 OpenAI-Compatible Endpoints
EndpointDescription
POST /v1/chat/completionsChat completions (supports stream: true)
POST /v1/completionsText completions
POST /v1/embeddingsEmbeddings
POST /v1/moderationsContent moderation
POST /v1/files / GET /v1/filesFile upload / list
POST /v1/images/generationsImage generation
POST /v1/videos/generationsVideo generation
POST /v1/audio/speechText-to-speech
POST /v1/audio/transcriptionsSpeech transcription (multipart)
POST /v1/audio/translationsSpeech translation (multipart)
GET /v1/modelsList all model pool names

Request body limit is 64MB.

⁠curl Examples
# Non-streaming
curl http://localhost:6661/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "vanex", "messages": [{"role": "user", "content": "Hello"}]}'

# Streaming (SSE)
curl -N http://localhost:6661/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "vanex", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
⁠Python (openai SDK)
from openai import OpenAI

client = OpenAI(base_url="http://localhost:6661/v1", api_key="not-needed")
resp = client.chat.completions.create(
    model="vanex",   # use the model pool name
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
⁠Node.js (openai SDK)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "http://localhost:6661/v1", apiKey: "not-needed" });
const resp = await client.chat.completions.create({
  model: "vanex",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);

The model field takes a model pool name, case-insensitive. When model is omitted, the global default route (the overall highest-priority one) is used.

⁠4.2 Anthropic Native Endpoint
POST /anthropic/v1/messages
curl http://localhost:6661/anthropic/v1/messages \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "vanex",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hi"}]
  }'
⁠4.3 Gemini Native Endpoint
POST /gemini/v1beta/models/{model}:generateContent
POST /gemini/v1beta/models/{model}:streamGenerateContent?alt=sse
curl http://localhost:6661/gemini/v1beta/models/vanex:generateContent \
  -H 'Content-Type: application/json' \
  -d '{"contents": [{"parts": [{"text": "Write a poem"}]}]}'
⁠4.4 Debugging Response Header

Successful responses carry an x-vanex-route: <pool>,<route_id>,<attempt_index> header, useful for identifying which upstream route actually served the request.


⁠5. Admin Panel

Visit http://localhost:6661/ in a browser to open the built-in static admin page (no login required), where you can:

  • View real-time status of all pools and routes (active / cooldown / disabled)
  • CRUD providers and fetch a provider's upstream model list
  • CRUD model pools
  • Attach routes, adjust priorities, enable/disable routes
  • Switch routing mode (auto scheduling / pin a specific route)

All UI operations call the REST APIs in the next section, so scripted operations can simply use curl.


⁠6. Admin API Reference

The admin API has no authentication — do not expose it directly to the public internet (see Security Recommendations⁠).

⁠6.1 Status
GET /admin/status

Returns every pool's candidate routes (sorted by priority) with real-time status: active (available) / cooldown (cooling down, with remaining seconds) / disabled (disabled, with reason).

⁠6.2 Providers
GET    /admin/providers              # list
POST   /admin/providers              # create
PATCH  /admin/providers/:id          # update (all fields)
DELETE /admin/providers/:id          # delete (cascades to its routes)
GET    /admin/providers/:id/models   # fetch the provider's upstream model list

Create/update request body:

FieldTypeRequiredDescription
namestringyesProvider name (unique)
endpointstringyesUpstream base URL, e.g. https://api.openai.com/v1
protocolstringnoopenai (default) / anthropic / gemini / dashscope / ark
use_proxyboolnoWhether to use the VANEX_PROXY proxy, default true
api_keystringnoUpstream API key, defaults to empty
curl -X POST http://localhost:6661/admin/providers \
  -H 'Content-Type: application/json' \
  -d '{"name":"my-gemini","endpoint":"https://generativelanguage.googleapis.com","protocol":"gemini","use_proxy":false,"api_key":"AIza..."}'
# => 201 {"id": 2}
⁠6.3 Model Pools
GET    /admin/pools          # list (including route details per pool)
POST   /admin/pools          # create
PATCH  /admin/pools/:id      # rename (the default pool vanex is forbidden)
DELETE /admin/pools/:id      # delete (the default pool vanex is forbidden; cascades to routes)
curl -X POST http://localhost:6661/admin/pools \
  -H 'Content-Type: application/json' \
  -d '{"name": "my-claude"}'
⁠6.4 Routes
POST   /admin/pools/:pool_id/routes   # attach a route
PATCH  /admin/routes/:id              # change priority
DELETE /admin/routes/:id              # delete a route

Attach request body:

FieldTypeRequiredDescription
provider_idintyesProvider ID
upstream_modelstringnoReal upstream model name; if omitted, the client's model is passed through
priorityintnoPriority, default 100, lower value = higher priority
curl -X POST http://localhost:6661/admin/pools/1/routes \
  -H 'Content-Type: application/json' \
  -d '{"provider_id": 2, "upstream_model": "gemini-2.5-pro", "priority": 20}'
⁠6.5 Settings & Routing Mode
GET  /admin/settings/:key        # read a setting
POST /admin/settings             # write a setting (upsert)

The special setting active_route controls the global routing mode:

# auto: priority-based scheduling + failover (default)
curl -X POST http://localhost:6661/admin/settings \
  -H 'Content-Type: application/json' \
  -d '{"key": "active_route", "value": "auto"}'

# Pin to a specific route (by route ID); only that route is used
curl -X POST http://localhost:6661/admin/settings \
  -H 'Content-Type: application/json' \
  -d '{"key": "active_route", "value": "3"}'

This setting is persisted in the settings table of the database, survives restarts, and is synced to memory immediately upon write.


⁠7. Routing & Fault Tolerance

⁠Scheduling Flow
  1. Match the request's model against pool names (case-insensitive) and take all of the pool's routes, sorted by ascending priority;
  2. Filter out routes that are: disabled (active=false), in cooldown, or have reached the consecutive-failure threshold;
  3. If a pinned route is set (active_route ≠ auto), keep only that route;
  4. Try them in order; on retryable errors, automatically fail over to the next one.
⁠Error Handling Policy
Upstream ResponseHandling
2xxReturn success and reset the route's failure counter
429Enter cooldown: seconds taken from the upstream Retry-After header or VANEX_COOLDOWN_SECS, capped at VANEX_COOLDOWN_MAX_SECS; then fail over
5xx / 403Increment failure counter and fail over to the next route
Connection errorCooldown + increment failure counter, then fail over
Consecutive failures ≥ VANEX_FAILURE_THRESHOLDAuto-disable the route and persist it (recording disabled_reason); manual re-enable via the admin panel/API is required
No route availableReturn 429 no route available, retry later

Note: the gateway's own errors are returned in a format normalized to the inbound protocol (OpenAI endpoints return OpenAI-style errors; Anthropic/Gemini native endpoints return their respective formats), so clients need no special handling.


⁠8. Data Storage & Backup

All data lives in a single SQLite file (vanexlite.db by default), with 4 tables:

TableContent
providersProviders (name/endpoint/protocol/use_proxy/api_key)
poolsModel pools
model_api_keysRoutes (pool_id/provider_id/upstream_model/priority/active/disabled_reason)
settingsKey-value settings (e.g. active_route)
# Backup (you can copy the file while running, but with SQLite WAL a sqlite3 backup or a stopped service is safer)
cp vanexlite.db vanexlite.db.bak

# Safer online backup
sqlite3 vanexlite.db ".backup /backup/vanexlite-$(date +%F).db"

# Restore: stop the service → overwrite the file → restart
sqlite3 vanexlite.db ".restore /backup/vanexlite-2026-08-23.db"

SQLite uses a single-writer model. Ensure only one vanexlite process uses a given database file.


⁠9. Operations Guide

⁠Logging

Logging goes through tracing (stderr), defaulting to the info level:

RUST_LOG=debug ./vanexlite          # debugging
RUST_LOG=info ./vanexlite 2>&1 | tee vanexlite.log

Key log events: request received, forwarding to upstream, rate limited, route in cooldown, failover to next route, route auto-disabled after repeated failures.

⁠Systemd Service Example
[Unit]
Description=Vanex Lite LLM Gateway
After=network.target

[Service]
Type=simple
WorkingDirectory=/opt/vanexlite
ExecStart=/opt/vanexlite/vanexlite
Restart=always
RestartSec=5
Environment="VANEX_PORT=6661"
Environment="VANEX_DB_PATH=/opt/vanexlite/vanexlite.db"
Environment="RUST_LOG=info"

[Install]
WantedBy=multi-user.target

The service supports graceful shutdown on Ctrl+C / SIGINT (waiting up to 30 seconds for in-flight requests).

⁠Health Check
curl -f http://localhost:6661/admin/status
⁠Security Recommendations

The Lite edition's admin API has no authentication and is intended for intranet use only. If public exposure is required, choose one of:

  • Nginx reverse proxy + Basic Auth / IP allowlist;
  • Firewall restrictions on source subnets (e.g. ufw allow from 192.168.1.0/24 to any port 6661);
  • SSH tunnel access (ssh -L 6661:localhost:6661 user@server).

⁠10. FAQ

Q: Port already in use (Failed to bind)? Change the port: VANEX_PORT=6662 ./vanexlite.

Q: The request returns Model not found? The model field must match a model pool name (case-insensitive). Use GET /v1/models or GET /admin/pools to see existing pools.

Q: The request returns no route available, retry later? All routes in the pool are in cooldown or disabled. Check GET /admin/status for the cause; auto-disabled routes must be re-enabled manually (enable them in the admin panel, or recreate the route).

Q: Database is locked error? The same database file is being opened by multiple processes. Ensure only one vanexlite instance is running.

Q: I changed a provider/route but it didn't take effect? Admin API writes automatically hot-reload the in-memory route table, so no restart is needed; if you modified the database directly with sqlite3, a process restart is required.

Q: The container can't write to the database? This is a mount directory permission issue — the container runs as uid 1000: chown -R 1000:1000 <mount-directory>.

Q: How do I change the port while keeping the image's health check working? The health check uses ${VANEX_PORT:-6661}, so changing the VANEX_PORT environment variable adapts automatically; remember to update the container port mapping accordingly.

Tag summary

Content type

Image

Digest

sha256:210182819…

Size

10 MB

Last updated

about 1 month ago

docker pull vanexllm/vanex-lite