Boreas measures the raw performance of LLM inference engines — latency, throughput, capacity, stability — against any OpenAI-compatible
/v1/chat/completions endpoint (llama.cpp, vLLM, Ollama, …). It benchmarks the serving layer, not the quality of the answers.
docker run -d --name boreas -p 8070:8070 redteamsfr/boreas:latest
Results appear when the run completes: latency percentiles, RPS, error breakdown, generator health.
Verify the instance:
curl -s http://localhost:8070/healthz
The demo ships without authentication so it can be tried in one command. If you expose the port beyond localhost, set a token:
docker run -d -p 8070:8070 -e BOREAS_TOKEN="$(openssl rand -base64 32)" \
redteamsfr/boreas:latest
Content type
Image
Digest
sha256:f9746ad78…
Size
16.6 MB
Last updated
13 days ago
docker pull redteamsfr/boreas