Air-gapped LLM inference benchmarking tool (llama-benchy) on python:3.12-slim
1.6K
An LLM inference benchmarking tool for OpenAI-compatible endpoints (vLLM, SGLang, and similar), packaged for fully offline / air-gapped use.
llama-benchy drives an OpenAI-compatible /v1 API and measures inference
performance across prompt-processing (prefill) and token-generation (decode)
workloads. You sweep combinations of:
--pp — prompt (prefill) lengths in tokens, e.g. 512 1024 2048--tg — tokens to generate (decode length), e.g. 32 128--depth — pre-filled context depth, e.g. 0 2048 4096--runs — repetitions per case--concurrency — parallel request count--enable-prefix-caching / --no-cache)Results can be emitted in Markdown (default) or saved to a file via
--save-result. Prompts are built from a corpus (a Project Gutenberg book,
pre-cached in the image) and tokenized with a local tokenizer so prompt lengths
are exact.
This image is built for offline operation: HuggingFace/transformers run in
offline mode at runtime, the gpt2 fallback tokenizer and the default corpus
book are baked in at build time, and no network fetches to HF/PyPI/Gutenberg
happen when the container runs.
docker pull n8500x/llama-benchy
The entrypoint wraps llama-benchy and auto-injects --tokenizer when a
tokenizer is mounted at /tokenizer. A typical run against an OpenAI-compatible
server:
docker run --rm --network host \
-v /srv/tokenizers/llama3:/tokenizer:ro \
n8500x/llama-benchy \
--base-url http://localhost:8000/v1 \
--api-key EMPTY \
--model meta-llama/Llama-3.1-8B-Instruct \
--pp 512 1024 2048 --tg 32 128 --depth 0 \
--runs 3 --concurrency 1 --format md
Show help / version (no tokenizer required):
docker run --rm n8500x/llama-benchy --help
docker run --rm n8500x/llama-benchy --version
If no tokenizer is mounted and none is passed, you can still use a pre-cached
tokenizer (e.g. --tokenizer gpt2).
The repo ships two Makefiles that wrap the docker run invocation. Configure
once by copying .env.example to .env and setting your endpoint and
TOKENIZER_DIR.
Benchmarking (Makefile):
make bench # run with current .env config
make bench CONFIG=vllm.env # use an alternate config file
make bench-save # run and save results under ./results/
make quick # single small sanity run
make sweep # wide pp/tg/depth/concurrency sweep
make cache-test # prefix-caching comparison
Test harness (Makefile.benchtestharness) — runs suites of cases, collects
logs, and prints a colored pass/fail summary:
make -f Makefile.benchtestharness smoke # quick sanity (~30s)
make -f Makefile.benchtestharness regression # standard suite (~5min)
make -f Makefile.benchtestharness full # full sweep (~30min)
make -f Makefile.benchtestharness stress # high-concurrency stress
make -f Makefile.benchtestharness report # pretty-print latest results
make -f Makefile.benchtestharness check-env # validate config before running
python:3.12-slimllama-benchy (installed from PyPI at build time), plus
transformers for tokenizationbash, ca-certificates, curl, wget, jqgpt2 fallback tokenizer and the default
Project Gutenberg corpus book (so runtime is fully offline)1000, GID 0 (root group) — compatible with
OpenShift/OCP arbitrary-UID execution| Variable | Default | Purpose |
|---|---|---|
TOKENIZER_PATH | /tokenizer | Directory the entrypoint reads a mounted tokenizer from |
HF_HUB_OFFLINE | 1 | Force HuggingFace Hub offline |
TRANSFORMERS_OFFLINE | 1 | Force transformers offline |
HF_HOME | /home/benchy/.cache/huggingface | HuggingFace cache location |
HOME | /home/benchy | Home dir (holds the pre-seeded corpus cache) |
TLS verification is disabled image-wide via a sitecustomize.py shim (urllib,
requests, httpx, aiohttp, urllib3) plus GIT_SSL_NO_VERIFY / PIP_TRUSTED_HOST
for build-time installs — intended for environments behind a TLS-intercepting
proxy.
.env keys (for the Makefiles)| Key | Example | Purpose |
|---|---|---|
BASE_URL | http://vllm-server:8000/v1 | OpenAI-compatible endpoint |
API_KEY | EMPTY | API key (passed as --api-key / -e API_KEY) |
MODEL | meta-llama/Llama-3.1-8B-Instruct | Model id (omit to auto-detect) |
TOKENIZER_DIR | /srv/tokenizers/llama3 | Host tokenizer dir, mounted at /tokenizer:ro |
PP TG DEPTH RUNS CONCURRENCY | see .env.example | Benchmark sweep parameters |
FORMAT | md | Output format |
EXTRA_ARGS | --no-cache --skip-coherence | Extra flags appended verbatim |
/tokenizer — declared VOLUME; mount a host directory of tokenizer
files here (read-only recommended). The entrypoint injects
--tokenizer /tokenizer automatically when it's non-empty./workspace/results — mount a host dir here to collect saved results
(used by make bench-save and the harness with --save-result).docker build -t n8500x/llama-benchy:latest .
Behind a proxy, pass build args:
docker build \
--build-arg HTTP_PROXY=http://proxy:8080 \
--build-arg HTTPS_PROXY=http://proxy:8080 \
--build-arg NO_PROXY=localhost,127.0.0.1 \
-t n8500x/llama-benchy:latest .
Via Make (builds a timestamped tag plus latest, and can push):
make build # docker build → n8500x/llama-benchy:<timestamp> + :latest
make push # push most recent local tag + latest to Docker Hub
make release # build + push
To bake in additional tokenizers, add AutoTokenizer.from_pretrained(...) lines
to the Dockerfile's pre-cache step (network is available at build time only).
/tokenizer (or pass
--tokenizer with a pre-cached one such as gpt2); otherwise the entrypoint
fails fast with an error.--network host so the container can reach a
server on the host's localhost. Adjust BASE_URL and networking for your
own topology.1000, GID 0 with group-writable
work dirs, so it runs under OpenShift's arbitrary-UID policy without changes.Content type
Image
Digest
sha256:6fea96c45…
Size
126.6 MB
Last updated
5 months ago
docker pull n8500x/llama-benchy