Sign inSign up

n8500x/llama-benchy

By n8500x

•Updated 5 months ago

Air-gapped LLM inference benchmarking tool (llama-benchy) on python:3.12-slim

Image
0

1.6K

n8500x/llama-benchy repository overview

⁠llama-benchy

An LLM inference benchmarking tool for OpenAI-compatible endpoints (vLLM, SGLang, and similar), packaged for fully offline / air-gapped use.

⁠Overview

llama-benchy drives an OpenAI-compatible /v1 API and measures inference performance across prompt-processing (prefill) and token-generation (decode) workloads. You sweep combinations of:

  • --pp — prompt (prefill) lengths in tokens, e.g. 512 1024 2048
  • --tg — tokens to generate (decode length), e.g. 32 128
  • --depth — pre-filled context depth, e.g. 0 2048 4096
  • --runs — repetitions per case
  • --concurrency — parallel request count
  • prefix-caching behaviour (--enable-prefix-caching / --no-cache)

Results can be emitted in Markdown (default) or saved to a file via --save-result. Prompts are built from a corpus (a Project Gutenberg book, pre-cached in the image) and tokenized with a local tokenizer so prompt lengths are exact.

This image is built for offline operation: HuggingFace/transformers run in offline mode at runtime, the gpt2 fallback tokenizer and the default corpus book are baked in at build time, and no network fetches to HF/PyPI/Gutenberg happen when the container runs.

⁠Quick start

docker pull n8500x/llama-benchy

The entrypoint wraps llama-benchy and auto-injects --tokenizer when a tokenizer is mounted at /tokenizer. A typical run against an OpenAI-compatible server:

docker run --rm --network host \
  -v /srv/tokenizers/llama3:/tokenizer:ro \
  n8500x/llama-benchy \
    --base-url http://localhost:8000/v1 \
    --api-key EMPTY \
    --model meta-llama/Llama-3.1-8B-Instruct \
    --pp 512 1024 2048 --tg 32 128 --depth 0 \
    --runs 3 --concurrency 1 --format md

Show help / version (no tokenizer required):

docker run --rm n8500x/llama-benchy --help
docker run --rm n8500x/llama-benchy --version

If no tokenizer is mounted and none is passed, you can still use a pre-cached tokenizer (e.g. --tokenizer gpt2).

⁠Using the Makefiles

The repo ships two Makefiles that wrap the docker run invocation. Configure once by copying .env.example to .env and setting your endpoint and TOKENIZER_DIR.

Benchmarking (Makefile):

make bench                   # run with current .env config
make bench CONFIG=vllm.env   # use an alternate config file
make bench-save              # run and save results under ./results/
make quick                   # single small sanity run
make sweep                   # wide pp/tg/depth/concurrency sweep
make cache-test              # prefix-caching comparison

Test harness (Makefile.benchtestharness) — runs suites of cases, collects logs, and prints a colored pass/fail summary:

make -f Makefile.benchtestharness smoke       # quick sanity (~30s)
make -f Makefile.benchtestharness regression  # standard suite (~5min)
make -f Makefile.benchtestharness full        # full sweep (~30min)
make -f Makefile.benchtestharness stress      # high-concurrency stress
make -f Makefile.benchtestharness report      # pretty-print latest results
make -f Makefile.benchtestharness check-env   # validate config before running

⁠What's inside

  • Base image: python:3.12-slim
  • Tool: llama-benchy (installed from PyPI at build time), plus transformers for tokenization
  • System packages: bash, ca-certificates, curl, wget, jq
  • Pre-cached at build time: the gpt2 fallback tokenizer and the default Project Gutenberg corpus book (so runtime is fully offline)
  • Runs as non-root UID 1000, GID 0 (root group) — compatible with OpenShift/OCP arbitrary-UID execution

⁠Configuration

⁠Environment variables (set in the image)
VariableDefaultPurpose
TOKENIZER_PATH/tokenizerDirectory the entrypoint reads a mounted tokenizer from
HF_HUB_OFFLINE1Force HuggingFace Hub offline
TRANSFORMERS_OFFLINE1Force transformers offline
HF_HOME/home/benchy/.cache/huggingfaceHuggingFace cache location
HOME/home/benchyHome dir (holds the pre-seeded corpus cache)

TLS verification is disabled image-wide via a sitecustomize.py shim (urllib, requests, httpx, aiohttp, urllib3) plus GIT_SSL_NO_VERIFY / PIP_TRUSTED_HOST for build-time installs — intended for environments behind a TLS-intercepting proxy.

⁠.env keys (for the Makefiles)
KeyExamplePurpose
BASE_URLhttp://vllm-server:8000/v1OpenAI-compatible endpoint
API_KEYEMPTYAPI key (passed as --api-key / -e API_KEY)
MODELmeta-llama/Llama-3.1-8B-InstructModel id (omit to auto-detect)
TOKENIZER_DIR/srv/tokenizers/llama3Host tokenizer dir, mounted at /tokenizer:ro
PP TG DEPTH RUNS CONCURRENCYsee .env.exampleBenchmark sweep parameters
FORMATmdOutput format
EXTRA_ARGS--no-cache --skip-coherenceExtra flags appended verbatim
⁠Volumes
  • /tokenizer — declared VOLUME; mount a host directory of tokenizer files here (read-only recommended). The entrypoint injects --tokenizer /tokenizer automatically when it's non-empty.
  • /workspace/results — mount a host dir here to collect saved results (used by make bench-save and the harness with --save-result).

⁠Build

docker build -t n8500x/llama-benchy:latest .

Behind a proxy, pass build args:

docker build \
  --build-arg HTTP_PROXY=http://proxy:8080 \
  --build-arg HTTPS_PROXY=http://proxy:8080 \
  --build-arg NO_PROXY=localhost,127.0.0.1 \
  -t n8500x/llama-benchy:latest .

Via Make (builds a timestamped tag plus latest, and can push):

make build     # docker build → n8500x/llama-benchy:<timestamp> + :latest
make push      # push most recent local tag + latest to Docker Hub
make release   # build + push

To bake in additional tokenizers, add AutoTokenizer.from_pretrained(...) lines to the Dockerfile's pre-cache step (network is available at build time only).

⁠Notes

  • Offline by design: at runtime the container never fetches from HuggingFace, PyPI, or Gutenberg. You must mount a tokenizer at /tokenizer (or pass --tokenizer with a pre-cached one such as gpt2); otherwise the entrypoint fails fast with an error.
  • TLS verification is disabled throughout the image for use behind TLS-intercepting proxies. Do not use this image where certificate validation is a security requirement.
  • Networking: examples use --network host so the container can reach a server on the host's localhost. Adjust BASE_URL and networking for your own topology.
  • Non-root / OCP: the image runs as UID 1000, GID 0 with group-writable work dirs, so it runs under OpenShift's arbitrary-UID policy without changes.

Tag summary

Content type

Image

Digest

sha256:6fea96c45…

Size

126.6 MB

Last updated

5 months ago

docker pull n8500x/llama-benchy