High-performance LLM inference server — 2x concurrency over vanilla SGLang. By iterate.ai.
10K+
High-performance LLM inference server - 2x the concurrent capacity on the same GPU, with confidential computing built in.
Built by iterate.ai, Dream big, build fast.
Lifeboat is an LLM inference server built for the concurrency and long-running context that agentic workloads demand. It runs on infrastructure you already own - roughly 2x the concurrent capacity on the same GPU, no per-token cloud costs, and data stays inside your perimeter.
The gain comes from adaptive KV-cache management, fairness-aware scheduling and memory tiering - not from quantizing model weights. Weights stay BF16; only KV-cache values use FP8 E4M3.
It ships as a single container with a management UI, an OpenAI-compatible API
(plus Anthropic /v1/messages), a token-aware load balancer, multi-node cluster
management, and optional GPU confidential computing for protecting data in use.
It also runs where there is no GPU at all - the Lite image (~720 MB against 17 GB) serves GGUF models on CPU or Vulkan, and a native desktop app runs it with no Docker at all. Same console, same API, same load balancer.
| Baseline | Lifeboat + FP8 KV | ||
|---|---|---|---|
| KV cache capacity | 284 K tokens | 568 K tokens | 2.0x |
| Max concurrency (100 % success) | 1,024 | 2,048 | 2.0x |
| Throughput @ 1024 concurrent | 6,304 tok/s | 9,611 tok/s | 1.52x |
| Throughput @ 2048 concurrent | 4,965 tok/s | 8,714 tok/s | 1.76x |
| Success @ 2048 concurrent | 62 % | 100 % | baseline fails |
Baseline is a stock build of the upstream tensor engine, not Lifeboat with its optimizations off. Qwen3-30B-A3B (MoE, ~57 GB BF16) on one RTX PRO 6000 Blackwell (96 GB), 4,096-token prompts / 512-token completions, prefix caching disabled both sides. KV pool size dominates and is unrecorded here: given headroom the two converge. Harness and the raw result files these figures are read out of.
pip install lifeboat
lifeboat engine install
lifeboat up # console on http://127.0.0.1:8001
The quickest way to evaluate Lifeboat, and how to run it on small hardware:
console, model registry, load balancer and the full OpenAI-compatible API, with
no container and no accelerator. lifeboat doctor reports the board's cores,
memory, GPU and the largest model it can serve before you download anything. Python 3.10-3.13
on macOS (Apple Silicon), Linux x86-64/ARM64 and Windows x64 -
PyPI,
quickstart.
Edge and embedded hardware is a first-class target. The Linux engines link
glibc 2.27 and compile every CPU variant in, selected at runtime - so one
artifact runs from a 2018 distribution to a current one, and from an Atom to a
Zen 5. Nothing is -march=native.
| Board | Works on | Comfortable model |
|---|---|---|
| Raspberry Pi 4 / 5 (64-bit) | armv8.0 CPU backend | 1.5-4B at 4-bit |
| NVIDIA Jetson Orin (JetPack 5/6) | CUDA, selected automatically | up to 4B at 4-bit |
| Mini-PC, NUC, thin client, embedded x86 (Atom, N100) | x64 baseline, no AVX; Vulkan iGPU | 1.5-8B at 4-bit |
Measured against other servers, same hardware and same model file, out of the box: 1.94x llama.cpp on a 64-core server, 2.01x ollama on a Jetson, 1.48x on a Mac. Harness and raw results.
It carries the GGUF engine only: the 2x concurrency, FP8 KV cache, speculative decoding and safetensors weights stay with the container, as do Intel Macs (no wheel).
A native app with a tray icon, for running models on a laptop or workstation. It bundles the GGUF engine and the same control plane, console and API as the container. GPU offload uses Metal on Apple Silicon and Vulkan on Windows and Linux, which covers NVIDIA, AMD and Intel with one download.
The macOS and Windows builds are code-signed (macOS also notarized); the Linux builds are not, and ship a checksum file instead.
| Platform | GPU offload |
|---|---|
| macOS Apple Silicon (13+) | Metal, plus MLX |
| macOS Intel (13+) | Metal; GGUF only, no MLX |
Linux x64 / arm64 - .deb or tarball | Vulkan |
| Windows x64 - signed installer | Vulkan; needs the free .NET 8 Desktop Runtime |
Exact filenames and a SHA256SUMS for every build are on the releases page.
Linux builds need glibc 2.31+ (Debian 11+, Ubuntu 20.04+, RHEL 9+). They are built in an old-glibc container and verified on a clean Debian 11 before release, not on the machine that made them.
macOS - open the .dmg and drag Lifeboat to Applications. Signed with
a Developer ID, notarized by Apple and stapled, so there is no security
warning and it works offline.
Windows - run the setup.exe; per-user, so no administrator prompt.
Linux - sudo apt install ./lifeboat-desktop_2.2.52_amd64.deb, or unpack
the tarball with sudo tar -C / -xzf Lifeboat-2.2.52-linux-x86_64.tar.gz.
Use apt install ./file.deb rather than dpkg -i: the tray binding and the
Vulkan loader are recommended packages and dpkg skips them. You get
/opt/lifeboat plus lifeboat-core (server and CLI, runs headless) and
lifeboat-tray on PATH.
Per-platform requirements, GPU support, upgrading and uninstalling: docs/desktop.md.
macOS and Windows artifacts are code-signed. Linux has no equivalent - verify
with sha256sum -c SHA256SUMS-linux.txt --ignore-missing.
The tray menu has Open Console, Show Log and Quit - no Dock icon, by design. The console at http://127.0.0.1:30800 asks you to create the administrator account on first visit.
It refuses a safetensors download on macOS or Windows before it starts rather than after an hour - that format needs the tensor engine (Linux, CUDA or ROCm only). See Which model format.
Lifeboat ships three images. Pick by what the host has - see Tags:
| Host | Tag | Size |
|---|---|---|
| NVIDIA GPU | iterateai/lifeboat:latest | ~17 GB |
| AMD Instinct GPU | iterateai/lifeboat:latest-rocm | ~28 GB |
| No GPU, Intel/AMD integrated, or AWS Graviton | iterateai/lifeboat:lite | ~720 MB |
# OPTIONAL. Without it, the console asks you to create the administrator
# account on first visit - Lifeboat ships with no credentials at all.
sudo mkdir -p /etc/lifeboat
sudo install -m 600 /dev/stdin /etc/lifeboat/admin-pass <<< 'CHANGE-ME'
docker run --device nvidia.com/gpu=all --name lifeboat -p 8001:8001 \
-v lifeboat-data:/opt/lifeboat/.lifeboat \
-v /path/to/models:/models \
-v /etc/lifeboat/admin-pass:/run/secrets/admin-pass:ro \
-e LIFEBOAT_ADMIN_PASSWORD_FILE=/run/secrets/admin-pass \
-e TMPDIR=/opt/lifeboat/.lifeboat/tmp \
--security-opt no-new-privileges \
--cap-drop ALL \
--cap-add NET_BIND_SERVICE --cap-add SETUID --cap-add SETGID --cap-add CHOWN \
--read-only --tmpfs /tmp:size=2G --tmpfs /run:size=10M \
-dt --restart=always \
iterateai/lifeboat:latest
--device nvidia.com/gpu=allis preferred over--gpus all. It bakes the device list into the OCI spec via CDI, so GPU access survives cgroup re-application - the cause of theFailed to initialize NVML: Unknown Errorthat otherwise appears ~24 h after start. Rundocker run --rm --privileged --pid=host iterateai/lifeboat:latest host-setuponce per GPU node; it is idempotent, so re-run it after a driver upgrade.
There is no container toolkit and no --gpus flag on ROCm - the GPU is
passed through as two device nodes, /dev/kfd (the compute node) and /dev/dri
(the render nodes):
docker run --name lifeboat -p 8001:8001 \
-v lifeboat-data:/opt/lifeboat/.lifeboat \
-v /path/to/models:/models \
-v /etc/lifeboat/admin-pass:/run/secrets/admin-pass:ro \
-e LIFEBOAT_ADMIN_PASSWORD_FILE=/run/secrets/admin-pass \
-e TMPDIR=/opt/lifeboat/.lifeboat/tmp \
--cap-drop ALL \
--cap-add NET_BIND_SERVICE --cap-add SETUID --cap-add SETGID --cap-add CHOWN \
--read-only --tmpfs /tmp:size=2G --tmpfs /run:size=10M \
--device=/dev/kfd --device=/dev/dri \
--ipc=host \
-dt --restart=always \
iterateai/lifeboat:latest-rocm
Confirm the GPU is visible from inside the container:
docker exec lifeboat rocm-smi
SETUID/SETGID/CHOWNare required on both vendors. The entrypoint starts as root, fixes volume ownership, thensetprivs down to the unprivilegedlifeboatuser - without them the drop fails at PID 1. Everything else is still dropped. Or pass--user lifeboat:lifeboatand omit all three: the entrypoint skips the drop, which is what the Compose files do.
--ipc=hostis required on ROCm, not optional hardening. Without it the engine aborts at its first GPU allocation with an HSAMemory in useerror naming neither IPC nor shared memory. Some hosts also need--security-opt seccomp=unconfined. If the device nodes are not readable by the unprivileged user, add--group-add video --group-add render.
No accelerator runtime, so no toolkit, no device mapping and no driver floor. Serves GGUF on the CPU, or on a Vulkan-capable GPU (Intel Arc/iGPU, AMD Radeon, older NVIDIA) if you pass the render node through:
docker run --name lifeboat -p 8001:8001 \
-v lifeboat-data:/opt/lifeboat/.lifeboat \
-v /path/to/models:/models \
-e TMPDIR=/opt/lifeboat/.lifeboat/tmp \
--cap-drop ALL --cap-add NET_BIND_SERVICE \
--cap-add SETUID --cap-add SETGID --cap-add CHOWN \
--read-only --tmpfs /tmp:size=2G --tmpfs /run:size=10M \
--device=/dev/dri -dt --restart=always \
iterateai/lifeboat:lite
Drop --device=/dev/dri on a host with no GPU: Docker refuses to start a
container whose devices: entry does not exist.
What Lite does not have. The optimization suite above hooks the tensor engine's scheduler, so none of it applies - no FP8 KV cache, no fairness-aware scheduling, no elastic memory, no speculative decoding. You keep the whole control plane: load balancer and all four routing modes, capacity gate and queue, sticky sessions, the full OpenAI and Anthropic surface, model registry, licensing and console. Use the engine's own knobs (
--cache-type-k q8_0,--parallel).
Sizing. Decode speed is bounded by memory bandwidth, not core count, so no
setting changes it. On 8 GB, 1.5B-4B at 4-bit is comfortable and 14B will
not fit. GET /api/hardware/profile sizes the actual host, reading a container
memory limit from the cgroup rather than /proc.
Then open http://localhost:8001 and add a model from the dashboard.
http://localhost:8001http://localhost:8001/v1/chat/completionshttp://localhost:8001/v1/messageshttp://localhost:8001/api/versionhttp://localhost:8001/metricsLifeboat ships with no credentials. The first visit asks you to create the administrator account, which becomes the superadmin - there is no default password to look up or change. Setting
LIFEBOAT_ADMIN_PASSWORD(or..._FILE) before the first start creates it from that instead, for unattended installs.
/tmpmust be at least ~2 GB on the GPU images. Kernels JIT-compile on first use into$TMPDIR; a small tmpfs fails as a misleading compiler error or a bare segfault. PointingTMPDIRat the data volume also persists that cache, so the first-launch compile is paid once.
One compose file per target. The AMD file sets the device nodes, groups and
ipc: host; the NVIDIA file sets up CDI; the Lite file defaults GPU
passthrough to a harmless self-mapping:
# NVIDIA
curl -fsSLO https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/docker-compose.yaml
# AMD / ROCm
curl -fsSL -o docker-compose.yaml \
https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/docker-compose.rocm.yaml
# CPU / Vulkan (Lite)
curl -fsSL -o docker-compose.yaml \
https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/docker-compose.cpu.yaml
docker compose up -d
docker compose logs -f lifeboat
Or let the installer choose - it preflights the host, detects the accelerator
(Lite when there is no GPU), downloads the matching compose file and writes
.env:
curl -fsSL https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/get-lifeboat.sh | bash
# preflight only, change nothing
curl -fsSL https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/get-lifeboat.sh | bash -s -- --check
NVIDIA - :latest
| Architectures | linux/amd64, linux/arm64 (multi-arch manifest - resolves automatically) |
| GPUs | Ampere, Ada, Hopper, Blackwell (sm_80, 86, 89, 90, 120, 121) |
| Verified on | RTX 3090 Ti, RTX PRO 6000 Blackwell, GB10 / DGX Spark, H100 |
| Driver | >= 580.65.06 - the image ships the CUDA 13 runtime (torch cu130), NVIDIA's documented minimum for it. Older drivers need a cu128 build; contact [email protected] |
| Container runtime | Docker >= 25 with the NVIDIA Container Toolkit (CDI enabled) |
| Compose file | docker-compose.yaml |
AMD - :latest-rocm
| Architectures | linux/amd64 (ROCm publishes no arm64 build) |
| GPUs | Instinct MI210 / MI250 (CDNA2), MI300X / MI325X (CDNA3), MI350X / MI355X (CDNA4) |
| Verified on | Instinct MI210 (64 GB, ROCm 7.2) |
| Driver | ROCm >= 6.3 with the amdgpu kernel module loaded |
| Container runtime | Docker >= 25 - no container toolkit needed; the GPU is passed as /dev/kfd + /dev/dri |
| Compose file | docker-compose.rocm.yaml (requires ipc: host; the supplied file sets it) |
CPU / Vulkan - :lite
| Architectures | linux/amd64, linux/arm64 (multi-arch manifest) |
| Accelerator | none required. Vulkan offload on Intel Arc / iGPU, AMD Radeon, older NVIDIA - pass /dev/dri |
| Verified on | x86-64 and arm64 hosts, CPU-only and with Vulkan |
| Driver | none. CPU variants are selected at runtime, so it runs on any x86-64 back to SSE - nothing is -march=native |
| Container runtime | Docker >= 25. No toolkit |
| Compose file | docker-compose.cpu.yaml |
| Engine | GGUF only. The optimization suite does not apply - see the Lite note above |
AWS Graviton uses
:lite, not:latest.:latest's arm64 half pulls and runs, but is built for NVIDIA arm64 parts (Grace-Hopper, DGX Spark) - on Graviton that is 17 GB of unusable CUDA.
Host, any image
| OS | Ubuntu 22.04+, RHEL 9+, Rocky 9+, Amazon Linux 2023 |
| VRAM | >= 8 GB per model on the GPU images. Lite uses system RAM instead |
| Disk | >= 50 GB free for a GPU image, ~3 GB for Lite, plus model weights |
| Port | 8001/tcp - dashboard and API |
:latestis the NVIDIA image and will not run on an AMD GPU. A manifest selects on CPU architecture only, so one tag cannot serve both vendors and no architecture check catches the mistake - both arelinux/amd64. AMD uses:latest-rocm; no GPU uses:lite.
Some capabilities are generation-dependent and configured automatically: FP8 needs CDNA3+ on AMD, FP4 needs CDNA4 or Blackwell, and GPU confidential computing is NVIDIA-only (auto-disabled on AMD, with the reason reported).
The GPU images ship two engines - a high-throughput tensor engine and a GGUF engine - and Lifeboat picks per model, including GGUF vision (mmproj) and quantized architectures the tensor engine cannot load. 168+ architectures, plus a Transformers fallback. Lite, the desktop app and pip carry GGUF only.
| Where you are running | Use |
|---|---|
Linux + NVIDIA or AMD Instinct, via :latest / :latest-rocm | safetensors (or GGUF) |
:lite, pip, or any Windows / macOS desktop | GGUF |
| Apple Silicon desktop | GGUF (default) or MLX - both run on the GPU |
Safetensors needs the tensor engine - Linux with CUDA or ROCm only. The other shapes say so before a download starts, not after an hour of bandwidth.
There are three separate free routes, and they are not the same thing:
After the grace period, starting an inference server requires a license - the dashboard, model management and audit log keep working, and the activation dialog is one click from wherever you hit the limit.
# Activate non-interactively during install
curl -fsSL https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/get-lifeboat.sh \
| bash -s -- --license-key LB-XXXX-XXXX-XXXX-XXXX
Air-gapped sites install first, read the Cluster ID off the License page,
and request a signed offline .lic bound to that ID, which they upload through
the activation dialog. One license binds to one cluster.
All licensing - free tier, trial, purchase, upgrades and offline files - goes through iterate.ai/lifeboat. That is the only address you need.
The license is a runtime gate, not a pull gate - activation happens in the
running container, never at docker pull. No credentials are needed to pull
this image.
The chart is published as a rolling tarball, so no repository checkout is
needed. It deploys the rolling image, so re-running helm upgrade picks up the
newest build.
helm install lifeboat \
https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/helm/lifeboat-latest.tgz \
--namespace lifeboat --create-namespace \
--set admin.password='CHANGE-ME' \
--set gpu.count=1 # add --set gpu.vendor=amd|lite for AMD / CPU
Or add it as a Helm repo so helm upgrade tracks new chart versions:
helm repo add lifeboat https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/helm
helm repo update
helm install lifeboat lifeboat/lifeboat \
--namespace lifeboat --create-namespace \
--set admin.password='CHANGE-ME'
helm show values <chart> lists everything. admin.password is required
(or admin.existingSecret) - the chart refuses to render without it. There is
deliberately no license flag: activation is a runtime API call, so activate
in the dashboard once the pod is up.
gpu.vendor takes nvidia (default), amd or lite, and switches
the image, the requested device resource, hostIPC, the seccomp profile and the
pod's supplemental groups together - plus it gates OFF the NVIDIA-only runtime
class and CDI annotation, which would otherwise make an AMD or CPU-only pod fail
admission for a GPU it never asked for.
NVIDIA clusters require nvidia-device-plugin >= 0.15 with
deviceListStrategy=cdi-annotations, or the NVIDIA GPU Operator with CDI
enabled. AMD clusters require the AMD GPU device plugin (or the AMD GPU
Operator). lite requires neither.
The pod runs as UID 999 (
securityContext.runAsUser) with all capabilities dropped. Do not remove that setting - the image has noUSERdirective, and root withoutCAP_DAC_OVERRIDEcannot traverse/opt/lifeboat.
Newest container image: 2.2.53. The two
ship on their own cadences, so the numbers are not always the same.
| Tag | Target | Contents |
|---|---|---|
latest | NVIDIA | Multi-arch manifest - resolves to amd64 or arm64 automatically |
X.Y.Z | NVIDIA | Pinned release, multi-arch (e.g. 2.2.53) |
X.Y.Z-amd64 / X.Y.Z-arm64 | NVIDIA | Architecture-specific, for pinning a single platform |
latest-rocm | AMD | Rolling AMD/ROCm image (amd64). Tensor-engine kernels built for MI210 / MI250 (CDNA2, gfx90a) |
amd | AMD | Same image as latest-rocm, pinned name |
X.Y.Z-rocm | AMD | Pinned AMD release (e.g. 2.2.46-rocm) |
amd-mi300x | AMD | Tensor-engine kernels built for MI300X / MI325X (CDNA3, gfx942) |
amd-mi355x | AMD | Tensor-engine kernels built for MI350X / MI355X (CDNA4, gfx950) |
lite | CPU / Vulkan | Rolling Lite image, multi-arch (amd64 + arm64). ~720 MB |
latest-lite | CPU / Vulkan | Same image as lite |
X.Y.Z-lite | CPU / Vulkan | Pinned Lite release (e.g. 2.2.46-lite) |
X.Y.Z-lite-amd64 / -arm64 | CPU / Vulkan | Architecture-specific Lite |
Version tags are immutable; latest, latest-rocm and lite roll forward.
The AMD images are tagged separately on purpose: tensor-engine kernels are compiled per CDNA generation (the FP8 number format differs between CDNA3 and CDNA4), so one image targets one family - pick the tag that names your card. The GGUF engine covers all three, so only the tensor-engine path is generation-specific, and the container warns at startup if it is running on a GPU it was not built for.
Stated plainly, because a customer should not have to discover it with
tcpdump.
A licensed install sends a heartbeat every 6 hours: the license key, the cluster ID, a pod count, the version, and a short hardware description (CPU model, GPU model, memory) sent once - after that the heartbeat carries only a hash of it, so re-sending is unnecessary. Steady state is about 160 bytes. That heartbeat is also how a license binding is verified, which is a contractual obligation rather than telemetry.
An unlicensed install sends a smaller message once a day, so an evaluation is visible to support. It carries no key.
No model names, no prompts, no completions, no token counts, no request metadata. Enforced in code: the snapshot builder omits those fields, and a test walks every key of the real payload against a forbidden list, so adding one breaks the build.
LIFEBOAT_TELEMETRY=off disables the hardware and census reporting.Lifeboat includes third-party open-source software; the applicable license texts and attribution notices ship with the product. Interplay®, Generate® and AgentOne™ are trademarks of iterate.ai, protected by issued U.S. patents and others pending.
Content type
Image
Digest
sha256:3f410eb5c…
Size
7.8 GB
Last updated
11 days ago
docker pull iterateai/lifeboat