Sign inSign up

joshxt/ezlocalai

By joshxt

•Updated about 14 hours ago

Image
0

10K+

joshxt/ezlocalai repository overview

⁠ezlocalai

GitHub Dockerhub

ezlocalai is an easy set up artificial intelligence server that allows you to easily run multimodal artificial intelligence from your computer. It is designed to be as easy as possible to get started with running local models. It automatically handles downloading the model of your choice and configuring the server based on your CPU, RAM, and GPU specifications. It also includes OpenAI Style⁠ endpoints for easy integration with other applications using ezlocalai as an OpenAI API proxy with any model. Additional functionality is built in for voice cloning text to speech and a voice to text for easy voice communication as well as image generation and video generation entirely offline after the initial setup.

⁠Prerequisites

Additional Linux Prerequisites

Install the CLI and start ezlocalai with a single command:

pip install ezlocalai
ezlocalai start

It will take several minutes to download the models on the first run. Once running, access the API at http://localhost:8091⁠.

⁠CLI Commands
# Start with defaults (auto-detects GPU, uses Qwen3-VL-4B)
ezlocalai start

# Start with a specific model
ezlocalai start --model unsloth/gemma-3-4b-it-GGUF

# Start with custom options
ezlocalai start --model unsloth/Qwen3-VL-4B-Instruct-GGUF \
                --uri http://localhost:8091 \
                --api-key my-secret-key \
                --ngrok <your-ngrok-token>

# Other commands
ezlocalai stop      # Stop the container
ezlocalai restart   # Restart the container
ezlocalai status    # Check if running and show configuration
ezlocalai logs      # Show container logs (use -f to follow)
ezlocalai update    # Pull/rebuild latest images

# Send prompts directly from the CLI
ezlocalai prompt "Hello, world!"
ezlocalai prompt "What's in this image?" -image ./photo.jpg
ezlocalai prompt "Explain quantum computing" -m unsloth/Qwen3-VL-4B-Instruct-GGUF -temp 0.7
⁠CLI Options
OptionDefaultDescription
--model, -munsloth/Qwen3-VL-4B-Instruct-GGUFHuggingFace GGUF model(s), comma-separated
--urihttp://localhost:8091Server URL
--api-keyNoneAPI key for authentication
--ngrokNonengrok token for public URL
⁠Prompt Command Options
OptionDefaultDescription
-m, --modelAuto-detectedModel to use for the prompt
-temp, --temperatureModel defaultTemperature for response generation (0.0-2.0)
-tp, --top-pModel defaultTop-p (nucleus) sampling parameter (0.0-1.0)
-image, --imageNonePath to local image file or URL to include with prompt
-stats, --statsOffShow statistics (tokens, speed, timing) after response

For additional options (Whisper, image model, etc.), edit ~/.ezlocalai/.env:

⁠Data Persistence

All data is stored in ~/.ezlocalai/:

DirectoryContents
~/.ezlocalai/data/models/Downloaded GGUF model files
~/.ezlocalai/data/hf/HuggingFace cache
~/.ezlocalai/data/voices/Voice cloning samples
~/.ezlocalai/data/outputs/Generated images/audio
~/.ezlocalai/.envYour configuration

Models persist across container updates - you won't re-download them when updating the CLI or rebuilding the CUDA image.

⁠Benchmarks

Performance tested on Intel i9-12900KS + RTX 4090 (24GB):

ModelSizeSpeedNotes
Qwen3-VL-4B4B~210 tok/sVision-capable, great for chat
Qwen3-Coder-30B30B (MoE)~65 tok/sCoding model, hot-swappable

Both models pre-calibrate at startup and hot-swap in ~1 second.

⁠Distributed Fallback / Multi-Machine Setup

ezlocalai supports a distributed fallback system where multiple instances can fall back to each other when local resources (VRAM/RAM) are exhausted, or fall back to any OpenAI-compatible API. This enables:

  • Load balancing: When one machine is busy, requests automatically route to another
  • Redundancy: If one server is overloaded, the fallback handles requests
  • Resource optimization: Each machine handles what it can, forwarding the rest
  • Hybrid deployment: Mix local ezlocalai instances with cloud APIs
⁠Configuration

Set these environment variables in your .env file or pass them to the container:

# Fallback server URL - can be another ezlocalai instance OR any OpenAI-compatible API
FALLBACK_SERVER=http://192.168.1.100:8091  # Another ezlocalai instance
# Or use a cloud provider:
# FALLBACK_SERVER=https://api.openai.com/v1

# Authentication for the fallback server
FALLBACK_API_KEY=your-api-key

# Optional: Override model for OpenAI-compatible fallback (pass-through by default)
# If not set, the originally requested model is passed through to the fallback server
# FALLBACK_MODEL=gpt-4o-mini

# Combined memory threshold (VRAM + RAM) in GB - fallback triggers when below this
# Models can offload to system RAM, so combined memory is more accurate than VRAM alone
FALLBACK_MEMORY_THRESHOLD=8.0

The system automatically detects whether FALLBACK_SERVER points to another ezlocalai instance or an OpenAI-compatible API by checking for the /v1/resources endpoint. If it's another ezlocalai server, full endpoint forwarding is used (preserving the original request). Otherwise, it falls back to standard OpenAI API calls, passing through the originally requested model (or using FALLBACK_MODEL if set as an override).

⁠Example: Two-Machine Setup

Machine A (Primary with RTX 4090):

EZLOCALAI_URL=http://0.0.0.0:8091
EZLOCALAI_API_KEY=shared-key
FALLBACK_SERVER=http://machine-b:8091
FALLBACK_API_KEY=shared-key

Machine B (Fallback with RTX 3080):

EZLOCALAI_URL=http://0.0.0.0:8091
EZLOCALAI_API_KEY=shared-key
FALLBACK_SERVER=http://machine-a:8091
FALLBACK_API_KEY=shared-key

Both machines fall back to each other - creating a resilient two-node cluster.

⁠Example: Local + Cloud Hybrid

Run a local ezlocalai with OpenAI as the fallback:

EZLOCALAI_URL=http://0.0.0.0:8091
FALLBACK_SERVER=https://api.openai.com/v1
FALLBACK_API_KEY=sk-your-openai-key
FALLBACK_MODEL=gpt-4o-mini
⁠Monitoring Fallback Status

Check the fallback status via API:

# Get resource status including fallback info
curl http://localhost:8091/v1/resources

# Check fallback availability and models
curl http://localhost:8091/v1/fallback/status
⁠What Gets Forwarded

When fallback is triggered to another ezlocalai instance, these endpoints are automatically forwarded:

  • /v1/chat/completions - Chat completions (including streaming)
  • /v1/completions - Text completions
  • /v1/embeddings - Text embeddings
  • /v1/audio/transcriptions - Speech-to-text
  • /v1/audio/speech - Text-to-speech
  • /v1/audio/music - Music generation
  • /v1/images/generations - Image generation
  • /v1/images/edits - Image editing (image + text to image)
  • /v1/videos/generations - Video generation

For OpenAI-compatible APIs, only chat completions and embeddings are forwarded.

⁠On-Demand LLM Residency

DEFAULT_MODEL accepts a comma-separated model list. With LLM_MODEL_RESIDENCY=auto, ezlocalai estimates each configured model's GPU footprint at startup. Models remain resident together when they fit; when models assigned to the same GPU exceed its usable VRAM, only the first model is loaded at startup. Requesting another configured model unloads the idle resident model, loads the requested model, and leaves it resident until a different model is requested.

DEFAULT_MODEL=unsloth/Qwen3.8-27B-GGUF,unsloth/Qwen3.6-35B-A3B-MTP-GGUF
LLM_MODEL_RESIDENCY=auto
LLM_MODEL_RESIDENCY_MARGIN_GB=1.5

Use LLM_MODEL_RESIDENCY=resident to force all configured models to load together or LLM_MODEL_RESIDENCY=swap to force one-at-a-time loading. Swap mode serializes local text work so an active model is never unloaded mid-generation. The worker continues to advertise every configured model, but its heartbeat marks all swap-dependent model slots occupied while the resident model is in use.

Model-only load/unload timings are exposed under model_lifecycle in GET /v1/resources. To alternate configured models and print those timings:

python benchmark_model_lifecycle.py \
  --models unsloth/Qwen3.8-27B-GGUF,unsloth/Qwen3.6-35B-A3B-MTP-GGUF \
  --rounds 2

⁠Qwen3.8-27B Performance Tuning

Router failures are retained independently of worker registration in ROUTER_ERROR_FILE (default /data/router-errors.json, on the router's data volume). ROUTER_ERROR_ARCHIVE_MAX bounds the archive (default 2,000 events). The dashboard and /v1/router/errors show the newest 100 archived/live events; the HTML dashboard displays 50, with UTC dates and offline labels. This archive survives pruning, deregistration and router restarts; it cannot recover errors already discarded by an older router. Protect the data volume as error messages may include upstream diagnostics. If a crashed process registers with a new worker ID, recent failures are restored by its stable label and URL so it cannot bypass the circuit-breaker cooldown. A transfer interruption is not proof of an OOM or native crash: correlate its timestamp with the affected worker's container exit status and logs. Tunnel interruptions now preserve the underlying error; LLM failover retries only before assistant output starts, otherwise it emits an explicit stream error rather than silently completing or duplicating output.

Qwen3.8-27B automatically uses MTP, through xllamacpp 2026.9.10809, with one inference slot per instance, three draft tokens and a 0.1 draft probability threshold. DFlash2 remains opt-in with LLM_SPECULATIVE_TYPE=dflash2; only that backend downloads the revision-pinned Inco Q4_K_M draft⁠ (about 1.1 GB of weights, plus its runtime buffers) into the shared HF cache. The target model and sampling settings are unchanged: draft tokens are verified by the target, not accepted unconditionally. Physical prompt batches are hardware- and context-aware: 24 GB cards use an ubatch of 1024 through 200K context and 512 above 200K, while 32 GB cards use 1024. Other MTP model families retain their conservative defaults.

The host-RAM prompt cache is context-aware, but its requested capacity is capped to the smaller of 25% of total RAM or available RAM minus 4 GiB. This prevents a cache that grows during long-context traffic from OOM-killing a memory-constrained worker. The effective and requested sizes are reported in /v1/resources under model_lifecycle.loaded_llm_runtime. Tune the guard with LLM_PROMPT_CACHE_RAM_MARGIN_MIB and LLM_PROMPT_CACHE_MAX_RAM_FRACTION. LLM_PROMPT_CACHE_ALLOW_UNSAFE=true restores uncapped behavior, including for explicit LLM_PROMPT_CACHE_MIB values, and should only be used after verifying host-RAM headroom under long prompts.

The CUDA image includes a pinned native hotfix for DFlash's large-image cache exhaustion (failed to process mtmd chunk). It preserves full-resolution target vision input without enlarging the KV cache. See hotfix details and native build instructions⁠. Verify a deployed CUDA worker with:

docker exec ezlocalai python -c 'from xllamacpp._ezlocalai_hotfix import HOTFIX; print(HOTFIX)'

It should report dflash-pinned-image-positions-v1. The upstream version number alone does not identify the patched build. Test an idle worker's vision stream with python benchmark_vision.py --url http://localhost:8091 --stream (set EZLOCALAI_API_KEY if authentication is enabled).

All values remain operator-overridable:

LLM_SPECULATIVE_TYPE=auto  # auto, dflash2, mtp, none
MTP_SPEC_DRAFT_N_MAX=auto  # Qwen3.8-27B: T4=2; other cards=3
MTP_SPEC_DRAFT_P_MIN=auto  # Qwen3.8-27B: 0.1
KV_CACHE_TYPE=auto        # q4_0; q8_0 remains an explicit precision opt-in
DFLASH_SPEC_DRAFT_N_MAX=auto  # T4: 2; 3090: 3; others: 4; explicit 1..7 overrides
DFLASH_SPEC_DRAFT_P_MIN=0.0
LLM_BATCH_SIZE=auto
LLM_UBATCH_SIZE=auto

GPU-aware defaults for Qwen3.8-27B (explicit settings take precedence):

Worker GPUMTP draft maximumMTP p-minTarget K/V cache
RTX 3090 / 3090 Ti30.1q4_0
RTX 409030.1q4_0
RTX 509030.1q4_0
Tesla / NVIDIA T420.1q4_0
NVIDIA A100 (40/80 GB or MIG)30.1q4_0
NVIDIA H100 (PCIe/SXM or MIG)30.1q4_0

These are starting profiles, not measured optima for every workload. To tune a mixed fleet from one configuration, use card-specific overrides such as MTP_SPEC_DRAFT_N_MAX_3090=3, MTP_SPEC_DRAFT_N_MAX_5090=4, or KV_CACHE_TYPE_5090=q8_0. The same suffixes work for T4, A100, and H100 (for example MTP_SPEC_DRAFT_N_MAX_T4=2, KV_CACHE_TYPE_A100=q8_0, and DFLASH_SPEC_DRAFT_N_MAX_H100=4). Card-specific overrides beat global overrides. KV_CACHE_TYPE=auto selects the table; an existing explicit KV_CACHE_TYPE=q4_0 still keeps Q4 on every card unless overridden per card. Other model families and GPU types retain Q4 by default; Jetson keeps its explicit F16 setting. The memory planner uses the resolved cache precision. The local Q3 / 3090 Ti comparison⁠ found mixed DFlash gains, including regressions on short thinking requests. The MTP three-token baseline is shared across larger cards; T4 starts with two. Higher values require on-card validation, not extrapolation from free VRAM. Batch/ubatch remain hardware- and context-aware as described above; explicit operator values are preserved.

Deployment: remove an explicit LLM_SPECULATIVE_TYPE=dflash2 override or set it to auto/mtp, then rebuild/restart each worker. Explicit MTP_SPEC_DRAFT_* and KV_CACHE_TYPE* values still win; set them to auto (or remove them) to adopt these defaults. The router does not choose the native decoding backend. DFlash's optional starting lengths are 2 on T4, 3 on 3090, and 4 on the other cards.

The Colab notebook probes GPU and host memory inside its isolated server environment. For Qwen3.8-27B Q3_K_XL, its starting settings are:

Idle Colab GPUInstancesContext per instanceAuto batch / physical batchTotal host prompt cache
T4 (16 GB)18,192512 / 128Disabled
A100 (40 GB)1262,1444,096 / 1,024Available RAM / 8, at most 8 GiB
A100 (80 GB)3262,144Memory-dependent / 1,024Available RAM / 8, at most 8 GiB, divided among instances
H100 (80 GB)3262,144Memory-dependent / 1,024Available RAM / 8, at most 8 GiB, divided among instances

With MTP or DFlash enabled, N_PARALLEL=3 loads three independent copies of the model, each with native n_parallel=1 and the full LLM_MAX_TOKENS context. Requests reserve an instance through completion or stream cancellation; additional requests queue until one becomes free. The API exposes one public model name. Updated routers use the advertised independent instances even with ROUTER_BUSY_SLOT_FALLBACK=false. Both the worker and router need this update. Without speculative decoding, N_PARALLEL retains native shared-slot behavior. For speculative decoding, N_PARALLEL=0 still selects one instance.

The A100/H100 notebook default assumes approximately 22 GiB per Qwen3.8-27B Q3_K_XL instance and reserves 4 GiB of headroom: at least 70 GiB free selects three; 48–70 GiB selects two; smaller allocations select one. This gives full 80 GB A100 and H100 runtimes three instances while keeping 40 GB and MIG allocations memory-gated. Other model/quant choices retain one instance. Override with inference_overrides = {"N_PARALLEL": "1"}. The residency planner counts every copy; an overcommitted configuration falls back to swapping with one usable slot. Replicas increase concurrency and share GPU compute and bandwidth; they do not promise faster individual requests.

These Colab profiles are conservative starting points, not benchmarks on those GPUs. Context follows free, visible VRAM: below 20 GiB uses 8,192 tokens, 20–32 GiB uses 230,000, and at least 32 GiB uses 262,144. The 230K setting reflects the reported Qwen3.8-27B Q3_K_XL workload on 24 GB cards; 40 GB cards use the full context. MIG partitions and busy devices therefore do not inherit full-card memory budgets. Batch sizes also respond to free memory; T4 drops to 256 / 64 below 12 GiB free. The 27B model plus its vision projector is a tight fit on T4; CPU offload can still be necessary and depends on available system RAM. Keep other services disabled. The notebook prints its selection; inference_overrides can set explicit context, batch, cache, or speculative-decoding values. Context and host-cache limits here apply only to the notebook; existing server context/cache settings are unchanged.

For this 27B model, target KV at 262,144 tokens is approximately 4.5 GiB with Q4 versus 8.5 GiB with Q8 (excluding recurrent state, weights, draft and compute buffers). Q8 therefore needs about 4 GiB extra. It is a precision upgrade, not a guaranteed speed improvement. Keep Q4 on 24 GB cards at long context. A smaller main-model quant also does not guarantee faster prefill or decode: kernel choice, acceptance, cache traffic and prompt content matter.

An isolated sweep on an idle worker can be run with:

python benchmark_speculative.py --context 220000 --tokens 256 \
  --draft-lengths 3,4,5,7 --prompt-chars 0,120000,360000 --repeats 2 \
  --kv-cache q4_0

The sweep includes no-speculation and MTP controls, real code prefixes, actual prompt-token counts, prefill/decode timing, and greedy output hashes. Use the same immutable --prompt-file for all runs/cards. Allocated context alone is not a long-context benchmark: compare actual prompt lengths. Do not run this alongside the resident server on the same GPU. Greedy hash checks are a smoke test, not a proof of quality or of production-sampling throughput. Use --sampling-profile thinking or --sampling-profile instruct for the actual ezlocalai sampling settings; non-greedy runs do not compare output hashes. For a focused follow-up, --backends mtp,dflash2 skips the no-speculation control.

Single-GPU NVIDIA workers can also A/B test llama.cpp's experimental concurrent CUDA-stream optimization with GGML_CUDA_GRAPH_OPT=1. It primarily targets token-generation throughput rather than prompt processing. Results vary by GPU and model, so leave it disabled unless a representative decode benchmark shows a repeatable improvement.

Existing MTP_SPEC_DRAFT_* settings apply only with the MTP backend; they do not tune DFlash2. Set LLM_SPECULATIVE_TYPE=dflash2 for an A/B comparison or none to disable speculation. Other model families never receive the 27B draft. DFLASH_MODEL_FILE selects another quant from the same pinned repository; DFLASH_MODEL_PATH can point to an already downloaded compatible draft.

Benchmark representative prompts on each card: more speculation is not always faster, and a draft that forces target layers onto CPU can erase the benefit. Explicit LLM_UBATCH_SIZE values are attempted first; the resilient loader retries smaller physical batches if model initialization runs out of GPU memory.

⁠Qwen TTS with llama.cpp

Local Qwen TTS now uses a persistent, private stdio worker built against the same llama.cpp revision as xllamacpp. It uses upstream libmtmd for the speaker encoder, talker, code predictor and vocoder. Standard images no longer install qwen-tts or its dedicated FlashAttention wheel. Torch/Transformers remain dependencies of other media models; they are not used for TTS synthesis.

The default retains the Qwen3-TTS-12Hz-0.6B-Base model, using the llama.cpp-compatible Q8_0 talker and F16 codec conversion⁠. The old QWEN_TTS_MODEL=Qwen/Qwen3-TTS-12Hz-0.6B-Base value is migrated automatically. This is a backend/quantization change, not a claim of identical audio or verified perceptual-quality parity with the old implementation.

Voice .wav references, WAV responses, chunked PCM streaming, audio caching, and the existing voice/LLM slot-sharing policy are preserved. The native worker stays loaded between chunks and close() waits for process exit before the LLM can reclaim the slot. Cloning is audio-only (x-vector): .txt transcripts are preserved on disk but no longer condition generation. auto language uses script detection for Russian, Chinese, Japanese and Korean, otherwise English; specify the language explicitly for other supported languages (German, Italian, Portuguese, Spanish and French). Mixed-language synthesis should be checked on your actual voice samples.

QWEN_TTS_MODEL=Mouserat/qwen3-tts-0.6b-base-gguf
QWEN_TTS_CONTEXT_SIZE=4096
QWEN_TTS_THREADS=20
QWEN_TTS_TIMEOUT=300
QWEN_TTS_MAX_NEW_TOKENS=320  # audio frames, not text tokens

QWEN_TTS_MODEL_FILE and QWEN_TTS_MMPROJ_FILE accept filenames in a custom GGUF repository or local file paths; both files must be compatible with upstream libmtmd. QWEN_TTS_REVISION pins a custom repository revision. The old dtype, attention, transcript and non-streaming settings do not control the native model. QWEN_TTS_GENERATE_KWARGS supports max_new_tokens, temperature, top_k, top_p and repetition_penalty; unsupported options fail explicitly.

CUDA images build kernels for 3090/4090/5090 (SM 86/89/120). Native installs need CMake, a C++ compiler, and the CUDA toolkit for GPU synthesis:

python scripts/build_tts.py --cuda --build-dir /absolute/path/tts-build
export QWEN_TTS_BIN=/absolute/path/tts-build/bin/ezlocalai-tts

Omit --cuda for a CPU build. Model download happens during precache without loading a GPU model. Changing the backend uses a new audio-cache namespace.

⁠Embeddings

ezlocalai serves /v1/embeddings with a dedicated GGUF embedding model, independent of TEXT_SERVER. By default it uses Qwen3-Embedding-0.6B Q8_0 with a 32k context:

EMBEDDING_ENABLED=true
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B-GGUF
EMBEDDING_MODEL_ALIAS=Qwen3-Embedding-0.6B
EMBEDDING_QUANT_TYPE=Q8_0
EMBEDDING_CONTEXT_LENGTH=10000
EMBEDDING_N_PARALLEL=1
EMBEDDING_GPU_LAYERS=auto
EMBEDDING_KV_CACHE_TYPE=f16

Embedding context defaults to 10,000 tokens per instance. Existing explicit EMBEDDING_CONTEXT_LENGTH=8192 settings must be changed to adopt the extra headroom. Native allocation can round this up for alignment. Oversized inputs return an actionable HTTP 400 rather than a retryable 500; split documents into smaller chunks instead of retrying unchanged input. Inputs are never silently truncated. More context increases memory use per embedding instance, so check headroom when running multiple embedders alongside the LLM.

When using the router, workers advertise the embedding capability only when EMBEDDING_ENABLED=true, so embedding requests route to workers that can serve them and the dashboard shows the active embedding model. EMBEDDING_N_PARALLEL controls how many full-context embedding model instances are loaded and reported to the router; each instance keeps the full EMBEDDING_CONTEXT_LENGTH per request instead of splitting the context across internal xllamacpp slots. With EMBEDDING_GPU_LAYERS=auto, ezlocalai estimates the 32k embedding cache footprint for each instance and partially offloads layers to CPU when VRAM is tight.

⁠Image And Video

Image and video workers are opt-in so setting IMG_MODEL or VIDEO_MODEL alone does not warm-load or advertise those capabilities:

IMAGE_ENABLED=false
IMG_MODEL=
VIDEO_ENABLED=false
VIDEO_MODEL=unsloth/LTX-2.3-GG

Tag summary

Content type

Image

Digest

sha256:0441e2afc…

Size

1.7 GB

Last updated

about 15 hours ago

docker pull joshxt/ezlocalai