Efficient LLM inference with .oom format - 2x smaller than GGUF
2.9K
Efficient LLM inference with .oom format - 2x smaller than GGUF
OomLlama is a Rust-powered LLM inference engine that uses the .oom (OomLlama Model) format. It achieves 2x smaller model sizes than GGUF Q4 through Q2 quantization with lazy layer loading.
| Model | GGUF (Q4) | OOM (Q2) |
|---|---|---|
| 70B | ~40 GB | ~20 GB |
| 32B | ~20 GB | ~10 GB |
| 7B | ~4 GB | ~2.5 GB |
# CLI usage
docker run jtmeent/oomllama generate "What is the meaning of life?"
# API server
docker run -p 8000:8000 jtmeent/oomllama:api
# With model volume
docker run -v ~/.cache/oomllama:/models jtmeent/oomllama list
from oomllama import OomLlama
llm = OomLlama("humotica-32b")
response = llm.generate("Hello!")
print(response)
latest, 0.6.0 - CLI toolapi - REST API server (port 8000)| Variable | Default | Description |
|---|---|---|
MODEL_NAME | humotica-32b | Model to load |
MODEL_PATH | auto | Custom model path |
GPU_ID | none | CUDA GPU ID |
PORT | 8000 | API port (api tag only) |
When using the api tag:
POST /generate - Generate textPOST /chat - Chat completionGET /models - List modelsGET /info - Model infoGET /health - Health checkGET /docs - Swagger UI+--------------------------------------+
| Header: OOML (magic) + metadata |
+--------------------------------------+
| Tensors: Q2 quantized (2 bits/weight)|
| - Scale + Min per 256-weight block |
| - 68 bytes per block |
+--------------------------------------+
For GPU inference with bundled CUDA, download the wheel directly:
pip install https://brein.jaspervandemeent.nl/downloads/oomllama-0.6.0-cuda.whl
One Love, One fAmIly
Built by Humotica AI Lab - Jasper, Claude, Gemini
Content type
Image
Digest
sha256:d974e2c83…
Size
36.5 MB
Last updated
6 months ago
docker pull humotica/oomllama