Sign inSign up

humotica/oomllama

By humotica

•Updated 6 months ago

Efficient LLM inference with .oom format - 2x smaller than GGUF

Image
Machine learning & AI
Developer tools
0

2.9K

humotica/oomllama repository overview

⁠OomLlama

Efficient LLM inference with .oom format - 2x smaller than GGUF

PyPI Docker

⁠What is OomLlama?

OomLlama is a Rust-powered LLM inference engine that uses the .oom (OomLlama Model) format. It achieves 2x smaller model sizes than GGUF Q4 through Q2 quantization with lazy layer loading.

ModelGGUF (Q4)OOM (Q2)
70B~40 GB~20 GB
32B~20 GB~10 GB
7B~4 GB~2.5 GB

⁠Quick Start

# CLI usage
docker run jtmeent/oomllama generate "What is the meaning of life?"

# API server
docker run -p 8000:8000 jtmeent/oomllama:api

# With model volume
docker run -v ~/.cache/oomllama:/models jtmeent/oomllama list

⁠Python Usage

from oomllama import OomLlama

llm = OomLlama("humotica-32b")
response = llm.generate("Hello!")
print(response)

⁠Tags

  • latest, 0.6.0 - CLI tool
  • api - REST API server (port 8000)

⁠Environment Variables

VariableDefaultDescription
MODEL_NAMEhumotica-32bModel to load
MODEL_PATHautoCustom model path
GPU_IDnoneCUDA GPU ID
PORT8000API port (api tag only)

⁠API Endpoints

When using the api tag:

  • POST /generate - Generate text
  • POST /chat - Chat completion
  • GET /models - List models
  • GET /info - Model info
  • GET /health - Health check
  • GET /docs - Swagger UI

⁠The .oom Format

+--------------------------------------+
| Header: OOML (magic) + metadata      |
+--------------------------------------+
| Tensors: Q2 quantized (2 bits/weight)|
| - Scale + Min per 256-weight block   |
| - 68 bytes per block                 |
+--------------------------------------+
⁠Key Features
  • Q2 Quantization: 2-bit weights with per-block scale/min
  • Lazy Layer Loading: Only active layer in memory
  • Interleaved RoPE: Proper Qwen model support (no gibberish!)
  • CUDA Support: GPU inference via Candle

⁠CUDA Version

For GPU inference with bundled CUDA, download the wheel directly:

pip install https://brein.jaspervandemeent.nl/downloads/oomllama-0.6.0-cuda.whl

⁠Credits

  • Model Format: Gemini IDD & Root AI (Humotica AI Lab)
  • Quantization: OomLlama.rs by Humotica
  • Interleaved RoPE Fix: Root AI & Jasper
  • Base Models: Meta (Llama), Alibaba (Qwen)

One Love, One fAmIly

Built by Humotica AI Lab - Jasper, Claude, Gemini

Tag summary

Content type

Image

Digest

sha256:d974e2c83…

Size

36.5 MB

Last updated

6 months ago

docker pull humotica/oomllama