Author & Maintainer: Massimo Lauri
Hugging Face: massimolauri/NomadCoder-4B-GGUF
Docker Hub: massimolauri/nomadcoder-4b
License: Apache-2.0
NomadCoder-4B is a high-efficiency 4-billion parameter coding LLM optimized for 100% CPU execution with native support for up to 128k context window (131072 tokens).
It features Massimo Lauri's Engram Weight Folding, integrating static N-gram associative memory patterns inspired by DeepSeek Engram (arXiv:2601.07372) directly into transformer layers 1 and 15. This container runs natively via llama.cpp server with full OpenAI-compatible API support on port 10200.
Run the server immediately with a single Docker command:
docker run -d \
-p 10200:10200 \
--name nomadcoder \
massimolauri/nomadcoder-4b:latest
The server is ready in seconds and listens on http://localhost:10200.
Use NomadCoder-4B with standard OpenAI SDKs, curl, or coding assistants (Cursor, Continue.dev, OpenWebUI, LibreChat):
curl http://localhost:10200/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "NomadCoder-4B",
"messages": [
{"role": "system", "content": "You are NomadCoder, an expert programming assistant."},
{"role": "user", "content": "Write a thread-safe LRU cache in Rust with unit tests."}
],
"temperature": 0.2,
"max_tokens": 1024
}'
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:10200/v1",
api_key="none"
)
response = client.chat.completions.create(
model="NomadCoder-4B",
messages=[
{"role": "user", "content": "Explain how N-gram associative memory improves code generation."}
]
)
print(response.choices[0].message.content)
| Variable | Default | Description |
|---|---|---|
PORT | 10200 | HTTP port for the server |
HOST | 0.0.0.0 | Network bind address |
THREADS | 16 | CPU threads allocated (adjust to physical CPU cores) |
CTX_SIZE | 131072 | Context window size (128k native support) |
N_GPU_LAYERS | 0 | GPU offload layers (0 = pure CPU execution) |
MODEL_ALIAS | NomadCoder-4B | Model name exposed to the OpenAI API |
docker run -d \
-p 10200:10200 \
-e THREADS=8 \
-e CTX_SIZE=65536 \
massimolauri/nomadcoder-4b:latest
On modern CPUs, token generation speed is bounded by memory bandwidth:
| Hardware Architecture | Memory Bandwidth | Generation Speed | With N-Gram Speculative Decoding |
|---|---|---|---|
| DDR4 (Dual-channel 3200 MT/s) | ~50 GB/s | 14.8 - 15.2 tok/s | 22 - 26 tok/s |
| DDR5 (Dual-channel 5600-6400 MT/s) | ~85 - 100 GB/s | 30 - 38 tok/s | 50 - 60 tok/s |
| DDR5 (High-speed / Quad-channel) | ~130+ GB/s | 42 - 48 tok/s | 65 - 75+ tok/s |
Note: Quantization format is Q4_K_M (4-bit native SIMD register alignment), which preserves 99.2% FP16 accuracy while maximizing CPU vector execution.
latest: Latest official build with NomadCoder-4B-Q4_K_M.ggufq4_k_m: Specific 4-bit medium releaseContent type
Image
Digest
sha256:091d67819…
Size
2.8 GB
Last updated
18 days ago
docker pull massimolauri/nomadcoder-4b