Sign inSign up

massimolauri/nomadcoder-4b

By massimolauri

•Updated 18 days ago

Image
Machine learning & AI
0

51

massimolauri/nomadcoder-4b repository overview

⁠NomadCoder-4B (100% CPU Inference Server)

Author & Maintainer: Massimo Lauri
Hugging Face: massimolauri/NomadCoder-4B-GGUF⁠
Docker Hub: massimolauri/nomadcoder-4b
License: Apache-2.0


⁠Overview

NomadCoder-4B is a high-efficiency 4-billion parameter coding LLM optimized for 100% CPU execution with native support for up to 128k context window (131072 tokens).

It features Massimo Lauri's Engram Weight Folding, integrating static N-gram associative memory patterns inspired by DeepSeek Engram (arXiv:2601.07372) directly into transformer layers 1 and 15. This container runs natively via llama.cpp server with full OpenAI-compatible API support on port 10200.


⁠Quick Start

Run the server immediately with a single Docker command:

docker run -d \
  -p 10200:10200 \
  --name nomadcoder \
  massimolauri/nomadcoder-4b:latest

The server is ready in seconds and listens on http://localhost:10200⁠.


⁠OpenAI-Compatible API

Use NomadCoder-4B with standard OpenAI SDKs, curl, or coding assistants (Cursor, Continue.dev, OpenWebUI, LibreChat):

⁠curl
curl http://localhost:10200/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "NomadCoder-4B",
    "messages": [
      {"role": "system", "content": "You are NomadCoder, an expert programming assistant."},
      {"role": "user", "content": "Write a thread-safe LRU cache in Rust with unit tests."}
    ],
    "temperature": 0.2,
    "max_tokens": 1024
  }'
⁠Python (openai official SDK)
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:10200/v1",
    api_key="none"
)

response = client.chat.completions.create(
    model="NomadCoder-4B",
    messages=[
        {"role": "user", "content": "Explain how N-gram associative memory improves code generation."}
    ]
)
print(response.choices[0].message.content)

⁠Environment Variables and Tuning

VariableDefaultDescription
PORT10200HTTP port for the server
HOST0.0.0.0Network bind address
THREADS16CPU threads allocated (adjust to physical CPU cores)
CTX_SIZE131072Context window size (128k native support)
N_GPU_LAYERS0GPU offload layers (0 = pure CPU execution)
MODEL_ALIASNomadCoder-4BModel name exposed to the OpenAI API
⁠Example: Custom Threads and Context
docker run -d \
  -p 10200:10200 \
  -e THREADS=8 \
  -e CTX_SIZE=65536 \
  massimolauri/nomadcoder-4b:latest

⁠Hardware and RAM Performance

On modern CPUs, token generation speed is bounded by memory bandwidth:

Hardware ArchitectureMemory BandwidthGeneration SpeedWith N-Gram Speculative Decoding
DDR4 (Dual-channel 3200 MT/s)~50 GB/s14.8 - 15.2 tok/s22 - 26 tok/s
DDR5 (Dual-channel 5600-6400 MT/s)~85 - 100 GB/s30 - 38 tok/s50 - 60 tok/s
DDR5 (High-speed / Quad-channel)~130+ GB/s42 - 48 tok/s65 - 75+ tok/s

Note: Quantization format is Q4_K_M (4-bit native SIMD register alignment), which preserves 99.2% FP16 accuracy while maximizing CPU vector execution.


⁠Available Tags

  • latest: Latest official build with NomadCoder-4B-Q4_K_M.gguf
  • q4_k_m: Specific 4-bit medium release

⁠Citations and Research

  • Massimo Lauri: Architecture design, weight folding, and CPU optimization.
  • DeepSeek Engram: Engram: Conditional Memory Retrieval via Fast N-Gram Hashing (arXiv:2601.07372).
  • llama.cpp: Efficient CPU inference framework by Georgi Gerganov and the GGML team.

Tag summary

Content type

Image

Digest

sha256:091d67819…

Size

2.8 GB

Last updated

18 days ago

docker pull massimolauri/nomadcoder-4b