Mixed-precision Q8K/Q4K quantization for Microsoft Phi-3 Mini (3.8B parameters) using Rust and the Candle ML framework. Attention-critical layers keep 8-bit precision (Q8K) while less-sensitive MLP down-projections drop to 4-bit (Q4K), compressing BF16 weights from ~7.6 GB to ~2 GB (74% reduction) with minimal accuracy loss. CPU-only, no GPU required.
The name "nibble" references the 4-bit Q4K blocks — a nibble is exactly 4 bits.
| Author | Artem Ryzhov |
| Website | https://www.ryzhov.website |
| [email protected] / [email protected] | |
| artem-ryzhov-90560a2a0 | |
| GitHub | ryzhov-artem |
| Docker Hub | artemr87 |
Interactive Chat (Standalone)
docker run -it --rm artemr87/phi3-nibble:latest
With Resource Limits (Recommended)
docker run -it --rm \
--memory=4g \
--cpus=4 \
artemr87/phi3-nibble:latest
<|system|>, <|user|>, <|assistant|> tokens)| Layer | Format | Bits/Weight | Rationale |
|---|---|---|---|
| Attention Q/K/V/O | Q8K | 8 | High information density |
| MLP gate_proj, up_proj | Q8K | 8 | Activation-critical |
| MLP down_proj | Q4K | 4 | Less sensitive, 50% block size reduction |
| Embeddings, norms | F32 | 32 | Preserved unquantized |
Parameters: 3.8 billion
Layers: 32 transformer blocks
Hidden Size: 3072
Intermediate Size: 8192
Attention Heads: 32 (all heads used for K/V)
Head Dimension: 96
Max Sequence Length: 4096 tokens
Vocabulary Size: 32,064 tokens
Forward Pass:
token embeddings → 32x (RMSNorm → multi-head attention with RoPE + KV cache → residual → RMSNorm → SiLU-gated MLP → residual) → final RMSNorm → lm_head → logits
# Cache Memory Formula
cache_memory_mb = (
num_layers * 2 * # K and V caches
batch_size * # Always 1 for inference
num_kv_heads * # 32 for Phi-3
seq_len * # Current sequence length
head_dim * # 96
4 # FP32 bytes
) / (1024 * 1024)
# Example: 1000 tokens cached
# 32 × 2 × 1 × 32 × 1000 × 96 × 4 = 750 MB
$ docker run -it --rm artemr87/phi3-nibble:latest
Loading Phi-3 Mixed Q8K/Q4K model...
Loading from 2 shards...
Shard 1: /app/shard1.safetensors
Shard 2: /app/shard2.safetensors
Phi-3 Mini Mixed Q8K/Q4K Conversational AI
Commands: 'exit' to quit | 'reset' to clear history
You: Explain the difference between Q8K and Q4K quantization.
Assistant: Q8K and Q4K are block-based quantization formats that compress
floating-point model weights into fixed-point integers:
Q8K (8-bit): Each block of 256 elements is stored as 8-bit integers
with per-block scale and minimum values. This preserves most of the
original precision while halving the storage compared to FP16. Used for
attention-critical layers where accuracy matters most.
Q4K (4-bit): Same block size (256 elements) but uses only 4 bits per
weight, achieving 75% compression from FP16. Some precision is lost, but
for less-sensitive layers like MLP down-projections, the quality impact
is minimal.
The "mixed" strategy applies Q8K where precision matters (attention,
gate/up projections) and Q4K where storage savings outweigh accuracy
costs (down projections).
[Pos: 198 | Cache: 128 tok (96.0 MB) | Speed: 2.8 t/s | Hist: 3 msgs]
You: Write a Python function to calculate Fibonacci numbers.
Assistant: Here's an efficient implementation using iteration:
def fibonacci(n: int) -> int:
if n <= 1:
return n
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
This runs in O(n) time and O(1) space.
[Pos: 312 | Cache: 256 tok (192.0 MB) | Speed: 3.1 t/s | Hist: 5 msgs]
You: reset
Conversation history cleared!
You: exit
Goodbye!
| Resource | Minimum | Recommended | Notes |
|---|---|---|---|
| RAM | 4GB | 6GB+ | Model: ~2GB, Cache: ~750MB/1K tokens |
| CPU | 2 cores (x86_64) | 4+ cores | AMD/Intel with AVX2 |
| Storage | 5GB | 8GB | Container + model shards |
| Architecture | x86_64 (AMD64) | — | ARM64 not currently supported |
| Model | Params | Size | Quant | Speed |
|---|---|---|---|---|
| TinyLlama Q8K | 1.1B | 1.3GB | Q8K | ~8 t/s |
| Phi-3 Mini Mixed (this) | 3.8B | ~2GB | Q8K/Q4K | ~3 t/s |
| Phi-3 Mini Q8K | 3.8B | 4.2GB | Q8K | ~3 t/s |
The mixed Q8K/Q4K strategy achieves roughly the same inference speed as pure Q8K while cutting model size in half — Q4K layers are smaller but require slightly more computation for dequantization.
| Command | Action |
|---|---|
exit | Quit the conversation and terminate the container |
reset | Clear conversation history and KV-cache |
Resources:
External Links:
Related Papers:
Apache License 2.0
Base Model License: Microsoft Phi-3 (MIT License)
Built with Hugging Face Candle (Apache 2.0), Microsoft Phi-3 (MIT), SafeTensors (Apache 2.0)
Made with care by Artem Ryzhov
Content type
Image
Digest
sha256:1fd99b621…
Size
3.5 GB
Last updated
6 months ago
docker pull artemr87/phi3-nibble:1.0.1