Sign inSign up

artemr87/phi3-nibble

By artemr87

•Updated 6 months ago

Image
0

141

artemr87/phi3-nibble repository overview

⁠Phi-3 Mini Mixed Q8K/Q4K — phi-nibble

Docker Pulls Docker Image Size License

Mixed-precision Q8K/Q4K quantization for Microsoft Phi-3 Mini (3.8B parameters) using Rust and the Candle ML framework. Attention-critical layers keep 8-bit precision (Q8K) while less-sensitive MLP down-projections drop to 4-bit (Q4K), compressing BF16 weights from ~7.6 GB to ~2 GB (74% reduction) with minimal accuracy loss. CPU-only, no GPU required.

The name "nibble" references the 4-bit Q4K blocks — a nibble is exactly 4 bits.


⁠Contact & Support


⁠Quick Start

Interactive Chat (Standalone)

docker run -it --rm artemr87/phi3-nibble:latest

With Resource Limits (Recommended)

docker run -it --rm \
  --memory=4g \
  --cpus=4 \
  artemr87/phi3-nibble:latest

⁠Features

⁠Mixed-Precision Quantization
  • Per-layer format routing: Q8K for attention Q/K/V/O + MLP gate/up, Q4K for MLP down
  • Optional column permutation (L2-norm magnitude reordering) for reduced quantization error
  • Permutation vectors stored alongside blocks and applied at inference time
  • ~2 GB packed SafeTensors from ~7.6 GB BF16 original (74% compression)
⁠Conversational AI
  • Multi-turn chat with context retention
  • Sliding window history (3072 tokens) with automatic FIFO pruning
  • Phi-3 chat format compliance (<|system|>, <|user|>, <|assistant|> tokens)
  • Adaptive sampling: temperature 0.6 / top-k 50 / top-p 0.9 for chat; temperature 0.35 / top-k 25 / top-p 0.95 for code queries
  • Repeat penalty (1.15) over last 64 tokens
⁠Performance
  • KV-cache for efficient attention computation
  • Rotary position embeddings (RoPE, theta=10000) for long-range dependencies
  • Grouped-query attention (32 heads, 32 KV heads)
  • SiLU-gated feed-forward network
  • Memory-mapped loading for fast startup
  • Streaming token output with real-time speed reporting (tokens/sec)

⁠Quantization Strategy

LayerFormatBits/WeightRationale
Attention Q/K/V/OQ8K8High information density
MLP gate_proj, up_projQ8K8Activation-critical
MLP down_projQ4K4Less sensitive, 50% block size reduction
Embeddings, normsF3232Preserved unquantized
⁠Q8K Block Structure
  • 256 elements per block (QK_K constant)
  • 8-bit integers with per-block scale/min
  • Block-wise permutation (64-element blocks) for improved locality
⁠Q4K Block Structure
  • 256 elements per block
  • 4-bit integers with per-block scale/min
  • 50% storage reduction vs Q8K
⁠Optional Permutation
  • Column-wise reordering by L2-norm magnitude prior to quantization
  • Clusters similar magnitudes within local blocks
  • Permutation indices stored in packed SafeTensors and applied during forward pass
  • Preserves cache locality for CPU inference

⁠Model Architecture (Phi-3 Mini 4K Instruct)

Parameters:           3.8 billion
Layers:               32 transformer blocks
Hidden Size:          3072
Intermediate Size:    8192
Attention Heads:      32 (all heads used for K/V)
Head Dimension:       96
Max Sequence Length:   4096 tokens
Vocabulary Size:      32,064 tokens

Forward Pass:

token embeddings → 32x (RMSNorm → multi-head attention with RoPE + KV cache → residual → RMSNorm → SiLU-gated MLP → residual) → final RMSNorm → lm_head → logits

⁠Memory Management
# Cache Memory Formula
cache_memory_mb = (
    num_layers * 2 *           # K and V caches
    batch_size *               # Always 1 for inference
    num_kv_heads *             # 32 for Phi-3
    seq_len *                  # Current sequence length
    head_dim *                 # 96
    4                          # FP32 bytes
) / (1024 * 1024)

# Example: 1000 tokens cached
# 32 × 2 × 1 × 32 × 1000 × 96 × 4 = 750 MB

⁠Example Session

$ docker run -it --rm artemr87/phi3-nibble:latest

Loading Phi-3 Mixed Q8K/Q4K model...
Loading from 2 shards...
  Shard 1: /app/shard1.safetensors
  Shard 2: /app/shard2.safetensors

Phi-3 Mini Mixed Q8K/Q4K Conversational AI
Commands: 'exit' to quit | 'reset' to clear history

You: Explain the difference between Q8K and Q4K quantization.
Assistant: Q8K and Q4K are block-based quantization formats that compress
floating-point model weights into fixed-point integers:

Q8K (8-bit): Each block of 256 elements is stored as 8-bit integers
with per-block scale and minimum values. This preserves most of the
original precision while halving the storage compared to FP16. Used for
attention-critical layers where accuracy matters most.

Q4K (4-bit): Same block size (256 elements) but uses only 4 bits per
weight, achieving 75% compression from FP16. Some precision is lost, but
for less-sensitive layers like MLP down-projections, the quality impact
is minimal.

The "mixed" strategy applies Q8K where precision matters (attention,
gate/up projections) and Q4K where storage savings outweigh accuracy
costs (down projections).

[Pos: 198 | Cache: 128 tok (96.0 MB) | Speed: 2.8 t/s | Hist: 3 msgs]

You: Write a Python function to calculate Fibonacci numbers.
Assistant: Here's an efficient implementation using iteration:

def fibonacci(n: int) -> int:
    if n <= 1:
        return n
    a, b = 0, 1
    for _ in range(2, n + 1):
        a, b = b, a + b
    return b

This runs in O(n) time and O(1) space.

[Pos: 312 | Cache: 256 tok (192.0 MB) | Speed: 3.1 t/s | Hist: 5 msgs]

You: reset
Conversation history cleared!

You: exit
Goodbye!

⁠Hardware Requirements

ResourceMinimumRecommendedNotes
RAM4GB6GB+Model: ~2GB, Cache: ~750MB/1K tokens
CPU2 cores (x86_64)4+ coresAMD/Intel with AVX2
Storage5GB8GBContainer + model shards
Architecturex86_64 (AMD64)—ARM64 not currently supported

⁠Performance Comparison

ModelParamsSizeQuantSpeed
TinyLlama Q8K1.1B1.3GBQ8K~8 t/s
Phi-3 Mini Mixed (this)3.8B~2GBQ8K/Q4K~3 t/s
Phi-3 Mini Q8K3.8B4.2GBQ8K~3 t/s

The mixed Q8K/Q4K strategy achieves roughly the same inference speed as pure Q8K while cutting model size in half — Q4K layers are smaller but require slightly more computation for dequantization.


⁠Commands

CommandAction
exitQuit the conversation and terminate the container
resetClear conversation history and KV-cache

⁠Learn More

Resources:

External Links:

Related Papers:

  • Phi-3 Technical Report — Microsoft Research
  • Quantization Survey — arXiv:2103.13630
  • RoPE Embeddings — arXiv:2104.09864

⁠License

Apache License 2.0

  • Commercial use allowed
  • Modification allowed
  • Distribution allowed
  • Private use allowed
  • Must include original license
  • Must state changes

Base Model License: Microsoft Phi-3 (MIT License)


⁠Roadmap

  • ARM64 support (Apple M-series native build)
  • REST API mode (HTTP server for production deployments)
  • LoRA adapter support for domain-specific fine-tuning
  • Batch inference (multiple prompts simultaneously)
  • Model merging (combine Phi-3 variants)

Built with Hugging Face Candle⁠ (Apache 2.0), Microsoft Phi-3⁠ (MIT), SafeTensors⁠ (Apache 2.0)

Made with care by Artem Ryzhov

Tag summary

Content type

Image

Digest

sha256:1fd99b621…

Size

3.5 GB

Last updated

6 months ago

docker pull artemr87/phi3-nibble:1.0.1