📦 Phi-3 Mini Q8K - Advanced CPU Conversational AI
📞 Contact & Support
| 👤 Author | Artem Ryzhov |
| 🌐 Website | https://www.ryzhov.website |
|
[email protected] [email protected] | |
| artem-ryzhov-90560a2a0 | |
| 💻 GitHub | ryzhov-artem |
Advanced Q8K quantization implementation for Microsoft Phi-3 Mini (3.8B parameters) using Rust and the Candle ML framework. Delivers 2-3 tokens/second on consumer CPUs with 4.2GB model size (46% reduction from 7.6GB) while maintaining <1% accuracy loss.
Quick Start Interactive Chat (Standalone)
docker run -it --rm artemr87/phi3-q8k:latest
With Resource Limits (Recommended)
docker run -it --rm
--memory=6g
--cpus=4
artemr87/phi3-q8k:latest
✨ Features:
Advanced AI Capabilities: Multi-turn conversational AI with context retention Sliding window context (3072 tokens) for long conversations Automatic history pruning to maintain performance Chat format compliance with Phi-3's <|system|>, <|user|>, <|assistant|> tokens Custom system prompts for role specialization
Performance Optimizations:
Q8K quantization (8-bit block quantization with permutation) KV-cache for efficient attention computation Block-wise permutation (64-element blocks) for locality Rotary position embeddings (RoPE) for long-range dependencies Grouped-query attention (32 heads, 32 KV heads) SiLU activation in feed-forward layers
🔧 Technical Implementation:
Packed safetensors format (single-file deployment) Memory-mapped loading for instant startup Dual validation (matmul + direct reconstruction) RMSE tracking (Root Mean Square Error < 1e-3) Adaptive QR decomposition for permutation optimization CPU-only inference (no CUDA/ROCm required)
🛠️ Sampling & Generation:
Temperature control (default: 0.6 for balanced creativity) Top-K sampling (k=50) for diversity Top-P (nucleus) sampling (p=0.9) for coherence Repeat penalty (1.15) to reduce repetition Repeat-last-N (64 tokens) for local penalty Token streaming for real-time output
📊 Example Session:
$ docker run -it --rm artemr87/phi3-q8k:latest
Loading Phi-3 Q8K model... Loading from 2 shards... Shard 1: /app/shard1.safetensors (2.8GB) Shard 2: /app/shard2.safetensors (1.5GB)
Phi-3 Mini Q8K Conversational AI Commands: 'exit' to quit | 'reset' to clear history
You: What is quantum computing? Assistant: Quantum computing is a type of computation that harnesses quantum mechanical phenomena like superposition and entanglement to process information in fundamentally different ways than classical computers. Unlike classical bits (0 or 1), quantum bits (qubits) can exist in multiple states simultaneously, enabling exponential speedups for certain problems like cryptography and optimization.
[Pos: 182 | Cache: 128 tok (96.0 MB) | Speed: 2.4 t/s | Hist: 3 msgs]
You: How does it compare to classical computers? Assistant: Quantum computers excel at specific tasks: • Factoring large numbers (Shor's algorithm) • Searching databases (Grover's algorithm) • Simulating molecules and materials
However, classical computers remain superior for most everyday tasks like word processing, web browsing, and sequential calculations. Quantum computers are not replacements but complementary tools for specialized problems.
[Pos: 365 | Cache: 256 tok (192.0 MB) | Speed: 2.1 t/s | Hist: 5 msgs]
You: reset Conversation history cleared!
You: exit Goodbye!
💻 Hardware Requirements:
Resource Minimum Recommended Notes RAM 6GB 8GB+ Model: 4.2GB, Cache: ~750MB/1K tokens CPU 2 cores (x86_64) 4+ cores AMD/Intel with AVX2 Storage 5GB 10GB Container + model shards Speed ~1.5 t/s ~3.0 t/s Tokens per second Architecture Support ✅ x86_64 (AMD64): Primary target ⚠️ ARM64: Not currently supported (use native x86_64 hardware)
🔬 Technical Deep Dive: Quantization Strategy Original FP16 Model: 7.6 GB (3.8B params × 2 bytes) Q8K Quantized: 4.2 GB (46% reduction) Accuracy Loss: <1% (RMSE < 1e-3) Compression Ratio: 1.81×
Q8K Block Structure:
256 elements per block (QK_K constant) 8-bit integers with per-block scale/min Block-wise permutation for improved compression Attention layer caching (shared permutations for Q/K/V/O) Permutation Strategies Block-wise permutation (64-element blocks)
Clusters similar magnitudes within local blocks Preserves cache locality for CPU inference Used by default for all layers SVD-inspired ranking (column variance)
Sorts by importance (high variance = high importance) Optional for attention layers QR pivot decomposition (partial QR)
Adaptive: 25%-100% QR based on matrix size Optimizes column ordering for numerical stability Model Architecture (Phi-3 Mini)
Parameters: 3.8 billion Layers: 32 transformer blocks Hidden Size: 3072 Intermediate Size: 8192 Attention Heads: 32 (all heads used for K/V) Head Dimension: 96 Max Sequence Length: 4096 tokens Vocabulary Size: 32,064 tokens
Attention Mechanism:
Grouped-query attention (GQA) with 32 heads RoPE embeddings (θ = 10,000) for position encoding KV-cache with automatic reset on reset command Causal masking for autoregressive generation Feed-Forward Network SiLU activation (Swish variant): x * sigmoid(x) Gated FFN: silu(gate) * up structure Gate-up projection: 3072 → 8192 Down projection: 8192 → 3072 Memory Management
cache_memory_mb = ( num_layers * 2 * # K and V caches batch_size * # Always 1 for inference num_kv_heads * # 32 for Phi-3 seq_len * # Current sequence length head_dim * # 96 4 # FP32 bytes ) / (1024 * 1024)
Sliding Window Context Max history: 3072 tokens (75% of 4096 context) Pruning strategy: FIFO (oldest messages removed first) System prompt: Always retained (pinned to position 0) Trigger: Automatic when token count > max_history_tokens
Commands & Controls Interactive Commands exit - Quit the conversation and terminate container reset - Clear conversation history and KV-cache (restarts from scratch)
📈 Performance Benchmarks Speed Comparison (4-core x86_64 CPU) Model Parameters Size Speed Use Case TinyLlama Q8K 1.1B 1.3GB ~8 t/s Chat, simple tasks Phi-3 Mini Q8K 3.8B 4.2GB ~3 t/s Advanced reasoning Llama-3 8B 8B 9GB ~1.5 t/s Complex analysis
📖 Learn More 🎓 Resources Live Demo: ryzhov.website/phi3-chat - Try the model in your browser Technical Article: Complete Implementation Guide - 30-min deep dive Author Profile: Artem Ryzhov - Full-stack Node.js & Rust developer More Projects: Articles Hub - ML, backend, systems programming 🔗 External Links Candle Framework: github.com/huggingface/candle Phi-3 Model: huggingface.co/microsoft/Phi-3-mini-4k-instruct SafeTensors: github.com/huggingface/safetensors GGML/GGUF: github.com/ggerganov/llama.cpp 📚 Related Papers Phi-3 Technical Report: Microsoft Research Quantization Survey: arXiv:2103.13630 RoPE Embeddings: arXiv:2104.09864
✅ Commercial Use: This container is free for personal, academic, and commercial use under Apache 2.0 license 🛠️ Custom Deployment: Available for consulting on model optimization and deployment 📄 License Apache License 2.0
✅ Commercial use ✅ Modification ✅ Distribution ✅ Private use ⚠️ Must include original license ⚠️ Must state changes Base Model License: Microsoft Phi-3 (MIT License)
See LICENSE file for full details.
Roadmap for future:
ARM64 support (native build for Apple M-series) 4-bit quantization (Q4K for 2x further compression) Batch inference (process multiple prompts simultaneously) REST API mode (HTTP server for production deployments) Fine-tuning support (LoRA adapters for domain-specific tasks) Model merging (combine multiple Phi-3 variants)
🔗 External Links Candle Framework: github.com/huggingface/candle TinyLlama Model: huggingface.co/TinyLlama GGUF Format: github.com/ggerganov/llama.cpp
🙏 Acknowledgments Built with:
Hugging Face Candle (Apache 2.0) TinyLlama (Apache 2.0) SafeTensors (Apache 2.0)
Made with ❤️ by Artem Ryzhov
Content type
Image
Digest
sha256:47ee258a1…
Size
3.9 GB
Last updated
8 months ago
docker pull artemr87/phi3-q8k