Sign inSign up

mutazai/nvfp4-cuda13-sota-quantization

By mutazai

•Updated 10 months ago

NVFP4 Quantization Pipeline | CUDA 13.0 | 72% Memory Reduction | <0.3% Loss

Image
Machine learning & AI
Developer tools
Data science
0

2.5K

mutazai/nvfp4-cuda13-sota-quantization repository overview

⁠🦌 ELK-AI | NVFP4 Quantization Toolkit

Quantization | NVFP4 Ready | Blackwell Native

NVFP4 Ready - Parameters Varies VRAM vLLM CUDA 13 Blackwell H100 A100

Mutaz Al Awamleh⁠ | ELK-AI⁠


⁠🧠 About

Complete NVFP4 quantization toolkit for converting models to 4-bit precision. Includes calibration tools and validation scripts.


⁠🚀 Why This Image?

We rebuilt the entire inference stack from scratch with CUDA 13.0.

ProblemOur Solution
FlashInfer compilation 2+ hrsPre-compiled for SM80-SM121
Mamba SSM CUDA mismatchesPre-built for all architectures
50+ undocumented env varsBattle-tested configuration
Days of CUDA graph tuningOptimized out of the box

Result: From WEEKS of setup to 30 SECONDS.


⁠💻 Hardware Requirements

RequirementMinimum
VRAMVaries
GPUAny NVIDIA GPU
CUDA12.0+

⁠📦 Specs

SpecValue
Parameters-
QuantizationNVFP4 Ready
Size11.8 GB (was 15.0 GB)
ContextModel-dependent
SpeedOptimized

⁠🏗️ 7-Layer Optimization Stack

+-------------------------------------------------------------+
|           ELK-AI GROUND-UP OPTIMIZATION STACK               |
+-------------------------------------------------------------+
|  Layer 7: NVFP4              | 21% memory reduction        |
|  Layer 6: FP8 KV-Cache          | 2x context length          |
|  Layer 5: FlashInfer 0.2.6      | 3x faster decoding         |
|  Layer 4: CUDA Graphs           | 6x faster warmup           |
|  Layer 3: Mamba SSM 2.2.4       | State-space models         |
|  Layer 2: vLLM V1 Engine        | Optimal batching           |
|  Layer 1: CUDA 13.0 + SM121     | Native Blackwell           |
+-------------------------------------------------------------+
|  Base: NGC PyTorch 25.11 | cuBLAS 12.9 | TensorRT 10.x      |
+-------------------------------------------------------------+

⁠🚀 Quick Start

docker run --gpus all -p 8000:8000 elkaioptimization/nvfp4-cuda13-sota-quantization:1.0

⁠🦌 More ELK-AI Models

ModelSizeLink
Qwen3-VL-2B2.1 GBDocker Hub⁠
Qwen3-VL-4B4.2 GBDocker Hub⁠
Qwen3-VL-8B8.4 GBDocker Hub⁠
Nemotron3-30B31.5 GBDocker Hub⁠
Nemotron-VL-12B12.6 GBDocker Hub⁠
Devstral-24B53.8 GBDocker Hub⁠
vLLM Base11.8 GBDocker Hub⁠
vLLM Blackwell11.8 GBDocker Hub⁠

🦌 ELK-AI — From weeks of setup to 30 seconds

LinkedIn⁠ | Website⁠ | Docker Hub⁠

Tag summary

Content type

Image

Digest

sha256:2ca88e0b1…

Size

11 GB

Last updated

10 months ago

docker pull mutazai/nvfp4-cuda13-sota-quantization