NVFP4 Quantization Pipeline | CUDA 13.0 | 72% Memory Reduction | <0.3% Loss
2.5K
Quantization | NVFP4 Ready | Blackwell Native
Complete NVFP4 quantization toolkit for converting models to 4-bit precision. Includes calibration tools and validation scripts.
We rebuilt the entire inference stack from scratch with CUDA 13.0.
| Problem | Our Solution |
|---|---|
| FlashInfer compilation 2+ hrs | Pre-compiled for SM80-SM121 |
| Mamba SSM CUDA mismatches | Pre-built for all architectures |
| 50+ undocumented env vars | Battle-tested configuration |
| Days of CUDA graph tuning | Optimized out of the box |
Result: From WEEKS of setup to 30 SECONDS.
| Requirement | Minimum |
|---|---|
| VRAM | Varies |
| GPU | Any NVIDIA GPU |
| CUDA | 12.0+ |
| Spec | Value |
|---|---|
| Parameters | - |
| Quantization | NVFP4 Ready |
| Size | 11.8 GB (was 15.0 GB) |
| Context | Model-dependent |
| Speed | Optimized |
+-------------------------------------------------------------+
| ELK-AI GROUND-UP OPTIMIZATION STACK |
+-------------------------------------------------------------+
| Layer 7: NVFP4 | 21% memory reduction |
| Layer 6: FP8 KV-Cache | 2x context length |
| Layer 5: FlashInfer 0.2.6 | 3x faster decoding |
| Layer 4: CUDA Graphs | 6x faster warmup |
| Layer 3: Mamba SSM 2.2.4 | State-space models |
| Layer 2: vLLM V1 Engine | Optimal batching |
| Layer 1: CUDA 13.0 + SM121 | Native Blackwell |
+-------------------------------------------------------------+
| Base: NGC PyTorch 25.11 | cuBLAS 12.9 | TensorRT 10.x |
+-------------------------------------------------------------+
docker run --gpus all -p 8000:8000 elkaioptimization/nvfp4-cuda13-sota-quantization:1.0
| Model | Size | Link |
|---|---|---|
| Qwen3-VL-2B | 2.1 GB | Docker Hub |
| Qwen3-VL-4B | 4.2 GB | Docker Hub |
| Qwen3-VL-8B | 8.4 GB | Docker Hub |
| Nemotron3-30B | 31.5 GB | Docker Hub |
| Nemotron-VL-12B | 12.6 GB | Docker Hub |
| Devstral-24B | 53.8 GB | Docker Hub |
| vLLM Base | 11.8 GB | Docker Hub |
| vLLM Blackwell | 11.8 GB | Docker Hub |
🦌 ELK-AI — From weeks of setup to 30 seconds
Content type
Image
Digest
sha256:2ca88e0b1…
Size
11 GB
Last updated
10 months ago
docker pull mutazai/nvfp4-cuda13-sota-quantization