vLLM NVFP4 CUDA 13 - Blackwell-optimized vLLM for NVFP4/FP8 serving
2.6K
Foundation image with pre-compiled FlashInfer, Mamba SSM, and NVFP4 support. Build your own optimized models.
We rebuilt the entire inference stack from scratch with CUDA 13.0.
| Problem | Our Solution |
|---|---|
| FlashInfer compilation 2+ hrs | Pre-compiled for SM80-SM121 |
| Mamba SSM CUDA mismatches | Pre-built for all architectures |
| 50+ undocumented env vars | Battle-tested configuration |
| Days of CUDA graph tuning | Optimized out of the box |
Result: From WEEKS of setup to 30 SECONDS.
| Requirement | Minimum |
|---|---|
| VRAM | Varies |
| GPU | Any NVIDIA GPU |
| CUDA | 12.0+ |
| Spec | Value |
|---|---|
| Parameters | - |
| Quantization | NVFP4 Ready |
| Size | 11.8 GB (was 15.0 GB) |
| Context | Model-dependent |
| Speed | Optimized |
+-------------------------------------------------------------+
| ELK-AI GROUND-UP OPTIMIZATION STACK |
+-------------------------------------------------------------+
| Layer 7: NVFP4 | 21% memory reduction |
| Layer 6: FP8 KV-Cache | 2x context length |
| Layer 5: FlashInfer 0.2.6 | 3x faster decoding |
| Layer 4: CUDA Graphs | 6x faster warmup |
| Layer 3: Mamba SSM 2.2.4 | State-space models |
| Layer 2: vLLM V1 Engine | Optimal batching |
| Layer 1: CUDA 13.0 + SM121 | Native Blackwell |
+-------------------------------------------------------------+
| Base: NGC PyTorch 25.11 | cuBLAS 12.9 | TensorRT 10.x |
+-------------------------------------------------------------+
docker run --gpus all -p 8000:8000 elkaioptimization/vllm-nvfp4-cuda-13:2.5.0
| Model | Size | Link |
|---|---|---|
| Qwen3-VL-2B | 2.1 GB | Docker Hub |
| Qwen3-VL-4B | 4.2 GB | Docker Hub |
| Qwen3-VL-8B | 8.4 GB | Docker Hub |
| Nemotron3-30B | 31.5 GB | Docker Hub |
| Nemotron-VL-12B | 12.6 GB | Docker Hub |
| Devstral-24B | 53.8 GB | Docker Hub |
| vLLM Base | 11.8 GB | Docker Hub |
| vLLM Blackwell | 11.8 GB | Docker Hub |
🦌 ELK-AI — From weeks of setup to 30 seconds
Content type
Image
Digest
sha256:d2135f8a5…
Size
27.3 GB
Last updated
10 months ago
docker pull elkaioptimization/vllm-nvfp4-cuda-13