Community User
ELK-AI
Dubai, UAE
Displaying 1 to 15 of 15 repositories
🦌+🦅 2x Faster • 3x Smaller • 6x Savings • <1% Loss | NVFP4 Quantized
9m
10K+
SOTA LLM Inference | Latest CUDA 13.0 | Blackwell 6-20x | H100/A100 2-6x faster
9m
10K+
🦌 Qwen3-VL-32B NVFP4 | Vision-Language | 20GB (was 62GB) | <0.3% accuracy loss
9m
479
Nemotron3-30B NVFP4 Quantized - First NVFP4 for NVIDIA Nemotron-3
9m
10K+
Zero-Config TensorRT-LLM | Pre-loaded Nemotron-3-Nano-30B | By Mutaz Al Awamleh - ELK-AI
9m
598
Devstral-Small-2-24B FP8 - Blackwell-optimized Mistral coding model
9m
558
NVIDIA Nemotron-VL-12B NVFP4 - Blackwell-optimized multimodal vision-language model
9m
2.2K
Qwen3-VL-8B-Thinking NVFP4 W4A16 - First NVFP4 quantization for Blackwell
9m
490
Qwen3-VL-4B-Thinking NVFP4 W4A16 - First NVFP4 quantization for Blackwell
9m
486
Qwen3-VL-2B-Thinking NVFP4 W4A16 - First NVFP4 quantization for Blackwell
9m
501
vLLM Blackwell NVFP4 Optimized - Blackwell-optimized vLLM for NVFP4/FP8
9m
4.8K
NVFP4 Quantization Pipeline | CUDA 13.0 | 72% Memory Reduction | <0.3% Loss
10m
2.5K