A collection of optimized Docker images for the llama.cpp inference engine, built for different hardware configurations.
This repository provides pre-built Docker images for running LLaMA models with various hardware acceleration options. The images are optimized for performance and include different CUDA compute capabilities to support various NVIDIA GPU architectures.
llamacpp-<llama-version>-cpu
Tag: llamacpp-<llama-version>-cuda75
Tag: llamacpp-<llama-version>-cuda89
Tag: llamacpp-<llama-version>-cuda120
Pull and run any of the images:
# CPU-only version
docker run -p 11434:11434 snowgoons/llamacpp:cpu
# CUDA 8.9 version (for RTX 40xx series)
docker run -p 11434:11434 --gpus all snowgoons/llamacpp:cuda89
The images are configured to:
11434/usr/local/ai/models for model storagedocker run -p 11434:11434 \
-v /path/to/models:/usr/local/ai/models \
--gpus all \
snowgoons/llamacpp:cuda89 \
--host 0.0.0.0 \
--port 11434 \
-hf "your-model-name" \
-ngl 50
nvidia-docker2 or Docker Engine v20.10+Content type
Image
Digest
sha256:52dd77597…
Size
4 GB
Last updated
15 days ago
docker pull snowgoons/llamacpp:dev-1789660702-v0.4.1-cuda120