Production-ready Docker image for llama-cpp-python server with CUDA 13.1.0 GPU acceleration.
1.5K
A Docker image for running llama-cpp-python server with CUDA acceleration. This image provides a production-ready environment for serving Large Language Models (LLMs) with GPU acceleration.
docker run -d \
--gpus all \
-p 8080:8080 \
-v /path/to/models:/llama/models \
-v /path/to/config:/llama/config \
ametnes/llama-cpp:0.3.16-cuda13.1.0
docker run -d \
--gpus all \
-p 8080:8080 \
-e CONFIG_PATH=/llama/config/custom-settings.json \
-v /path/to/models:/llama/models \
-v /path/to/config:/llama/config \
ametnes/llama-cpp:0.3.16-cuda13.1.0
| Variable | Default | Description |
|---|---|---|
CONFIG_PATH | /llama/config/settings.json | Path to the llama-cpp server configuration file |
| Path | Description |
|---|---|
/llama/models | Directory containing LLM model files |
/llama/config | Directory containing configuration files |
| Port | Description |
|---|---|
8080 | HTTP API server port |
For optimal GPU performance, ensure the following environment variables are set:
NVIDIA_VISIBLE_DEVICES=all
NVIDIA_DRIVER_CAPABILITIES=compute,utility
LD_LIBRARY_PATH=/usr/local/cuda-13.1/compat:/usr/local/cuda/lib64:/usr/local/nvidia/lib64
apiVersion: apps/v1
kind: Deployment
metadata:
name: llama-cpp
spec:
replicas: 1
selector:
matchLabels:
app: llama-cpp
template:
metadata:
labels:
app: llama-cpp
spec:
containers:
- name: llama-cpp
image: ametnes/llama-cpp:0.3.16-cuda13.1.0
ports:
- containerPort: 8080
env:
- name: CONFIG_PATH
value: /llama/config/settings.json
- name: NVIDIA_VISIBLE_DEVICES
value: "all"
- name: NVIDIA_DRIVER_CAPABILITIES
value: "compute,utility"
resources:
limits:
nvidia.com/gpu: 1
requests:
nvidia.com/gpu: 1
volumeMounts:
- name: models
mountPath: /llama/models
- name: config
mountPath: /llama/config
volumes:
- name: models
persistentVolumeClaim:
claimName: llama-models-pvc
- name: config
configMap:
name: llama-config
The server expects a JSON configuration file. Example settings.json:
{
"model": "/llama/models/llama-2-7b-chat.Q4_K_M.gguf",
"n_ctx": 4096,
"n_gpu_layers": 35,
"host": "0.0.0.0",
"port": 8080
}
To build this image from the Dockerfile:
docker build -t ametnes/llama-cpp:0.3.16-cuda13.1.0 \
--build-arg LLAMA_VERSION=0.3.16 \
-f cuda.Dockerfile .
| Argument | Default | Description |
|---|---|---|
BASE_DEVEL | nvidia/cuda:13.1.0-devel-ubuntu24.04 | CUDA development base image |
BASE_RUNTIME | nvidia/cuda:13.1.0-runtime-ubuntu24.04 | CUDA runtime base image |
LLAMA_VERSION | 0.3.16 | Version of llama-cpp-python to install |
UID | 1000 | User ID for the llama user |
GID | 1000 | Group ID for the llama user |
CONFIG_PATH | /llama/config/settings.json | Default config file path |
nvidia/cuda:13.1.0-runtime-ubuntu24.04llama (UID 1000)/llamaEnsure:
NVIDIA_VISIBLE_DEVICES environment variable is setThe container runs as UID 1000. Ensure mounted volumes have appropriate permissions:
chown -R 1000:1000 /path/to/models /path/to/config
This image includes llama-cpp-python which is licensed under the MIT License.
For issues and questions, please open an issue in the repository.
Content type
Image
Digest
sha256:86608ab0e…
Size
1.5 GB
Last updated
9 months ago
docker pull ametnes/llama-cpp:0.3.16-cuda13.1.0