Sign inSign up

ametnes/llama-cpp

By ametnes

•Updated 9 months ago

Production-ready Docker image for llama-cpp-python server with CUDA 13.1.0 GPU acceleration.

Image
Machine learning & AI
Data science
Web servers
0

1.5K

ametnes/llama-cpp repository overview

⁠llama-cpp CUDA Server

A Docker image for running llama-cpp-python⁠ server with CUDA acceleration. This image provides a production-ready environment for serving Large Language Models (LLMs) with GPU acceleration.

⁠Features

  • CUDA 13.1.0 support with NVIDIA GPU acceleration
  • llama-cpp-python 0.3.16 with server capabilities
  • Python 3.12 on Ubuntu 24.04
  • Multi-stage build for optimized image size
  • Non-root user execution (UID 1000)
  • Configurable via environment variables

⁠Requirements

  • Docker with NVIDIA Container Toolkit installed
  • NVIDIA GPU with CUDA 13.1+ compatible drivers
  • Kubernetes cluster with NVIDIA device plugin (for Kubernetes deployments)

⁠Quick Start

⁠Basic Usage
docker run -d \
  --gpus all \
  -p 8080:8080 \
  -v /path/to/models:/llama/models \
  -v /path/to/config:/llama/config \
  ametnes/llama-cpp:0.3.16-cuda13.1.0
⁠With Custom Config Path
docker run -d \
  --gpus all \
  -p 8080:8080 \
  -e CONFIG_PATH=/llama/config/custom-settings.json \
  -v /path/to/models:/llama/models \
  -v /path/to/config:/llama/config \
  ametnes/llama-cpp:0.3.16-cuda13.1.0

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
CONFIG_PATH/llama/config/settings.jsonPath to the llama-cpp server configuration file
⁠Volumes
PathDescription
/llama/modelsDirectory containing LLM model files
/llama/configDirectory containing configuration files
⁠Ports
PortDescription
8080HTTP API server port

⁠GPU Configuration

For optimal GPU performance, ensure the following environment variables are set:

NVIDIA_VISIBLE_DEVICES=all
NVIDIA_DRIVER_CAPABILITIES=compute,utility
LD_LIBRARY_PATH=/usr/local/cuda-13.1/compat:/usr/local/cuda/lib64:/usr/local/nvidia/lib64

⁠Kubernetes Example

apiVersion: apps/v1
kind: Deployment
metadata:
  name: llama-cpp
spec:
  replicas: 1
  selector:
    matchLabels:
      app: llama-cpp
  template:
    metadata:
      labels:
        app: llama-cpp
    spec:
      containers:
      - name: llama-cpp
        image: ametnes/llama-cpp:0.3.16-cuda13.1.0
        ports:
        - containerPort: 8080
        env:
        - name: CONFIG_PATH
          value: /llama/config/settings.json
        - name: NVIDIA_VISIBLE_DEVICES
          value: "all"
        - name: NVIDIA_DRIVER_CAPABILITIES
          value: "compute,utility"
        resources:
          limits:
            nvidia.com/gpu: 1
          requests:
            nvidia.com/gpu: 1
        volumeMounts:
        - name: models
          mountPath: /llama/models
        - name: config
          mountPath: /llama/config
      volumes:
      - name: models
        persistentVolumeClaim:
          claimName: llama-models-pvc
      - name: config
        configMap:
          name: llama-config

⁠Configuration File Format

The server expects a JSON configuration file. Example settings.json:

{
  "model": "/llama/models/llama-2-7b-chat.Q4_K_M.gguf",
  "n_ctx": 4096,
  "n_gpu_layers": 35,
  "host": "0.0.0.0",
  "port": 8080
}

⁠Building from Source

To build this image from the Dockerfile:

docker build -t ametnes/llama-cpp:0.3.16-cuda13.1.0 \
  --build-arg LLAMA_VERSION=0.3.16 \
  -f cuda.Dockerfile .
⁠Build Arguments
ArgumentDefaultDescription
BASE_DEVELnvidia/cuda:13.1.0-devel-ubuntu24.04CUDA development base image
BASE_RUNTIMEnvidia/cuda:13.1.0-runtime-ubuntu24.04CUDA runtime base image
LLAMA_VERSION0.3.16Version of llama-cpp-python to install
UID1000User ID for the llama user
GID1000Group ID for the llama user
CONFIG_PATH/llama/config/settings.jsonDefault config file path

⁠Image Details

  • Base Image: nvidia/cuda:13.1.0-runtime-ubuntu24.04
  • Python Version: 3.12
  • User: llama (UID 1000)
  • Working Directory: /llama
  • Exposed Port: 8080

⁠Troubleshooting

⁠CUDA Not Detected

Ensure:

  1. NVIDIA Container Toolkit is installed on the host
  2. GPU resources are properly requested in Kubernetes
  3. NVIDIA_VISIBLE_DEVICES environment variable is set
  4. CUDA driver version is compatible with CUDA 13.1
⁠Permission Errors

The container runs as UID 1000. Ensure mounted volumes have appropriate permissions:

chown -R 1000:1000 /path/to/models /path/to/config

⁠License

This image includes llama-cpp-python which is licensed under the MIT License.

⁠Support

For issues and questions, please open an issue in the repository.

Tag summary

Content type

Image

Digest

sha256:86608ab0e…

Size

1.5 GB

Last updated

9 months ago

docker pull ametnes/llama-cpp:0.3.16-cuda13.1.0