Sign inSign up

snowgoons/llamacpp

By snowgoons

•Updated 15 days ago

Llama.CPP Docker images

Image
0

9.1K

snowgoons/llamacpp repository overview

⁠LLaMA.cpp Docker Images

A collection of optimized Docker images for the llama.cpp⁠ inference engine, built for different hardware configurations.

⁠Overview

This repository provides pre-built Docker images for running LLaMA models with various hardware acceleration options. The images are optimized for performance and include different CUDA compute capabilities to support various NVIDIA GPU architectures.

⁠Available Images

⁠CPU-Only Build
  • Tag: llamacpp-<llama-version>-cpu
    • Description: Pure CPU inference build, compatible with any system
    • Usage: Best for systems without NVIDIA GPUs or when running on CPUs
⁠CUDA Builds
  • Tag: llamacpp-<llama-version>-cuda75

    • Description: CUDA build optimized for compute capability 7.5 (RTX 20xx series)
    • Usage: For older NVIDIA GPUs like RTX 2080, 2070, 2060
  • Tag: llamacpp-<llama-version>-cuda89

    • Description: CUDA build optimized for compute capability 8.9 (RTX 40xx series)
    • Usage: For newer NVIDIA GPUs like RTX 4090, 4080, 4070
  • Tag: llamacpp-<llama-version>-cuda120

    • Description: CUDA build optimized for compute capability 12.0 (RTX 50xx series)
    • Usage: For latest NVIDIA GPUs like RTX 5090, 5000 series

⁠Quick Start

Pull and run any of the images:

# CPU-only version
docker run -p 11434:11434 snowgoons/llamacpp:cpu

# CUDA 8.9 version (for RTX 40xx series)
docker run -p 11434:11434 --gpus all snowgoons/llamacpp:cuda89

⁠Configuration

The images are configured to:

  • Run on port 11434
  • Use the Qwen/Qwen2.5-7B-Instruct-GGUF model by default
  • Mount /usr/local/ai/models for model storage
  • Accept additional command-line arguments
⁠Customization Example
docker run -p 11434:11434 \
  -v /path/to/models:/usr/local/ai/models \
  --gpus all \
  snowgoons/llamacpp:cuda89 \
  --host 0.0.0.0 \
  --port 11434 \
  -hf "your-model-name" \
  -ngl 50

⁠Usage Notes

  • GPU Support: CUDA builds require NVIDIA GPU with appropriate drivers and nvidia-docker2 or Docker Engine v20.10+
  • Model Loading: Models are loaded from Hugging Face Hub by default, but can be mounted via volumes for local models
  • Performance: The CUDA builds provide significant speedup over CPU inference

Tag summary

Content type

Image

Digest

sha256:52dd77597…

Size

4 GB

Last updated

15 days ago

docker pull snowgoons/llamacpp:dev-1789660702-v0.4.1-cuda120