Sign inSign up

samueltallet/alpine-llama-cpp-server

By samueltallet

•Updated 2 days ago

A lightweight LLaMA.cpp HTTP server image based on Alpine.

Image
Machine learning & AI
Web servers
2

10K+

samueltallet/alpine-llama-cpp-server repository overview

Alpine LLaMA is an ultra-compact Docker image (less than 14 MB), providing a LLaMA.cpp⁠ HTTP server for language model inference.

Docker Image Size‍ Raspberry Pi Friendly

⁠Use cases

This Docker image is particularly suited for:

  • Environments with limited disk space or low bandwidth.
  • Servers that cannot do GPU-accelerated inference, e.g. a CPU-only VPS or a Raspberry Pi.

⁠Quick start

You can host an HTTP inference server who leverages the LFM2.5 350M⁠ language model with:

docker run --name alpine-llama -p 80:8080 -e LLAMA_API_KEY=sk-xxxx -e LLAMA_ARG_MODEL_URL=https://huggingface.co/LiquidAI/LFM2.5-350M-GGUF/resolve/main/LFM2.5-350M-Q4_K_M.gguf samueltallet/alpine-llama-cpp-server

Once the GGUF model file is downloaded (and cached in the Docker container filesystem), you can query the OpenAI-compatible Chat Completions API endpoint exposed.

For example, you can classify the sentiment of a feedback with:

curl -s http://127.0.0.1/v1/chat/completions \
  -H 'Authorization: Bearer sk-xxxx' \
  -d '{
    "messages": [
      { "role": "user", "content": "Classify as exactly one word (Positive, Neutral, or Negative) the sentiment of this feedback: This application doesn''t work for all scenarios, but I think it has potential." }
    ],
    "temperature": 0,
    "max_tokens": 2
  }' | jq '.choices[0].message.content'
# > "Neutral"

Notes for the above script:

  • If you run docker remotely, replace 127.0.0.1 with your server IP.
  • In production, use your own strong secret key instead of sk-xxxx.
  • Install jq⁠ with sudo apt install jq on Debian-based systems.

⁠More examples

See the GitHub repository README⁠ for more examples (structured output and summarization).

⁠Configuration

You can pass environment variables to the Docker container to configure the Alpine LLaMA server:

Environment VariableDescriptionExample Value
LLAMA_ARG_HF_REPOHugging Face (HF) repository of a modelbartowski/Llama-3.2-1B-Instruct-GGUF
LLAMA_ARG_HF_FILEand model file to use in this HF repositoryLlama-3.2-1B-Instruct-Q4_K_M.gguf
LLAMA_ARG_MODELor path to a model file in your hard disk/home/you/LLMs/Llama-3.2-1B.gguf
LLAMA_ARG_MODEL_URLor URL to download the model file from.https://your.host/Llama-3.2-1B.gguf
LLAMA_API_KEYKey for authenticating HTTP API requests.sk-n5V9UAJt6wRFfZQ4eDYk37uGzbKXdpNj
LLAMA_ARG_ALIASAlias of the model in HTTP API requests.Llama-3.2-1B

An exhaustive list of these variables can be found in the official LLaMA.cpp server documentation⁠.

⁠License

Project licensed under MIT. See the LICENSE⁠ file for details.

© 2026 Samuel Tallet

Tag summary

Content type

Image

Digest

sha256:898345bb2…

Size

12.9 MB

Last updated

2 days ago

docker pull samueltallet/alpine-llama-cpp-server