A lightweight LLaMA.cpp HTTP server image based on Alpine.
10K+
Alpine LLaMA is an ultra-compact Docker image (less than 14 MB), providing a LLaMA.cpp HTTP server for language model inference.
This Docker image is particularly suited for:
You can host an HTTP inference server who leverages the LFM2.5 350M language model with:
docker run --name alpine-llama -p 80:8080 -e LLAMA_API_KEY=sk-xxxx -e LLAMA_ARG_MODEL_URL=https://huggingface.co/LiquidAI/LFM2.5-350M-GGUF/resolve/main/LFM2.5-350M-Q4_K_M.gguf samueltallet/alpine-llama-cpp-server
Once the GGUF model file is downloaded (and cached in the Docker container filesystem), you can query the OpenAI-compatible Chat Completions API endpoint exposed.
For example, you can classify the sentiment of a feedback with:
curl -s http://127.0.0.1/v1/chat/completions \
-H 'Authorization: Bearer sk-xxxx' \
-d '{
"messages": [
{ "role": "user", "content": "Classify as exactly one word (Positive, Neutral, or Negative) the sentiment of this feedback: This application doesn''t work for all scenarios, but I think it has potential." }
],
"temperature": 0,
"max_tokens": 2
}' | jq '.choices[0].message.content'
# > "Neutral"
Notes for the above script:
127.0.0.1 with your server IP.sk-xxxx.sudo apt install jq on Debian-based systems.See the GitHub repository README for more examples (structured output and summarization).
You can pass environment variables to the Docker container to configure the Alpine LLaMA server:
| Environment Variable | Description | Example Value |
|---|---|---|
LLAMA_ARG_HF_REPO | Hugging Face (HF) repository of a model | bartowski/Llama-3.2-1B-Instruct-GGUF |
LLAMA_ARG_HF_FILE | and model file to use in this HF repository | Llama-3.2-1B-Instruct-Q4_K_M.gguf |
LLAMA_ARG_MODEL | or path to a model file in your hard disk | /home/you/LLMs/Llama-3.2-1B.gguf |
LLAMA_ARG_MODEL_URL | or URL to download the model file from. | https://your.host/Llama-3.2-1B.gguf |
LLAMA_API_KEY | Key for authenticating HTTP API requests. | sk-n5V9UAJt6wRFfZQ4eDYk37uGzbKXdpNj |
LLAMA_ARG_ALIAS | Alias of the model in HTTP API requests. | Llama-3.2-1B |
An exhaustive list of these variables can be found in the official LLaMA.cpp server documentation.
Project licensed under MIT. See the LICENSE file for details.
© 2026 Samuel Tallet
Content type
Image
Digest
sha256:898345bb2…
Size
12.9 MB
Last updated
2 days ago
docker pull samueltallet/alpine-llama-cpp-server