ARM64-optimized llama.cpp server running Gemma 3 270M model. Perfect for Raspberry Pi and ARM SBCs
986
A lightweight, ARM64-optimized Docker container running Google's Gemma 3 270M model via llama.cpp server. This image provides a complete local AI inference solution designed specifically for ARM64 devices like Raspberry Pi 4, Rock 4C+, and other ARM-based single board computers.
Developed with expertise from Supreet Sethi, AI specialist with deep knowledge in RAG (Retrieval-Augmented Generation) systems and edge AI deployment.
Gemma 3 270M is arguably the best model available at its size class. While most models at this scale are underwhelming, Gemma 3 270M delivers exceptional performance despite its compact footprint. It's a well-rounded generative model that can easily run on resource-constrained ARM64 devices while providing surprisingly capable text generation, making it ideal for edge AI deployments.
# Pull and run the container
docker run -d -p 8080:8080 --name gemma-server djinn/llama-gemma-server:latest
# Test the API
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Hello, how are you?"}],
"max_tokens": 100
}'
# Access web interface
open http://localhost:8080
POST /v1/chat/completions
Content-Type: application/json
{
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing simply."}
],
"max_tokens": 200,
"temperature": 0.7,
"stream": false
}
POST /v1/completions
Content-Type: application/json
{
"prompt": "The capital of France is",
"max_tokens": 50,
"temperature": 0.7
}
GET /v1/models # List available models
GET /props # Server properties
GET /health # Health check
Add "stream": true to any completion request for real-time token streaming.
LLAMA_SERVER_HOST - Bind address (default: 0.0.0.0)LLAMA_SERVER_PORT - Port number (default: 8080)LLAMA_SERVER_TIMEOUT - Request timeout in secondsLLAMA_N_CTX - Context window size (default: model's trained context)LLAMA_N_THREADS - CPU threads to use (default: auto-detect)To use a different model, mount it as a volume:
docker run -d -p 8080:8080 \
-v /path/to/your/model.gguf:/app/model.gguf \
--name custom-gemma \
djinn/llama-gemma-server:latest \
./bin/llama-server -m model.gguf --host 0.0.0.0
# Allocate more memory and CPU
docker run -d -p 8080:8080 \
--memory=6g \
--cpus=4 \
--name gemma-tuned \
djinn/llama-gemma-server:latest
import requests
def chat_with_gemma(message):
response = requests.post(
"http://localhost:8080/v1/chat/completions",
json={
"messages": [{"role": "user", "content": message}],
"max_tokens": 150,
"temperature": 0.7
}
)
return response.json()["choices"][0]["message"]["content"]
result = chat_with_gemma("What are the benefits of edge AI?")
print(result)
const axios = require('axios');
async function chatWithGemma(message) {
const response = await axios.post('http://localhost:8080/v1/chat/completions', {
messages: [{role: 'user', content: message}],
max_tokens: 150,
temperature: 0.7
});
return response.data.choices[0].message.content;
}
chatWithGemma("Explain machine learning in simple terms")
.then(result => console.log(result));
# Simple chat
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Write a haiku about technology"}],
"max_tokens": 100
}'
# Streaming response
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Tell me a story"}],
"max_tokens": 200,
"stream": true
}'
# Check container logs
docker logs gemma-server
# Common issues:
# - Insufficient memory (need 4GB+)
# - Port 8080 already in use
# - Wrong architecture (x86_64 instead of ARM64)
--memory=6g# Monitor memory usage
docker stats gemma-server
# Reduce memory usage by limiting context:
docker run -d -p 8080:8080 \
yourusername/llama-gemma-server:latest \
./bin/llama-server -m gemma-3-270m-it-Q8_0.gguf \
--host 0.0.0.0 --ctx-size 2048
-O2 optimization, tuned for Cortex-A72version: '3.8'
services:
gemma-server:
image: yourusername/llama-gemma-server:latest
ports:
- "8080:8080"
deploy:
resources:
limits:
memory: 6G
cpus: '4'
reservations:
memory: 2G
restart: unless-stopped
healthcheck:
test: ["CMD", "wget", "--quiet", "--tries=1", "--spider", "http://localhost:8080/health"]
interval: 30s
timeout: 10s
retries: 3
server {
listen 80;
server_name your-domain.com;
location / {
proxy_pass http://localhost:8080;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_buffering off; # Important for streaming
}
}
This image is built from the official llama.cpp repository with ARM64-specific optimizations. To report issues or contribute improvements, please visit the project repository.
latest - Latest stable build with Gemma 3 270Marm64 - Explicit ARM64 tagv1.0 - Stable release versionsdev - Development builds (may be unstable)For production deployments, use specific version tags rather than latest.
Content type
Image
Digest
sha256:48aafe463…
Size
293.9 MB
Last updated
about 1 year ago
docker pull djinn/llama-gemma-server