Sign inSign up

djinn/llama-gemma-server

By djinn

•Updated about 1 year ago

ARM64-optimized llama.cpp server running Gemma 3 270M model. Perfect for Raspberry Pi and ARM SBCs

Image
Developer tools
Web servers
0

986

djinn/llama-gemma-server repository overview

⁠Llama.cpp Gemma 3 270M Server

A lightweight, ARM64-optimized Docker container running Google's Gemma 3 270M model via llama.cpp server. This image provides a complete local AI inference solution designed specifically for ARM64 devices like Raspberry Pi 4, Rock 4C+, and other ARM-based single board computers.

Developed with expertise from Supreet Sethi, AI specialist with deep knowledge in RAG (Retrieval-Augmented Generation) systems and edge AI deployment.

⁠Why Gemma 3 270M?

Gemma 3 270M is arguably the best model available at its size class. While most models at this scale are underwhelming, Gemma 3 270M delivers exceptional performance despite its compact footprint. It's a well-rounded generative model that can easily run on resource-constrained ARM64 devices while providing surprisingly capable text generation, making it ideal for edge AI deployments.

⁠Features

  • ARM64 Optimized: Built specifically for ARM64 architecture with Cortex-A72 optimizations
  • Lightweight: Minimal Alpine Linux base with only essential runtime dependencies
  • Ready-to-Run: Pre-loaded with Gemma 3 270M Q8_0 quantized model
  • OpenMP Acceleration: Multi-threaded inference using all available CPU cores
  • REST API: Compatible with OpenAI API format for easy integration
  • Web Interface: Built-in chat interface accessible via browser
  • Resource Efficient: Designed to run smoothly on devices with 4GB+ RAM

⁠Quick Start

# Pull and run the container
docker run -d -p 8080:8080 --name gemma-server djinn/llama-gemma-server:latest

# Test the API
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello, how are you?"}],
    "max_tokens": 100
  }'

# Access web interface
open http://localhost:8080

⁠Supported Platforms

⁠Tested ARM64 Devices
  • Raspberry Pi 4 (4GB/8GB models recommended)
  • Rock 4C+
  • NVIDIA Jetson Nano/Xavier
  • Apple Silicon Macs (M1/M2 for development)
  • ARM64 Cloud Instances (AWS Graviton, etc.)
⁠System Requirements
  • Architecture: ARM64/AArch64 only
  • Memory: 4GB RAM minimum, 8GB recommended
  • Storage: 2GB free space for container and model
  • CPU: ARM Cortex-A72 or newer recommended

⁠API Endpoints

POST /v1/chat/completions
Content-Type: application/json

{
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain quantum computing simply."}
  ],
  "max_tokens": 200,
  "temperature": 0.7,
  "stream": false
}
⁠Text Completions
POST /v1/completions
Content-Type: application/json

{
  "prompt": "The capital of France is",
  "max_tokens": 50,
  "temperature": 0.7
}
⁠Server Information
GET /v1/models          # List available models
GET /props              # Server properties
GET /health             # Health check
⁠Streaming Support

Add "stream": true to any completion request for real-time token streaming.

⁠Configuration Options

⁠Environment Variables
  • LLAMA_SERVER_HOST - Bind address (default: 0.0.0.0)
  • LLAMA_SERVER_PORT - Port number (default: 8080)
  • LLAMA_SERVER_TIMEOUT - Request timeout in seconds
  • LLAMA_N_CTX - Context window size (default: model's trained context)
  • LLAMA_N_THREADS - CPU threads to use (default: auto-detect)
⁠Custom Model

To use a different model, mount it as a volume:

docker run -d -p 8080:8080 \
  -v /path/to/your/model.gguf:/app/model.gguf \
  --name custom-gemma \
  djinn/llama-gemma-server:latest \
  ./bin/llama-server -m model.gguf --host 0.0.0.0
⁠Performance Tuning
# Allocate more memory and CPU
docker run -d -p 8080:8080 \
  --memory=6g \
  --cpus=4 \
  --name gemma-tuned \
  djinn/llama-gemma-server:latest

⁠Performance Expectations

⁠Raspberry Pi 4 (8GB)
  • Tokens/second: 3-6 tokens/sec
  • Memory usage: ~1.5-2GB
  • CPU usage: 60-80% across all cores
  • Response latency: 2-5 seconds for typical queries
⁠Rock 4C+
  • Tokens/second: 4-8 tokens/sec
  • Memory usage: ~1.5-2GB
  • CPU usage: 50-70% across all cores
  • Response latency: 1-3 seconds for typical queries

⁠Integration Examples

⁠Python Client
import requests

def chat_with_gemma(message):
    response = requests.post(
        "http://localhost:8080/v1/chat/completions",
        json={
            "messages": [{"role": "user", "content": message}],
            "max_tokens": 150,
            "temperature": 0.7
        }
    )
    return response.json()["choices"][0]["message"]["content"]

result = chat_with_gemma("What are the benefits of edge AI?")
print(result)
⁠Node.js Client
const axios = require('axios');

async function chatWithGemma(message) {
    const response = await axios.post('http://localhost:8080/v1/chat/completions', {
        messages: [{role: 'user', content: message}],
        max_tokens: 150,
        temperature: 0.7
    });
    
    return response.data.choices[0].message.content;
}

chatWithGemma("Explain machine learning in simple terms")
    .then(result => console.log(result));
⁠cURL Examples
# Simple chat
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Write a haiku about technology"}],
    "max_tokens": 100
  }'

# Streaming response  
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Tell me a story"}],
    "max_tokens": 200,
    "stream": true
  }'

⁠Troubleshooting

⁠Container Won't Start
# Check container logs
docker logs gemma-server

# Common issues:
# - Insufficient memory (need 4GB+)
# - Port 8080 already in use
# - Wrong architecture (x86_64 instead of ARM64)
⁠Slow Performance
  • Ensure you're running on ARM64 hardware (not emulated)
  • Increase memory allocation: --memory=6g
  • Reduce concurrent requests
  • Check CPU temperature (thermal throttling)
⁠Memory Issues
# Monitor memory usage
docker stats gemma-server

# Reduce memory usage by limiting context:
docker run -d -p 8080:8080 \
  yourusername/llama-gemma-server:latest \
  ./bin/llama-server -m gemma-3-270m-it-Q8_0.gguf \
  --host 0.0.0.0 --ctx-size 2048
⁠API Errors
  • 504 Gateway Timeout: Increase request timeout or reduce max_tokens
  • Memory allocation error: Container needs more RAM
  • Model loading failed: Check available disk space

⁠Build Information

⁠Optimizations
  • CPU Target: ARM Cortex-A72 with ARMv8-A+CRC instructions
  • SIMD: NEON vectorization enabled
  • Threading: OpenMP parallel processing
  • Quantization: Q8_0 format for optimal size/quality balance
  • Compiler: GCC with -O2 optimization, tuned for Cortex-A72
⁠Model Details
  • Base Model: Google Gemma 3 270M Instruct
  • Quantization: Q8_0 (8-bit quantization)
  • File Size: ~270MB
  • Context Window: 8192 tokens
  • Vocabulary: 256,000 tokens

⁠Security Considerations

  • Container runs as non-root user
  • No external network access required (air-gapped compatible)
  • Model and inference entirely local (no data sent to external services)
  • API has no authentication by default - add reverse proxy for production

⁠Production Deployment

⁠Docker Compose
version: '3.8'
services:
  gemma-server:
    image: yourusername/llama-gemma-server:latest
    ports:
      - "8080:8080"
    deploy:
      resources:
        limits:
          memory: 6G
          cpus: '4'
        reservations:
          memory: 2G
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "wget", "--quiet", "--tries=1", "--spider", "http://localhost:8080/health"]
      interval: 30s
      timeout: 10s
      retries: 3
⁠Behind Reverse Proxy (nginx)
server {
    listen 80;
    server_name your-domain.com;
    
    location / {
        proxy_pass http://localhost:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_buffering off;  # Important for streaming
    }
}

⁠Contributing

This image is built from the official llama.cpp repository with ARM64-specific optimizations. To report issues or contribute improvements, please visit the project repository.

⁠License

  • llama.cpp: MIT License
  • Gemma Model: Gemma Terms of Use (Google)
  • Container: MIT License

⁠Tags and Versions

  • latest - Latest stable build with Gemma 3 270M
  • arm64 - Explicit ARM64 tag
  • v1.0 - Stable release versions
  • dev - Development builds (may be unstable)

For production deployments, use specific version tags rather than latest.

Tag summary

Content type

Image

Digest

sha256:48aafe463…

Size

293.9 MB

Last updated

about 1 year ago

docker pull djinn/llama-gemma-server