Sign inSign up

kylefoxaustin/openwebui-ollama

By kylefoxaustin

•Updated over 1 year ago

A containerized OpenWebUI+ollama. Run locally. CPU-only and GPU+CPU options

Image
Machine learning & AI
1

1.3K

kylefoxaustin/openwebui-ollama repository overview

⁠OpenWebUI with Ollama Docker Images

Docker images combining OpenWebUI⁠ with Ollama⁠ in a single container for a seamless AI development experience.

⁠Table of Contents

⁠Overview

These Docker images provide a combined deployment of OpenWebUI and Ollama in a single container, managed by supervisord. This approach offers several advantages over the traditional multi-container setup:

  • Simplified deployment - Only one container to manage
  • Reduced configuration complexity - No need to configure network communication between containers
  • Shared resources - More efficient resource utilization
  • Consistent state - Both applications start and stop together

The images are available in both CPU and GPU variants to suit different hardware configurations, with support for both Intel/AMD (x86_64) and ARM64 architectures.

⁠System Requirements

⁠Intel/AMD (x86_64) Requirements
⁠CPU Version
  • Architecture: x86_64 only (Intel or AMD CPUs)
  • Minimum: 4 CPU cores, 8GB RAM
  • Recommended: 8+ CPU cores, 16GB+ RAM
  • At least 10GB free disk space (more needed for models)
⁠GPU Version
  • Architecture: x86_64 only (Intel or AMD CPUs)
  • Minimum: NVIDIA GPU with 4GB VRAM, CUDA 11.7+
  • Recommended: NVIDIA GPU with 8GB+ VRAM
  • NVIDIA drivers 525.60.13 or later
  • NVIDIA Container Toolkit⁠ installed
  • At least 10GB free disk space (more needed for models)
⁠ARM64 Requirements
  • Architecture: ARM64-based device (NXP i.MX8, i.MX 9, 64bit with 64bit OS)

  • CPU Version:

    • 8GB+ RAM recommended
    • At least 10GB free disk space
  • GPU Version (requires an NVIDIA GPU):

    • At least 8GB RAM (16GB+ recommended for larger models)
    • At least 10GB free disk space
    • Note: this container will default to CPU only if a GPU is not detected

⁠Available Images

⁠Intel/AMD (x86_64) Images
  • kylefoxaustin/openwebui-ollama:latest-cpu - CPU version
  • kylefoxaustin/openwebui-ollama:latest-gpu - GPU version with NVIDIA CUDA support
⁠ARM64 Images
  • kylefoxaustin/openwebui-ollama:arm64-cpu - ARM64 CPU version (NXP i.MX 8, i.MX 9 etc)
  • kylefoxaustin/openwebui-ollama:arm64-gpu - ARM64 GPU version (ARMv8 64bit core with NVIDIA GPU)

⁠Quick Start

⁠Intel/AMD (x86_64) Quick Start
⁠CPU Version
docker run -d \
  --name openwebui \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest-cpu
⁠GPU Version
docker run -d \
  --name openwebui-gpu \
  --gpus all \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest-gpu
⁠ARM64 Quick Start
⁠CPU Version (NXP i.MX 8, i.MX 9, etc)
docker run -d \
  --name openwebui-arm \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:arm64-cpu
⁠GPU Version (ARMv8 class 64bit core with NVIDIA GPU)
docker run -d \
  --name openwebui-arm-gpu \
  --runtime nvidia \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  -v /usr/local/cuda:/usr/local/cuda \
  -e OLLAMA_HOST=0.0.0.0 \
  -e OLLAMA_NUM_PARALLEL=1 \
  -e OLLAMA_MAX_QUEUE=1 \
  kylefoxaustin/openwebui-ollama:arm64-gpu

Access the web interface at: http://localhost:8080⁠

⁠Usage Scenarios

⁠Coexisting with Local Ollama Installation

If you already have Ollama running on your host machine, you'll need to map the container's Ollama port to a different host port:

docker run -d \
  --name openwebui \
  -p 8080:8080 \
  -p 11435:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest
⁠Running CPU and GPU Containers Simultaneously

To run both CPU and GPU containers at the same time, use different port mappings:

# CPU Container
docker run -d \
  --name openwebui-cpu \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-cpu-data:/root/.ollama \
  -v openwebui-cpu-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest-cpu

# GPU Container
docker run -d \
  --name openwebui-gpu \
  --gpus all \
  -p 8081:8080 \
  -p 11435:11434 \
  -v ollama-gpu-data:/root/.ollama \
  -v openwebui-gpu-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest-gpu

Access the interfaces at:

⁠Using with External Ollama

To use OpenWebUI with an external Ollama instance (e.g., running on another server or container):

docker run -d \
  --name openwebui-only \
  -p 8080:8080 \
  -e OLLAMA_BASE_URL=http://<ollama-host>:11434 \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest

Replace <ollama-host> with the hostname or IP address of your Ollama server.

⁠Environment Variables

VariableDescriptionDefault
OLLAMA_HOSTHost for Ollama to listen on0.0.0.0
PORTPort for OpenWebUI to listen on8080
HOSTHost for OpenWebUI to listen on0.0.0.0
OLLAMA_BASE_URLURL for OpenWebUI to connect to Ollamahttp://localhost:11434
NVIDIA_VISIBLE_DEVICES(GPU only) Controls which GPUs are visibleall
NVIDIA_DRIVER_CAPABILITIES(GPU only) Required NVIDIA capabilitiescompute,utility
OLLAMA_NUM_PARALLELConcurrent request processing1
OLLAMA_MAX_QUEUEMaximum queued requests5
OLLAMA_GPU_LAYERS(Jetson only) Number of model layers to offload to GPU20

⁠Data Persistence

The following volumes are used for data persistence:

  • /root/.ollama: Ollama models and configuration
  • /app/backend/data: OpenWebUI data (conversations, settings, etc.)

For data backup, you can simply create archives of these volumes:

# Create a backup directory
mkdir -p ~/openwebui-backups

# Backup Ollama data
docker run --rm -v ollama-data:/data -v ~/openwebui-backups:/backup \
  ubuntu tar czf /backup/ollama-data-$(date +%Y%m%d).tar.gz -C /data .

# Backup OpenWebUI data
docker run --rm -v openwebui-data:/data -v ~/openwebui-backups:/backup \
  ubuntu tar czf /backup/openwebui-data-$(date +%Y%m%d).tar.gz -C /data .

⁠Troubleshooting

⁠Common Issues
  1. Port Conflicts: If you see "address already in use" errors, you likely have another service using the same port. Use alternative ports as shown in the usage scenarios.

  2. GPU not detected: Ensure your NVIDIA drivers are properly installed and the NVIDIA Container Toolkit is set up correctly. Test with:

    docker run --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi
    
  3. Container crashes: Check logs with:

    docker logs openwebui
    

    For more detailed logs:

    # Ollama logs
    docker exec -it openwebui cat /var/log/supervisor/ollama.err.log
    docker exec -it openwebui cat /var/log/supervisor/ollama.out.log
    
    # OpenWebUI logs
    docker exec -it openwebui cat /var/log/supervisor/openwebui.err.log
    docker exec -it openwebui cat /var/log/supervisor/openwebui.out.log
    
    # Supervisor logs
    docker exec -it openwebui cat /var/log/supervisor/supervisord.log
    
  4. Models not loading: The first time you pull a model might take some time. Check the Ollama logs:

    docker exec -it openwebui cat /var/log/supervisor/ollama.err.log
    

    You can directly pull models with:

    docker exec -it openwebui ollama pull <model-name>
    
  5. Web UI not accessible: Make sure that the internal Ollama instance is properly running:

    docker exec -it openwebui curl -s http://localhost:11434/api/tags
    

    Check if the OpenWebUI process is running:

    docker exec -it openwebui supervisorctl status
    
  6. Out of memory errors: Larger models require substantial RAM and VRAM. Try a smaller model or increase your container's memory limit:

    docker update --memory 16G --memory-swap 32G openwebui
    
  7. Slow model performance: For GPU containers, make sure CUDA is properly detected:

    docker exec -it openwebui-gpu nvidia-smi
    
⁠ARM64-Specific Issues
⁠ARM64 CPU Version
  1. Package Installation Failures: Some Python packages may not have ARM64 wheels available. If you encounter build errors, try modifying the requirements or building packages from source.

  2. Performance Issues: ARM CPUs are typically less powerful than x86_64 CPUs. Consider using smaller models optimized for less powerful hardware.

⁠Jetson GPU Troubleshooting
  1. GPU Not Detected: Ensure your Jetson device has the proper NVIDIA drivers installed and that you're using the --runtime nvidia flag when running the container.

  2. Internal Server Errors (HTTP 500): This often indicates that the model is overwhelming the GPU. Solutions include:

    • Reduce GPU layers: Lower the OLLAMA_GPU_LAYERS value to offload fewer layers to the GPU
    • Mount CUDA libraries: Ensure -v /usr/local/cuda:/usr/local/cuda is present
    • Limit parallelism: Use -e OLLAMA_NUM_PARALLEL=1
    • Control queue depth: Add -e OLLAMA_MAX_QUEUE=1
  3. Slow Model Loading or Timeouts: Jetson devices have limited GPU memory and bandwidth:

    • Use smaller quantized models (e.g., Llama3-8B-Q4, TinyLlama)
    • Increase timeouts with -e OLLAMA_LOAD_TIMEOUT=10m
    • For Nano, consider sticking with CPU-only mode for larger models

⁠Advanced Configuration

⁠Docker Compose

For more complex setups, you can use Docker Compose. Here's an example configuration:

version: '3.8'

services:
  openwebui:
    image: kylefoxaustin/openwebui-ollama:latest-gpu
    container_name: openwebui
    restart: unless-stopped
    ports:
      - "8080:8080"
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
      - openwebui-data:/app/backend/data
    environment:
      - OLLAMA_HOST=0.0.0.0
      - PORT=8080
      - HOST=0.0.0.0
      - OLLAMA_BASE_URL=http://localhost:11434
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  ollama-data:
  openwebui-data:

Save this to docker-compose.yml and run with:

docker-compose up -d
⁠Resource Limits

To control CPU and memory usage when running your container:

docker run -d \
  --name openwebui \
  --cpus 4 \
  --memory 8G \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest
⁠Custom Network Configuration

To place your container on a specific network:

# Create a custom network
docker network create ai-network

# Run the container on that network
docker run -d \
  --name openwebui \
  --network ai-network \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest

⁠Security Considerations

These containers are designed for development and testing purposes. If deploying in a production environment, consider the following security measures:

  1. Do not expose the container to the public internet without proper authentication and TLS encryption.

  2. Use a reverse proxy like Nginx or Traefik with proper SSL/TLS termination.

  3. Run containers with limited privileges:

    docker run -d \
      --name openwebui \
      --security-opt=no-new-privileges \
      --cap-drop=ALL \
      -p 8080:8080 \
      -p 11434:11434 \
      -v ollama-data:/root/.ollama \
      -v openwebui-data:/app/backend/data \
      kylefoxaustin/openwebui-ollama:latest
    
  4. Consider network isolation using Docker networks to limit container communication.

  5. Regularly update the images to get the latest security patches.

⁠Updating

To update to the latest version:

# Pull the latest images
docker pull kylefoxaustin/openwebui-ollama:latest
docker pull kylefoxaustin/openwebui-ollama:latest-gpu

# Restart your containers
docker stop openwebui
docker rm openwebui
docker run -d \
  --name openwebui \
  -p 8080:8080 \
  -p 11434:11434 \
  -v ollama-data:/root/.ollama \
  -v openwebui-data:/app/backend/data \
  kylefoxaustin/openwebui-ollama:latest

⁠Performance Tuning

⁠Intel/AMD Performance
⁠CPU Performance

For better CPU performance:

  1. Allocate more CPU cores:

    docker run -d --cpus 8 ... kylefoxaustin/openwebui-ollama:latest
    
  2. Enable CPU optimization:

    docker run -d --cpuset-cpus="0-7" ... kylefoxaustin/openwebui-ollama:latest
    
⁠GPU Performance

For better GPU performance:

  1. Select specific GPUs if you have multiple:

    docker run -d --gpus '"device=0,1"' ... kylefoxaustin/openwebui-ollama:latest-gpu
    
  2. Increase shared memory:

    docker run -d --shm-size=8g ... kylefoxaustin/openwebui-ollama:latest-gpu
    
  3. Optimize for specific CUDA capabilities:

    docker run -d \
      -e NVIDIA_DRIVER_CAPABILITIES=compute,utility,video \
      ... kylefoxaustin/openwebui-ollama:latest-gpu
    
⁠ARM64 Performance
⁠Optimizing for Different Jetson Models

Each Jetson platform has different capabilities requiring specific tuning:

Jetson Nano (4GB):

  • Best with CPU-only container for most models
  • For GPU usage, limit to very small models with high quantization (TinyLlama, Q4)
  • Set OLLAMA_GPU_LAYERS=5 to minimize GPU memory usage

Jetson Xavier:

  • Can handle medium-sized models with Q4 quantization
  • Set OLLAMA_GPU_LAYERS=15 for balanced performance
  • Limit to 1-2 parallel processes

Jetson Orin Nano:

  • Works well with 7B-8B class models
  • Try OLLAMA_GPU_LAYERS=20 as a starting point
  • Can handle some parallelism with -e OLLAMA_NUM_PARALLEL=2

Jetson Orin AGX:

  • Can run larger models (up to 13B with quantization)
  • Effective with OLLAMA_GPU_LAYERS=20 for stability
  • Can handle higher parallelism depending on model size

⁠Testing

To test your deployment, the repository includes a testing script that verifies container startup, connectivity, and functionality.

# Clone the repository
git clone https://github.com/kylefoxaustin/openwebui-ollama.git
cd openwebui-ollama/tools

# Make the script executable
chmod +x test_script_cpu_gpu_containers.sh

# Run the test (update username/image name as needed)
./test_script_cpu_gpu_containers.sh

This script will:

  1. Detect your system architecture (x86_64 or ARM64)
  2. Test both CPU and GPU images if available
  3. Verify container startup and service availability
  4. Test GPU accessibility (for GPU containers)
  5. Provide a comprehensive test report
⁠Tag and Push Script

If you build your own versions of these images, you can use the included tag and push script:

# Make the script executable
chmod +x tag_push.sh

# Edit the script to update your Docker Hub username
# Then run the script to tag and push your images
./tag_push.sh

⁠License

These Docker images combine OpenWebUI and Ollama, each with their respective licenses. See the original projects for more information.


Maintained by kylefoxaustin⁠

Last updated: April 2025

Tag summary

Content type

Image

Digest

sha256:ace7e542c…

Size

2.6 GB

Last updated

over 1 year ago

docker pull kylefoxaustin/openwebui-ollama:arm64-gpu