Sign inSign up

visitsb/ollama-ipex

By visitsb

•Updated over 2 years ago

Run Ollama on Intel GPU

Image
Machine learning & AI
Data science
0

10K+

visitsb/ollama-ipex repository overview

⁠ollama-ipex - Run Ollama on Intel GPU

The Ollama Docker image⁠ provides NVIDIA support⁠, but no support for Intel GPUs.

To support Intel GPUs, this docker image bundles-

  • IPEX-LLM⁠ (a PyTorch library for running LLM on Intel CPU and GPU e.g. local PC with iGPU, discrete GPU such as Arc, Flex and Max with very low latency)
  • Ollama⁠
⁠Start Ollama-IPEX server

Similar to standard Ollama Docker image⁠ a default ollama server will spin up on port 11434.

docker run --rm --tty --interactive \
           --env SYCL_PI_LEVEL_ZERO_USE_IMMEDIATE_COMMANDLISTS=1 \
           --env SYCL_CACHE_PERSISTENT=1 \
           --device /dev/dri/card0 \
           --device /dev/dri/renderD128 \
           --memory="32g" \
           --shm-size="16g" \
           --publish 11434:11434 \
           --name ollama-ipex \
        visitsb/ollama-ipex:latest

Alternatively, you can also use docker-compose.yml shown below. For additional environment specific variables, please refer to runtime configurations 1⁠, 2⁠ for Intel Arc™ A-Series Graphics and Intel Data Center GPU Flex, Intel Data Center GPU Max Series and 3⁠ for Ollama.

services:
  ollama-ipex:
    container_name: ollama-ipex
    image: visitsb/ollama-ipex:latest
    ports:
      - "11434:11434/tcp" # Ollama API
    volumes:
      - /models:/root/.models:rw
    cpuset: "0-3"
    shm_size: "16G"
    deploy:
      resources:
        limits:
          cpus: '1.00'
          memory: 32G
        reservations:
          cpus: '0.50'
          memory: 8G
    environment:
      # Make sure all layers of your model are running on Intel GPU
      - OLLAMA_NUM_GPU=999
      - OLLAMA_KEEP_ALIVE=-1
      - OLLAMA_DEBUG=1
      # Use total memory as free memory
      - ZES_ENABLE_SYSMAN=1
      # Use GPU acceleration
      - SYCL_CACHE_PERSISTENT=1
      # For environment variables specific to
      # Intel Arc™ A-Series Graphics and Intel Data Center GPU Flex and 
      # Intel Data Center GPU Max Series, please refer to links for 
      # Runtime configurations.
    devices:
      - "/dev/dri/card0:/dev/dri/card0"
      - "/dev/dri/renderD128:/dev/dri/renderD128"
⁠Pull models

In another window, you can interact with ollama now-

# Pull an Ollama model locally from https://ollama.com/library
docker exec --tty --interactive ollama-ipex /usr/local/bin/ollama pull llama3

# List Ollama models
docker exec --tty --interactive ollama-ipex /usr/local/bin/ollama list
NAME            ID              SIZE    MODIFIED
wizardlm2:7b    c9b1aff820f2    4.1 GB  About an hour ago
gemma:7b        a72c7f4d0a15    5.0 GB  2 hours ago
llama3:latest   365c0bd3c000    4.7 GB  3 hours ago
mistral:latest  61e88e884507    4.1 GB  3 hours ago
phi3:latest     a2c89ceaed85    2.3 GB  4 hours ago
⁠Use models interactively
docker exec --tty --interactive ollama /usr/local/bin/ollama run phi3
>>> /show info
Model details:
Family              llama
Parameter Size      4B
Quantization Level  Q4_K_M

>>> how much is 2 plus 2?
2 plus 2 equals 4. This is a basic arithmetic addition problem. The number 2 added to another 2 gives you the sum of 4.

>>> Send a message (/? for help)
⁠Troubleshoot and verify if Intel GPU is being utilized

Use intel_gpu_top⁠ to check GPU utilization. Anytime an ollama model is loaded and questions are being answered, you should see stats updated in your Intel GPU utlization. In my case, using standard Ollama Docker image⁠ results in 300-400% CPU, 0% GPU utilization when running llama3. In contrast, when using this container with Intel UHD 630 Graphics a 100% CPU, 98.95% GPU utilization was observed.

For troubleshooting, you should begin by checking whether your host machine correctly detects your Intel GPU. If your server supports iGPU, you will see Kernel modules: i915, otherwise the message will not be there if your Intel GPU isn't correctly detected by your host OS.

# lspci -v -s $(lspci | grep VGA | cut -d" " -f 1 | head -n1)
00:02.0 VGA compatible controller: Intel Corporation CoffeeLake-H GT2 [UHD Graphics 630] (prog-if 00 [VGA controller])
        Subsystem: Apple Inc. CoffeeLake-H GT2 [UHD Graphics 630]
        Flags: bus master, fast devsel, latency 0, IRQ 78
        Memory at 80000000 (64-bit, non-prefetchable) [size=16M]
        Memory at a0000000 (64-bit, prefetchable) [size=256M]
        I/O ports at 3000 [size=64]
        Expansion ROM at 000c0000 [virtual] [disabled] [size=128K]
        Capabilities: [40] Vendor Specific Information: Len=0c <?>
        Capabilities: [70] Express Root Complex Integrated Endpoint, MSI 00
        Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable- 64bit-
        Capabilities: [d0] Power Management version 2
        Capabilities: [100] Process Address Space ID (PASID)
        Capabilities: [200] Address Translation Service (ATS)
        Capabilities: [300] Page Request Interface (PRI)
        Kernel driver in use: i915
        Kernel modules: i915

Additionally ls -la /dev/dri should show an output similar to one below-

# ls -la /dev/dri
total 0
drwxr-xr-x  3 root root        100 May 20 01:43 .
drwxr-xr-x 18 root root       4120 May 21 01:17 ..
drwxr-xr-x  2 root root         80 May 20 01:43 by-path
crw-rw----  1 root video  226,   0 May 20 01:43 card0
crw-rw----  1 root render 226, 128 May 20 01:43 renderD12

Lastly, the environment variable OLLAMA_DEBUG=1 also shows additional logs when running the container. You can use the logs to see messages about SYCL devices to verify whether your Intel GPU is being detected correctly.

⁠Further Reference

Below sources were used as reference when creating this docker image.

Tag summary

Content type

Image

Digest

sha256:431c30954…

Size

11.3 GB

Last updated

over 2 years ago

docker pull visitsb/ollama-ipex