GPU-ready base image for AI apps with Diffusers, PyTorch+CUDA 1.3, Gradio/Flask, and OpenCV.
1.1K
A minimal, GPU-ready base image for building AI apps with 🤗 Diffusers, PyTorch+CUDA, Gradio/Flask, and OpenCV.
Ships with a preconfigured Python venv at /opt/venv, small native deps, and tini as PID 1.
Defaults to sleep infinity so you can exec in during development—or override CMD in derived images.
diffusers/diffusers-pytorch-cuda@sha256:e706cb... (PyTorch + CUDA userland)/opt/venv (on PATH)libglib2.0-0, libgl1, ffmpeg, tini/opt/venv)
diffusers, transformersflask, flask-cors, requestsopencv-python-headlessgradioEXPOSE 5000tini for clean signal handlingsleep infinity (easy to override)docker build -t uaisoftwareinc/ai-cuda:latest .
nvidia-smi works on hostYou can run the container and bind a mount to work inside of it.
docker run --gpus all -p 5000:5000 -v ~/.cache/huggingface:/root/.cache/huggingface -v $(pwd):/data uaisoftwareinc/ai-cuda:latest
Create an app image that extends this base and sets a real runtime CMD:
# Dockerfile.app
FROM uaisoftwareinc/ai-cuda:latest
WORKDIR /data
# Optional: cache layer for deps
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Your code
COPY . .
EXPOSE 5000
# Flask example (make sure your app binds 0.0.0.0:5000)
CMD ["python", "app.py"]
Build & run:
docker build -f Dockerfile.app -t my-app:latest .
docker run --rm -it --gpus all -p 5000:5000 my-app:latest
Tip: if you prefer one-file simplicity, you can directly replace the CMD ["sleep","infinity"] with your app command in the base Dockerfile—but keeping a dedicated base image makes extension cleaner.
app.pyfrom flask import Flask, jsonify
import torch, diffusers
app = Flask(__name__)
@app.get("/health")
def health():
return jsonify({
"cuda": torch.cuda.is_available(),
"devices": torch.cuda.device_count()
})
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)
import gradio as gr
import torch
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained("runwayml/stable-diffusion-v1-5").to("cuda")
def infer(prompt):
return pipe(prompt, num_inference_steps=20).images[0]
gr.Interface(fn=infer, inputs="text", outputs="image").launch(server_name="0.0.0.0", server_port=5000)
Run with:
docker run --rm -it --gpus all -p 5000:5000 my-app:latest
Add more Python libs (in a derived Dockerfile):
FROM uaisoftwareinc/ai-cuda:latest
RUN pip install --no-cache-dir accelerate xformers safetensors
Pin versions in requirements.txt for reproducibility.
Cache models by mounting Hugging Face cache:
-v ~/.cache/huggingface:/root/.cache/huggingface
Alternate CMDs for dev vs prod:
sleep infinity and docker exec into the container.CMD ["python", "app.py"] (or your server command).services:
dev:
image: uaisoftwareinc/ai-cuda:latest
command: ["sleep", "infinity"]
ports: ["5000:5000"]
volumes:
- ~/.cache/huggingface:/root/.cache/huggingface
device_requests:
- driver: nvidia
count: all
capabilities: [gpu]
tty: true
stdin_open: true
--gpus all (or compose device_requests) and that nvidia-smi works on the host.0.0.0.0, not 127.0.0.1.opencv-python-headless to keep it lean. Only switch to opencv-python if you need GUI windows.tini is already set as ENTRYPOINT for clean shutdowns./opt/venv ensures Python deps are isolated and easy to extend.sleep infinity) for painless dev, but is trivial to turn into a runnable app by overriding CMD.Content type
Image
Digest
sha256:22da63864…
Size
5.7 GB
Last updated
10 months ago
docker pull uaisoftwareinc/ai-cuda