This project dockerizes Stable Diffusion 3.5 model inference and exposes image generation API
1.7K
A Docker-containerized REST API for text-to-image generation using Stable Diffusion 3.5 Large Turbo. Provides an OpenAI-compatible /v1/images/generations endpoint for seamless integration with existing tools and workflows.
Before running this API, ensure you have:
Docker with NVIDIA Container Toolkit installed
# Verify Docker installation
docker --version
# Verify NVIDIA Container Toolkit
nvidia-smi
docker run --rm --gpus all nvidia/cuda:12.6.0-runtime-ubuntu22.04 nvidia-smi
NVIDIA GPU with CUDA support (recommended: 16GB+ VRAM for SD 3.5 Large Turbo)
Hugging Face Account with access to the Stable Diffusion 3.5 Large Turbo model
# 1. Build the Docker image
docker build -t sd-api .
# 2. Run with your Hugging Face token and API key
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=hf_your_token_here \
-e API_KEY=your_api_key_here \
-v ~/.cache/huggingface:/root/.cache/huggingface \
sd-api
# 3. Test the API
curl -X POST http://localhost:8000/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt": "A beautiful sunset over mountains", "n": 1, "size": "1024x1024"}'
A pre-built Docker image is available on Docker Hub:
docker pull blumfontein/docker-stable-diffusion
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
blumfontein/docker-stable-diffusion
Docker Hub: https://hub.docker.com/r/blumfontein/docker-stable-diffusion
docker build -t sd-api:latest .
docker build -t sd-api:v1.0.0 .
docker build --build-arg BUILDKIT_INLINE_CACHE=1 -t sd-api:latest .
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
sd-api
Cache downloaded models to avoid re-downloading on container restart:
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
-v ~/.cache/huggingface:/root/.cache/huggingface \
sd-api
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
-e MODEL_ID=stabilityai/stable-diffusion-3.5-large \
-v ~/.cache/huggingface:/root/.cache/huggingface \
sd-api
Adjust GPU memory allocation (useful for managing memory with other GPU applications):
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
-e GPU_MEMORY_UTILIZATION=0.7 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
sd-api
docker run -d --name sd-api --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e API_KEY=your_api_key_here \
-v ~/.cache/huggingface:/root/.cache/huggingface \
sd-api
# View logs
docker logs -f sd-api
# Stop container
docker stop sd-api
| Variable | Required | Default | Description |
|---|---|---|---|
HUGGING_FACE_HUB_TOKEN | Yes | - | Your Hugging Face access token for model download |
HF_TOKEN | No | - | Alternative to HUGGING_FACE_HUB_TOKEN for Hugging Face authentication |
API_KEY | No | - | API key for Bearer token authentication. If not set, authentication is disabled |
MODEL_ID | No | stabilityai/stable-diffusion-3.5-large-turbo | Hugging Face model identifier |
HOST | No | 0.0.0.0 | Server bind address |
PORT | No | 8000 | Server port |
GPU_MEMORY_UTILIZATION | No | 0.9 | GPU memory utilization (0.0 to 1.0) - percentage of available GPU memory to allocate for the model |
ENABLE_CACHE_DIT | No | true | Enable DIT caching optimization for improved performance |
GENERATION_TIMEOUT | No | 120 | Image generation timeout in seconds |
MAX_WORKERS | No | 1 | Number of concurrent generation workers |
The API supports Bearer token authentication via the Authorization header. Authentication is enabled by setting the API_KEY environment variable.
API_KEY is set: All requests to the /v1/images/generations endpoint must include a valid Authorization: Bearer <token> header.API_KEY is not set: Authentication is disabled and all requests are allowed (development mode).curl -X POST http://localhost:8000/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt": "A beautiful sunset over mountains"}'
| HTTP Status | Scenario | Message |
|---|---|---|
| 401 | Missing header | Authorization header is required |
| 401 | Invalid format | Invalid Authorization header format. Expected 'Bearer <token>' |
| 401 | Wrong token | Invalid API key |
Note: The
/healthand/docsendpoints do not require authentication.
GET /health
Check the health status of the API and model loading state.
curl http://localhost:8000/health
Response:
{
"status": "healthy",
"model_loaded": true
}
| Status | Description |
|---|---|
healthy | Model is loaded and ready to serve requests |
degraded | Service is running but model is not loaded |
POST /v1/images/generations
Generate images from a text prompt using Stable Diffusion 3.5.
Request Body:
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
prompt | string | Yes | - | Text description of the image (1-4096 characters) |
n | integer | No | 1 | Number of images to generate (1-4) |
size | string | No | 1024x1024 | Image dimensions (512x512, 768x768, 1024x1024) |
response_format | string | No | b64_json | Response format (b64_json or url) |
Example Request:
curl -X POST http://localhost:8000/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "A beautiful sunset over mountains, digital art, high quality",
"n": 1,
"size": "1024x1024",
"response_format": "b64_json"
}'
Example Response:
{
"created": 1706450400,
"data": [
{
"b64_json": "iVBORw0KGgoAAAANSUhEUgAABAAAAAQA..."
}
]
}
Save Generated Image:
# Generate and save to file
curl -s -X POST http://localhost:8000/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt": "A majestic mountain landscape"}' \
| jq -r '.data[0].b64_json' | base64 -d > output.png
GET /docs
Interactive Swagger UI documentation.
http://localhost:8000/docs
GET /redoc
ReDoc API documentation.
http://localhost:8000/redoc
The API returns OpenAI-compatible error responses:
{
"error": "Error message describing what went wrong",
"code": "error_code",
"param": null
}
| HTTP Status | Code | Description |
|---|---|---|
| 401 | authentication_error | Missing, invalid, or incorrect Bearer token |
| 400 | validation_error | Invalid request parameters |
| 500 | gpu_out_of_memory | GPU ran out of memory during generation |
| 500 | generation_failed | Image generation failed |
| 500 | internal_error | Unexpected server error |
| 503 | model_not_loaded | Model is not yet loaded (wait for startup) |
Create a docker-compose.yml:
version: '3.8'
services:
sd-api:
build: .
ports:
- "8000:8000"
environment:
- HUGGING_FACE_HUB_TOKEN=${HUGGING_FACE_HUB_TOKEN}
- API_KEY=${API_KEY}
- MODEL_ID=stabilityai/stable-diffusion-3.5-large-turbo
volumes:
- huggingface-cache:/root/.cache/huggingface
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 300s
volumes:
huggingface-cache:
Run with:
export HUGGING_FACE_HUB_TOKEN=your_token_here
export API_KEY=your_api_key_here
docker-compose up -d
import base64
import requests
# Generate an image
response = requests.post(
"http://localhost:8000/v1/images/generations",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"prompt": "A cyberpunk city at night, neon lights, rain",
"n": 1,
"size": "1024x1024"
}
)
# Save the image
data = response.json()
image_b64 = data["data"][0]["b64_json"]
image_bytes = base64.b64decode(image_b64)
with open("generated_image.png", "wb") as f:
f.write(image_bytes)
print("Image saved to generated_image.png")
This API is compatible with the OpenAI Python SDK:
from openai import OpenAI
# Point to local SD API
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="your-api-key" # Must match the API_KEY environment variable
)
response = client.images.generate(
prompt="A serene Japanese garden with cherry blossoms",
n=1,
size="1024x1024"
)
# Access the base64 image data
image_b64 = response.data[0].b64_json
Symptom: Container fails to start with authentication errors.
Solution: Verify your HUGGING_FACE_HUB_TOKEN is correct and has access to the model:
# Test token access
curl -H "Authorization: Bearer YOUR_TOKEN" \
https://huggingface.co/api/models/stabilityai/stable-diffusion-3.5-large-turbo
Symptom: Generation fails with CUDA out of memory errors.
Solutions:
GPU_MEMORY_UTILIZATION=0.7 or lower (default is 0.9)
docker run --gpus all -p 8000:8000 \
-e HUGGING_FACE_HUB_TOKEN=your_token_here \
-e GPU_MEMORY_UTILIZATION=0.7 \
sd-api
512x512 or 768x768n=1)Symptom: Health check shows model_loaded: false.
Solution: Wait for model download and loading to complete. First startup can take 5-10 minutes depending on internet speed. Monitor with:
docker logs -f sd-api
Symptom: Container fails to use GPU or falls back to CPU.
Solutions:
Verify NVIDIA Container Toolkit installation:
docker run --rm --gpus all nvidia/cuda:12.6.0-runtime-ubuntu22.04 nvidia-smi
Ensure Docker daemon has GPU support:
sudo systemctl restart docker
Check NVIDIA drivers:
nvidia-smi
.
├── Dockerfile # Docker image definition
├── requirements.txt # Python dependencies
├── README.md # This documentation
└── app/
├── __init__.py # Package initialization
├── main.py # FastAPI application and endpoints
├── models.py # Pydantic request/response models
└── generator.py # Model loading and image generation
This project provides an inference API for the Stable Diffusion 3.5 Large Turbo model. Usage of the model is subject to the Stability AI License.
Contributions are welcome! Please ensure any changes:
Content type
Image
Digest
sha256:8179a2527…
Size
6.1 GB
Last updated
7 months ago
docker pull blumfontein/docker-stable-diffusion