HuggingFace Embeddings microservice in python (AMD GPUs)
2.2K
HuggingFace embeddings microservice exposing both Ollama-compatible and OpenAI-compatible REST APIs. Uses HuggingFaceEmbeddings under the hood with automatic GPU detection.
/api/embed, /api/embeddings, /api/tags, /api/version)/v1/embeddings, /v1/models)/gpu) — NVIDIA (nvidia-ml-py), AMD (rocm-smi), Apple Silicon (system_profiler)Requires Python 3.11+ — Follow the guide for your platform and hardware:
| Platform | Hardware | Guide |
|---|---|---|
| Windows | CPU only | guides/Windows_CPU.md |
| Windows | NVIDIA CUDA | guides/Windows_CUDA.md |
| Linux | CPU only | guides/Linux_CPU.md |
| Linux | NVIDIA CUDA | guides/Linux_CUDA.md |
| Linux | AMD ROCm | guides/Linux_ROCM.md |
| macOS | CPU only | guides/macOS_CPU.md |
| macOS | Apple Silicon (MPS) | guides/macOS_MPS.md |
Each guide includes setup, verification, and validation steps.
Pick the image matching your hardware:
| Hardware | Compose File | Image |
|---|---|---|
| CPU only | compose-cpu.yaml | jchristn77/hfembed-cpu |
| NVIDIA CUDA | compose-cuda.yaml | jchristn77/hfembed-cuda |
| AMD ROCm | compose-rocm.yaml | jchristn77/hfembed-rocm |
Note: Apple Silicon GPU (MPS) is not available in Docker. Docker Desktop runs a Linux VM without Metal passthrough. Use compose-cpu.yaml on macOS, or run locally for MPS acceleration.
# Start (pick matching compose file)
docker compose -f compose-cpu.yaml up -d
docker compose -f compose-cuda.yaml up -d
docker compose -f compose-rocm.yaml up -d
Build locally (optional):
build-docker-cpu.bat v1.0.0
build-docker-cuda.bat v1.0.0
build-docker-rocm.bat v1.0.0
The service listens on port 8000 by default.
| Environment Variable | Default | Description |
|---|---|---|
HFEMBED_PORT | 8000 | Port to listen on (can also use -p / --port CLI argument) |
EMBEDDINGS_PRELOAD_MODELS | (none) | Space-delimited model names to preload at startup |
DEFAULT_MODEL | all-MiniLM-L6-v2 | Default model when none specified in request |
HF_TOKEN | (none) | HuggingFace API token for private models |
HF_HOME | (system default) | Directory for caching downloaded models |
CORS_ORIGINS | * | Comma-separated allowed CORS origins |
REQUEST_TIMEOUT | 300 | Request timeout in seconds |
After starting the Docker container, run these commands to verify:
# Connectivity
curl http://localhost:8000/
# GPU/device info
curl http://localhost:8000/gpu
# Single embedding
curl -X POST http://localhost:8000/api/embed \
-H "Content-Type: application/json" \
-d '{"model":"all-MiniLM-L6-v2","input":"Hello world"}'
# Batch embeddings
curl -X POST http://localhost:8000/api/embed \
-H "Content-Type: application/json" \
-d '{"model":"all-MiniLM-L6-v2","input":["Hello world","Goodbye world"]}'
curl.exe http://localhost:8000/
curl.exe http://localhost:8000/gpu
curl.exe -X POST http://localhost:8000/api/embed -H "Content-Type: application/json" -d "{\"model\":\"all-MiniLM-L6-v2\",\"input\":\"Hello world\"}"
curl.exe -X POST http://localhost:8000/api/embed -H "Content-Type: application/json" -d "{\"model\":\"all-MiniLM-L6-v2\",\"input\":[\"Hello world\",\"Goodbye world\"]}"
# Generate embeddings (batch)
curl -X POST http://localhost:8000/api/embed \
-H "Content-Type: application/json" \
-d '{"model": "all-MiniLM-L6-v2", "input": ["Hello world", "Goodbye world"]}'
# Generate embedding (single, legacy)
curl -X POST http://localhost:8000/api/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "all-MiniLM-L6-v2", "prompt": "Hello world"}'
# List models
curl http://localhost:8000/api/tags
# Generate embeddings
curl -X POST http://localhost:8000/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-hf-token" \
-d '{"model": "all-MiniLM-L6-v2", "input": ["Hello world", "Goodbye world"]}'
# List models
curl http://localhost:8000/v1/models
See REST_API.md for complete API documentation.
Content type
Image
Digest
sha256:5c1bfd14c…
Size
5.2 GB
Last updated
8 months ago
docker pull jchristn77/hfembed-rocm