Sign inSign up

jchristn77/hfembed-cuda

By jchristn77

•Updated 8 months ago

HuggingFace Embeddings microservice in python (NVIDIA and CUDA)

Image
Machine learning & AI
Data science
0

2.0K

jchristn77/hfembed-cuda repository overview

⁠hfembed

HuggingFace embeddings microservice exposing both Ollama-compatible and OpenAI-compatible REST APIs. Uses HuggingFaceEmbeddings⁠ under the hood with automatic GPU detection.

⁠Features

  • Ollama-compatible API (/api/embed, /api/embeddings, /api/tags, /api/version)
  • OpenAI-compatible API (/v1/embeddings, /v1/models)
  • Automatic GPU detection: CUDA (NVIDIA), ROCm (AMD), MPS (Apple Silicon), CPU fallback
  • Comprehensive live GPU metrics (/gpu) — NVIDIA (nvidia-ml-py), AMD (rocm-smi), Apple Silicon (system_profiler)
  • System metrics (CPU, RAM) via psutil
  • Model preloading at startup
  • Configurable via environment variables
  • Docker images for CPU, CUDA (NVIDIA), and ROCm (AMD)
  • Persistent model cache via volume mount

⁠Quick Start

⁠Option 1: Run Locally (venv)

Requires Python 3.11+ — Follow the guide for your platform and hardware:

PlatformHardwareGuide
WindowsCPU onlyguides/Windows_CPU.md⁠
WindowsNVIDIA CUDAguides/Windows_CUDA.md⁠
LinuxCPU onlyguides/Linux_CPU.md⁠
LinuxNVIDIA CUDAguides/Linux_CUDA.md⁠
LinuxAMD ROCmguides/Linux_ROCM.md⁠
macOSCPU onlyguides/macOS_CPU.md⁠
macOSApple Silicon (MPS)guides/macOS_MPS.md⁠

Each guide includes setup, verification, and validation steps.

⁠Option 2: Run with Docker

Pick the image matching your hardware:

HardwareCompose FileImage
CPU onlycompose-cpu.yamljchristn77/hfembed-cpu
NVIDIA CUDAcompose-cuda.yamljchristn77/hfembed-cuda
AMD ROCmcompose-rocm.yamljchristn77/hfembed-rocm

Note: Apple Silicon GPU (MPS) is not available in Docker. Docker Desktop runs a Linux VM without Metal passthrough. Use compose-cpu.yaml on macOS, or run locally for MPS acceleration.

# Start (pick matching compose file)
docker compose -f compose-cpu.yaml up -d
docker compose -f compose-cuda.yaml up -d
docker compose -f compose-rocm.yaml up -d

Build locally (optional):

build-docker-cpu.bat v1.0.0
build-docker-cuda.bat v1.0.0
build-docker-rocm.bat v1.0.0

The service listens on port 8000 by default.

⁠Configuration

Environment VariableDefaultDescription
HFEMBED_PORT8000Port to listen on (can also use -p / --port CLI argument)
EMBEDDINGS_PRELOAD_MODELS(none)Space-delimited model names to preload at startup
DEFAULT_MODELall-MiniLM-L6-v2Default model when none specified in request
HF_TOKEN(none)HuggingFace API token for private models
HF_HOME(system default)Directory for caching downloaded models
CORS_ORIGINS*Comma-separated allowed CORS origins
REQUEST_TIMEOUT300Request timeout in seconds

⁠Validate Setup (Docker)

After starting the Docker container, run these commands to verify:

⁠Linux/macOS
# Connectivity
curl http://localhost:8000/

# GPU/device info
curl http://localhost:8000/gpu

# Single embedding
curl -X POST http://localhost:8000/api/embed \
  -H "Content-Type: application/json" \
  -d '{"model":"all-MiniLM-L6-v2","input":"Hello world"}'

# Batch embeddings
curl -X POST http://localhost:8000/api/embed \
  -H "Content-Type: application/json" \
  -d '{"model":"all-MiniLM-L6-v2","input":["Hello world","Goodbye world"]}'
⁠Windows
curl.exe http://localhost:8000/
curl.exe http://localhost:8000/gpu
curl.exe -X POST http://localhost:8000/api/embed -H "Content-Type: application/json" -d "{\"model\":\"all-MiniLM-L6-v2\",\"input\":\"Hello world\"}"
curl.exe -X POST http://localhost:8000/api/embed -H "Content-Type: application/json" -d "{\"model\":\"all-MiniLM-L6-v2\",\"input\":[\"Hello world\",\"Goodbye world\"]}"

⁠Usage Examples

⁠Ollama-compatible
# Generate embeddings (batch)
curl -X POST http://localhost:8000/api/embed \
  -H "Content-Type: application/json" \
  -d '{"model": "all-MiniLM-L6-v2", "input": ["Hello world", "Goodbye world"]}'

# Generate embedding (single, legacy)
curl -X POST http://localhost:8000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model": "all-MiniLM-L6-v2", "prompt": "Hello world"}'

# List models
curl http://localhost:8000/api/tags
⁠OpenAI-compatible
# Generate embeddings
curl -X POST http://localhost:8000/v1/embeddings \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-hf-token" \
  -d '{"model": "all-MiniLM-L6-v2", "input": ["Hello world", "Goodbye world"]}'

# List models
curl http://localhost:8000/v1/models

⁠API Reference

See REST_API.md⁠ for complete API documentation.

Tag summary

Content type

Image

Digest

sha256:d39d20144…

Size

4.4 GB

Last updated

8 months ago

docker pull jchristn77/hfembed-cuda