A lightweight, production-ready FastAPI microservice that serves a CTranslate2 int8-quantized BGE-small embedding model. Send it text, get back a 384-dimensional, L2-normalized embedding vector.
The model (~33MB) is baked into the image — no downloads at runtime, no external dependencies. Tokenization uses the lightweight tokenizers library (no PyTorch / Transformers), keeping the image small (~600MB) and cold starts fast (~2s).
docker run -p 8080:8080 dixisouls/embedding-service:latest
The service is ready when GET /health/ready returns 200. Interactive API docs are at http://localhost:8080/docs.
| Variable | Default | Description |
|---|---|---|
PORT | 8080 | Listen port |
HOST | 0.0.0.0 | Bind address |
LOG_LEVEL | INFO | Log level (loguru) |
MAX_LENGTH | 512 | Tokenizer truncation length |
MAX_BATCH_SIZE | 256 | Max texts per /embed/batch request |
INTER_THREADS | 1 | CTranslate2 inter-op threads |
INTRA_THREADS | 0 | CTranslate2 intra-op threads (0 = auto) |
docker run -p 9000:9000 -e PORT=9000 -e MAX_BATCH_SIZE=512 dixisouls/embedding-service:latest
| Method | Path | Description |
|---|---|---|
| GET | / | Service banner |
| GET | /health/live | Liveness — process is up |
| GET | /health/ready | Readiness — model loaded (503 until ready) |
| GET | /info (/model_info) | Model metadata |
| POST | /embed | Single text → embedding |
| POST | /embed/batch | List of texts → embeddings (single forward pass) |
| GET | /docs | Interactive Swagger UI |
GET /infoReturns model metadata.
{
"model": "bge_ct2_int8",
"backend": "ctranslate2",
"quantization": "int8",
"dim": 384,
"max_length": 512,
"max_batch_size": 256,
"model_size_mb": 33.21
}
POST /embedEmbed a single string.
Request body
{ "text": "Introduction to Machine Learning" }
Response (200)
{
"embedding": [0.0123, -0.0456, "... 384 floats total ..."],
"dim": 384,
"latency_ms": 4.9
}
Errors: 422 if text is empty, 503 if the model is not yet loaded.
POST /embed/batchEmbed many strings in a single forward pass. Each sequence is mean-pooled over its real token length (padding excluded) and L2-normalized. Output order matches input order.
Request body
{ "texts": ["Introduction to Machine Learning", "History of Ancient Rome"] }
Response (200)
{
"embeddings": [[0.0123, "..."], [-0.0078, "..."]],
"count": 2,
"dim": 384,
"latency_ms": 7.4
}
Errors: 422 if the list is empty or contains blank strings, 413 if it exceeds
MAX_BATCH_SIZE, 503 if the model is not yet loaded.
curl# 1. Start the container
docker run -d --name embedder -p 8080:8080 dixisouls/embedding-service:latest
# 2. Wait for readiness
until curl -sf localhost:8080/health/ready >/dev/null; do sleep 1; done
# 3. Inspect the model
curl -s localhost:8080/info
# 4. Single embedding
curl -s -X POST localhost:8080/embed \
-H 'content-type: application/json' \
-d '{"text":"Introduction to Machine Learning"}'
# 5. Batch embedding
curl -s -X POST localhost:8080/embed/batch \
-H 'content-type: application/json' \
-d '{"texts":["Introduction to Machine Learning","History of Ancient Rome","Organic Chemistry"]}'
# 6. Stop and remove
docker rm -f embedder
Content type
Image
Digest
sha256:faeb5af39…
Size
159 MB
Last updated
4 months ago
docker pull dixisouls/embedding-service