Sign inSign up

mlachgar/vision-forge-api

By mlachgar

•Updated 6 months ago

Image
0

3.2K

mlachgar/vision-forge-api repository overview

⁠Vision Forge API

Vision Forge API is a FastAPI service for image tagging powered by SigLIP. It scores uploaded images against config-driven canonical tag sets and prediction profiles, with bearer API key authentication and a small admin surface for key management and config reloads.

⁠Features

  • FastAPI + Uvicorn service
  • SigLIP image/text scoring
  • Config-driven tag sets, profiles, and prompt templates
  • Bearer API key auth with admin and predict roles
  • /predict scoring with cached text embeddings
  • Prompt-level reranking for canonical tags
  • Light set-balancing when a profile spans multiple tag sets
  • Admin endpoints for API key CRUD and live config reload
  • Docker image variants for CPU and GPU builds

⁠Runtime Layout

The container uses two directories:

⁠/config

Configuration directory. The image ships with sample config files here at build time, and the path can be overridden with VISION_FORGE_CONFIG_DIR.

Expected files:

  • auth.yaml
  • settings.yaml
  • tag_sets.yaml
  • profiles.yaml
  • prompts.yaml

Typical responsibilities:

  • auth.yaml defines token prefix, token length, and default roles
  • settings.yaml defines the app name, prediction limits, embedding location, model cache location, and SigLIP model id
  • tag_sets.yaml defines canonical tag groups
  • profiles.yaml defines prediction profiles and which tag sets they use
  • prompts.yaml defines prompt templates for canonical tags
⁠/data

Writable runtime storage. The path can be overridden with VISION_FORGE_DATA_DIR.

Expected content:

  • api_keys.json - persisted API keys and roles used by the auth cache
  • embeddings/text_embeddings.json - cached text embeddings for canonical tags
  • embeddings/metadata.json - embedding cache metadata
  • model_cache/ - Hugging Face / Transformers model cache

The service will create or refresh missing embedding cache entries on startup when needed.

⁠API Surface

  • GET /health
  • GET /tag-sets
  • GET /profiles
  • POST /predict
  • GET /admin/api-keys
  • POST /admin/api-keys
  • PATCH /admin/api-keys/{name}
  • DELETE /admin/api-keys/{name}
  • POST /admin/reload

The prediction endpoint expects a multipart image upload and supports these query parameters:

  • limit
  • min_score
  • profile
  • tag_sets
  • extra_tags

Returned scores are normalized to the 0.0..1.0 range.

⁠Prediction Behavior

The prediction pipeline does the following:

  1. Encodes the image once with SigLIP
  2. Scores candidate canonical tags and any extra tags
  3. Reranks the strongest canonical candidates using the best matching prompt similarity
  4. Filters by min_score
  5. Sorts by score descending
  6. Applies the requested limit
  7. Balances results lightly across tag sets when a profile spans multiple sets

⁠Docker Images

Published variants:

  • cpu-lite
  • cpu-full
  • gpu-lite
  • gpu-full

Release builds publish both floating variant tags and versioned tags. For example, a release such as v1.2.3 publishes:

  • 1.2.3-cpu-lite
  • 1.2.3-cpu-full
  • 1.2.3-gpu-lite
  • 1.2.3-gpu-full
  • cpu-lite
  • cpu-full
  • gpu-lite
  • gpu-full
  • latest for cpu-full only

⁠Usage

Run the container with runtime data mounted at /data:

docker run --rm -it \
  -p 8000:8000 \
  -v "$PWD/data:/data" \
  -e VISION_FORGE_DEVICE=cpu \
  mlachgar/vision-forge-api:cpu-full

If you want to override the bundled config, mount your own directory at /config and point VISION_FORGE_CONFIG_DIR to it:

docker run --rm -it \
  -p 8000:8000 \
  -v "$PWD/config:/config:ro" \
  -v "$PWD/data:/data" \
  -e VISION_FORGE_CONFIG_DIR=/config \
  -e VISION_FORGE_DEVICE=cpu \
  mlachgar/vision-forge-api:cpu-full

Tag summary

Content type

Image

Digest

sha256:33520a352…

Size

4.7 GB

Last updated

6 months ago

docker pull mlachgar/vision-forge-api