Sign inSign up

buluma/tgdl-ml

By buluma

•Updated 5 months ago

CLIP embeddings, OCR, face detection and WD14 tagging sidecar for telegram-media-downloader

Image
0

416

buluma/tgdl-ml repository overview

⁠tgdl-ml

Machine-learning sidecar for telegram-media-downloader⁠.

Built on top of the Immich ML runtime. Provides:

  • CLIP embeddings — semantic image search (ViT-B/32 and others)
  • Face detection & recognition — clustering into people groups
  • OCR — text extraction via PaddleOCR
  • WD14 tagging — Danbooru-style image tags via SmilingWolf/wd-v1-4-convnext-tagger-v2
⁠Usage

See the main project docker-compose.yml⁠ for the full setup.

Enable with:

TGDL_ML_URL=http://tgdl-ml:3800 docker compose --profile tgdl-ml up -d
⁠Endpoints
EndpointDescription
POST /embed-imageCLIP image embedding
POST /embed-textCLIP text embedding
POST /detectFace detection
POST /detect/batchBatch face detection
POST /ocrText extraction
POST /tagCLIP-based image tagging
POST /tag-wd14WD14 Danbooru tagger
GET /healthHealth check
GET /infoCapabilities & model info

Tag summary

Content type

Image

Digest

sha256:d09566673…

Size

446.6 MB

Last updated

5 months ago

docker pull buluma/tgdl-ml