Open-source RAG prompt compression middleware. Cut token costs ~50% with <3% accuracy loss.
8.7K
Open-source prompt compression middleware for RAG pipelines.
Sits between your vector DB and LLM - compresses retrieved chunks using LLMLingua-2, cutting token count by ~50% with less than 3% accuracy loss.
docker run -p 8000:8000 itsaryanchauhan/winnow
POST /v1/compress → compress a single context POST /v1/compress/batch → compress multiple contexts GET /health → health check
GitHub: https://github.com/itsaryanchauhan/Winnow Docs: https://itsaryanchauhan-winnow.hf.space/docs
Content type
Image
Digest
sha256:327967aae…
Size
427.6 MB
Last updated
7 months ago
docker pull itsaryanchauhan/winnow