A simple cache proxy of OpenAI compatible API for both streaming or not
3.4K
LLM Cache Proxy is a FastAPI-based application that serves as a caching layer for OpenAI's API. It intercepts requests to the OpenAI API, caches responses, and serves cached responses for identical requests, potentially reducing API costs and improving response times.
# Run the container
docker run -p 9999:9999 \
-v $(pwd)/data:/app/data \
so2liu/llm-cache-proxy
OR
# Run the container
docker run -p 9999:9999 \
-e OPENAI_API_KEY=your_api_key_here \
-e OPENAI_BASE_URL=https://api.openai.com/v1 \
-e VERBOSE=false
-v $(pwd)/data:/app/data \
so2liu/llm-cache-proxy
The proxy is now available at http://localhost:9999
Use /cache/chat/completions for cached requests
Use /chat/completions for disable caching requests
More detail: https://github.com/so2liu/llm-cache-server
Content type
Image
Digest
sha256:a844f271f…
Size
150.3 MB
Last updated
12 months ago
docker pull so2liu/llm-cache-proxy