A local reranker service with a Jina compatible API.
2.8K
A local reranker service with a Jina compatible API.
This project provides a FastAPI-based web service that implements a reranking API endpoint (/v1/rerank) compatible with the Jina AI Rerank API. It allows you to host a reranking model entirely on your own infrastructure for enhanced privacy and performance.
/v1/rerank endpoint structuresentence-transformers library for PyTorch backendlifespan for resource managementPyTorch Backend:
MLX Backend (Apple Silicon only):
# Clone the repository
git clone https://github.com/olafgeibig/local-reranker.git
cd local-reranker
# Create virtual environment
uv venv
source .venv/bin/activate
# Install dependencies
uv pip install -e ".[dev]"
The CLI supports both modern subcommands and legacy arguments for backward compatibility.
# Start server with subcommand
cli serve --backend <backend_type> [options]
# Show configuration
cli config show
# Old-style arguments still work
cli --backend <backend_type> --model <model> --host <host> --port <port>
pytorch: PyTorch-based reranker (default, cross-platform)mlx: MLX-based reranker (Apple Silicon optimized)--backend: Backend type to use (default: pytorch)--model: Model name to use (overrides reranker default)--host: Host to bind server to (default: 0.0.0.0)--port: Port to bind server to (default: 8010)--log-level: Uvicorn log level (debug, info, warning, error, critical; default: info)--reload: Enable auto-reload for developmentPyTorch Backend (default):
cli serve --backend pytorch --model jinaai/jina-reranker-v2-base-multilingual
MLX Backend (Apple Silicon):
cli serve --backend mlx --model jinaai/jina-reranker-v3-mlx
Development Mode:
cli serve --backend pytorch --reload --log-level debug
Configuration Management:
cli config show
# Start with default settings
cli serve
# Start with custom settings
cli serve --backend mlx --host 0.0.0.0 --port 8080
Configuration Management:
cli config show
# Start with default settings
cli serve
# Start with custom settings
cli serve --reranker mlx --host 0.0.0.0 --port 8080
# Ensure virtual environment is active
uvicorn local_reranker.api:app --host 0.0.0.0 --port 8010 --reload
# From project root directory
uv run uvicorn local_reranker.api:app --host 0.0.0.0 --port 8010 --reload
The server will start, and the first time it runs, it will download the default reranker model:
jina-reranker-v2-base-multilingual (~1.4GB)jina-reranker-v3-mlx (~1.2GB, Apple Silicon optimized)Model download may take some time depending on your internet connection.
Once the server is running, you can send requests to the /v1/rerank endpoint. Here's an example using curl:
curl -X POST "http://localhost:8010/v1/rerank" \
-H "Content-Type: application/json" \
-d '{
"model": "jina-reranker-v2-base-multilingual",
"query": "What are the benefits of using FastAPI?",
"documents": [
"FastAPI is a modern, fast (high-performance) web framework for building APIs with Python 3.7+ based on standard Python type hints.",
"Django is a high-level Python Web framework that encourages rapid development and clean, pragmatic design.",
"The key features are: Fast, Fast to code, Fewer dependencies, Intuitive, Easy, Short, Robust, Standards-based.",
"Flask is a micro web framework written in Python."
],
"top_n": 3,
"return_documents": true
}'
model: (Currently ignored by API, uses the configured default) The name of the reranker modelquery: The search query stringdocuments: A list of strings or dictionaries ({"text": "..."}) to be reranked against the querytop_n: (Optional) The maximum number of results to returnreturn_documents: (Optional, default false) Whether to include document text in resultsClone the repository:
git clone https://github.com/olafgeibig/local-reranker.git
cd local-reranker
Create a virtual environment:
uv venv
source .venv/bin/activate
Install development dependencies:
uv pip install -e ".[dev]"
Verify installation:
# Test CLI works
cli config show
# Test server starts
cli serve --backend pytorch --help
Tests are implemented using pytest. To run tests:
# Ensure virtual environment is active
python -m pytest
# Or using uv run
uv run pytest
# Run specific test categories
uv run pytest -m "not integration" # Skip integration tests
uv run pytest -m "integration" # Only integration tests
uv run pytest -m "slow" # Only slow tests
The project uses modern development tools:
# Run linting
uv run ruff check
# Run type checking
uv run mypy src/
# Run both
uv run ruff check && uv run mypy src/
MLX not found:
# Ensure you're on Apple Silicon
uname -m # Should show arm64
# Install MLX dependencies
uv add mlx mlx-lm safetensors
Model download fails:
# Check internet connection
# Try manual download
huggingface-cli download jinaai/jina-reranker-v3-mlx
Performance issues:
# Check MLX is using GPU (if available)
python -c "import mlx; print(mlx.metal.is_available())"
# Monitor memory usage
top -o mem | grep python
If you're having trouble with MLX backend configuration, try these explicit CLI commands:
# Force MLX backend with explicit model
cli serve --backend mlx --model jinaai/jina-reranker-v3-mlx
# Check current configuration
cli config show
# Use development mode for debugging
cli serve --backend mlx --reload --log-level debug
Model download fails:
# Check internet connection
# Try manual download
huggingface-cli download jinaai/jina-reranker-v3-mlx
Performance issues:
# Check MLX is using GPU (if available)
python -c "import mlx; print(mlx.metal.is_available())"
# Monitor memory usage
top -o mem | grep python
The application uses pydantic-settings for configuration management. You can set the following environment variables to override defaults:
# Force MLX backend
export RERANKER_RERANKER_TYPE=mlx
# Custom model name
export RERANKER_MODEL_NAME=custom-mlx-model
# Custom host and port
export RERANKER_HOST=0.0.0.0
export RERANKER_PORT=8080
# Enable debug logging
export RERANKER_LOG_LEVEL=debug
# Enable auto-reload
export RERANKER_RELOAD=true
Note: Using the CLI command line options is recommended over environment variables for clarity.
local-reranker/
├── src/local_reranker/
│ ├── __init__.py
│ ├── api.py # FastAPI application
│ ├── cli.py # Command line interface
│ ├── config.py # Configuration management
│ ├── models.py # Pydantic models
│ ├── reranker.py # Base reranker interface
│ ├── reranker_pytorch.py # PyTorch implementation
│ ├── reranker_mlx.py # MLX implementation
│ └── utils.py # Utility functions
├── tests/ # Test suite
├── pyproject.toml # Project configuration
└── README.md # This file
git checkout -b feature/amazing-feature)git commit -m 'Add amazing feature')git push origin feature/amazing-feature)This project is licensed under the MIT License - see the LICENSE file for details.
Content type
Image
Digest
sha256:3fcf51bce…
Size
5.3 GB
Last updated
30 days ago
docker pull dheaps/local-reranker