Hardware Requirements: The API server runs two instances of the 4B model simultaneously in bfloat16 (FP16): one instance in standard HuggingFace for Agent B's intuition, and one instance in vLLM for Agent A's generation. Therefore, it requires approximately 20-24 GB of VRAM (e.g., an RTX 3090, 4090, or A10G).
The repository includes a Dockerfile that allows you to run the LatentBridge OpenAI-compatible API server using vLLM. It exposes port 8000.
Run the container (requires NVIDIA Container Toolkit for --gpus all):
docker run -it --rm --gpus all -p 8000:8000 --name latent-server massimolauri/latentbridge:qwen3.5-4b
The server acts like a standard OpenAI endpoint, but with a custom agent_b_prompt parameter. This allows you to give the complex reasoning context to Agent B (the latent intuition) and only the final question to Agent A (the generator).
Example via curl:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.5-4B",
"messages": [{"role": "user", "content": "Solve the problem and give only the final answer."}],
"agent_b_prompt": "Think step-by-step. 2+2 is 4, then multiply by 3 to get 12.",
"temperature": 0.7
}'
Once running, the API will be available at http://localhost:8000/v1/chat/completions.
Content type
Image
Digest
sha256:477aa0248…
Size
11 GB
Last updated
2 months ago
docker pull massimolauri/latentbridge:qwen3.5-4b