Sign inSign up

massimolauri/latentbridge

By massimolauri

•Updated 2 months ago

Vllm Latent Bridge Server

Image
0

38

massimolauri/latentbridge repository overview

⁠Running the API Server via Docker

Hardware Requirements: The API server runs two instances of the 4B model simultaneously in bfloat16 (FP16): one instance in standard HuggingFace for Agent B's intuition, and one instance in vLLM for Agent A's generation. Therefore, it requires approximately 20-24 GB of VRAM (e.g., an RTX 3090, 4090, or A10G).

The repository includes a Dockerfile that allows you to run the LatentBridge OpenAI-compatible API server using vLLM. It exposes port 8000.

Run the container (requires NVIDIA Container Toolkit for --gpus all):

docker run -it --rm --gpus all -p 8000:8000 --name latent-server massimolauri/latentbridge:qwen3.5-4b
⁠Using the Multi-Agent Prompt

The server acts like a standard OpenAI endpoint, but with a custom agent_b_prompt parameter. This allows you to give the complex reasoning context to Agent B (the latent intuition) and only the final question to Agent A (the generator).

Example via curl:

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.5-4B",
    "messages": [{"role": "user", "content": "Solve the problem and give only the final answer."}],
    "agent_b_prompt": "Think step-by-step. 2+2 is 4, then multiply by 3 to get 12.",
    "temperature": 0.7
  }'

Once running, the API will be available at http://localhost:8000/v1/chat/completions.

Tag summary

Content type

Image

Digest

sha256:477aa0248…

Size

11 GB

Last updated

2 months ago

docker pull massimolauri/latentbridge:qwen3.5-4b