WIP: Bonsai 2 27B (Prism ML's ternary Qwen3.8-27B) on vLLM, for RTX PRO 6000
247
Work in progress. Expect rough edges and breaking changes.
Serves Bonsai 2 27B, Prism ML's ternary Qwen3.8-27B, on vLLM with custom CUDA kernels. Unofficial; not affiliated with Prism ML. Built and tested for the RTX PRO 6000 Blackwell only.
docker run --rm --gpus all --ipc=host -p 8000:8000 -v bonsai:/cache fraserpricee/bonsai-vllm:20260918
Source, details and issues: https://github.com/fraserprice/bonsai-vllm
Weights: https://huggingface.co/fraserprice/Ternary-Bonsai-2-27B-vllm
Apache 2.0. Created using Bonsai by Prism ML; built from Qwen3.8-27B; image built on vLLM.
Content type
Image
Digest
sha256:a40d14f5e…
Size
8.8 GB
Last updated
11 days ago
docker pull fraserpricee/bonsai-vllm:20260918