CUDA llama.cpp server with TurboQuant KV cache and speculative decoding support.
8.0K
Docker image for llama-server built from the AtomicBot-ai fork of llama.cpp with support for TurboQuant KV cache and speculative decoding.
This image was built for CUDA 13 / NVIDIA GPUs and is used as an OpenAI-compatible local inference server.
Fork source:
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
Image source commit:
2e81dc5f634501c744b69a65a8eeb84ba42e82ee
dexogen/atomic-llama-cpp-turboquant:latest
dexogen/atomic-llama-cpp-turboquant:cuda13
dexogen/atomic-llama-cpp-turboquant:2e81dc5
All tags currently point to the same image.
llama-server from AtomicBot-ai/atomic-llama-cpp-turboquantThis is not an official llama.cpp image.
It is a convenience build of the AtomicBot-ai fork, published so it can be pulled directly without rebuilding locally.
The image expects NVIDIA GPU runtime support on the host.
docker pull dexogen/atomic-llama-cpp-turboquant:cuda13
A minimal run usually looks like this:
docker run --rm --gpus all \
-p 8080:8080 \
-v /path/to/models:/models \
dexogen/atomic-llama-cpp-turboquant:cuda13 \
--host 0.0.0.0 \
--port 8080 \
--model /models/model.gguf \
--n-gpu-layers 999 \
--ctx-size 262144 \
--cache-type-k turbo3 \
--cache-type-v turbo3
For speculative decoding, pass the appropriate draft/MTP model arguments supported by this fork and your target model.
This image contains software built from the upstream fork listed above. Refer to the source repository for licensing and upstream details.
Content type
Image
Digest
sha256:e5bb390b5…
Size
3.9 GB
Last updated
5 months ago
docker pull dexogen/atomic-llama-cpp-turboquant