Sign inSign up

dexogen/atomic-llama-cpp-turboquant

By dexogen

•Updated 5 months ago

CUDA llama.cpp server with TurboQuant KV cache and speculative decoding support.

Image
Machine learning & AI
1

8.0K

Tags for dexogen/atomic-llama-cpp-turboquant

Sort by

TAG

Last pushed 5 months by dexogen

docker pull dexogen/atomic-llama-cpp-turboquant:latest
DigestOS/ARCHCompressed size

0bc516c11db2

linux/amd64

3.93 GB

TAG

Last pushed 5 months by dexogen

docker pull dexogen/atomic-llama-cpp-turboquant:cuda13
DigestOS/ARCHCompressed size

0bc516c11db2

linux/amd64

3.93 GB

TAG

Last pushed 5 months by dexogen

docker pull dexogen/atomic-llama-cpp-turboquant:2e81dc5
DigestOS/ARCHCompressed size

0bc516c11db2

linux/amd64

3.93 GB