A fast, lightweight Rust + CUDA inference server for small language models on a single NVIDIA GPU.
137
Sort by
Tags cannot be overwritten in this Repository. View in settings
TAG
Last pushed 5 months by danielatzep
| Digest | OS/ARCH | Compressed size |
|---|---|---|
a3ea71f839f1 | linux/amd64 | 2.02 GB |