https://github.com/ztxz16/fastllm
4.1K
fastllm is a high-performance LLMs inference library implemented in C++ with no backend dependencies (e.g. PyTorch).
It enables hybrid inference of MOE models, achieving 20+ tps on consumer-grade single GPUs (e.g., 4090) for DeepSeek R1 671B INT4 model inference.
Content type
Image
Digest
sha256:b27d13899…
Size
7.2 GB
Last updated
15 days ago
docker pull garenleeasa/ftllm:v0.1.8.2