Community User
Displaying 1 to 7 of 7 repositories
Qwen3.5-35B-A3B-FP8 on NVIDIA GB10 Blackwell - 262K context, 53 tok/s, MTP
5m
10K+
1
Functionary v3.2 tool router for CCA — ARM64+CUDA (DGX Spark GB10)
6m
640
vLLM for NVIDIA GB10 DGX Spark - SM121 Blackwell, FlashInfer + Triton MoE
6m
9.9K
1
Qwen3.5-9B-FP8 with thinking + tool calling — CCA note-taker and tool orchestrator.
6m
1.1K
Qwen3-Embedding-8B via vLLM pooling — 4096-dim embeddings for CCA code intelligence.
6m
1.4K
LLM inference benchmark tool with web dashboard for OpenAI-compatible APIs
7m
4.2K