A catalogue of local model performance
Sloptimized.
Fastest measured local inference result per language model. Speed, not model quality. Complete catalogue ↗Optimizations ↗
- 20 models
- 338 runs
- 33 techniques
Best result per model
Prefill and decode are tokens/second. Context is depth / window; n/m means not measured in the winning run.
| Model | Quant | Backend | Hardware | Context | Prefill (tok/s) | Status | ||
|---|---|---|---|---|---|---|---|---|
| GLM-5.2 | 744B / 40B | INT4 g64 + INT8 MTP | Colibri | CPU + V100 (2×) | 0 / 4K | n/m | 1.2 | Provisional |
| DeepSeek V4 Flash 0731 | 284B / 13B | UD-Q8_K_XL | llama.cpp | CPU + V100 (2×) | 0 / 4K | 15.9 | 4.1 | Valid |
| MiniMax M2.5 | 229B / 10B | Q3_K_M | llama.cpp | V100 (4×) | 23827 / 32K | 352.1 | 21.3 | Valid |
| Step 3.5 Flash | 197B / 11B | Q4_K_M | llama.cpp | V100 (4×) | 24699 / 32K | n/m | 27.2 | Valid |
| Qwen3.5 122B-A10B | 125B / 10B | Q3_K_M | llama.cpp | V100 (2×) | 0 / 1K | 240.1 | 44.3 | Provisional |
| Mistral Small 4 119B 2603 | 119B | UD-IQ3_XXS | llama.cpp | V100 (2×) | 0 / 1K | 422.6 | 75.1 | Valid |
| Laguna S 2.1 | 118B / 8B | Q4_K_M | llama.cpp | V100 (3×) | 0 | n/m | 38.9 | Valid |
| gpt-oss-120B | 117B / 5B | MXFP4 | llama.cpp | V100 (2×) | 0 / 1K | 627.4 | 99.9 | Provisional |
| Qwen3-Coder-Next 80B-A3B | 80B / 3B | Q4_K_M | llama.cpp | V100 (2×) | 0 / 1K | 270.5 | 79.6 | Provisional |
| Qwen2.5 72B Instruct | 72B | Q4_K_M | llama.cpp | V100 (2×) | 8K / 32K | 118.0 | 8.6 | Valid |
| Dolphin 2.2 70B | 70B | Q3_K_M | llama.cpp | V100 (2×) | 8K / 32K | 127.9 | 11.5 | Provisional |
| Qwen3.6 35B-A3B | 35B / 3B | Q4_K_M | llama.cpp | V100 | 0 / 4K | 76.8 | 101.5 | Provisional |
| Laguna XS 2.1 | 33B / 3B | Q4_K_M | llama.cpp | V100 | 0 / 4K | 275.6 | 104.6 | Provisional |
| Gemma 4 31B IT | 31B | UD-Q4_K_XL | llama.cpp | V100 | — | 258.9 | 35.6 | Valid |
| Muse Glimmer 30B | 30B | K-Quant-17GB | llama.cpp | V100 | 0 / 4K | 79.3 | 45.4 | Valid |
| Qwen3.6 27B | 28B | UD-Q4_K_XL | llama.cpp | V100 | — / 4K | 84.3 | 39.4 | Valid |
| Qwen3.8 27B | 27B | Q4_K_M | llama.cpp | V100 | 20 / 4K | 36.9 | 55.0 | Valid |
| Gemma 4 26B-A4B IT | 25B / 4B | UD-Q4_K_XL | llama.cpp | V100 | — | 776.5 | 81.0 | Valid |
| Bonsai 8B | 8B | Q1_0 | llama.cpp | V100 | — | 1338.8 | 123.3 | Valid |
| TinyLlama OpenOrca 1.1B | 1B | Q4_K_M | llama.cpp | V100 | 0 / 4K | 4150.9 | 384.1 | Valid |
Valid means no known correctness or sample-count issue. Provisional means a known correctness problem or too few samples. Both compete here. Hardware, context, cache state, and quantization can make rows unlike; use the complete catalogue for rigorous comparisons.