LocalAI
Popular local LLM speed results
Check prefill, time to first token, and generation speed for 20 popular hardware and model combinations using recommended settings.
Where to start
Choose the combination closest to your device or purchase shortlist. Other combinations remain available in the calculator without separate search pages.
Q4_K_M · 4K
Mac mini M4 24GB with Gemma 4 12B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Gemma 4 12B on Mac mini M4 24GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Mac mini M4 32GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on Mac mini M4 32GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Mac mini M4 Pro 48GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on Mac mini M4 Pro 48GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Mac mini M4 Pro 64GB with Qwen3.6 35B-A3B (MoE): prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.6 35B-A3B (MoE) on Mac mini M4 Pro 64GB with recommended settings and the same 4K input.
Q4_K_M · 4K
MBP M5 Pro 64GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on MBP M5 Pro 64GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Studio M4 Max 128GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on Studio M4 Max 128GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Studio M4 Max 128GB with Qwen3.6 35B-A3B (MoE): prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.6 35B-A3B (MoE) on Studio M4 Max 128GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Studio M3 Ultra 512GB with Qwen3.8-Flash-Next: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8-Flash-Next on Studio M3 Ultra 512GB with recommended settings and the same 4K input.
Q4_K_M · 4K
Studio M5 Ultra 512GB with Qwen3.8-Flash-Next: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8-Flash-Next on Studio M5 Ultra 512GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX 3090 24GB with Gemma 4 12B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Gemma 4 12B on RTX 3090 24GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX 3090 24GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on RTX 3090 24GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX 4090 24GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on RTX 4090 24GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX 5090 32GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on RTX 5090 32GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX 5090 32GB with Gemma 4 26B-A4B (MoE): prefill and generation speed
Experience prefill, time to first token, and generation speed for Gemma 4 26B-A4B (MoE) on RTX 5090 32GB with recommended settings and the same 4K input.
Q4_K_M · 4K
RTX PRO 6000 96GB with Qwen3.8-Flash-Next: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8-Flash-Next on RTX PRO 6000 96GB with recommended settings and the same 4K input.
Q4_K_M · 4K
DGX Spark 128GB with Qwen3.6 35B-A3B (MoE): prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.6 35B-A3B (MoE) on DGX Spark 128GB with recommended settings and the same 4K input.
Q4_K_M · 4K
2× DGX Spark 256GB with Qwen3.8-Flash-Next: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8-Flash-Next on 2× DGX Spark 256GB with recommended settings and the same 4K input.
Q4_K_M · 4K
2× RTX 3090 48GB (NVLink) with Qwen3.6 35B-A3B (MoE): prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.6 35B-A3B (MoE) on 2× RTX 3090 48GB (NVLink) with recommended settings and the same 4K input.
Q4_K_M · 4K
2× RTX 5090 64GB (PCIe) with Qwen3.8-Flash-Next: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8-Flash-Next on 2× RTX 5090 64GB (PCIe) with recommended settings and the same 4K input.
Q4_K_M · 4K
Ryzen AI Max+ 395 128GB with Qwen3.8 27B: prefill and generation speed
Experience prefill, time to first token, and generation speed for Qwen3.8 27B on Ryzen AI Max+ 395 128GB with recommended settings and the same 4K input.