Prefill and token generation speed
How fast will local AI run on my device?
See which local LLMs fit on your Mac or GPU and how fast they run. Compare prefill, token generation, memory, power, price, and TCO under consistent conditions.
New to local LLMs? Start with setup and PC requirements →Last updated
307 GB/s
27B
14.7 ~ 17.1 tok/sDecode speed
Model fully loaded (17.56GB / 56GB)
First token (TTFT)
8.6 to 10.0 secondsContext 4,096 input tokens included
Prefill speed
412 ~ 479 tok/s8.6 to 9.9 seconds (Estimate from user-submitted records)
Total time (E2E)
Approximately 15.4 to 17.9 seconds117 output tokens included
Press Play to process the input and generate the answer.
Ready
Expand settings Settings4-bit · 4K · Baseline decodeCollapse
QuantizationNeed 17.56GB
Input context4K tokens
Baseline
Baseline
Sample question:
117 output tokensCheck configuration
MBP M5 Pro 64GB · Memory 64GB
Qwen3.8 27B Q4_K_M configuration. Check that the retailer lists the same memory option.
Already own it? Open the setup guideWould another device feel faster?
Compare the same model including prefill, or just token generation.
What changes with a longer document?Expand input length, waiting time, speed and memory
Detailed analysis and all 45 hardware configurations▼
Popular configurations
Speed simulations by device and model
Qwen3.8 27B on Mac mini M4 32GB: response speed
5.1 ~ 5.7 tok/s · 16.4 ~ 35.5 s
Qwen3.8 27B: is a Mac mini M4 Pro 48GB worth it?
12.3 ~ 13.7 tok/s · 7.8 ~ 14.3 s
Qwen3.8 27B on Mac Studio M4 Max 128GB
27.9 ~ 60.4 tok/s · 5.8 to 10.5 seconds
Qwen3.8 27B on RTX 3090: do you need an upgrade?
42.3 ~ 46.9 tok/s · 4.4 to 7.5 seconds