Prefill and token generation speed
RTX PRO 6000 96GB with Qwen3.8-Flash-Next: prefill and generation speed
Expand settings Settings4-bit · 4K · Baseline decodeCollapse
Check memory
Compare what changes with more memory.
Qwen3.8-Flash-Next Q4_K_M configuration. Check that the retailer lists the same memory option.
Already own it? Open the setup guideWould another device feel faster?
Compare the same model including prefill, or just token generation.
Detailed analysis and all 45 hardware configurations▼
* Memory is calculated for the selected 4,096 input tokens . The model's native context is 262,144 tokens. KV-cache use grows with context length.
Qwen3.8-Flash-Nextkeeps all weights (125B, 106.31GB) in memory, while each decode step uses the routed 6B active experts(3.54GB) for the active-weight bandwidth estimate.
Hardware and model specifications
Simulation assumptions
All 45devices at 4-bit (Q4_K_M)
Select a row to switch devices| Hardware | available | Bandwidth | DeepSeek-V4.1-Flash | Qwen3.8 27B | Qwen3.8-Flash-Next | Gemma 4 12B | Gemma 4 26B-A4B | Qwen3.6 35B-A3B | GLM-5.3-Flash | DeepSeek-V4-Flash-0731 |
|---|---|---|---|---|---|---|---|---|---|---|
| MBP M5 Pro 64GB | 56G | 307G/s | OOM | 14.7~17.1 | OOM | 31.3~34.8 | 28~68.2 | 23.5~56.3 | OOM | OOM |
| MBA M5 16GB | 13G | 153G/s | OOM | OOM | OOM | 14.7~16.5 | 12.3~19.1 (offload) | OOM | OOM | OOM |
| MBA M5 24GB | 20G | 153G/s | OOM | 6.5~7.3 | OOM | 14.7~16.5 | 11.2~31 | 15.6~24.2 (offload) | OOM | OOM |
| MBA M5 32GB | 27G | 153G/s | OOM | 6.5~7.3 | OOM | 14.7~16.5 | 11.2~31 | 9.4~25.6 | OOM | OOM |
| MBA M4 16GB | 13G | 120G/s | OOM | OOM | OOM | 11.2~12.6 | 9.6~15 (offload) | OOM | OOM | OOM |
| MBA M4 24GB | 20G | 120G/s | OOM | 5~5.6 | OOM | 11.2~12.6 | 9.3~27.3 | 12.2~19 (offload) | OOM | OOM |
| MBP M5 Pro 48GB | 41G | 307G/s | OOM | 13.9~15.4 | OOM | 31.3~34.8 | 28~68.2 | 23.5~56.3 | OOM | OOM |
| MBP M5 Max 64GB | 56G | 614G/s | OOM | 28.5~31.6 | OOM | 64.4~71.3 | 51.4~111.6 | 43.1~92.1 | OOM | OOM |
| Mac mini M4 16GB | 13G | 120G/s | OOM | OOM | OOM | 11.6~12.9 | 9.6~15 (offload) | OOM | OOM | OOM |
| Mac mini M4 24GB | 20G | 120G/s | OOM | 5.1~5.7 | OOM | 11.6~12.9 | 9.3~27.3 | 12.2~19 (offload) | OOM | OOM |
| Mac mini M4 32GB | 27G | 120G/s | OOM | 5.1~5.7 | OOM | 11.6~12.9 | 9.3~27.3 | 7.8~22.5 | OOM | OOM |
| Mac mini M4 Pro 24GB | 20G | 273G/s | OOM | 12.3~13.7 | OOM | 27.8~30.9 | 23.4~55.8 | 27.8~43.2 (offload) | OOM | OOM |
| Mac mini M4 Pro 48GB | 41G | 273G/s | OOM | 12.3~13.7 | OOM | 27.8~30.9 | 23.4~55.8 | 19.6~46.1 | OOM | OOM |
| Mac mini M4 Pro 64GB | 56G | 273G/s | OOM | 12.3~13.7 | OOM | 27.8~30.9 | 23.4~55.8 | 19.6~46.1 | OOM | OOM |
| Mac mini M6 16GB | 13G | 153G/s | OOM | OOM | OOM | 15.2~16.9 | 12.3~19.1 (offload) | OOM | OOM | OOM |
| Mac mini M6 24GB | 20G | 170G/s | OOM | 7.5~8.3 | OOM | 16.9~18.8 | 13.1~34.7 | 17.3~26.9 (offload) | OOM | OOM |
| Mac mini M6 32GB | 27G | 170G/s | OOM | 7.5~8.3 | OOM | 16.9~18.8 | 13.1~34.7 | 11~28.7 | OOM | OOM |
| Mac mini M5 Pro 24GB | 20G | 307G/s | OOM | 13.9~15.4 | OOM | 31.3~34.8 | 28~68.2 | 31.2~48.6 (offload) | OOM | OOM |
| Mac mini M5 Pro 48GB | 41G | 307G/s | OOM | 13.9~15.4 | OOM | 31.3~34.8 | 28~68.2 | 23.5~56.3 | OOM | OOM |
| Mac mini M5 Pro 64GB | 56G | 307G/s | OOM | 13.9~15.4 | OOM | 31.3~34.8 | 28~68.2 | 23.5~56.3 | OOM | OOM |
| Studio M4 Max 64GB | 56G | 546G/s | OOM | 25.3~28.1 | OOM | 57.2~63.4 | 46.7~99.2 | 39.2~81.9 | OOM | OOM |
| Studio M4 Max 128GB | 114G | 546G/s | OOM | 25.3~28.1 | 34.2~48.1 | 57.2~63.4 | 46.7~99.2 | 39.2~81.9 | OOM | OOM |
| Studio M5 Max 64GB | 56G | 614G/s | OOM | 28.5~31.6 | OOM | 64.4~71.3 | 51.4~111.6 | 43.1~92.1 | OOM | OOM |
| Studio M5 Max 128GB | 114G | 614G/s | OOM | 28.5~31.6 | 38.5~54 | 64.4~71.3 | 51.4~111.6 | 43.1~92.1 | OOM | OOM |
| RTX 3090 24GB | 22.5G | 936G/s | OOM | 42.3~46.9 | OOM | 95.5~106.1 | 64.1~148.7 | 58.2~146.3 | OOM | OOM |
| RTX 4090 24GB | 22.5G | 1008G/s | OOM | 38.3~51.9 | OOM | 107.1~118.5 | 69~160.2 | 62.7~157.5 | OOM | OOM |
| RTX 5090 32GB | 30.5G | 1792G/s | OOM | 59.1~79.9 | OOM | 198~218.3 | 122.7~284.8 | 111.5~280 | OOM | OOM |
| Ryzen AI Max+ 395 128GB | 96G | 256G/s | OOM | 11.2~12.5 | 13~20.2 (offload) | 25.4~28.3 | 17.5~40.7 | 15.9~40 | OOM | OOM |
| MBP M5 Max 128GB | 114G | 614G/s | OOM | 28.5~31.6 | 38.5~54 | 64.4~71.3 | 51.4~111.6 | 43.1~92.1 | OOM | OOM |
| Studio M5 Ultra 256GB | 236G | 1200G/s | OOM | 54.2~60.2 | 73.2~103.1 | 122.4~136 | 88.8~167.3 | 74.5~138.2 | 81.3~90.3 | 112.5~125 |
| Studio M5 Ultra 512GB | 480G | 1200G/s | OOM | 54.2~60.2 | 73.2~103.1 | 122.4~136 | 88.8~167.3 | 74.5~138.2 | 81.3~90.3 | 112.5~125 |
| RTX 5060 Ti 16GB | 15G | 448G/s | OOM | 5.6~5.7 (offload) | OOM | 47.6~52.7 | 19.4~29.6 (offload) | 15.1~19.5 (offload) | OOM | OOM |
| NVIDIA A40 48GB | 46G | 696G/s | OOM | 31.4~34.9 | 8~8.2 (offload) | 71~78.9 | 47.7~110.6 | 43.3~108.8 | OOM | OOM |
| RTX PRO 6000 96GB | 93G | 1792G/s | OOM | 87.6~96.6 | 118.5~165.4 (offload) | 198~218.3 | 122.7~284.8 | 111.5~280 | 6.6~6.6 (offload) | 9.2~9.2 (offload) |
| MBP M5 32GB | 27G | 153G/s | OOM | 6.7~7.5 | OOM | 15.2~16.9 | 11.2~31 | 9.4~25.6 | OOM | OOM |
| MBP M4 Pro 48GB | 41G | 273G/s | OOM | 12.3~13.7 | OOM | 27.8~30.9 | 23.4~55.8 | 19.6~46.1 | OOM | OOM |
| MBP M4 Max 128GB | 114G | 546G/s | OOM | 25.3~28.1 | 34.2~48.1 | 57.2~63.4 | 46.7~99.2 | 39.2~81.9 | OOM | OOM |
| Studio M3 Ultra 512GB | 480G | 819G/s | OOM | 37~41.1 | 50~70.3 | 83.5~92.8 | 88.8~130.2 | 74.5~107.5 | 55.5~61.6 | 76.8~85.3 |
| MBP M3 Max 128GB | 114G | 400G/s | OOM | 18.6~20.6 | 25.1~35.2 | 41.9~46.5 | 32.7~74.4 | 27.4~61.4 | OOM | OOM |
| MBP M3 Pro 36GB | 30G | 150G/s | OOM | 6.6~7.3 | OOM | 14.9~16.6 | 11.2~31 | 9.4~25.6 | OOM | OOM |
| DGX Spark 128GB | 116G | 273G/s | OOM | 12.3~13.7 | 16.7~23.4 | 27.8~30.9 | 18.7~43.4 | 17~42.7 | OOM | OOM |
| 2× DGX Spark 256GB | 232G | 546G/s | OOM | 12.3~19.7 | 16.7~33.8 | 27.8~44.5 | 18.7~62.5 | 17~61.4 | 18.5~29.6 | 25.6~41 |
| 2× RTX 3090 48GB (NVLink) | 45G | 1872G/s | OOM | 57.5~79.8 | 8.2~8.4 (offload) | 129.8~180.3 | 87.2~252.9 | 79.2~248.6 | OOM | OOM |
| 2× RTX 5090 64GB (PCIe) | 61G | 3584G/s | OOM | 87.6~135.3 | 17.5~18.2 (offload) | 198~305.6 | 122.7~398.7 | 111.5~392 | 5.1~5.2 (offload) | 7.1~7.1 (offload) |
| 2× A40 96GB (NVLink) | 92G | 1392G/s | OOM | 42.7~58.6 | 57.8~100.4 (offload) | 96.5~132.5 | 64.8~185.8 | 58.9~182.7 | 3.4~3.4 (offload) | 4.9~4.9 (offload) |
Popular configurations
Frequently compared hardware and models
RTX PRO 6000 96GB × Qwen3.8-Flash-Next
118.5 ~ 165.4 tok/s · 0.98 to 1.6 seconds
Studio M3 Ultra 512GB × Qwen3.8-Flash-Next
45.9 ~ 62 tok/s · 3.3 to 4.5 seconds
Studio M5 Ultra 512GB × Qwen3.8-Flash-Next
45.9 ~ 79.8 tok/s · 2.6 to 4.5 seconds
2× DGX Spark 256GB × Qwen3.8-Flash-Next
16.7 ~ 33.8 tok/s · 1.7 to 4.3 seconds