See the speed before choosing hardware

Prefill and token generation speed

RTX PRO 6000 96GB with Qwen3.8-Flash-Next: prefill and generation speed

1792 GB/s
Active 6B / Sangju 180B
118.5 ~ 165.4 tok/sDecode speed
PLE table RAM offload (VRAM 77.01GB + RAM 30.12GB)
First token (TTFT)
0.98 to 1.6 secondsContext 4,096 input tokens included
Prefill speed
2,545 ~ 4,208 tok/s0.97 to 1.6 seconds (Estimation based on public benchmarks)
Total time (E2E)
Approximately 1.7 to 2.6 seconds117 output tokens included
Press Play to process the input and generate the answer.
Ready
Expand settings 4-bit · 4K · Baseline decode
QuantizationNeed 107.13GB
Input context4K tokens
Baseline
Baseline
PLE n-gram table: system RAMplaced automatically. Applied alongside decode optimization.
Sample question:
117 output tokens

Check memory

Compare what changes with more memory.

Qwen3.8-Flash-Next Q4_K_M configuration. Check that the retailer lists the same memory option.

Already own it? Open the setup guide

Would another device feel faster?

Compare the same model including prefill, or just token generation.

Detailed analysis and all 45 hardware configurations
Memory allocation: Qwen3.8-Flash-Next (Q4_K_M)Core VRAM 77.01GB + PLE table RAM 30.12GB
Core weights 76.19GB■ PLE table RAM 30.12GB■ KV cache (4K) 0.42GB■ Runtime 0.4GB■ OS/display 3GB

* Memory is calculated for the selected 4,096 input tokens . The model's native context is 262,144 tokens. KV-cache use grows with context length.

How MoE uses memory

Qwen3.8-Flash-Nextkeeps all weights (125B, 106.31GB) in memory, while each decode step uses the routed 6B active experts(3.54GB) for the active-weight bandwidth estimate.

Hardware and model specifications

Memory bandwidth1792 GB/s
Fast memory96 GB
Model parameters125B (active 6B)
Native context262,144 tokens

Simulation assumptions

Effective bandwidth factor78% ~ 86%
KV cache (4K)About 0.42 GB
Decode estimationMoE architecture calibration heuristic
Prefill calibrationEstimation based on public benchmarks
Prefill computeTransferred benchmark model
Acceleration multiplier1.0x (automatic PLE RAM placement, no decode acceleration)
AssumptionsSingle request · warm · uncached · no queue/network delay

All 45devices at 4-bit (Q4_K_M)

Select a row to switch devices
HardwareavailableBandwidthDeepSeek-V4.1-FlashQwen3.8 27BQwen3.8-Flash-NextGemma 4 12BGemma 4 26B-A4BQwen3.6 35B-A3BGLM-5.3-FlashDeepSeek-V4-Flash-0731
MBP M5 Pro 64GB56G307G/sOOM14.7~17.1OOM31.3~34.828~68.223.5~56.3OOMOOM
MBA M5 16GB13G153G/sOOMOOMOOM14.7~16.512.3~19.1 (offload)OOMOOMOOM
MBA M5 24GB20G153G/sOOM6.5~7.3OOM14.7~16.511.2~3115.6~24.2 (offload)OOMOOM
MBA M5 32GB27G153G/sOOM6.5~7.3OOM14.7~16.511.2~319.4~25.6OOMOOM
MBA M4 16GB13G120G/sOOMOOMOOM11.2~12.69.6~15 (offload)OOMOOMOOM
MBA M4 24GB20G120G/sOOM5~5.6OOM11.2~12.69.3~27.312.2~19 (offload)OOMOOM
MBP M5 Pro 48GB41G307G/sOOM13.9~15.4OOM31.3~34.828~68.223.5~56.3OOMOOM
MBP M5 Max 64GB56G614G/sOOM28.5~31.6OOM64.4~71.351.4~111.643.1~92.1OOMOOM
Mac mini M4 16GB13G120G/sOOMOOMOOM11.6~12.99.6~15 (offload)OOMOOMOOM
Mac mini M4 24GB20G120G/sOOM5.1~5.7OOM11.6~12.99.3~27.312.2~19 (offload)OOMOOM
Mac mini M4 32GB27G120G/sOOM5.1~5.7OOM11.6~12.99.3~27.37.8~22.5OOMOOM
Mac mini M4 Pro 24GB20G273G/sOOM12.3~13.7OOM27.8~30.923.4~55.827.8~43.2 (offload)OOMOOM
Mac mini M4 Pro 48GB41G273G/sOOM12.3~13.7OOM27.8~30.923.4~55.819.6~46.1OOMOOM
Mac mini M4 Pro 64GB56G273G/sOOM12.3~13.7OOM27.8~30.923.4~55.819.6~46.1OOMOOM
Mac mini M6 16GB13G153G/sOOMOOMOOM15.2~16.912.3~19.1 (offload)OOMOOMOOM
Mac mini M6 24GB20G170G/sOOM7.5~8.3OOM16.9~18.813.1~34.717.3~26.9 (offload)OOMOOM
Mac mini M6 32GB27G170G/sOOM7.5~8.3OOM16.9~18.813.1~34.711~28.7OOMOOM
Mac mini M5 Pro 24GB20G307G/sOOM13.9~15.4OOM31.3~34.828~68.231.2~48.6 (offload)OOMOOM
Mac mini M5 Pro 48GB41G307G/sOOM13.9~15.4OOM31.3~34.828~68.223.5~56.3OOMOOM
Mac mini M5 Pro 64GB56G307G/sOOM13.9~15.4OOM31.3~34.828~68.223.5~56.3OOMOOM
Studio M4 Max 64GB56G546G/sOOM25.3~28.1OOM57.2~63.446.7~99.239.2~81.9OOMOOM
Studio M4 Max 128GB114G546G/sOOM25.3~28.134.2~48.157.2~63.446.7~99.239.2~81.9OOMOOM
Studio M5 Max 64GB56G614G/sOOM28.5~31.6OOM64.4~71.351.4~111.643.1~92.1OOMOOM
Studio M5 Max 128GB114G614G/sOOM28.5~31.638.5~5464.4~71.351.4~111.643.1~92.1OOMOOM
RTX 3090 24GB22.5G936G/sOOM42.3~46.9OOM95.5~106.164.1~148.758.2~146.3OOMOOM
RTX 4090 24GB22.5G1008G/sOOM38.3~51.9OOM107.1~118.569~160.262.7~157.5OOMOOM
RTX 5090 32GB30.5G1792G/sOOM59.1~79.9OOM198~218.3122.7~284.8111.5~280OOMOOM
Ryzen AI Max+ 395 128GB96G256G/sOOM11.2~12.513~20.2 (offload)25.4~28.317.5~40.715.9~40OOMOOM
MBP M5 Max 128GB114G614G/sOOM28.5~31.638.5~5464.4~71.351.4~111.643.1~92.1OOMOOM
Studio M5 Ultra 256GB236G1200G/sOOM54.2~60.273.2~103.1122.4~13688.8~167.374.5~138.281.3~90.3112.5~125
Studio M5 Ultra 512GB480G1200G/sOOM54.2~60.273.2~103.1122.4~13688.8~167.374.5~138.281.3~90.3112.5~125
RTX 5060 Ti 16GB15G448G/sOOM5.6~5.7 (offload)OOM47.6~52.719.4~29.6 (offload)15.1~19.5 (offload)OOMOOM
NVIDIA A40 48GB46G696G/sOOM31.4~34.98~8.2 (offload)71~78.947.7~110.643.3~108.8OOMOOM
RTX PRO 6000 96GB93G1792G/sOOM87.6~96.6118.5~165.4 (offload)198~218.3122.7~284.8111.5~2806.6~6.6 (offload)9.2~9.2 (offload)
MBP M5 32GB27G153G/sOOM6.7~7.5OOM15.2~16.911.2~319.4~25.6OOMOOM
MBP M4 Pro 48GB41G273G/sOOM12.3~13.7OOM27.8~30.923.4~55.819.6~46.1OOMOOM
MBP M4 Max 128GB114G546G/sOOM25.3~28.134.2~48.157.2~63.446.7~99.239.2~81.9OOMOOM
Studio M3 Ultra 512GB480G819G/sOOM37~41.150~70.383.5~92.888.8~130.274.5~107.555.5~61.676.8~85.3
MBP M3 Max 128GB114G400G/sOOM18.6~20.625.1~35.241.9~46.532.7~74.427.4~61.4OOMOOM
MBP M3 Pro 36GB30G150G/sOOM6.6~7.3OOM14.9~16.611.2~319.4~25.6OOMOOM
DGX Spark 128GB116G273G/sOOM12.3~13.716.7~23.427.8~30.918.7~43.417~42.7OOMOOM
2× DGX Spark 256GB232G546G/sOOM12.3~19.716.7~33.827.8~44.518.7~62.517~61.418.5~29.625.6~41
2× RTX 3090 48GB (NVLink)45G1872G/sOOM57.5~79.88.2~8.4 (offload)129.8~180.387.2~252.979.2~248.6OOMOOM
2× RTX 5090 64GB (PCIe)61G3584G/sOOM87.6~135.317.5~18.2 (offload)198~305.6122.7~398.7111.5~3925.1~5.2 (offload)7.1~7.1 (offload)
2× A40 96GB (NVLink)92G1392G/sOOM42.7~58.657.8~100.4 (offload)96.5~132.564.8~185.858.9~182.73.4~3.4 (offload)4.9~4.9 (offload)

Popular configurations

View all 20