Response speed and acceleration
Local AI speed report — September 2026
When you open the speed table, you will naturally look for the highest number first. But whether that equipment is the best choice for me is another matter. This table looks at not only the speed of writing answers, but also the wait before starting. Rather than memorizing the rankings, read it as a resource that leaves you with two candidates.
Equipment speed table of the month
Qwen3.6 35B-A3B · input 4,096 tokens · output 302 tokens · saved estimate on 2026-09-05. The superiority of equipment with overlapping ranges is not confirmed.
| equipment | execution settings | First token | Decode | complete complete |
|---|---|---|---|---|
| RTX 5090 32GB | llama.cpp CUDA · Q4_K_M | 0.582 s | 111.5 ~ 280 tok/s | About 2.47 s |
| 2× RTX 3090 48GB (NVLink) | llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces | 0.975 s | 79.2 ~ 248.6 tok/s | About 3.48 s |
| MBP M5 Max 128GB | MLX 4-bit | 2.093 s | 43.1 ~ 92.1 tok/s | About 7.22 s |
| DGX Spark 128GB | llama.cpp · Q4_K_M | 2.033 s | 17 ~ 42.7 tok/s | About 14.41 s |
| Mac mini M4 Pro 64GB | MLX 4-bit | 4.473 s | 19.6 ~ 46.1 tok/s | About 15.41 s |
| Ryzen AI Max+ 395 128GB | llama.cpp HIP · Q4_K_M | 37.873 s | 15.9 ~ 40 tok/s | About 51.1 s |
Start with the workload behind the table
This record uses the 4-bit model of Qwen3.6 35B-A3B, with an input of 4,096 tokens and a response of 302 tokens. It is not a collection of best results with a free mix of questions or models. This is a fixed condition for examining differences in equipment, assuming that the same work is assigned.
Numbers are calculation results using specifications and benchmark corrections. The values are not measured because I personally own all the equipment. We did not apply unfounded MTP multipliers in bulk, and the multiple equipment configurations are estimates of the conditions under which the supported runtime divides the model.

First, narrow the choices by memory
Even if the creation speed appears to be high, if it does not contain the desired model and cache, it is not a candidate under the same conditions. In particular, you should not compare a Mac's total memory and the GPU's dedicated VRAM as if it were space used only by the model. Look at the available space reserved for the operating system and executable programs.
Large MoE models can be fast with a small amount of active computation, but require retaining full weights. Don't over-interpret that the equipment that was good for this model is good enough for other large models. If the model you want to use is different, change it in the calculator and check.

Then ask which part of the wait is longer
When inserting a document and receiving a short summary, the time until the first token is important. When writing long text with short instructions, decode speed has a longer impact. The two values are not always good together in one piece of equipment.
Try including the prefill in the same model and compare just the token generation again. Differences that seem large at first may become smaller in the work you do. Conversely, even if the total completion time is similar, it may feel more frustrating to start answering later.

Conditions for saying that the ranking has improved since a month ago
This page is a calculation history saved on September 5, 2026. The table will not be silently overwritten if the current calculator settings or correction values are changed. Comparisons of error corrections and new months should separate the change history so that readers of past links can understand the same information.
When comparing with other months in the future, you should first check whether the model, precision, and input length have been maintained. A change in the ranking of a table with different conditions cannot be directly called an improvement in equipment performance. Whether it is a runtime improvement or a modification of the calculation criteria, these are separate changes.
What if the fastest option costs too much?
Even if it ranks one level lower, I can leave it as a purchase candidate if it's good enough for my job. Conversely, if the price difference is small and you have to wait for a long reply every day, there is a reason to spend money on higher-end equipment. Rankings alone cannot replace that judgment.
Put the two devices you are interested in on the comparison screen and share the conditions. Rather than asking ‘Which one is better?’, you can get a more specific answer by asking ‘Would this difference be ok when processing documents of this length with this model?’ The purpose of this table is not to rush you into making a purchase, but to reduce the number of questions you have to ask next time.