Response speed and acceleration

Local AI speed report — September 2026

When you open the speed table, you will naturally look for the highest number first. But whether that equipment is the best choice for me is another matter. This table looks at not only the speed of writing answers, but also the wait before starting. Rather than memorizing the rankings, read it as a resource that leaves you with two candidates.

Equipment speed table of the month

Qwen3.6 35B-A3B · input 4,096 tokens · output 302 tokens · saved estimate on 2026-09-05. The superiority of equipment with overlapping ranges is not confirmed.

equipmentexecution settingsFirst tokenDecodecomplete complete
RTX 5090 32GBllama.cpp CUDA · Q4_K_M0.582 s111.5 ~ 280 tok/sAbout 2.47 s
2× RTX 3090 48GB (NVLink)llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces0.975 s79.2 ~ 248.6 tok/sAbout 3.48 s
MBP M5 Max 128GBMLX 4-bit2.093 s43.1 ~ 92.1 tok/sAbout 7.22 s
DGX Spark 128GBllama.cpp · Q4_K_M2.033 s17 ~ 42.7 tok/sAbout 14.41 s
Mac mini M4 Pro 64GBMLX 4-bit4.473 s19.6 ~ 46.1 tok/sAbout 15.41 s
Ryzen AI Max+ 395 128GBllama.cpp HIP · Q4_K_M37.873 s15.9 ~ 40 tok/sAbout 51.1 s

Start with the workload behind the table

This record uses the 4-bit model of Qwen3.6 35B-A3B, with an input of 4,096 tokens and a response of 302 tokens. It is not a collection of best results with a free mix of questions or models. This is a fixed condition for examining differences in equipment, assuming that the same work is assigned.

Numbers are calculation results using specifications and benchmark corrections. The values ​​are not measured because I personally own all the equipment. We did not apply unfounded MTP multipliers in bulk, and the multiple equipment configurations are estimates of the conditions under which the supported runtime divides the model.

Illustration of an understated monthly benchmark workbench checking multiple local AI machines under the same conditions.
The monthly speed table compares the perceived differences between devices at a glance by fixing the same model, quantization, and input length.

First, narrow the choices by memory

Even if the creation speed appears to be high, if it does not contain the desired model and cache, it is not a candidate under the same conditions. In particular, you should not compare a Mac's total memory and the GPU's dedicated VRAM as if it were space used only by the model. Look at the available space reserved for the operating system and executable programs.

Large MoE models can be fast with a small amount of active computation, but require retaining full weights. Don't over-interpret that the equipment that was good for this model is good enough for other large models. If the model you want to use is different, change it in the calculator and check.

Illustration of a prefill section that reads the input and a decode section that writes the answer in two directions.
I need to separate the speed at which I read long documents from the speed at which answer tokens follow each other to find the right equipment for my work.

Then ask which part of the wait is longer

When inserting a document and receiving a short summary, the time until the first token is important. When writing long text with short instructions, decode speed has a longer impact. The two values ​​are not always good together in one piece of equipment.

Try including the prefill in the same model and compare just the token generation again. Differences that seem large at first may become smaller in the work you do. Conversely, even if the total completion time is similar, it may feel more frustrating to start answering later.

Equipment configuration diagram showing a Mac, single GPU, dual GPU, and small AI system side by side on the same workbench.
As device count and display memory grow, the speed of a single request and the ability to handle multiple requests do not increase at the same rate.

Conditions for saying that the ranking has improved since a month ago

This page is a calculation history saved on September 5, 2026. The table will not be silently overwritten if the current calculator settings or correction values ​​are changed. Comparisons of error corrections and new months should separate the change history so that readers of past links can understand the same information.

When comparing with other months in the future, you should first check whether the model, precision, and input length have been maintained. A change in the ranking of a table with different conditions cannot be directly called an improvement in equipment performance. Whether it is a runtime improvement or a modification of the calculation criteria, these are separate changes.

What if the fastest option costs too much?

Even if it ranks one level lower, I can leave it as a purchase candidate if it's good enough for my job. Conversely, if the price difference is small and you have to wait for a long reply every day, there is a reason to spend money on higher-end equipment. Rankings alone cannot replace that judgment.

Put the two devices you are interested in on the comparison screen and share the conditions. Rather than asking ‘Which one is better?’, you can get a more specific answer by asking ‘Would this difference be ok when processing documents of this length with this model?’ The purpose of this table is not to rush you into making a purchase, but to reduce the number of questions you have to ask next time.