Methodology
Local LLM speed and cost calculation benchmark: Instead of concluding with a single number, compare ranges under the same conditions.
We explain under what conditions and correction criteria are calculated the prefill, first token time, decode speed, memory load, power, market price, and TCO for each local LLM device.
Basic comparison conditions
Representative comparisons on the equipment detail page are Q4 quantization, 4K context, and single user conditions. We first determine whether the model weights, execution overhead, and KV cache fit within actual available memory.
Prefill and Decode
Prefill is a step for processing input tokens, and decode is a step for sequentially generating answer tokens. The benchmark range consistent with hardware specifications is calibrated for each condition, and values that deviate from the overall flow are excluded from the representative range.
Power, price, TCO
Power is calculated from sustained local inference load ranges rather than instantaneous peaks. TCO compares the purchase price, expected sale price, electricity bill, and actual replacement subscription fee for one and two years.
limit
Actual results may vary depending on runtime version, cooling, quantization file, context length and multi-GPU communication. Therefore, it is recommended that you double-check the speed gauge with the selected model and exact memory configuration just before purchasing.
Updated September 2026