Substantial configuration · Officially released product · Updated 2026-09-06
Integrated memory 24GB · Usable space for model: approximately 20GB
MacBook Air M4 24GB Local LLM Speed/Memory/Electricity Rate
Configuration that secures memory space in the lower-priced M4 Air. You can compare the 12B class with some 27B Q4 models.
Core specifications
- available memory
- 20GB
- Memory bandwidth
- 120GB/s
- local LLM sustained load
- about 28W
- Korean price range (KRW)
- 165–225 × KRW 10,000
a person who fits well
- Utilizes 20GB of available memory
- Quiet finished product composition
- Qwen3.6 35B-A3B (MoE) level local inference
Check before purchasing
Since the operating system and KV cache use the same memory, headroom quickly decreases for long contexts or models above 20B.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B MLX 4-bit | 8GB | 11.2 ~ 12.6 tok/s | 12.8 ~ 31.1 s | 132 ~ 323 tok/s |
| Qwen3.6 35B-A3B (MoE) MLX 4-bit | 21.28GB | 12.2 ~ 19 tok/s | 5.7 ~ 16.3 s | 252 ~ 725 tok/s |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
Which local LLM can be run on the MBA M4 24GB?
Gemma 4 12B, Qwen3.6 35B-A3B (MoE) etc. can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.
How much computing power does the local LLM have on the MBA M4 24GB?
Expected to be around 28W under sustained load, with an actual range of 19–35W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.