Capacity priority · Officially released product · Updated 2026-09-06
27B goes in. This is a configuration that chooses capacity over speed.
Mac Mini M4 32GB Local LLM 27B Running Performance
The 27B Q4 class model in a small body is ideal when loading is a priority and creation speed can wait.
Core specifications
- available memory
- 27GB
- Memory bandwidth
- 120GB/s
- local LLM sustained load
- about 48W
- Korean price range (KRW)
- 180–240 × KRW 10,000
a person who fits well
- 27B Q4 execution
- low power private server
- Capacity priority users
Check before purchasing
Due to the bandwidth of the M4 base chip, decode is slower than the Pro series with the same memory.
What tasks is it suitable for?
| task | determination | reason |
|---|---|---|
| 27B conversation | fit | Q4 Loading is possible, but the long answer will have to wait. |
| personal documents | fit | It's great for working on documents all the time with low power consumption. |
| Create image | conditional | Memory is advantageous, but generation rates are lower than with dedicated GPUs. |
Configuration Selection Criteria
The first bottleneck encountered
Even after loading 27B, there is still memory, but answer generation takes longer due to the bandwidth of the basic M4.
Recommended purchase setup
M4 · Integrated memory 32GB · SSD 512GB or more
Criteria for spending more money
27B This makes sense if loading is the goal. If quick response is more important to you, look at the price difference with the M4 Pro 48GB.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B MLX 4-bit | 8GB | 11.6 ~ 12.9 tok/s | 12.8 ~ 31.1 s | 132 ~ 323 tok/s |
| Qwen3.8 27B MLX 4-bit | 17.56GB | 5.1 ~ 5.7 tok/s | 16.4 ~ 35.5 s | 116 ~ 252 tok/s |
| Qwen3.6 35B-A3B (MoE) MLX 4-bit | 21.28GB | 7.8 ~ 22.5 tok/s | 5.5 ~ 15.7 s | 263 ~ 757 tok/s |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
What LLM can be run on the Mac mini M4 32GB?
Compare Gemma 4 12B, Qwen3.8 27B, Qwen3.6 35B-A3B (MoE), and other models with 4K input and the recommended format for each device. Memory requirements vary with the model and context length.
What is the local LLM power of the Mac mini M4 32GB?
Expected to be around 48W under sustained load, with an actual range of 32–63W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.