No-frills · Officially released product · Updated 2026-09-06

A practical model that uses 12B comfortably. From 27B onwards, conditions apply.

Mac Mini M4 24GB local LLM performance and appropriate model

This is a good balance when you frequently use the 12B class model and want to keep the main body price and power usage down.

Core specifications

available memory
20GB
Memory bandwidth
120GB/s
local LLM sustained load
about 45W
Korean price range (KRW)
150–235 × KRW 10,000

a person who fits well

  • Class 12B daily work
  • Personal document summary
  • small desk environment

Check before purchasing

For class 27B, memory space quickly decreases depending on quantization and context length.

What tasks is it suitable for?

taskdeterminationreason
Document/RAGfitThere is room for 12B models and medium length documents.
coding assistantfitIt is ideal for quietly running small code models all the time.
27B conversationconditionalAlthough Q4 loading is possible, the margin is small for long contexts.

Configuration Selection Criteria

The first bottleneck encountered

120GB/s bandwidth limits perceived speed before memory capacity.

Recommended purchase setup

M4 · Integrated memory 24GB · SSD 512GB or more

Criteria for spending more money

If 12B is your main device, it is more reasonable than 32GB. If you use 27GB often, there is a reason to increase it to 32GB.

4K context · Q4 model-specific expected performance

Modelrequired memoryDecodeFirst tokenPrefill
Gemma 4 12B
MLX 4-bit
8GB11.6 ~ 12.9 tok/s12.8 ~ 31.1 s132 ~ 323 tok/s
Qwen3.6 35B-A3B (MoE)
MLX 4-bit
21.28GB12.2 ~ 19 tok/s5.7 ~ 16.3 s252 ~ 725 tok/s

Single-user expected range and may vary depending on runtime, cooling, and quantization files.

Frequently Asked Questions

What LLM can be run on the Mac mini M4 24GB?

Gemma 4 12B, Qwen3.6 35B-A3B (MoE) etc. can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.

What is the local LLM power of the Mac mini M4 24GB?

Expected to be around 45W under sustained load, with an actual range of 30–60W.

Are the displayed token speeds ground truth?

This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.