Substantial configurationUpdated September 2026

Mac mini M4 Pro 24GB Integrated memory 24GB · Usable space for model: approximately 20GB

This configuration prioritizes the creation speed of the 12B model with high bandwidth within the 24GB loading limit.

Hardware core specifications

available memory
20GB
Memory bandwidth
273GB/s
Estimated sustained LLM load
About 86W
Korean price range (KRW)
200 × KRW 10,000–330 × KRW 10,000

a person who fits well

  • Utilizes 20GB of available memory
  • Quiet finished product composition
  • Qwen3.6 35B-A3B (MoE) level local inference

In this case skip it

Since the operating system and KV cache use the same memory, headroom quickly decreases for long contexts or models above 20B.

Recommended format for each device 4K input

The model you will experience with this equipment

MiniCPM5-2B

Required 2.07GB · spare time

MLX 4-bit

Decode

131.9 ~ 146.6 tok/s

First token

1.7 to 3.8 seconds

speed tuning

Qwen3.6 35B-A3B (MoE)

recommended

Required 21.28GB · tight

MLX 4-bit

Decode

27.8 ~ 43.2 tok/s

First token

2.8 ~ 6.5 s

speed tuning

Actual speed may vary depending on runtime version, cooling status, context length and quantization file.

On sales pages, look at memory before chips.

The model and speed assumptions above change if the product is not the 24GB configuration.

Frequently Asked Questions

What LLM can be run on the Mac mini M4 Pro 24GB?
You can compare MiniCPM5-2B, Qwen3.6 35B-A3B (MoE), etc. with 4K input and recommended formats. Memory requirements vary depending on model and context length.
What is the local LLM computing power of the Mac mini M4 Pro 24GB?
Expected to consume approximately 86W under sustained load, with an actual range estimated between 58–115W.
Are the displayed token speeds ground truth?
These ranges normalize collected benchmark measurements to the same conditions. Only combinations without matching measurements are estimated from nearby measured values. Results can vary by runtime, cooling, model file, and context length.