35B An Jeong-kwonUpdated September 2026

MacBook Pro M5 Max 64GB Integrated memory 64GB · Usable space for model approximately 56GB

This is the 64GB M5 Max configuration for users who want both high memory bandwidth and portability.

Hardware core specifications

available memory
56GB
Memory bandwidth
614GB/s
Estimated sustained LLM load
About 120W
Korean price range (KRW)
680 × KRW 10,000–780 × KRW 10,000

a person who fits well

  • Utilizes 56GB of available memory
  • Quiet finished product composition
  • K2 Horizon MoVA 36B-A4B-level local inference

In this case skip it

Unified memory is advantageous for loading large models, but CUDA-specific tools and multi-GPU workflows should be explored separately.

Recommended format for each device 4K input

The model you will experience with this equipment

K2 Horizon 0.9B

Required 1.03GB · spare time

MLX 4-bit

Decode

857.3 ~ 950 tok/s

First token

0.41 to 0.80 seconds

K2 Horizon MoVA 36B-A4B

recommended

Required 21.94GB · spare time

MLX 4-bit

Decode

192.5 ~ 213.3 tok/s

First token

2.2 to 4.5 seconds

Actual speed may vary depending on runtime version, cooling status, context length and quantization file.

On sales pages, look at memory before chips.

The model and speed assumptions above change if the product is not the 64GB configuration.

Frequently Asked Questions

What LLM can be run on the MBP M5 Max 64GB?
You can compare K2 Horizon 0.9B, K2 Horizon MoVA 36B-A4B, etc. with 4K input and recommended formats. Memory requirements vary depending on model and context length.
What is the power consumption of a local LLM on an MBP M5 Max 64GB?
Expected to consume approximately 120W under sustained load, with an actual range estimated between 85–140W.
Are the displayed token speeds ground truth?
These ranges normalize collected benchmark measurements to the same conditions. Only combinations without matching measurements are estimated from nearby measured values. Results can vary by runtime, cooling, model file, and context length.