Executable programs and extensions

Run local AI with MLX on Apple Silicon: frameworks and model tools explained

Check your Mac’s chip and the model files you need first. Having MLX does not automatically port every CUDA script or model.

On an Apple Silicon Mac, MLX1 is the compute framework, MLX-LM2 is a language-model tool, Metal3 is the GPU4 execution foundation, and PyTorch MPS5 is a backend for a different framework. GUI apps are another option. Unified memory6 is shared by the CPU7 and GPU, and is also used by the OS, apps, model and KV cache8, so the installed capacity is not all available to the model. Check Intel Mac support, CUDA9 code and model-specific compatibility before choosing a route.

Requirements and key details
  • MLX is an array and machine-learning framework; MLX-LM runs language models. Metal and PyTorch MPS are different kinds of components.
  • Unified memory is shared by the CPU and GPU, but it does not add capacity. The operating system, apps, model and KV cache all use it.
  • Confirm that the required MLX checkpoint and features are supported. CUDA code and Intel Mac setups do not automatically become Apple Silicon routes.

What changes when you use MLX on Apple Silicon?

To run a language model on an Apple Silicon Mac, first check About This Mac for an M-series chip. This guide follows one path: load an MLX checkpoint10 with MLX-LM and run a short chat. The commands do not convert Intel Mac setups, CUDA-only code or models without MLX weights. The tool names below matter because each plays a different part in that workflow.

A few blocks remain in a wooden tray beside a closed laptop, while other blocks sit in small bowls.
The operating system and open apps use the same memory too.

The four names refer to different jobs

In this workflow, MLX provides compute on Apple silicon’s CPU and GPU, while MLX-LM is the Python package that loads a language model for inference11 or fine-tuning. To start a short chat, you also need a checkpoint MLX-LM can read and a command to run it. An original model with a familiar name is not enough if its repository lacks MLX weights or a compatible setup.

MLX-LM is a Python package for inference and fine-tuning of language models. You can specify a model and run it once in a terminal, or open a chat session with mlx_lm.chat. A useful distinction is that MLX provides the compute framework while MLX-LM is a tool that loads language models. If the Hugging Face repository you choose does not contain MLX weights and a compatible setup, the existence of a similarly named original model does not guarantee that it will run.

Metal is Apple’s low-level technology for graphics and parallel computation on macOS. It is part of the execution foundation MLX uses for GPU operations on a Mac; Metal itself is neither a model catalog nor a chat interface. PyTorch MPS, in contrast, is a backend that lets PyTorch12 use Apple GPUs. Even on the same Mac, a PyTorch model and an MLX model may use different packages, weight formats and runtime13 paths.

An app with a graphical interface, such as LM Studio, is another category of option. A GUI app may guide you through downloading a model and chatting, and may include or let you select an engine such as MLX. When comparing terminal-based MLX-LM with a GUI app, first check whether it supports the model format you need, which settings it exposes and how it manages conversations. Features and model support can vary by app version.

A notebook, a laptop turned away, a metal ruler and hand tool, another notebook and a closed folder sit on a wooden desk.
A framework, model tool and GPU execution foundation are different layers.

Why you should check the chip first

The Mac installation and optimization path for MLX assumes Apple Silicon. Check the chip listed in About This Mac. If it shows an Intel processor, following Apple Silicon MLX commands will not give you the same execution path. In that case, compare tools in the llama.cpp family that support GGUF14, or a GUI engine that explicitly supports Intel Macs. The product name ‘Mac’ alone does not mean the same acceleration features are available.

CUDA may also come to mind. If a CUDA command uses .to("cuda") or names an NVIDIA GPU, changing that part to mlx does not make the code usable on Apple Silicon. The MLX project also documents a Linux CUDA backend, but that is a separately configured MLX execution path. It does not automatically convert a PyTorch CUDA model to an MLX model, nor does it mean CUDA-specific extensions and operations work unchanged on a Mac. Check the model implementation, weight format, kernels and dependencies separately.

Three separated work areas contain two laptops and a desktop with a visible graphics card.
Check the model files and runtime path for each machine.

Unified memory still needs room for other work

The CPU and GPU in Apple Silicon share a physical memory pool rather than using separate VRAM15 chips. MLX arrays can live in this shared memory, which makes it possible to reduce copying arrays between the CPU and GPU. This design can help when building model inference, but it does not increase the Mac’s installed memory. macOS and other apps use memory first, and model weights, intermediate calculations and the KV cache that grows during a conversation use the same resource.

A Mac with 24 GB of unified memory cannot assign all 24 GB to model weights. Looking only at the space needed when the model loads is not enough, either. A long question needs more working space while the input is processed, and the KV cache grows as the answer is generated. If a browser or image editor is open too, less capacity remains for the model. So do not decide that a model will work on a Mac from its quantized file size alone; consider actual memory pressure with other apps open and the length of the task.

Quantization16 represents model values with fewer bits and can reduce memory used by the weights. A quantized model still needs temporary memory while loading, space for calculations and room for the KV cache. Even if two files use the same model name and 4-bit label, their results and requirements can differ with the quantization method, model architecture and runtime tool. If memory is tight, choosing a smaller compatible checkpoint may help, but check the resulting answer quality and supported features yourself.

Check that your model has an MLX path

Open the model card and check for an MLX checkpoint. Then check the architecture in MLX-LM’s supported-model list and, if needed, run a short text-generation test. A PyTorch version on Hugging Face does not establish that an MLX version exists. Before using a community conversion, check its conversion date, source model version, tokenizer settings and license. If you need image input or tool calls17, test each feature separately after basic generation succeeds.

MLX-LM supports many language models, but not every language model. Even within the same family, a model card may require a special tokenizer implementation or permission to execute remote code. Review the source and code before enabling trust for code you do not know, and leave that permission off when it is not needed. If support for an architecture or feature is unclear, first verify basic text generation, then check features such as image input or tool calls against that repository and the current MLX-LM version.

Choose a terminal or GUI for the task

For someone who wants to try a model quickly or load it directly from Python, MLX-LM is a straightforward starting point. You can choose a model listed as a default in its official repository and run it with mlx_lm.generate or the interactive mlx_lm.chat. If you prefer selecting models by clicking and managing local API18 connections or conversation history in one place, a GUI app that supports an MLX engine may be more convenient. But a graphical app that uses MLX internally does not necessarily expose every MLX-LM option or every Hugging Face model.

For a developer already using PyTorch MPS, the natural next step may be checking whether the existing PyTorch code and model work on MPS. The same hardware does not let us rank MLX and PyTorch MPS performance in advance. Model implementation, operations, library versions, input size and memory state can change the result. Compare the same model family, similar quantization and the same input and output lengths where possible, while recognizing that implementation differences remain between the two paths. If you cannot match conditions, record what differed. Comparing tok/s from different models or settings can confuse a condition difference with a runtime speed difference.

If your aim is to move a CUDA project built for Linux and an NVIDIA GPU directly to a Mac, do not expect MLX to convert it automatically. First look for a supported MLX checkpoint and an example for running it. If none exists, keep the CUDA environment or check another path that supports the model, such as llama.cpp or a GUI app. MLX-LM documentation alone also cannot establish support for an image or speech model. The MLX ecosystem has separate tools beyond language models, and each tool’s supported architectures and features need to be checked separately.

If you already have an Apple Silicon Mac and want to try a supported language model from a terminal or Python, MLX-LM is a practical starting point. Confirm that the model has an MLX checkpoint and that memory remains available after accounting for quantization, context length19 and the apps you plan to keep open. If you want installation and model selection on screen, look at a GUI app that offers an MLX engine. An Intel Mac cannot follow the Apple Silicon MLX guide unchanged, and CUDA models and scripts do not move automatically. Check the chip, model checkpoint, required features and actual free memory before choosing a tool.

Terminology notes

  1. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  2. MLX-LM — A package for loading, running, and fine-tuning language models with MLX. It is distinct from the MLX framework and from other MLX-based apps.

    Back to the text
  3. Metal — Apple’s low-level technology for graphics and parallel GPU computation. It is not itself a model-selection or chat app.

    Back to the text
  4. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  5. PyTorch MPS — The PyTorch backend for computation on Apple GPUs. It is separate from MLX, and required operations and models must be supported on MPS.

    Back to the text
  6. Unified memory — An architecture where the CPU and GPU share one physical memory pool. It does not increase total memory capacity; available capacity depends on the system.

    Back to the text
  7. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  8. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  9. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text
  10. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  11. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  12. PyTorch — A software framework for building and running AI models. Check the compatible PyTorch version and hardware support along with the model.

    Back to the text
  13. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  14. GGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.

    Back to the text
  15. VRAM — Memory used by a graphics card’s GPU for model weights and intermediate values. It is distinct from system RAM.

    Back to the text
  16. Quantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.

    Back to the text
  17. Tool call — A structured request from a model for an external function such as reading a file, searching, or running a command. The agent runtime and its permission settings decide whether the request is actually executed.

    Back to the text
  18. API — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.

    Back to the text
  19. Context window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.

    Back to the text