Model setup recipes
Run a 3B model with MLX-LM and read its speed: loading, prefill, decode and KV cache
Downloading and loading a model, reading a long prompt and generating an answer are different waits. Record them under the same conditions.
Run mlx-community/Llama-3.2-3B-Instruct-4bit, the model named as the MLX-LM1 repository’s default, in a virtual Python environment on an Apple Silicon Mac. The 3B 4-bit model is a documented starting point, not a guarantee that it fits every Mac. Separate the initial download and model load, prefill2 for reading a long prompt, decode3 for generating tokens4, and the KV cache5 that grows as the conversation continues. Keep the model, prompt and output limit fixed, and record cold starts separately from warm sessions.
If the first run is slow, is the model slow?
The first time you run a model on an Apple Silicon Mac, it may download the files, load the weights, process the input and then generate an answer. If all of that appears as one wait, it is hard to know which part took the time. When you ask a follow-up while the model is still running, the downloaded files and active process can be reused. So separate the first-run wait from the speed of the next answer in an ongoing conversation.
This guide uses mlx-community/Llama-3.2-3B-Instruct-4bit, which the official MLX-LM repository names as its default model. Its model card and configuration identify a Llama 3.2 3B Instruct architecture and 4-bit quantization6. This model ID is available for reproducing an example listed in the official documentation. Behavior from this example does not represent the performance of all other models; larger models and other architectures need separate checks.
The English prompt in the command is only to check that installation and basic generation work. It is not a test of Korean answer quality, so evaluate the model separately with the language and material you plan to use.
First create an isolated environment that supports MLX-LM
MLX-LM is a Python package, so installing it in a project-specific virtual environment7 keeps it separate from other Python work. The commands below create that environment, install the official package and check basic generation with a short input. If installation fails, check the current official guide and your environment.
The current official MLX8 installation guide requires Apple Silicon, a native Python 3.10 or later, and macOS 14.0 or later. Python running as an Intel process through Rosetta is not a native environment, so check the Python architecture if no package is found.
The command’s --max-tokens 128 sets an upper limit on generated answer tokens. It does not set a 128-token context limit that includes the prompt. Changing it also changes the permitted answer length and total generation time, so leave it fixed during an initial comparison. When this command ends, running it again starts a new process. To observe follow-up responses in a session where the model remains loaded, run mlx_lm.chat and ask questions within that one process.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install mlx-lm
mlx_lm.generate --model mlx-community/Llama-3.2-3B-Instruct-4bit --prompt "Explain why prompt length affects generation time in two sentences." --max-tokens 1283B 4-bit is a starting size, not a memory guarantee
The 3B in the model name indicates its approximate parameter scale; 4-bit describes the precision used to represent the weights. Quantization can reduce memory needed for the weights, but running the model also requires metadata and temporary working space. Generating an answer adds a KV cache, and a longer input can require more memory during prefill. Disk space for downloading the files and unified memory9 during execution are different resources.
Do not turn this into a rule such as ‘3B means it runs on an 8 GB Mac.’ Macs differ in unified-memory capacity and the amount used by macOS, and open programs such as browsers, development tools and video-call apps also take space. Before running the model, check memory pressure in Activity Monitor and see which apps are open. Save any work, then begin with a short prompt. If memory pressure remains high or an app freezes, shorten the context and close other apps to check again. If the process crashes, first free more memory or choose a smaller compatible model instead of repeating the same run.

Reading the prompt and writing the answer take different time
The model first reads the prompt as tokens and computes internal representations. This input-processing stage is called prefill. It explains why a request to summarize a long document may take longer to process than a short question. The model then generates output tokens one by one in the decode stage. A long document can delay the first token, while a request for a long answer can keep generation running for longer. Both may feel like a ‘slow response,’ but they have different causes and settings to adjust.
For repeated trials, keep the model, prompt and output limit the same. Record the first run separately as a cold start that includes downloading the files and loading the model. Then observe a first question and follow-up question in the same live mlx_lm.chat session to distinguish the wait from restarting the process from work that continues within a session. However, a follow-up may include earlier conversation context, so it is not the same prefill trial if its input length differs. Do not mix values from a fresh session and an ongoing session in one column.
If generation reports prompt-processing and generation speeds, record each in its own input-token or output-token terms. Do not write down only the speed number: note whether the model was newly loaded, how long the input was, the output limit and which other apps were open. Also record the MLX-LM version, macOS, Mac chip and memory configuration so you can recreate the conditions later. Before comparing with a public benchmark, check whether the model ID, quantization, input length and measured interval match.

Context and KV cache change the memory trade-off
Even with the same model, the KV cache—which stores key and value information from earlier tokens—can grow as a conversation gets longer. A successful short question therefore does not prove that the same memory headroom will remain when you add several long documents. Output length matters too. Allowing a long answer adds generated tokens over time, increasing the cache and work time. Start with a short input and short output, then increase one condition at a time to see where pressure begins.
MLX-LM’s --prefill-step-size controls the chunk size used to process long prompts, while --max-kv-size sets a limit for supported cache configurations. The official documentation explains that reducing the prefill chunk can lower peak memory while processing the input, but may slow prompt processing. Lowering the limit for a rotating KV cache can reduce RAM10 use, but may affect answer quality if older context is not retained. The values below are examples of how to use the options, not universal recommendations for every Mac or model.
When comparing settings, do not change both values at once. First adjust only --prefill-step-size and check prefill time and memory pressure. Return to the original setting, then change --max-kv-size separately and compare context retention and answer quality on the same question. These are not performance buttons to enable automatically; they are trade-offs among memory, processing time and context retention.
If you need to preserve a long conversation, first establish a baseline with a shorter input before limiting the cache, then compare answer quality with the same question. A short or interrupted answer, or one that forgets an earlier instruction, is not evidence of a speed improvement. If the result no longer meets the context or quality you need, restore the cache limit and consider a smaller model or a machine with more memory headroom.
mlx_lm.generate --model mlx-community/Llama-3.2-3B-Instruct-4bit --prompt "Summarize this short note." --max-tokens 128 --prefill-step-size 512 --max-kv-size 4096How do you decide whether the speed suits your work?
Judge the waits you actually encounter often rather than relying on one total run time. If you frequently close the app or restart the computer and launch a fresh model process, you will often wait for the weights to load again. If you keep the model running in one session and ask many questions, you will more often notice prefill and decode on follow-ups. Asking for a one- or two-sentence answer to a long report and requesting a long draft from a short question also spend time in different stages. One average speed figure cannot describe all these jobs.
If practical, run the same conditions three times after warm-up and use the median. This reduces the influence of a single unusually fast or slow run. Record cold start, warm session, prefill, decode and peak memory separately, along with the model, input, output limit, version and other open apps. You can then repeat the same comparison after changing a setting.
In this example, start by completing a short chat with MLX-LM’s official 3B 4-bit checkpoint11. If memory pressure stays stable, increase the real context or output length one at a time and check both answer quality and memory after each change. If memory runs short in the short test or your required model is unsupported, stop there and choose a smaller supported model or another engine. To compare speed, record initial loading, follow-up responses within the session and available memory separately with the same input. Those results show whether the setup can handle your work.

Terminology notes
MLX-LM — A package for loading, running, and fine-tuning language models with MLX. It is distinct from the MLX framework and from other MLX-based apps.
Back to the textPrefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.
Back to the textDecode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the textKV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.
Back to the textQuantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.
Back to the textPython virtual environment — An isolated space for installing Python packages per project. It helps reduce version conflicts and is not a virtual machine.
Back to the textMLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.
Back to the textUnified memory — An architecture where the CPU and GPU share one physical memory pool. It does not increase total memory capacity; available capacity depends on the system.
Back to the textSystem RAM — System memory that temporarily holds data while programs run. It differs from storage and from a discrete GPU’s VRAM.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the text