Model setup recipes
MLX-JACCL on Two Macs: Distributed Setup and Validation
Check the distributed-run prerequisites before connecting Macs.
Connecting Macs does not automatically combine their memory into one pool. MLX1 JACCL provides inter-device communication when OS, Thunderbolt/RDMA topology, host configuration, and the model’s distributed path are supported. Configure two hosts, verify a two-rank smoke test, then inspect model placement and compare against a single-Mac baseline using the same workload.
First, distinguish distributed execution from pooled memory
It is tempting to think that if a large model file will not fit on one Mac, connecting two Macs simply adds their memory together. Distributed inference2 instead places parts of a model’s tensors on each device and exchanges data between devices during computation. The result depends on which tensors go where, whether the runtime3 supports that partition, and whether communication costs erase the compute gains. Adding storage or unified-memory figures does not predict whether a model will run.
This guide checks the JACCL backend in Apple MLX and the distributed-generation path in MLX-LM4. JACCL is a communication layer, not a model-sharding strategy by itself. MLX 0.30.6 announced JACCL and MLX-LM distributed-server changes, but it does not provide a universal speedup or a compatibility table for every Mac combination. This is a setup and validation procedure, not a reproduced benchmark.
Before starting, confirm that both Macs use Apple silicon, can reach each other over the same local network, and have MLX/MLX-LM versions that provide the JACCL path. Also verify macOS 26.3 or later, the requirement stated for the MLX 0.30.6 JACCL update. On managed devices or public Wi-Fi, do not change network permissions or firewall policy without approval; follow your administrator’s instructions.

Check both Macs’ versions and network first
The JACCL path connects two Mac minis with Thunderbolt cables, creates a hostfile with MLX’s RDMA setup tool, then uses `mlx.launch` for a small distributed test followed by an MLX-LM example. The official MLX guide says automatic setup checks SSH between nodes, RDMA availability, and the Thunderbolt mesh. If the connection is not a full mesh, the network is Ethernet-only, or the OS is unsupported, do not force this JACCL path; use MLX’s documented ring/Ethernet approach or return to one Mac.
The two Macs must resolve each other’s hostnames, allow passwordless SSH, and have Python and the test script at the same path. JACCL’s `--auto-setup` requires passwordless sudo on each machine and changes Thunderbolt network configuration. On managed Macs where that is not permitted, omit `--auto-setup`, review the commands it prints, and get administrator approval. The installed MLX must meet the macOS 26.3-or-later requirement for JACCL.
The procedure below can change network and kernel configuration, so run it on dedicated test machines. Do not start with a large model; first confirm that the process group has two ranks. Do not move on to real inference until both logs show ranks 0 and 1 and `size=2`.
# Run on both Macs; replace the hostnames with names configured for SSH.
sw_vers
python3 --version
python3 -m pip show mlx mlx-lm
ssh mac-mini-2 'python3 --version && python3 -m pip show mlx mlx-lm'
mlx.distributed_config --verbose --hosts mac-mini-1,mac-mini-2 --over thunderbolt --dot
mlx.distributed_config --verbose --backend jaccl --hosts mac-mini-1,mac-mini-2 --over thunderbolt --auto-setup --output hosts.jsonVerify distributed execution with a small test first
Once the hostfile is created, start with a small rank check using the format in the official MLX docs, and confirm both machines have the same Python and test script. JACCL needs RDMA-device information connecting the two ranks in the hostfile. If you edit the generated file, verify the actual RDMA device names and rank order for your cabling.
Begin with a small runnable model and a short prompt. Check logs to see where model files are downloaded, which rank each process starts as, and whether all hosts connect. Producing an answer is not sufficient proof that distributed execution worked. Verify per-device memory use and placement. If an error occurs, stop both processes and compare ports, addresses, and versions one item at a time.
If the rank check prints `rank=0 size=2` and `rank=1 size=2`, it confirms that a two-process distributed group was created—but it does not yet prove that the model was sharded. Next run the MLX-LM chat example from the MLX distributed guide and inspect logs to see whether the model is split across both devices. Confirm from official docs that the model architecture supports the requested distributed partition; a launcher starting successfully is not proof of sharding.
mlx.launch --verbose --backend jaccl --hostfile hosts.json -- python3 -c 'import mlx.core as mx; g=mx.distributed.init(backend="jaccl"); print(f"rank={g.rank()} size={g.size()}")'
mlx.launch --verbose --backend jaccl --hostfile hosts.json -- python3 -m mlx_lm chat --model mlx-community/DeepSeek-R1-0528-4bit
Compare one and two Macs with the same prompt
When creating a baseline, keep model version and weights, quantization5, prompt, maximum output length, and sampling settings fixed. Run once on one Mac and once through the two-Mac distributed path. Record the first model load separately as a cold start because it includes download and initialization; repeat requests after setup in a warm state.
Record more than total completion time. Include time to first token6, output-token7 count and generation rate, memory and GPU/CPU use on each Mac, network traffic, and errors or retries. If answer content differs, check that the random seed, sampling, tokenizer, and template match before comparing speed. Distributed execution changes both compute placement and communication, so gains can vary with input and output length.
For example, suppose a short answer takes longer when it crosses the network, while a difference appears only on long answers. In that case, do not decide based on one average; record short interactive questions and long batch jobs separately. This is an illustrative assumption about how to compare workloads, not a measurement from MLX.
| Stage | Record | Pass criterion |
|---|---|---|
| Connection | OS, package, address, and rank logs | Both processes discover each other |
| Placement | Per-device memory and workload logs | Intended model split and device participation are confirmed |
| Correctness | Output and token count for the same prompt | Answer quality and completion state are acceptable |
| Performance | TTFT, completion time, memory, and transfer volume | The workload benefit outweighs operational complexity |
Check execution feasibility separately from operational benefit.
Connection
- Record
- OS, package, address, and rank logs
- Pass criterion
- Both processes discover each other
Placement
- Record
- Per-device memory and workload logs
- Pass criterion
- Intended model split and device participation are confirmed
Correctness
- Record
- Output and token count for the same prompt
- Pass criterion
- Answer quality and completion state are acceptable
Performance
- Record
- TTFT, completion time, memory, and transfer volume
- Pass criterion
- The workload benefit outweighs operational complexity

If it fails, isolate network and sharding conditions in order
If the other machine is not discovered, check that both Macs are currently on the same network, that the router allows device-to-device traffic, and that firewall, VPN, or security software is not blocking connections. Do not simply disable the firewall; check operating-system block logs or administrator policy. Find port numbers in the current MLX command and official docs rather than opening ports by guesswork.
If connections work but initialization fails, compare package versions, macOS requirements, rank count, and address configuration. If memory is insufficient, separate the actual weight size and per-device placement from context/KV-cache and other-app use. A faster network alone does not distribute weights evenly across devices.
If the run works but is slow, identify which stage is delayed. When a model is small enough that communication costs exceed the compute savings—or the prompt is so short that distributed setup dominates—one Mac may be faster. If the distributed server is not stable or speed differences are inconsistent, keep the single-Mac path as the operational default.
Official implementation docs and release notes
This guide does not include device-specific performance measurements. Check JACCL support and runtime options in the official docs, release notes, and help output for your installed MLX/MLX-LM versions. Revalidate with a small example if macOS or the network environment changes.
Terminology notes
MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.
Back to the textInference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the textMLX-LM — A package for loading, running, and fine-tuning language models with MLX. It is distinct from the MLX framework and from other MLX-based apps.
Back to the textQuantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.
Back to the textTime to First Token — The time from sending a request until its first output token arrives. It is separate from the generation rate of later tokens.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the text