Model setup recipes

MLX-JACCL on Two Macs: Distributed Setup and Validation

Check the distributed-run prerequisites before connecting Macs.

Connecting Macs does not automatically combine their memory into one pool. MLX1 JACCL provides inter-device communication when OS, Thunderbolt/RDMA topology, host configuration, and the model’s distributed path are supported. Configure two hosts, verify a two-rank smoke test, then inspect model placement and compare against a single-Mac baseline using the same workload.

Requirements and key details
  • JACCL is a communication backend, not automatic support for every Mac, OS, or network topology.
  • Match MLX/MLX-LM versions, macOS, network interfaces, and model sharding support.
  • Inspect logs to confirm device placement before measuring throughput and latency.
  • Compare one and two Macs with the same prompt and model settings.

First, distinguish distributed execution from pooled memory

It is tempting to think that if a large model file will not fit on one Mac, connecting two Macs simply adds their memory together. Distributed inference2 instead places parts of a model’s tensors on each device and exchanges data between devices during computation. The result depends on which tensors go where, whether the runtime3 supports that partition, and whether communication costs erase the compute gains. Adding storage or unified-memory figures does not predict whether a model will run.

This guide checks the JACCL backend in Apple MLX and the distributed-generation path in MLX-LM4. JACCL is a communication layer, not a model-sharding strategy by itself. MLX 0.30.6 announced JACCL and MLX-LM distributed-server changes, but it does not provide a universal speedup or a compatibility table for every Mac combination. This is a setup and validation procedure, not a reproduced benchmark.

Before starting, confirm that both Macs use Apple silicon, can reach each other over the same local network, and have MLX/MLX-LM versions that provide the JACCL path. Also verify macOS 26.3 or later, the requirement stated for the MLX 0.30.6 JACCL update. On managed devices or public Wi-Fi, do not change network permissions or firewall policy without approval; follow your administrator’s instructions.

Two Mac desktops connected by a short cable
A cable connection does not by itself mean inference is distributed.

Check both Macs’ versions and network first

The JACCL path connects two Mac minis with Thunderbolt cables, creates a hostfile with MLX’s RDMA setup tool, then uses `mlx.launch` for a small distributed test followed by an MLX-LM example. The official MLX guide says automatic setup checks SSH between nodes, RDMA availability, and the Thunderbolt mesh. If the connection is not a full mesh, the network is Ethernet-only, or the OS is unsupported, do not force this JACCL path; use MLX’s documented ring/Ethernet approach or return to one Mac.

The two Macs must resolve each other’s hostnames, allow passwordless SSH, and have Python and the test script at the same path. JACCL’s `--auto-setup` requires passwordless sudo on each machine and changes Thunderbolt network configuration. On managed Macs where that is not permitted, omit `--auto-setup`, review the commands it prints, and get administrator approval. The installed MLX must meet the macOS 26.3-or-later requirement for JACCL.

The procedure below can change network and kernel configuration, so run it on dedicated test machines. Do not start with a large model; first confirm that the process group has two ranks. Do not move on to real inference until both logs show ranks 0 and 1 and `size=2`.

Check both Macs and configure the JACCL connection
# Run on both Macs; replace the hostnames with names configured for SSH.
sw_vers
python3 --version
python3 -m pip show mlx mlx-lm
ssh mac-mini-2 'python3 --version && python3 -m pip show mlx mlx-lm'
mlx.distributed_config --verbose --hosts mac-mini-1,mac-mini-2 --over thunderbolt --dot
mlx.distributed_config --verbose --backend jaccl --hosts mac-mini-1,mac-mini-2 --over thunderbolt --auto-setup --output hosts.json
Replace the hostnames with your SSH hosts. Auto-setup changes Thunderbolt network configuration; use dedicated test machines.

Verify distributed execution with a small test first

Once the hostfile is created, start with a small rank check using the format in the official MLX docs, and confirm both machines have the same Python and test script. JACCL needs RDMA-device information connecting the two ranks in the hostfile. If you edit the generated file, verify the actual RDMA device names and rank order for your cabling.

Begin with a small runnable model and a short prompt. Check logs to see where model files are downloaded, which rank each process starts as, and whether all hosts connect. Producing an answer is not sufficient proof that distributed execution worked. Verify per-device memory use and placement. If an error occurs, stop both processes and compare ports, addresses, and versions one item at a time.

If the rank check prints `rank=0 size=2` and `rank=1 size=2`, it confirms that a two-process distributed group was created—but it does not yet prove that the model was sharded. Next run the MLX-LM chat example from the MLX distributed guide and inspect logs to see whether the model is split across both devices. Confirm from official docs that the model architecture supports the requested distributed partition; a launcher starting successfully is not proof of sharding.

Verify ranks, then run the official MLX-LM example
mlx.launch --verbose --backend jaccl --hostfile hosts.json -- python3 -c 'import mlx.core as mx; g=mx.distributed.init(backend="jaccl"); print(f"rank={g.rank()} size={g.size()}")'
mlx.launch --verbose --backend jaccl --hostfile hosts.json -- python3 -m mlx_lm chat --model mlx-community/DeepSeek-R1-0528-4bit
The first command verifies that both ranks start. The second follows the distributed chat example in the MLX guide. Use node logs to confirm that this model/runtime combination supports sharding and how the model is placed.
Model-weight blocks and a connection are shown across two computers
Verify device participation and communication in both hosts’ logs.

Compare one and two Macs with the same prompt

When creating a baseline, keep model version and weights, quantization5, prompt, maximum output length, and sampling settings fixed. Run once on one Mac and once through the two-Mac distributed path. Record the first model load separately as a cold start because it includes download and initialization; repeat requests after setup in a warm state.

Record more than total completion time. Include time to first token6, output-token7 count and generation rate, memory and GPU/CPU use on each Mac, network traffic, and errors or retries. If answer content differs, check that the random seed, sampling, tokenizer, and template match before comparing speed. Distributed execution changes both compute placement and communication, so gains can vary with input and output length.

For example, suppose a short answer takes longer when it crosses the network, while a difference appears only on long answers. In that case, do not decide based on one average; record short interactive questions and long batch jobs separately. This is an illustrative assumption about how to compare workloads, not a measurement from MLX.

Check execution feasibility separately from operational benefit.
StageRecordPass criterion
ConnectionOS, package, address, and rank logsBoth processes discover each other
PlacementPer-device memory and workload logsIntended model split and device participation are confirmed
CorrectnessOutput and token count for the same promptAnswer quality and completion state are acceptable
PerformanceTTFT, completion time, memory, and transfer volumeThe workload benefit outweighs operational complexity

Check execution feasibility separately from operational benefit.

Connection

Record
OS, package, address, and rank logs
Pass criterion
Both processes discover each other

Placement

Record
Per-device memory and workload logs
Pass criterion
Intended model split and device participation are confirmed

Correctness

Record
Output and token count for the same prompt
Pass criterion
Answer quality and completion state are acceptable

Performance

Record
TTFT, completion time, memory, and transfer volume
Pass criterion
The workload benefit outweighs operational complexity
A model loaded on one computer compared with blocks split across two
Compare distributed execution with a single-Mac baseline on the same workload.

If it fails, isolate network and sharding conditions in order

If the other machine is not discovered, check that both Macs are currently on the same network, that the router allows device-to-device traffic, and that firewall, VPN, or security software is not blocking connections. Do not simply disable the firewall; check operating-system block logs or administrator policy. Find port numbers in the current MLX command and official docs rather than opening ports by guesswork.

If connections work but initialization fails, compare package versions, macOS requirements, rank count, and address configuration. If memory is insufficient, separate the actual weight size and per-device placement from context/KV-cache and other-app use. A faster network alone does not distribute weights evenly across devices.

If the run works but is slow, identify which stage is delayed. When a model is small enough that communication costs exceed the compute savings—or the prompt is so short that distributed setup dominates—one Mac may be faster. If the distributed server is not stable or speed differences are inconsistent, keep the single-Mac path as the operational default.

Official implementation docs and release notes

This guide does not include device-specific performance measurements. Check JACCL support and runtime options in the official docs, release notes, and help output for your installed MLX/MLX-LM versions. Revalidate with a small example if macOS or the network environment changes.

Terminology notes

  1. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  2. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  3. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  4. MLX-LM — A package for loading, running, and fine-tuning language models with MLX. It is distinct from the MLX framework and from other MLX-based apps.

    Back to the text
  5. Quantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.

    Back to the text
  6. Time to First Token — The time from sending a request until its first output token arrives. It is separate from the generation rate of later tokens.

    Back to the text
  7. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text