There are two modes:
| Mode | Split | Main use |
|---|---|---|
| Tensor parallelism | Routed experts and per-layer work across two Macs, or two Sparks for V4.1 Q2 | Resident inference with lower per-token work on each GPU |
| Pipeline parallelism | Complete layer ranges across several machines | Fit larger models and overlap long prefills |
These are separate from tensor parallelism across CUDA cards. Network protocols have no authentication or encryption. Use trusted machines and a trusted network; run the same commit on every peer. Model paths and artifacts must agree. Update all TP peers together when changing versions.
This is a 50/50 split with exactly one worker. Do not pass --layers.
Routed experts are sharded; attention partitioning depends on the model. Both
GPUs work on the same token and exchange partial results. This can reduce generation
latency, but the gain depends on the model, link, and comparison setup.
Two 128 GB Macs are useful for V4 Flash Q4/MXFP4 or GLM 5.3 Flash Q4. GLM 5.2 IQ2_XXS is another tested capacity setup. V4.1 Flash Q2 also runs on two 128 GB Macs, with disk-only Engram tables. A larger quant may need larger machines even though its tensor layout is supported.
Use a Thunderbolt cable. RDMA requires an active verbs device with an IPv4-mapped GID; a working ping alone does not establish that.
rdma_ctl status
ibv_devinfo -vAddresses must be on the cabled member interfaces, not only the Thunderbolt bridge. For example, after checking which interfaces are active:
# Machine A, example member interface en1.
sudo ifconfig en1 inet 10.99.0.2/30 alias
# Machine B, example member interface en6.
sudo ifconfig en6 inet 10.99.0.1/30 aliasFor the large tested shards on otherwise idle 128 GB Macs, the setup raised the per-boot GPU wired-memory limit on both machines:
sudo sysctl iogpu.wired_limit_mb=120000This is specific to that memory configuration. It grants a larger GPU budget; it does not create more RAM. Check model, context, and system headroom before raising a limit on your machine.
Download the same model on both machines. For GLM 5.3 Flash Q4:
./download_model.sh glm53-q4Start the worker first; it retries while the coordinator loads:
# Machine B.
./ds4 --tensor-parallel --role worker \
--coordinator 10.99.0.2 9911 --transport rdma --ctx 8192
# Machine A.
./ds4 --tensor-parallel --role coordinator \
--listen 10.99.0.2 9911 --transport rdma --ctx 8192The verbs device and GID are selected automatically. If ambiguous, specify
--rdma-device and --rdma-gid-index from ibv_devinfo. Use --transport tcp
on both peers when RDMA is unavailable. Do not probe the waiting coordinator
with curl or nc: it may treat the connection as a worker handshake.
Keep workers running in a terminal or managed session and retain both logs. Do not treat repeated handshake or RDMA timeouts as successful QA merely because a retry works.
The coordinator can be ds4, ds4-agent, ds4-server, or ds4-bench;
workers run ds4. For models with vision support, pass the same --vision FILE
to both for image input.
For GLM MTP, enable --mtp on both. For DeepSeek DSpark, both need the
matching support model and DSpark options. V4.1 supports vision but not
speculative decoding.
TP disk-cache restore currently rebuilds the exact saved token prefix on both ranks rather than restoring the coordinator alone. Expect prefill on restore. See speculation and serving.
V4.1 Flash Q2 text inference supports one GPU per rank, with a 50/50 expert
split. Each Spark holds about 81 GiB of weights, plus context and runtime
buffers. Engram tables stay on disk. Do not add --ssd-streaming or
--cuda-tensor-parallel: those select different memory/execution modes.
Build the same commit with make cuda-spark on both machines and download
ds41f-q2 on both. RDMA needs the libibverbs development headers at build time,
its runtime library, and an active RoCEv2 link. Check ibv_devinfo -v and use
the addresses of the directly connected ports, not the management network.
For example, with the coordinator at 172.31.250.1 on the direct link:
# Worker.
./ds4 --cuda -m gguf/DeepSeek-V4.1-Flash-Q2.gguf --ctx 32768 \
--tensor-parallel --role worker --coordinator 172.31.250.1 9911 \
--transport rdma
# Coordinator.
./ds4 --cuda -m gguf/DeepSeek-V4.1-Flash-Q2.gguf --ctx 32768 \
--tensor-parallel --role coordinator --listen 172.31.250.1 9911 \
--transport rdmaThe link address selects the matching verbs device and GID. If selection is
ambiguous, use --rdma-device and --rdma-gid-index. The transport stages
CUDA results through host memory; this is RoCE, not GPUDirect. TCP is also
available with --transport tcp on both peers.
The coordinator can be ds4-agent, ds4-server or ds4-bench. Match context
sizes on both sides. With five or more ready sessions, the server batches decode
across both GPUs. Smaller groups run in order, which is faster at those sizes.
Use --batched-session 8 on the server; allow memory for all eight contexts.
Vision, DSpark and other model/quant layouts are not supported by CUDA network TP.
Each process maps only its assigned layers, retaining that slice of the KV
state. Layer ranges are inclusive. N:output includes the final layer and
output head. Activations travel from one stage to the next over TCP.
For Flash Q4 on two machines, download ds4f-q4 on both, then start each side.
Replace the example address with your coordinator's reachable address:
# Machine A.
./ds4 --role coordinator --layers 0:19 --listen 10.99.0.2 9911
# Machine B.
./ds4 --role worker --layers 20:output --coordinator 10.99.0.2 9911Normally give the output head to the final worker. With several workers, choose non-overlapping ranges covering the entire model. Workers register their ranges with the coordinator; intermediate workers forward activations directly to the next stage.
Long prefill chunks can occupy different stages simultaneously. A single generation stream cannot use that overlap: each token must finish the route before the next one is sampled. Use pipeline mode primarily for capacity and long-prefill throughput, not as a guaranteed decode speedup.
For two 512 GB Mac Studios, use the split artifacts:
# Machine A.
./download_model.sh pro-q4-layers00-30
./ds4 -m gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf \
--role coordinator --layers 0:30 --listen 10.99.0.2 9911
# Machine B.
./download_model.sh pro-q4-layers31-output
./ds4 -m gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf \
--role worker --layers 31:output --coordinator 10.99.0.2 9911These downloads do not change ds4flash.gguf. Startup is expensive because
each side must make its model slice resident.
Keep default chunk sizes first. --dist-prefill-window N controls the number
of chunks in flight; --dist-prefill-chunk N overrides the session-derived
chunk size. --debug shows route and per-hop timings.
Activations use 32-bit transport by default. --dist-activation-bits 16 halves
the payload; 8 is more aggressive. These change numerical precision on the
wire, not weights or KV storage. Validate output when changing them.
A disconnected worker invalidates the route. In-flight work can fail; later requests need a complete route before proceeding, and the coordinator can replay the saved token prefix to rebuild worker state. Pipeline snapshots serialize all layer slices into one payload and redistribute them when loaded.
For protocol details, see ds4_distributed.c and ds4_tp.c.