Reference · rapid-mlx 0.15.6 · ← Back to README

Hardware tiers

What to serve for your Mac's unified memory. Ask the CLI — it detects your RAM and prints both picks with the command to run them:

$ rapid-mlx recipe
Recommended for this 18.0 GB Mac (18 GB tier)

1. Smart — qwen3.5-9b-4bit
   8.7 GB RAM · 82% capability · ~36 tok/s
   rapid-mlx serve qwen3.5-9b-4bit

2. Fast — qwen3.5-4b-4bit
   6.0 GB RAM · 78% capability · ~61 tok/s
   rapid-mlx serve qwen3.5-4b-4bit

Every tier has a smart pick (most capable model that fits) and a fast pick (highest throughput). recipe --max-ram 32 shows another tier; the same picks drive install.sh and the desktop app. The table below has every tier, with decode speeds that each name the machine, engine version and date they were measured on.

Source of truth. One table, two front doors. The picks below mirror the engine's recommendation policy in rapid-mlx 0.15.6 (rapid_mlx/model_recommendations.json) — the same ordered list install.sh's quick-start banner prints for your RAM bracket (smart first, fast second) and the same picks the desktop app's picker surfaces. Numbers come from three measured runs, each labeled: the October 2026 M4 Pro 48 GB run (0.15.4), the installer's recommendation qualification on an M2 Pro 32 GB (0.12.7, 2026-08-07, marked †) and an M3 Ultra 256 GB run (0.12.15, 2026-08-18, marked ‡). A number is never copied between machines.

The RAM → picks map

RAM tier Smart pick Fast pick Measured decode
48 GB+ qwen3.8-27b-4bit qwen3.6-35b-4bit 24.1 / 104.5 tok/s §
32 GB qwen3.8-27b-4bit qwen3.5-4b-4bit 24.1 / 85.8 tok/s §
24 GB bonsai-27b-2bit qwen3.5-4b-4bit 17.5 ‡ / 85.8 tok/s §
18 GB qwen3.5-9b-4bit qwen3.5-4b-4bit 63.0 / 85.8 tok/s §
16 GB qwen3.5-4b-4bit lfm2.5-1b-4bit 85.8 / 314.0 tok/s §
8 GB lfm2.5-2.6b-4bit lfm2.5-1b-4bit 93.5 † / 314.0 tok/s §

What the marks mean. § measured via rapid-mlx serve on a Mac mini M4 Pro 48 GB, rapid-mlx 0.15.4, 2026-10-02 — median of 3, temperature 0, raw data published (smart pick listed first, then the fast pick). † M2 Pro 32 GB, 0.12.7, 2026-08-07. ‡ M3 Ultra 256 GB, 0.12.15, 2026-08-18 — expect less on smaller chips. Where a pick has no mark its tier relies on the other row's measurement; nothing here is estimated except the one case flagged in the 48 GB section.

48 GB and up · Mac Studio / high-spec mini

Smart: rapid-mlx serve qwen3.8-27b-4bit  ·  Fast: rapid-mlx serve qwen3.6-35b-4bit

The 27B dense reasoner is the most capable model the policy places on this tier, and the 35B-A3B MoE is its speed counterpart — on the October M4 Pro run they measured 24.1 and 104.5 tok/s decode respectively, both with 4-stream aggregate close to their single-stream rate. The 35B MoE's 60 tok/s figure in the installer's own dataset is the engine's reviewed estimate, not a measurement (its only 32 GB-host run aborted with swap) — the M4 Pro number above is the measured one. Qwen3.8 27B also has an experimental tensorfold profile on this tier; it carries no fixed speed claim and stays opt-in.

Above this tier the policy stays with these two picks. The bigger MoEs are still servable — see the 192–256 GB flags.

32 GB · Mac mini / maxed MacBook

Smart: rapid-mlx serve qwen3.8-27b-4bit  ·  Fast: rapid-mlx serve qwen3.5-4b-4bit

The same 27B, one tier down: the October run measured it at 25.3 GB Metal peak through a 2K-prompt workload, which fits a 32 GB Mac with the OS and an editor running. MTP speculative decoding on by default, vision input, qwen3 thinking separation — the measured page has the full picture. The fast pick, qwen3.5-4b, measured 85.8 tok/s on the same October run.

24 GB · Mac mini / MacBook Pro

Smart: rapid-mlx serve bonsai-27b-2bit  ·  Fast: rapid-mlx serve qwen3.5-4b-4bit

Ternary Bonsai 27B, packed to 2-bit MLX: a 27B-class model at a 13 GB weight footprint (17.5 tok/s measured on the M3 Ultra qualification, ‡) is why this tier gets something far more capable than its memory would normally buy. It is slower than the other picks — the trade is quality per gigabyte, not speed. Its second-generation successor, Ternary Bonsai 2 27B, measured 24.4 GB Metal peak on the October run — that wants 32 GB, so the policy keeps the first generation here.

18 GB · MacBook Air / Pro

Smart: rapid-mlx serve qwen3.5-9b-4bit  ·  Fast: rapid-mlx serve qwen3.5-4b-4bit

Qwen 3.5 9B in 4-bit MLX — the smallest tier where a coding agent starts to work reliably: real instruction following and tool calling. The October run measured it at 63.0 tok/s (8.0 GB Metal peak); the fast pick is the same family's 4B at 85.8 tok/s. The hermes / qwen3 parser pair is auto-selected.

16 GB · MacBook Air / base MacBook Pro

Smart: rapid-mlx serve qwen3.5-4b-4bit  ·  Fast: rapid-mlx serve lfm2.5-1b-4bit

At a 6.0 GB weight footprint the 4B is the fast, general-purpose pick for a 16 GB Mac — 85.8 tok/s on the October run, the same Qwen 3.5 family and parser pair as the tier above, sized to fit alongside the OS and a browser. LFM2.5-1B is the throughput option: 314.0 tok/s decode and 618 tok/s aggregate on 4 streams in the same run, 1.8 GB peak — a basic-chat specialist.

8 GB · MacBook Air / base mini

Smart: rapid-mlx serve lfm2.5-2.6b-4bit  ·  Fast: rapid-mlx serve lfm2.5-1b-4bit

LFM2.5-2.6B, a 2.6B model whose 30 layers are 22 short-convolution blocks and only 8 GQA. Those 8 attention layers are why it belongs here: the KV cache costs roughly 16 KB per token, so a 32K conversation adds about half a gigabyte on top of 1.6 GB of weights. 128K context; 93.5 tok/s measured on the M2 Pro qualification (†).

What it is good at, and what it is not. It is post-trained for tool use and instruction following and beats models several times its size on those. Liquid publishes it as not recommended for agentic coding and knowledge-heavy tasks, and we agree — if you are driving a coding agent, this is the wrong model and the 18 GB tier (qwen3.5-9b-4bit) is where that starts working.

Its weights live in the 4bit/ folder of a repo that ships eight quantizations side by side; rapid-mlx fetches only that folder, so the download is 1.6 GB rather than the repo's ~19 GB.

Also fits: a real reasoner. ling-3.0-tiny-4bit (4.2 GB weights) is inclusionAI's 7.9B mixture-of-experts with only 1.3B active per token — thinking streamed as reasoning_content, native tool calling, 131K context. It is the strongest reasoning option that fits this tier; pick it over LFM2.5 when you want chain-of-thought rather than fast instruction following. It has no measured numbers on either of our runs, so it gets none here.

Above the table · the 192–256 GB flags

The policy's tiers stop at 48 GB, but the catalog serves bigger machines: GLM-5.3 Flash (320B/18B active), DeepSeek V4.1 Flash REAP, MiMo-V2.6 Flash (309B/15B active) and Qwen3.8 Flash-Next (180B/6B active) want a 192–256 GB Mac Studio. Their measured numbers live on their own pages — GLM-5.3 Flash, DeepSeek V4.1 Flash, MiMo-V2.6 Flash — each citing the machine, engine version and date it was measured on.

Check your machine

rapid-mlx doctor reports your chip, memory, macOS version, free disk and Hugging Face cache size, then checks packages and optional extras. Its first section on an M3 Pro 18 GB with rapid-mlx 0.15.5:

$ rapid-mlx doctor
◆ System
  ✓ Apple Silicon (Apple M3 Pro, 18 GB)
  ✓ macOS 15.6.1 (Darwin 24.6.0)
  ✓ Free disk: 94 GB

Next steps