Hardware tiers
What to serve for your Mac's unified memory. Ask the CLI —
it detects your RAM and prints both picks with the command to run them:
$ rapid-mlx recipe
Recommended for this 18.0 GB Mac (18 GB tier)
1. Smart — qwen3.5-9b-4bit
8.7 GB RAM · 82% capability · ~36 tok/s
rapid-mlx serve qwen3.5-9b-4bit
2. Fast — qwen3.5-4b-4bit
6.0 GB RAM · 78% capability · ~61 tok/s
rapid-mlx serve qwen3.5-4b-4bit
Every tier has a smart pick (most capable model that fits) and a
fast pick (highest throughput). recipe --max-ram 32
shows another tier; the same picks drive install.sh and
the desktop app. The table below has every tier, with decode speeds
that each name the machine, engine version and date they were measured
on.
rapid_mlx/model_recommendations.json) — the same ordered
list install.sh's quick-start banner prints for your RAM
bracket (smart first, fast second) and the same picks the desktop
app's picker surfaces. Numbers come from three measured runs, each
labeled: the October
2026 M4 Pro 48 GB run (0.15.4), the installer's recommendation
qualification on an M2 Pro 32 GB (0.12.7, 2026-08-07, marked
†) and an M3 Ultra 256 GB run
(0.12.15, 2026-08-18, marked ‡). A number is
never copied between machines.
The RAM → picks map
| RAM tier | Smart pick | Fast pick | Measured decode |
|---|---|---|---|
| 48 GB+ | qwen3.8-27b-4bit |
qwen3.6-35b-4bit |
24.1 / 104.5 tok/s § |
| 32 GB | qwen3.8-27b-4bit |
qwen3.5-4b-4bit |
24.1 / 85.8 tok/s § |
| 24 GB | bonsai-27b-2bit |
qwen3.5-4b-4bit |
17.5 ‡ / 85.8 tok/s § |
| 18 GB | qwen3.5-9b-4bit |
qwen3.5-4b-4bit |
63.0 / 85.8 tok/s § |
| 16 GB | qwen3.5-4b-4bit |
lfm2.5-1b-4bit |
85.8 / 314.0 tok/s § |
| 8 GB | lfm2.5-2.6b-4bit |
lfm2.5-1b-4bit |
93.5 † / 314.0 tok/s § |
What the marks mean.
§ measured via rapid-mlx serve on a
Mac mini M4 Pro 48 GB, rapid-mlx 0.15.4, 2026-10-02 — median of 3,
temperature 0, raw data
published (smart pick listed first, then the fast pick).
† M2 Pro 32 GB, 0.12.7, 2026-08-07.
‡ M3 Ultra 256 GB, 0.12.15, 2026-08-18 —
expect less on smaller chips. Where a pick has no mark its tier relies on
the other row's measurement; nothing here is estimated except the one
case flagged in the 48 GB section.
48 GB and up · Mac Studio / high-spec mini
Smart: rapid-mlx serve qwen3.8-27b-4bit
· Fast: rapid-mlx serve qwen3.6-35b-4bit
The 27B dense reasoner is the most capable model the policy places on this tier, and the 35B-A3B MoE is its speed counterpart — on the October M4 Pro run they measured 24.1 and 104.5 tok/s decode respectively, both with 4-stream aggregate close to their single-stream rate. The 35B MoE's 60 tok/s figure in the installer's own dataset is the engine's reviewed estimate, not a measurement (its only 32 GB-host run aborted with swap) — the M4 Pro number above is the measured one. Qwen3.8 27B also has an experimental tensorfold profile on this tier; it carries no fixed speed claim and stays opt-in.
Above this tier the policy stays with these two picks. The bigger MoEs are still servable — see the 192–256 GB flags.
32 GB · Mac mini / maxed MacBook
Smart: rapid-mlx serve qwen3.8-27b-4bit
· Fast: rapid-mlx serve qwen3.5-4b-4bit
The same 27B, one tier down: the October run measured it at 25.3 GB Metal peak through a 2K-prompt workload, which fits a 32 GB Mac with the OS and an editor running. MTP speculative decoding on by default, vision input, qwen3 thinking separation — the measured page has the full picture. The fast pick, qwen3.5-4b, measured 85.8 tok/s on the same October run.
24 GB · Mac mini / MacBook Pro
Smart: rapid-mlx serve bonsai-27b-2bit
· Fast: rapid-mlx serve qwen3.5-4b-4bit
Ternary Bonsai 27B, packed to 2-bit MLX: a 27B-class model at a 13 GB weight footprint (17.5 tok/s measured on the M3 Ultra qualification, ‡) is why this tier gets something far more capable than its memory would normally buy. It is slower than the other picks — the trade is quality per gigabyte, not speed. Its second-generation successor, Ternary Bonsai 2 27B, measured 24.4 GB Metal peak on the October run — that wants 32 GB, so the policy keeps the first generation here.
18 GB · MacBook Air / Pro
Smart: rapid-mlx serve qwen3.5-9b-4bit
· Fast: rapid-mlx serve qwen3.5-4b-4bit
Qwen 3.5 9B in 4-bit MLX — the smallest tier where a coding agent starts
to work reliably: real instruction following and tool calling. The
October run measured it at 63.0 tok/s (8.0 GB Metal peak); the
fast pick is the same family's 4B at 85.8 tok/s. The
hermes / qwen3 parser pair is auto-selected.
16 GB · MacBook Air / base MacBook Pro
Smart: rapid-mlx serve qwen3.5-4b-4bit
· Fast: rapid-mlx serve lfm2.5-1b-4bit
At a 6.0 GB weight footprint the 4B is the fast, general-purpose pick for a 16 GB Mac — 85.8 tok/s on the October run, the same Qwen 3.5 family and parser pair as the tier above, sized to fit alongside the OS and a browser. LFM2.5-1B is the throughput option: 314.0 tok/s decode and 618 tok/s aggregate on 4 streams in the same run, 1.8 GB peak — a basic-chat specialist.
8 GB · MacBook Air / base mini
Smart: rapid-mlx serve lfm2.5-2.6b-4bit
· Fast: rapid-mlx serve lfm2.5-1b-4bit
LFM2.5-2.6B, a 2.6B model whose 30 layers are 22 short-convolution blocks and only 8 GQA. Those 8 attention layers are why it belongs here: the KV cache costs roughly 16 KB per token, so a 32K conversation adds about half a gigabyte on top of 1.6 GB of weights. 128K context; 93.5 tok/s measured on the M2 Pro qualification (†).
What it is good at, and what it is not. It is post-trained for
tool use and instruction following and beats models several times its
size on those. Liquid publishes it as
not recommended for agentic coding and knowledge-heavy tasks,
and we agree — if you are driving a coding agent, this is the wrong
model and the 18 GB tier (qwen3.5-9b-4bit) is where
that starts working.
Its weights live in the 4bit/ folder of a repo that ships
eight quantizations side by side; rapid-mlx fetches only that folder,
so the download is 1.6 GB rather than the repo's ~19 GB.
Also fits: a real reasoner.
ling-3.0-tiny-4bit (4.2 GB weights) is inclusionAI's
7.9B mixture-of-experts with only 1.3B active per token — thinking
streamed as reasoning_content, native tool calling, 131K
context. It is the strongest reasoning option that fits this tier;
pick it over LFM2.5 when you want chain-of-thought rather than fast
instruction following. It has no measured numbers on either of our
runs, so it gets none here.
Above the table · the 192–256 GB flags
The policy's tiers stop at 48 GB, but the catalog serves bigger machines: GLM-5.3 Flash (320B/18B active), DeepSeek V4.1 Flash REAP, MiMo-V2.6 Flash (309B/15B active) and Qwen3.8 Flash-Next (180B/6B active) want a 192–256 GB Mac Studio. Their measured numbers live on their own pages — GLM-5.3 Flash, DeepSeek V4.1 Flash, MiMo-V2.6 Flash — each citing the machine, engine version and date it was measured on.
Check your machine
rapid-mlx doctor reports your chip, memory, macOS version,
free disk and Hugging Face cache size, then checks packages and
optional extras. Its first section on an M3 Pro 18 GB with
rapid-mlx 0.15.5:
$ rapid-mlx doctor ◆ System ✓ Apple Silicon (Apple M3 Pro, 18 GB) ✓ macOS 15.6.1 (Darwin 24.6.0) ✓ Free disk: 94 GB