Phi
4 MLX aliases · Microsoft's Phi-3.5 / Phi-4 line — small, dense, well-aligned.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve phi-4-mini-4bit
Microsoft's Phi line — small, dense, well-aligned. We ship the Phi-3.5-mini baseline plus the Phi-4 generation including the mini-reasoning checkpoint that streams reasoning_content via the deepseek_r1 reasoning parser.
- family
- Phi
- aliases
- 4
- lines
- 1
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 4 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Phi-3.5 / Phi-4 · 4 aliases
Phi-3.5-mini baseline + Phi-4 generation (mini / 14B / mini-reasoning). The mini-reasoning checkpoint streams reasoning_content via the deepseek_r1 reasoning parser.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| phi-3.5-mini-4bit | mlx-community/Phi-3.5-mini-instruct-4bit | — | — | spec | 128K | — | CDN |
| phi-4-14b-4bit | mlx-community/phi-4-4bit | hermes | — | spec | 16K | 4.6 | CDN |
| phi-4-mini-4bit | mlx-community/phi-4-mini-instruct-4bit | hermes | — | spec | 128K | 5.7 | CDN |
| phi-4-mini-reasoning-4bit | lmstudio-community/Phi-4-mini-reasoning-MLX-4bit | hermes | deepseek_r1 | spec | 128K | — | CDN |
Notes & caveats
- Phi-4 14B 4-bit is the best quality-per-RAM model in the registry for under-20 GB models.
- Use
phi-4-mini-reasoning-4bitwhen you want CoT streams in <6 GB of RAM.
Context is read from the config.json of the exact build each alias pulls. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.