DeepSeek
8 MLX aliases · V4-Flash / V4.1-Flash frontier line + R1 reasoning-distilled + Coder MoE.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve deepseek-r1-8b-4bitrapid-mlx serve deepseek-coder-v2-lite-16b-4bitrapid-mlx serve deepseek-v4-flash-4bit
DeepSeek's open-weights line on rapid-mlx covers the R1 reasoning-distilled models and the Coder MoE — both with reasoning_content streaming where appropriate. The frontier 1T-class V4-Flash family lives on its own hero deep dive (see DeepSeek V4-Flash) because the CSA + HCA sparse-attention path needed a vendored kernel and deserves a dedicated writeup.
- family
- DeepSeek
- aliases
- 8
- lines
- 4
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 6 of the 8 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Lines in this family
- DeepSeek V4-Flash (4 aliases)
- DeepSeek V4.1 Flash (REAP) (1 alias)
- DeepSeek R1 (2 aliases)
- DeepSeek Coder (1 alias)
DeepSeek V4-Flash · 4 aliases
The frontier line — the most capable open-weights model in the catalog, in 2/4/8-bit plus the 0731 MXFP4 checkpoint. Reasoning is parsed by a dedicated deepseek_v4 parser, warm starts reuse a persisted prefix cache, and the checkpoint-native DSpark speculative heads can draft for it. The CSA + HCA sparse-attention path needed a vendored kernel — the full writeup lives on the V4-Flash hero page.
parser: deepseek
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| deepseek-v4-flash-0731-mxfp4 | Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX | deepseek_v4_0731 | deepseek_v4 | moe | — | 51.8 | HF |
| deepseek-v4-flash-2bit | mlx-community/DeepSeek-V4-Flash-2bit-DQ | deepseek | deepseek_v4 | spec | 1M | 51.8 | CDN |
| deepseek-v4-flash-4bit | mlx-community/DeepSeek-V4-Flash-4bit | deepseek | deepseek_v4 | spec | 1M | 51.8 | CDN |
| deepseek-v4-flash-8bit | mlx-community/DeepSeek-V4-Flash-8bit | deepseek | deepseek_v4 | spec | 1M | 51.8 | CDN |
Notes & caveats
- Start from the hero page for RAM sizing — these are single-node-large checkpoints.
DeepSeek V4.1 Flash (REAP) · 1 alias
An experimental 2-bit REAP build of V4.1 Flash with a narrow DSpark K4 speculative sidecar: greedy generation only, requests serialize, no tools or images, and startup fails closed below its 224 GB memory floor. Pull it first (rapid-mlx pull deepseek-v41-flash-reap-2bit) — the measured numbers are on the model page.
parser: —
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| deepseek-v41-flash-reap-2bit | rapid-mlx/DeepSeek-V4.1-Flash-REAP-2bit-MLX | — | deepseek_v4 | moe | — | — | HF |
DeepSeek R1 · 2 aliases
Reasoning-distilled — chain-of-thought weights distilled into compact models that punch above their parameter count on reasoning benchmarks. The 8B variant uses the deepseek_v3 tool parser; the 32B variant uses the more recent deepseek parser. Both stream reasoning_content via the deepseek_r1 reasoning parser.
parser: deepseek
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| deepseek-r1-32b-4bit | mlx-community/DeepSeek-R1-Distill-Qwen-32B-4bit | — | deepseek_r1_distill | spec | 128K | 11 | CDN |
| deepseek-r1-8b-4bit | mlx-community/DeepSeek-R1-0528-Qwen3-8B-4bit | deepseek_v3 | deepseek_r1 | spec | 128K | 10.3 | CDN |
Notes & caveats
- The 32B-4bit quant is the better fit for a 32 GB Mac; the 8B-4bit drops to 16 GB.
DeepSeek Coder · 1 alias
Code-specialised MoE line, kept as a single 16B-Lite alias for users who want a DeepSeek-flavoured coder alongside the Qwen3 Coder family.
parser: —
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| deepseek-coder-v2-lite-16b-4bit | mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit-mlx | deepseek_v3 | — | moe · spec | 160K | 2.7 | CDN |
Notes & caveats
- Lite-MoE means good active-parameter speed-to-quality on a 32 GB+ Mac.
Context is read from the config.json of the exact build each alias pulls. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.
Notes & caveats
- For the frontier 1T-class line see the dedicated DeepSeek V4-Flash hero page.