Troubleshooting
Run the doctor first — it checks your Mac, Python, packages, optional extras, cache and network, and can apply verified repairs:
$ rapid-mlx doctor # full check $ rapid-mlx doctor --fix --dry-run # show the repair plan $ rapid-mlx doctor --fix # apply verified repairs only $ rapid-mlx doctor --verbose # the probe detail behind each line
--deep adds dependency, DNS and route probes;
--json prints a machine-readable report to attach to an
issue. The symptoms below are the common ones that are not a broken
install.
Performance
Much slower than the published numbers
-
The model is thinking first. A request that asks for reasoning
(
enable_thinking,reasoning_effort,reasoning_max_tokens, or an agent that sends them) spends tokens onreasoning_contentbefore the answer. Plain chat completions answer directly unless asked. Onserve,--no-thinkingturns reasoning off for every request. -
The model doesn't fit. Under memory pressure macOS swaps and
decode falls to a few tokens per second. Check the
hardware tiers or
rapid-mlx recipefor a model that fits. - Different machine. Published numbers name the Mac they were measured on; compare against the leaderboard entries for your chip and RAM.
Slow first response, fast follow-ups
The first request prefills the whole prompt; later turns reuse it from
the prefix cache. For long agent prompts, a stable system prompt keeps
that reuse working — and --pin-system-prompt keeps it
cached under memory pressure.
Memory and startup
Out of memory, or very slow (under 5 tok/s)
The model is too big for your RAM. Pick a smaller quant or model — the smart and fast pick for each tier:
| RAM | Smart | Fast |
|---|---|---|
| 48 GB+ | qwen3.8-27b-4bit | qwen3.6-35b-4bit |
| 32 GB | qwen3.8-27b-4bit | qwen3.5-4b-4bit |
| 24 GB | bonsai-27b-2bit | qwen3.5-4b-4bit |
| 18 GB | qwen3.5-9b-4bit | qwen3.5-4b-4bit |
| 16 GB | qwen3.5-4b-4bit | lfm2.5-1b-4bit |
| 8 GB | lfm2.5-2.6b-4bit | lfm2.5-1b-4bit |
Long contexts grow the KV cache: --context-length caps
the per-request window, and --kv-cache-dtype int8 halves
the cache at some decode speed.
HTTP 503 "Server is at its Metal memory limit"
The request would push Metal memory past the server's limit, so it is
refused before it starts instead of risking a crash. The message gives
the memory the request needs and the limit. While other requests are
running, retry after they finish, or send a shorter prompt or a smaller
max_tokens. Finished requests give up their working memory;
reusable prompt cache stays until memory is needed and is evicted
before a request is refused, and an idle server re-measures free
memory before it refuses anything. If
the message says no request is running, the memory is held by the
server itself — restart it, or serve a smaller model. A higher
--gpu-memory-utilization raises the limit when the message
suggests it.
503 "Server is busy (max concurrent requests reached)" is
the separate concurrency cap (--max-concurrent-requests,
default 256); retry after the Retry-After delay.
Pre-flight disk-space check stops serve
rapid-mlx serve stops before downloading when the model is
larger than the free space on the disk that holds the Hugging Face
cache. If your cache really lives on another drive (via
HF_HOME) that has room, add --force-disk-check.
To free space, list cached models with rapid-mlx ls and
remove one with rapid-mlx rm <alias>.
A Hugging Face model is refused before download
For models outside the catalog, serve and
pull check the format and architecture first. GGUF and
.bin-only repos cannot run on MLX; an unsupported
architecture is refused with the reason. --request files a
support request (repo id, architecture and format only). See
Bring your own model for converting a
safetensors model with rapid-mlx import.
Output
Empty responses
Usually a parser mismatch: a --reasoning-parser on a model
that does not think moves everything into
reasoning_content. Remove the flag, or pass
--no-reasoning-parser to skip the alias's automatic choice.
rapid-mlx info <alias> shows which parser the alias
selects.
Tool calls arrive as plain text
Catalog aliases select the right tool-call parser on their own. For a
model served by its org/name, pass
--enable-auto-tool-choice --tool-call-parser <name>;
the parser names are listed in rapid-mlx serve --help.
"parameters not found in model" at startup
Normal for vision-language models served on the text path — the vision
weights are skipped. To use images, install the vision extra
(pip install "rapid-mlx[vision]") and drop
--no-mllm / --text-only if you set it.
Raw <think> text in the answer
--no-thinkingonserveturns off reasoning in the prompt template and the parser.--no-reasoning-parseronly skips the automatic parser choice: the model still thinks, and the<think>text stays incontentinstead of moving toreasoning_content.
Servers and ports
Is a server already running?
rapid-mlx ps lists every running server — PID, port, model
and uptime.
Port already in use
Without --port, serve takes the first free
port from 8000 to 8009. An explicit --port never falls
back: pick another, or find the process with
lsof -i :<port>. With --host 0.0.0.0, the
start-up check also probes 127.0.0.1, because macOS lets a
wildcard listener share a port with a loopback one.
Platform and shell
"Rapid-MLX requires Apple Silicon"
MLX runs only on Apple Silicon (M1 and later). Intel Macs, Linux and Windows are not supported.
"Rapid-MLX requires macOS 14 (Sonoma) or later"
MLX publishes its wheels for macOS 14 and later only, so Rapid-MLX needs
macOS 14 (Sonoma) or newer. On macOS 13 an older installer can pass its
own version check and then fail while installing mlx.
Upgrade macOS to install.
Tab completion doesn't work
The CLI ships shell completion; add the hook once:
# add to ~/.zshrc or ~/.bashrc $ eval "$(register-python-argcomplete rapid-mlx)"
Still stuck?
Open an issue at
github.com/raullenchai/Rapid-MLX/issues
with the output of rapid-mlx doctor --json, or ask in the
Discord
(rapid-mlx feedback opens it).