Reference · rapid-mlx 0.15.6 · ← Back to README

Troubleshooting

Run the doctor first — it checks your Mac, Python, packages, optional extras, cache and network, and can apply verified repairs:

$ rapid-mlx doctor                 # full check
$ rapid-mlx doctor --fix --dry-run # show the repair plan
$ rapid-mlx doctor --fix           # apply verified repairs only
$ rapid-mlx doctor --verbose       # the probe detail behind each line

--deep adds dependency, DNS and route probes; --json prints a machine-readable report to attach to an issue. The symptoms below are the common ones that are not a broken install.

Performance

Much slower than the published numbers

Slow first response, fast follow-ups

The first request prefills the whole prompt; later turns reuse it from the prefix cache. For long agent prompts, a stable system prompt keeps that reuse working — and --pin-system-prompt keeps it cached under memory pressure.

Memory and startup

Out of memory, or very slow (under 5 tok/s)

The model is too big for your RAM. Pick a smaller quant or model — the smart and fast pick for each tier:

RAMSmartFast
48 GB+qwen3.8-27b-4bitqwen3.6-35b-4bit
32 GBqwen3.8-27b-4bitqwen3.5-4b-4bit
24 GBbonsai-27b-2bitqwen3.5-4b-4bit
18 GBqwen3.5-9b-4bitqwen3.5-4b-4bit
16 GBqwen3.5-4b-4bitlfm2.5-1b-4bit
8 GBlfm2.5-2.6b-4bitlfm2.5-1b-4bit

Long contexts grow the KV cache: --context-length caps the per-request window, and --kv-cache-dtype int8 halves the cache at some decode speed.

HTTP 503 "Server is at its Metal memory limit"

The request would push Metal memory past the server's limit, so it is refused before it starts instead of risking a crash. The message gives the memory the request needs and the limit. While other requests are running, retry after they finish, or send a shorter prompt or a smaller max_tokens. Finished requests give up their working memory; reusable prompt cache stays until memory is needed and is evicted before a request is refused, and an idle server re-measures free memory before it refuses anything. If the message says no request is running, the memory is held by the server itself — restart it, or serve a smaller model. A higher --gpu-memory-utilization raises the limit when the message suggests it.

503 "Server is busy (max concurrent requests reached)" is the separate concurrency cap (--max-concurrent-requests, default 256); retry after the Retry-After delay.

Pre-flight disk-space check stops serve

rapid-mlx serve stops before downloading when the model is larger than the free space on the disk that holds the Hugging Face cache. If your cache really lives on another drive (via HF_HOME) that has room, add --force-disk-check. To free space, list cached models with rapid-mlx ls and remove one with rapid-mlx rm <alias>.

A Hugging Face model is refused before download

For models outside the catalog, serve and pull check the format and architecture first. GGUF and .bin-only repos cannot run on MLX; an unsupported architecture is refused with the reason. --request files a support request (repo id, architecture and format only). See Bring your own model for converting a safetensors model with rapid-mlx import.

Output

Empty responses

Usually a parser mismatch: a --reasoning-parser on a model that does not think moves everything into reasoning_content. Remove the flag, or pass --no-reasoning-parser to skip the alias's automatic choice. rapid-mlx info <alias> shows which parser the alias selects.

Tool calls arrive as plain text

Catalog aliases select the right tool-call parser on their own. For a model served by its org/name, pass --enable-auto-tool-choice --tool-call-parser <name>; the parser names are listed in rapid-mlx serve --help.

"parameters not found in model" at startup

Normal for vision-language models served on the text path — the vision weights are skipped. To use images, install the vision extra (pip install "rapid-mlx[vision]") and drop --no-mllm / --text-only if you set it.

Raw <think> text in the answer

Servers and ports

Is a server already running?

rapid-mlx ps lists every running server — PID, port, model and uptime.

Port already in use

Without --port, serve takes the first free port from 8000 to 8009. An explicit --port never falls back: pick another, or find the process with lsof -i :<port>. With --host 0.0.0.0, the start-up check also probes 127.0.0.1, because macOS lets a wildcard listener share a port with a loopback one.

Platform and shell

"Rapid-MLX requires Apple Silicon"

MLX runs only on Apple Silicon (M1 and later). Intel Macs, Linux and Windows are not supported.

"Rapid-MLX requires macOS 14 (Sonoma) or later"

MLX publishes its wheels for macOS 14 and later only, so Rapid-MLX needs macOS 14 (Sonoma) or newer. On macOS 13 an older installer can pass its own version check and then fail while installing mlx. Upgrade macOS to install.

Tab completion doesn't work

The CLI ships shell completion; add the hook once:

# add to ~/.zshrc or ~/.bashrc
$ eval "$(register-python-argcomplete rapid-mlx)"

Still stuck?

Open an issue at github.com/raullenchai/Rapid-MLX/issues with the output of rapid-mlx doctor --json, or ask in the Discord (rapid-mlx feedback opens it).

Next steps