Models · family

GLM

5 MLX aliases · Zhipu's GLM line · GLM-5.3-Flash + GLM-5.2 (REAP) + GLM-4.7 + GLM-4.5-Air.

Pick one

One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.

Zhipu's GLM line — Chinese-frontier open-weights with strong English performance. rapid-mlx serves GLM-5.3-Flash (320B total, 18B active, experimental, 192 GB+ Macs), GLM-5.2 (REAP-pruned), GLM-4.7, and GLM-4.5-Air with our custom glm47 tool-call parser; GLM-5.3 uses the glm5 reasoning parser, the others glm4.

family
GLM
aliases
5
lines
1
install
rapid-mlx serve <alias>
OpenAI base URL
http://localhost:8000/v1

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 2 of the 5 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

GLM-4 · 5 aliases

GLM-4.5-Air + GLM-4.7-Flash. Both use the glm47 tool-call parser and the glm4 reasoning parser.

parser: glm47

aliashf repotool parserreasoningflagscontextAA indexget it
glm-5.2-reap50pipenetwork/GLM-5.2-REAP50-MLX-4bitglm47glm4moe——HF
glm4.5-air-4bitmlx-community/GLM-4.5-Air-4bitglm47glm4spec128K16.7CDN
glm4.7-9b-4bitmlx-community/GLM-4.7-Flash-4bitglm47glm4spec198K23.3CDN
glm5.3-flash-4bitVontra/GLM-5.3-Flash-MLX-4bit-MTPglm47glm5hybrid · moe——HF
glm5.3-flash-tensorfoldVontra/GLM-5.3-Flash-MLX-4bit-MTP—glm5hybrid · moe——HF

Notes & caveats

Context is read from the config.json of the exact build each alias pulls. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.

Where next