Outcome: A controlled comparison of trailbrake, Rapid-MLX, and oMLX separated a small serving-overhead improvement from differences in how servers report streaming speed.
Tokens per second looks like a simple metric until servers measure different intervals. One server may emit every token; another may buffer much of a response before sending its first visible chunk. Comparing only the interval after that chunk can reward buffering rather than faster generation.
I needed to understand where my narrow Qwen3 runtime, trailbrake, actually helped.
I built the runtime and comparison tooling, then recorded model revisions, weight hashes, sampler settings, output length, warmups, repeated runs, memory use, and output checks. The protocol used a persistent HTTP client that drained each response and checked exact usage and output length before reusing the connection.
The July 18, 2026 record used an M5 Max, Qwen3-8B-4bit, a fixed code prompt, 112 generated tokens, two warmups, and seven measured runs. Cache and speculation settings were controlled.
| Recorded median | trailbrake 0.2 | Rapid-MLX 0.10.12 | oMLX 0.5.1 |
|---|---|---|---|
| End-to-end response time | 1,164.573 ms | 1,197.366 ms | 1,224.467 ms |
| End-to-end output rate | 96.173 tok/s | 93.539 tok/s | 91.468 tok/s |
All three returned the same raw continuation. The full trailbrake token sequence also matched the independent MLX-LM reference.
I reported 2.82% higher end-to-end throughput than Rapid-MLX and 5.14% higher than oMLX for this exact workload. The decode-only difference from Rapid-MLX was just 1.91%, which the record treats as core parity rather than a kernel-performance win.
The oMLX response arrived in three buffered bursts. Its post-first-chunk decode rate was therefore excluded from the direct decode comparison; client-observed total time remained comparable.
This is a dated, single-request Qwen3 workload. It does not establish an overall advantage across model families, batching, long contexts, or each competitor's broader feature set. No new hardware benchmark was run to prepare this case study.
The full protocol contains model revisions, hashes, launch commands, and the comparison command. The measurement client provides the reproduction entry point.