Skip to content

Latest commit

 

History

History
38 lines (21 loc) · 2.62 KB

File metadata and controls

38 lines (21 loc) · 2.62 KB

Measuring local inference fairly

Outcome: A controlled comparison of trailbrake, Rapid-MLX, and oMLX separated a small serving-overhead improvement from differences in how servers report streaming speed.

The problem

Tokens per second looks like a simple metric until servers measure different intervals. One server may emit every token; another may buffer much of a response before sending its first visible chunk. Comparing only the interval after that chunk can reward buffering rather than faster generation.

I needed to understand where my narrow Qwen3 runtime, trailbrake, actually helped.

What I owned

I built the runtime and comparison tooling, then recorded model revisions, weight hashes, sampler settings, output length, warmups, repeated runs, memory use, and output checks. The protocol used a persistent HTTP client that drained each response and checked exact usage and output length before reusing the connection.

The comparison

The July 18, 2026 record used an M5 Max, Qwen3-8B-4bit, a fixed code prompt, 112 generated tokens, two warmups, and seven measured runs. Cache and speculation settings were controlled.

Recorded median trailbrake 0.2 Rapid-MLX 0.10.12 oMLX 0.5.1
End-to-end response time 1,164.573 ms 1,197.366 ms 1,224.467 ms
End-to-end output rate 96.173 tok/s 93.539 tok/s 91.468 tok/s

All three returned the same raw continuation. The full trailbrake token sequence also matched the independent MLX-LM reference.

The decision

I reported 2.82% higher end-to-end throughput than Rapid-MLX and 5.14% higher than oMLX for this exact workload. The decode-only difference from Rapid-MLX was just 1.91%, which the record treats as core parity rather than a kernel-performance win.

The oMLX response arrived in three buffered bursts. Its post-first-chunk decode rate was therefore excluded from the direct decode comparison; client-observed total time remained comparable.

Limits and reproduction

This is a dated, single-request Qwen3 workload. It does not establish an overall advantage across model families, batching, long contexts, or each competitor's broader feature set. No new hardware benchmark was run to prepare this case study.

The full protocol contains model revisions, hashes, launch commands, and the comparison command. The measurement client provides the reproduction entry point.

All case studies