A reproducible deployment guide for the embedded Qwen4-Exp Lightning MTP path in oMLX on a 128 GB M4 Max Mac Studio.
The repository now covers the full progression from the original non-speculative profile to the qualified depth-6 deployment:
- Fix the stale external-drafter toggle and enable the checkpoint's embedded MTP tensors.
- Increase adaptive native MTP from depth 3 to depth 6.
- Use aggressive Burst Decode with balanced prefill priority.
- Re-run the complete 152-transcript Spark Bench suite instead of assuming probe quality transfers.
- Mac Studio: Apple M4 Max, 128 GB unified memory
- macOS 26.5.2
- oMLX 0.6.4 with custom kernels
- Model:
Jundot/Qwen3.8-Flash-Next-oQ4e-mtp - Quant: oQ mixed precision, 4-bit overall, group size 64, MLX safetensors
- Model footprint reported by oMLX: approximately 103.94 GB
- Native embedded MTP: enabled, adaptive maximum depth 6
- External VLM-MTP drafter: disabled
- Burst Decode: aggressive
- Prefill priority: balanced
- Maximum concurrent requests: 1
- Prefill memory guard: custom 120 GB ceiling (still bounded by the lower Metal working-set limit)
- Benchmark requests: thinking OFF, uncapped, temperature 0.3
Both runs used Spark Bench v6.7.1, 76 scenarios × 2 repeats, thinking OFF, uncapped requests, the same Mac and checkpoint, and the same grader revision.
| Metric | Original MTP-off deployment | Current MTP6 deployment | Change |
|---|---|---|---|
| TrueScore | 91.8 | 91.9 | +0.1 |
| Capability | 92.0 | 91.8 | -0.2 |
| Operational | 82.7 | 82.3 | -0.4 |
| Pass@1 | 97.4% | 97.4% | unchanged |
| Pass@K | 94.7% | 94.7% | unchanged |
| Median turn latency | 6.56s | 6.77s | +0.21s |
| Transport errors | 0 | 0 | unchanged |
| Reasoning leakage | 0/152 | 0/152 | unchanged |
| Qualification | passed | passed | clean |
The correct conclusion is quality held, not that MTP improved model quality. TrueScore moved by only +0.1 while capability and operational sub-scores moved slightly in the other direction.
The current run independently passed:
- 152/152 parsed transcripts and a complete 76×2 matrix
- Golden gate 12/12
- Tool-call preflight
- 0 transport errors
- 0 reasoning characters
- Clean quarantine checks
The decode sweep held prefill priority at context while changing MTP depth and Burst Decode. It used frozen prompts, temperature 0, seed 12345, thinking OFF, omitted request token caps, runtime telemetry, output hashes, and executable validation. Prefill priority was optimized separately afterward.
| Probe | Previous public MTP3 | MTP6 + aggressive Burst | Change | Validation |
|---|---|---|---|---|
| 1,200-line sequence, 5,999 tokens | 70.44 tok/s | 83.06 tok/s | +17.9% | byte-identical |
| Python LRU module | 65.12 tok/s | 71.07 tok/s | +9.1% | both compiled; 27/27 and 30/30 generated tests passed |
| Structured prime JSON | 75.83 tok/s | 82.67 tok/s | +9.0% | exact schema, values and checksum |
The structured figure is the mean of two runs per arm. Tool-call decode rates are intentionally excluded from the headline because a 27-token completion overstates scheduler/accounting changes. Balanced prefill was then selected independently because it admitted the tested 15,416-token prompt; the final deployment combines aggressive Burst Decode with balanced prefill.
Fresh-restart peaks are not a forever steady-state guarantee. Immediately after the complete 152-transcript MTP6 benchmark, without restarting oMLX, follow-up probes measured:
- Sequence: 69.14 tok/s, byte-identical
- Code: 53.74 tok/s, compiled and passed 28/28 generated tests
- Structured JSON: 62.72 tok/s mean across five exact-valid runs
- Tool calls: 5/5 valid
This post-suite slowdown is preserved as a real deployment caveat. It does not invalidate the 91.9 full-suite score, but it means the roughly 83 tok/s decode result should be described as a controlled fresh-state measurement, not guaranteed sustained performance after hours of inference.
The original profile had embedded mtp.* tensors but native MTP was off. A stale external VLM-MTP toggle was on without a drafter and was ignored.
These controls select different paths:
mtp_enabled: native embedded Lightning MTP used by this checkpointvlm_mtp_enabled: external VLM assistant-drafter path, which also requiresvlm_mtp_draft_model
The current model profile is:
{
"mtp_enabled": true,
"mtp_num_draft_tokens": 6,
"vlm_mtp_enabled": false,
"enable_thinking": false,
"thinking_budget_enabled": false
}The tested server scheduler settings are:
{
"server": {
"burst_decode_mode": "aggressive"
},
"scheduler": {
"max_concurrent_requests": 1,
"chunked_prefill": false,
"prefill_priority": "balanced"
},
"memory": {
"prefill_memory_guard": true,
"memory_guard_tier": "custom",
"memory_guard_custom_ceiling_gb": 120.0
}
}A checkpoint name, embedded draft weights, or a benchmark label is not activation proof. Verify the effective settings and a positive runtime activation line.
brew tap jundot/omlx
brew install jundot/omlx/omlx --with-custom-kernel
hf download Jundot/Qwen3.8-Flash-Next-oQ4e-mtp \
--local-dir "$HOME/mlm/Jundot--Qwen3.8-Flash-Next-oQ4e-mtp"Start oMLX once so it discovers the model directory. The directory name becomes the default OpenAI-compatible model ID unless you configure an alias.
The scripts back up each settings file, change only the documented fields, write atomically, and read the result back.
python3 scripts/configure_model.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
--mode mtp \
--depth 6
python3 scripts/configure_server.py \
--burst-decode aggressive \
--prefill-priority balanced \
--max-concurrent-requests 1 \
--memory-guard-ceiling-gb 120
brew services restart jundot/omlx/omlxVerify the requested depth and the positive runtime activation signature:
python3 scripts/verify_activation.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
--expected-depth 6Expected log signature:
Qwen4-Exp Lightning MTP enabled for <MODEL_PATH> (checkpoint layout: mtp.)
Run sequentially; do not benchmark two profiles against the same accelerator concurrently.
# Previous public profile
python3 scripts/configure_model.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --mode mtp --depth 3
brew services restart jundot/omlx/omlx
python3 scripts/ab_probe.py \
--phase mtp-depth3 --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp
# Current profile
python3 scripts/configure_model.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --mode mtp --depth 6
python3 scripts/configure_server.py \
--burst-decode aggressive --prefill-priority balanced
brew services restart jundot/omlx/omlx
python3 scripts/verify_activation.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --expected-depth 6
python3 scripts/ab_probe.py \
--phase mtp-depth6 --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtpab_probe.py omits max_tokens and max_completion_tokens. A diagnostic client deadline is optional and separate from the uncapped benchmark contract. Raw SSE events, usage, TTFT, finish reason, content hashes, and reasoning leakage are saved under evidence/<phase>/probes/.
Balanced prefill was selected because it preserved the tested long-context path:
- 15,416 prompt tokens accepted
- 215–220 prompt tok/s in repeated tests
- 70–72s TTFT
- Correct one-token acknowledgement
The speed-priority profile was faster on short prompts but rejected the same 15,416-token request under the memory guard. Decode speed, prefill speed, TTFT, and context admission are separate dimensions.
Even with balanced prefill, a larger generated long-context probe can still hit the memory guard. Do not blindly raise iogpu.wired_limit_mb; treat Metal headroom as a reversible experiment and preserve OS/swap margin.
python3 scripts/configure_model.py \
--model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
--mode baseline
brew services restart jundot/omlx/omlxYou can also restore the timestamped backup path printed by either configuration script.
scripts/ab_probe.py: streamed decode/tool/structured probes with raw SSE capture.scripts/prefill_probe.py: tokenizer-reported prompt throughput and TTFT ladder.scripts/validate_semantics.py: exact sequence, structured JSON and tool-call checks.scripts/validate_code_probe.py: compile and execute generated Python test suites.scripts/runtime_behavior_probe.py: serial versus queued-concurrent behavior.scripts/response_format_probe.py: OpenAIjson_schemaandjson_objectcompatibility.scripts/verify_spark_bench_run.py: complete 76×2 matrix and qualification audit.
Diagnostic client deadlines are explicit and do not change the uncapped Spark Bench contract.
results/ab-results.jsonpreserves the original MTP-off → MTP3 activation A/B.results/optimized-deployment-results.jsoncontains the qualified MTP6 score, old/new comparison, controlled depth sweep, prefill findings, and post-suite stability measurements.results/activation-proof.txtcontains the positive activation signature with the machine path removed.docs/METHODOLOGY.mddefines the contracts and caveats.
- Jundot/oMLX and its oQ quantization/runtime work
- The Qwen model authors
- Spark Bench
This repository contains configuration, scripts, and measurement metadata only. It does not redistribute model weights or third-party code.
The original material in this repository is released under the MIT License. oMLX and the model retain their own licenses.