Skip to content

About

Verified Qwen3.8 Flash Next oMLX Lightning MTP deployment guide for Apple Silicon

Topics

Resources

Stars

14 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Qwen3.8 Flash Next on Apple Silicon with oMLX Lightning MTP

A reproducible deployment guide for the embedded Qwen4-Exp Lightning MTP path in oMLX on a 128 GB M4 Max Mac Studio.

The repository now covers the full progression from the original non-speculative profile to the qualified depth-6 deployment:

  1. Fix the stale external-drafter toggle and enable the checkpoint's embedded MTP tensors.
  2. Increase adaptive native MTP from depth 3 to depth 6.
  3. Use aggressive Burst Decode with balanced prefill priority.
  4. Re-run the complete 152-transcript Spark Bench suite instead of assuming probe quality transfers.

Current qualified deployment

  • Mac Studio: Apple M4 Max, 128 GB unified memory
  • macOS 26.5.2
  • oMLX 0.6.4 with custom kernels
  • Model: Jundot/Qwen3.8-Flash-Next-oQ4e-mtp
  • Quant: oQ mixed precision, 4-bit overall, group size 64, MLX safetensors
  • Model footprint reported by oMLX: approximately 103.94 GB
  • Native embedded MTP: enabled, adaptive maximum depth 6
  • External VLM-MTP drafter: disabled
  • Burst Decode: aggressive
  • Prefill priority: balanced
  • Maximum concurrent requests: 1
  • Prefill memory guard: custom 120 GB ceiling (still bounded by the lower Metal working-set limit)
  • Benchmark requests: thinking OFF, uncapped, temperature 0.3

Old versus new

Full-suite quality

Both runs used Spark Bench v6.7.1, 76 scenarios × 2 repeats, thinking OFF, uncapped requests, the same Mac and checkpoint, and the same grader revision.

Metric Original MTP-off deployment Current MTP6 deployment Change
TrueScore 91.8 91.9 +0.1
Capability 92.0 91.8 -0.2
Operational 82.7 82.3 -0.4
Pass@1 97.4% 97.4% unchanged
Pass@K 94.7% 94.7% unchanged
Median turn latency 6.56s 6.77s +0.21s
Transport errors 0 0 unchanged
Reasoning leakage 0/152 0/152 unchanged
Qualification passed passed clean

The correct conclusion is quality held, not that MTP improved model quality. TrueScore moved by only +0.1 while capability and operational sub-scores moved slightly in the other direction.

The current run independently passed:

  • 152/152 parsed transcripts and a complete 76×2 matrix
  • Golden gate 12/12
  • Tool-call preflight
  • 0 transport errors
  • 0 reasoning characters
  • Clean quarantine checks

Controlled decode probes

The decode sweep held prefill priority at context while changing MTP depth and Burst Decode. It used frozen prompts, temperature 0, seed 12345, thinking OFF, omitted request token caps, runtime telemetry, output hashes, and executable validation. Prefill priority was optimized separately afterward.

Probe Previous public MTP3 MTP6 + aggressive Burst Change Validation
1,200-line sequence, 5,999 tokens 70.44 tok/s 83.06 tok/s +17.9% byte-identical
Python LRU module 65.12 tok/s 71.07 tok/s +9.1% both compiled; 27/27 and 30/30 generated tests passed
Structured prime JSON 75.83 tok/s 82.67 tok/s +9.0% exact schema, values and checksum

The structured figure is the mean of two runs per arm. Tool-call decode rates are intentionally excluded from the headline because a 27-token completion overstates scheduler/accounting changes. Balanced prefill was then selected independently because it admitted the tested 15,416-token prompt; the final deployment combines aggressive Burst Decode with balanced prefill.

Long-session stability

Fresh-restart peaks are not a forever steady-state guarantee. Immediately after the complete 152-transcript MTP6 benchmark, without restarting oMLX, follow-up probes measured:

  • Sequence: 69.14 tok/s, byte-identical
  • Code: 53.74 tok/s, compiled and passed 28/28 generated tests
  • Structured JSON: 62.72 tok/s mean across five exact-valid runs
  • Tool calls: 5/5 valid

This post-suite slowdown is preserved as a real deployment caveat. It does not invalidate the 91.9 full-suite score, but it means the roughly 83 tok/s decode result should be described as a controlled fresh-state measurement, not guaranteed sustained performance after hours of inference.

What changed

The original profile had embedded mtp.* tensors but native MTP was off. A stale external VLM-MTP toggle was on without a drafter and was ignored.

These controls select different paths:

  • mtp_enabled: native embedded Lightning MTP used by this checkpoint
  • vlm_mtp_enabled: external VLM assistant-drafter path, which also requires vlm_mtp_draft_model

The current model profile is:

{
  "mtp_enabled": true,
  "mtp_num_draft_tokens": 6,
  "vlm_mtp_enabled": false,
  "enable_thinking": false,
  "thinking_budget_enabled": false
}

The tested server scheduler settings are:

{
  "server": {
    "burst_decode_mode": "aggressive"
  },
  "scheduler": {
    "max_concurrent_requests": 1,
    "chunked_prefill": false,
    "prefill_priority": "balanced"
  },
  "memory": {
    "prefill_memory_guard": true,
    "memory_guard_tier": "custom",
    "memory_guard_custom_ceiling_gb": 120.0
  }
}

A checkpoint name, embedded draft weights, or a benchmark label is not activation proof. Verify the effective settings and a positive runtime activation line.

Install

brew tap jundot/omlx
brew install jundot/omlx/omlx --with-custom-kernel

hf download Jundot/Qwen3.8-Flash-Next-oQ4e-mtp \
  --local-dir "$HOME/mlm/Jundot--Qwen3.8-Flash-Next-oQ4e-mtp"

Start oMLX once so it discovers the model directory. The directory name becomes the default OpenAI-compatible model ID unless you configure an alias.

Apply the optimized profile

The scripts back up each settings file, change only the documented fields, write atomically, and read the result back.

python3 scripts/configure_model.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
  --mode mtp \
  --depth 6

python3 scripts/configure_server.py \
  --burst-decode aggressive \
  --prefill-priority balanced \
  --max-concurrent-requests 1 \
  --memory-guard-ceiling-gb 120

brew services restart jundot/omlx/omlx

Verify the requested depth and the positive runtime activation signature:

python3 scripts/verify_activation.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
  --expected-depth 6

Expected log signature:

Qwen4-Exp Lightning MTP enabled for <MODEL_PATH> (checkpoint layout: mtp.)

Reproduce the depth comparison

Run sequentially; do not benchmark two profiles against the same accelerator concurrently.

# Previous public profile
python3 scripts/configure_model.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --mode mtp --depth 3
brew services restart jundot/omlx/omlx
python3 scripts/ab_probe.py \
  --phase mtp-depth3 --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp

# Current profile
python3 scripts/configure_model.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --mode mtp --depth 6
python3 scripts/configure_server.py \
  --burst-decode aggressive --prefill-priority balanced
brew services restart jundot/omlx/omlx
python3 scripts/verify_activation.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp --expected-depth 6
python3 scripts/ab_probe.py \
  --phase mtp-depth6 --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp

ab_probe.py omits max_tokens and max_completion_tokens. A diagnostic client deadline is optional and separate from the uncapped benchmark contract. Raw SSE events, usage, TTFT, finish reason, content hashes, and reasoning leakage are saved under evidence/<phase>/probes/.

Prefill and context behavior

Balanced prefill was selected because it preserved the tested long-context path:

  • 15,416 prompt tokens accepted
  • 215–220 prompt tok/s in repeated tests
  • 70–72s TTFT
  • Correct one-token acknowledgement

The speed-priority profile was faster on short prompts but rejected the same 15,416-token request under the memory guard. Decode speed, prefill speed, TTFT, and context admission are separate dimensions.

Even with balanced prefill, a larger generated long-context probe can still hit the memory guard. Do not blindly raise iogpu.wired_limit_mb; treat Metal headroom as a reversible experiment and preserve OS/swap margin.

Roll back native MTP

python3 scripts/configure_model.py \
  --model Jundot--Qwen3.8-Flash-Next-oQ4e-mtp \
  --mode baseline

brew services restart jundot/omlx/omlx

You can also restore the timestamped backup path printed by either configuration script.

Diagnostic scripts

  • scripts/ab_probe.py: streamed decode/tool/structured probes with raw SSE capture.
  • scripts/prefill_probe.py: tokenizer-reported prompt throughput and TTFT ladder.
  • scripts/validate_semantics.py: exact sequence, structured JSON and tool-call checks.
  • scripts/validate_code_probe.py: compile and execute generated Python test suites.
  • scripts/runtime_behavior_probe.py: serial versus queued-concurrent behavior.
  • scripts/response_format_probe.py: OpenAI json_schema and json_object compatibility.
  • scripts/verify_spark_bench_run.py: complete 76×2 matrix and qualification audit.

Diagnostic client deadlines are explicit and do not change the uncapped Spark Bench contract.

Evidence

  • results/ab-results.json preserves the original MTP-off → MTP3 activation A/B.
  • results/optimized-deployment-results.json contains the qualified MTP6 score, old/new comparison, controlled depth sweep, prefill findings, and post-suite stability measurements.
  • results/activation-proof.txt contains the positive activation signature with the machine path removed.
  • docs/METHODOLOGY.md defines the contracts and caveats.

Credits

This repository contains configuration, scripts, and measurement metadata only. It does not redistribute model weights or third-party code.

License

The original material in this repository is released under the MIT License. oMLX and the model retain their own licenses.

About

Verified Qwen3.8 Flash Next oMLX Lightning MTP deployment guide for Apple Silicon

Topics

Resources

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages