Abstract
Multi-token-prediction heads inherit an expensive shape: a model-width transformer block followed by an output projection over the target's full vocabulary. On Qwen3.5-9B that means scoring 248,320 vocabulary rows at every draft step. This study asked whether a separately trained, half-width MTP block with its own 32,768-token output head could keep enough acceptance to speed up the full speculative-decoding path. The distillation worked on its own terms: the best StudentSV checkpoint reached 0.507 recursive chain acceptance against the teacher's 0.593 (85.5%) at about an eighth of the draft compute, with slots 1 to 3 each above the teacher. The study still concluded against it. A verifier-as-oracle audit sorted every compression into two families. Fidelity cuts (precision) cost at most 1.9 acceptance-retention points even under severe distribution shift. Function cuts (pruning, low-rank, and the distilled student) keep 97% on calibration-style text and lose 10 to 19 points where the text actually shifts, because parameters that serve rare content collect no mass on any finite calibration sample. The recipe that survives needs no training: trim the vocabulary to the model's own 32k most used tokens, quantize to NVFP4, verify against the full vocabulary. Measured end to end in the memra engine on one RTX PRO 6000, it delivers 1.8x and 2.2x over plain decoding on a short-prompt suite and 2.0x and 2.7x on a long-context suite, at 9B and 27B.
Concluded result, measured wall time
- end-to-end over plain decoding, 9B and 27B
- 1.8x–2.7x
- trimmed NVFP4 draft head at 9B
- 212 MB
- off-distribution tax on function cuts
- 10–19 pt
- tax on precision cuts
- ≤1.9 pt
1. Research question
Large-vocabulary speculative drafters can spend a surprising amount of time moving and multiplying the output head. FR-Spec and VocabTrim showed that a static high-frequency subset can cut that cost while the full target still protects output correctness. In my first Qwen3.5-9B measurements on the llama.cpp path, trimming removed roughly 85% of the draft LM-head kernel cost but produced only a 1 to 3% end-to-end gain: the MTP transformer block remained, and that serving path was not draft-bound. Section 4 records how this number reversed once the same recipe ran in an engine whose critical loop is the draft path.
So the question is not only which logits the drafter should score. It is:
Can a small block and a small vocabulary be trained as one deployment-specific drafter, keeping enough acceptance to beat the inherited co-trained head on wall time and memory?
The project started as teacher-to-student distillation, and the first ablations seemed to reject it: equal-budget KD lost to plain cross-entropy twice. That verdict was a loss bug, not a finding; section 4 documents the reversal. The final recipe was distillation first: soft cross-entropy against precomputed teacher logits, then a short CE reinforcement.
Two compression axes in one drafter
2. Experimental design
Subject and interface
The subject is Qwen3.5-9B with its co-trained MTP block. Each prediction receives the current token embedding and the predecessor trunk hidden state, projects the pair into the draft block, and scores the next token. The experiment holds the target trunk fixed and replaces only this proposal path.
The student
StudentSV is a 212M-parameter drafter: a half-width draft block and an independent 32,768-row LM head initialized from frequently used rows of the target head. The shortlist was calibrated separately for code and conversation; the main run was code-weighted.
Training and controls
- 960,000 training samples drawn from a mixed corpus of about seven million tokens.
- Cross-entropy training without a teacher forward pass in the main arm.
- A full-vocabulary half-width baseline, to separate block compression from vocabulary compression.
- The co-trained MTP head measured both in the engine and in a windowed PyTorch reproduction.
- Held-out code and generation-distribution replay sets, with no evaluation text in training.
- Training on the model's own outputs beat a generic corpus by 17 points at a third of the steps (0.548 against 0.376 held-out acceptance). The training distribution should match the serving distribution.
- Chain training: four recursive draft slots against precomputed teacher top-64 logits, soft cross-entropy over the full 32,768-row draft vocabulary.
Metrics
Slot-0 teacher-forced replay acceptance (how often the drafter's first proposal matches the target's greedy token on a fixed hidden-state track) is the training diagnostic. Self-drafting chain acceptance over four recursive slots is the deployment-shaped metric: each slot drafts from the previous slot's own output, which is how speculation runs. End-to-end wall time was measured later in the memra engine (section 3.3).
3. Results
3.1 The student
| Draft head | Held-out code | Generation distribution | Est. draft FLOPs |
|---|---|---|---|
| Engine teacher, full context | 0.678 | 0.348 | 1.00x |
| PyTorch teacher, 2k window | 0.675 | 0.371 | 1.00x |
| StudentSV, 960k samples | 0.527 | 0.166 | ~0.13x |
| StudentSV, 60k samples | 0.337 | 0.105 | ~0.13x |
| Half-width, full vocabulary, 60k | 0.373 | n/a | ~0.35x |
Acceptance was still buying scale
The scaled StudentSV run kept 78% of the windowed teacher's code acceptance at about an eighth of the draft FLOPs. The jump from 60k to 960k samples was large and the training loss was still falling, so the scaling lever was real. It did not establish a speedup by itself.
The generation-distribution score was much weaker, and the later audit explained why the gap was structural rather than fixable by another run: only about 20% of the corpus was conversational, the chat shortlist covered 94.6% of target tokens against 98.9% on code, and no finite calibration sample collects mass on the parameters that serve rare content. That mechanism closed the student direction.
3.2 Chain distillation
| Checkpoint | Slot 0 | Slot 1 | Slot 2 | Slot 3 | Chain |
|---|---|---|---|---|---|
| Teacher, engine chain | 0.740 | 0.598 | 0.520 | 0.470 | 0.593 |
| Staged CE then KD, ~24k steps | 0.679 | 0.563 | 0.468 | 0.405 | 0.449 |
| KD only, 12k steps | 0.656 | 0.563 | 0.488 | 0.425 | 0.449 |
| + CE reinforcement, 4k | 0.665 | 0.570 | 0.500 | 0.459 | 0.457 |
| KD only, 24k steps | 0.706 | 0.598 | 0.527 | 0.448 | 0.490 |
| + CE reinforcement, 4k | 0.708 | 0.611 | 0.548 | 0.477 | 0.507 |
Distilling directly from precomputed teacher logits, with no CE warmup, matched the staged recipe at half the steps, and doubling the distillation schedule before the CE reinforcement set the best result: 0.507 chain acceptance, 85.5% of the teacher's 0.593. Slots 1 to 3 each beat the teacher (+0.013, +0.028, +0.007). The whole remaining gap was slot 0, which chain acceptance compounds. The 24k distillation curve was near saturation, so the steps lever was spent. The remaining levers, more own-output data and a wider block, were never pulled: the off-distribution audit closed the direction first.
The student tracks the teacher into depth
3.3 The recipe end to end
The recipe that survived needs no training: trim the head's vocabulary to the model's own 32k most used tokens, quantize it, and keep full-vocabulary verification, so outputs stay distribution-exact by construction. The head's per-token reads drop from 1.34 GB to 0.21 GB. Measured in memra on one RTX PRO 6000 (96 GB) at K=4, on an 80-prompt short suite (MT-Bench) and a 20-prompt long suite of 1.9k to 5.9k-token prompts:
| Draft head | 9B short | 9B long | 27B short | 27B long |
|---|---|---|---|---|
| Co-trained head, unmodified | 1.52x | 1.72x | 2.02x | 2.42x |
| Trimmed, Q8 | 1.78x | 2.04x | 2.15x | 2.59x |
| Trimmed, NVFP4 | 1.82x | 1.98x | 2.15x | 2.67x |
| Trimmed, NVFP4 projection with a Q4_0 block | 1.81x | 2.09x | 2.04x | 2.70x |
Greedy aggregate cells, speedup over plain decoding on the same suite. Sampled cells (temperature 0.7, three seeds) track greedy with a 0.1 to 0.2x penalty. A trimmed 4-bit head is the fastest or tied-fastest arm in every cell, and every arm gains more at 27B than at 9B. The fastest cell, 27B long, reaches 128.6 tok/s against 47.7 for plain decoding. Trimming costs 0 to 3 acceptance points because the hot vocabulary covers 95 to 99% of generated tokens; 4-bit quantization costs another 1.4 to 5.3, and the bandwidth saving buys that back in every cell but one: at 9B long, Q8's higher acceptance beats all-NVFP4 by 3%.
| Compressed head | In-distribution | Same register | Domain shift | Shift tax (pt) |
|---|---|---|---|---|
| Q8 | 99.8 | 100.0 | 99.9 | −0.1 |
| Q4_0 | 99.4 | 99.5 | 97.7 | 1.7 |
| NVFP4 | 99.3 | 100.1 | 97.5 | 1.9 |
| Pruned, 25% of channels kept | 97.2 | 95.0 | 80.4 | 16.8 |
| Pruned, 50% kept | 98.4 | 97.9 | 88.2 | 10.2 |
| Low-rank, data-free SVD | 89.9 | 85.1 | 71.1 | 18.7 |
| Low-rank, calibrated SVD | 90.5 | 87.0 | 73.7 | 16.8 |
Acceptance retention (%) against the reference head on identical states, 2,500 positions per corpus. The domain-shift corpus is German, Chinese, Hebrew and Russian Wikipedia, PubMed abstracts and C4 web text, where the reference head's own accuracy drops from 0.88 to 0.36. Fidelity cuts stay within noise. Function cuts pay 10 to 19 points, and same-register evaluation shows only a tenth of that.
Composition with a quantized base. On a Q4_K_M build of Qwen3.6-27B, which decodes 1.49x faster than the BF16 base by itself, a trimmed mixed-precision head reached 137.5 tok/s on the long suite: 2.88x over BF16 plain decoding. That number combines two separate gains, so it is not the head's speedup and is not a headline here.
3.4 An audit before deployment
Shifted test data is what a deployment rarely has. Running a compressed head beside its reference on in-distribution text gives two numbers: how often they disagree, and how much each disagreement hurts. Harm per disagreement separates the families (0.14 for fidelity cuts against 0.28 to 0.54 for function cuts), and the disagreement rate ranks the off-distribution tax exactly on the domain-shift suite (Spearman 1.0, exact p = 0.0002). It takes minutes of forward passes.
4. Negative results and changed decisions
- RefutedQuantization alone breaks MTP agreement.
Clean NVFP4 against BF16 controls found no degradation, so the original premise of healing the head was closed.
- RefutedA distilled student holds up off-distribution.
Function cuts, the student included, keep 97% on calibration-style text and lose 10 to 19 points under real shift. Six zero-training selection schemes for pruning all failed the same way, including one calibrated on the pruned head's own failures.
- RevisedEarly KD losses were a loss bug, not a verdict.
The first distillation runs log-softmaxed only the teacher's top-64 subset and left the rest of the draft vocabulary unconstrained, so pure KD collapsed and CE looked better. Fixed soft-CE distillation over the full vocabulary beat the staged recipe on every slot at two thirds of the steps.
- RevisedTrim the inherited head and stop.
Trim-only on the llama.cpp path saved most LM-head kernel work but only 1 to 3% end to end, because the transformer block remained and that serving path was not draft-bound. The verdict reversed in the memra engine: with the draft path on the critical loop, trim plus NVFP4 reached 1.8x to 2.7x across both models and suites. The lever was real; the first harness hid it.
- RetractedAn early 2.68x headline.
A single-prompt sampled cell failed the aggregate protocol: one prompt swings 12 acceptance points across seeds. It was retracted, and every number on this page is a suite aggregate.
5. What this does not establish
- Wall time comes from one GPU type (RTX PRO 6000) in one engine (memra). Memory accounting is an estimate.
- The evidence covers two scales of one model family. Other families are untested.
- The StudentSV FLOP ratio is an architectural estimate.
- The student was undertrained and below the acceptance-retention threshold the project wanted.
- The student's 2k training and evaluation window says nothing about long-context behavior.
6. Artifacts and release plan
The harness and the paper are public at github.com/avifenesh/hqmtp: training and evaluation code, fixed corpora, the run ledgers behind every number here, the negative-result verdicts, and the paper draft (paper/PAPER.md). The engine is memra, and the trim-only motivation is public in llama.cpp issue #25187.
- Available now Harness, protocol, run ledgers, verdicts and the paper draft.
- Next Checkpoints and extracted training data with a provenance manifest.
- Paper v1.0 Review rounds closed, PDF and BibTeX.
References
- W. Zhao et al. FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling. ACL 2025.
- R. Goel et al. VocabTrim: Vocabulary Pruning for Efficient Speculative Decoding in LLMs. 2025.
- H. Cai et al. FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction. 2025.
Changes
Substantive result changes raise the version and stay listed here. No preliminary number is silently promoted to a final one.
- 0.6, 4 Oct 2026Headline aligned with the written paper: 1.8x to 2.7x end to end across 9B and 27B and both suites, three-seed aggregates. Added the end-to-end and robustness tables and the audit. The earlier 2.88x figure is a composed cell (trimmed head on a quantized base, against a BF16 baseline); it now sits in section 3.3 with that context.
- 0.5, 5 Aug 2026Study concluded: function cuts pay an off-distribution tax that fidelity cuts do not, and the trimmed head wins.