Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong, and Yulun Zhang
"Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction", arXiv 2026
Project · Training data · Evaluation
- 2026-09-25: This repo (code and data) is released.
Abstract: Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher.
Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
- Release checkpoints.
src/
├── training/ # Training configuration, launcher, and checkpoint export
├── verl/ # Training core, JSD, Qwen3.5 integration, and CDPruner
├── data/ # LT-OPD-14K preparation
└── evaluation/ # Benchmark inference and scoring
If you want to use the codes, please:
cd srcUse Python 3.12 and a CUDA environment. Install the package and its training dependencies:
pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu126
pip install -e '.[train,eval]'
pip install flash-attn==2.8.3 --no-build-isolationLT-OPD-14K contains 14,000 examples from OneThinker, PixMo, LLaVA, TextVQA, and Vision-OPD. The Hugging Face release keeps the training order and ships every image in content-addressed tar shards, with per-sample provenance and SHA-256 recorded in media.jsonl.
python -m data.prepare --dataset yyy051007/LT-OPD-14K --output data/LT-OPD-14KPreparation verifies the release checksums and all 14,000 images before writing the training files. Pass --source-media to reuse an image directory you already have.
Download the base model, then launch training:
hf download Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--local-dir models/Qwen3.5-4B
python -m training.train \
--model models/Qwen3.5-4B \
--data-dir data/LT-OPD-14K \
--output outputs/lt-opdThe default configuration uses eight GPUs and full-parameter training, including the visual encoder and merger.
| Setting | Value |
|---|---|
| Training updates | 175 |
| Batch | 80 questions × 8 student rollouts |
| Teacher | Frozen Qwen3.5-4B with all visual tokens |
| Objective | Token-level JSD |
| Visual-token curriculum | 25% for 14 updates, cosine decay, 5% for the final 75 updates |
Set --gpus and --nodes for your hardware. Use --resume to continue a saved run.
Export the trained checkpoint:
python -m training.export \
--checkpoint /path/to/saved_checkpoint \
--base-model models/Qwen3.5-4B \
--output outputs/lt-opd/exportThe evaluation suite covers V*Bench, HRBench-4K, GQA, MMMU, MMBench, MME, POPE, TextVQA, and OCRBench. Inference and scoring are separate commands; dataset preparation, protocol details, and scorer versions are listed in the evaluation guide.
We present the performance of LT-OPD compared with previous SOTA methods.
Main Results (click to expand)
- Results in Tab.2 of the main paper
- Results in Fig. 4 of the main paper (compared with RL algorithms)
If you find our dataset and code helpful in your research or work, please cite the following paper.
@article{li2026fewer,
title={Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction},
author={Li Junxian and Yang Ruixuan and Zhang Tianao and Xu Tiange and Dong Weisheng and Zhang Yulun},
journal={arXiv preprint arXiv:2609.32353},
year={2026}
}Built on verl, SDPO, and CDPruner. Benchmark scorers retain their upstream attribution and versions.




