Omni·Embed-MiniBinding modalities without forgetting via dense distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly,
Ivan Laptev, Hisham Cholakkal
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Omni-Embed-Mini is a compact omni-modal embedding model that maps text, images, audio, speech, video, and documents into a unified embedding space without modifying its text backbone. At just 935M parameters, it preserves text retrieval quality while delivering strong cross-modal retrieval at a fraction of the size of existing omni-modal embedders.
Abstract
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone’s own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7× to 9.5× smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.
Shared space
Six modalities, one teacher: the backbone's own caption embedding
Every media sample is paired with a dense cascaded caption. The teacher target is that caption embedded by the identical frozen backbone that processes the student, so no auxiliary teacher, no projection head, and no drift in the text geometry. Pick a modality to play a real training row from the released corpus.
Media
heysquad · 3.9s
row heysquad_057946_a01744cffd4b
Original caption
“What are organic plants understood to be?”
Sampled user prompt
describe what you hear
Dense cascaded caption · teacher target
A young adult female with a high-pitched, clear, and articulate voice delivers the line in a standard, neutral accent. Her speaking style is measured and calm, with a moderate pace and consistent volume, conveying a neutral, informative affect. She says <transcription> What are organic plants understood to be? </transcription>. The audio is clean and crisp, likely recorded in a quiet indoor space with no discernible background noise or reverberation.
454 characters, shown in full
System prompt · Speech
Listen to this speech audio. Write a single flowing paragraph covering: (1) the speaker (age group, gender, voice pitch/tone, clarity, accent); (2) speaking style (pace, loudness, emotion); (3) a transition into the transcription using a casual phrase followed by the marker [TRANSCRIPTION]; (4) audio quality, recording environment, background sounds. Do not write out the actual words spoken. The placeholder is post-processed into a transcription span around the source transcript.
Each panel plays an actual row from the released caption corpus, with its own media, source caption and cascaded dense caption. Audio and video are transcoded for the web; nothing else is edited.
Corpus composition
227.7k effective rows survive caption and duration filtering, from 590,858 total rows across 16 source datasets.
Mean caption length
Visual documents run about six times longer, because the expert-analyst captioner enumerates layout, headings, tables and charts.
Demo
Query in any modality, retrieve across all six
One query, six modalities, one index. The row is an illustration; run your own against the live index in the hosted demo.
Results
Size versus performance
Ranked against every other omni embedder. At 0.9B our model is the smallest open omni embedder. The 2.3B variant reaches an overall-modality average of 51.39: it leads the omni block on video and is competitive on image and visual documents. A detailed breakdown can be found on the leaderboard.
Overall
Mean of the six per-modality scores
Bars are the mean of each model’s six per-modality scores. The per-modality breakdown, parameter counts, full per-task tables and the all-class ranking, including text-only, vision-language and CLAP encoders, are on the leaderboard.
Ablations
Four design choices
A frozen backbone, dense captions, hybrid hard-negative mining and backbone LoRA, each isolated and measured on its own.
Adding five modalities leaves the untrained paths where they were
0.9B vs Qwen3-Embedding-0.6B
2.3B vs Qwen3-VL-Embedding-2B
The 0.9B's text path touches only unchanged backbone weights, never the trained projectors or LoRA, so training cannot move it: every checkpoint returns 49.57, across all three seeds and all four mining arms. The 2.3B leaves its vision path untrained too, so image, video and visual-doc behaviour is inherited rather than learned.
Caption density is the most important ingredient
Dense cascaded captions lift every modality. Speech gains least (+2.68), where source transcripts already carry most of the content; video gains most (+14.06), where short clip captions miss nearly every temporal event.
Two miners are needed; either one alone leaves a slack axis
The ordering is text-only < no-mining < media-only < both. Text-only mining on its own hurts: the student sees harder text negatives but no harder media ones, which biases the projector. Visual documents are the least sensitive modality, all four arms within 0.6 nDCG@10; caption quality moves that one, not negative sampling.
Unfreezing the backbone helps at one scale and hurts at the other
0.9B
2.3B
At 0.9B, unfreezing lifts every media modality but costs 3.13 points of text. At 2.3B the same change degrades everything, worst on the two paths Phase 2 trains (speech 10.53, audio 8.40). Text improved at neither scale, so we keep the backbone frozen: it is the only setting that behaved consistently across both backbones we tried.
The dense-caption, mining and frozen-vs-LoRA panels are measured on ablation checkpoints. Backbone invariance compares the released checkpoints against the stock backbones they are built on.
Architecture
One frozen backbone, two passes, one tied teacher
The same media enters from both sides. On the left it takes the student path through the modality encoders and projectors; on the right it becomes a dense cascaded caption. Both meet in the identical frozen backbone, and gradient reaches only the projectors and the encoder LoRA adapters.
| Variant | Inference | Backbone | Vision | Audio encoders |
|---|---|---|---|---|
| Omni-Embed-Mini-0.9B | 935.35M | Qwen3-Embedding-0.6B, hidden 1024 | ViT extracted from Qwen3.5-0.8B | Whisper-small + Dasheng-base |
| Omni-Embed-Mini-2.3B | 2440.21M | Qwen3-VL-Embedding-2B, native vision | backbone's own vision module | Whisper-small + Dasheng-base |
Citation
Cite Omni-Embed-Mini
If you find Omni-Embed-Mini useful in your research, please cite our paper.
@inproceedings{kurpath-etal-2026-omni-embed-mini,
title = "{O}mni-{E}mbed-{M}ini: Binding Modalities Without Forgetting via Dense Distillation",
author = "Kurpath, Mohammed Irfan and Kaithakkodan, Jaseel Muhammad and Mullappilly, Sahal Shaji and Laptev, Ivan and Cholakkal, Hisham",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2026",
year = "2026",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/2610.02148"
}