EMNLP 2026 Findings

Omni·Embed-MiniBinding modalities without forgetting via dense distillation

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly,

Ivan Laptev, Hisham Cholakkal

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)

Omni-Embed-Mini is a compact omni-modal embedding model that maps text, images, audio, speech, video, and documents into a unified embedding space without modifying its text backbone. At just 935M parameters, it preserves text retrieval quality while delivering strong cross-modal retrieval at a fraction of the size of existing omni-modal embedders.

Abstract

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone’s own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7× to 9.5× smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.

Shared space

Six modalities, one teacher: the backbone's own caption embedding

Every media sample is paired with a dense cascaded caption. The teacher target is that caption embedded by the identical frozen backbone that processes the student, so no auxiliary teacher, no projection head, and no drift in the text geometry. Pick a modality to play a real training row from the released corpus.

two rows per modality, straight from the corpus
Speech
46.7k rows · 21% of corpusmean caption 694 chars

Media

heysquad · 3.9s

row heysquad_057946_a01744cffd4b

Original caption

“What are organic plants understood to be?”

Sampled user prompt

describe what you hear

Dense cascaded caption · teacher target

A young adult female with a high-pitched, clear, and articulate voice delivers the line in a standard, neutral accent. Her speaking style is measured and calm, with a moderate pace and consistent volume, conveying a neutral, informative affect. She says <transcription> What are organic plants understood to be? </transcription>. The audio is clean and crisp, likely recorded in a quiet indoor space with no discernible background noise or reverberation.

454 characters, shown in full

System prompt · Speech

Listen to this speech audio. Write a single flowing paragraph covering: (1) the speaker (age group, gender, voice pitch/tone, clarity, accent); (2) speaking style (pace, loudness, emotion); (3) a transition into the transcription using a casual phrase followed by the marker [TRANSCRIPTION]; (4) audio quality, recording environment, background sounds. Do not write out the actual words spoken. The placeholder is post-processed into a transcription span around the source transcript.

Each panel plays an actual row from the released caption corpus, with its own media, source caption and cascaded dense caption. Audio and video are transcoded for the web; nothing else is edited.

Corpus composition

227.7k effective rows survive caption and duration filtering, from 590,858 total rows across 16 source datasets.

Speech: 46.7k rowsAudio: 45.2k rowsImage: 86.1k rowsVideo: 21.3k rowsVisual doc: 28.3k rows
227.7krows
Speech21%46.7k rows
Audio20%45.2k rows
Image38%86.1k rows
Video9%21.3k rows
Visual doc12%28.3k rows

Mean caption length

Visual documents run about six times longer, because the expert-analyst captioner enumerates layout, headings, tables and charts.

Speech: 694 ch · 106wAudio: 582 ch · 87wImage: 745 ch · 126wVideo: 732 ch · 127wVisual doc: 4,157 ch · 654w
1,382modality mean
Speech10%694 ch · 106w
Audio8%582 ch · 87w
Image11%745 ch · 126w
Video11%732 ch · 127w
Visual doc60%4,157 ch · 654w

Demo

Query in any modality, retrieve across all six

Open the demo
Image
Video
Audio
Speech
Visual doc
Text

One query, six modalities, one index. The row is an illustration; run your own against the live index in the hosted demo.

Results

Size versus performance

Ranked against every other omni embedder. At 0.9B our model is the smallest open omni embedder. The 2.3B variant reaches an overall-modality average of 51.39: it leads the omni block on video and is competitive on image and visual documents. A detailed breakdown can be found on the leaderboard.

Overall

Mean of the six per-modality scores

×100
0
20
40
60
36.3
51.4
47.3
48.3
49.3
47.5
52.9
49.3
49.5
Omni-Embed-Mini-0.9B
Omni-Embed-Mini-2.3B
BidirLM-Omni-2.5B
omni-embed-nemotron-3B
e5-omni-3B
LCO-Embedding-Omni-3B
e5-omni-7B
LCO-Embedding-Omni-7B
gemini-embedding-2
Left to rightsmaller to larger model, with the one undisclosed size last

Bars are the mean of each model’s six per-modality scores. The per-modality breakdown, parameter counts, full per-task tables and the all-class ranking, including text-only, vision-language and CLAP encoders, are on the leaderboard.

Ablations

Four design choices

A frozen backbone, dense captions, hybrid hard-negative mining and backbone LoRA, each isolated and measured on its own.

Backbone invariance

Adding five modalities leaves the untrained paths where they were

shareddifferenceBackbone alone → Omni-Embed-Mini

0.9B vs Qwen3-Embedding-0.6B

Text
49.4549.57
+0.12

2.3B vs Qwen3-VL-Embedding-2B

Text
47.0947.94
+0.85
Image
64.8065.67
−0.87
Video
55.1855.72
−0.54
Visual doc
58.1058.29
−0.19

The 0.9B's text path touches only unchanged backbone weights, never the trained projectors or LoRA, so training cannot move it: every checkpoint returns 49.57, across all three seeds and all four mining arms. The 2.3B leaves its vision path untrained too, so image, video and visual-doc behaviour is inherited rather than learned.

Dense captions

Caption density is the most important ingredient

shareddifferenceOriginal captions → Dense captions
Overall
23.7830.79
+7.01
Speech
41.7544.43
+2.68
Audio
31.0233.72
+2.70
Image
17.6326.23
+8.60
Video
4.7418.80
+14.06

Dense cascaded captions lift every modality. Speech gains least (+2.68), where source transcripts already carry most of the content; video gains most (+14.06), where short clip captions miss nearly every temporal event.

Hybrid mining

Two miners are needed; either one alone leaves a slack axis

mean over the 5 media modalities
Text onlytext FAISS
30.20
No miningin-batch only
31.47
Media onlyper-modality FAISS
32.07
Text + mediamain recipe
34.03

The ordering is text-only < no-mining < media-only < both. Text-only mining on its own hurts: the student sees harder text negatives but no harder media ones, which biases the projector. Visual documents are the least sensitive modality, all four arms within 0.6 nDCG@10; caption quality moves that one, not negative sampling.

Frozen vs LoRA

Unfreezing the backbone helps at one scale and hurts at the other

shareddifferenceBackbone + LoRA → Frozen backboneΔ = frozen − LoRA; red is a real regression, amber is under a point

0.9B

Text
46.4449.57
+3.13
Speech
44.4347.09
−2.66
Audio
33.7234.83
−1.11
Image
26.2330.04
−3.81
Video
18.8024.46
−5.66
Visual doc
46.9851.04
−4.06

2.3B

Text
47.2947.94
+0.65
Speech
31.0241.55
+10.53
Audio
25.5733.97
+8.40
Image
64.3464.80
+0.46
Video
54.6055.18
+0.58
Visual doc
58.0258.10
+0.08

At 0.9B, unfreezing lifts every media modality but costs 3.13 points of text. At 2.3B the same change degrades everything, worst on the two paths Phase 2 trains (speech 10.53, audio 8.40). Text improved at neither scale, so we keep the backbone frozen: it is the only setting that behaved consistently across both backbones we tried.

The dense-caption, mining and frozen-vs-LoRA panels are measured on ablation checkpoints. Backbone invariance compares the released checkpoints against the stock backbones they are built on.

Architecture

One frozen backbone, two passes, one tied teacher

The same media enters from both sides. On the left it takes the student path through the modality encoders and projectors; on the right it becomes a dense cascaded caption. Both meet in the identical frozen backbone, and gradient reaches only the projectors and the encoder LoRA adapters.

STUDENT PATHTEACHER PATHFrozen backboneQwen3-Embedding-0.6Bcausal, hidden 1024NO GRADIENTstudent passtext + spliced mediateacher passcaption only, cachedEOS poolSpeechAudioImageVideoVisual docTextDual audio encoderWhisper-smallDasheng-base128 tokens eachQwen3.5 ViT2×2 spatial mergervideo → 196 queriesshared by all threeTemporal interleavew₁ d₁ w₂ d₂ …256 tokens per 30 s chunkProjectors1-D down-projectorSwiGLU residualto hidden 1024no projector, no LoRA, no gradientSpeechAudioImageVideoVisual docCascaded captionerQwen3-Omni-30B-A3BDense captionone paragraph, per modalitythe only teacher signalMatryoshka SigLIP + cosine distillationℒ = ℒcontr + λ ℒdist, λ = 1MRL dims 128 / 256 / 512 / 1024, weighted 1/√dstudent and teacher share weights, so cosine alone sufficesŝ studentt̂ teacherOnline hybrid minertext FAISS · 10×/epoch · 50k captionsmedia FAISS · 5×/epoch · 10k per modality5 negatives, 0.92 cosine false-negative cutre-embeds through the live modelnegativesTRAINING SCHEDULE020%100%P1projectors onlyP2+ LoRA on Whisper, Dasheng, ViT; miner on
Speech and audio are both encoded twice, by Whisper on mel and by Dasheng on the raw waveform; the two streams emit 128 tokens each and are interleaved in time rather than mixed by a learned bottleneck, which costs no extra parameters and is intended to encourage the backbone’s causal attention to attend jointly to semantic speech content and acoustic texture within local context windows. Image, document and video share one ViT, with video additionally pooled by cross-attention against a learnable 196-query budget so sequence length is independent of frame count. Because the teacher embedding is produced by the same frozen weights, it is constant across training and is computed once and cached; only the media-side index drifts, and only toward harder negatives.
VariantInferenceBackboneVisionAudio encoders
Omni-Embed-Mini-0.9B935.35MQwen3-Embedding-0.6B, hidden 1024ViT extracted from Qwen3.5-0.8BWhisper-small + Dasheng-base
Omni-Embed-Mini-2.3B2440.21MQwen3-VL-Embedding-2B, native visionbackbone's own vision moduleWhisper-small + Dasheng-base

Citation

Cite Omni-Embed-Mini

If you find Omni-Embed-Mini useful in your research, please cite our paper.

@inproceedings{kurpath-etal-2026-omni-embed-mini,
    title = "{O}mni-{E}mbed-{M}ini: Binding Modalities Without Forgetting via Dense Distillation",
    author = "Kurpath, Mohammed Irfan and Kaithakkodan, Jaseel Muhammad and Mullappilly, Sahal Shaji and Laptev, Ivan and Cholakkal, Hisham",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2026",
    year = "2026",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/2610.02148"
}