NTH

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

AuthorsMohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal

AffiliationsIvan Laptev, Hisham Cholakkal · Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)

October 2, 2026 2 min read
Watch on YouTube
The one-line take

A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.

Key results

49.57
0.9B text retrieval

MTEB-v2 BEIR-8 nDCG@10 score with the text backbone frozen

227704
Filtered training rows

Effective non-text samples used after caption and duration filtering

51.39
2.3B overall modality average

Average performance across text, speech, audio, image, video, and visual-document evaluations

55.18
2.3B video score

MMEB-V2 average hit@1 score for video

7.01
Dense-caption gain

Point reduction in the four-modality mean when dense captions are replaced by original captions

What the paper found

Omni-Embed-Mini addresses the multimodal expansion dilemma: adding media inputs often damages text retrieval or requires models with billions of parameters. Its 0.9B variant maps text, speech, audio, images, video, and visual documents into one cosine space while keeping the Qwen3-Embedding-0.6B text backbone frozen and bit-identical. Each media example is paired with a dense caption generated by Qwen3-Omni-30B-A3B-Instruct; the frozen backbone embeds that caption as a teacher target, while lightweight projectors and phased LoRA adapters train the media encoders. The objective combines Matryoshka SigLIP contrastive learning, cosine self-distillation, and an online hybrid hard-negative miner. Training uses 227704 filtered media rows, and the 0.9B model reaches 49.57 nDCG@10 on the 8-task MTEB-v2 BEIR subset without changing text retrieval. Its 2.3B Qwen3-VL-Embedding-2B variant achieves a 51.39 overall-modality average, including 55.18 on MMEB-V2 video, and edges out closed Gemini-embedding-2 at 49.51 overall. Dense captions matter substantially: replacing them with original captions reduces the four-modality mean by 7.01 points. The approach is smaller than open systems such as NVIDIA’s omni-embed-nemotron-3B and e5-omni, while extending OpenAI-style Matryoshka representations across six modalities, although speech and general audio remain weaker than its vision and text results.

Original abstract

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com

Read the original paper

More in Multimodal AI

Browse all 61 papers →
01Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
02Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis
03Multimodal

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.

Read analysis