Your Embedding Model is SMARTer Than You Think
AuthorsJianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee
SMART turns ordinary single-vector embedding models into stronger multimodal retrievers by reusing their hidden states for late-interaction retrieval without needing full retraining.
Key results
Controlled code-marker binding evaluation set
Pairwise accuracy on the toy benchmark
MAXSIM over final-layer hidden states on the toy benchmark
MMEB-V2 before SMART
MMEB-V2 with inference-only SMART
MMEB-V2 baseline average
What the paper found
Your Embedding Model is SMARTer Than You Think, from UW-Madison, Korea University, and NetApp, introduces SMART, a single-to-multi adaptation for retrieval transformers that turns off-the-shelf single-vector multimodal embedders into late-interaction retrievers without retraining from scratch. The core insight is that contrastive supervision on a pooled <eot> embedding already shapes the geometry of the model’s remaining hidden states, so SMART reuses those frozen states with ColBERT-style MAXSIM and combines them with the original pooled cosine score in a simple hybrid objective. On a controlled 5×5 code–marker binding benchmark with 1000 queries, the pooled single-vector score reaches only 31.9% pairwise accuracy, while late interaction over final-layer hidden states reaches 56.8%. On MMEB-V2, inference-only SMART improves the Qwen3-VL-Embedding-8B average from 78.83% to 79.34%, and the 2B model from 74.87% to 75.77%, while also lifting VLM2Vec-V2.0 from 64.50% to 67.04%. With a lightweight frozen-backbone adapter trained for 1 hour and 50 minutes on eight A6000 GPUs, Qwen3-VL-Embedding-2B surpasses jina-embeddings-v4 by 0.34 points on the visdoc subset. In a LoRA conversion study, a SMART-based single-vector-to-multi-vector conversion trained in 9.5 hours trails a from-scratch multi-vector model by only 0.63 average points, showing that SMART can recover local evidence, improve retrieval, and cut training compute by about 20%.
Original abstract
Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine-grained, local evidence critical for dense retrieval tasks. Multi-vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi-vector capabilities of standard single-vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late-interaction over these frozen hidden states during inference, SMART acts as a plug-and-play upgrade that consistently improves performance across diverse modalities, improving even the state-of-the-art models further on MMEB-V2. We also reveal SMART's superior performance, as simple lightweight post-training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single-vector model to outperform SoTA multi-vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval. We open source our code and weights at https://github.com/HanSolo9682/SMART.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.