NTH

Your Embedding Model is SMARTer Than You Think

AuthorsJianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee

June 14, 2026 2 min read
Watch on YouTube
The one-line take

SMART turns ordinary single-vector embedding models into stronger multimodal retrievers by reusing their hidden states for late-interaction retrieval without needing full retraining.

Key results

1000
toy benchmark queries

Controlled code-marker binding evaluation set

31.9%
single-vector accuracy

Pairwise accuracy on the toy benchmark

56.8%
late-interaction accuracy

MAXSIM over final-layer hidden states on the toy benchmark

78.83%
Qwen3-VL-Embedding-8B average

MMEB-V2 before SMART

79.34%
Qwen3-VL-Embedding-8B + SMART average

MMEB-V2 with inference-only SMART

64.50%
VLM2Vec-V2.0 average

MMEB-V2 baseline average

What the paper found

Your Embedding Model is SMARTer Than You Think, from UW-Madison, Korea University, and NetApp, introduces SMART, a single-to-multi adaptation for retrieval transformers that turns off-the-shelf single-vector multimodal embedders into late-interaction retrievers without retraining from scratch. The core insight is that contrastive supervision on a pooled <eot> embedding already shapes the geometry of the model’s remaining hidden states, so SMART reuses those frozen states with ColBERT-style MAXSIM and combines them with the original pooled cosine score in a simple hybrid objective. On a controlled 5×5 code–marker binding benchmark with 1000 queries, the pooled single-vector score reaches only 31.9% pairwise accuracy, while late interaction over final-layer hidden states reaches 56.8%. On MMEB-V2, inference-only SMART improves the Qwen3-VL-Embedding-8B average from 78.83% to 79.34%, and the 2B model from 74.87% to 75.77%, while also lifting VLM2Vec-V2.0 from 64.50% to 67.04%. With a lightweight frozen-backbone adapter trained for 1 hour and 50 minutes on eight A6000 GPUs, Qwen3-VL-Embedding-2B surpasses jina-embeddings-v4 by 0.34 points on the visdoc subset. In a LoRA conversion study, a SMART-based single-vector-to-multi-vector conversion trained in 9.5 hours trails a from-scratch multi-vector model by only 0.63 average points, showing that SMART can recover local evidence, improve retrieval, and cut training compute by about 20%.

Original abstract

Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine-grained, local evidence critical for dense retrieval tasks. Multi-vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi-vector capabilities of standard single-vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late-interaction over these frozen hidden states during inference, SMART acts as a plug-and-play upgrade that consistently improves performance across diverse modalities, improving even the state-of-the-art models further on MMEB-V2. We also reveal SMART's superior performance, as simple lightweight post-training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single-vector model to outperform SoTA multi-vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval. We open source our code and weights at https://github.com/HanSolo9682/SMART.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis