EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass
AuthorsHaomin Peng, Yongkang Li, Zhaoxiang Liu, Xiaojie Jin, Shiguo Lian, Yunchao Wei, Xinggang Wang
EOVSAM turns SAM 3's slow exhaustive vocabulary search into a fast single-pass open-vocabulary segmenter while preserving competitive accuracy.
Key results
Maximum inference acceleration over vanilla SAM 3, measured as a multiplicative factor.
Open-vocabulary semantic segmentation score on ADE20K A-150.
Open-vocabulary panoptic quality on ADE20K.
A-150 inference speed on a single NVIDIA RTX 3090 at resolution 512.
What the paper found
EOVSAM, from researchers at Huazhong University of Science and Technology, China Unicom, and Beijing Jiaotong University, redesigns Meta’s SAM 3 for open-vocabulary segmentation without repeatedly running one text prompt per category. The method removes prompt cross-attention, converts SAM 3 into a prompt-free mask generator, and uses NVIDIA’s C-RADIOv4 to extract SAM 3 localization features alongside SigLIP 2 semantic features. Its central contribution, Attentional Aggregation, turns decoder attention maps into differentiable region embeddings, jointly optimizing mask localization and text-based recognition instead of classifying masks after generation. Trained on COCO Panoptic, EOVSAM reaches 39.0 mIoU on ADE20K A-150 and 30.9 PQ for open-vocabulary panoptic segmentation on ADE20K, while accelerating inference by up to 338 times over vanilla SAM 3. The same checkpoint remains effective at lower resolutions, delivering 8.76 FPS at resolution 512 on A-150 without retraining, compared with 3.77 FPS at resolution 1152, while retaining competitive accuracy. The experiments also show that inheriting SAM 3 weights is important for localization and that Attentional Aggregation prevents severe out-of-distribution recognition collapse, establishing a single-pass alternative to SAM 3’s vocabulary-scale multi-pass inference.
Original abstract
Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocabulary performance through exhaustive vocabulary traversal, yet suffers from prohibitive computational overhead as target categories scale. In this paper, we propose an Efficient Open-Vocabulary segmentation framework with SAM 3 (EOVSAM), which adapts SAM 3 for single-pass prediction. EOVSAM removes prompt conditioning to turn SAM 3 into an efficient mask generator and introduces a new Attentional Aggregation strategy to optimize open-vocabulary classification end-to-end. This formulation avoids the multi-stage pipelines and post-processing heuristics commonly used by existing methods, while mitigating the closed-set collapse that can arise when classification is optimized directly. EOVSAM consistently improves segmentation accuracy over vanilla SAM 3 on all evaluated datasets and accelerates inference by up to 338$\times$. Furthermore, EOVSAM maintains high accuracy at lower resolutions while achieving even more remarkable inference speeds. Experiments on standard semantic and panoptic segmentation benchmarks show that EOVSAM combines competitive or state-of-the-art accuracy with a substantial speed advantage over existing open-vocabulary segmentation models. Code and models are available at https://github.com/hustvl/EOVSAM.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.