SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
AuthorsZhixiong Zhang, Yizhuo Li, Shuangrui Ding, Yuhang Zang, Shengyuan Ding, Long Xing, Yibin Wang, Qiaosheng Zhang, Jiaqi Wang
Resources
SetCon improves referring segmentation by having the model predict shared semantic concepts for whole target sets, then refining them into subgroups for stronger image and video grounding.
Key results
The two-stage annotation pipeline augments reasoning segmentation datasets with hierarchical semantic supervision for training SetCon.
SetCon achieves state-of-the-art performance on gRefCOCO val, surpassing the previous best by 3.3 gIoU.
SetCon achieves state-of-the-art performance on MUSE val, surpassing the previous best by 12.1 gIoU.
On the MeViS benchmark, SetCon sets a new state of the art with a 10.9 J&F improvement.
On Ref-SeCVOS, SetCon sets a new state of the art with a 12.4 J&F improvement.
What the paper found
SetCon reframes open-ended referring segmentation as explicit set-level concept prediction instead of the standard per-target [SEG] token interface used by LVLM-based models such as LISA, PixelLM, and Sa2VA. The paper shows that flat [SEG] sequences degrade sharply as target count increases and that learned [SEG] embeddings align more with spatial position than semantic category, causing duplicated or missing masks. SetCon replaces these tokens with natural-language concept spans generated by a Qwen3-VL-8B-Instruct backbone and uses their hidden states as semantic conditions for joint mask-set decoding with SAM 3, adding a hierarchical decomposition that first predicts one global set concept and then finer sub-category concepts. To train this interface, the authors build a two-stage annotation pipeline with Qwen3-VL-235B-A22B, producing 236,396 samples and 784,809 concept phrases, with hierarchical supervision over 80% multi-subcategory cases. On image benchmarks, SetCon reaches state-of-the-art performance on gRefCOCO and MUSE, improving by 3.3 gIoU and 12.1 gIoU on their validation sets, respectively, while remaining competitive on RefCOCO/+/g and achieving the best average on ReasonSeg. On video, it sets new state of the art on seven benchmarks, including +10.9 J&F on MeViS and +12.4 J&F on Ref-SeCVOS. The key novelty is that semantic concepts, not special tokens, become the interface between language reasoning and mask prediction, making multi-instance and cross-category grounding both more interpretable and more stable.
Original abstract
Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-category groups, or open-ended target sets remains challenging. Previous Large Vision Language Model (LVLM)-based methods represent referred targets with one or more special tokens sequentially, treating multiple targets as separate outputs rather than a coherent set and offering little incentive to capture set-level properties such as completeness and mutual exclusivity. We reformulate open-ended referring segmentation as explicit set-level concept prediction and propose Set-Concept Segmentation (SetCon), which uses LVLM-generated natural-language concepts, instead of segmentation-specific tokens, as semantic conditions for joint mask-set decoding. A hierarchical semantic decomposition first predicts a shared set-level concept defining the target scope and then refines it into fine-grained concept groups aligned with target subsets. To support this, a two-stage annotation pipeline augments existing reasoning segmentation datasets with hierarchical semantic supervision (236k samples, 784k concept phrases). SetCon achieves state-of-the-art results on image benchmarks (+3.3 gIoU on gRefCOCO, +12.1 gIoU on MUSE), with margins that grow as the number of referred targets increases. The concept interface also transfers to video under a detect-and-track setting, yielding new state-of-the-art results on seven referring video benchmarks, including +10.9 J&F on MeViS and +12.4 J&F on Ref-SeCVOS.
Read the original paper