NTH

InstructSAM: Segment Any Instance with Any Instructions

AuthorsYuqian Yuan, Wentong Li, Zhaocheng Li, Yutong Lin, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang, Wenqiao Zhang

June 14, 2026 2 min read
Watch on YouTube
The one-line take

InstructSAM lets users give natural-language instructions to segment multiple instances in one shot by bridging a vision-language model with SAM3 through learnable instance queries.

Key results

500K
Inst2 Seg training QA pairs

Large-scale instruction-mask training set size

3328
Inst2 Seg benchmark instructions

Manually verified evaluation instructions

986
Inst2 Seg images

Benchmark image count

31.5
InstructSAM Inst2 Seg mAP

Overall instance-level performance

8.3
InstructSAM vs SAM3-Agent mAP gain

Improvement over SAM3-Agent-Qwen2.5-VL-3B on Inst2 Seg

1.1
InstructSAM inference time

Average Inst2 Seg latency in seconds

What the paper found

InstructSAM is a unified instruction-driven segmentation framework from Zhejiang University and Nanjing University of Aeronautics and Astronautics that closes the gap between complex natural-language intent and instance-level mask prediction by turning segmentation into a set-prediction problem. Instead of making a vision-language model such as Qwen3-VL or Gemini emit [SEG] tokens autoregressively, it injects a bank of learnable instance queries into the LLM, lets them interact through a hybrid-attention scheme, and projects the resulting instruction-conditioned embeddings into SAM3’s detector query space for single-pass multi-instance segmentation. The paper also introduces Inst2 Seg, a new benchmark with 500K QA pairs for training and 3,328 manually verified instructions over 986 images, covering single-target, multi-target, no-target, and reasoning cases. Built on a Qwen3-VL-2B backbone with SAM3 initialization, the 2B-scale InstructSAM achieves 31.5 mAP and 60.4 gIoU on Inst2 Seg, outperforming multi-round SAM3-Agent-Qwen2.5-VL-3B by 8.3 mAP and 11.7 gIoU, while also improving ReasonSeg cIoU by 5.2 on the test set. It reaches 57.3 mAP and 68.3 cIoU on gRefCOCO val, and its latency on Inst2 Seg is 1.1 s versus 29.6 s for SAM3-Agent-Qwen3-VL-2B, showing that explicit query-based reasoning can be both more accurate and far faster than agentic pipelines.

Original abstract

In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven instance segmentation as a set-structured query prediction problem and propose an explicit reasoning-to-instance query interface that elegantly bridges a vision-language model (VLM) and SAM3. Specifically, a bank of learnable instance queries is injected into the VLM and contextualized with instruction and visual information, enabling each query to serve as an instance-aware slot. A hybrid-attention mechanism further promotes interaction among these queries, visual tokens, and instruction tokens, improving instance enumeration and reducing duplicate predictions. The resulting LLM-conditioned queries are projected into SAM3's detector query space to drive accurate multi-instance segmentation in a single forward pass. This design equips SAM3 with high-level instruction understanding, compositional reasoning, and instance-level set prediction without modifying its core architecture. To support training and evaluation, we further construct Inst2Seg, a high-quality and large-scale instruction-based instance segmentation dataset and benchmark that couples free-form instructions with instance-level masks. Extensive experiments show that only 2B-scale InstructSAM achieves strong results across complex instruction-driven and phrase-level referring segmentation benchmarks, outperforming prior end-to-end methods and SAM3's agentic pipeline while enabling efficient single-pass multi-instance prediction.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis