PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
AuthorsSara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid
Resources
PANORAMA teaches vision-language models to describe an entire image while precisely linking every phrase to the pixels it refers to.
Key results
Human-annotated images in the panoptic grounded captioning benchmark.
Average image-pixel coverage by grounded panoptic regions.
Total multimodal training samples spanning captioning and segmentation capabilities.
Overall grounding score on the PanoCaps test benchmark.
Point reduction caused by removing mask proposal selection.
Zero-shot overall generalized grounding score on GroundingSuite.
What the paper found
PANORAMA addresses panoptic grounded captioning, where a vision-language model must describe foreground objects and background regions while attaching every referring phrase to pixel-level masks. The researchers introduce PanoCaps, a human-annotated benchmark with 3.5K images, 99.1% average pixel coverage, and phrase-mask alignments spanning diverse scenes. Their PANORAMA model uses Qwen3-VL to generate captions and contextualized [SEG] representations, projects each phrase into a 256-dimensional concept vector, and conditions the pretrained SAM 3 segmenter to produce phrase-specific mask proposals. A learned scorer then selects zero, one, or multiple proposals, separating semantic referent identification from boundary prediction and supporting plural references without collapsing instances. Trained within a 957K-sample mixture, PANORAMA-4B achieves 45.6 gPQ on PanoCaps, exceeding SAMTok-8B at 44.5 and outperforming generalist systems including Gemini 2.5 Pro and Qwen3-VL-235B-A22B on joint grounding quality. Ablations show that replacing proposal selection with direct mask decoding reduces gPQ by 7.5 points and AP50 by 20.0 points, while removing phrase-conditioned proposals costs 1.8 gPQ. The model also reaches 75.5 overall gIoU on the unseen GroundingSuite benchmark, demonstrating transfer to single-region, multi-region, and no-target grounding. Training uses LoRA and NVIDIA H100 GPUs, while the pretrained SAM 3 proposal machinery remains largely fixed.
Original abstract
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.