NTH

Vision as Unified Multimodal Generation

AuthorsXiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, Siyi Xie, Xin You, Jinsheng Quan, Zhongang Cai, Haiwen Diao, Ziwei Liu, Lei Yang, Dahua Lin, Quan Wang

July 12, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows how a single multimodal model can handle detection, segmentation, depth, OCR, and more by turning all vision tasks into text-and-image generation.

Key results

50M
SN-VC-50M

released generated and curated computer-vision instruction-response examples

980
SigLIP2 resolution

maximum input resolution used for fine spatial conditioning

10
max views

maximum multi-view inputs sampled per training example

56.6
COCO-Common F1@mIoU

structured visual understanding performance

80.5
RefCOCOg test cIoU

referring benchmark result on structured visual understanding

4.0
NYUv2 abs rel

depth estimation result

What the paper found

Vision as Unified Multimodal Generation, from SenseTime Research and collaborators at Nanyang Technological University, reframes computer vision as a single multimodal generation problem inside a unified multimodal model rather than a collection of task-specific heads. Built on the Bagel-7B-MoT backbone, SenseNova-Vision converts heterogeneous supervision into instruction-response pairs spanning text, image, and mixed text-image outputs, and trains on the SenseNova-Vision Corpus, including the released SN-VC-50M subset with 50 million converted examples. The model covers structured perception, dense geometry, segmentation, and multi-view 3D, using standard cross-entropy for text outputs and rectified-flow for visual outputs, with SigLIP2 inputs up to 980 pixels and multi-view samples capped at 10 views. On structured visual understanding, it matches the best detector on COCO-Common at 56.6 and reaches 80.5 on RefCOCOg test; on dense geometry it obtains 4.0 abs rel and 98.1 delta-1 on NYUv2 depth, plus 12.8 mean error and 68.9 percent delta-11.25 on ScanNet normals. For segmentation, it reaches 63.2 gIoU on reasoning segmentation and 73.9 mIoU on interactive segmentation with point prompts. In multi-view geometry, it records 87.9 F1 on 7Scenes reconstruction and 80.1 AUC@30 on CO3Dv2 camera pose. The paper’s central claim is that native text-and-image generation can unify symbolic records, dense maps, masks, point clouds, and camera poses without architectural specialization, while preserving broad multimodal abilities such as 79.0 on MMVP and 0.85 on GenEval.

Original abstract

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis