NTH

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

AuthorsHaozhan Shen, Tiancheng Zhao, Kangjia Zhao, Jianwei Yin

June 5, 2026 2 min read
Watch on YouTube
The one-line take

The paper asks whether vision-language models or video generation models better support spatial intelligence, and finds they encode complementary strengths that can be combined for stronger representations.

Key results

92.08
ScanNet20 mAP avg gain for VLMs

Average semantic tagging mAP for frozen VLMs on ScanNet20, compared against 69.89 for VGMs.

87.28
ScanNet20 APmid avg gain for VLMs

Average APmid for frozen VLMs on ScanNet20, compared against 58.63 for VGMs.

22.66
Instance grouping T-mIoU avg for VLMs

Average multi-view instance grouping T-mIoU for frozen VLMs, compared against 13.24 for VGMs.

11.23
Instance grouping T-SR avg for VLMs

Average instance grouping T-SR for frozen VLMs, compared against 4.35 for VGMs.

0.152
VGM point-map error avg

Average 3D geometry point-map error for VGMs on DL3DV, lower than the VLM average of 0.223.

0.615
Fusion model AUC@30

Feature-level fusion of WAN2.1-T2V-14B and Qwen3-VL-8B achieved this camera pose AUC@30 on the 3D geometry task.

What the paper found

This paper, from Zhejiang University and Om AI Research, asks which pretraining paradigm better supports spatial intelligence: Vision-Language Models such as OpenAI-style language-aligned backbones Qwen3-VL, Qwen2.5-VL, and InternVL3, or Video Generation Models such as WAN, CogVideoX, OpenSora-2.0, and Aether. The authors freeze each model and probe intermediate features with identical lightweight heads on three spatial tasks: semantic tagging on ScanNet20, multi-view instance grouping on ScanNet masks, and 3D geometry prediction on DL3DV using VGGT-generated depth, point maps, and camera poses. The results show a clean division of strengths. VLMs are dramatically stronger at semantics, raising ScanNet20 mAP from 69.89 to 92.08 on average and APmid from 58.63 to 87.28, while also outperforming VGMs on instance grouping, with average T-mIoU improving from 13.24 to 22.66 and T-SR from 4.35 to 11.23. In contrast, VGMs encode geometry more directly, cutting point-map error from 0.223 to 0.152, lowering AbsRel from 0.113 to 0.072, and raising camera AUC@30 from 0.330 to 0.527. A naive feature concatenation of WAN2.1-T2V-14B and Qwen3-VL-8B already combines these advantages, reaching 92.30 mAP, 23.70 T-mIoU, 0.042 AbsRel, and 0.615 AUC@30, which suggests the two representation families are complementary rather than competing.

Original abstract

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation backbones: Vision-Language Models (VLMs), which use language supervision to align visual observations with semantic concepts, and Video Generation Models (VGMs), which learn from temporally evolving visual worlds. However, it still remains unclear which pre-training scheme provides a better representation substrate for spatial intelligence. In this paper, we present the first systematic frozen-feature probing study of VLMs and VGMs across three representative axes of spatial intelligence: semantic tagging, instance grouping, and 3D geometry prediction. Using the lightweight probe, our framework enables a controlled comparison of what information is already encoded in frozen representations from two model families. Experimental results reveal a clear complementarity: VLMs are stronger at semantic tagging and instance grouping, while VGMs provide more accessible signals for dense geometry and camera motion. Moreover, a naive fusion of the two already yields a representation that excels at both geometry and semantics, suggesting a promising direction for building stronger spatial-intelligence backbones by effectively integrating features from both model families. Our code is available at \href{https://github.com/om-ai-lab/Probing-VLM-VGM}{https://github.com/om-ai-lab/Probing-VLM-VGM}.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis