Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
AuthorsHaozhan Shen, Tiancheng Zhao, Kangjia Zhao, Jianwei Yin
The paper asks whether vision-language models or video generation models better support spatial intelligence, and finds they encode complementary strengths that can be combined for stronger representations.
Key results
Average semantic tagging mAP for frozen VLMs on ScanNet20, compared against 69.89 for VGMs.
Average APmid for frozen VLMs on ScanNet20, compared against 58.63 for VGMs.
Average multi-view instance grouping T-mIoU for frozen VLMs, compared against 13.24 for VGMs.
Average instance grouping T-SR for frozen VLMs, compared against 4.35 for VGMs.
Average 3D geometry point-map error for VGMs on DL3DV, lower than the VLM average of 0.223.
Feature-level fusion of WAN2.1-T2V-14B and Qwen3-VL-8B achieved this camera pose AUC@30 on the 3D geometry task.
What the paper found
This paper, from Zhejiang University and Om AI Research, asks which pretraining paradigm better supports spatial intelligence: Vision-Language Models such as OpenAI-style language-aligned backbones Qwen3-VL, Qwen2.5-VL, and InternVL3, or Video Generation Models such as WAN, CogVideoX, OpenSora-2.0, and Aether. The authors freeze each model and probe intermediate features with identical lightweight heads on three spatial tasks: semantic tagging on ScanNet20, multi-view instance grouping on ScanNet masks, and 3D geometry prediction on DL3DV using VGGT-generated depth, point maps, and camera poses. The results show a clean division of strengths. VLMs are dramatically stronger at semantics, raising ScanNet20 mAP from 69.89 to 92.08 on average and APmid from 58.63 to 87.28, while also outperforming VGMs on instance grouping, with average T-mIoU improving from 13.24 to 22.66 and T-SR from 4.35 to 11.23. In contrast, VGMs encode geometry more directly, cutting point-map error from 0.223 to 0.152, lowering AbsRel from 0.113 to 0.072, and raising camera AUC@30 from 0.330 to 0.527. A naive feature concatenation of WAN2.1-T2V-14B and Qwen3-VL-8B already combines these advantages, reaching 92.30 mAP, 23.70 T-mIoU, 0.042 AbsRel, and 0.615 AUC@30, which suggests the two representation families are complementary rather than competing.
Original abstract
Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation backbones: Vision-Language Models (VLMs), which use language supervision to align visual observations with semantic concepts, and Video Generation Models (VGMs), which learn from temporally evolving visual worlds. However, it still remains unclear which pre-training scheme provides a better representation substrate for spatial intelligence. In this paper, we present the first systematic frozen-feature probing study of VLMs and VGMs across three representative axes of spatial intelligence: semantic tagging, instance grouping, and 3D geometry prediction. Using the lightweight probe, our framework enables a controlled comparison of what information is already encoded in frozen representations from two model families. Experimental results reveal a clear complementarity: VLMs are stronger at semantic tagging and instance grouping, while VGMs provide more accessible signals for dense geometry and camera motion. Moreover, a naive fusion of the two already yields a representation that excels at both geometry and semantics, suggesting a promising direction for building stronger spatial-intelligence backbones by effectively integrating features from both model families. Our code is available at \href{https://github.com/om-ai-lab/Probing-VLM-VGM}{https://github.com/om-ai-lab/Probing-VLM-VGM}.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.