Video Generation Models are General-Purpose Vision Learners
AuthorsLetian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
Resources
This paper argues that video generation models can double as powerful general vision learners, enabling a single pretrained model to tackle diverse perception tasks with strong data efficiency.
Key results
Human-centric videos generated for multi-task supervision.
WAN 2.1-based generalist model evaluated across perception tasks.
Average score of the 14B WAN 2.1 model across Sintel, KITTI, and ETH3D.
1.3B model processing 81 frames on one v6e TPU.
1.3B feed-forward model on one v6e TPU.
Comparable performance to leading systems with up to this much less training data.
What the paper found
Researchers at Google DeepMind, the University of Toronto, UCL, Oxford, MIT, and Lund University propose GenCeption, arguing that large-scale text-to-video generation can serve as a general-purpose pretraining objective for computer vision. Built on the open-weights WAN 2.1 diffusion model, GenCeption converts iterative denoising into a single-step, feed-forward predictor: clean video latents enter a DiT at timestep zero, and text prompts select outputs such as depth, surface normals, segmentation, camera pose, and 2D or 3D keypoints. Dense targets share an RGB-space representation and a unified L2 loss, while sparse predictions use lightweight learnable tokens. Training relies mainly on 7,500 synthetic human videos rendered from 800 RenderPeople assets and 200 CMU motion-capture motions. Across benchmarks, the 14B generalist reaches an average depth AbsRel of 0.071, competitive with specialized systems including DepthAnything 3, D4RT, VGGT-Ω, SAM 3, Sapiens, David, Genmo, and Lotus-2, while outperforming V-JEPA and VideoMAE V2 under comparable settings. The model also matches leading depth and geometry systems using 7× to 500× less training data, and scales with model size and dataset volume. Removing WAN 2.1’s 50-step sampling yields 5.92s inference at 13.6 FPS for the 1.3B version on one v6e TPU. Despite training primarily on synthetic human footage, GenCeption transfers to real videos, multiple instances, animals, robots, and other unseen articulated objects, suggesting that video generators encode reusable spatiotemporal and physical-world priors rather than serving only as synthesis engines.
Original abstract
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.