OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
AuthorsJiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
AffiliationsUniversity of California, Berkeley · Duke University · Impossible Research · Carnegie Mellon University · University of Washington · Elorian
Resources
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.
Key results
OmniTaskonomy includes 19 image-to-image generation tasks.
The taxonomy spans 25 image-to-text visual understanding capabilities.
The transfer benchmark retains 9444 labeled evaluation examples.
Jigsaw generation improves 2D ordering accuracy by 7.9 percentage points.
Mean gradient alignment correlates with mean transfer gain at r = 0.795.
What the paper found
OmniTaskonomy investigates when image-to-image generation supervision improves image-to-text visual understanding in unified multimodal models. Using BAGEL-7B-MoT and paired Jigsaw and Zoom-In tasks from VisGym, the study finds that sequential I2I pretraining followed by I2T fine-tuning transfers more reliably than mixing objectives from the start, with benefits scaling as I2I data grows. With 100k I2I examples and only 1k I2T examples, Zoom-In performance approaches training with 10k I2T examples alone. The paper introduces OmniTaskonomy, a capability taxonomy covering 19 I2I tasks and 25 understanding capabilities, evaluated across 9444 examples drawn from BLINK, CV-Bench, MMStar, MMT-Bench, MMVP, VStarBench, and RealWorldQA. Transfer is selective: Jigsaw improves 2D ordering by 7.9 percentage points, while Euclidean depth and surface-normal prediction improve metric 3D relation by 3.8 and 3.4 percentage points; 2.5D segmentation also improves category recognition by 1.2 percentage points. To explain these patterns, the authors measure PCA-projected gradient cosine alignment in BAGEL’s shared understanding pathway, finding the strongest alignment in early pre-attention RMSNorm layers. Alignment correlates with transfer at r = 0.795 across capabilities and r = 0.529 across 133 source-target pairs. The taxonomy was constructed and validated with Gemini, GPT-5.6, and Claude Opus 4.8, suggesting that generation objectives should be selected according to shared visual capabilities and optimization compatibility rather than treated as universally beneficial.
Original abstract
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
The study shows that video models may retain the right physical understanding even after generating the wrong motion—and that targeted internal writes can bring that knowledge back.