NTH

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

AuthorsJunlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis

August 9, 2026 2 min read
Watch on YouTube
The one-line take

This study maps how vision and language learn from each other and shows that unifying them early can make multimodal foundation models both stronger and more compute-efficient.

Key results

70%
Language share in optimal mix

The recommended language-understanding-generation mixture is 70/25/5.

5%
Generation share in optimal mix

Visual generation uses only 5% of the optimized pretraining mixture.

13.5B
Scaled MoE model size

The recipe is validated with 13.5B-parameter MoE models.

2T
Scaled training tokens

The large-scale models are trained on 2T tokens.

54.31%
Language accuracy

Accuracy achieved by the asymmetric recipe in the scaled comparison.

43.08%
Visual understanding average

Average visual-understanding score achieved by the asymmetric recipe.

What the paper found

Researchers Junlin Han and colleagues at Meta’s FAIR and Reality Labs, with the University of Oxford, investigate the mechanics of unified multimodal pretraining using Llama-3-like Transformers and Transfusion, trained on DCLM language data, roughly 350M Shutterstock-Image pairs, and a controlled CLEVR benchmark. Their experiments show asymmetric knowledge flow: language broadly boosts visual understanding and generation, visual understanding strongly supports generation, while visual generation provides little backward transfer. In CLEVR, low-level color and shape concepts fail to transfer zero-shot, but structural concepts such as spatial relations, size, and counting can move from understanding into generation. Modality synergy depends on complexity and architecture: simple tasks reinforce each other, whereas complex distributions create capacity competition; shared attention and normalization combined with modality-specific feed-forward networks, as used in MoE-style designs, preserve synergy while isolating computation. Training vision from the beginning is crucial: delayed alignment produces “vision laziness,” with weaker visual activations and reduced attention to image tokens, while sequential curricula underperform simultaneous joint training even with 12.5% replay. A 1T-token mixing search identifies a 70/25/5 language-understanding-generation recipe, and scale tests train 13.5B MoE models on 2T tokens. Relative to a balanced recipe, the asymmetric model reaches 54.31% language accuracy, 43.08% average visual-understanding accuracy, and 0.482 GenEval, despite allocating only 5% of tokens to generation.

Original abstract

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis