Scalable Visual Pretraining for Language Intelligence
AuthorsYiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
Resources
This paper argues that training on the visual form of documents, not just extracted text, can improve foundation models by preserving information like layouts, equations, and figures.
Key results
Maximum VP gain over matched text pretraining across four backbones
Approximate VP tokens from the scientific-PDF corpus
Approximate parsed-text tokens for the matched baseline
VP gain relative to the text-pretraining improvement over the base model
Post-VP alignment score, rising from 0.631
VP retrieval performance, rising from 64.0%
What the paper found
A team led by Shanghai Artificial Intelligence Laboratory, with researchers from the University of Science and Technology of China, Zhejiang University, and Shanghai Jiao Tong University, presents Visual Pretraining, or VP, as an alternative to converting scientific documents into plain text. Instead of OCR, VP renders PDF pages as images, removes blank patches, preserves raster-order layout, and trains a shared language model to predict the next frozen visual latent with an InfoNCE objective, jointly with text next-token prediction. Across Qwen3.5, Qwen3, Meta’s Llama 3.2 Vision, and Llama 3.1, VP improves scientific reasoning over matched text pretraining: GPQA Diamond gains up to 3.22 points, while the same corpus requires only 20B visual tokens versus 80B parsed text tokens. Normalized downstream gains reach 2.88x on AIME-25, 2.02x on GPQA, and 1.27x on MMLU-Pro. VP also strengthens cross-modal representations without image-text pairing, raising paired visual-text cosine similarity from 0.631 to 0.907, and improves image-to-text retrieval R@1 from 64.0% to 99.0%. The largest benefits occur on pages dense with equations, figures, tables, and complex layouts, supporting the paper’s central claim that textualization discards reasoning-relevant structure. A decoder-free visual-latent objective delivers these gains more efficiently than pixel reconstruction, positioning raw scientific pages as compact, unlabeled supervision for scalable language and multimodal intelligence.
Original abstract
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.