NTH

Scalable Visual Pretraining for Language Intelligence

AuthorsYiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen

July 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that training on the visual form of documents, not just extracted text, can improve foundation models by preserving information like layouts, equations, and figures.

Key results

3.22
GPQA Diamond improvement

Maximum VP gain over matched text pretraining across four backbones

20B
Visual token budget

Approximate VP tokens from the scientific-PDF corpus

80B
Text token budget

Approximate parsed-text tokens for the matched baseline

2.88x
AIME-25 normalized gain

VP gain relative to the text-pretraining improvement over the base model

0.907
Paired visual-text cosine similarity

Post-VP alignment score, rising from 0.631

99.0%
Image-to-text retrieval R@1

VP retrieval performance, rising from 64.0%

What the paper found

A team led by Shanghai Artificial Intelligence Laboratory, with researchers from the University of Science and Technology of China, Zhejiang University, and Shanghai Jiao Tong University, presents Visual Pretraining, or VP, as an alternative to converting scientific documents into plain text. Instead of OCR, VP renders PDF pages as images, removes blank patches, preserves raster-order layout, and trains a shared language model to predict the next frozen visual latent with an InfoNCE objective, jointly with text next-token prediction. Across Qwen3.5, Qwen3, Meta’s Llama 3.2 Vision, and Llama 3.1, VP improves scientific reasoning over matched text pretraining: GPQA Diamond gains up to 3.22 points, while the same corpus requires only 20B visual tokens versus 80B parsed text tokens. Normalized downstream gains reach 2.88x on AIME-25, 2.02x on GPQA, and 1.27x on MMLU-Pro. VP also strengthens cross-modal representations without image-text pairing, raising paired visual-text cosine similarity from 0.631 to 0.907, and improves image-to-text retrieval R@1 from 64.0% to 99.0%. The largest benefits occur on pages dense with equations, figures, tables, and complex layouts, supporting the paper’s central claim that textualization discards reasoning-relevant structure. A decoder-free visual-latent objective delivers these gains more efficiently than pixel reconstruction, positioning raw scientific pages as compact, unlabeled supervision for scalable language and multimodal intelligence.

Original abstract

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis