NTH

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

AuthorsYichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min

July 12, 2026 3 min read
Watch on YouTube
The one-line take

PixelRAG shows that for web-based retrieval-augmented generation, using screenshots instead of extracted text can improve both answer quality and efficiency across several QA and agent tasks.

Key results

7,134,778
Wikipedia corpus size

articles rendered into the screenshot datastore

30M
Wikipedia tile count

screenshot tiles indexed for Wikipedia

2
Wikipedia offline build time

days to complete the full pipeline on one machine

78.8
SimpleQA accuracy

P IXEL RAG end-to-end accuracy on SimpleQA

83.8
SimpleQA Recall@3

P IXEL RAG evidence recall on SimpleQA

3x
Token cost reduction

maximum reduction from image compression

What the paper found

P IXEL RAG, from UC Berkeley, Princeton, EPFL, Databricks, and Renmin University, argues that web retrieval should operate in pixels rather than parsed text. The system renders Wikipedia and news pages into screenshot tiles, indexes them with Qwen3-VL-Embedding-2B, and feeds retrieved images directly to a vision-language reader, eliminating HTML-to-text parsing. On the full 7,134,778-article Wikipedia corpus, this produces about 30M tiles and an offline pipeline that completes in about 2 days on one machine. Across six benchmarks, including Natural Questions, NQ-Tables, SimpleQA, MMSearch, Encyclopedic VQA, and LiveVQA, pixel retrieval beats both no-retrieval and text-based RAG; on SimpleQA it reaches 78.8 accuracy and 83.8 Recall@3, compared with 71.6 and 77.4 for Trafilatura. The gains are strongest on structured content, where tables and infoboxes are preserved visually instead of being flattened, and on EVQA it improves accuracy by up to 18.1% over text baselines. The paper also introduces synthetic contrastive training with LLM-generated queries, dynamic hard-negative mining, and LoRA fine-tuning of both the ViT and language backbone. Beyond accuracy, pixel-space RAG enables compression: at 2× downsampling, a fine-tuned reader nearly matches native resolution while cutting visual tokens roughly in half, and the authors report up to 3× token cost reduction overall. The broader claim is that as VLMs like Qwen3.5, Qwen3.6, and Llama 4 improve, screenshot-native retrieval becomes both more accurate and more efficient than text-centric pipelines.

Original abstract

Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting. We introduce PixelRAG, a new retrieval-augmented method that represents websites in their native visual form and performs retrieval and reading entirely in pixel space, enabling an end-to-end architecture that eliminates text abstraction. PixelRAG is, to our knowledge, the first pipeline to operate over a full Wikipedia corpus in this form, scaling to a datastore of 30 million screenshot images with an efficient visual retrieval index. Built on an existing visual embedding model (i.e., Qwen3-VL-Embedding), PixelRAG further fine-tunes this model on screenshot data with carefully curated contrastive training data. Retrieved screenshots are then fed directly as pixel inputs to a VLM, without intermediate text conversion. PixelRAG consistently outperforms both no-retrieval and text-based RAG baselines, most surprisingly on widely studied text-centric tasks such as NQ and SimpleQA. It also achieves strong gains on multimodal open-domain QA (e.g., MMSearch), benchmarks over noisy news corpora (e.g., LiveVQA), and agentic benchmarks (e.g., MoNaCo), improving accuracy by up to 18.1% over text-based baselines. Finally, pixel representations enable a new efficiency lever for RAG through image compression, achieving up to 3x token cost reduction at lower resolutions while maintaining accuracy. Our results challenge the necessity of text representations in web retrieval, suggesting that web RAG can operate directly in the web's native visual form while improving both performance and efficiency.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis