NTH

DataComp-VLM: Improved Open Datasets for Vision-Language Models

AuthorsMatteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian Böther, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy

July 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper creates a large open benchmark for testing how to build better vision-language training data, and shows that mixing the right kinds of data can matter more than filtering.

Key results

160
dataset count

public datasets in the DCVLM pool

6T
pool size

multimodal tokens in the DCVLM corpus

8B
model scale

largest VLM scale in the benchmark

200B
training budget

token budget for the x-large scale

63.6
core suite accuracy

DCVLM-BASELINE score on the 33-task Core set at 8B and 200B tokens

5.4
gain over FineVision

absolute percentage-point improvement on the Core set

What the paper found

DataComp-VLM, from collaborators across Google, Google DeepMind, Stanford University, LAION, and the Tübingen AI Center, reframes vision-language model training as a controlled data-curation problem and shows that mixing strategy dominates filtering. The benchmark, DCVLM, aggregates 160 public datasets into a 6T-token multimodal pool spanning image-caption pairs, interleaved documents, text-only corpora, and multimodal instruction data, then evaluates models from 1B to 8B parameters on up to 52 benchmarks across 9 domains. Across more than 1,000 experiments, common quality filters such as CLIP-score, text-quality classifiers, and multimodal perplexity give little or no gain once the pool is already curated, with the best filtering result only +0.8pp. In contrast, an instruction-heavy recipe of 10% image-caption, 5% multimodal documents, 15% text-only, and 70% instruction-tuning data scales best, and pretraining rankings transfer to supervised fine-tuning with Pearson r=0.99 across 54 SFT runs. The resulting DCVLM-BASELINE trains an 8B model on 200B tokens to 63.6% on the 33-task Core suite, beating FineVision by +5.4pp, while a 4B model trained on 100B tokens already surpasses an 8B FineVision model trained on 200B tokens, implying a 4× compute reduction. The paper’s main takeaway is that, for modern autoregressive VLMs, data composition is a stronger lever than post hoc filtering, especially at larger scales.

Original abstract

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis