DataComp-VLM: Improved Open Datasets for Vision-Language Models
AuthorsMatteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian Böther, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy
Resources
This paper creates a large open benchmark for testing how to build better vision-language training data, and shows that mixing the right kinds of data can matter more than filtering.
Key results
public datasets in the DCVLM pool
multimodal tokens in the DCVLM corpus
largest VLM scale in the benchmark
token budget for the x-large scale
DCVLM-BASELINE score on the 33-task Core set at 8B and 200B tokens
absolute percentage-point improvement on the Core set
What the paper found
DataComp-VLM, from collaborators across Google, Google DeepMind, Stanford University, LAION, and the Tübingen AI Center, reframes vision-language model training as a controlled data-curation problem and shows that mixing strategy dominates filtering. The benchmark, DCVLM, aggregates 160 public datasets into a 6T-token multimodal pool spanning image-caption pairs, interleaved documents, text-only corpora, and multimodal instruction data, then evaluates models from 1B to 8B parameters on up to 52 benchmarks across 9 domains. Across more than 1,000 experiments, common quality filters such as CLIP-score, text-quality classifiers, and multimodal perplexity give little or no gain once the pool is already curated, with the best filtering result only +0.8pp. In contrast, an instruction-heavy recipe of 10% image-caption, 5% multimodal documents, 15% text-only, and 70% instruction-tuning data scales best, and pretraining rankings transfer to supervised fine-tuning with Pearson r=0.99 across 54 SFT runs. The resulting DCVLM-BASELINE trains an 8B model on 200B tokens to 63.6% on the 33-task Core suite, beating FineVision by +5.4pp, while a 4B model trained on 100B tokens already surpasses an 8B FineVision model trained on 200B tokens, implying a 4× compute reduction. The paper’s main takeaway is that, for modern autoregressive VLMs, data composition is a stronger lever than post hoc filtering, especially at larger scales.
Original abstract
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.