PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
AuthorsRim Assouel, Amir Bar, Michal Drozdzal, Adriana Romero-Soriano
Resources
This paper shows that adding simple procedurally generated geometric tasks can significantly improve multimodal language models’ ability to understand spatial and fine-grained visual details.
Key results
Instruction tuning LLaVA-v1.5-style MLLMs with PGT improved the What’s Up benchmark by up to 20.0 points.
PGT improved CV-Bench-2D by up to 13.3 points in the instruction-tuning setting.
State-of-the-art MLLMs were finetuned with 5k procedurally generated samples rendered on neutral gray backgrounds.
The ablation on scaling real semantic training data compares the smaller LLaVA-Instruct-1.5 regime to Cambrian-7M.
What the paper found
PGT, short for Procedurally Generated Tasks, is a data-centric method from Mila, Universite de Montreal, and Meta FAIR that improves visual grounding in multimodal large language models by overlaying unambiguous geometric primitives onto training images and turning them into dense, verifiable supervision signals. The paper targets three failure modes in MLLMs: spatial relations, abstract counting, and relative distance or depth reasoning. Instead of adding human annotations or changing model architecture, PGT procedurally generates questions about colored boxes, circles, and labeled points directly on existing datasets such as LLaVA-Instruct-1.5 and Cambrian-7M, preserving the original semantic content while forcing the model to attend to image geometry rather than language priors. Across four backbones in instruction tuning—Vicuna-1.5-7B, Llama-3-8B, Qwen-2.5-7B, and Qwen-2.5-14B—PGT raises What’s Up by up to 20.0 points and CV-Bench-2D by up to 13.3 points, with consistent gains on CV-Bench-3D, BLINK, and TallyQA and little or no regression on general benchmarks like GQA, POPE, and MMSTAR. When fine-tuning state-of-the-art models such as Qwen-2.5-VL and InternVL3 with only 5k PGT samples, the method still improves spatial and depth benchmarks, sometimes outperforming specialized real-data mixes and RL-based baselines like ViGoRL and SpatialLadder. Ablations show that removing the relative-distance task sharply degrades 3D/depth performance, indicating that 2D geometric comparison may be the core transferable primitive behind emergent depth understanding.
Original abstract
Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Generated Tasks (PGT), a simple data-driven framework that serves a dual purpose: inducing fine-grained visual understanding and acting as a low-cost diagnostic tool to identify the source of perception failures. By overlaying unambiguous geometric primitives on images, PGT generate additional dense supervision that disentangles visual grounding capability from semantic priors. Extensive experiments on relational, quantitative, and 3D/depth understanding benchmarks show that PGT yields remarkable gains across diverse architectures. Instruction tuning MLLMs on LLaVA-v1.5-Instruct augmented with PGT data results in improvements of up to +20% on the What'sUp benchmark and +13.3% on CV-Bench-2D, while maintaining general perception capabilities. Moreover, finetuning state-of-the-art MLLMs on PGT data leads to boosts of up to +5.5% on What'sUp and +8.3% on CV-Bench-2D. These findings demonstrate that PGT effectively address the bottleneck of fine-grained perception, revealing that many spatial reasoning deficits stem from inadequate supervision signals rather than inherent architectural or resolution limitations.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.