NTH
AI research

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models

AuthorsYujun Tong, Dongliang Chang, Zijin Yin, Xintong Liu, Yuanchen Fang, Zhanyu Ma

May 19, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that making a multimodal model imagine visual details first can sometimes help it understand images better, revealing a surprising generation-to-understanding loop.

Key results

1,595 samples
VisThink-Bench size

A curated evaluation suite spanning 34 subtasks from 12 benchmarks, used to test G→U on visually grounded reasoning.

34 subtasks
VisThink-Bench coverage

The benchmark organizes tasks into perceptual, logical reasoning, and spatial reasoning categories.

12 benchmarks
Benchmarks in VisThink-Bench

The suite is drawn from 12 existing benchmarks to map specific visual edits to comprehension gains.

1.2%
MMStar gain

BAGEL+G→U improves performance on MMStar relative to the vanilla BAGEL baseline.

4.2%
HallusionBench gain

BAGEL+G→U improves performance on HallusionBench relative to the vanilla BAGEL baseline.

R2 = 0.27, p < 0.01
Correlation with editing quality

Better generative fidelity, measured by semantic consistency and perceptual quality, correlates with larger downstream understanding gains.

What the paper found

Reversing the Flow proposes Generation-to-Understanding, or G→U, a zero-shot mechanism in which a unified multimodal model uses its own image generation as an internal reasoning step before answering a visual question. Built and tested on BAGEL-7B, the method first prompts the model to create a “visual thought” through controlled edits such as deblurring, denoising, outpainting, zoom-in, object removal, or novel-view synthesis, then feeds the generated image back into the understanding pathway alongside the original image and question. On VisThink-Bench, a curated suite of 1,595 samples spanning 34 subtasks from 12 benchmarks, this reversed information flow consistently improves multimodal reasoning, with gains above 10 percent on spatial, perceptual, and logic-heavy tasks such as 3D height estimation, illusion reasoning, and color recognition. Across standard benchmarks, BAGEL+G→U improves MMStar by 1.2 percent, HallusionBench by 4.2 percent, MMBench by 1.8 percent, and R-Bench by 1.6 percent, showing stronger robustness without retraining or extra parameters. The paper also quantifies the mechanism: better generative fidelity correlates with larger downstream gains, with semantic consistency and perceptual quality explaining improvement at R2 = 0.27, p < 0.01. However, the gains are bounded by the model’s synthesis quality; text-heavy and symbolic tasks benefit less, and autonomous prompt writing often produces plausible but poorly targeted edits. The central contribution is conceptual as much as empirical: the authors show that in large multimodal models, generation can function as a perceptual hypothesis generator, but only when the imagined image remains grounded, faithful, and task-aligned.

Original abstract

The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BAGEL, BLIP3o achieves remarkable progress; In practice, however, this unification remains one-directional: understanding routinely guides generation, yet how and why generation can support understanding is rarely investigated. We revisit this asymmetry and propose Generation-to-Understanding (G2U) synergy, where visual generation becomes an explicit intermediate reasoning step. Our framework enables a model to perform controlled generative acts, such as detail enhancement, context expansion or structural visualisation, to produce self-generated visual thoughts, which are then fed back into the model to refine perception without retraining or external tools. Through a comprehensive evaluation on twelve benchmarks, this reversed information flow consistently improves multimodal understanding. We show that generative fidelity bounds perceptual gain and that distinct families of edit prompts govern transfer efficiency. We further analyse whether models can decide what to imagine. While they can produce plausible edits, these self-generated visual thoughts lack stable task alignment, revealing that current large multimodal models fall short of true self-reflection. This work exposes a missing mechanism in unified cognition and suggests that imagination is not the end of understanding but its beginning.

Read the original paper