NTH

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

AuthorsXin Gao, Cheng Yang, Chufan Shi, Taylor Berg-Kirkpatrick

July 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks whether fixing a model’s text knowledge also fixes its image generation, and finds that the answer is usually no unless the edit is made more reasoning-aware.

Key results

2971
edit subjects

U NI KE benchmark size in edit subjects

5535
evaluation instances

U NI KE total multi-stage evaluation instances

92%
text-side efficacy

approximate best text-only edit success reported

18.5%
best direct VQA accuracy

best overall image-generation verification under direct generation

18.6%
max gain

largest VQA improvement from reasoning-augmented editing

34.82%
projection retention

Ovis-U1 top-1536 singular directions capture this share of edit perturbation energy

What the paper found

This paper asks whether text-only knowledge edits in unified multimodal models actually transfer to image generation, and the answer is mostly no. The authors introduce U NI KE, the first benchmark for cross-modality knowledge editing in UMMs, built with 2,971 edit subjects and 5,535 evaluation instances spanning attribute and relation edits, then test three representative models—Ovis-U1, BLIP3o-4B, and OmniGen2—using parameter editors MEMIT, PMET, and AlphaEdit. The central finding is a sharp modality gap: text-side edit efficacy reaches about 92% in some settings, but the best direct image-generation VQA accuracy is only 18.5%, showing that successful factual rewriting in text does not reliably change generated images. A reasoning-augmented protocol that first forces the model to verbalize the edited fact improves every model-editor pair, with gains as large as 18.6 percentage points, but it still leaves a large gap. Mechanistically, the paper traces the failure to a conditioning-pathway bottleneck: for Ovis-U1, edit perturbations are heavily attenuated by a frozen Linear(4096, 1536) projection, retaining only about 34.82% of perturbation energy in the top 1,536 directions, whereas reasoning expands the effective conditioning shift by 8.56× and yields the largest visual gains. Overall, the study shows that cross-modal knowledge editing requires more than changing text outputs; the edit must survive the language-to-vision interface and align with the image generator’s conditioning pathway.

Original abstract

Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world applications, effectively updating internal knowledge becomes critical. While knowledge editing has matured for text-only models, it remains unclear whether edits that successfully modify textual outputs also transfer to image generation in UMMs. To study this question, we introduce UniKE, the first benchmark for cross-modality knowledge editing in UMMs, comprising 2,971 edit subjects spanning attribute and relation edits. Using VQA-based visual verification, we reveal a striking modality gap: text-side efficacy can reach approximately 92%, whereas the best overall VQA accuracy under direct image generation is only 18.5%. We further propose Reasoning-augmented Parameter Editing, which explicitly activates edited knowledge before generation and improves overall VQA accuracy for all evaluated model-editor pairs, with gains up to 18.6 percentage points. Mechanistic analysis shows that this gap is associated with partial alignment between edited textual representations and the conditioning pathways for visual generation, where edits sufficient for text outputs may remain too weak or misaligned to steer image synthesis. These findings show that textual knowledge edits do not guarantee reliable cross-modality transfer and motivate modality-aware editing methods. Our code and data are available at https://github.com/gxx27/UniKE.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis