NTH

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

AuthorsTongyao Zhu, Wei Hern Lim, Min-Yen Kan

September 14, 2026 2 min read
Watch on YouTube
The one-line take

This study shows that coding models often fix bugs by changing too much, and that targeted training can make their edits smaller, safer, and easier to review.

Key results

400
BigCodeBench repair problems

Controlled evaluation problems used to inject localized AST-level corruptions.

0.195
Generic excess Levenshtein distance

Average frontier-model excess edit distance before preservation prompting.

0.131
Prompted excess Levenshtein distance

Average excess edit distance after adding an explicit preservation instruction.

26.6%
Added cognitive complexity reduction

Reduction produced by preservation prompting.

0.782
RL out-of-domain Pass@1

Qwen3-4B-Instruct-2507 performance after reinforcement learning on held-out corruption types.

0.050
RL out-of-domain excess Levenshtein distance

Excess edit distance achieved alongside reinforcement-learning repair performance.

What the paper found

This paper argues that code-repair quality has a second dimension beyond passing tests: edit fidelity, meaning whether a model changes only what the bug requires. The authors build a controlled benchmark from 400 BigCodeBench Python problems by injecting localized AST-level corruptions, so every task has a known minimal repair. Across frontier systems including OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.7, Google’s Gemini 3.1 Pro, and DeepSeek V3.2, functionally correct patches often contain unnecessary validation, data-flow rewrites, contract changes, and added control-flow complexity. A one-line off-by-one error, for example, can be fixed minimally while GPT-5.4 adds 60 lines that tests do not require. Using normalized token-level Levenshtein distance and added cognitive complexity, a preservation instruction reduces average excess edit distance from 0.195 to 0.131, cuts added cognitive complexity by 26.6%, and increases Pass@1 by 2.3 percentage points. Larger models and reasoning modes do not reliably produce smaller patches. The paper then tests whether minimal editing can be learned with Qwen3-4B-Instruct-2507: supervised fine-tuning overfits familiar corruption patterns, while reinforcement learning using execution feedback plus edit distance reaches 0.782 out-of-domain Pass@1 with only 0.050 excess edit distance and preserves broader coding performance on LiveCodeBench v6. The central conclusion is that over-editing is a measurable, steerable maintenance failure: models should be evaluated not only on whether repairs work, but also on whether they preserve the surrounding implementation.

Original abstract

Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis