When Models Edit Too Much: On the Fidelity of Minimal Code Edits
AuthorsTongyao Zhu, Wei Hern Lim, Min-Yen Kan
Resources
This study shows that coding models often fix bugs by changing too much, and that targeted training can make their edits smaller, safer, and easier to review.
Key results
Controlled evaluation problems used to inject localized AST-level corruptions.
Average frontier-model excess edit distance before preservation prompting.
Average excess edit distance after adding an explicit preservation instruction.
Reduction produced by preservation prompting.
Qwen3-4B-Instruct-2507 performance after reinforcement learning on held-out corruption types.
Excess edit distance achieved alongside reinforcement-learning repair performance.
What the paper found
This paper argues that code-repair quality has a second dimension beyond passing tests: edit fidelity, meaning whether a model changes only what the bug requires. The authors build a controlled benchmark from 400 BigCodeBench Python problems by injecting localized AST-level corruptions, so every task has a known minimal repair. Across frontier systems including OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.7, Google’s Gemini 3.1 Pro, and DeepSeek V3.2, functionally correct patches often contain unnecessary validation, data-flow rewrites, contract changes, and added control-flow complexity. A one-line off-by-one error, for example, can be fixed minimally while GPT-5.4 adds 60 lines that tests do not require. Using normalized token-level Levenshtein distance and added cognitive complexity, a preservation instruction reduces average excess edit distance from 0.195 to 0.131, cuts added cognitive complexity by 26.6%, and increases Pass@1 by 2.3 percentage points. Larger models and reasoning modes do not reliably produce smaller patches. The paper then tests whether minimal editing can be learned with Qwen3-4B-Instruct-2507: supervised fine-tuning overfits familiar corruption patterns, while reinforcement learning using execution feedback plus edit distance reaches 0.782 out-of-domain Pass@1 with only 0.050 excess edit distance and preserves broader coding performance on LiveCodeBench v6. The central conclusion is that over-editing is a measurable, steerable maintenance failure: models should be evaluated not only on whether repairs work, but also on whether they preserve the surrounding implementation.
Original abstract
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.