To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
AuthorsAmir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia, Gopi Krishnan Rajbahadur, Ahmed E. Hassan
Resources
LLMs can find the code that needs changing but often avoid deleting it, so this paper measures the problem and explores training models to remove code more reliably.
Key results
Mean recall for Claude Opus 4.5 on tasks all five SWE-bench Verified models solved.
Share of passing model patches that retained developer-removed logic behind added control flow.
Deletion-only tasks mined from real Python and JavaScript repository commits.
Highest deletion-compliant success rate, achieved by Claude Opus 4.8.
Percentage-point reduction in incomplete deletion after adding deletion examples to post-training.
Percentage-point improvement on SWE-bench Verified after deletion-focused post-training.
What the paper found
This paper from Queen’s University identifies deletion avoidance in large language models: a systematic tendency to preserve code that a human developer removes, often by wrapping it in a guard or fallback, a pattern the authors call Guard-and-Go. Analyzing OpenAI’s GPT-5, Anthropic’s Claude Opus 4.5, DeepSeek-related Kimi-K2, GLM-4.6, and Salesforce SAGE on SWE-bench Verified, the best mean deletion recall on tasks all five models solved was only 71.7%, while Guard-and-Go appeared in 29.0% of passing patches; these patches were larger than the developer patch in 61.1% of cases. To isolate deletion from localization and replacement work, the authors introduce CanItDelete, a 200-task benchmark of deletion-only edits mined from Python and JavaScript repositories, where success ranged from 18.0% to 79.0%. Supplying exact deletion spans improved performance substantially, but also exposed over-deletion and boundary-control failures. A deletion-sensitive retrofit of 34 SWE-bench Verified tasks reduced resolution from 63.2% to 41.9%, showing that ordinary tests often accept patches that retain code intended for removal. Finally, adding 12,821 deletion examples to a 7B model’s post-training mixture reduced incomplete deletion by 13.9 percentage points and improved SWE-bench Verified by 5.3 percentage points, suggesting deletion avoidance is an undertrained, partially correctable behavior rather than an intrinsic limitation.
Original abstract
Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. We identify one concrete source: deletion avoidance, the systematic tendency to retain code that an intended edit requires removing. Across the five leading models on the official SWE-bench Verified leaderboard, deletion recall against the developer patch reaches at most 71.7% even on tasks all five solve, and models reach the right file for over 92% of required deletions but cut the exact line in under 52% of cases. Instead, 29.0% of passing patches wrap the targeted code in a guard or fallback, a pattern we call Guard-and-Go. Such patches pass because the original tests rarely check removal: when we retrofit 34 Verified tasks with tests that fail if the targeted code remains, four frontier models spanning closed and open weights fall from 63.2% to 41.9%. Because real repairs mix removal with addition, we curate CanItDelete, a benchmark of 200 tasks mined from real commits whose entire required edit is deletion. Even with the addition work gone, the best model still fails one task in five, and smaller open models fall to 18.0%. We then ablate GPT-5.6 Sol under four cumulative prompts; success moves little until we supply the exact lines, which nearly eliminate incomplete deletion yet raise success only to 80.5% because the model then deletes beyond the spans or adds code instead. Finally, through a pilot study we show one potential fix: teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.