Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study
AuthorsPriyansh Trivedi, Olivier Schmitt
Resources
This study shows that while cleaner code may not help coding agents finish more tasks, it can make them work more efficiently with fewer tokens and fewer back-and-forth file visits.
Key results
hidden-test coding tasks across the six pairs
Claude Sonnet 4.6 runs across both sides and 10 repetitions per task
dataset-level cleaner-minus-messier footprint change
dataset-level cleaner-minus-messier footprint change
dataset-level cleaner-minus-messier navigation change
What the paper found
This paper from SonarSource asks whether code cleanliness affects autonomous coding agents when the repository and task are held constant. The authors build a controlled minimal-pair benchmark on the Harbor framework with six behaviorally equivalent repository pairs, three public and three private, and 33 hidden-test tasks routed through cleanliness hotspots or cross-module seams. Cleanliness is proxied by SonarQube issue density and cognitive complexity. Using Claude Code with Claude Sonnet 4.6, they run 660 trials and find that pass rate is essentially unchanged between cleaner and messier code, but operational footprint is not: cleaner repositories reduce input tokens by 7.1%, output tokens by 8.5%, reasoning characters by 11.1%, conversation messages by 7.0%, and file revisitations by 33.8%. The strongest behavioral shift is revisitation, which drops about a third on cleaner code, suggesting the agent commits earlier and reopens edited files less often. Track-level analysis shows multi-module tasks account for most token savings, while cognitive-hotspot tasks mainly change how work is distributed across files. A comment-normalization ablation indicates that suppression markers are not the main driver of the effect, and the cleaner-side advantage persists even after equalizing comments. The central takeaway is that maintainability principles still matter for AI-driven development: cleaner code does not make Claude Code more accurate here, but it makes the agent cheaper and less navigationally uncertain.
Original abstract
As autonomous coding agents see rapid adoption, their evaluation has primarily focused on task completion rates holding the target codebase fixed. This leaves a critical question unanswered: does the structural and stylistic quality, or ``cleanliness'' of the underlying code affect an agent's ability to navigate and modify it? To isolate the effect of code cleanliness from agent capability, we introduce an evaluation protocol built around minimal pairs: repositories that match on architecture, dependencies, and external behaviour, but differ on static-analysis rule violations and cognitive complexity. The pairs are constructed in both directions, by agent pipelines that either degrade a clean repository or clean a messy one. We author 33 tasks across six such pairs, evaluated through hidden tests at the application's public surface. Across 660 trials with Claude Code, code cleanliness does not change the agent's pass rate. However, it substantially alters the agent's operational footprint: agents working on cleaner code use 7 to 8% fewer tokens and reduce file revisitations by 34%. Our findings suggest that traditional maintainability principles remain highly relevant in the era of AI-driven development, shaping the computational cost and navigational efficiency of coding agents. Code cleanliness joins model choice, harness, and prompting as a factor that materially affects agent behaviours.
Read the original paper