How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring
AuthorsZhihao Lin, Mingyi Zhou, Yizhuo Yang, Li Li
Resources
This paper shows that giving code agents simple structural hints from static analysis can make them more predictable, more stable across runs, and slightly better at finding the right code.
Key results
Anchor-Topo vs. 0.8321 baseline
Anchor-Topo average interaction rounds vs. 35.3 baseline
Anchor-Topo vs. 0.6187 baseline
Tagged runs on Verified raise non-exploratory guidance from 0.15–0.18
What the paper found
This paper studies whether LLM code agents need richer repository structure than plain keyword search, using OpenAI’s Codex with GPT-5.1-codex on SWE-bench Lite and SWE-bench Verified. The authors introduce CodeAnchor, which runs lightweight static analysis offline with PyCG and AST extractors, then injects deterministic anchors as plain-text comments encoding calls, inheritance, imports, containment, data-flow, and configuration usage directly into source files. The key finding is the deterministic anchoring effect: static structure improves navigation discipline and reproducibility more than raw reasoning power. On SWE-bench Lite, bidirectional call/inheritance tags raise Func@5 from 0.8321 to 0.8540, cut rounds from 35.3 to 33.7, and increase Pass@1 on single-run localization from 0.742 to 0.776. On SWE-bench Verified, Func@5 rises from 0.6187 to 0.6308 and rounds fall from 42.4 to 40.9. The gains are scale-sensitive: dense annotations add little over basic topology and often cost more context, while inverse-only links help large hub-heavy repositories. Trajectory analysis shows tags increase link-following rate from 0.15–0.18 to 0.21–0.24 and roughly halve run-to-run variance, confirming that static structure makes code-agent behavior more predictable, not just more accurate.
Original abstract
LLM-based code agents navigate repositories through keyword search but miss the structural relationships, such as call graphs, inheritance hierarchies, and configuration dependencies, that define how software actually works. This makes agent navigation stochastic and difficult to reproduce across runs. We investigate whether lightweight static analysis can provide deterministic anchors for these agents: stable structural facts injected as plain-text comments that constrain probabilistic exploration and make navigation more predictable. Starting from a strong baseline, Codex from OpenAI, we systematically inject varying granularities of structural annotations and measure their effects on localization, trajectory behavior, and run-to-run stability. Our study identifies what we call the deterministic anchoring effect: static structure helps less by making agents "smarter" and more by making their navigation disciplined and reproducible. Three observations support this finding: (1) Anchoring works: lightweight call/inheritance topology improves function-level localization (+2.2pp Func@5) and shortens trajectories (-1.6 interaction rounds); (2) Anchoring is scale-sensitive: the optimal granularity and directionality depend on repository characteristics, where denser semantics show diminishing returns and hub-heavy projects benefit from inverse-only links that expose "who-calls-me" without forward edges; (3) Anchoring stabilizes: tags raise link-following rate from 0.15-0.18 to 0.21-0.24, roughly halve run-to-run variance, and improve single-run reliability (Pass@1 +3.4 pp) on medium-scale repositories, at the cost of roughly 10% more input tokens. These observations suggest practical guidelines: default to lightweight topology on medium projects, prune forward edges in large repositories, and reserve dense tags for implicit-dependency cases.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.