LLM Agents Can See Code Repositories
AuthorsDongjian Ma, Silin Chen, Yufei Yang, Yulin Shi, Yanfu yan, Xiaodong Gu
Resources
This paper shows that coding agents work better when they can 'see' a repository’s structure, and that combining visual graphs with text can cut token use while preserving or improving bug-fixing performance.
Key results
largest accuracy drop in vision-only mode across evaluated models
largest token-cost inflation in vision-only mode
best reduction from multimodal repository context on GPT-5-mini
largest cost reduction from multimodal integration on GPT-5.1
What the paper found
LLM Agents Can See Code Repositories is a systematic study from Shanghai Jiao Tong University and Zhejiang University showing that multimodal large language models can use visual repository structure as a useful complement to text for software issue resolution. The authors build SeeRepo, which converts AST-derived repository graphs into Graphviz images over four relations—contains, imports, invokes, and inherits—and evaluate GPT-5-mini, GPT-5.1, Doubao-Seed-2.0-Lite, and Kimi K2.5 on SWE-bench Verified, SWE-Rebench Leaderboard, and SWE-QA. The key result is a sharp boundary between vision-only and multimodal access: replacing text with graph images drops Pass@1 by as much as 34.1 points and raises cost by up to 268%, while adding visual structure alongside text cuts input tokens by up to 26% and cost by up to 46% without sacrificing accuracy. On SWE-bench Verified, GPT-5-mini reaches 55.4% Pass@1 with SeeRepo versus 55.0% text-only, GPT-5.1 drops slightly to 48.8% but halves cost to 0.0975, Kimi K2.5 improves to 70.6%, and Doubao-Seed-2.0-Lite rises to 52.0%. The best visual design is a graph layout with adaptive hop depth, and visualization is most effective during fault localization, not repair or validation, where late-stage structure adds noise rather than signal.
Original abstract
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Yet most agents consume repositories almost entirely as text, which differs from how human developers use visual structure such as folder hierarchies and dependency relationships to orient themselves in large codebases. With multimodal large language models (MLLMs), it is an open question whether agents can effectively benefit from visual representations of repositories. This paper presents the first systematic empirical study of visual repository representations for LLM-based agents on repository-level issue resolution. We evaluate four recent multimodal models. Our results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries. In contrast, integrating visual graphs of repository structure as a supplementary modality alongside standard text interfaces helps agents understand structure more efficiently: input token consumption decreases by up to 26% while issue-resolution accuracy is maintained or improved. Visualization is most useful during fault localization and when the agent autonomously controls exploration depth. These findings point to a practical hybrid text-and-vision design for next-generation coding agents.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.