NTH

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

AuthorsMd Shohel Arman, Igor Molybog

AffiliationsDaffodil International University, Dhaka, Bangladesh · HawAII, University of Hawai‘i at Mānoa, Honolulu, HI, USA

October 4, 2026 2 min read
Watch on YouTube
The one-line take

Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.

Key results

0.88
Roundtrip density-fidelity correlation

Pearson correlation between description density and regeneration fidelity across 11 SWE-bench Verified fixtures.

0.5
Held-out baseline fidelity

Mean held-out fidelity before prompt optimization.

1.0
Held-out optimized fidelity

Mean held-out fidelity after optimizing the description-writing prompt.

0.08
Source-withheld issue-only score

Mean fraction of issue-resolution tests passed without documentation.

0.71
Source-withheld optimized-documentation score

Mean fraction of issue-resolution tests passed with optimized documentation.

33
SWE-ContextBench Lite issue-only resolution

Tasks resolved with issue-only context, versus 29 with compact descriptions and 30 with retrieved context.

What the paper found

This paper introduces a roundtrip benchmark for coding-agent documentation: a model describes a source file, another model regenerates it without seeing the code, and the regenerated implementation is scored against the original unit tests. Using 11 single-file fixtures filtered from SWE-bench Verified, the study finds that completeness—not verbosity—determines fidelity: descriptions containing the same facts perform equally well at 38 words and 664 words, while fidelity correlates with description density at 0.88. A hill-climbing prompt optimizer using Gemini 2.5 Pro as the proposer reaches 1.0 regeneration fidelity and improves held-out fidelity from 0.5 to 1.0, by explicitly requiring exact imports, literals, signatures, defaults, exceptions, and return paths. The downstream result is sharply conditional. When the source file is withheld, optimized documentation raises mean issue-resolution test-pass fraction from 0.08 to 0.71 across the fixtures, showing that compact specifications can substitute for unavailable code. When source is present, however, documentation is redundant or distracting: on SWE-ContextBench Lite with Gemini 3.8 Flash, issue-only resolution reaches 33 tasks, compared with 29 using compact descriptions and 30 using retrieved context. This null persists across Qwen 3.6, Gemini 3.1 Pro, Gemini 3.8 Flash, and ten repositories. The comparison with LangChain’s OpenWiki reinforces the distinction: human-oriented reference documentation records observable behavior, whereas reconstruction-oriented descriptions must preserve hidden contracts and internal details. The central conclusion is a boundary condition: documentation helps when code cannot fit in context, but usually does not improve agents that can already inspect the source.

Original abstract

We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.

Read the original paper

More in Code Generation

Browse all 43 papers →
01Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis
03Code Generation

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo

CodeMidas turns existing codebases into scalable, automatically verified RL environments that train coding agents to perform better across diverse software tasks.

Read analysis