NTH

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

AuthorsBowen Qin, Yi Xie

August 4, 2026 3 min read
Watch on YouTube
The one-line take

This paper benchmarks how well coding agents find the right repository context before they try to write a patch.

Key results

427
Benchmark samples

Total Agent Retrieval Bench samples across positive and no-gold retrieval settings.

25
Repository coverage

Number of repositories represented in the benchmark.

0.2379
Qwen3-4B weighted MRR

Best sample-weighted MRR across the 345 positive samples.

0.7029
Qwen3-8B Recall@20

Best sample-weighted Recall@20 across positive retrieval tasks.

0.2713
RRF hybrid MRR

MRR achieved by fusing Qwen3-8B and RepoMap rankings.

0.3967
RRF seed final File F1

Final File F1 in the 45-sample Codex GPT-5.5 seed-intervention pilot.

What the paper found

Agent Retrieval Bench, from Bowen Qin at the National University of Singapore and Yi Xie at Peking University, isolates a stage that conventional coding-agent evaluations often hide: finding the repository files an agent needs before it edits. Its 427-sample benchmark spans 25 repositories and four workflow-grounded tasks—code2test, comment2context, trace2code, and edit2ripple—plus natural and counterfactual no-gold cases, evaluated on frozen pre-fix commits to prevent leakage. Across 345 positive samples, Qwen3-Embedding-4B achieves the best sample-weighted MRR at 0.2379, Qwen3-Embedding-8B leads Recall@20 at 0.7029, and Aider-style RepoMap leads budgeted context yield at 8k tokens with 0.3788. The task winners differ sharply: RepoMap dominates trace2code because failure traces often expose tests rather than root-cause implementation files, while embeddings are stronger for other workflow signals. Reciprocal-rank fusion of Qwen3-8B and RepoMap improves MRR to 0.2713 and Recall@20 to 0.7331, demonstrating complementary semantic and structural signals. Logged OpenAI GPT-5.4-mini strict-context trajectories still miss every gold file on 35.2% of samples, while Codex CLI GPT-5.4 and GPT-5.5 use roughly 6.2 and 6.5 context events per sample. In a 45-sample Codex GPT-5.5 seed-intervention pilot, RRF retrieval seeds reach final File F1 of 0.3967 versus 0.3437 for random context, but oracle context reaches 0.6337, revealing substantial headroom. Selective abstention helps when wrong-repository controls are included, yet fails on the 50 natural no-gold cases, showing that confidence calibration remains unsolved.

Original abstract

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file semantic similarity. The benchmark covers four positive-retrieval tasks: code2test, comment2context, trace2code, and edit2ripple; a fifth subset evaluates selective retrieval using natural evidence-backed no-gold cases and counterfactual wrong-repository controls. Agent Retrieval Bench contains 427 samples across 25 repositories: 345 positive examples, 50 natural no-gold examples, and 32 counterfactual controls. The corpus includes 308 base-commit snapshots, 392,000 files, and 7.9 million chunks. We evaluate lexical retrieval, RepoMap, open-source embeddings, selective abstention, and logged agent context selection. No single retrieval family dominates: Qwen3-Embedding-4B has the best sample-weighted MRR on positive samples, Qwen3-Embedding-8B the best Recall@20, and RepoMap the best budgeted context yield at 8K tokens, with task-level winners differing substantially. Selective thresholds calibrated with counterfactual controls do not improve selective success on natural no-gold cases, revealing a calibration gap. Logged trajectories also miss every gold file on 27-35 percent of samples. A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis