XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models
AuthorsDung Le Quang, Dong Cao Van, Nam Le Hai, Linh Ngo Van, Anh M. T. Bui, Phuong T. Nguyen
XREPOTEST tests whether today’s LLMs can generate genuinely useful unit tests inside real multilingual software repositories.
Key results
Repository-level functions spanning Rust, Go, Julia, PHP, and Ruby.
The benchmark covers Rust, Go, Julia, PHP, and Ruby.
Models include Claude 4.5, GPT-5.2, DeepSeek V4-pro, Qwen, and Llama-3.3.
The strongest models achieve at most this test pass rate across the benchmark.
Average fraction of generated tests that compile successfully.
Passing suites that do not directly invoke the intended focal function.
What the paper found
XR EPOTEST introduces a repository-level benchmark for multilingual unit-test generation across 5 programming languages—Rust, Go, Julia, PHP, and Ruby—using 3642 focal functions from real projects rather than isolated snippets. Its Docker-based evaluation runs native test frameworks and measures compilation success, test pass rate, line coverage, mutation score, and a new Invocation Rate, which checks whether tests directly call the intended function. Across 14 models, including Anthropic’s Claude 4.5, OpenAI’s GPT-5.2 and GPT-OSS, DeepSeek V4-pro, Alibaba’s Qwen models, and Meta’s Llama-3.3, repository-level testing remains difficult: even the strongest systems reach at most 27% test pass rate under standard context, while the average compilation success rate is 57.4%. Claude 4.5 Sonnet leads many language settings, but GPT-5.2 is especially strong on Go and Ruby. File-level context is the most consistently useful augmentation, raising Claude 4.5 Sonnet’s Go pass rate from 23.65% to 34.14% and coverage from 43.53% to 52.63%, although richer context can reduce invocation reliability. The dominant failure is API hallucination, followed by incorrect assertions and test logic. Crucially, 9.7% of suites that pass nevertheless fail Invocation Rate, showing that conventional pass and coverage metrics can reward tests that exercise wrappers, mocks, or unrelated APIs. Agentic execution and iterative repair improve pass rates, but can also lower coverage and invocation, so XR EPOTEST argues for evaluating correctness, behavioral targeting, and fault detection together.
Original abstract
Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness. We introduce XREPOTEST, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby. XREPOTEST evaluates tests under realistic repository constraints using a containerized execution framework and multiple context augmentation strategies, including file-level, LSP-based, and retrieval-based context. Beyond standard metrics such as test pass rate and coverage, we propose Invocation Rate (IR) to assess whether generated tests meaningfully exercise the intended functionality. Experiments with 14 state-of-the-art LLMs, including Claude 4.5, GPT-5.2, DeepSeek V4-Pro, and Qwen families, reveal a substantial gap between standalone and repository-level performance, as well as trade-offs between richer context and test reliability. Overall, XREPOTEST provides a challenging and informative benchmark to advance scalable and robust unit test generation in realistic software environments. The dataset and code are publicly available at: https://github.com/solis-team/XRepoTest
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.