NTH

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

AuthorsDung Le Quang, Dong Cao Van, Nam Le Hai, Linh Ngo Van, Anh M. T. Bui, Phuong T. Nguyen

September 5, 2026 2 min read
Watch on YouTube
The one-line take

XREPOTEST tests whether today’s LLMs can generate genuinely useful unit tests inside real multilingual software repositories.

Key results

3642
Benchmark focal functions

Repository-level functions spanning Rust, Go, Julia, PHP, and Ruby.

5
Programming languages

The benchmark covers Rust, Go, Julia, PHP, and Ruby.

14
Evaluated LLMs

Models include Claude 4.5, GPT-5.2, DeepSeek V4-pro, Qwen, and Llama-3.3.

27%
Maximum standard-context TPR

The strongest models achieve at most this test pass rate across the benchmark.

57.4%
Average compilation success rate

Average fraction of generated tests that compile successfully.

9.7%
TPR-pass suites failing IR

Passing suites that do not directly invoke the intended focal function.

What the paper found

XR EPOTEST introduces a repository-level benchmark for multilingual unit-test generation across 5 programming languages—Rust, Go, Julia, PHP, and Ruby—using 3642 focal functions from real projects rather than isolated snippets. Its Docker-based evaluation runs native test frameworks and measures compilation success, test pass rate, line coverage, mutation score, and a new Invocation Rate, which checks whether tests directly call the intended function. Across 14 models, including Anthropic’s Claude 4.5, OpenAI’s GPT-5.2 and GPT-OSS, DeepSeek V4-pro, Alibaba’s Qwen models, and Meta’s Llama-3.3, repository-level testing remains difficult: even the strongest systems reach at most 27% test pass rate under standard context, while the average compilation success rate is 57.4%. Claude 4.5 Sonnet leads many language settings, but GPT-5.2 is especially strong on Go and Ruby. File-level context is the most consistently useful augmentation, raising Claude 4.5 Sonnet’s Go pass rate from 23.65% to 34.14% and coverage from 43.53% to 52.63%, although richer context can reduce invocation reliability. The dominant failure is API hallucination, followed by incorrect assertions and test logic. Crucially, 9.7% of suites that pass nevertheless fail Invocation Rate, showing that conventional pass and coverage metrics can reward tests that exercise wrappers, mocks, or unrelated APIs. Agentic execution and iterative repair improve pass rates, but can also lower coverage and invocation, so XR EPOTEST argues for evaluating correctness, behavioral targeting, and fault detection together.

Original abstract

Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness. We introduce XREPOTEST, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby. XREPOTEST evaluates tests under realistic repository constraints using a containerized execution framework and multiple context augmentation strategies, including file-level, LSP-based, and retrieval-based context. Beyond standard metrics such as test pass rate and coverage, we propose Invocation Rate (IR) to assess whether generated tests meaningfully exercise the intended functionality. Experiments with 14 state-of-the-art LLMs, including Claude 4.5, GPT-5.2, DeepSeek V4-Pro, and Qwen families, reveal a substantial gap between standalone and repository-level performance, as well as trade-offs between richer context and test reliability. Overall, XREPOTEST provides a challenging and informative benchmark to advance scalable and robust unit test generation in realistic software environments. The dataset and code are publicly available at: https://github.com/solis-team/XRepoTest

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis