NTH

CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test

AuthorsZhangyi Hu, Chenhui Liu, Tian Huang, Jindong Li, Yang Yang, Jiemin Wu, Zining Zhong, Menglin Yang, Yutao Yue

May 27, 2026 3 min read
Watch on YouTube
The one-line take

CoSPlay lets an LLM improve its own code and tests at inference time, using self-generated unit tests and execution signals to boost code generation without ground-truth labels.

Key results

22.1% to 33.2%
Qwen2.5-7B-Instruct average BoN

On the four benchmarks LiveBench, LiveCodeBench v2, CodeContests, and CodeForces, CoSPlay with cluster selection improves average Best-of-N accuracy from the Qwen2.5-7B-Instruct baseline to the reported final value.

14.6% to 78.3%
Qwen2.5-7B-Instruct average UT accuracy

The self-generated unit tests become much more accurate after CoSPlay’s cooperative self-play loop, as reported for Qwen2.5-7B-Instruct across the four-benchmark suite.

5.7% absolute
CURE-7B BoN gain

When applied to the RLVR-trained CURE-7B model, CoSPlay with cluster selection further improves Best-of-N accuracy by the reported absolute amount.

What the paper found

CoSPlay introduces a GT-free, training-free test-time scaling framework for code generation that jointly improves candidate programs and self-generated unit tests through cooperative self-play. Instead of relying on ground-truth unit tests for RLVR-style training or naïvely sampling more code and tests, it first uses PlanSearch-style natural-language exploration to generate diverse solution plans and then derives failure-oriented UT attack ideas from those plans. It initializes a code pool and a UT pool, then iteratively updates both via an execution matrix: UT pass counts estimate test reliability, code pass counts estimate code quality, all-failing codes are pruned, low-support UTs are regenerated to break spurious code–test coupling, high-support non-trivial UTs are used to repair failing codes, and trivial UTs are replaced. When Best-of-N ties remain, CoSPlay applies execution-consensus clustering over random valid inputs and selects the largest reliable output-signature cluster. On four benchmarks—LiveBench, LiveCodeBench v2, CodeContests, and CodeForces—CoSPlay on Qwen2.5-7B-Instruct raises average UT accuracy from 14.6% to 78.3% and BoN from 22.1% to 33.2%, matching or slightly exceeding CURE-7B, while also improving CURE-7B by 5.7% absolute BoN. It generalizes across 7B, 14B, and frontier-scale models such as DeepSeek-V3.2-685B and Gemini-2.0-Flash, and outperforms scaled TTS baselines under comparable token budgets.

Original abstract

Recently, Reinforcement Learning with Verifiable Rewards (RLVR) and Test-Time Scaling (TTS) have advanced LLM code generation through executable verification. Yet Ground-Truth Unit Tests (GT UTs) remain a bottleneck: SOTA RLVR methods require them for costly training, while existing TTS methods lose competitiveness without them. This motivates GT-free TTS, where existing methods directly use self-generated UTs to refine and select code candidates. Yet such UTs are often noisy or spuriously coupled with wrong code, and UT quality in turn cannot be validated without reliable code. The key challenge is therefore to jointly improve both. To this end, we present CoSPlay, a GT-free, training-free framework that jointly improves codes and UTs through cooperative self-play. It first explores diverse solution ideas and identifies their potential failure modes to produce discriminative UT ideas. It then uses bidirectional pass-count signals from the Code-UT execution matrix to iteratively prune or fix weak codes and refresh or replace unreliable UTs, letting the two pools co-evolve. Finally, when multiple codes remain tied at the highest pass count, it picks the final code from the largest output-consensus cluster, since correct codes agree on the same inputs while wrong codes diverge. Experiments on four challenging benchmarks show that CoSPlay on Qwen2.5-7B-Instruct improves average BoN from 22.1% to 33.2% and UT accuracy from 14.6% to 78.3%, matching or surpassing the RLVR model CURE-7B. When applied to CURE-7B, it further improves BoN by 5.7%. CoSPlay also generalizes across diverse backbones and outperforms GT-free TTS baselines under comparable token budgets, with continued gains as the budget scales up. These results suggest a scalable inference strategy for competitive code generation without any GT data.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis