NTH

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

AuthorsJiasheng Zheng, Boxi Cao, Boxi Yu, Yuzhong Zhang, Jialun Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun

June 29, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a new way to create harder, more original coding tasks for training LLMs, helping reinforcement learning with verifiable rewards scale beyond hand-crafted or heuristic data.

Key results

28.91
Originality

ADR synthetic-data originality score, compared with Educational Instruct's 6.04

71.89
Difficulty

ADR synthetic-data difficulty score

84.36
Diversity

ADR synthetic-data diversity score

81.36
Test Quality

ADR synthetic-data test quality score

25.37%
LCB-v5 Pass@1

ADR on Qwen2.5-Coder-7B-Instruct

4.79%
Pass@8 gain

ADR training-dynamics improvement over the base model

What the paper found

This paper, from the Chinese Academy of Sciences and collaborating universities, introduces Atomic Decomposition and Recombination (ADR), a fully automated framework for scaling code Reinforcement Learning with Verifiable Rewards by generating tasks from atomic logical elements rather than heuristic seed expansion. ADR first extracts a compact element schema from seed problems, then uses controlled recombination, template-based synthesis, execution-grounded validation, and adversarial solution-space refinement to produce novel, solvable, and harder verifiable code tasks. The paper formalizes a four-dimensional quality taxonomy—originality, difficulty, diversity, and test quality—and reports that ADR reaches 28.91 originality, 71.89 difficulty, 84.36 diversity, and 81.36 test quality, versus Educational Instruct’s 6.04 originality. On LiveCodeBench v5 and v6, ADR trained on Qwen2.5-Coder-7B-Instruct achieves 25.37% and 26.14% Pass@1, outperforming KodCode’s 22.75% and 23.57%, and the same approach lifts Qwen3-8B to 35.85% and 31.43%. In training dynamics, ADR delivers a +4.79% Pass@8 gain, compared with only +0.60% for the strongest baseline, showing that the method expands the model’s capability frontier rather than merely increasing sample density. The framework also transfers to tool usage and data science, reaching 41.67% on BigCodeBench and 42.44% on DS-1000.

Original abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of RLVR is severely constrained by the scarcity of sufficiently challenging verifiable code tasks that target near the model's edge of competence. Prior studies often rely on heuristic seed expansions for data synthesis, which severely limits both novelty and difficulty. Consequently, the training value of such data fails to scale proportionally with the size of its synthesis. To this end, we propose Atomic Decomposition and Recombination (ADR), a novel framework that generates verifiable code tasks via decomposition into atomic elements and controlled recombination, thereby enabling the generation of genuinely novel and challenging verifiable code tasks. Experiments and analysis demonstrate that ADR achieves superior originality, difficulty, diversity, and test quality over existing baselines, and consistently delivers greater improvements in code ability across RLVR in diverse downstream domains, including algorithmic programming, tool usage, and data science. Our work sheds light on a new paradigm for novel code task synthesis and scalable RLVR training.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis