Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
AuthorsJiasheng Zheng, Boxi Cao, Boxi Yu, Yuzhong Zhang, Jialun Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
Resources
This paper introduces a new way to create harder, more original coding tasks for training LLMs, helping reinforcement learning with verifiable rewards scale beyond hand-crafted or heuristic data.
Key results
ADR synthetic-data originality score, compared with Educational Instruct's 6.04
ADR synthetic-data difficulty score
ADR synthetic-data diversity score
ADR synthetic-data test quality score
ADR on Qwen2.5-Coder-7B-Instruct
ADR training-dynamics improvement over the base model
What the paper found
This paper, from the Chinese Academy of Sciences and collaborating universities, introduces Atomic Decomposition and Recombination (ADR), a fully automated framework for scaling code Reinforcement Learning with Verifiable Rewards by generating tasks from atomic logical elements rather than heuristic seed expansion. ADR first extracts a compact element schema from seed problems, then uses controlled recombination, template-based synthesis, execution-grounded validation, and adversarial solution-space refinement to produce novel, solvable, and harder verifiable code tasks. The paper formalizes a four-dimensional quality taxonomy—originality, difficulty, diversity, and test quality—and reports that ADR reaches 28.91 originality, 71.89 difficulty, 84.36 diversity, and 81.36 test quality, versus Educational Instruct’s 6.04 originality. On LiveCodeBench v5 and v6, ADR trained on Qwen2.5-Coder-7B-Instruct achieves 25.37% and 26.14% Pass@1, outperforming KodCode’s 22.75% and 23.57%, and the same approach lifts Qwen3-8B to 35.85% and 31.43%. In training dynamics, ADR delivers a +4.79% Pass@8 gain, compared with only +0.60% for the strongest baseline, showing that the method expands the model’s capability frontier rather than merely increasing sample density. The framework also transfers to tool usage and data science, reaching 41.67% on BigCodeBench and 42.44% on DS-1000.
Original abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of RLVR is severely constrained by the scarcity of sufficiently challenging verifiable code tasks that target near the model's edge of competence. Prior studies often rely on heuristic seed expansions for data synthesis, which severely limits both novelty and difficulty. Consequently, the training value of such data fails to scale proportionally with the size of its synthesis. To this end, we propose Atomic Decomposition and Recombination (ADR), a novel framework that generates verifiable code tasks via decomposition into atomic elements and controlled recombination, thereby enabling the generation of genuinely novel and challenging verifiable code tasks. Experiments and analysis demonstrate that ADR achieves superior originality, difficulty, diversity, and test quality over existing baselines, and consistently delivers greater improvements in code ability across RLVR in diverse downstream domains, including algorithmic programming, tool usage, and data science. Our work sheds light on a new paradigm for novel code task synthesis and scalable RLVR training.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.