Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
AuthorsHao Xiang, Qiaoyu Tang, Le Yu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Bowen Yu, Peng Wang, Hongyu Lin, Dayiheng Liu
Resources
The paper shows that you can build bigger reasoning-training tasks by composing smaller verifiable environments like LEGO bricks, improving LLM reasoning generalization with far less manual environment design.
Key results
Initial verifiable environments used to build composite tasks
Main RL training scale for each configuration
RACES-trained average across six unseen benchmarks
RACES-trained average across six unseen benchmarks
Qwen3-4B-Instruct-2507 trained with only 50 base environments
What the paper found
This paper introduces RACES, short for Recursive Automated Composition for Environment Scaling, a framework for reinforcement learning with verifiable environments that treats each environment like a LEGO brick: if one environment’s output type matches another’s input type, they can be recursively composed into new executable tasks. RACES standardizes environments as a four-tuple of input sampler, deterministic mapper, natural-language descriptor, and verifier, then instantiates four operators—SEQUENTIAL, PARALLEL, SORT, and SELECT—to induce different reasoning behaviors such as state carryover, multi-thread tracking, order recovery, and distractor discrimination. The authors build a pool of 300 environments and train Qwen3-14B and DeepSeek-R1-Distill-Qwen-14B with GRPO on 32× NVIDIA A100 80GB GPUs using 12,800 training instances for 300 steps. On six unseen benchmarks, RACES improves DeepSeek-R1-Distill-Qwen-14B from 48.2 to 51.3 average score and Qwen3-14B from 58.8 to 61.1, with notable gains on IFEval, LongBench-v2, Enigmata, LiveCodeBench, and AIME. The method is also more sample-efficient: using only 50 base environments, RACES reaches 50.8 on Qwen3-4B-Instruct-2507, surpassing individual-environment RL trained on 300 environments at 50.4, while deeper compositions show a non-monotonic difficulty curve with best average performance at composition size 5. The key result is that compositional structure, not just more isolated tasks, produces stronger reasoning generalization.
Original abstract
Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research demonstrates that scaling environment quantity improves RL performance, existing manual or individual construction methods suffer from linear scaling limits, thereby hindering scalable reasoning generalization. This paper introduces RACES (\textbf{R}ecursive \textbf{A}utomated \textbf{C}omposition for \textbf{E}nvironment \textbf{S}caling), a framework that conceptualizes verifiable environments as composable building blocks that can be recursively assembled. The key insight is that when the codomain (output type) of one environment matches the domain (input type) of another, they can be automatically fused into a new verifiable environment, enabling recursive composition. RACES is implemented with 300 individual environments and defines a set of composition operators (\textsc{SEQUENTIAL}, \textsc{PARALLEL}, \textsc{SORT}, and \textsc{SELECT}) that induce diverse reasoning patterns. Extensive experiments show that RL training on these composite environments consistently enhances reasoning generalization. Specifically, RACES improves DeepSeek-R1-Distill-Qwen-14B by an average of 3.1 points (from 48.2 to 51.3) and boosts Qwen3-14B performance from 58.8 to 61.1 on six benchmarks, which are unseen during the construction of training environments. Moreover, RACES achieves performance comparable to training on 300 individual environments using only 50 base environments, demonstrating significant efficiency in environment utilization.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.