NTH

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization

AuthorsHao Xiang, Qiaoyu Tang, Le Yu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Bowen Yu, Peng Wang, Hongyu Lin, Dayiheng Liu

June 11, 2026 2 min read
Watch on YouTube
The one-line take

The paper shows that you can build bigger reasoning-training tasks by composing smaller verifiable environments like LEGO bricks, improving LLM reasoning generalization with far less manual environment design.

Key results

300
environment pool size

Initial verifiable environments used to build composite tasks

12,800
training instances

Main RL training scale for each configuration

51.3
DeepSeek-R1-Distill-Qwen-14B avg score

RACES-trained average across six unseen benchmarks

61.1
Qwen3-14B avg score

RACES-trained average across six unseen benchmarks

50.8
50-env RACES avg score

Qwen3-4B-Instruct-2507 trained with only 50 base environments

What the paper found

This paper introduces RACES, short for Recursive Automated Composition for Environment Scaling, a framework for reinforcement learning with verifiable environments that treats each environment like a LEGO brick: if one environment’s output type matches another’s input type, they can be recursively composed into new executable tasks. RACES standardizes environments as a four-tuple of input sampler, deterministic mapper, natural-language descriptor, and verifier, then instantiates four operators—SEQUENTIAL, PARALLEL, SORT, and SELECT—to induce different reasoning behaviors such as state carryover, multi-thread tracking, order recovery, and distractor discrimination. The authors build a pool of 300 environments and train Qwen3-14B and DeepSeek-R1-Distill-Qwen-14B with GRPO on 32× NVIDIA A100 80GB GPUs using 12,800 training instances for 300 steps. On six unseen benchmarks, RACES improves DeepSeek-R1-Distill-Qwen-14B from 48.2 to 51.3 average score and Qwen3-14B from 58.8 to 61.1, with notable gains on IFEval, LongBench-v2, Enigmata, LiveCodeBench, and AIME. The method is also more sample-efficient: using only 50 base environments, RACES reaches 50.8 on Qwen3-4B-Instruct-2507, surpassing individual-environment RL trained on 300 environments at 50.4, while deeper compositions show a non-monotonic difficulty curve with best average performance at composition size 5. The key result is that compositional structure, not just more isolated tasks, produces stronger reasoning generalization.

Original abstract

Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research demonstrates that scaling environment quantity improves RL performance, existing manual or individual construction methods suffer from linear scaling limits, thereby hindering scalable reasoning generalization. This paper introduces RACES (\textbf{R}ecursive \textbf{A}utomated \textbf{C}omposition for \textbf{E}nvironment \textbf{S}caling), a framework that conceptualizes verifiable environments as composable building blocks that can be recursively assembled. The key insight is that when the codomain (output type) of one environment matches the domain (input type) of another, they can be automatically fused into a new verifiable environment, enabling recursive composition. RACES is implemented with 300 individual environments and defines a set of composition operators (\textsc{SEQUENTIAL}, \textsc{PARALLEL}, \textsc{SORT}, and \textsc{SELECT}) that induce diverse reasoning patterns. Extensive experiments show that RL training on these composite environments consistently enhances reasoning generalization. Specifically, RACES improves DeepSeek-R1-Distill-Qwen-14B by an average of 3.1 points (from 48.2 to 51.3) and boosts Qwen3-14B performance from 58.8 to 61.1 on six benchmarks, which are unseen during the construction of training environments. Moreover, RACES achieves performance comparable to training on 300 individual environments using only 50 base environments, demonstrating significant efficiency in environment utilization.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →