From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
AuthorsChao Chen, Chengzu Li, Zhiwei Li, Yinhong Liu, Zhijiang Guo
Resources
This paper turns an LLM from a student into a coach: it studies its own failures and redesigns the training environment to help itself learn better in reinforcement learning.
Key results
Aggregate valid rate for Qwen3-4B + GRPO + Ours on the 3-agent benchmark
Aggregate optimal rate for Qwen3-4B + GRPO + Ours on the 3-agent benchmark
Aggregate valid rate for Qwen3-4B + GRPO + Ours on the 4-agent benchmark
Aggregate optimal rate for Qwen3-4B + GRPO + Ours on the 4-agent benchmark
Aggregate valid rate for Qwen3-4B + GRPO + Ours on the 5-agent benchmark
Aggregate optimal rate for Qwen3-4B + GRPO + Ours on the 5-agent benchmark
What the paper found
This paper, from LARK at HKUST(GZ) and the University of Cambridge, reframes RL training as a closed-loop environment-design problem: after each stage, the current Qwen3-4B checkpoint becomes an “LLM-as-Environment-Engineer” that reads structured failure breakdowns, validation history, and training context, then rewrites the next environment generator instead of selecting examples. The authors introduce MAPF-FrozenLake, a controllable multi-agent path-finding testbed with 3×3 to 10×10 grids and generator knobs for data_ratio, hole_ratio, and wait_ratio, trained with GRPO under an adaptive reward that combines strict path-validity checks and a length penalty. On 3-, 4-, and 5-agent generalization benchmarks, the Qwen3-4B + GRPO + Ours system achieves the strongest aggregate performance, reaching 51.67 valid rate and 31.67 optimal rate on 3-agent tasks, 33.14/21.33 on 4-agent tasks, and 18.67/11.00 on 5-agent tasks, while outperforming larger proprietary LLMs including GPT-5.4, Grok-4.2, Gemini-3.1-Pro, and Kimi-K2.5 as environment designers. The analysis shows the key mechanism is evidence-driven adaptation: successful designs preserve what already works, target specific failure modes, and concentrate training near the competence frontier rather than monotonically increasing difficulty. An ablation further shows that “bookkeeping-only” training details outperform full RL-detail prompts, and that the current RL checkpoint is a better environment engineer than the untrained base model, indicating that policy learning improves self-diagnosis.
Original abstract
Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environments between stages, requiring practitioners to heuristically infer which configuration will best improve the current policy. To automate this process, we propose the LLM-as-Environment-Engineer framework in which the current policy model analyzes failure trajectories together with contextual information and proposes modifications to the next-stage training environment configuration. We also introduce MAPF-FrozenLake, a controllable testbed whose generator exposes multi-dimensional environment configurations, making it suitable for studying and benchmarking environment redesign. On this testbed, we condition the environment engineer on structured summaries of policy behavior, failure cases, and environment statistics, from which it produces the configuration for the next training stage. With Qwen3-4B as the backbone, our framework achieves the strongest aggregate performance on our benchmarks, outperforming larger proprietary LLMs (e.g., GPT, Gemini) and fixed-environment training baselines. We further analyze which forms of context are most effective, finding that successful environment updates rely on failure evidence and preserve configurations that already work. Interestingly, the current RL checkpoint serves as a better environment engineer than the original base model, suggesting that policy learning improves the model's ability to diagnose its remaining weaknesses.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.