One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
AuthorsSimon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
Resources
Training AI agents against just one simulated user can make them brittle, so this paper uses diverse simulated users to help agents generalize to real people.
Key results
Maximum held-out success improvement over single-simulator RL.
Further held-out gain reported for Co-Training.
Qwen3-4B-Instruct held-out retail success rate.
Qwen3-4B-Instruct held-out airline success rate.
Real-user task outcome, compared with 0.43 for single-simulator RL.
What the paper found
The paper identifies simulator collapse as a structural failure in multi-agent reinforcement learning: when an agent is trained against one frozen, mode-collapsed LLM user, group-relative REINFORCE or GRPO updates reward narrow strategies that exploit that simulator’s dominant behavior. Training reward can rise while held-out performance and policy entropy collapse, causing poor transfer to unseen models and real users. The study evaluates Qwen3-4B-Instruct and Qwen3-8B agents against simulators including OpenAI’s GPT-5-mini, Anthropic’s Claude Haiku 4.5, and Google’s Gemini-3-Flash across Persuasion for Good, τ 2 -bench, and CooperBench. It proposes Verbalized Sampling, which samples one response from a simulator-generated distribution of five plausible replies at each turn, and Co-Training, which updates the agent and simulator together; Population Co-Training additionally samples from a FIFO pool of recent simulator checkpoints. Verbalized Sampling improves held-out success by up to 9%, while Co-Training reaches gains of 14%; on Qwen3-4B-Instruct, Population Co-Training achieves 62.2% retail and 45.7% airline success on τ 2 -bench. In a human study, Co-Training raises τ 2 -bench task outcome to 0.70, compared with 0.43 for single-simulator RL. The released SCOPE framework unifies simulator rotation, self-play, checkpoint populations, and dual-model co-training, supporting the paper’s central conclusion that environment diversity is as important as policy diversity for real-world generalization.
Original abstract
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.