NTH

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

AuthorsSimon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi

August 18, 2026 2 min read
Watch on YouTube
The one-line take

Training AI agents against just one simulated user can make them brittle, so this paper uses diverse simulated users to help agents generalize to real people.

Key results

9%
Verbalized Sampling held-out improvement

Maximum held-out success improvement over single-simulator RL.

14%
Co-Training held-out improvement

Further held-out gain reported for Co-Training.

62.2%
Population Co-Training τ 2 -bench Retail

Qwen3-4B-Instruct held-out retail success rate.

45.7%
Population Co-Training τ 2 -bench Airline

Qwen3-4B-Instruct held-out airline success rate.

0.70
Human τ 2 -bench outcome with Co-Training

Real-user task outcome, compared with 0.43 for single-simulator RL.

What the paper found

The paper identifies simulator collapse as a structural failure in multi-agent reinforcement learning: when an agent is trained against one frozen, mode-collapsed LLM user, group-relative REINFORCE or GRPO updates reward narrow strategies that exploit that simulator’s dominant behavior. Training reward can rise while held-out performance and policy entropy collapse, causing poor transfer to unseen models and real users. The study evaluates Qwen3-4B-Instruct and Qwen3-8B agents against simulators including OpenAI’s GPT-5-mini, Anthropic’s Claude Haiku 4.5, and Google’s Gemini-3-Flash across Persuasion for Good, τ 2 -bench, and CooperBench. It proposes Verbalized Sampling, which samples one response from a simulator-generated distribution of five plausible replies at each turn, and Co-Training, which updates the agent and simulator together; Population Co-Training additionally samples from a FIFO pool of recent simulator checkpoints. Verbalized Sampling improves held-out success by up to 9%, while Co-Training reaches gains of 14%; on Qwen3-4B-Instruct, Population Co-Training achieves 62.2% retail and 45.7% airline success on τ 2 -bench. In a human study, Co-Training raises τ 2 -bench task outcome to 0.70, compared with 0.43 for single-simulator RL. The released SCOPE framework unifies simulator rotation, self-play, checkpoint populations, and dual-model co-training, supporting the paper’s central conclusion that environment diversity is as important as policy diversity for real-world generalization.

Original abstract

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →