Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
AuthorsZhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong
Resources
This paper introduces a more stable asynchronous reinforcement learning method for training agentic LLMs, using single-rollout updates to better handle long-horizon tasks like coding and reasoning.
Key results
SAO accuracy on the reasoning benchmark
SAO accuracy on the reasoning benchmark
SAO accuracy on the reasoning benchmark
SAO accuracy on the reasoning benchmark
SAO accuracy on the coding benchmark
SAO trains stably for around this many steps
What the paper found
Single-Rollout Asynchronous Optimization (SAO) is a reinforcement-learning method from Tsinghua University, deployed in the agentic RL pipeline for GLM-5.2, that targets the main failure mode of asynchronous LLM training: policy lag and unstable off-policy updates. Instead of GRPO-style group sampling, SAO uses one rollout per prompt and trains immediately when a trajectory finishes, then stabilizes updates with direct double-sided token-level importance sampling, strict clipping and masking, a faster critic schedule with K = 2 value updates per policy update, frozen-attention value-model tuning, and a skip-observation token-level GAE for multi-turn agent traces. On reasoning benchmarks with Qwen3-30B-A3B, SAO reaches 97.3 on AIME2025, 74.8 on BeyondAIME, 88.3 on HMMT Nov 2025, and 74.0 on IMOAnswerBench, beating GRPO and remaining stable for around 1000 training steps; on SWE-Bench Verified it improves from 27.0 with GRPO plus DIS to 29.8. The paper also shows that the single-rollout design adapts better than running-mean baselines in a simulated online learning task with shifting style rewards, where GLM-4.7 is used as the judge and SAO rapidly realigns after each preference shift.
Original abstract
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framework does not naturally fit asynchronous agentic training. In this paper, we present Single-rollout Asynchronous Optimization (SAO) to address the stability and off-policy challenges in asynchronous RL. To reduce off-policy effects and improve generalization, we replace group-wise sampling with single-rollout sampling, that is, using one rollout per prompt. We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. We also demonstrate that single-rollout RL is particularly effective in a simulated online learning setting, where the model must adapt to changing evolving environments. To this end, SAO is successfully deployed in the agentic RL pipeline for training the open GLM-5.2 model (750B-A40B).
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.