NTH

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

AuthorsZhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong

July 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a more stable asynchronous reinforcement learning method for training agentic LLMs, using single-rollout updates to better handle long-horizon tasks like coding and reasoning.

Key results

97.3
AIME2025

SAO accuracy on the reasoning benchmark

74.8
BeyondAIME

SAO accuracy on the reasoning benchmark

88.3
HMMT Nov 2025

SAO accuracy on the reasoning benchmark

74.0
IMOAnswerBench

SAO accuracy on the reasoning benchmark

29.8
SWE-Bench Verified

SAO accuracy on the coding benchmark

1000
training steps

SAO trains stably for around this many steps

What the paper found

Single-Rollout Asynchronous Optimization (SAO) is a reinforcement-learning method from Tsinghua University, deployed in the agentic RL pipeline for GLM-5.2, that targets the main failure mode of asynchronous LLM training: policy lag and unstable off-policy updates. Instead of GRPO-style group sampling, SAO uses one rollout per prompt and trains immediately when a trajectory finishes, then stabilizes updates with direct double-sided token-level importance sampling, strict clipping and masking, a faster critic schedule with K = 2 value updates per policy update, frozen-attention value-model tuning, and a skip-observation token-level GAE for multi-turn agent traces. On reasoning benchmarks with Qwen3-30B-A3B, SAO reaches 97.3 on AIME2025, 74.8 on BeyondAIME, 88.3 on HMMT Nov 2025, and 74.0 on IMOAnswerBench, beating GRPO and remaining stable for around 1000 training steps; on SWE-Bench Verified it improves from 27.0 with GRPO plus DIS to 29.8. The paper also shows that the single-rollout design adapts better than running-mean baselines in a simulated online learning task with shifting style rewards, where GLM-4.7 is used as the judge and SAO rapidly realigns after each preference shift.

Original abstract

Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framework does not naturally fit asynchronous agentic training. In this paper, we present Single-rollout Asynchronous Optimization (SAO) to address the stability and off-policy challenges in asynchronous RL. To reduce off-policy effects and improve generalization, we replace group-wise sampling with single-rollout sampling, that is, using one rollout per prompt. We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. We also demonstrate that single-rollout RL is particularly effective in a simulated online learning setting, where the model must adapt to changing evolving environments. To this end, SAO is successfully deployed in the agentic RL pipeline for training the open GLM-5.2 model (750B-A40B).

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →