SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
AuthorsJialong Liu, Yuling Shi, Ning Yang, Xiaodong Gu, Zuchao Li
SRPO teaches language models to turn their own mistakes into step-by-step learning signals, improving long-horizon reasoning and agent performance with relatively little training.
Key results
Qwen3-8B performance after SRPO training
Long-horizon shopping-agent benchmark result
Interactive household-task benchmark result
Real-world software-engineering benchmark result
SRPO uses 3.8× fewer total training FLOPs than GRPO
What the paper found
SRPO, or Self-Reflective Policy Optimization, addresses the credit-assignment failure of PPO and GRPO on long-horizon reasoning, where a terminal success signal provides only sparse episode-level supervision. Using a two-stage process, the model first reviews a completed trajectory and compresses its diagnosis into a 2–5-point reflection patch, then prepends that patch to the original prompt through reset-with-memory. The reflection-conditioned policy becomes a temporary teacher, while the unmodified policy generates on-policy rollouts and is trained with teacher-forced, per-token reverse-KL rewards and a clipped PPO objective. This converts sparse feedback from O(1) information per episode into O(T) token-level signals without external critics, reward models, or larger teachers. On Qwen3-8B, SRPO reaches 73.3% on AIME’24, 64.7% on WebShop, 76.8% on ALFWorld, and 31.2% on SWE-Bench-Lite, while using 3.8× fewer total FLOPs than GRPO. The method also generalizes across Qwen3-1.5B, Qwen3-32B, and Llama-3.1-8B-Instruct. Ablations show that compact, semantically aligned reflections, reverse KL, and state resetting are essential; verbose or mismatched reflections largely eliminate the gains. Reflection quality was additionally assessed with GPT-4, supporting SRPO’s central claim that self-generated hindsight guidance can internalize corrective reasoning without requiring a stronger external model.
Original abstract
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framework that internalizes this capability. SRPO enables LLMs to analyze their own completed trajectories, synthesize errors into concise "reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals. This process effectively transforms sparse terminal supervision into dense, token-level learning signals without requiring external critics, separate reward models, or larger teacher models. We demonstrate that SRPO achieves state-of-the-art performance across mathematical reasoning and long-horizon agentic benchmarks with exceptional data efficiency. Using a Qwen3-8B base model, SRPO attains 73.3% on AIME'24 using only 8% (0.08x) of the training FLOPs required by scaled supervised fine-tuning, while significantly improving success rates on WebShop (64.7%), ALFWorld (76.8%), and SWE-Bench-Lite (31.2%). Code is available at https://github.com/Galleons2029/SRPO
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.