ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
AuthorsJinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Resources
ReflectRL turns expert mistakes into useful training signals by teaching language models to reflect on flawed solutions before learning to solve problems directly.
Key results
Expert failure trajectories used for reflective training.
In-distribution and out-of-distribution reasoning benchmarks.
Qwen2.5-Math-7B with ReflectRL.
Percentage-point gain for Qwen2.5-Math-7B over GRPO.
Approximate seconds per step after the transition, versus around 20 for GRPO.
Entropy maintained at step 250, while GRPO falls below 0.03.
What the paper found
ReflectRL tackles sparse rewards in on-policy reasoning by reusing failed expert traces, called Golden Negative Trajectories, as structured error-analysis prompts rather than examples to imitate. Built from DeepSeek-R1 failures, the OpenR1-GNT-69k dataset contains 69k trajectories. The method elicits Reflective Reasoning—identify, repair, and solve from a flawed trace—then uses a cosine Reflective-to-Direct Policy Transition to remove dependence on external traces at inference. It plugs into RLVR methods such as GRPO and DAPO, as well as On-Policy Distillation, without auxiliary losses, new trainable parameters, or online expert queries. Across 9 benchmarks and Qwen2.5 and Llama-3.1 backbones spanning 1.5B to 8B parameters, Qwen2.5-Math-7B with GRPO reaches an in-distribution average of 42.4, up 5.4 percentage points, and an out-of-distribution average of 40.0, up 19.1 percentage points. ReflectRL also shortens successful reasoning from over 800 tokens to about 420 tokens and reduces update time from around 20 seconds to approximately 13 seconds per step. Mechanistically, the useful signal comes from a coherent valid prefix plus a localized error region; shuffled traces and answer-only context underperform. Training remains exploratory: at step 250, ReflectRL maintains entropy around 0.15 while GRPO falls below 0.03. The experiments use NVIDIA H200 GPUs and show that high-quality mistakes from models such as DeepSeek-R1 can teach Qwen and Llama systems how to diagnose and correct reasoning failures.
Original abstract
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.