Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
AuthorsChangdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li
Resources
This paper shows that RL post-training itself can provide a free, step-level signal for judging LLM agents, enabling better test-time scaling, uncertainty estimation, and failure analysis without training a separate reward model.
Key results
Average selected-trajectory success rate across BFCLv4-MT, WebShop, AgentDojo, and τ2-Airline
Average selected-trajectory success rate across BFCLv4-MT, WebShop, AgentDojo, and τ2-Airline
Uncertainty quantification on greedy trajectories
Uncertainty quantification on greedy trajectories
Uncertainty quantification on greedy trajectories
Uncertainty quantification on greedy trajectories
What the paper found
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents, from the University of Wisconsin–Madison and Argonne National Laboratory, argues that RL post-training already contains a free process-level signal for agent evaluation, eliminating dedicated process reward model training. The paper derives “progress advantage,” the log-probability ratio between an RL-trained policy and its reference policy, and proves that under a stochastic MDP it exactly recovers the optimal advantage function, while the exact implicit reward fails to telescope because of stochastic environment transitions. This signal is applied without task-specific training to test-time scaling, uncertainty quantification, and failure attribution on BFCLv4-MT, WebShop, AgentDojo, τ2-bench, and Who & When across Gemma4, Qwen3.5, Qwen3, and Olmo3 families. In best-of-8 selection it reaches an average success rate of 38.8 on Gemma4-4B and 62.1 on Qwen3.5-9B, exceeding training-based baselines such as WildReward-8B and ThinkPRM-14B; on τ2-bench AUROC it attains 0.865 on Airline and 0.690 on Retail with Gemma4-4B, and 0.720 and 0.678 with Qwen3.5-9B. For a task-specific WebShop comparison, progress advantage reaches 35.0 success rate versus 33.0 for AgentPRM-7B, and a GRPO-tuned variant reaches 38.0.
Original abstract
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov decision process, which we term progress advantage -- log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This formulation makes the resulting signal annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, it consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.