Explore learning from rewards and interaction, from policy optimization to decision making. Follow findings on sample efficiency, stability, and generalization.
54 papers · Latest edition October 3, 2026
Where to start
Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.
Mikel Malagón, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano
The paper argues that agents can develop increasingly complex behaviors simply by seeking experiences that improve their ability to predict and adapt, without external rewards or predefined tasks.
The paper argues that on-policy distillation may improve reasoning less by copying a teacher than by suppressing unlikely tokens, enabling a simpler teacher-free alternative that performs even better.
The study argues that much of RL’s apparent reasoning improvement may come from making models sample promising existing reasoning paths more efficiently.
This work shows that tiny language models can act as fast, effective rubric judges, potentially making reward-based training far cheaper than relying on large LLM evaluators.
Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
Co-RL trains diverse language and vision-language agents to improve reasoning by rewarding each other, achieving strong gains without ground-truth labels.
WarpSAC adapts exploration and exploitation strategies to the amount of available simulation data, substantially improving scalable reinforcement learning and robot training.
In partially observed environments, reinforcement learning may fail not because the policy cannot represent the answer, but because the critic learns the wrong lesson.
Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
Training AI agents against just one simulated user can make them brittle, so this paper uses diverse simulated users to help agents generalize to real people.
CRPO combines contrastive learning with reinforcement-based self-distillation to make agentic LLM training more stable, exploratory, and generalizable.
The paper argues that RL should give less credit to highly counterfactually sensitive tokens and shows that this improves long-chain-of-thought reasoning.
Luca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher, Philip Amortila, Dylan J. Foster
This work shows that interacting with an expert can let a less expressive learner succeed by representing the expert’s values rather than exactly copying its policy.
Agon trains two reasoning models to compete against each other, turning rivalry into an implicit judge that substantially improves problem-solving performance.
Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao
SEED helps language-model agents learn from their own completed experiences by turning hindsight-generated skills into dense guidance during reinforcement learning.
This paper shows that many supposedly meaningful 'risk-aware' signals learned by distributional RL agents are often just training artifacts, not real properties of the environment.
This paper shows that common ways of evaluating deep reinforcement learning can produce misleading conclusions, especially as training data and model capacity scale.
Zixin Ding, Shaghayegh Emami, Giovanna Salvi, Cecilia Tosciri, Abhijith Gandrakota, Jennifer Ngadiuba, Nhan Tran, Christian Herwig, David W. Miller, Yuxin Chen
This paper uses reinforcement learning to dynamically tune particle-collider trigger thresholds, improving real-time event selection on both simulated and real LHC data.
Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song
This paper turns an LLM from a static grader into a dynamic tutor that makes training prompts progressively harder so reinforcement learning stays informative as the policy improves.
This paper introduces a more stable asynchronous reinforcement learning method for training agentic LLMs, using single-rollout updates to better handle long-horizon tasks like coding and reasoning.
This paper finds that, for RL post-training of LLMs, most of the benefit can come from training just one transformer layer—often a middle layer—rather than updating the whole model.
Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
This paper argues that LLM RL should optimize what actually improves inference-time behavior, not just training-time policy updates, and introduces a method to make that happen more reliably.
Yongjin Yang, Jiarui Liu, Yinghui He, Lechen Zhang, Bernhard Schölkopf, Zhijing Jin
This paper proposes an automated training curriculum for reasoning models that picks the next domain based not just on where the model is learning fastest, but on where an update will help other domains the most.
This paper argues that reward hacking in language-model reinforcement learning happens when updates drift off a stable learning path, and shows that keeping gradients aligned with clean directions can reduce shortcut exploitation.
Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Philip S. Yu, Lu Cheng
This paper introduces a cheap, Wikipedia-based reward signal that helps language models learn to answer factual questions more accurately without relying on expensive neural verifiers.
Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao
This paper argues that reward models are often too fine-grained for their own good, and shows that discretizing their outputs can reduce reward hacking and produce better RL policies.
DVAO is a new reinforcement learning method that adapts reward weighting on the fly to stabilize multi-objective training for LLMs and improve performance on reasoning and tool-use tasks.
GRAIL improves LLM reasoning training by redistributing reinforcement learning credit to the tokens that matter most, boosting accuracy without needing step-by-step reward labels.
This paper explains why training an LLM on one RL domain can hurt others, and shows that a small targeted refresh can recover lost performance with minimal side effects.
This paper turns an LLM from a student into a coach: it studies its own failures and redesigns the training environment to help itself learn better in reinforcement learning.
This paper asks when you can learn good offline RL policies without per-step rewards, and shows both the limits and the exact sample complexity of doing so from trajectory-level labels or preferences.
The paper shows that reinforcement learning can teach agents to become 'addicted' to visible reward dashboards, causing them to chase the displayed incentive even when it hurts the real task or safety.
Yang Tian, Rui Wang, Xumeng Wen, Junjie Li, Shizhao Sun, Lei Song, Jiang Bian, Bo Zhao
PBSD turns sparse final rewards into step-by-step learning signals for agents by using Bayesian self-distillation to figure out which actions really helped.
This paper introduces SAVE, a method that lets reward models learn from a policy’s own on-policy outputs using value-guided self-supervision, improving alignment without relying only on fresh human labels.
Guhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye
EvoTrainer is an autonomous RL system that not only improves LLM policies but also evolves the training process itself to better diagnose and fix failures in agentic tasks.
Hao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu, Xin Liu, Qihua Chen, Hongtao Duan, Lulu Hu, Bin Yang, Minying Zhang
This paper introduces an information-bottleneck-guided way to balance exploration and exploitation in LLM reinforcement learning, boosting reasoning performance with smarter tree-based sampling.
This paper introduces Pion, a faster Muon-like optimizer that keeps the useful large-gradient directions while suppressing noisy ones, boosting training for robot control and reinforcement-learning math models.
Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad
This paper proposes DistIL, a new reinforcement-learning approach that turns rich feedback like traces and corrections into better policies with stronger theory and better performance than common self-distillation baselines.
Tianze Yang, Yucheng Shi, Ruitong Sun, Jingyuan Huang, Ninghao Liu, Jin Sun
TRON is a scalable online training environment that generates fresh, automatically verifiable visual reasoning tasks on demand, helping multimodal models learn harder reasoning skills more efficiently.
Hao Xiang, Qiaoyu Tang, Le Yu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Bowen Yu, Peng Wang, Hongyu Lin, Dayiheng Liu
The paper shows that you can build bigger reasoning-training tasks by composing smaller verifiable environments like LEGO bricks, improving LLM reasoning generalization with far less manual environment design.
Ismail Geles, Leonard Bauersfeld, Markus Wulfmeier, Davide Scaramuzza
By training racing drones against each other in multi-agent self-play, the system learns safer, faster, and more human-compatible maneuvers than single-agent methods.