NTH
Research collection

Reinforcement Learning research

Explore learning from rewards and interaction, from policy optimization to decision making. Follow findings on sample efficiency, stability, and generalization.

54 papers · Latest edition October 3, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Reinforcement Learning papers

Newest editions first.

05Rl

MInTRL: Off-policy Intervention can boost On-policy RL

Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong

MInTRL helps reinforcement-learning systems discover better solutions by making small, targeted corrections during otherwise on-policy rollouts.

Read analysis
07Rl

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Fanrui Zhang, Ruixue Ding, Qiang Zhang, Xi Chen, Boli Chen, Shihang Wang, Qiuchen Wang, Hongmin Zhan, Jinxin Bian, Li xingchao, Peijin Zheng, Hao cheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha

ARISE-RL trains agents to improve themselves by generating increasingly challenging tasks and learning from fine-grained rubric-based feedback.

Read analysis
08Rl

Efficient Exploration Is Enough

Mikel Malagón, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano

The paper argues that agents can develop increasingly complex behaviors simply by seeking experiences that improve their ability to predict and adapt, without external rewards or predefined tasks.

Read analysis
17Rl

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi

Training AI agents against just one simulated user can make them brittle, so this paper uses diverse simulated users to help agents generalize to real people.

Read analysis
22Rl

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Yu Wang, Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, Fuli Feng

CAST teaches LLM agents to make better long-horizon decisions by using a game solver’s changing state values as turn-by-turn guidance.

Read analysis
26Rl

Intelligence from Learnable Novelty

Yanbo Zhang, Michael Levin

A single differentiable measure of learnable novelty may unify exploration, computation, and unsupervised abstraction across intelligent systems.

Read analysis
27Rl

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao

SEED helps language-model agents learn from their own completed experiences by turning hindsight-generated skills into dense guidance during reinforcement learning.

Read analysis
30Rl

Understanding Reasoning from Pretraining to Post-Training

Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov

Using chess and math as controlled laboratories, the paper shows how pretraining sets the stage for—and shapes what RL can achieve in—LLM reasoning.

Read analysis
31Rl

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

Zixin Ding, Shaghayegh Emami, Giovanna Salvi, Cecilia Tosciri, Abhijith Gandrakota, Jennifer Ngadiuba, Nhan Tran, Christian Herwig, David W. Miller, Yuxin Chen

This paper uses reinforcement learning to dynamically tune particle-collider trigger thresholds, improving real-time event selection on both simulated and real LHC data.

Read analysis
32Rl

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song

This paper turns an LLM from a static grader into a dynamic tutor that makes training prompts progressively harder so reinforcement learning stays informative as the policy improves.

Read analysis
39Rl

Discretizing Reward Models

Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao

This paper argues that reward models are often too fine-grained for their own good, and shows that discretizing their outputs can reduce reward hacking and produce better RL policies.

Read analysis
51Rl

Reinforcement Learning from Rich Feedback with Distributional DAgger

Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

This paper proposes DistIL, a new reinforcement-learning approach that turns rich feedback like traces and corrections into better policies with stronger theory and better performance than common self-distillation baselines.

Read analysis