Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
AuthorsEzgi Korkmaz
Resources
This paper shows that common ways of evaluating deep reinforcement learning can produce misleading conclusions, especially as training data and model capacity scale.
Key results
Arcade Learning Environment benchmark used for data-scarce evaluation.
Frame-training scale used for asymptotic comparisons.
ALE 100K median score for the dueling architecture.
ALE 100K median score for C51.
Original DRQ gain over DER before direct comparison with dueling.
Gain over the core dueling baseline in the paper’s direct comparison.
What the paper found
Ezgi Korkmaz’s paper challenges a central assumption in deep reinforcement learning: that algorithms ranking highly after extensive training will rank highly when data are scarce. Its theory proves a non-monotonic relationship between performance and sample complexity: lower-capacity models with larger inherent Bellman error can outperform more expressive distributional methods in the low-data regime, while the ordering reverses asymptotically. Large-scale experiments in the Arcade Learning Environment compare the 100K benchmark with 200M-frame training, revisiting the Atari research tradition associated with DeepMind, including DQN, C51, QRDQN, IQN, and the dueling architecture. In ALE 100K, dueling achieves a human-normalized median score of 0.2304, versus 0.0941 for C51, 0.0820 for QRDQN, and 0.0528 for IQN, showing that higher-capacity value-distribution models can be disadvantaged by limited data. The paper also exposes evaluation bias in data-efficient reinforcement learning: DRQ’s original comparisons reported an 82% gain over DER, but direct comparison with the core dueling baseline reduces that gain to 11%, while one reproduction performs 15% below dueling. Korkmaz argues that benchmarks must include core algorithms, match methods across capacity and data regimes, report direct baselines, and scrutinize dataset-selection assumptions, because conclusions drawn from ALE 100K can systematically misdirect subsequent research.
Original abstract
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.