NTH

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

AuthorsEzgi Korkmaz

July 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that common ways of evaluating deep reinforcement learning can produce misleading conclusions, especially as training data and model capacity scale.

Key results

100K
Low-data training budget

Arcade Learning Environment benchmark used for data-scarce evaluation.

200M
High-data training budget

Frame-training scale used for asymptotic comparisons.

0.2304
Dueling human-normalized median

ALE 100K median score for the dueling architecture.

0.0941
C51 human-normalized median

ALE 100K median score for C51.

82%
DRQ reported gain

Original DRQ gain over DER before direct comparison with dueling.

11%
DRQ direct-comparison gain

Gain over the core dueling baseline in the paper’s direct comparison.

What the paper found

Ezgi Korkmaz’s paper challenges a central assumption in deep reinforcement learning: that algorithms ranking highly after extensive training will rank highly when data are scarce. Its theory proves a non-monotonic relationship between performance and sample complexity: lower-capacity models with larger inherent Bellman error can outperform more expressive distributional methods in the low-data regime, while the ordering reverses asymptotically. Large-scale experiments in the Arcade Learning Environment compare the 100K benchmark with 200M-frame training, revisiting the Atari research tradition associated with DeepMind, including DQN, C51, QRDQN, IQN, and the dueling architecture. In ALE 100K, dueling achieves a human-normalized median score of 0.2304, versus 0.0941 for C51, 0.0820 for QRDQN, and 0.0528 for IQN, showing that higher-capacity value-distribution models can be disadvantaged by limited data. The paper also exposes evaluation bias in data-efficient reinforcement learning: DRQ’s original comparisons reported an 82% gain over DER, but direct comparison with the core dueling baseline reduces that gain to 11%, while one reproduction performs 15% below dueling. Korkmaz argues that benchmarks must include core algorithms, match methods across capacity and data regimes, report direct baselines, and scrutinize dataset-selection assumptions, because conclusions drawn from ALE 100K can systematically misdirect subsequent research.

Original abstract

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →