NTH

Efficient Exploration Is Enough

AuthorsMikel Malagón, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano

September 12, 2026 2 min read
Watch on YouTube
The one-line take

The paper argues that agents can develop increasingly complex behaviors simply by seeking experiences that improve their ability to predict and adapt, without external rewards or predefined tasks.

Key results

512
Evaluation horizon

Interaction steps used to estimate exploration efficiency.

128
OpenES population

Agents evolved in each optimization generation.

2000
OpenES generations

Evolutionary optimization duration.

16
Experimental repetitions

Repeated runs used for the empirical evaluation.

0.0125
RandColors right-room irreducible error

Best achievable prediction error for the higher-entropy room.

0.0025
RandColors left-room irreducible error

Best achievable prediction error for the lower-entropy room; the right-room error is 5 times higher.

What the paper found

This paper proposes a definition of exploration based not on uniform environment coverage or immediate prediction surprise, but on generating experience that most improves a world model’s global generalization. Its Expected Cumulative Error, or ECE, evaluates the predictive loss of models learned from an agent’s interaction history, making lower ECE more efficient. Theoretically, under Markovian dynamics and a tabular Dirichlet-Multinomial model, optimal explorers can be fully deterministic and prioritize informative, low-entropy, learnable state-action regions before harder or noisier ones; absorbing states are avoided because redundant visits waste the interaction budget. The analysis shows that prediction-error reduction scales with visitation count through variance and bias terms, with the latter decaying faster. Empirically, the SmallWorld suite uses GRU agents, feed-forward world models trained with AdamW, and OpenES to minimize a Monte Carlo ECE estimate across Empty, Blocks, Maze, and RandColors. With a 512-step horizon, a population of 128 agents, 2000 generations, and 16 repetitions, optimization produces structured behaviors without extrinsic rewards: helix-like navigation, symmetry exploitation, cyclic maze traversal, and an automatic curriculum from deterministic corridors to stochastic rooms. In RandColors, the right room has irreducible error 0.0125 versus 0.0025 for the left room, a 5 times difference, so agents explore easier regions first and later allocate more time to the harder region. The result challenges reward-centered and noise-driven exploration by showing that generalizable experience alone can induce survival-oriented and progressively complex behavior.

Original abstract

This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →