Efficient Exploration Is Enough
AuthorsMikel Malagón, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano
Resources
The paper argues that agents can develop increasingly complex behaviors simply by seeking experiences that improve their ability to predict and adapt, without external rewards or predefined tasks.
Key results
Interaction steps used to estimate exploration efficiency.
Agents evolved in each optimization generation.
Evolutionary optimization duration.
Repeated runs used for the empirical evaluation.
Best achievable prediction error for the higher-entropy room.
Best achievable prediction error for the lower-entropy room; the right-room error is 5 times higher.
What the paper found
This paper proposes a definition of exploration based not on uniform environment coverage or immediate prediction surprise, but on generating experience that most improves a world model’s global generalization. Its Expected Cumulative Error, or ECE, evaluates the predictive loss of models learned from an agent’s interaction history, making lower ECE more efficient. Theoretically, under Markovian dynamics and a tabular Dirichlet-Multinomial model, optimal explorers can be fully deterministic and prioritize informative, low-entropy, learnable state-action regions before harder or noisier ones; absorbing states are avoided because redundant visits waste the interaction budget. The analysis shows that prediction-error reduction scales with visitation count through variance and bias terms, with the latter decaying faster. Empirically, the SmallWorld suite uses GRU agents, feed-forward world models trained with AdamW, and OpenES to minimize a Monte Carlo ECE estimate across Empty, Blocks, Maze, and RandColors. With a 512-step horizon, a population of 128 agents, 2000 generations, and 16 repetitions, optimization produces structured behaviors without extrinsic rewards: helix-like navigation, symmetry exploitation, cyclic maze traversal, and an automatic curriculum from deterministic corridors to stochastic rooms. In RandColors, the right room has irreducible error 0.0125 versus 0.0025 for the left room, a 5 times difference, so agents explore easier regions first and later allocate more time to the harder region. The result challenges reward-centered and noise-driven exploration by showing that generalizable experience alone can induce survival-oriented and progressively complex behavior.
Original abstract
This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.