Intelligence from Learnable Novelty
AuthorsYanbo Zhang, Michael Levin
Resources
A single differentiable measure of learnable novelty may unify exploration, computation, and unsupervised abstraction across intelligent systems.
Key results
Epiplexity ranked rule 110 highest across all 88 rules.
Gradient ascent produced coherent traveling and colliding solitons.
The unsupervised encoder mapped images into 64-dimensional unit-normalized codes.
Both linear and 5-nearest-neighbor probes reached 0.89 accuracy.
The epiplexity bonus beat the task-reward baseline in 9 of 10 environments.
Reinforcement-learning results were reported after 600,000 training steps.
What the paper found
In “Intelligence from Learnable Novelty,” Yanbo Zhang and Michael Levin of the Allen Discovery Center at Tufts and Harvard’s Wyss Institute propose epiplexity, the portion of surprise that a bounded learner can convert into a reusable model. Unlike raw novelty search, which can reward an unlearnable noisy signal, or surprise minimization, which favors a static “dark room,” epiplexity isolates learnable structure through minimum description length. Their estimator uses a frozen random reservoir and a closed-form ridge-regression readout, making the score cheap, deterministic, and differentiable. As a measurement, it ranked rule 110 highest across all 88 locally unique elementary cellular automata, recovering its status as a Turing-complete, computation-supporting system without supervision. As an optimization objective, gradient ascent trained a neural cellular automaton for 2,000 steps, producing traveling and colliding solitons. Applied to an unsupervised 64-dimensional MNIST encoder, the objective organized representations around digit classes: both a linear probe and a 5-nearest-neighbor classifier reached 0.89 accuracy without labels. In reinforcement learning, PPO received epiplexity as an intrinsic reward alongside task reward and improved performance in 9 of 10 environments after 600,000 training steps, while avoiding the catastrophic failures seen with a raw state-magnitude bonus. The central claim is that complexity generation, abstraction, and exploration can emerge from maximizing one observer-relative quantity, although the fixed reservoir limits the method to structure it can currently learn.
Original abstract
Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most content in a dark room. Both failures have a single cause: each objective treats as one quantity the surprise a learner can convert into knowledge and the surprise it never can. Here we show that the learnable part of that information, which we call learnable novelty, yields the seemingly disparate projections of intelligence, and we give a closed-form estimator of it built on a cheap and differentiable reservoir computer. Used as a measure, with no supervision of any kind, the estimator recovers decades of complexity classification, ranking the Turing-complete rule~110 highest among the elementary cellular automata. Used as an objective, its gradient carries a neural cellular automaton from simple dynamics into a regime of solitons, the traveling, colliding structures by which rule~110 computes, as well as organizes the representation of an image encoder around the ten digit classes of MNIST, fully unsupervised: no label ever enters training. Handed to a reinforcement-learning agent as an intrinsic reward, it supplies the exploration that task rewards lack, improving on the task baseline in nine of ten environments and collapsing in none. Complexity generation, abstraction, and exploration, ordinarily pursued with unrelated objectives in separate fields, thus emerge from ascent on one differentiable quantity, and the projections of intelligence gain a common quantitative footing.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.