Fractal basins trap latent reasoning
AuthorsJeffrey Lai, Anthony Bao, John Quinn, William Gilpin
Resources
The study argues that when AI reasoning gets harder, models become temporarily chaotic and get trapped near almost-correct solutions.
Key results
Parameter count cited for recurrent models that surpassed larger models on ARC-AGI.
Parameter scale exceeded by the larger language models in the ARC-AGI comparison.
Additional computing resources consumed by adversarial prompts versus similar benign prompts.
Parameter count of the looped language model fine-tuned for Countdown arithmetic.
Resolution used for population-level sampling of two-dimensional initial-state slices.
Minimum number of valid slices used per model-task setting for correlation analysis.
What the paper found
This paper argues that reasoning slowdowns in AI are not merely software inefficiencies but signatures of transient chaos. Modeling latent-state reasoning as a discrete dynamical system, the researchers vary initial hidden states and map each state to its convergence time, revealing fractal basins whose complexity rises with task difficulty. The effect appears across Equilibrium Reasoners, Fixed-Point Reasoning Models, Parcae, and Tiny Recursive Model on Sudoku-Extreme, Maze-Hard, Countdown arithmetic, and ARC-AGI-1. Hard instances contain weakly unstable saddle points corresponding to nearly correct but invalid solutions—such as Sudoku grids with repeated digits or maze dead ends—which scatter nearby trajectories and produce long, unpredictable reasoning routes. Basin entropy correlates strongly with the number of reasoning loops, while the fast Lyapunov indicator identifies boundaries between divergent solution paths and tracks how often a trace changes candidate answers. A training study on integer linear systems shows that fractal structure emerges at a bifurcation where incorrect fixed points become saddles and the model first acquires multistep reasoning; only the core variables requiring Gaussian elimination exhibit positive finite-time Lyapunov exponents. The analysis uses 200×200 initial-state grids and at least 300 valid slices per model-task setting. The findings also contextualize overthinking: it can nearly double inference cost, while adversarial prompts can consume 10x more compute. Finally, the paper notes that compact recurrent systems with 7M parameters have surpassed language models exceeding 10B parameters on ARC-AGI, and examines Parcae as a 140M-parameter example, suggesting that reasoning capability depends on navigating—and escaping—complex latent landscapes rather than scale alone.
Original abstract
Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widely observed to reason for longer on harder tasks, but the general mechanism responsible for these slowdowns is unknown. Here, we show that reasoning models exhibit transient chaos, a physical consequence of the computational complexity of difficult tasks. As a consequence, we show that diverse leading reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks like Sudoku and maze solving, visual puzzles, and mathematical logic. We show that transient chaos emerges due to reasoning becoming trapped for extended durations near saddle points, which we show correspond to nearly-correct attempted solutions of the underlying problem. Our results show that reasoning slowdowns are an inevitable consequence of problem hardness in modern artificial intelligence models, and establish reasoning traces as a rich new class of dynamical system.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.