NTH

Fractal basins trap latent reasoning

AuthorsJeffrey Lai, Anthony Bao, John Quinn, William Gilpin

September 15, 2026 2 min read
Watch on YouTube
The one-line take

The study argues that when AI reasoning gets harder, models become temporarily chaotic and get trapped near almost-correct solutions.

Key results

7M
Compact recurrent model size

Parameter count cited for recurrent models that surpassed larger models on ARC-AGI.

10B
Large language model comparison

Parameter scale exceeded by the larger language models in the ARC-AGI comparison.

10x
Adversarial compute amplification

Additional computing resources consumed by adversarial prompts versus similar benign prompts.

140M
Parcae model size

Parameter count of the looped language model fine-tuned for Countdown arithmetic.

200x200
Basin-map resolution

Resolution used for population-level sampling of two-dimensional initial-state slices.

300
Valid basin slices

Minimum number of valid slices used per model-task setting for correlation analysis.

What the paper found

This paper argues that reasoning slowdowns in AI are not merely software inefficiencies but signatures of transient chaos. Modeling latent-state reasoning as a discrete dynamical system, the researchers vary initial hidden states and map each state to its convergence time, revealing fractal basins whose complexity rises with task difficulty. The effect appears across Equilibrium Reasoners, Fixed-Point Reasoning Models, Parcae, and Tiny Recursive Model on Sudoku-Extreme, Maze-Hard, Countdown arithmetic, and ARC-AGI-1. Hard instances contain weakly unstable saddle points corresponding to nearly correct but invalid solutions—such as Sudoku grids with repeated digits or maze dead ends—which scatter nearby trajectories and produce long, unpredictable reasoning routes. Basin entropy correlates strongly with the number of reasoning loops, while the fast Lyapunov indicator identifies boundaries between divergent solution paths and tracks how often a trace changes candidate answers. A training study on integer linear systems shows that fractal structure emerges at a bifurcation where incorrect fixed points become saddles and the model first acquires multistep reasoning; only the core variables requiring Gaussian elimination exhibit positive finite-time Lyapunov exponents. The analysis uses 200×200 initial-state grids and at least 300 valid slices per model-task setting. The findings also contextualize overthinking: it can nearly double inference cost, while adversarial prompts can consume 10x more compute. Finally, the paper notes that compact recurrent systems with 7M parameters have surpassed language models exceeding 10B parameters on ARC-AGI, and examines Parcae as a 140M-parameter example, suggesting that reasoning capability depends on navigating—and escaping—complex latent landscapes rather than scale alone.

Original abstract

Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widely observed to reason for longer on harder tasks, but the general mechanism responsible for these slowdowns is unknown. Here, we show that reasoning models exhibit transient chaos, a physical consequence of the computational complexity of difficult tasks. As a consequence, we show that diverse leading reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks like Sudoku and maze solving, visual puzzles, and mathematical logic. We show that transient chaos emerges due to reasoning becoming trapped for extended durations near saddle points, which we show correspond to nearly-correct attempted solutions of the underlying problem. Our results show that reasoning slowdowns are an inevitable consequence of problem hardness in modern artificial intelligence models, and establish reasoning traces as a rich new class of dynamical system.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis