NTH

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

AuthorsQucheng Gao, Zuyi Yang, Xiao Chen

August 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that self-attention can spontaneously organize token representations into stable clusters once attention becomes sufficiently sharp.

Key results

1
Dimension-token scaling

The analyzed thermodynamic regime uses d/N=1, equivalently d=N.

0.8
Macroscopic-cluster formation sharpness

At beta=0.8, simulations show broad alignment followed by tight macroscopic-cluster formation.

log N
Static REM sharpness scaling

The static random-energy-model benchmark requires beta scaling as log N, unlike the dynamical transition at beta=O(1).

What the paper found

This paper analyzes a minimal normalized self-attention system as a dynamical process in which token similarities create attention weights and those weights reshape the representations. In the scaling regime d=N, the central quantity is the overlap gap: when tokens are more similar to members of their own cluster than to any competing cluster by a finite amount, inter-cluster attention is exponentially suppressed as dimension grows. The result is a high-dimensional manifold of locally attracting clustered fixed points, spanning a few macroscopic groups to an extensive number of microscopic fragments; internal distortions contract, while collective cluster rotations remain neutral. Starting from normalized Gaussian vectors, the dynamics exhibit a finite-sharpness attention-condensation transition: low beta produces diffuse attention and rank collapse, intermediate beta yields a macroscopic cluster coexisting with condensed microscopic groups, and higher beta produces broadly distributed microscopic fragmentation. For a nearly rank-collapsed parent cluster, the control parameter is effective sharpness alpha=beta epsilon^2; at beta=0.8, simulations show broad collective alignment followed by tight-cluster formation. At finite N, weak leakage causes eventual coarsening, but fragmented-state lifetimes grow exponentially with beta sqrt(N), making the order of limits essential. Unlike static random-energy-model softmax, which requires beta scaling as log N, feedback dynamically generates finite overlap gaps and enables condensation at beta=O(1). The analysis also explains why linear attention does not condense through the same mechanism and suggests implications for architectures such as ChatGPT and other transformers; the manuscript notes that OpenAI’s ChatGPT, specifically GPT-5.5, assisted with drafting and exploratory calculations.

Original abstract

Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis