Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention
AuthorsQucheng Gao, Zuyi Yang, Xiao Chen
Resources
This paper argues that self-attention can spontaneously organize token representations into stable clusters once attention becomes sufficiently sharp.
Key results
The analyzed thermodynamic regime uses d/N=1, equivalently d=N.
At beta=0.8, simulations show broad alignment followed by tight macroscopic-cluster formation.
The static random-energy-model benchmark requires beta scaling as log N, unlike the dynamical transition at beta=O(1).
What the paper found
This paper analyzes a minimal normalized self-attention system as a dynamical process in which token similarities create attention weights and those weights reshape the representations. In the scaling regime d=N, the central quantity is the overlap gap: when tokens are more similar to members of their own cluster than to any competing cluster by a finite amount, inter-cluster attention is exponentially suppressed as dimension grows. The result is a high-dimensional manifold of locally attracting clustered fixed points, spanning a few macroscopic groups to an extensive number of microscopic fragments; internal distortions contract, while collective cluster rotations remain neutral. Starting from normalized Gaussian vectors, the dynamics exhibit a finite-sharpness attention-condensation transition: low beta produces diffuse attention and rank collapse, intermediate beta yields a macroscopic cluster coexisting with condensed microscopic groups, and higher beta produces broadly distributed microscopic fragmentation. For a nearly rank-collapsed parent cluster, the control parameter is effective sharpness alpha=beta epsilon^2; at beta=0.8, simulations show broad collective alignment followed by tight-cluster formation. At finite N, weak leakage causes eventual coarsening, but fragmented-state lifetimes grow exponentially with beta sqrt(N), making the order of limits essential. Unlike static random-energy-model softmax, which requires beta scaling as log N, feedback dynamically generates finite overlap gaps and enables condensation at beta=O(1). The analysis also explains why linear attention does not condense through the same mechanism and suggests implications for architectures such as ChatGPT and other transformers; the manuscript notes that OpenAI’s ChatGPT, specifically GPT-5.5, assisted with drafting and exploratory calculations.
Original abstract
Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.