Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
AuthorsBurc Gokden
Resources
This work argues that a power-law, graph-based form of attention can mathematically subsume standard attention and may effectively collapse to it during inference.
Key results
Measured relative fluctuations reached 0.000001 to 0.00000000001 on near-critical workloads.
The strongest audited checkpoint was trained on 41B tokens.
The generator A had numerical rank 1 across the audited instances.
ALM had median numerical rank 62.5 of 64 despite floating-point zero determinants.
Caching the collapsed operator produced a reported 3× speedup.
Scoring protocols agreed within 0.00005 per item on sampled TruthfulQA evaluations.
What the paper found
This paper formalizes Power Law Graph Attention, or PLGA, as an input-generated bilinear generalization of scaled dot-product attention: a query Gram matrix passes through a shared residual metric learner, a strictly positive tensor, an elementwise power law, and a learned operator G. Standard attention is recovered exactly when G equals the identity, while rotary position embeddings impose a commutant condition for preserving relative-position dependence. The central empirical finding is inference collapse: on an audited 110M-parameter checkpoint trained on 41B tokens, deductive tensors fluctuate only from 0.000001 to 0.00000000001 relatively, the generator A has rank 1, and its learned operator can be cached, producing a reported 3× speedup. However, the nonlinear positive tensor ALM remains near full rank, with median numerical rank 62.5 of 64, so floating-point zero determinants do not imply low rank. The proposed explanation combines rotary averaging, statistical concentration, and strong contraction in the learned row map, but only the algebraic components are proved; criticality and self-organized criticality remain hypotheses. On sampled TruthfulQA items, sequential and one-pass block scoring agreed within 0.00005 per item, while the paper cautions that its worst-case perturbation bound is far too loose to certify decoding margins. The work also discloses using Anthropic's Claude Fable 5 and OpenAI's GPT-5.6-Sol for drafting and verification, not as experimental models.
Original abstract
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at $G_{LM}=I$; $A_{LM}$ and $A_P$ are strictly entrywise positive, with Perron-Frobenius structure on $A_{LM}$; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of $10^{-6}$ and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within $5\times 10^{-5}$ per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.