NTH

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

AuthorsBurc Gokden

August 15, 2026 3 min read
Watch on YouTube
The one-line take

This work argues that a power-law, graph-based form of attention can mathematically subsume standard attention and may effectively collapse to it during inference.

Key results

0.000001
Relative deductive-output fluctuation lower scale

Measured relative fluctuations reached 0.000001 to 0.00000000001 on near-critical workloads.

41B
Training token count

The strongest audited checkpoint was trained on 41B tokens.

1
Generator numerical rank

The generator A had numerical rank 1 across the audited instances.

62.5
ALM median numerical rank

ALM had median numerical rank 62.5 of 64 despite floating-point zero determinants.

3
Cached inference speedup

Caching the collapsed operator produced a reported 3× speedup.

0.00005
TruthfulQA block-versus-sequential tolerance

Scoring protocols agreed within 0.00005 per item on sampled TruthfulQA evaluations.

What the paper found

This paper formalizes Power Law Graph Attention, or PLGA, as an input-generated bilinear generalization of scaled dot-product attention: a query Gram matrix passes through a shared residual metric learner, a strictly positive tensor, an elementwise power law, and a learned operator G. Standard attention is recovered exactly when G equals the identity, while rotary position embeddings impose a commutant condition for preserving relative-position dependence. The central empirical finding is inference collapse: on an audited 110M-parameter checkpoint trained on 41B tokens, deductive tensors fluctuate only from 0.000001 to 0.00000000001 relatively, the generator A has rank 1, and its learned operator can be cached, producing a reported 3× speedup. However, the nonlinear positive tensor ALM remains near full rank, with median numerical rank 62.5 of 64, so floating-point zero determinants do not imply low rank. The proposed explanation combines rotary averaging, statistical concentration, and strong contraction in the learned row map, but only the algebraic components are proved; criticality and self-organized criticality remain hypotheses. On sampled TruthfulQA items, sequential and one-pass block scoring agreed within 0.00005 per item, while the paper cautions that its worst-case perturbation bound is far too loose to certify decoding margins. The work also discloses using Anthropic's Claude Fable 5 and OpenAI's GPT-5.6-Sol for drafting and verification, not as experimental models.

Original abstract

The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at $G_{LM}=I$; $A_{LM}$ and $A_P$ are strictly entrywise positive, with Perron-Frobenius structure on $A_{LM}$; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of $10^{-6}$ and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within $5\times 10^{-5}$ per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis