Ask Self, Ask Others: Relation Is All You Need
AuthorsYuting Ge, Pengju Yang, Mingkai Nie
Resources
This paper replaces conventional attention with explicit Self-and-Exchange relations, aiming to preserve language-model quality while improving efficiency.
Key results
Full Relation lowers final validation NLL versus MHA by 0.0412 at 10M parameters.
Full Relation lowers final validation NLL versus MHA by 0.0151 at 30M parameters.
Full Relation lowers final validation NLL versus MHA by 0.0310 at 100M parameters.
The fixed-context speedup range is 3.60–4.41× versus materialized Full Relation.
FlashRelation reaches 76.4–84.9% of PyTorch FlashAttention throughput.
Hybrid Relation uses 75% Linear Relation layers, or nine of twelve layers.
What the paper found
This paper proposes Relation, a token-mixing primitive that changes the Transformer order from scores directly determining attention flow to relation being constructed first and flow derived afterward. Its Self–Exchange Relation separates each token’s self-evidence from exchange evidence with prior tokens: Self uses a bounded sigmoid mapping, Exchange uses SiLU, and a causal Relation matrix is normalized only after a logarithmic count correction. Multi-Head Relation adds alternating adjacent-head Givens rotations to mix information states. In matched decoder-only models trained on TinyStories and the SmolLM corpus at 10M, 30M, and 100M parameters, Full Relation reduces final validation NLL versus MHA by 0.0412, 0.0151, and 0.0310, respectively. FlashRelation is an exact tiled implementation that avoids materializing the full relation and flow matrices; at context length 1024 it is 3.60–4.41× faster than the materialized reference and reaches 76.4–84.9% of PyTorch FlashAttention throughput on scale-matched workloads using an NVIDIA RTX 5090. Linear Relation replaces token-wise history with a recurrent state, reducing token-mixing complexity to O(Td²/H) and enabling fixed-size decode history, while Hybrid Relation combines nine Linear Relation layers with three Full Relation layers—75% Linear Relation—and achieves an NLL of 1.2780 in a 31.97M-parameter model. The result is a relation-first alternative to standard attention, complementary to systems such as FlashAttention and architectures such as Kimi Linear, though experiments remain limited to models up to approximately 100M parameters.
Original abstract
Attention directly derives normalized information flow from pairwise scores. We introduce Relation, an alternative token-mixing primitive that first organizes pairwise evidence into explicit Self and Exchange relations and derives information flow afterward. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and a KV-style Relation Cache. Across matched decoder-only models at approximately 10M, 30M, and 100M parameters, Full Relation achieves lower final validation NLL than MHA at all three scales. In a fixed-context reference benchmark, FlashRelation is 3.60-4.41x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. Hybrid Relation uses 75% Linear Relation layers and achieves strong language-modeling quality. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.