NTH

Ask Self, Ask Others: Relation Is All You Need

AuthorsYuting Ge, Pengju Yang, Mingkai Nie

August 23, 2026 2 min read
Watch on YouTube
The one-line take

This paper replaces conventional attention with explicit Self-and-Exchange relations, aiming to preserve language-model quality while improving efficiency.

Key results

0.0412
10M NLL improvement

Full Relation lowers final validation NLL versus MHA by 0.0412 at 10M parameters.

0.0151
30M NLL improvement

Full Relation lowers final validation NLL versus MHA by 0.0151 at 30M parameters.

0.0310
100M NLL improvement

Full Relation lowers final validation NLL versus MHA by 0.0310 at 100M parameters.

4.41
FlashRelation speedup

The fixed-context speedup range is 3.60–4.41× versus materialized Full Relation.

84.9%
FlashAttention throughput fraction

FlashRelation reaches 76.4–84.9% of PyTorch FlashAttention throughput.

75%
Hybrid Linear layer proportion

Hybrid Relation uses 75% Linear Relation layers, or nine of twelve layers.

What the paper found

This paper proposes Relation, a token-mixing primitive that changes the Transformer order from scores directly determining attention flow to relation being constructed first and flow derived afterward. Its Self–Exchange Relation separates each token’s self-evidence from exchange evidence with prior tokens: Self uses a bounded sigmoid mapping, Exchange uses SiLU, and a causal Relation matrix is normalized only after a logarithmic count correction. Multi-Head Relation adds alternating adjacent-head Givens rotations to mix information states. In matched decoder-only models trained on TinyStories and the SmolLM corpus at 10M, 30M, and 100M parameters, Full Relation reduces final validation NLL versus MHA by 0.0412, 0.0151, and 0.0310, respectively. FlashRelation is an exact tiled implementation that avoids materializing the full relation and flow matrices; at context length 1024 it is 3.60–4.41× faster than the materialized reference and reaches 76.4–84.9% of PyTorch FlashAttention throughput on scale-matched workloads using an NVIDIA RTX 5090. Linear Relation replaces token-wise history with a recurrent state, reducing token-mixing complexity to O(Td²/H) and enabling fixed-size decode history, while Hybrid Relation combines nine Linear Relation layers with three Full Relation layers—75% Linear Relation—and achieves an NLL of 1.2780 in a 31.97M-parameter model. The result is a relation-first alternative to standard attention, complementary to systems such as FlashAttention and architectures such as Kimi Linear, though experiments remain limited to models up to approximately 100M parameters.

Original abstract

Attention directly derives normalized information flow from pairwise scores. We introduce Relation, an alternative token-mixing primitive that first organizes pairwise evidence into explicit Self and Exchange relations and derives information flow afterward. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and a KV-style Relation Cache. Across matched decoder-only models at approximately 10M, 30M, and 100M parameters, Full Relation achieves lower final validation NLL than MHA at all three scales. In a fixed-context reference benchmark, FlashRelation is 3.60-4.41x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. Hybrid Relation uses 75% Linear Relation layers and achieves strong language-modeling quality. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis