NTH

Graph Machine: Towards Better Pretraining via Edges

AuthorsLintai Hou

September 14, 2026 2 min read
Watch on YouTube
The one-line take

Graph Machine replaces much of a Transformer with dynamically routed pointer-like edges, achieving near-preserved language-modeling quality while retrieving only a few tokens per head.

Key results

75%
Dense layers replaced

Share of Qwen3-0.6B dense Transformer layers replaced with GM sparse layers.

15.7B
Pretraining tokens

Tokens used to pretrain the Graph Language Machines from scratch.

0.098%
Two-position KV access

Dense causal KV-access fraction when retrieving 2 of 4,096 positions per KV head.

0.195%
Four-position KV access

Dense causal KV-access fraction when retrieving 4 of 4,096 positions per KV head.

0.003
Hyperion-K16-R3-S loss change

Final test-loss reduction relative to Qwen3.

19%
Referral-plus-attention compute reduction

Reduction achieved by Hyperion-K16-R3-S relative to Qwen3.

What the paper found

Graph Machine, or GM, proposes a fourth sequence-modeling regime: it keeps an O(n) state like a Transformer but uses dynamically addressed sparse access instead of scanning the full history. Its state stores node features plus differentiable, pointer-like edge indices and weights. A referral mechanism composes multi-hop neighborhoods by approximating sparse adjacency-matrix powers, while Sparse Edge Attention combines edge weights with query–key scores as a product of experts. In a controlled pretraining study, researchers replaced 75% of the dense Transformer layers in Qwen3-0.6B with GM layers and trained from scratch on 15.7B tokens from FineWeb-Edu. At a 4,096-token context, retrieving only 2 or 4 positions per KV head corresponds to 0.098% or 0.195% of dense causal KV access. The strongest configuration, Hyperion-K16-R3-S, reduced final test loss by 0.003 relative to Qwen3 while using 19% less referral-plus-attention compute; the result suggests that most dense attention can be replaced by learned relational traversal at this scale. The main caveat is systems efficiency: the prototype was several times slower than Qwen3 on an H100, so custom kernels and larger-scale evaluations remain necessary.

Original abstract

We introduce the Graph Machine (GM), an architecture that maintains an $O(n)$-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves $O(n)$ complexity in its sparse layers without restricting the potentially accessible state size to $O(1)$. Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best model marginally improves loss.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis