Graph Machine: Towards Better Pretraining via Edges
AuthorsLintai Hou
Resources
Graph Machine replaces much of a Transformer with dynamically routed pointer-like edges, achieving near-preserved language-modeling quality while retrieving only a few tokens per head.
Key results
Share of Qwen3-0.6B dense Transformer layers replaced with GM sparse layers.
Tokens used to pretrain the Graph Language Machines from scratch.
Dense causal KV-access fraction when retrieving 2 of 4,096 positions per KV head.
Dense causal KV-access fraction when retrieving 4 of 4,096 positions per KV head.
Final test-loss reduction relative to Qwen3.
Reduction achieved by Hyperion-K16-R3-S relative to Qwen3.
What the paper found
Graph Machine, or GM, proposes a fourth sequence-modeling regime: it keeps an O(n) state like a Transformer but uses dynamically addressed sparse access instead of scanning the full history. Its state stores node features plus differentiable, pointer-like edge indices and weights. A referral mechanism composes multi-hop neighborhoods by approximating sparse adjacency-matrix powers, while Sparse Edge Attention combines edge weights with query–key scores as a product of experts. In a controlled pretraining study, researchers replaced 75% of the dense Transformer layers in Qwen3-0.6B with GM layers and trained from scratch on 15.7B tokens from FineWeb-Edu. At a 4,096-token context, retrieving only 2 or 4 positions per KV head corresponds to 0.098% or 0.195% of dense causal KV access. The strongest configuration, Hyperion-K16-R3-S, reduced final test loss by 0.003 relative to Qwen3 while using 19% less referral-plus-attention compute; the result suggests that most dense attention can be replaced by learned relational traversal at this scale. The main caveat is systems efficiency: the prototype was several times slower than Qwen3 on an H100, so custom kernels and larger-scale evaluations remain necessary.
Original abstract
We introduce the Graph Machine (GM), an architecture that maintains an $O(n)$-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves $O(n)$ complexity in its sparse layers without restricting the potentially accessible state size to $O(1)$. Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best model marginally improves loss.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.