xHC: Expanded Hyper-Connections
AuthorsXiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan
Resources
xHC expands Transformer memory through many residual streams while using sparse updates and optimized memory traffic to make large-scale LLM training more efficient.
Key results
Default xHC residual-stream expansion rate.
Only four of the 16 streams receive sparse residual mixing and write-back.
xHC downstream average, compared with 44.8 for mHC.
xHC downstream average, compared with 50.5 for mHC.
Compute required by the vanilla baseline relative to xHC.
Per-sublayer traffic in units of hidden dimension C, reduced from 73.5C for full xHC.
What the paper found
“xHC: Expanded Hyper-Connections,” from researchers at Shanghai Jiao Tong University, Xiaohongshu Inc., Peking University, USTC, and CUHK, extends Hyper-Connections beyond the usual four residual streams. The authors identify two mHC bottlenecks: write-back carries too little diverse information as stream count grows, while dense residual-mixing generation scales as O(N³C). xHC addresses these with temporal feature augmentation—three causal depthwise convolutions using kernel sizes 4, 8, and 12 after MoE MLP layers—plus sparse updates: it maintains N=16 streams but routes writes and residual mixing through only k=4 active streams while preserving dense reads. On DeepSeekMoE-style 18B and 28B MoE models, xHC raises average downstream scores from 44.8 to 48.8 and from 50.5 to 53.6, respectively, across MMLU, BBH, GSM8K, HumanEval, CMMLU, and related benchmarks. Scaling-law fits show vanilla and mHC require 1.50× and 1.19× xHC’s compute to reach the same loss. For deployment, xHC-Flash amortizes full-state operations across sublayers, cutting memory traffic from 73.5C to 40C with xHC-Flash-4sub while retaining nearly all performance; the method also remains effective with the Muon optimizer.
Original abstract
Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at $N{=}4$. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasing training cost. We attribute this limitation to two bottlenecks: insufficient write-back information for an expanding number of streams and residual-mixing generation whose cost scales cubically with $N$. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC combines temporal feature augmentation for richer write-back with a sparse residual-stream architecture that updates only $k=4$ of the $N=16$ streams while retaining dense access to the full residual state. Across 18B and 28B MoE models, xHC delivers strong and consistent downstream improvements. On an 18B MoE model, xHC improves the average downstream score by 4.0 points over mHC, while adding only modest training FLOPs over the vanilla baseline. Scaling-law experiments show that the vanilla and mHC require $1.50\times$ and $1.19\times$ the compute of xHC, respectively, to reach the same loss. Practical large-$N$ training also requires controlling memory traffic from the expanded residual state. We therefore introduce xHC-Flash, which reduces the per-sublayer memory traffic from $73.5C$ to $40C$, comparable to the $34C$ required by mHC at $N{=}4$, while retaining the gains of full xHC. Together, xHC and xHC-Flash make large-$N$ residual-stream expansion effective and practical for LLM pre-training.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.