NTH
Research collection

Transformers research

Explore transformer architectures, training methods, and sequence modeling. Follow changes to attention, scaling, and computational efficiency.

42 papers · Latest edition October 8, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Transformers papers

Newest editions first.

03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis
05Transformer

Trading Depth for Time in Recurrent Transformers

Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen

This paper shows that letting a Transformer spend extra computation on latent thought steps can approach the quality of doubling its depth while using substantially fewer parameters.

Read analysis
15Transformer

Time-Aware Tranformer-Based Prediction Model for AECOPD

Weihao Qu, Ling Zheng, Dongyang Wang, Jiacun Wang, Haowen Pan

A time-aware Transformer uses daily ventilator data to predict COPD flare-ups earlier without waiting for delayed clinical measurements.

Read analysis
16Transformer

Algebraic Decomposition Theory for Transformer Length Generalization

Andy Yang, Blerta Veseli, Corentin Barloy, Michaël Cadilhac, Andreas Krebs, Charles Paperman, Howard Straubing, Michael Hahn

This paper develops a mathematical theory explaining exactly when transformers can recognize patterns longer than those seen during training.

Read analysis
17Transformer

Ask Self, Ask Others: Relation Is All You Need

Yuting Ge, Pengju Yang, Mingkai Nie

This paper replaces conventional attention with explicit Self-and-Exchange relations, aiming to preserve language-model quality while improving efficiency.

Read analysis
18Transformer

Maglev: Sliding Recurrent Memory

Bo Liu, Qiang Liu

Maglev teaches a lightweight recurrent Transformer to preserve long-range context in a compact memory, enabling faster long-context generation without full attention at inference.

Read analysis
20Transformer

Full-bandwidth transformer

Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford

A transformer that feeds its hidden thoughts back into the next decoding step appears to gain better performance and shorter reasoning with little extra cost.

Read analysis
22Transformer

Multi-Head Attention Residuals

Cheng Luo, Zefan Cai, Junjie Hu

MHAR lets different feature groups retrieve different layers of a transformer's history, improving language-model quality with little added cost.

Read analysis
25Transformer

A Controlled Study of Attention-Only Transformers

Henry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee

A large controlled study finds that attention alone can nearly reproduce transformer performance, though feed-forward layers still provide an advantage for recalling knowledge stored in the weights.

Read analysis
27Transformer

xHC: Expanded Hyper-Connections

Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan

xHC expands Transformer memory through many residual streams while using sparse updates and optimized memory traffic to make large-scale LLM training more efficient.

Read analysis
28Transformer

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang

DeepLoop makes Transformers deeper by reusing the same blocks repeatedly and adjusts residual scaling to keep this recurrent computation stable.

Read analysis
29Transformer

Explaining Attention with Program Synthesis

Amiri Hayes, Belinda Li, Jacob Andreas

This paper turns transformer attention heads into small executable Python programs, offering a new way to explain and partially replace opaque neural behavior with human-readable code.

Read analysis
30Transformer

Chiaroscuro Attention: Spending Compute in the Dark

Prateek Kumar Sikdar

This paper asks when transformers really need full attention, and shows that a hybrid of spectral mixing and attention can cut compute while outperforming a strong baseline on some text tasks.

Read analysis
31Transformer

Rethinking the Role of Efficient Attention in Hybrid Architectures

Ziqing Qiao, Yinuo Xu, Chaojun Xiao, Zhou Su, Zihan Zhou, Yingfa Chen, Xiaoyue Xu, Xu Han, Zhiyuan Liu

This paper shows that in hybrid language models, efficient attention mostly shapes how quickly long-context skills emerge, while full attention does the heavy lifting for retrieval.

Read analysis
32Transformer

Variable-Width Transformers

Zhaofeng Wu, Oliver Sieberling, Shawn Tan, Rameswar Panda, Yury Polyanskiy, Yoon Kim

This paper proposes making transformers wide at the edges and narrow in the middle, showing that smarter width allocation can improve language model efficiency and performance.

Read analysis
33Transformer

Forget Attention: Importance-Aware Attention Is All You Need

Soohyeong Shin, Yeongwook Yang

SISA is a new attention design that lets state-space models guide which tokens matter inside softmax attention, aiming to combine global retrieval with better prioritization in language models.

Read analysis
34Transformer

Functional Attention: From Pairwise Affinities to Functional Correspondences

Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers

This paper turns attention from pairwise token matching into a functional map over continuous fields, aiming to make transformer-style operator learning more compact, resolution-invariant, and better suited for PDEs and other scientific tasks.

Read analysis
35Transformer

Sparsely gated tiny linear experts

Simon Schug

This paper makes transformer feedforward layers both smaller and more interpretable by replacing dense expert blocks with sparsely selected single-neuron linear experts.

Read analysis
37Transformer

Parallax: Parameterized Local Linear Attention for Language Modeling

Yifei Zuo, Dhruv Pai, Zhichen Zeng, Alec Dewulf, Shuming Hu, Zhaoran Wang

Parallax is a new attention design for language models that makes local linear attention practical at scale and claims better training quality and faster decoding than FlashAttention-style baselines.

Read analysis
39Transformer

Transformers Can Learn Posterior Predictive Distributions In-Context

Gyeonghun Kang, Changwoo J. Lee, Xiang Cheng

This paper shows that transformers can go beyond point estimates and learn full Bayesian predictive distributions in context, with theory explaining why architecture choices like normalization and attention depth matter.

Read analysis
40Transformer

How Many Different Outputs Can a Transformer Generate?

Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan

This paper shows that transformers can only generate a surprisingly limited set of outputs, with the reachable sequence space growing linearly with prompt length and shrinking sharply beyond a threshold.

Read analysis
42Transformer

LT2: Linear-Time Looped Transformers

Chunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, Hanjie Chen

This paper makes looped transformers much cheaper by swapping in linear or sparse attention, showing you can keep quality while dramatically cutting compute.

Read analysis