Universal interpolation for deep residual self-attention networks
AuthorsSibylle Marcotte, Joan Bruna
AffiliationsDepartment of Computer Science · New York University
This work shows that remarkably small, frozen attention systems can still transform arbitrary token sequences into one another simply by choosing how—and how long—to apply them.
Key results
Gaussian-initialized noncausal single-head blocks suffice when r≥d, on a full-measure connected domain.
The two-block universal interpolation theorem assumes d≥3.
What the paper found
This paper proves that deep residual self-attention can exactly interpolate sequence batches even when its attention parameters are frozen in advance: fitting a task requires choosing only which blocks to apply, their signs, and their durations. The central result shows that two Gaussian-initialized, noncausal single-head blocks suffice for every batch size and sequence length, provided the head width r is at least the token dimension d. With probability one, the pair works on an open, connected set covering all but a measure-zero portion of the admissible inputs, at both continuous and finite depth. A separate result gives universal interpolation across the full admissible domain using 2Nnd frozen blocks. The analysis uses Lie brackets and control-theoretic rank conditions to show how composing fixed transformations creates the directions needed to reach any target. Causal masking imposes a sharp limitation: the first tokens evolve through a shared linear map, preserving their linear relations, so arbitrary interpolation fails when the batch contains more than d sequences; interpolation is recovered when input and target share those relations. The main noncausal theorem assumes d≥3. The paper also discloses that OpenAI GPT-6 Astra assisted with writing, code, and proofs.
Original abstract
Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to any other collection of $N$ sequences of $n$ tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.
Read the original paperMore in Attention Mechanisms
Browse all 21 papers →Can Computation from Earlier Problems Help LLMs Solve New Ones?
Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
STAIR helps language models reuse useful computation from earlier questions, improving performance on later problems with only a tiny trainable module.
Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
Xiaoran Liu, Ziwei He, Xipeng Qiu
This paper explains why different hybrid attention designs succeed or fail at long-context modeling and introduces a method that extends context length efficiently without additional training.
CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.