NTH

Universal interpolation for deep residual self-attention networks

AuthorsSibylle Marcotte, Joan Bruna

AffiliationsDepartment of Computer Science · New York University

October 11, 2026 2 min read
Watch on YouTube
The one-line take

This work shows that remarkably small, frozen attention systems can still transform arbitrary token sequences into one another simply by choosing how—and how long—to apply them.

Key results

2
Frozen blocks in main theorem

Gaussian-initialized noncausal single-head blocks suffice when r≥d, on a full-measure connected domain.

3
Minimum token dimension in main theorem

The two-block universal interpolation theorem assumes d≥3.

What the paper found

This paper proves that deep residual self-attention can exactly interpolate sequence batches even when its attention parameters are frozen in advance: fitting a task requires choosing only which blocks to apply, their signs, and their durations. The central result shows that two Gaussian-initialized, noncausal single-head blocks suffice for every batch size and sequence length, provided the head width r is at least the token dimension d. With probability one, the pair works on an open, connected set covering all but a measure-zero portion of the admissible inputs, at both continuous and finite depth. A separate result gives universal interpolation across the full admissible domain using 2Nnd frozen blocks. The analysis uses Lie brackets and control-theoretic rank conditions to show how composing fixed transformations creates the directions needed to reach any target. Causal masking imposes a sharp limitation: the first tokens evolve through a shared linear map, preserving their linear relations, so arbitrary interpolation fails when the batch contains more than d sequences; interpolation is recovered when input and target share those relations. The main noncausal theorem assumes d≥3. The paper also discloses that OpenAI GPT-6 Astra assisted with writing, code, and proofs.

Original abstract

Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to any other collection of $N$ sequences of $n$ tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.

Read the original paper

More in Attention Mechanisms

Browse all 21 papers →
01Attention

Can Computation from Earlier Problems Help LLMs Solve New Ones?

Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song

STAIR helps language models reuse useful computation from earlier questions, improving performance on later problems with only a tiny trainable module.

Read analysis
03Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis