Functional Attention: From Pairwise Affinities to Functional Correspondences
AuthorsJiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers
This paper turns attention from pairwise token matching into a functional map over continuous fields, aiming to make transformer-style operator learning more compact, resolution-invariant, and better suited for PDEs and other scientific tasks.
Key results
Evaluated on six PDE benchmarks: Darcy, Airfoil, Pipe, Navier-Stokes, Elasticity, and Plasticity.
Reported improvement over Transolver across the PDE benchmarks.
Relative L2 error on the Darcy benchmark, improved from 0.57 for Transolver.
Relative L2 error on the Navier-Stokes benchmark, improved from 9.44 for Transolver.
3D RNA segmentation accuracy on xyz coordinates.
What the paper found
Functional Attention From Pairwise Affinities to Functional Correspondences, by Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, and Daniel Cremers, reframes transformer attention for operator learning as a linear map between learned function spaces rather than a softmax over discrete tokens. The method, called F UNC ATTN, borrows from the functional maps framework in geometry processing and replaces pairwise affinities with a compact k × k spectral transport operator estimated by Tikhonov-regularized least squares, using learned adaptive bases instead of fixed Fourier modes. This makes the attention layer resolution-invariant, globally structured, and theoretically stable; the paper proves local Lipschitz continuity with the regularization parameter λ controlling the bound. Empirically, on six PDE benchmarks including Darcy, Airfoil, Pipe, Navier-Stokes, Elasticity, and Plasticity, F UNC ATTN reaches state-of-the-art relative L2 errors on five tasks and improves over Transolver by 6% to 26.3%, for example dropping Darcy error from 0.57 to 0.42 and Navier-Stokes from 9.44 to 8.00. It also improves 3D RNA segmentation accuracy to 89.0% on xyz coordinates, beats prior methods on a difficult triangular-domain Darcy notch problem with 0.64% relative L2 error, and generalizes best in out-of-distribution AirfRANS airfoil tests, reaching 23.4% error under unseen Reynolds numbers and 13.3% under unseen angles of attack. The key novelty is that attention is computed as an optimal operator in a learned spectral space, which yields linear scaling in sequence length and strong robustness across discretizations and geometries.
Original abstract
Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although transformer-based operators are popular, they often rely on token-wise attention. These methods treat continuous fields as discrete tokens and usually ignore the global functional structure. We introduce \emph{Functional Attention}, which reinterprets attention as a functional correspondence between adaptive bases. Inspired by geometric functional maps, our method replaces softmax affinities with structured linear operators. This yields a compact, generalizable, resolution-invariant representation that explicitly captures global dependencies. Experiments demonstrate that \emph{Functional Attention} can match state-of-the-art performance in many operator learning tasks, including solving PDEs, 3D segmentation, and regression, while remaining robust to varying discretizations. Project page is available at https://github.com/xjffff/FUNCATTN.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.