Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
AuthorsJaehee Seo, Jisu Kim
AffiliationsDepartment of Statistics · Seoul National University, Korea
Resources
This work shows, in theory, how transformers can adapt to data living on locally different geometric structures and still achieve statistically optimal in-context prediction.
Key results
Synthetic heterogeneous-manifold benchmark ambient dimension.
Number of in-context observations per prompt.
Heterogeneous manifold components in the benchmark.
Training-task budget for the strongest reported checkpoint.
Latent-target MSE in units of 10^-3 at the 1000000-prompt budget.
Transformer MSE reduction relative to the geometry-informed LP-CV benchmark.
What the paper found
This paper asks whether in-context learning can remain statistically optimal when data lie on a growing mixture of locally different manifolds rather than a single Euclidean space. It models components with varying intrinsic dimensions, Hölder smoothness, sampling masses, separation, reach, and bounded covariate noise, then derives an aggregate minimax risk that preserves each component’s effective sample size. The proposed estimator first fits a higher-order local tangent graph, projects observations into intrinsic coordinates, and applies chartwise local-polynomial regression. A structure-informed softmax transformer with ReLU feed-forward networks compiles this geometry-first procedure into one forward pass, achieving approximation error smaller than any fixed inverse polynomial, logarithmic depth, and polynomial resource growth. The theory also separates in-context error, transformer approximation, finite-task meta-training, and optimization error, showing that near empirical-risk minimization reaches the same aggregate rate with sufficient training tasks. This provides a mathematical account of the kind of few-shot adaptation associated with ChatGPT and other large language models, without analyzing a particular commercial model. In a synthetic benchmark with ambient dimension 64, context size 40000, and 12 heterogeneous components, an empirical softmax workspace predictor trained on 1000000 prompts achieved test MSE 0.505 in units of 10^-3, 26.75% lower than the geometry-informed LP-CV benchmark; the model used 286401 trainable parameters.
Original abstract
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.