NTH

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

AuthorsJaehee Seo, Jisu Kim

AffiliationsDepartment of Statistics · Seoul National University, Korea

September 28, 2026 2 min read
Watch on YouTube
The one-line take

This work shows, in theory, how transformers can adapt to data living on locally different geometric structures and still achieve statistically optimal in-context prediction.

Key results

64
Ambient dimension

Synthetic heterogeneous-manifold benchmark ambient dimension.

40000
Context size

Number of in-context observations per prompt.

12
Mixture components

Heterogeneous manifold components in the benchmark.

1000000
Pretraining prompts

Training-task budget for the strongest reported checkpoint.

0.505
Transformer test MSE

Latent-target MSE in units of 10^-3 at the 1000000-prompt budget.

26.75%
Improvement over LP-CV

Transformer MSE reduction relative to the geometry-informed LP-CV benchmark.

What the paper found

This paper asks whether in-context learning can remain statistically optimal when data lie on a growing mixture of locally different manifolds rather than a single Euclidean space. It models components with varying intrinsic dimensions, Hölder smoothness, sampling masses, separation, reach, and bounded covariate noise, then derives an aggregate minimax risk that preserves each component’s effective sample size. The proposed estimator first fits a higher-order local tangent graph, projects observations into intrinsic coordinates, and applies chartwise local-polynomial regression. A structure-informed softmax transformer with ReLU feed-forward networks compiles this geometry-first procedure into one forward pass, achieving approximation error smaller than any fixed inverse polynomial, logarithmic depth, and polynomial resource growth. The theory also separates in-context error, transformer approximation, finite-task meta-training, and optimization error, showing that near empirical-risk minimization reaches the same aggregate rate with sufficient training tasks. This provides a mathematical account of the kind of few-shot adaptation associated with ChatGPT and other large language models, without analyzing a particular commercial model. In a synthetic benchmark with ambient dimension 64, context size 40000, and 12 heterogeneous components, an empirical softmax workspace predictor trained on 1000000 prompts achieved test MSE 0.505 in units of 10^-3, 26.75% lower than the geometry-informed LP-CV benchmark; the model used 286401 trainable parameters.

Original abstract

Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis