NTH

Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

AuthorsArman Adibi, Alireza Jafari, Mohammad Ghavamzadeh, Hadi Daneshmand

September 18, 2026 2 min read
Watch on YouTube
The one-line take

This work argues that transformers can generate data in context by internally simulating diffusion-like sampling processes.

Key results

16
Controlled sampler depth

GPT-2-style transformer layers used in the two-moons and smile experiments

32
Hidden embedding dimension

Embedding width of the controlled GPT-2-style model

What the paper found

This paper expands in-context learning from prediction to generation by asking whether a frozen transformer can act as a sampler: given independent samples from an unknown distribution, can it produce a new sample without parameter updates or an estimated score network? The theory answers yes. Standard softmax attention computes Gaussian-mixture responsibility weights and their weighted empirical mean, while feedforward layers perform Euler updates, allowing a transformer to implement closed-form diffusion exactly, including a smoothed variant that reduces memorization. A second theorem shows approximation of a finite Estimation-Free Sampling procedure, combining particle gradient descent with reverse-time transport, so transformers can execute both score-based and estimation-free samplers in context. Experiments with a GPT-2-style model show compositional out-of-distribution generation: training excludes a smile-shaped component, yet smile points in the prompt lead to new samples along that unseen curve. In a 16-layer, 32-dimensional model trained on two-moons and related point clouds, hidden states first spread toward a more uniform geometry and later recover the prompt’s structured distribution. The same U-shaped layerwise pattern appears in normalized embeddings from OpenAI’s GPT-2, Meta’s Llama-3.3-70B-Instruct, Cerebras-GPT, and Qwen2.5, measured with RBF MMD2 and an interacting-particle energy. However, the authors emphasize an expressivity-versus-mechanism gap: these theorems prove that transformers can implement the algorithms, not that pretrained models literally do so, and smaller Qwen2.5 variants show a weaker intermediate transport phase.

Original abstract

A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis