Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
AuthorsArman Adibi, Alireza Jafari, Mohammad Ghavamzadeh, Hadi Daneshmand
Resources
This work argues that transformers can generate data in context by internally simulating diffusion-like sampling processes.
Key results
GPT-2-style transformer layers used in the two-moons and smile experiments
Embedding width of the controlled GPT-2-style model
What the paper found
This paper expands in-context learning from prediction to generation by asking whether a frozen transformer can act as a sampler: given independent samples from an unknown distribution, can it produce a new sample without parameter updates or an estimated score network? The theory answers yes. Standard softmax attention computes Gaussian-mixture responsibility weights and their weighted empirical mean, while feedforward layers perform Euler updates, allowing a transformer to implement closed-form diffusion exactly, including a smoothed variant that reduces memorization. A second theorem shows approximation of a finite Estimation-Free Sampling procedure, combining particle gradient descent with reverse-time transport, so transformers can execute both score-based and estimation-free samplers in context. Experiments with a GPT-2-style model show compositional out-of-distribution generation: training excludes a smile-shaped component, yet smile points in the prompt lead to new samples along that unseen curve. In a 16-layer, 32-dimensional model trained on two-moons and related point clouds, hidden states first spread toward a more uniform geometry and later recover the prompt’s structured distribution. The same U-shaped layerwise pattern appears in normalized embeddings from OpenAI’s GPT-2, Meta’s Llama-3.3-70B-Instruct, Cerebras-GPT, and Qwen2.5, measured with RBF MMD2 and an interacting-particle energy. However, the authors emphasize an expressivity-versus-mechanism gap: these theorems prove that transformers can implement the algorithms, not that pretrained models literally do so, and smaller Qwen2.5 variants show a weaker intermediate transport phase.
Original abstract
A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.