NTH

Simplex Diffusion Models

AuthorsJustin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

AffiliationsGoogle DeepMind · EPFL · UCL Gatsby

October 1, 2026 2 min read
Watch on YouTube
The one-line take

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Key results

11.8M
TinyGSM dataset size

Synthetic math problems used to train the code-generation models.

49.0%
TinyGSM SDM accuracy

Accuracy at 512 steps without Self-Conditioning, compared with 45.8% for masked diffusion with Self-Conditioning.

32.1%
Distilled SDM GSM8K accuracy

GSM8K accuracy after distillation to 8 sampling steps.

21.4%
Distilled discrete diffusion GSM8K accuracy

Comparison result using 128 sampling steps.

99.1%
Sudoku accuracy

SDM accuracy with Self-Conditioning in 180 sampling steps.

17.0
OpenWebText GenPPL

Generative perplexity at 5.46 unigram entropy under GPT-2-large evaluation.

What the paper found

Simplex Diffusion Models address information collapse in discrete diffusion by representing every intermediate token state as a probability vector on the simplex, rather than repeatedly sampling a single category. A Dirichlet forward process preserves uncertainty, while a closed-form reverse bridge enables a DDIM-like sampler with tunable stochasticity through the churn parameter, avoiding the ordinary differential equation integration required by Dirichlet Flow Matching. Training uses a standard cross-entropy objective, and the framework connects discrete and continuous diffusion: high temperature approaches categorical diffusion, while low temperature approaches deterministic interpolation with Gaussian fluctuations. On TinyGSM, an 11.8M-example synthetic math dataset generated with GPT-3.5 and evaluated on GSM8K using the SmolLM tokenizer, SDMs reached 49.0% accuracy at 512 steps without Self-Conditioning, versus 45.8% for masked diffusion with Self-Conditioning. Distillation reduced generation to 8 steps while solving 32.1% of GSM8K problems, exceeding distilled discrete diffusion’s 21.4% at 128 steps. On Sudoku, SDMs with Self-Conditioning achieved 99.1% accuracy, and on OpenWebText they reached 17.0 GenPPL at 5.46 unigram entropy under GPT-2-large evaluation, close to real validation statistics. The results suggest that continuous belief states can improve discrete generation, although experiments remain limited to relatively small models.

Original abstract

Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis