Simplex Diffusion Models
AuthorsJustin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
AffiliationsGoogle DeepMind · EPFL · UCL Gatsby
Resources
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.
Key results
Synthetic math problems used to train the code-generation models.
Accuracy at 512 steps without Self-Conditioning, compared with 45.8% for masked diffusion with Self-Conditioning.
GSM8K accuracy after distillation to 8 sampling steps.
Comparison result using 128 sampling steps.
SDM accuracy with Self-Conditioning in 180 sampling steps.
Generative perplexity at 5.46 unigram entropy under GPT-2-large evaluation.
What the paper found
Simplex Diffusion Models address information collapse in discrete diffusion by representing every intermediate token state as a probability vector on the simplex, rather than repeatedly sampling a single category. A Dirichlet forward process preserves uncertainty, while a closed-form reverse bridge enables a DDIM-like sampler with tunable stochasticity through the churn parameter, avoiding the ordinary differential equation integration required by Dirichlet Flow Matching. Training uses a standard cross-entropy objective, and the framework connects discrete and continuous diffusion: high temperature approaches categorical diffusion, while low temperature approaches deterministic interpolation with Gaussian fluctuations. On TinyGSM, an 11.8M-example synthetic math dataset generated with GPT-3.5 and evaluated on GSM8K using the SmolLM tokenizer, SDMs reached 49.0% accuracy at 512 steps without Self-Conditioning, versus 45.8% for masked diffusion with Self-Conditioning. Distillation reduced generation to 8 steps while solving 32.1% of GSM8K problems, exceeding distilled discrete diffusion’s 21.4% at 128 steps. On Sudoku, SDMs with Self-Conditioning achieved 99.1% accuracy, and on OpenWebText they reached 17.0 GenPPL at 5.46 unigram entropy under GPT-2-large evaluation, close to real validation statistics. The results suggest that continuous belief states can improve discrete generation, although experiments remain limited to relatively small models.
Original abstract
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Register Tokens for Bounded-State Reasoning in Diffusion Language Models
Albert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala
The work gives diffusion language models a small set of learned memory tokens so they can preserve reasoning across multiple chunks without retaining the generated text.