NTH

From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models

AuthorsYuchen Liang, Ness Shroff, Yingbin Liang

May 30, 2026 3 min read
Watch on YouTube
The one-line take

This paper speeds up discrete diffusion sampling by using Gibbs-style corrections, making text and music generation much faster without extra training.

Key results

O(polylog(ε−1))
GADD sampling complexity

The paper proves Gibbs-Accelerated Discrete Diffusion achieves logarithmic dependence on target TV accuracy ε for uniform-rate discrete diffusion models.

256
Music sequence length

The conditional music experiment uses sequence length d = 256.

What the paper found

From Scores to Gibbs Correctors introduces Gibbs-Accelerated Discrete Diffusion, or GADD, a predictor-corrector sampler for uniform-rate discrete diffusion models that converts the concrete score function directly into Gibbs posterior conditionals, eliminating any extra training beyond standard score estimation. The central novelty is theoretical: GADD is proved to achieve overall sampling complexity O(polylog(ε−1)), the first diffusion-based sampler for uniform-rate discrete models with logarithmic dependence on target TV accuracy ε, whereas prior deterministic samplers such as Euler, Tweedie τ-leaping, and higher-order Runge-Kutta variants remain polynomial in ε−1. The method alternates an Euler predictor with a random-scan Gibbs corrector whose posterior is computed from the score estimator via closed-form normalization, so one forward score pass yields all token-wise conditional probabilities. The analysis replaces Girsanov-based arguments with an induction framework that tracks error propagation across predictor steps and explicitly handles inaccurate correctors, then shows how the diffusion warm start reduces both Gibbs mixing error and score-estimation error. On synthetic spiky distributions, WikiText103 with a SEDD Uniform model, and zero-shot conditional music completion on Lakh pianoroll data, GADD consistently improves sample quality and wall-clock efficiency over vanilla Euler and CTMC correctors; for example, in text sampling at NFE 256 it lowers generative perplexity from 275.4 for vanilla Euler to 149.6, while also beating the θ-Trapezoidal sampler. The paper also proves that CTMC correctors still incur O(poly(ε−1)) complexity, highlighting why Gibbs-based correction is asymptotically superior in this setting.

Original abstract

Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, but, especially for uniform-rate models, they often require many steps to generate a single sample. Existing acceleration methods either rely on training additional quantities or suffer from slow mixing. In this work, we propose a novel Gibbs-based corrector for discrete diffusion models, termed Gibbs-Accelerated Discrete Diffusion (GADD). GADD leverages the structure of the concrete score function to construct Gibbs posterior likelihoods directly, without requiring any additional training beyond standard score estimation. We show that GADD achieves an overall sampling complexity of $\mathcal{O}(\mathrm{polylog} (\varepsilon^{-1}))$, yielding the first such rate for diffusion-based samplers for uniform-rate discrete diffusion models. We also conduct numerical experiments demonstrating the practical advantages of GADD across synthetic data, zero-shot text sampling, and zero-shot conditional music generation. These results corroborate the theory and show that GADD consistently improves sample quality and wall-clock efficiency over standard baselines, including vanilla Euler methods and CTMC correctors. Beyond this, our theoretical analysis introduces a novel framework for analyzing predictor-corrector methods in discrete diffusion models, which may be of independent interest. Unlike existing approaches that rely on the Girsanov change-of-measure technique, our method is based on an induction argument that tracks error propagation across predictor iterations while accounting for inaccuracies in the corrector updates.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis