Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion
AuthorsJaihoon Kim, Taehoon Yoon, Prin Phunyaphibarn, Seungjun Kim, Morteza Mardani, Minhyuk Sung
Resources
This paper speeds up reward-guided sampling in discrete diffusion models by learning a cheap twist function, making aligned generation more practical for text, DNA, proteins, and large language models.
Key results
A learned twist head attached to the diffusion backbone adds less than 5% additional computational overhead versus a single forward pass.
In some setups, evaluating the learned twist function adds as little as 0.5% runtime overhead.
Regulatory DNA sequence design is evaluated on an enhancer activity dataset consisting of 700,000 DNA sequences.
What the paper found
Contrastive Distribution Matching, or CDM, addresses a core bottleneck in discrete diffusion reward alignment: Twisted Sequential Monte Carlo needs an optimal twist function, but in discrete state spaces that twist is usually estimated by expensive Monte Carlo over clean samples, which makes inference slow, especially when the reward is costly to evaluate. The paper’s key idea is to amortize that cost by learning the twist with a contrastive forward-KL objective, so the learned twist can be evaluated in a single lightweight head attached to the diffusion backbone, adding less than 5% runtime overhead, as low as 0.5% in some setups. CDM replaces regression-style Soft Value training with positive and negative samples: positive samples come from the reward-tilted target, negative samples from the current approximation, and the gradient explicitly upweights high-reward regions while suppressing poor ones. A diffusion-native trick reuses clean positive samples through the closed-form forward kernel, enabling multiple updates per sample and reducing reward calls. Across toxic text generation on OpenWebText, regulatory DNA sequence design on a 700,000-sequence enhancer dataset, protein designability with DPLM-2 and ESMFold-based self-consistency, and diffusion LLM alignment with LLaDA-8B-Instruct on RewardBench, CDM consistently beats Best-of-N, SMC with M=1 or 4, and Soft Value under matched wall-clock time, while remaining compatible with fine-tuned proposals such as d1 and DRAKES and mitigating their mode collapse.
Original abstract
Discrete diffusion models have emerged as powerful frameworks for generating structured categorical data. However, efficiently sampling from reward-tilted distributions remains a fundamental challenge. While Twisted Sequential Monte Carlo (SMC) offers asymptotic exactness for this task, estimating the optimal twist function in discrete state spaces necessitates costly Monte Carlo approximations, resulting a severe computational bottleneck at inference. To overcome this limitation, we introduce Contrastive Distribution Matching (CDM), a novel framework that amortizes the cost of SMC inference by learning a parameterized twist function via positive and negative samples. For efficient training, we reformulate the gradient estimator to leverage the closed-form forward kernels of discrete diffusion models. In practice, evaluating our learned twist function incurs less than 5% additional computational overhead compared to a single forward pass of the base model. Through extensive empirical evaluations, we demonstrate that CDM consistently outperforms existing baselines under matched wall-clock time. We validate the effectiveness and versatility of our approach across a diverse range of applications, including toxic text generation, regulatory DNA sequence design, protein designability, and diffusion large language model alignment.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.