Reinforcing Few-step Generators via Reward-Tilted Distribution Matching
AuthorsYushi Huang, Xiangxin Zhou, Ruoyu Wang, Chi Zhang, Jun Zhang, Tianyu Pang
This paper makes fast image generators better at following human preferences by blending distillation and reinforcement learning into a new two-stage training recipe.
Key results
RTDMD achieves this score on Stable Diffusion 3-Medium with 4-step generation, outperforming prior few-step baselines.
RTDMD reaches this PickScore on Stable Diffusion 3-Medium with 4-step generation.
RTDMD reaches this HPSv2 score on Stable Diffusion 3-Medium with 4-step generation.
RTDMD achieves this Aesthetic Score on Stable Diffusion 3-Medium and exceeds the teacher on this metric.
RTDMD achieves this ImageReward score on Stable Diffusion 3-Medium and exceeds the teacher on this metric.
What the paper found
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching introduces RTDMD, a two-stage method for aligning 4-step flow-based text-to-image generators with human preferences while preserving the teacher prior. The key theoretical move is to minimize KL divergence to a reward-tilted teacher distribution p̃ψ(x) ∝ pψ(x)exp(βr(x)), which cleanly decomposes into distribution matching plus reward maximization. In the cold-start stage, Ambient-Consistent Distribution Matching Distillation (AC-DMD) re-derives DMD on subintervals [tk,1] for noisy intermediate latents produced by coefficient-preserving sampling (CPS), then stabilizes the fake score model with a consistency regularizer from Consistent Diffusion. In the reinforcement stage, the paper derives a hybrid policy gradient for the mixed stochastic-deterministic sampler: GRPO-style updates for the K−1 noisy transitions, direct backpropagation through the final deterministic step, and a variance-reduced step-subset GRPO (SubGRPO) estimator with shared noise. On Stable Diffusion 3-Medium, RTDMD reaches CLIPScore 0.3161, PickScore 22.86, HPSv2 0.3211, Aesthetic Score 5.9642, and ImageReward 1.3024, outperforming prior few-step methods such as GDMD and Rdm; on FLUX.2 4B it achieves the best results on 7 of 9 metrics and even beats the full FLUX.2 9B on most benchmarks, showing that preference-aligned distillation can narrow the quality gap caused by aggressive step reduction.
Original abstract
Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward-Tilted Distribution Matching Distillation (RTDMD), a two-stage framework that unifies distribution matching distillation with reward-guided reinforcement learning for few-step flow generators. We show that minimizing the KL divergence to a reward-tilted teacher distribution naturally decomposes into a distribution matching term and a reward maximization term. In the first stage, we introduce Ambient-Consistent Distribution Matching Distillation (AC-DMD), which performs subinterval-wise distribution matching and augments the fake score objective with a consistency regularizer to help the fake score model track the shifting generator distribution under limited updates. In the second stage, we jointly optimize both terms: for the reward maximization term, we derive a hybrid policy gradient that combines a GRPO-style estimator for the stochastic intermediate transitions with direct reward backpropagation through the deterministic final step, and further introduce step-subset GRPO (SubGRPO) to reduce variance. Experiments on SD3, SD3.5, and FLUX.2 demonstrate that RTDMD establishes new state-of-the-art results across preference, aesthetic, and compositional metrics with only 4 inference steps, outperforming previous few-step text-to-image generation methods. Code and models are available at https://github.com/Harahan/RTDMD.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.