NTH

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

AuthorsYushi Huang, Xiangxin Zhou, Ruoyu Wang, Chi Zhang, Jun Zhang, Tianyu Pang

May 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes fast image generators better at following human preferences by blending distillation and reinforcement learning into a new two-stage training recipe.

Key results

0.3161
SD3-M CLIPScore

RTDMD achieves this score on Stable Diffusion 3-Medium with 4-step generation, outperforming prior few-step baselines.

22.86
SD3-M PickScore

RTDMD reaches this PickScore on Stable Diffusion 3-Medium with 4-step generation.

0.3211
SD3-M HPSv2

RTDMD reaches this HPSv2 score on Stable Diffusion 3-Medium with 4-step generation.

5.9642
SD3-M Aesthetic Score

RTDMD achieves this Aesthetic Score on Stable Diffusion 3-Medium and exceeds the teacher on this metric.

1.3024
SD3-M ImageReward

RTDMD achieves this ImageReward score on Stable Diffusion 3-Medium and exceeds the teacher on this metric.

What the paper found

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching introduces RTDMD, a two-stage method for aligning 4-step flow-based text-to-image generators with human preferences while preserving the teacher prior. The key theoretical move is to minimize KL divergence to a reward-tilted teacher distribution p̃ψ(x) ∝ pψ(x)exp(βr(x)), which cleanly decomposes into distribution matching plus reward maximization. In the cold-start stage, Ambient-Consistent Distribution Matching Distillation (AC-DMD) re-derives DMD on subintervals [tk,1] for noisy intermediate latents produced by coefficient-preserving sampling (CPS), then stabilizes the fake score model with a consistency regularizer from Consistent Diffusion. In the reinforcement stage, the paper derives a hybrid policy gradient for the mixed stochastic-deterministic sampler: GRPO-style updates for the K−1 noisy transitions, direct backpropagation through the final deterministic step, and a variance-reduced step-subset GRPO (SubGRPO) estimator with shared noise. On Stable Diffusion 3-Medium, RTDMD reaches CLIPScore 0.3161, PickScore 22.86, HPSv2 0.3211, Aesthetic Score 5.9642, and ImageReward 1.3024, outperforming prior few-step methods such as GDMD and Rdm; on FLUX.2 4B it achieves the best results on 7 of 9 metrics and even beats the full FLUX.2 9B on most benchmarks, showing that preference-aligned distillation can narrow the quality gap caused by aggressive step reduction.

Original abstract

Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward-Tilted Distribution Matching Distillation (RTDMD), a two-stage framework that unifies distribution matching distillation with reward-guided reinforcement learning for few-step flow generators. We show that minimizing the KL divergence to a reward-tilted teacher distribution naturally decomposes into a distribution matching term and a reward maximization term. In the first stage, we introduce Ambient-Consistent Distribution Matching Distillation (AC-DMD), which performs subinterval-wise distribution matching and augments the fake score objective with a consistency regularizer to help the fake score model track the shifting generator distribution under limited updates. In the second stage, we jointly optimize both terms: for the reward maximization term, we derive a hybrid policy gradient that combines a GRPO-style estimator for the stochastic intermediate transitions with direct reward backpropagation through the deterministic final step, and further introduce step-subset GRPO (SubGRPO) to reduce variance. Experiments on SD3, SD3.5, and FLUX.2 demonstrate that RTDMD establishes new state-of-the-art results across preference, aesthetic, and compositional metrics with only 4 inference steps, outperforming previous few-step text-to-image generation methods. Code and models are available at https://github.com/Harahan/RTDMD.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis