Representation Distribution Matching for One-Step Visual Generation
AuthorsLan Feng, Wuyang Li, Eloi Zablocki, Matthieu Cord, Alexandre Alahi
Resources
This paper shows how to train strong one-step image generators by matching feature distributions from multiple frozen encoders, achieving state-of-the-art quality and even compressing a four-step model into a single step.
Key results
iRDM state-of-the-art one-step ImageNet result
previous best one-step generator on the same metric
iRDM versus the prior best one-step generator
FLUX.2 [klein] post-trained one-step model
original four-step FLUX.2 [klein]
FLUX.2 [klein] post-trained one-step model
What the paper found
Representation Distribution Matching, or RDM, reframes one-step visual generation as direct distribution matching in frozen pretrained feature spaces, and the paper’s key result is that the classical MMD becomes highly effective once it is estimated correctly: the generated side should use large fresh batches above 2048, while the data side should be frozen once with a 4096-landmark Nyström reference over the full 1.28M-image ImageNet training set. The authors show that matching a single encoder is gameable, so improved RDM balances a diverse battery of 14 encoders with constrained proportional Lagrangian weighting and evaluates with SWr14, an independent sliced-Wasserstein metric over the same panel. On ImageNet-256, iRDM reaches a new one-step state of the art at SWr14 1.30, improving over the prior best 2.05, and PickScore prefers it over the previous best one-step generator on 71.2% of matched samples. The same recipe post-trains Black Forest Labs’ FLUX.2 [klein], a 4B four-step text-to-image model, into a one-step generator that in 90 H200 GPU-hours surpasses the original on GenEval, 0.826 versus 0.794, and on PickScore, 22.76 versus 22.58, with the joint image-text objective clearly outperforming an image-marginal ablation.
Original abstract
We elucidate the design space of Representation Distribution Matching (RDM), our name for the paradigm that trains a one-step image generator by matching generated and reference feature distributions under frozen pretrained encoders. We identify two design axes, how the distributions are compared and the representations they are compared in, and controlled studies along them yield three findings. First, the classical MMD, which could not train convincing generators a decade ago, becomes a strong and scalable objective once estimated right. Second, the generated batch is then the operative variable, with an optimum above 2048, far beyond customary batch sizes. Third, any single representation can be gamed, driven below the real score while images stay visibly fake, so we match against a balanced battery of encoders and evaluate with SW_r14, a Sliced-Wasserstein distance over 14 encoders that is independent of the training loss and resists gaming. Combining the preferred choices yields improved RDM (iRDM): it sets the one-step state of the art on ImageNet at SW_r14 1.30, corroborated by PickScore, a human-preference proxy our objective never optimizes, which prefers it over the prior best one-step generator on 71.2% of matched samples. The same recipe post-trains the four-step FLUX.2 [klein] into a one-step generator, surpassing the four-step version on GenEval, 0.826 to 0.794, and on PickScore, 22.76 to 22.58, in 90 H200 GPU-hours. Project page: https://alan-lanfeng.github.io/rdm/.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.