MEND: RL For Flow Models via Proximal Velocity Matching
AuthorsShreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
AffiliationsThe University of Texas at Austin Google University of Colorado Boulder
MEND makes flow-model reward tuning more efficient by moving samples only when the reward gain justifies the size of the change.
Key results
SD3.5-M comparison against Flow-GRPO.
Approximate update count for the comparison baseline.
Equal-budget result; ReFL scores 23.92 and DiffusionNFT 23.43.
GPU-hours for the SD3.5-M run.
What the paper found
MEND changes how flow models learn from reward scores: rather than reweighting samples or moving every sample along a reward gradient, it caps rewards within each prompt group and proposes short gradient-directed moves only for samples below the cap. A move is accepted only when its capped reward gain exceeds a quadratic distance cost; the model then learns the accepted changes through proximal velocity matching, without a KL penalty, frozen reference model, or advantage weights. On SD3.5-M, MEND beats Flow-GRPO on five of six evaluators after 100 updates, compared with about 4k updates for Flow-GRPO, at the same distance from base-model images. In equal-budget tests, it reaches PickScore 24.03, ahead of ReFL at 23.92 and DiffusionNFT at 23.43. The method also improves SD3-M and Z-Image-Turbo, and the SD3.5-M run takes 10.0 GPU-hours. Its guarantees apply to selected training targets, not to held-out rewards: the paper notes that longer training can over-optimize a reward and reduce other measures.
Original abstract
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
Read the original paperMore in Generative Models
Browse all 69 papers →Empirical Variational Autoencoder
Kaede Shiohara
EVA makes VAEs generate high-quality images and sounds faster by replacing their fixed Gaussian latent prior with a learned, self-predictive one.
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
Yitong Wang, Fangyun Wei, Jinjing Zhao, Sirui Zhang, Hongyang Zhang, Dong Chen, Bo Dai, Yan Lu
Compo turns poster creation from describing a design in words into arranging and specifying its elements directly on a spatial canvas.
Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Jiangshan Wang, Zeqiang Lai, Jiayi Guo, Xin Yang, Xin Huang, Jiarui Chen, Ziheng Ouyang, Chunchao Guo, Xiangyu Yue
Tex-Zero shows that high-quality native 3D textures may be learned from cleverly structured 2D images instead of costly real 3D assets.