NTH

MEND: RL For Flow Models via Proximal Velocity Matching

AuthorsShreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik

AffiliationsThe University of Texas at Austin Google University of Colorado Boulder

October 11, 2026 2 min read
Watch on YouTube
The one-line take

MEND makes flow-model reward tuning more efficient by moving samples only when the reward gain justifies the size of the change.

Key results

100
MEND updates

SD3.5-M comparison against Flow-GRPO.

4k
Flow-GRPO updates

Approximate update count for the comparison baseline.

24.03
MEND PickScore

Equal-budget result; ReFL scores 23.92 and DiffusionNFT 23.43.

10.0
Training compute

GPU-hours for the SD3.5-M run.

What the paper found

MEND changes how flow models learn from reward scores: rather than reweighting samples or moving every sample along a reward gradient, it caps rewards within each prompt group and proposes short gradient-directed moves only for samples below the cap. A move is accepted only when its capped reward gain exceeds a quadratic distance cost; the model then learns the accepted changes through proximal velocity matching, without a KL penalty, frozen reference model, or advantage weights. On SD3.5-M, MEND beats Flow-GRPO on five of six evaluators after 100 updates, compared with about 4k updates for Flow-GRPO, at the same distance from base-model images. In equal-budget tests, it reaches PickScore 24.03, ahead of ReFL at 23.92 and DiffusionNFT at 23.43. The method also improves SD3-M and Z-Image-Turbo, and the SD3.5-M run takes 10.0 GPU-hours. Its guarantees apply to selected training targets, not to held-out rewards: the paper notes that longer training can over-optimize a reward and reduce other measures.

Original abstract

Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.

Read the original paper

More in Generative Models

Browse all 69 papers →
01Generative Model

Empirical Variational Autoencoder

Kaede Shiohara

EVA makes VAEs generate high-quality images and sounds faster by replacing their fixed Gaussian latent prior with a learned, self-predictive one.

Read analysis