NTH

Exploring the Design Space of Reward Backpropagation for Flow Matching

AuthorsRuoyu Wang, Boye Niu, Xiangxin Zhou, Yushi Huang, Tongliang Liu, Chi Zhang

June 28, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes it cheaper and more stable to train image generators from human preferences by redesigning how reward gradients are backpropagated through flow-matching sampling trajectories.

Key results

0.4134
HPv2.1 FLUX.1-dev base to best

HPSv2.1 score for FlowBP-Lagrange on FLUX.1-dev, compared with the 0.3016 base model

69.88
GenEval overall FLUX.1-dev base to best

Overall compositional score for FlowBP-Lagrange on FLUX.1-dev, compared with the 63.25 base model

9B
FLUX.2-Klein-base size

Model size of the FLUX.2-Klein-base backbone

What the paper found

This paper from Tencent Hunyuan, Westlake University, and the University of Sydney studies direct reward backpropagation for text-to-image flow matching and identifies two failure modes in prior methods: full-trajectory activations are too expensive to store, and multi-step Jacobian chains destabilize gradients. The proposed FlowBP framework turns the backward trajectory into a surrogate-design problem, separating reward-model input, active-set selection, integration weights, and bridge coupling, and it recovers ReFL, DRaFT-LV, DRTune, and LeapAlign as special cases. Three new variants are introduced: FlowBP-Sparse reconstructs the endpoint with sparse Euler updates, FlowBP-Bridge adds a controlled one-Jacobian bridge, and FlowBP-Lagrange replaces LeapAlign’s single-velocity leap with higher-order Lagrange quadrature. Across SD3.5-M, FLUX.1-dev 12B, and FLUX.2-Klein-base 9B, the methods improve HPSv2.1, PickScore, ImageReward, UR-Align, and UR-IQ on most metrics; for example, on FLUX.1-dev, HPSv2.1 rises from 0.3016 to 0.4134 and GenEval improves from 63.25 to 69.88, while the authors report that memory scales with active-set size rather than rollout length and gradient chaining is bounded to at most one Jacobian factor.

Original abstract

Aligning text-to-image flow matching models with human preferences via direct reward backpropagation is sample-efficient but hampered by two well-known pathologies: activations cannot be stored across the full sampling trajectory at modern model scale, and chained Jacobian products across steps inflate the reward gradient as it travels back to early indices. Connector-based methods, such as LeapAlign, address these issues by replacing the full backward trajectory with a short pinned path, highlighting a useful decoupling between sampling and optimization. However, the quality of the resulting gradient depends on how accurately this short path approximates the full rollout, especially over long intervals. We propose FlowBP, a unified surrogate-trajectory framework that treats the backward trajectory itself as the design object. FlowBP keeps a no-gradient cached rollout for sampling, then builds a lightweight backward surrogate from cached and selectively re-forwarded velocities. This view separates four choices: the reward-model input, active set, integration weights, and bridge coupling, and recovers prior direct-gradient methods as particular settings. Within this framework, we instantiate three variants: FlowBP-Sparse uses sparse Euler reconstruction, FlowBP-Bridge adds controlled bridge coupling, and FlowBP-Lagrange raises the order of leap quadrature. All three bound memory by the active-set size and limit gradient chaining to at most one Jacobian factor. Across SD3.5-M, FLUX.1-dev, and FLUX.2-Klein-base on preference, quality, and compositional metrics, the three variants improve over direct-gradient baselines on most metrics.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis