FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows
AuthorsDaniel Gilo, Sven Elflein, Ido Sobol, Or Litany
Resources
FlowBender teaches generative models to look at their own mistakes and correct them on the fly, improving both constraint satisfaction and sample quality.
Key results
Best first-order FlowBender result on super-resolution
Best first-order FlowBender result on super-resolution
Zero-order FlowBender result on JPEG restoration
Zero-order FlowBender result on JPEG restoration
Best first-order FlowBender result on 3D texture generation on Objaverse
Best first-order FlowBender result on 3D texture generation on Objaverse
What the paper found
FlowBender, from Technion and NVIDIA with collaborators at the University of Toronto and the Vector Institute, reframes conditional diffusion and flow-matching as a closed-loop control problem: instead of treating the condition as a static cue or relying on hand-tuned inference-time gradients, it trains the model to consume its own alignment error and learn a non-linear correction policy. The method uses a two-pass sampling step, where an unguided look-ahead estimates the clean sample, the forward operator computes a feedback signal, and a refinement pass outputs the corrected velocity; it supports both first-order gradients for differentiable operators and a zero-order residual variant for black-box or non-differentiable operators such as JPEG compression. Across image translation, restoration, and 3D mesh texturing, FlowBender consistently improves fidelity and plausibility simultaneously: on super-resolution it reaches 44.07 PSNR and 98.96 SSIM with first-order feedback versus 34.35 and 96.88 for standard fine-tuning, while zero-order feedback attains 3.36 FID; on JPEG restoration it improves from 26.29 PSNR, 79.45 SSIM, 22.24 LPIPS, and 4.35 FID to 28.86, 83.13, 16.33, and 3.80; and on Objaverse texturing it lowers FID from 8.74 to 6.64 while raising M.PSNR from 21.91 to 26.39. The paper also shows that a prior-step shortcut reduces inference to as little as N+1 evaluations for N-step sampling, and an ablation finds pun = 0.1 is the best null-feedback probability, while orthogonal decomposition reveals that about 80% of the correction energy lies outside the guidance gradient, confirming the update is not just gradient steering in disguise.
Original abstract
Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task. For instance, a depth-conditioned model often produces images whose re-extracted depth disagrees with the input, even though the forward operator--the depth predictor defining the constraint--is available during both training and inference. Existing approaches generally fall into two categories: supervised models that treat the conditioning signal as a static cue and ignore alignment information at inference, and guidance-based methods that consult it through hand-tuned linear updates, typically trading fidelity to the condition against the plausibility of the generated sample. We argue that the fundamental gap in both paradigms is that the model is never trained to utilize its own alignment error. We introduce FlowBender, a closed-loop framework that treats this error as a first-class input, training the network to learn a correction policy conditioned on inference-time feedback. At each step, an unguided look-ahead pass estimates the clean signal, a task-specific deviation is computed via the forward operator, and a refinement pass consumes this signal to produce a corrected velocity. We propose several variants of FlowBender, including a gradient-based formulation for differentiable operators and a zero-order variant for non-differentiable settings such as JPEG compression. For efficient sampling, we introduce a prior-step shortcut that enables closed-loop correction at a minimal additional computational cost. Across image-to-image translation, restoration, and 3D mesh texturing, FlowBender consistently outperforms standard supervised baselines, alignment-loss-augmented training, and state-of-the-art inference-time guidance, improving fidelity and plausibility simultaneously rather than trading them against each other. Project page: https://flow-bender.github.io/
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.