NTH

FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows

AuthorsDaniel Gilo, Sven Elflein, Ido Sobol, Or Litany

June 20, 2026 3 min read
Watch on YouTube
The one-line take

FlowBender teaches generative models to look at their own mistakes and correct them on the fly, improving both constraint satisfaction and sample quality.

Key results

44.07
Super-resolution PSNR

Best first-order FlowBender result on super-resolution

98.96
Super-resolution SSIM

Best first-order FlowBender result on super-resolution

28.86
JPEG restoration PSNR

Zero-order FlowBender result on JPEG restoration

3.80
JPEG restoration FID

Zero-order FlowBender result on JPEG restoration

26.39
Objaverse M.PSNR

Best first-order FlowBender result on 3D texture generation on Objaverse

6.64
Objaverse FID

Best first-order FlowBender result on 3D texture generation on Objaverse

What the paper found

FlowBender, from Technion and NVIDIA with collaborators at the University of Toronto and the Vector Institute, reframes conditional diffusion and flow-matching as a closed-loop control problem: instead of treating the condition as a static cue or relying on hand-tuned inference-time gradients, it trains the model to consume its own alignment error and learn a non-linear correction policy. The method uses a two-pass sampling step, where an unguided look-ahead estimates the clean sample, the forward operator computes a feedback signal, and a refinement pass outputs the corrected velocity; it supports both first-order gradients for differentiable operators and a zero-order residual variant for black-box or non-differentiable operators such as JPEG compression. Across image translation, restoration, and 3D mesh texturing, FlowBender consistently improves fidelity and plausibility simultaneously: on super-resolution it reaches 44.07 PSNR and 98.96 SSIM with first-order feedback versus 34.35 and 96.88 for standard fine-tuning, while zero-order feedback attains 3.36 FID; on JPEG restoration it improves from 26.29 PSNR, 79.45 SSIM, 22.24 LPIPS, and 4.35 FID to 28.86, 83.13, 16.33, and 3.80; and on Objaverse texturing it lowers FID from 8.74 to 6.64 while raising M.PSNR from 21.91 to 26.39. The paper also shows that a prior-step shortcut reduces inference to as little as N+1 evaluations for N-step sampling, and an ablation finds pun = 0.1 is the best null-feedback probability, while orthogonal decomposition reveals that about 80% of the correction energy lies outside the guidance gradient, confirming the update is not just gradient steering in disguise.

Original abstract

Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task. For instance, a depth-conditioned model often produces images whose re-extracted depth disagrees with the input, even though the forward operator--the depth predictor defining the constraint--is available during both training and inference. Existing approaches generally fall into two categories: supervised models that treat the conditioning signal as a static cue and ignore alignment information at inference, and guidance-based methods that consult it through hand-tuned linear updates, typically trading fidelity to the condition against the plausibility of the generated sample. We argue that the fundamental gap in both paradigms is that the model is never trained to utilize its own alignment error. We introduce FlowBender, a closed-loop framework that treats this error as a first-class input, training the network to learn a correction policy conditioned on inference-time feedback. At each step, an unguided look-ahead pass estimates the clean signal, a task-specific deviation is computed via the forward operator, and a refinement pass consumes this signal to produce a corrected velocity. We propose several variants of FlowBender, including a gradient-based formulation for differentiable operators and a zero-order variant for non-differentiable settings such as JPEG compression. For efficient sampling, we introduce a prior-step shortcut that enables closed-loop correction at a minimal additional computational cost. Across image-to-image translation, restoration, and 3D mesh texturing, FlowBender consistently outperforms standard supervised baselines, alignment-loss-augmented training, and state-of-the-art inference-time guidance, improving fidelity and plausibility simultaneously rather than trading them against each other. Project page: https://flow-bender.github.io/

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis