Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching
AuthorsYoad Tewel, Yuval Atzmon, Gal Chechik, Lior Wolf
Resources
This paper shows how to train image and video editors without paired before-and-after data by bootstrapping guidance from the model itself and using flow matching to preserve structure.
Key results
In the video editing user study, the method won 75.3% of comparisons overall against Ditto.
On unseen 3D-CGI video inputs, the method won 85% of comparisons against Ditto.
What the paper found
Bootstrap Your Generator, from NVIDIA and Tel Aviv University authors Yoad Tewel, Yuval Atzmon, Gal Chechik, and Lior Wolf, proposes an unpaired training framework for flow matching visual editing that needs no source–target image pairs and no external reward model. The key idea is to fine-tune a pretrained text-to-image or text-to-video generator using only source captions, target captions, and inverse instructions: a frozen EMA copy of the model first bootstraps pseudo-targets by multi-step sampling, then the trainable editor is supervised by a semantic prior that aligns the edit direction between source and target prompts, plus a cycle-consistency loss that reconstructs the source after applying the inverse edit. To make cycle training work on flow-matching models, the paper introduces gradient routing with Straight-Through Estimation, so clean multi-step predictions condition the reverse pass while gradients flow through the noisy one-step state. Experiments on Wan2.2 video editing and FLUX.1-dev image editing show strong gains on long-tail, data-scarce edits: in a user study, the method wins 75.3% of comparisons against Ditto, a supervised baseline trained on one million video pairs, and it generalizes to unseen 3D-CGI videos with an 85% win rate. On GEdit-Bench, it matches or exceeds FLUX-Kontext and FlowEdit across categories, especially motion, human, and style edits, while ablations confirm that bootstrapping, directional regularization, cycle loss, and gradient routing are all necessary for preserving structure without collapsing to identity.
Original abstract
Modern generative models possess a deep understanding of visual content, yet training them for image editing typically requires massive datasets of paired examples. This limits scalability, especially for video editing where collecting paired data is prohibitively expensive. We propose Bootstrap Your Generator (ByG), a general framework for unpaired training of flow matching editing models. It leverages the base model's knowledge without any external signal. Our approach pairs instruction-following cues extracted from the frozen model with cycle-consistency for structure preservation. To make this tractable, we propose to route gradients from downstream losses over clean predictions to noisy training states. We demonstrate state-of-the-art results on challenging data-scarce image and video editing scenarios. Extensive evaluations and user studies show that our method effectively generalizes to unseen domains and outperforms supervised baselines trained on millions of samples. Analysis reveals that our gradient routing bridges the train-inference gap, and extracting semantic cues from a base model provides a robust training signal that obviates the need for external reward models.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.