NTH

DanceOPD: On-Policy Generative Field Distillation

AuthorsWei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua

July 8, 2026 2 min read
Watch on YouTube
The one-line take

DanceOPD teaches a single image generator to juggle text-to-image creation, editing, and guidance signals by learning from its own rollout states, helping one model do many generation tasks without the usual capability conflicts.

Key results

8.1%
T2I+Edit GEditBench gain

DanceOPD over the best reproduced OPD baseline

16.1%
Local+Global GEditBench gain

DanceOPD over the best competing composition baseline

9.9%
Realism reward gain

DanceOPD over off-policy distillation on SD3.5-M

85.3%
Student-to-teacher gap closed

For realism-field absorption

7.6%
CFG composition gain vs train-only absorption

Best measured composition improvement

22.8%
Dense-query degradation

Same-step dense supervision with K=2, G=3

What the paper found

DanceOPD is a ByteDance Seed and NUS-led on-policy generative field distillation method for flow-matching image generators that treats each capability source as a velocity field and composes text-to-image generation, local editing, global editing, realism absorption, and classifier-free guidance without collapsing them into a compromise model. The core idea is to hard-route each sample to exactly one teacher field, query that field on a stop-gradient state from the student’s own 16-step Euler ODE rollout, and supervise with a single low-noise semantic-side query using plain velocity MSE. On Z-Image, this design outperforms the best reproduced OPD baseline on GEditBench-EN by 8.1% in T2I-plus-edit composition and by 16.1% in local-plus-global edit composition, while preserving or slightly improving GenEval over the T2I source. For realism-field absorption on SD3.5-M, DanceOPD improves the realism reward by 9.9% over off-policy distillation and closes 85.3% of the student-to-teacher reward gap, with T2I remaining within 0.1% of off-policy distillation. For CFG absorption, the best measured composition improves over train-only absorption by 7.6% and over eval-only CFG by 1.4%, but excessive absorbed-plus-external guidance causes a 31.2% drop. Ablations show that hard routing beats soft all-teacher mixing by 15.2%, low-noise queries beat median- and high-noise queries by 23.7% and 19.5%, dense same-rollout querying drops performance by 22.8%, and plain velocity MSE is the most stable objective.

Original abstract

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis