NTH

AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation

AuthorsHaobo Li, Yanhong Zeng, Yunhong Lu, Jiapeng Zhu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Yujun Shen, Zhipeng Zhang

June 28, 2026 2 min read
Watch on YouTube
The one-line take

AAD-1 improves one-step video generation by using a smarter, non-mirrored discriminator and a staged training recipe to reduce motion collapse and make outputs more realistic.

Key results

94.34
VBench subject consistency

Best one-step autoregressive result on VBench-I2V

95.08
VBench background consistency

Best one-step autoregressive result on VBench-I2V

98.22
Motion smoothness

VBench-I2V score for the final Stage-III model

98.65
I2V subject faithfulness

VBench-I2V conditioning preservation score

62.81
DMD warm-up imaging quality

Without DMD warm-up, imaging quality in ablation

1.08
Static-video Dynamic Degree

Causal backbone with frame-wise logits collapses to static video

What the paper found

AAD-1, from SJTU, Ant Group, Tsinghua University, Zhejiang University, and Anyverse Dynamics, tackles one-step autoregressive image-to-video generation by breaking the symmetry between generator and discriminator: the causal generator preserves streaming inference, while a bidirectional discriminator with a video-level holistic logit can see the full spatiotemporal rollout and detect motion collapse and long-range drift that causal critics miss. The method combines a three-stage training recipe—ODE initialization, a self-rollout distribution-matching warm-up, and asymmetric adversarial refinement—on top of the public Wan 2.1 T2V 14B backbone, using sink tokens, a sliding window of 9, and a single sampling step per chunk. On VBench-I2V, the final one-step model reaches 94.34 subject consistency, 95.08 background consistency, 98.22 motion smoothness, and 98.65 I2V subject faithfulness, outperforming prior autoregressive baselines such as CausVid and Self Forcing while matching or exceeding the bidirectional Wan 2.1 reference on several consistency metrics. The ablations show why the design matters: without DMD warm-up, aesthetic quality drops to 53.63 and imaging quality to 62.81, while a causal frame-wise discriminator collapses to static video with a Dynamic Degree of 1.08; the best regularization is λ = 20, because λ = 0 causes training collapse and λ = 50 introduces grid artifacts. The paper also reports that full training takes about 3.5 days on 64 NVIDIA H20 GPUs, and one-step inference on a 14B model runs at 1.134 s and 14.33 FPS on a single H100 GPU.

Original abstract

We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt adversarial distillation but suffer from motion collapse and training instability, resulting in static videos. AAD-1 addresses these challenges through two key designs in architecture and training strategy. Our key architectural insight is to break the symmetry between generator and discriminator. While the generator remains causal to preserve autoregressive sampling capability, the discriminator attends bidirectionally over the full spatiotemporal context and produces a single holistic realism score for the entire video sequence. This asymmetric design enables the discriminator to effectively detect global temporal failures and long-range drift that cause motion collapse in autoregressive generation. To stabilize training, we introduce a phased strategy that first uses distribution matching to bootstrap a stable one-step generator, providing a warm-up phase that brings the student distribution closer to the teacher before adversarial distillation begins. Extensive experiments on VBench demonstrate that AAD-1 achieves state-of-the-art performance in one-step autoregressive video generation.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis