NTH

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

AuthorsYicheng Xiao, Wenxun Dai, Xinran Qin, Lin Song, Maoquan Zhang, Hang Xu, Yukang Chen, Yitong Li, Guohui Zhang, Yuan Zhang, Xuying Zhang, Tommy Zhang, Jianlong Yuan, Peihao Li, Shuai Lu, Siming Fu, Chuyang Zhao, Xin Han, Jie Huang, Wenbo Li, Guoqing Ma, Wei Huang, Xiaojuan Qi, Haoyang Huang, Nan Duan

August 7, 2026 2 min read
Watch on YouTube
The one-line take

JoyAI-Video-Edit aims to make high-quality, open-ended video editing run in real time by combining autoregressive generation with diffusion distillation.

Key results

16B
Model parameters

Size of the JoyAI-Video-Edit autoregressive diffusion model

2
Diffusion sampling steps

SA-DMD reduces deployment generation to a two-step process

30.19
End-to-end throughput

FPS for 720p editing on one Nvidia B200 GPU

3.60
OpenVE-Bench overall score

Overall score from the Gemini multimodal evaluation

229
LongV2VBench task count

One-minute long-video editing tasks in the introduced benchmark

3.30
LongV2VBench overall score

Overall editing-quality score on the long-video benchmark

What the paper found

The JD-affiliated Joy Future Academy team presents JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion system for open-ended video editing that generates causally, without future frames or a fixed duration. Its architecture combines an MLLM condition encoder, causal video VAE, and multimodal diffusion transformer, using chunk-wise attention, bounded sliding-window KV caching, and a global first-chunk sink to keep memory and per-chunk computation constant. Two key training techniques address autoregressive failure: Source-Anchored Distribution Matching Distillation, or SA-DMD, compresses diffusion into a two-step generator while anchoring each prediction to the aligned source chunk, and Long-Horizon Autoregressive Distillation trains on extended rollouts to reduce accumulated temporal drift. With FP8 quantization, operator fusion, and optimized VAE execution, the system reaches 30.19 FPS for end-to-end 720p editing on a single Nvidia B200 GPU. On the short-video OpenVE-Bench, judged by a Gemini multimodal evaluator, it scores 3.60 overall and outperforms streaming baselines while approaching offline systems such as Runway Aleph and Kuaishou’s Kling-3.0 Omni. The authors also introduce LongV2VBench, containing 229 one-minute editing tasks; JoyAI-Video-Edit scores 3.30 overall and ranks first across all five categories, demonstrating sustained quality over long streams rather than only short clips.

Original abstract

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis