Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
AuthorsJintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu
Resources
Vidu S2 brings real-time, interactive, editable, and potentially spatial AI-generated video to live applications.
Key results
Real-time Vidu S2-Avatar output resolution in pixels.
Maximum reported real-time generation rate in FPS.
Curated clips used to construct four video-editing task subsets.
Vidu S2-Avatar visual-quality score on StreamAV-Bench.
Vidu S2-Editing overall score on the joint editing benchmark.
Vidu S2-Editing score for unpaired virtual try-on, where lower is better.
What the paper found
Vidu S2 targets the gap between offline video generators such as Sora and Veo and genuinely interactive visual systems. It combines Vidu S2-Avatar, an audio-visual Diffusion Transformer for continuously generated digital characters, with Vidu S2-Editing, a causal streaming editor for style transfer, virtual try-on, subject replacement, and background replacement. Avatar generation reaches 720p at up to 42 FPS, accepts reference images dynamically during a stream, and follows large-motion instructions such as dancing. Its central training contribution is Self-Replay Forcing, which re-noises detached autoregressive rollouts and replays them through a gradient-connected causal pass, improving robustness to accumulated streaming errors; temporally ordered dense captions, preference optimization, and a one-step latent super-resolution Refiner further support control and fidelity. Editing uses frame-aligned attention so each output frame reads only its corresponding source frame while reference-image tokens propagate appearance consistently, preserving input motion and timing. An inference stack based on SageAttention, SpargeAttention, Sparse-Linear Attention, W8A8 GEMM, kernel fusion, CUDA Graphs, and Ulysses-style multi-GPU parallelism makes deployment practical on lower-cost GPUs. On StreamAV-Bench, Vidu S2-Avatar scores 3.370 for visual quality and 0.617 for audio-visual synchronization, outperforming listed baselines including Runway, HeyGen, and PixVerse in complementary evaluations. Vidu S2-Editing trains on 800,000 curated clips and achieves 4.26 on the joint OpenVE–RefVIE benchmark and 9.9515 VFIDI on ViViD virtual try-on. The system also converts generated or edited streams into synchronized left- and right-eye spatial video for VR headsets.
Original abstract
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.