NTH

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

AuthorsJintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu

September 18, 2026 2 min read
Watch on YouTube
The one-line take

Vidu S2 brings real-time, interactive, editable, and potentially spatial AI-generated video to live applications.

Key results

720
Avatar resolution

Real-time Vidu S2-Avatar output resolution in pixels.

42
Avatar frame rate

Maximum reported real-time generation rate in FPS.

800000
Editing training clips

Curated clips used to construct four video-editing task subsets.

3.370
StreamAV visual quality

Vidu S2-Avatar visual-quality score on StreamAV-Bench.

4.26
Joint OpenVE–RefVIE score

Vidu S2-Editing overall score on the joint editing benchmark.

9.9515
ViViD VFIDI

Vidu S2-Editing score for unpaired virtual try-on, where lower is better.

What the paper found

Vidu S2 targets the gap between offline video generators such as Sora and Veo and genuinely interactive visual systems. It combines Vidu S2-Avatar, an audio-visual Diffusion Transformer for continuously generated digital characters, with Vidu S2-Editing, a causal streaming editor for style transfer, virtual try-on, subject replacement, and background replacement. Avatar generation reaches 720p at up to 42 FPS, accepts reference images dynamically during a stream, and follows large-motion instructions such as dancing. Its central training contribution is Self-Replay Forcing, which re-noises detached autoregressive rollouts and replays them through a gradient-connected causal pass, improving robustness to accumulated streaming errors; temporally ordered dense captions, preference optimization, and a one-step latent super-resolution Refiner further support control and fidelity. Editing uses frame-aligned attention so each output frame reads only its corresponding source frame while reference-image tokens propagate appearance consistently, preserving input motion and timing. An inference stack based on SageAttention, SpargeAttention, Sparse-Linear Attention, W8A8 GEMM, kernel fusion, CUDA Graphs, and Ulysses-style multi-GPU parallelism makes deployment practical on lower-cost GPUs. On StreamAV-Bench, Vidu S2-Avatar scores 3.370 for visual quality and 0.617 for audio-visual synchronization, outperforming listed baselines including Runway, HeyGen, and PixVerse in complementary evaluations. Vidu S2-Editing trains on 800,000 curated clips and achieves 4.26 on the joint OpenVE–RefVIE benchmark and 9.9515 VFIDI on ViViD virtual try-on. The system also converts generated or edited streams into synchronized left- and right-eye spatial video for VR headsets.

Original abstract

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis