Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
AuthorsLianghua Huang, Zhifan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chenwei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Zoubin Bi
Resources
Wan-Streamer is a single multimodal foundation model built for real-time two-way audio-visual conversation with sub-second latency, replacing many separate pipeline modules.
Key results
Shortest streaming unit used for real-time interaction at 25 fps
Approximate signal-to-signal response latency of the thinker-performer system
Approximate end-to-end latency including 350 ms bidirectional network delay
Bidirectional network latency assumed in total interaction latency
Preliminary validation resolution for Wan-Streamer v0.1
What the paper found
Wan-Streamer v0.1 from the Alibaba Group is a native-streaming, end-to-end interactive foundation model built for real-time full-duplex audio-visual conversation, and its key novelty is that language, audio, and video are all modeled as both inputs and outputs inside a single Transformer rather than through cascaded ASR, TTS, avatar, or video-generation modules. The system uses strictly causal audio and video VAEs, causal multimodal encoders and decoders, and block-causal attention so each newly observed unit can immediately affect the next response while committed response latents become part of the interaction history. Training combines independent-task pretraining, duplex interaction training, and distillation from a stronger CFG teacher with rolling distillation and self-forcing to reduce train-test mismatch. At inference, a two-GPU thinker-performer pipeline overlaps perception, KV-cache update, decoding, and flow-matching latent generation, producing approximately 200 ms model-side response latency and approximately 550 ms total interaction latency with a 350 ms bidirectional network budget. The paper reports streaming units as short as 160 ms at 25 fps, and validates the v0.1 system at a preliminary 192p output resolution, positioning Wan-Streamer as a unified benchmark for low-latency multimodal dialogue with synchronized speech and visual response.
Original abstract
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. Wan-Streamer seamlessly models language, audio, and video as both input and output within a single Transformer, where the sequence is represented as interleaved visual, audio, and text input tokens together with visual, audio, and text output tokens, coordinated by block-causal attention for incremental streaming. Unlike cascaded interactive systems that rely on separate VAD, ASR, language, TTS, audio-driven animation, or video-generation modules, Wan-Streamer does not rely on external language, speech, avatar, or video-generation modules: perception, reasoning, generation, response timing, turn management, and cross-modal synchronization are learned jointly within one unified model, reducing pipeline latency and error accumulation. To support natural audio-visual responsiveness, we redesign the entire stack around streamability, including causal encoders, causal decoders, block-causal attention, and low-latency multimodal token scheduling, enabling streaming units as short as 160 ms at 25 fps. Wan-Streamer achieves approximately 200 ms model-side response latency and approximately 550 ms total interaction latency when combined with 350 ms bidirectional network latency, supporting sub-second duplex audio-visual communication. These results position Wan-Streamer as a unified, end-to-end, multimodal interactive foundation model for low-latency streaming interaction.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.