NTH

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

AuthorsLianghua Huang, Zhifan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chenwei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Zoubin Bi

June 25, 2026 2 min read
Watch on YouTube
The one-line take

Wan-Streamer is a single multimodal foundation model built for real-time two-way audio-visual conversation with sub-second latency, replacing many separate pipeline modules.

Key results

160
streaming unit

Shortest streaming unit used for real-time interaction at 25 fps

200
model-side latency

Approximate signal-to-signal response latency of the thinker-performer system

550
total interaction latency

Approximate end-to-end latency including 350 ms bidirectional network delay

350
network latency budget

Bidirectional network latency assumed in total interaction latency

192
output resolution

Preliminary validation resolution for Wan-Streamer v0.1

What the paper found

Wan-Streamer v0.1 from the Alibaba Group is a native-streaming, end-to-end interactive foundation model built for real-time full-duplex audio-visual conversation, and its key novelty is that language, audio, and video are all modeled as both inputs and outputs inside a single Transformer rather than through cascaded ASR, TTS, avatar, or video-generation modules. The system uses strictly causal audio and video VAEs, causal multimodal encoders and decoders, and block-causal attention so each newly observed unit can immediately affect the next response while committed response latents become part of the interaction history. Training combines independent-task pretraining, duplex interaction training, and distillation from a stronger CFG teacher with rolling distillation and self-forcing to reduce train-test mismatch. At inference, a two-GPU thinker-performer pipeline overlaps perception, KV-cache update, decoding, and flow-matching latent generation, producing approximately 200 ms model-side response latency and approximately 550 ms total interaction latency with a 350 ms bidirectional network budget. The paper reports streaming units as short as 160 ms at 25 fps, and validates the v0.1 system at a preliminary 192p output resolution, positioning Wan-Streamer as a unified benchmark for low-latency multimodal dialogue with synchronized speech and visual response.

Original abstract

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. Wan-Streamer seamlessly models language, audio, and video as both input and output within a single Transformer, where the sequence is represented as interleaved visual, audio, and text input tokens together with visual, audio, and text output tokens, coordinated by block-causal attention for incremental streaming. Unlike cascaded interactive systems that rely on separate VAD, ASR, language, TTS, audio-driven animation, or video-generation modules, Wan-Streamer does not rely on external language, speech, avatar, or video-generation modules: perception, reasoning, generation, response timing, turn management, and cross-modal synchronization are learned jointly within one unified model, reducing pipeline latency and error accumulation. To support natural audio-visual responsiveness, we redesign the entire stack around streamability, including causal encoders, causal decoders, block-causal attention, and low-latency multimodal token scheduling, enabling streaming units as short as 160 ms at 25 fps. Wan-Streamer achieves approximately 200 ms model-side response latency and approximately 550 ms total interaction latency when combined with 350 ms bidirectional network latency, supporting sub-second duplex audio-visual communication. These results position Wan-Streamer as a unified, end-to-end, multimodal interactive foundation model for low-latency streaming interaction.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis