Avatar V: Scaling Video-Reference Avatar Video Generation
AuthorsBenjamin Liang, Ce Chen, Desmond Lin, Ivan Somov, Jiajun Zhao, Jiewei Yuan, Jingfeng Zhang, Junhao Huang, Nik Nolte, Pedram Haqiqi, Penghan Wang, Rong Yan, Rui Zhang, Sam Prokopchuk, Sivan Wang, Viktor Goriachko, Yi Ren, Yuanming Li, Yutao Chen, Zhenhui Ye, Zhibin Hong, Zilong Nie, Zujin Guo
Resources
Avatar V is a large-scale system that uses reference videos, not just photos, to generate more realistic talking avatars with better identity, gestures, and lip-sync.
Key results
source corpus used for data curation
output of the pretraining pipeline
output of the avatar-specific fine-tuning pipeline
two-phase distillation speedup
inference steps after distillation
unoptimized baseline latency improvement
What the paper found
Avatar V, from HeyGen Research, is a production-scale video-reference avatar generator that replaces single-image identity conditioning with full reference-video conditioning so it can reproduce not only appearance but also behavioral traits such as talking rhythm, micro-expressions, and gestural style. Its core Diffusion Transformer uses Sparse Reference Attention to make reference conditioning scale almost linearly with reference length, while a dedicated motion stream closes the loop between motion prediction and generation, and an identity-aware super-resolution refiner restores facial detail at 1080p. The system is trained on a data engine that mines 50M raw videos into 100M+ pretraining clips and 10M+ avatar fine-tuning clips, then optimized through a five-stage pipeline: text-to-video pretraining, audio-to-video pretraining, personality supervised fine-tuning, two-phase distillation for over 10× acceleration, and RLHF alignment. At inference, Avatar V runs in 24 denoising steps, supports chunk-based generation for unlimited duration, and uses custom compiler, NVSHMEM, and infrastructure optimizations that cut latency by 3× over the unoptimized baseline and by 33% over torch.compile Inductor. On a 70-case cross-scene benchmark, it outperforms systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5, reaching the best SyncNet confidence at 8.97, the best face similarity at 0.840, and the highest human MOS for identity at 4.98 out of 5, while its Avatar Turing Test shows annotators still identify real footage correctly 77.8% of the time.
Original abstract
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dynamics, remains an open challenge. Existing methods predominantly condition on single static images, which provide insufficient identity information and cannot capture dynamic motion traits, while standard pixel-level objectives underserve the perceptually critical facial regions that determine avatar fidelity. We present Avatar V, a production-scale framework that addresses these limitations through video-reference-conditioned identity modeling. Rather than compressing identity into fixed-size embeddings, the model conditions directly on the full token sequence of a reference video, learning to reproduce both static identity attributes (facial geometry, skin texture) and dynamic behavioral patterns (talking rhythm, micro-expressions) through attention over the reference context. We introduce Sparse Reference Attention, an asymmetric mechanism achieving linear-complexity conditioning on arbitrarily long references; a motion representation stream enabling closed-loop talking style transfer; and an identity-aware super-resolution refiner inheriting the full reference conditioning. These are supported by a data engine curating 100M+ training clips from 50M raw videos, and a five-stage training pipeline with flow matching pre-training, personality fine-tuning, two-phase distillation (>10x acceleration), and RLHF alignment, deployed across thousands of GPUs. Avatar V generates 1080p videos of unlimited duration, achieving state-of-the-art identity preservation, lip synchronization, and generation quality on our cross-scene benchmark, consistently outperforming leading systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5 in both automated metrics and human evaluation.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.