NTH

Avatar V: Scaling Video-Reference Avatar Video Generation

AuthorsBenjamin Liang, Ce Chen, Desmond Lin, Ivan Somov, Jiajun Zhao, Jiewei Yuan, Jingfeng Zhang, Junhao Huang, Nik Nolte, Pedram Haqiqi, Penghan Wang, Rong Yan, Rui Zhang, Sam Prokopchuk, Sivan Wang, Viktor Goriachko, Yi Ren, Yuanming Li, Yutao Chen, Zhenhui Ye, Zhibin Hong, Zilong Nie, Zujin Guo

June 18, 2026 3 min read
Watch on YouTube
The one-line take

Avatar V is a large-scale system that uses reference videos, not just photos, to generate more realistic talking avatars with better identity, gestures, and lip-sync.

Key results

50M
raw videos

source corpus used for data curation

100M+
pretraining clips

output of the pretraining pipeline

10M+
avatar fine-tuning clips

output of the avatar-specific fine-tuning pipeline

10×
distillation acceleration

two-phase distillation speedup

24
denoising steps

inference steps after distillation

3×
latency reduction

unoptimized baseline latency improvement

What the paper found

Avatar V, from HeyGen Research, is a production-scale video-reference avatar generator that replaces single-image identity conditioning with full reference-video conditioning so it can reproduce not only appearance but also behavioral traits such as talking rhythm, micro-expressions, and gestural style. Its core Diffusion Transformer uses Sparse Reference Attention to make reference conditioning scale almost linearly with reference length, while a dedicated motion stream closes the loop between motion prediction and generation, and an identity-aware super-resolution refiner restores facial detail at 1080p. The system is trained on a data engine that mines 50M raw videos into 100M+ pretraining clips and 10M+ avatar fine-tuning clips, then optimized through a five-stage pipeline: text-to-video pretraining, audio-to-video pretraining, personality supervised fine-tuning, two-phase distillation for over 10× acceleration, and RLHF alignment. At inference, Avatar V runs in 24 denoising steps, supports chunk-based generation for unlimited duration, and uses custom compiler, NVSHMEM, and infrastructure optimizations that cut latency by 3× over the unoptimized baseline and by 33% over torch.compile Inductor. On a 70-case cross-scene benchmark, it outperforms systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5, reaching the best SyncNet confidence at 8.97, the best face similarity at 0.840, and the highest human MOS for identity at 4.98 out of 5, while its Avatar Turing Test shows annotators still identify real footage correctly 77.8% of the time.

Original abstract

Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dynamics, remains an open challenge. Existing methods predominantly condition on single static images, which provide insufficient identity information and cannot capture dynamic motion traits, while standard pixel-level objectives underserve the perceptually critical facial regions that determine avatar fidelity. We present Avatar V, a production-scale framework that addresses these limitations through video-reference-conditioned identity modeling. Rather than compressing identity into fixed-size embeddings, the model conditions directly on the full token sequence of a reference video, learning to reproduce both static identity attributes (facial geometry, skin texture) and dynamic behavioral patterns (talking rhythm, micro-expressions) through attention over the reference context. We introduce Sparse Reference Attention, an asymmetric mechanism achieving linear-complexity conditioning on arbitrarily long references; a motion representation stream enabling closed-loop talking style transfer; and an identity-aware super-resolution refiner inheriting the full reference conditioning. These are supported by a data engine curating 100M+ training clips from 50M raw videos, and a five-stage training pipeline with flow matching pre-training, personality fine-tuning, two-phase distillation (>10x acceleration), and RLHF alignment, deployed across thousands of GPUs. Avatar V generates 1080p videos of unlimited duration, achieving state-of-the-art identity preservation, lip synchronization, and generation quality on our cross-scene benchmark, consistently outperforming leading systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5 in both automated metrics and human evaluation.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis