NTH

ID-V2V: Identity-Preserving Video Restylization

AuthorsYuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu

August 2, 2026 3 min read
Watch on YouTube
The one-line take

ID-V2V edits a video's style and scene while preserving the people, expressions, gaze, and lip movements that make the original performance recognizable.

Key results

330,000
Relighting training pairs

Approximate paired face images generated from OLAT data using 67 participants.

40k
Video training set

Human-centric videos used to train ID-V2V.

0.701
Single-subject AdaFace

ID-V2V facial identity similarity score, versus 0.529 for WanAnimate.

0.668
Single-subject Exp-AU

Expression-preservation similarity score for ID-V2V.

85.62%
Two-subject facial likeness preference

User-study win rate for ID-V2V.

86.09%
Two-subject facial performance preference

User-study win rate for preserving facial performance.

What the paper found

ID-V2V, developed by researchers from Netflix, Adobe, and Eyeline Labs, addresses identity-preserving video restylization: propagating an edited first-frame’s changes to scene, lighting, or style across a source video while retaining facial likeness, expressions, eye gaze, and lip synchronization. Its central innovation is to decouple edit-driven synthesis from identity preservation. Built on VACE and Wan2.1, the model uses the edited keyframe and DepthAnything 2 depth sequences for coherent visual editing, while relit facial regions and DAViD-predicted face normal maps provide pixel-level appearance and geometric constraints. A dedicated relighting model creates training supervision from a single video, avoiding scarce paired restylization data; its relighting dataset contains approximately 330,000 image pairs from 67 participants, and the video model is trained on 40k human-centric videos at 81 frames per video. On single-subject videos, ID-V2V achieves an AdaFace identity score of 0.701, compared with 0.529 for WanAnimate, while reaching 0.668 on Exp-AU expression similarity. In user studies, it wins 75.00% of single-subject facial-likeness judgments and, for two-subject videos, wins 85.62% for facial likeness and 86.09% for facial-performance preservation. The method also supports multiple interacting subjects and can generalize facial relighting to bodies and backgrounds, although extreme colored lighting and large source-to-target geometry changes can cause residual lighting or edit inconsistencies.

Original abstract

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis