ID-V2V: Identity-Preserving Video Restylization
AuthorsYuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
ID-V2V edits a video's style and scene while preserving the people, expressions, gaze, and lip movements that make the original performance recognizable.
Key results
Approximate paired face images generated from OLAT data using 67 participants.
Human-centric videos used to train ID-V2V.
ID-V2V facial identity similarity score, versus 0.529 for WanAnimate.
Expression-preservation similarity score for ID-V2V.
User-study win rate for ID-V2V.
User-study win rate for preserving facial performance.
What the paper found
ID-V2V, developed by researchers from Netflix, Adobe, and Eyeline Labs, addresses identity-preserving video restylization: propagating an edited first-frame’s changes to scene, lighting, or style across a source video while retaining facial likeness, expressions, eye gaze, and lip synchronization. Its central innovation is to decouple edit-driven synthesis from identity preservation. Built on VACE and Wan2.1, the model uses the edited keyframe and DepthAnything 2 depth sequences for coherent visual editing, while relit facial regions and DAViD-predicted face normal maps provide pixel-level appearance and geometric constraints. A dedicated relighting model creates training supervision from a single video, avoiding scarce paired restylization data; its relighting dataset contains approximately 330,000 image pairs from 67 participants, and the video model is trained on 40k human-centric videos at 81 frames per video. On single-subject videos, ID-V2V achieves an AdaFace identity score of 0.701, compared with 0.529 for WanAnimate, while reaching 0.668 on Exp-AU expression similarity. In user studies, it wins 75.00% of single-subject facial-likeness judgments and, for two-subject videos, wins 85.62% for facial likeness and 86.09% for facial-performance preservation. The method also supports multiple interacting subjects and can generalize facial relighting to bodies and backgrounds, although extreme colored lighting and large source-to-target geometry changes can cause residual lighting or edit inconsistencies.
Original abstract
In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.