KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation
AuthorsJianjie Luo, Yiming Zhong, Haoming Shen, Yupeng Xiao, Zhenguo Yang
KeyID keeps video characters recognizable by drafting the action first and then repairing identity at strategically chosen keyframes.
Key results
KeyID’s identity-consistency score on VIP-200K.
KeyID’s ArcFace-based identity score on VIP-200K.
Improvement over the strongest baseline, TPIGE.
Improvement over TPIGE.
Sequential Action track score, earning second place.
Number of frames in the five-second output video.
What the paper found
KeyID is a training-free framework for identity-preserving video generation that separates motion synthesis from identity insertion. Its Reference-Aware Video Generation stage uses OpenAI’s GPT-5.4 for prompt enhancement, ERNIE-Image-Turbo and Qwen-Image-Edit-2511 for the initial scene, Wan2.2 with Prompt Relay for timeline-controlled video drafting, and LTX-2.3 for motion interpolation. Instead of forcing one diffusion process to satisfy detailed actions and preserve a face simultaneously, KeyID generates an identity-agnostic draft, samples local groups around five representative keyframes, applies FaceSwap correction, ranks candidates with ArcFace, and interpolates the corrected anchors across the sequence. On the VIP-200K benchmark, KeyID reaches a Face-Cur score of 0.633 and a Face-Arc score of 0.630, improving over the strongest baseline TPIGE by 28.7% and 33.2%, respectively, while achieving the best CLIPScore of 30.9. In ablations, adding keyframe editing raises Face-Cur from 0.279 to 0.596, while combining it with reference-aware drafting reaches 0.636 Face-Cur, 0.643 Face-Arc, and 30.9 CLIPScore. For sequential actions, the system produces an 81-frame draft and a 121-frame final video, and earns a challenge score of 1.81 for second place in Track 2, demonstrating that sparse identity correction can preserve facial fidelity without sacrificing prompt adherence or complex temporal actions.
Original abstract
Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enhancements, struggling to maintain rigid identity consistency during complex, long-sequence actions. To address these limitations, we propose KeyID, a training-free IPVG framework that decouples the synthesis of video dynamics from the injection of identity. Specifically, KeyID comprises two components: (1) Reference-Aware Video Generation, which produces an identity-agnostic video draft aligned with multiple references, and (2) Identity-Preserved Keyframe Editing, which integrates the target identity via sparse keyframe correction and subsequent motion interpolation. By shifting from dense frame-level supervision to sparse keyframe-level refinement, KeyID effectively resolves the capacity conflict between prompt adherence and identity fidelity. Crucially, our modular design allows seamless extension to multi-subject references and complex sequential action generation without additional training. KeyID outperforms prior works and is validated by automatic and human evaluations on the official challenge benchmark, ultimately securing the runner-up position in the Track 2 (Sequential Action) of the ACM Multimedia 2026 IPVG Grand Challenge. Source code is available at https://github.com/WISLab-GDUT/KeyID.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.