Memento: Reconstruct to Remember for Consistent Long Video Generation
AuthorsXuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong
Resources
Memento improves long video generation by teaching the model to reconstruct recurring subjects from memory, helping characters stay consistent across shots and scenes.
Key results
curated video sequences in the subject-aware dataset
total clips used for training
Wan2.2 14B backbone size
best reported long-term subject consistency score
best reported subject consistency across scene transitions
cross-shot consistency win rate versus HoloCine
What the paper found
Memento, from the Xiamen University and Baidu ERNIE Team collaboration, reframes long video generation as an identity-grounding problem: instead of only optimizing plausible next-shot continuation, it adds memory-based subject reconstruction so the model must recover a recurring subject from historical memory plus the global story caption. Built on the Wan2.2 14B video backbone, the system uses a dual-query memory bank in which story-conditioned queries retrieve long-range subject evidence and shot-conditioned queries retrieve short-range visual context, reducing competition between identity preservation and local coherence. Training uses a subject-aware cinematic curation pipeline over 2,033 video sequences and 20,227 clips, with Qwen3-VL-8B and ByteTrack generating pronoun-free story, shot, and reconstruction captions. On quantitative evaluation, Memento reaches 0.7338 inter-shot subject consistency and 0.7268 inter-scene subject consistency, both the best among StoryDiffusion+Wan2.2-I2V, StoryMem, and HoloCine, while also leading in story-level semantic consistency at 0.3063 and shot-level semantic consistency at 0.2893. A user study over 30 cases and 10 participants reports win rates of 57.3% to 69.0% for cross-shot consistency and 60.3% to 63.0% for prompt following versus the baselines. The paper also shows a 5-minute generation example, suggesting the memory design scales toward minute-level narrative video without architectural changes.
Original abstract
Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long-range subject evidence from short-range cues, Memento introduces a dual-query memory mechanism, where one query retrieves identity-relevant memory and the other selects short-context keyframes for coherent continuation. Additionally, a subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.