NTH

Memento: Reconstruct to Remember for Consistent Long Video Generation

AuthorsXuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong

June 22, 2026 2 min read
Watch on YouTube
The one-line take

Memento improves long video generation by teaching the model to reconstruct recurring subjects from memory, helping characters stay consistent across shots and scenes.

Key results

2033
training sequences

curated video sequences in the subject-aware dataset

20227
training clips

total clips used for training

14
model backbone

Wan2.2 14B backbone size

0.7338
inter-shot subject consistency

best reported long-term subject consistency score

0.7268
inter-scene subject consistency

best reported subject consistency across scene transitions

69.0%
user-study win rate

cross-shot consistency win rate versus HoloCine

What the paper found

Memento, from the Xiamen University and Baidu ERNIE Team collaboration, reframes long video generation as an identity-grounding problem: instead of only optimizing plausible next-shot continuation, it adds memory-based subject reconstruction so the model must recover a recurring subject from historical memory plus the global story caption. Built on the Wan2.2 14B video backbone, the system uses a dual-query memory bank in which story-conditioned queries retrieve long-range subject evidence and shot-conditioned queries retrieve short-range visual context, reducing competition between identity preservation and local coherence. Training uses a subject-aware cinematic curation pipeline over 2,033 video sequences and 20,227 clips, with Qwen3-VL-8B and ByteTrack generating pronoun-free story, shot, and reconstruction captions. On quantitative evaluation, Memento reaches 0.7338 inter-shot subject consistency and 0.7268 inter-scene subject consistency, both the best among StoryDiffusion+Wan2.2-I2V, StoryMem, and HoloCine, while also leading in story-level semantic consistency at 0.3063 and shot-level semantic consistency at 0.2893. A user study over 30 cases and 10 participants reports win rates of 57.3% to 69.0% for cross-shot consistency and 60.3% to 63.0% for prompt following versus the baselines. The paper also shows a 5-minute generation example, suggesting the memory design scales toward minute-level narrative video without architectural changes.

Original abstract

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long-range subject evidence from short-range cues, Memento introduces a dual-query memory mechanism, where one query retrieves identity-relevant memory and the other selects short-context keyframes for coherent continuation. Additionally, a subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis