NTH

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

AuthorsShuai Yang, Bingjie Gao, Ziwei Liu, Jiaqi Wang, Dahua Lin, Tong Wu

July 1, 2026 2 min read
Watch on YouTube
The one-line take

PermaVid helps edited videos stay coherent over time by using separate memory for appearance and geometry so later generations don’t drift after changes.

Key results

4k
UE-Mem videos

Synthetic long-video dataset size

1000
UE-Mem frames per video

Long revisiting trajectories in each video

100
UE-Mem scenes

Number of Unreal Engine scenes

27.87
Global edit CLIP-Vid

Best semantic consistency on the global-edit benchmark

0.8544
Local edit VBench-Avg

Best visual quality on the local-edit benchmark

What the paper found

PermaVid, developed by researchers from Shanghai Jiao Tong University, Stanford University, Nanyang Technological University, The Chinese University of Hong Kong, and the Shanghai Innovation Institute, tackles a failure mode in memory-based video generation: after an edit, older context becomes stale and causes the model to revert to pre-edit content. The core idea is to disentangle spatial context into an RGB memory for semantic appearance and a depth memory for geometric structure, then update and retrieve these banks differently for global edits, which invalidate all RGB context, and local edits, which only invalidate overlapping regions. Built on Wan2.1-14B with a VACE-initialized context branch, the system fuses mixed-modality references through a DiT video generator and trains in two stages on SpatialVid and the new UE-Mem dataset. UE-Mem contains 4k long videos with 1000 frames each across 100 Unreal Engine scenes, captured with revisiting trajectories and 6-DoF poses. On a 200-image benchmark with complex revisiting camera paths, PermaVid achieves the best global-edit scores, including 22.84 PSNR, 0.8703 SSIM, 0.2102 LPIPS, and 27.87 CLIP-Vid, and also leads local-edit evaluation with 22.332 PSNR, 0.8622 SSIM, 0.2369 LPIPS, and 0.8544 VBench-Avg. The paper’s ablation shows that removing disentangled context memory causes outdated RGB reuse and progressive semantic inconsistency, while the proposed retrieval overhead remains only at the millisecond level.

Original abstract

Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi-modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit-aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance-aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry-only structure disentangled from semantics. Building on this design, we introduce a memory-guided video generation model that performs multi-modal feature fusion under reference conditions drawn from mixed-modality memory contexts. Experiments demonstrate that our method maintains strong long-term semantic and structural consistency after edits, significantly outperforming state-of-the-art methods.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis