PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
AuthorsShuai Yang, Bingjie Gao, Ziwei Liu, Jiaqi Wang, Dahua Lin, Tong Wu
Resources
PermaVid helps edited videos stay coherent over time by using separate memory for appearance and geometry so later generations don’t drift after changes.
Key results
Synthetic long-video dataset size
Long revisiting trajectories in each video
Number of Unreal Engine scenes
Best semantic consistency on the global-edit benchmark
Best visual quality on the local-edit benchmark
What the paper found
PermaVid, developed by researchers from Shanghai Jiao Tong University, Stanford University, Nanyang Technological University, The Chinese University of Hong Kong, and the Shanghai Innovation Institute, tackles a failure mode in memory-based video generation: after an edit, older context becomes stale and causes the model to revert to pre-edit content. The core idea is to disentangle spatial context into an RGB memory for semantic appearance and a depth memory for geometric structure, then update and retrieve these banks differently for global edits, which invalidate all RGB context, and local edits, which only invalidate overlapping regions. Built on Wan2.1-14B with a VACE-initialized context branch, the system fuses mixed-modality references through a DiT video generator and trains in two stages on SpatialVid and the new UE-Mem dataset. UE-Mem contains 4k long videos with 1000 frames each across 100 Unreal Engine scenes, captured with revisiting trajectories and 6-DoF poses. On a 200-image benchmark with complex revisiting camera paths, PermaVid achieves the best global-edit scores, including 22.84 PSNR, 0.8703 SSIM, 0.2102 LPIPS, and 27.87 CLIP-Vid, and also leads local-edit evaluation with 22.332 PSNR, 0.8622 SSIM, 0.2369 LPIPS, and 0.8544 VBench-Avg. The paper’s ablation shows that removing disentangled context memory causes outdated RGB reuse and progressive semantic inconsistency, while the proposed retrieval overhead remains only at the millisecond level.
Original abstract
Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi-modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit-aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance-aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry-only structure disentangled from semantics. Building on this design, we introduce a memory-guided video generation model that performs multi-modal feature fusion under reference conditions drawn from mixed-modality memory contexts. Experiments demonstrate that our method maintains strong long-term semantic and structural consistency after edits, significantly outperforming state-of-the-art methods.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.