DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
AuthorsZhenhao Yang, Xiaoshi Wu, Zhengyao Lv, Xiaoyu Shi, Xintao Wang, Pengfei Wan, Kun Gai, Kwan-Yee K. Wong
Resources
DecMem introduces a split memory system to help video generators remember long-range details, aiming to produce more consistent minute-long worlds.
Key results
Pretrained video generation backbone used for DecMem
Memory bank initialization for quantitative evaluation
Chosen Sparse Global Memory retrieval setting
Frames per second reported for DecMem
What the paper found
DecMem from the Kling Team at Kuaishou Technology and the University of Hong Kong targets minute-long consistent world generation by replacing naïve dense long-context attention with a decoupled memory design. The core idea is to solve two failure modes in long-horizon video rollout—attention dispersion and linear cost growth—using Sparse Global Memory, which performs block-level top-k retrieval over history, and Anchored Local Memory, which restricts attention to the most recent 8 frames to stabilize short-range transitions. Built on a 1B latent diffusion transformer and trained on the WorldMem dataset, DecMem uses multimodal rotary position embeddings for camera pose, patch location, and time, then fuses global and local branches with a learned gate. On the WorldMem benchmark with 221 initial memory frames and 120 generated frames, it reaches 30.0785 PSNR, 0.0494 LPIPS, 9.8904 FID within the training window, and 25.2294 PSNR, 0.1006 LPIPS, 16.2667 FID under extrapolation, outperforming Oasis, MineWorld, and WorldMem. It also achieves 3.6496 FPS, nearly 2× faster than the strongest baseline, and in a 58-participant user study obtains 39.77% visual quality, 37.81% action controllability, and 42.12% spatio-temporal consistency preference. Ablations show that top-k = 80 is the best retrieval setting overall, while removing either Sparse Global Memory or Anchored Local Memory sharply degrades long-range consistency.
Original abstract
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains a key challenge. In this work, we move beyond explicit 3D memory and coarse frame-level implicit modeling, and propose a fine-grained, learnable, and scalable memory for consistent world generation. We first identify two fundamental limitations of naïve learnable memory architectures in long-horizon extrapolation, namely computational inefficiency and attention dispersion. Through a systematic analysis of attention dispersion, we propose DecMem, a decoupled memory architecture that employs Sparse Global Memory for efficient fine-grained access to global history and Anchored Local Memory for stable and high-quality extrapolation. Extensive experiments demonstrate that DecMem significantly outperforms current state-of-the-art methods. By ensuring precise and efficient long-term memory and achieving superior extrapolation capabilities, DecMem enables minute-level controllable long video generation with high fidelity and consistency.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.