NTH

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

AuthorsZhenhao Yang, Xiaoshi Wu, Zhengyao Lv, Xiaoyu Shi, Xintao Wang, Pengfei Wan, Kun Gai, Kwan-Yee K. Wong

June 16, 2026 2 min read
Watch on YouTube
The one-line take

DecMem introduces a split memory system to help video generators remember long-range details, aiming to produce more consistent minute-long worlds.

Key results

1B
Base model size

Pretrained video generation backbone used for DecMem

221
Initial memory frames

Memory bank initialization for quantitative evaluation

80
Best top-k retrieval blocks

Chosen Sparse Global Memory retrieval setting

3.6496
Inference speed

Frames per second reported for DecMem

What the paper found

DecMem from the Kling Team at Kuaishou Technology and the University of Hong Kong targets minute-long consistent world generation by replacing naïve dense long-context attention with a decoupled memory design. The core idea is to solve two failure modes in long-horizon video rollout—attention dispersion and linear cost growth—using Sparse Global Memory, which performs block-level top-k retrieval over history, and Anchored Local Memory, which restricts attention to the most recent 8 frames to stabilize short-range transitions. Built on a 1B latent diffusion transformer and trained on the WorldMem dataset, DecMem uses multimodal rotary position embeddings for camera pose, patch location, and time, then fuses global and local branches with a learned gate. On the WorldMem benchmark with 221 initial memory frames and 120 generated frames, it reaches 30.0785 PSNR, 0.0494 LPIPS, 9.8904 FID within the training window, and 25.2294 PSNR, 0.1006 LPIPS, 16.2667 FID under extrapolation, outperforming Oasis, MineWorld, and WorldMem. It also achieves 3.6496 FPS, nearly 2× faster than the strongest baseline, and in a 58-participant user study obtains 39.77% visual quality, 37.81% action controllability, and 42.12% spatio-temporal consistency preference. Ablations show that top-k = 80 is the best retrieval setting overall, while removing either Sparse Global Memory or Anchored Local Memory sharply degrades long-range consistency.

Original abstract

Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains a key challenge. In this work, we move beyond explicit 3D memory and coarse frame-level implicit modeling, and propose a fine-grained, learnable, and scalable memory for consistent world generation. We first identify two fundamental limitations of naïve learnable memory architectures in long-horizon extrapolation, namely computational inefficiency and attention dispersion. Through a systematic analysis of attention dispersion, we propose DecMem, a decoupled memory architecture that employs Sparse Global Memory for efficient fine-grained access to global history and Anchored Local Memory for stable and high-quality extrapolation. Extensive experiments demonstrate that DecMem significantly outperforms current state-of-the-art methods. By ensuring precise and efficient long-term memory and achieving superior extrapolation capabilities, DecMem enables minute-level controllable long video generation with high fidelity and consistency.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis