FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
AuthorsYu Lu, Junjie Yang, Piotr Koniusz, YuXin Song, Yi Yang
Resources
FadeMem makes long video generation more memory-efficient by keeping recent details crisp while compressing older context into a hierarchical cache that preserves coherence.
Key results
default number of historical KV entries per layer
default power-law distance warp used for memory consolidation
inference-time memory variant on the 60-second benchmark
lightly fine-tuned variant on the 60-second benchmark
best visual stability score reported for FadeMem
What the paper found
FadeMem, developed by researchers at Zhejiang University, UNSW, Data61/CSIRO, and Baidu Inc., is a distance-aware KV memory consolidation method for autoregressive video diffusion that targets the core bottleneck of long-rollout generation: bounded caches. The paper shows that frame correlations decay with temporal distance in a frequency-dependent way, with fine textures and local motion fading faster than scene layout and identity, and fits this behavior with an approximate power law, r*(t) ∝ t^-b. FadeMem turns that observation into a single ordered memory of at most M entries, defaulting to M = 12, where new KV blocks are inserted as fine-grained recent entries and older adjacent entries are progressively merged into coarser span-level anchors using a power-law temporal warp with β = 0.3. On the 60-second MovieGenBench setting at 480 × 832 and 16 FPS, the inference-only variant, FadeMem-TF, achieves a VBench-Long average of 80.45, while light fine-tuning lifts FadeMem-FT to 81.03, surpassing LongLive’s 80.55 and improving subject consistency, background consistency, motion smoothness, and imaging quality. In VLM-based evaluation with Gemini 3.1-Pro, FadeMem reaches the best visual stability score at 4.84. Ablations show that weighted-average consolidation outperforms nearest-state selection and max pooling, and that removing the first-frame anchor drops the average to 79.67, confirming that the method’s dense-near, sparse-far cache hierarchy is the key to preserving long-range identity and scene coherence under a fixed memory budget.
Original abstract
Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length. Existing bounded-cache methods reduce this cost with local windows, sink tokens, or compressed memory states, yet they usually assign fixed roles to different parts of the history. We propose FadeMem, a distance-aware KV memory consolidation mechanism that organizes historical KV blocks into a temporal hierarchy under a fixed cache budget. This design is motivated by frequency-dependent temporal decay: fine details decorrelate quickly, while coarse scene structure and identity remain useful over longer horizons. During generation, new history is inserted as fine-grained entries, while older adjacent entries are progressively merged under a power-law temporal allocation schedule, yielding a dense-near, sparse-far memory within one cache. Without architectural changes, FadeMem preserves recent context for short-term dynamics and compact long-range anchors for identity and scene coherence. Experiments show improved subject consistency, background stability, and temporal coherence over existing bounded-cache strategies.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.