Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
AuthorsYuxuan Bian, Zeyue Xue, Songchun Zhang, Shiyi Zhang, Weiyang Jin, Yaowei Li, Junhao Zhuang, Haoran Li, Jie Huang, Haoyang Huang, Nan Duan, Qiang Xu
Resources
This paper introduces a new memory system that lets video generators keep producing coherent frames for extremely long durations in real time, pushing toward truly infinite video generation.
Key results
Echo-Infinity reaches 18.5 FPS on the Wan2.1-T2V-1.3B backbone while generating long videos in real time.
The paper reports promising real-time rollouts over 24 hours with more than 1.3M frames generated.
On 30-second long-video evaluation, Echo-Infinity achieves 59.53% user preference, compared with 14.73% for Memorize-and-Generate.
On 240-second long-video evaluation, Echo-Infinity achieves 71.67% user preference, compared with 14.13% for ∞-RoPE.
What the paper found
Echo-Infinity, developed by researchers at The Chinese University of Hong Kong, Joy Future Academy at JD, Tsinghua, HKUST, HKU, Peking University, and USTC, tackles two bottlenecks in autoregressive video diffusion transformers: unbounded KV-cache growth and temporal RoPE overflow. The key idea is a learnable evolving memory built from Memory Queries, trainable tokens that are updated whenever old frames are evicted from the local window through cross-attention plus a sigmoid-gated residual, so the model filters, abstracts, and compresses arbitrary-length history at constant cost instead of relying on fixed window truncation or heuristic compression. Echo-Infinity also introduces a Unified Relative RoPE Recipe that anchors sink frames at temporal id 0 and keeps all active indices within the pretrained range during both training and inference, eliminating the train-test mismatch seen in prior test-time-only RoPE fixes such as ∞-RoPE and MemRoPE. Built on the Wan2.1-T2V-1.3B backbone with Causal Forcing and DMD distillation, the system reaches 18.5 FPS while generating more than 1.3 million frames over 24 hours on a single NVIDIA H100. On VBench-Long and MovieGen, it improves 30-second user preference to 59.53 percent versus 14.73 percent for Memorize-and-Generate, and 240-second user preference to 71.67 percent versus 14.13 percent for ∞-RoPE, while also achieving state-of-the-art 5-second and 60-second interactive results. Ablations show both the memory queries and unified relative RoPE are essential, with removing memory queries sharply reducing consistency and dynamic degree.
Original abstract
We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length history at constant cost. Existing methods mainly curate memory with predefined KV-cache schedules, fixed-ratio heuristic compression, or inference-time RoPE adaptation. These designs inevitably lose historical information and amplify compounding errors due to their limited cache window and ignorance of autoregressive generation noise. Inspired by human memory consolidation, Echo-Infinity replaces handcrafted memory curation with learnable Memory Query, which are updated by attention and a gating mechanism when past frames are evicted from the local window. The queries are optimized end-to-end with the video diffusion transformers (DiTs), forming an evolving memory that supports arbitrary compression ratios with constant computation independent of video length. They also act as a generalizable generation prior, improving quality even when only the optimized initial state is used. We further introduce Unified Relative RoPE Recipe, which anchors the sink frames to start from id 0 and lets the newest frame id grow at most to the DiTs' pretrained maximum temporal RoPE id throughout training and inference, freeing the model from the finite RoPE constraint and closing the train-test RoPE extrapolation gap. In long and short video generation, Echo-Infinity achieves state-of-the-art performance, and, to our knowledge, demonstrates promising 24-hour (>1.3 M frames) real-time rollouts for the first time, suggesting a practical path toward infinite video generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.