FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
AuthorsBo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan
AffiliationsNational University of Singapore · Harbin Institute of Technology (Shenzhen)
FrameMorrow helps long-video generators remember the right past frames by predicting what future content will need next.
Key results
Text-conditioned selector training data.
Point improvement from 84.95 to 89.30.
Observed across the tested long-video backbones.
What the paper found
FrameMorrow tackles long-horizon video drift by selecting past frames for what a generator will need next, rather than merely matching the current view. A causal Transformer predicts four prospective tokens from eligible history, recent context, and the known prompt or action; token-to-frame scores then retrieve four explicit historical frames. During training, Meta’s frozen DINOv2 compares history with realized future frames to provide listwise and pairwise ranking supervision, while inference uses no future observations. The selector plugs into each generator’s existing conditioning interface, including closed-source Seedance 2.0 and Kling O3. Evaluated across five benchmarks and 11 generative models, FrameMorrow was trained on 10K OpenVidHD videos and improved Self-Forcing consistency from 84.95 to 89.30, a gain of 4.35 points. On YuMe 1.5, action alignment rose from 0.883 to 0.912, while measured inference overhead was at most 2.1%. The results suggest a lightweight way to preserve relevant visual details and action alignment across extended generation without modifying the underlying generator.
Original abstract
Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.
Read the original paperMore in Generative Models
Browse all 66 papers →Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Jiangshan Wang, Zeqiang Lai, Jiayi Guo, Xin Yang, Xin Huang, Jiarui Chen, Ziheng Ouyang, Chunchao Guo, Xiangyu Yue
Tex-Zero shows that high-quality native 3D textures may be learned from cleverly structured 2D images instead of costly real 3D assets.
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu
LIFT lets users guide not just how a video camera moves, but exactly what should appear and where in a future view.
RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.