Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
AuthorsJiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou
Resources
MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.
Key results
Dataset size for multi-grid post-training.
More shots generated under the same token budget in a 1,616-frame video.
MovieGrid benchmark score.
MovieGrid benchmark score.
Frames generated by increasing the grid count from 16 to 64.
What the paper found
MovieGrid introduces a multi-grid post-training method for long-form multi-shot video generation. Instead of packing every shot onto one temporal axis, it divides a narrative into short video chunks, arranges them chronologically in a spatial grid, and jointly models the grid so each local temporal axis handles fewer transitions. Its MGLV dataset contains 54,281 grid videos derived from 1,000 long-form source videos, with character-aware story prompts generated using Qwen3-VL 8B. Training adapts the Wan2.2-5B backbone with Noise-Free Random-Grid Training, Grid Embedding, shared character tags, and Grid Boundary Loss; the method preserves selected chunks as clean visual context while denoising the others. Under the same token budget, MovieGrid generates 6.05 times more shots in a 1,616-frame video than Temporal Packing, and reaches 0.8970 intra-shot subject consistency and 0.6139 inter-shot subject consistency, exceeding the strongest cited baselines. Increasing the grid count from 16 to 64 produces 6,464 frames in one generation, while conditioning successive grids enables further story continuation. The system was trained on NVIDIA B200 GPUs, and Google is acknowledged as a supporter.
Original abstract
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.