SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
AuthorsZhida Zhang, Jie Ma, Zhan Peng, Haoxue Wu, Yang Han, Jun Liang, Jie Cao, Jing Li
Resources
SmartDirector lets video models generate more cinematic, better-paced stories by using multiple keyframes as narrative anchors across single-shot, multi-shot, and video-extension settings.
Key results
Internal diffusion model used as the base for the generation stage.
Wan-2.2-5B used as the super-resolution backbone.
Single-shot videos and 250 multi-shot videos in the evaluation benchmark.
SmartDirector achieves this FVD in the single-shot scenario.
SmartDirector achieves this FVD in the multi-shot scenario.
Human evaluation win rate for SmartDirector in overall quality.
What the paper found
SmartDirector is a two-stage cinematic video generation framework for arbitrary keyframe conditioning that targets single-shot synthesis, multi-shot narrative generation, and video extension. Built on an internal 32B diffusion backbone for Director-Gen and Wan-2.2-5B for Director-SR, it replaces direct latent keyframe insertion with a Multi-Chunk VAE that encodes each keyframe as the first frame of its own chunk, then applies full spatio-temporal attention in a DiT and a custom MC-RoPE temporal scheme to avoid boundary artifacts. The second stage uses high-resolution keyframes as semantic anchors to super-resolve low-resolution outputs, recovering faces, text, and other fine details. Trained on curated cinematic footage from copyright-free movies plus UltraVideo for super-resolution, SmartDirector is evaluated on a benchmark of 250 single-shot and 250 multi-shot videos at 24 FPS and native 1080p or higher. It sharply outperforms Dreamina Multiframes, cutting FVD from 226.85 to 41.12 in single-shot and from 251.83 to 65.65 in multi-shot, while raising Gemini-3-Pro semantic scores from 83.87 to 91.30 and from 59.32 to 88.48, with narrative coherence gains especially large. In human study, 30 participants judged 500 video pairs, and SmartDirector achieved a 54.73% win rate in multi-shot overall quality. As a standalone super-resolution method, it also reduces LPIPS versus SparkVSR on all four benchmarks, showing that keyframe-conditioned generation and restoration can be unified in one narrative-aware pipeline; related model and evaluation infrastructure comes from Google Gemini, Qwen3-VL, and NVIDIA training on 40 GPUs for 20,000 steps.
Original abstract
The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantly rely on sparse conditioning signals such as text prompts or first/last frames, which limits precise control over narrative structure and temporal pacing. In this paper, we propose SmartDirector, a framework that enhances the narrative capacity of video generation models through multiple keyframes. SmartDirector supports flexible generation scenarios including single-shot generation, multi-shot narrative synthesis, and video extension. The framework operates in two stages: Director-Gen generates a low-resolution video conditioned on the provided keyframes, and Director-SR refines the output by exploiting high-resolution keyframes as semantic anchors to recover fine-grained details. To enable robust multi-keyframe training, we construct a data pipeline that curates single-shot and multi-shot sequences from movies. Extensive experiments demonstrate that SmartDirector substantially outperforms existing state-of-the-art approaches. We will release the code to facilitate further research.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.