Geo-Align: Video Generation Alignment via Metric Geometry Reward
AuthorsZizun Li, Haoyu Guo, Runzhe Teng, Chunhua Shen, Tong He
Resources
This paper teaches video generators to follow camera motion more faithfully by rewarding them with geometry-based feedback instead of only supervised examples.
Key results
Wan2.1 backbone used through ReDirector
GRPO rollouts per condition
flow-matching inference during RL sampling
post-training optimization steps
NVIDIA A800 GPUs used for training
best result on DAVIS with ReCamMaster trajectories
What the paper found
Geo-Align is a reinforcement-learning framework for camera-controlled video retake, built on the pretrained Wan2.1 1.3B video model via ReDirector and optimized to follow novel camera trajectories without paired multi-view supervision. Its core novelty is a metric geometry reward: a frozen MapAnything 3D evaluator estimates rotation and translation from generated videos, then penalizes trajectory error with temporally increasing weights so late-frame drift matters more, while VideoAlign and HPSv3 preserve perceptual quality. To solve the scale ambiguity of real-world camera annotations, Geo-Align fuses CityWalk conditioning videos with target trajectories sampled from OmniWorld and rescaled by Truncated Gaussian Sampling into physically plausible motion ranges. Training uses GRPO with MixGRPO-style sliding-window sampling, group size 12, 25 denoising steps, and 140 RL iterations on 64 NVIDIA A800 GPUs for about 130 hours, updating only self-attention layers. On DAVIS with the 10 ReCamMaster trajectory categories, the method improves camera control and visual fidelity over ReDirector, reaching TransErr 0.0129 and RotErr 1.3645 while also raising Dyn-MEt3R to 0.8573; an ablation shows the full reward outperforms an aesthetic-only reward, reducing RotErr from 1.6082 to 1.3895 and improving MEt3R to 0.3082.
Original abstract
Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.