NTH

Geo-Align: Video Generation Alignment via Metric Geometry Reward

AuthorsZizun Li, Haoyu Guo, Runzhe Teng, Chunhua Shen, Tong He

June 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches video generators to follow camera motion more faithfully by rewarding them with geometry-based feedback instead of only supervised examples.

Key results

1.3B
base model size

Wan2.1 backbone used through ReDirector

12
group size

GRPO rollouts per condition

25
denoising timesteps

flow-matching inference during RL sampling

140
RL iterations

post-training optimization steps

64
GPU count

NVIDIA A800 GPUs used for training

0.0129
camera accuracy TransErr

best result on DAVIS with ReCamMaster trajectories

What the paper found

Geo-Align is a reinforcement-learning framework for camera-controlled video retake, built on the pretrained Wan2.1 1.3B video model via ReDirector and optimized to follow novel camera trajectories without paired multi-view supervision. Its core novelty is a metric geometry reward: a frozen MapAnything 3D evaluator estimates rotation and translation from generated videos, then penalizes trajectory error with temporally increasing weights so late-frame drift matters more, while VideoAlign and HPSv3 preserve perceptual quality. To solve the scale ambiguity of real-world camera annotations, Geo-Align fuses CityWalk conditioning videos with target trajectories sampled from OmniWorld and rescaled by Truncated Gaussian Sampling into physically plausible motion ranges. Training uses GRPO with MixGRPO-style sliding-window sampling, group size 12, 25 denoising steps, and 140 RL iterations on 64 NVIDIA A800 GPUs for about 130 hours, updating only self-attention layers. On DAVIS with the 10 ReCamMaster trajectory categories, the method improves camera control and visual fidelity over ReDirector, reaching TransErr 0.0129 and RotErr 1.3645 while also raising Dyn-MEt3R to 0.8573; an ablation shows the full reward outperforms an aesthetic-only reward, reducing RotErr from 1.6082 to 1.3895 and improving MEt3R to 0.3082.

Original abstract

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis