NTH

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

AuthorsShengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu

AffiliationsUC San Diego · University of Virginia · Meta · Amazon · LambdaProject page: https://jsxzs.github.io/LIFT/

October 9, 2026 2 min read
Watch on YouTube
The one-line take

LIFT lets users guide not just how a video camera moves, but exactly what should appear and where in a future view.

Key results

1.3B
LIFT model size

LIFT is built on the Wan2.1-Fun video model.

84.8%
Unseen-object clip share

Share of LIFT-Vista clips containing objects not visible in the initial frame.

0.51
Layout mIoU

LIFT score, compared with 0.41 for MagicMotion despite using only a final-frame layout.

8K
OPSD adaptation updates

Training sample updates used by LIFT in the sparse-layout adaptation comparison.

128K
SFT adaptation updates

Training sample updates used by the SFT baselines in the same comparison.

65.96%
User preference

Share of non-abstaining user-study choices that preferred LIFT.

What the paper found

LIFT gives image-to-video generation control over both camera movement and the composition of a future view, including regions hidden in the starting image. Users provide a camera trajectory and a final-frame layout of bounding boxes with local text prompts; the model generates the intervening video. Built on the 1.3B-parameter Wan2.1-Fun model, LIFT encodes camera paths as Plücker rays and layout maps as video latents. Its key training innovation is dual-mode on-policy self-distillation: a dense-layout teacher guides a shared student on its own rollout states, training it to work with either a camera path alone or a camera path plus only the final-frame layout. The method focuses distillation on the first 10 high-noise states, when global layout is established, and uses a flow-matching anchor to preserve video quality. The LIFT-Vista dataset draws on RealEstate10K, Sekai, and SpatialVID, with camera and object-layout annotations produced using tools including Depth Anything 3, SAM 3, and Qwen3-VL; 84.8% of its clips contain objects absent from the initial view. Against MagicMotion, which receives dense per-frame layouts, LIFT raises mIoU from 0.41 to 0.51 using only final-frame guidance. In the sparse-layout adaptation comparison, it reaches that result with 8K training sample updates, versus 128K for the SFT baselines. In a user study, LIFT received a 65.96% preference rate, supporting its value for interactive video creation.

Original abstract

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

Read the original paper

More in Generative Models

Browse all 66 papers →
03Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis