NTH

V-RAE: Rethinking Video Latent Spaces for Generation

AuthorsMinghui Guo, Shengqiong Wu, Hao Fei

August 16, 2026 2 min read
Watch on YouTube
The one-line take

V-RAE makes video generation more efficient and semantically meaningful by compressing videos into latents derived from pretrained vision models rather than pixel-focused autoencoders.

Key results

2.13
K600 reconstruction rFVD

V-RAE with V-JEPA 2.1 achieves 2.13 rFVD on K600.

117.86
UCF101 generation gFVD

The best V-RAE variant reaches 117.86 gFVD on UCF101.

19.16
K600 generation gFVD

The best V-RAE variant reaches 19.16 gFVD on K600.

6×
Convergence speedup

V-RAE converges up to 6× faster than VAE-based latent spaces.

89.13%
UCF101 semantic probing

V-RAE with DINOv3-L achieves 89.13% top-1 accuracy on UCF101.

0.919
K600 tFVD correlation

tFVD correlates with downstream gFVD at 0.919 on K600.

What the paper found

V-RAE rethinks video generation by replacing reconstruction-oriented VAE latents with semantically structured features from frozen vision encoders, including DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal-attention pooler removes redundancy, while a spatiotemporal Transformer decoder with 3D RoPE reconstructs temporally coherent video; only the pooler and decoder are trained. Using a DiT with rectified flow, V-RAE outperforms latent spaces from Wan2.2 VAE, CogVideoX VAE, HunyuanVideo VAE, Cosmos VAE, and AToken: the V-JEPA 2.1 variant reaches 2.13 rFVD on K600 and generation gFVD scores of 117.86 on UCF101 and 19.16 on K600, while converging up to 6× faster. Its compressed latents also preserve semantics, achieving 89.13% UCF101 top-1 accuracy with DINOv3-L versus 30.83% for the strongest VAE baseline. The paper shows that reconstruction quality is a poor proxy for generative utility: V-RAE introduces tFVD, which tests whether interpolated latent trajectories decode into coherent motion, and its correlation with downstream gFVD reaches 0.919 on K600, compared with 0.473 for rFVD. On Cityscapes future prediction, V-RAE further reduces gFVD to 111.36 versus 144.47 for Wan2.2 VAE, suggesting that semantic, temporally smooth latents support both generation and predictive world modeling.

Original abstract

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis