V-RAE: Rethinking Video Latent Spaces for Generation
AuthorsMinghui Guo, Shengqiong Wu, Hao Fei
Resources
V-RAE makes video generation more efficient and semantically meaningful by compressing videos into latents derived from pretrained vision models rather than pixel-focused autoencoders.
Key results
V-RAE with V-JEPA 2.1 achieves 2.13 rFVD on K600.
The best V-RAE variant reaches 117.86 gFVD on UCF101.
The best V-RAE variant reaches 19.16 gFVD on K600.
V-RAE converges up to 6× faster than VAE-based latent spaces.
V-RAE with DINOv3-L achieves 89.13% top-1 accuracy on UCF101.
tFVD correlates with downstream gFVD at 0.919 on K600.
What the paper found
V-RAE rethinks video generation by replacing reconstruction-oriented VAE latents with semantically structured features from frozen vision encoders, including DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal-attention pooler removes redundancy, while a spatiotemporal Transformer decoder with 3D RoPE reconstructs temporally coherent video; only the pooler and decoder are trained. Using a DiT with rectified flow, V-RAE outperforms latent spaces from Wan2.2 VAE, CogVideoX VAE, HunyuanVideo VAE, Cosmos VAE, and AToken: the V-JEPA 2.1 variant reaches 2.13 rFVD on K600 and generation gFVD scores of 117.86 on UCF101 and 19.16 on K600, while converging up to 6× faster. Its compressed latents also preserve semantics, achieving 89.13% UCF101 top-1 accuracy with DINOv3-L versus 30.83% for the strongest VAE baseline. The paper shows that reconstruction quality is a poor proxy for generative utility: V-RAE introduces tFVD, which tests whether interpolated latent trajectories decode into coherent motion, and its correlation with downstream gFVD reaches 0.919 on K600, compared with 0.473 for rFVD. On Cityscapes future prediction, V-RAE further reduces gFVD to 111.36 versus 144.47 for Wan2.2 VAE, suggesting that semantic, temporally smooth latents support both generation and predictive world modeling.
Original abstract
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.