VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
AuthorsZhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
Resources
VideoRAE turns video foundation-model features into efficient, semantically rich latent spaces that can accelerate both diffusion and autoregressive video generation.
Key results
Class-to-video generation with VideoRAE using V-JEPA 2 features.
Class-to-video generation with the continuous VideoRAE latent.
Approximate fold improvement over competing autoencoder baselines.
VideoRAE result in the 2B-scale text-to-video experiment.
Diffusion-transformer scale used for the controlled VideoUFO experiment.
Reconstruction score achieved with V-JEPA 2 features.
What the paper found
VideoRAE, from researchers at The Chinese University of Hong Kong, Shenzhen and collaborating institutions, repurposes frozen video foundation models for generation rather than using them only for understanding. It extracts hierarchical features from Meta’s V-JEPA 2 or VideoMAEv2, fuses them, and compresses them with a lightweight 1D self-attention projector. A local-and-global Representation Alignment, or REPA, objective preserves the teacher’s semantic structure during decoding, eliminating the need for KL regularization. The same latent framework supports continuous representations for Diffusion Transformers and discrete representations for autoregressive models through high-dimensional Multi-Codebook SimVQ. On UCF-101, VideoRAE reaches a class-to-video gFVD of 40 with an autoregressive generator and 93 with a DiT, while discrete reconstruction achieves an rFVD of 13 using V-JEPA 2 features. These semantic latents let downstream generators converge approximately 5× faster than competing autoencoder baselines. In a controlled 2B-scale text-to-video experiment built on OpenSora2 and trained on VideoUFO with NVIDIA H100 GPUs, replacing LTX-VAE with VideoRAE raises the VBench total score to 71.26, compared with 69.35 for LTX-VAE, while converging faster. The results argue that semantically structured latents, not maximum pixel-level fidelity, are the more effective foundation for scalable video generation.
Original abstract
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can limit the semantic and spatio-temporal structure captured by their latents. Meanwhile, Video Foundation Models (VFMs) such as V-JEPA 2 and VideoMAEv2 show strong video understanding capabilities, yet whether their frozen representations can be transformed into compact, reconstruction-capable, and generation-friendly video latents remains largely unexplored. We answer this question with VideoRAE, a representation autoencoder that leverages multi-scale hierarchical features from a frozen video foundation encoder and compresses them with a lightweight 1D self-attention projector. VideoRAE supports both continuous latents for Diffusion Transformers and discrete tokens for autoregressive models via multi-codebook high-dimensional quantization. During decoding, a local-and-global representation alignment objective with the frozen VFM teacher improves semantic preservation and enables training without KL regularization. Experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, it obtains state-of-the-art class-to-video gFVDs of 40 and 93 with AR and DiT generators, respectively, while converging approximately 5x faster than competing autoencoder baselines. In a controlled 2B-scale text-to-video study, replacing LTX-VAE with VideoRAE leads to faster convergence under comparable settings. These results validate frozen VFM representations as versatile and generation-friendly video latents. The model and code will be released on https://zhxie0117.github.io/VideoRAE.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.