NTH

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

AuthorsZhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang

July 25, 2026 2 min read
Watch on YouTube
The one-line take

VideoRAE turns video foundation-model features into efficient, semantically rich latent spaces that can accelerate both diffusion and autoregressive video generation.

Key results

40
UCF-101 discrete AR gFVD

Class-to-video generation with VideoRAE using V-JEPA 2 features.

93
UCF-101 continuous DiT gFVD

Class-to-video generation with the continuous VideoRAE latent.

5
Convergence acceleration

Approximate fold improvement over competing autoencoder baselines.

71.26
VBench total score

VideoRAE result in the 2B-scale text-to-video experiment.

2B
Text-to-video model scale

Diffusion-transformer scale used for the controlled VideoUFO experiment.

13
UCF-101 discrete reconstruction rFVD

Reconstruction score achieved with V-JEPA 2 features.

What the paper found

VideoRAE, from researchers at The Chinese University of Hong Kong, Shenzhen and collaborating institutions, repurposes frozen video foundation models for generation rather than using them only for understanding. It extracts hierarchical features from Meta’s V-JEPA 2 or VideoMAEv2, fuses them, and compresses them with a lightweight 1D self-attention projector. A local-and-global Representation Alignment, or REPA, objective preserves the teacher’s semantic structure during decoding, eliminating the need for KL regularization. The same latent framework supports continuous representations for Diffusion Transformers and discrete representations for autoregressive models through high-dimensional Multi-Codebook SimVQ. On UCF-101, VideoRAE reaches a class-to-video gFVD of 40 with an autoregressive generator and 93 with a DiT, while discrete reconstruction achieves an rFVD of 13 using V-JEPA 2 features. These semantic latents let downstream generators converge approximately 5× faster than competing autoencoder baselines. In a controlled 2B-scale text-to-video experiment built on OpenSora2 and trained on VideoUFO with NVIDIA H100 GPUs, replacing LTX-VAE with VideoRAE raises the VBench total score to 71.26, compared with 69.35 for LTX-VAE, while converging faster. The results argue that semantically structured latents, not maximum pixel-level fidelity, are the more effective foundation for scalable video generation.

Original abstract

Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can limit the semantic and spatio-temporal structure captured by their latents. Meanwhile, Video Foundation Models (VFMs) such as V-JEPA 2 and VideoMAEv2 show strong video understanding capabilities, yet whether their frozen representations can be transformed into compact, reconstruction-capable, and generation-friendly video latents remains largely unexplored. We answer this question with VideoRAE, a representation autoencoder that leverages multi-scale hierarchical features from a frozen video foundation encoder and compresses them with a lightweight 1D self-attention projector. VideoRAE supports both continuous latents for Diffusion Transformers and discrete tokens for autoregressive models via multi-codebook high-dimensional quantization. During decoding, a local-and-global representation alignment objective with the frozen VFM teacher improves semantic preservation and enables training without KL regularization. Experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, it obtains state-of-the-art class-to-video gFVDs of 40 and 93 with AR and DiT generators, respectively, while converging approximately 5x faster than competing autoencoder baselines. In a controlled 2B-scale text-to-video study, replacing LTX-VAE with VideoRAE leads to faster convergence under comparable settings. These results validate frozen VFM representations as versatile and generation-friendly video latents. The model and code will be released on https://zhxie0117.github.io/VideoRAE.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis