NTH

Beyond Pixels: From Video Priors to 4D Worlds

AuthorsZihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang

August 16, 2026 2 min read
Watch on YouTube
The one-line take

Latent-to-4D turns the hidden representations of video generators directly into stable dynamic 3D worlds without passing through RGB.

Key results

1143
Training clips

Annotated reconstruction clips used for final geometry-supervised training.

3
Compatible DiTs

Video diffusion transformers served by one unchanged checkpoint through the Wan VAE.

3.45
Text DINO-F1 gain

Upper end of the improvement over matched Wan-plus-4RC cascades on Text4D-200.

5.81
Image DINO-F1 gain

Improvement over the matched Wan2.2-plus-4RC cascade on I4D-200.

59.2%
Text geometry preference

Human preference for Latent-to-4D geometry fidelity over baselines.

72.1%
Image completeness preference

Human preference for Latent-to-4D completeness over baselines.

What the paper found

Beyond Pixels proposes Latent-to-4D, a method that converts a video generator’s final denoised VAE latent directly into dynamic geometry, cameras, and motion, bypassing the usual RGB decoding and reconstruction cascade. Its Latent-to-4D Alignment and Refinement, or L4AR, first matches the latent grid with a learned 3D convolution, then combines frame-wise and global spatiotemporal attention before a pretrained 4D decoder predicts world-space point maps and camera trajectories. Training uses roughly 1,143 annotated reconstruction clips while freezing the video generators, VAE, and main Transformer weights, allowing one checkpoint to operate across 3 compatible DiTs that share the Wan VAE, including Wan2.1 text-to-video and Wan2.2 image-to-video models. On the locked Text4D-200 and I4D-200 benchmarks, direct latent lifting improves projection-based DINO-F1 over matched Wan-plus-4RC baselines by 2.88–3.45 points for text conditioning and 5.81 points for image conditioning; CogVideoX-5B remains stronger on some RGB-reference metrics, so the gains are concentrated in geometric consistency rather than every measure. Human evaluations favor the method for geometry, completeness, temporal stability, and overall quality, with text-to-4D geometry preference reaching 59.2% and image-to-4D completeness reaching 72.1%. The central result is a reusable latent interface: compatible video models retain their original text, image, motion, pose, trajectory, and navigation controls while sharing a single geometry-supervised 4D pathway, though the evidence is limited to a common-VAE setting and projection metrics do not establish metric accuracy.

Original abstract

4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis