NTH

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

AuthorsYudong Jin, Tao Xie, Qihang Zhang, Zehong Shen, Zhen Xu, Yujun Shen, Hujun Bao, Xiaowei Zhou, Yinghao Xu

August 26, 2026 2 min read
Watch on YouTube
The one-line take

4DAnyone turns an ordinary monocular human video into a consistent, renderable 4D reconstruction by coordinating diffusion-generated views across time and camera angles.

Key results

5B
Base diffusion model size

Wan2.2-TI2V-5B is the video diffusion backbone.

40
Skeleton keypoints

The 3D-aware conditioning uses a compact 40-keypoint subset.

38k
MVGameHuman videos

MVGameHuman contributes 38k synchronized multi-view training videos.

24.33
Generated-video consistency PSNR

4DAnyone's score on DNA-Rendering.

24.15
4DGS reconstruction PSNR

Downstream FreeTimeGS reconstruction score on DNA-Rendering.

What the paper found

4DAnyone reconstructs a dynamic human avatar from a casual monocular video, even when camera intrinsics and poses are unknown, by generating reconstruction-grade videos from 16 target viewpoints and lifting them into 4D Gaussian Splatting with FreeTimeGS. Built on the 5B-parameter Wan2.2-TI2V-5B video diffusion model, it uses GVHMR to estimate 3D motion and conditions generation with depth-buffered, occlusion-aware skeletons based on 40 keypoints rather than unreliable dense depth. Its main contribution is scalable cross-view consistency: Reference Context Packing, or RCP, compresses accumulated reference views into a fixed-length context with O(1) complexity, while Target Context Routing, or TCR, cyclically regroups four-view target groups during high-noise denoising to propagate global structure, then fixes adjacent groups for low-noise detail refinement. Training combines the 38k-video MVGameHuman dataset with SynCamVideo, DNA-Rendering, Pexels, and TedTalk data. On DNA-Rendering, 4DAnyone reaches 24.33 PSNR for generated-video consistency, compared with 21.47 for the fine-tuned ReCamMaster baseline, and its downstream 4DGS reconstruction reaches 24.15 PSNR; it also generalizes to the out-of-distribution DyMVHumans benchmark. Ablations show that removing both RCP and TCR lowers consistency to 21.09 PSNR, while the full system handles diverse in-the-wild footage, though loose garments and incorrect pose estimates remain failure cases.

Original abstract

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as $O(N)$, weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with $O(1)$ reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis