World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
AuthorsHao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang, Yi Hua, Ben Mildenhall, Christoph Lassner, Narendra Ahuja, Gengshan Yang
Resources
World Tracing is a new way to turn images into 3D by keeping pixels aligned while also hallucinating the hidden parts of objects and scenes.
Key results
WT predicts six multilayer camera-space intersections per pixel.
WT-DiT total model size.
WT-O visible-surface depth on held-out objects.
WT-O visible-surface depth on held-out objects.
WT-S visible-surface depth on held-out 3D-FRONT scenes.
Mix-trained WT-S on NYU Depth V2 at 50 denoising steps.
What the paper found
World Tracing, developed by World Labs and the University of Illinois Urbana-Champaign, reframes image-to-3D as a pixel-aligned multilayer geometry problem: for each input pixel, the model predicts a stack of 6 camera-space XYZ intersections, with layer 0 reconstructing the visible surface and deeper layers completing occluded geometry beyond the visible. The system is instantiated as WT-DiT, a 1.7B-parameter flow-matching diffusion transformer built on a frozen MoGe ViT-L encoder, trained with an XYZ-only depth-filling objective, three-way attention factorization over layer-wise, ray-wise, and global tokens, and a mixed diffusion-time schedule that balances reconstruction and generative completion. On held-out object data, WT-O reaches 0.0149 MAE and 0.0243 RMSE on visible-surface depth while also improving complete geometry to 0.0213 L1, 0.00194 L2, and 0.898 F@0.05; on held-out 3D-FRONT scenes, WT-S posts 0.0102 MAE, 0.0215 RMSE, and 0.9867 F@0.05 at layer 0, plus 0.0216 All-L CD-L1 and 0.8951 All-L F@0.05 for full multilayer geometry. The model also extends to dynamics, achieving a 0.0105 mean CD-L2 across Obj.-Val, Truebone, and ActionBench, and its pixel-aligned representation enables training-free pose-aware textured-mesh generation, text-driven scene editing, and geometry-conditioned novel-view video synthesis. A mix-training ablation on real depth corpora further lifts WT-S to 0.0374 AbsRel on NYU and 0.0332 AbsRel on ETH3D-indoor, showing that the same multilayer architecture can absorb single-layer RGBD supervision without sacrificing occluded-geometry prediction.
Original abstract
Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate complete shapes that are often misaligned with the input. We introduce World Tracing, a generative pixel-aligned geometry representation that predicts 3D points aligned with observed pixels while completing geometry beyond the visible surface. For each input pixel, World Tracing predicts an ordered stack of camera-space 3D points, where the first layer represents the visible surface and subsequent layers represent front-to-back intersections with occluded surfaces. We instantiate this representation with a world-tracing diffusion transformer, WT-DiT, which treats multiple geometry layers as separate denoising tokens coupled through factorized and global attention. WT-DiT is trained with pixel-space flow matching and a mixed noise schedule that balances visible-surface reconstruction with occluded-geometry generation. World Tracing achieves strong performance on visible-surface reconstruction and complete geometry generation across object, scene, and dynamic benchmarks, outperforming both depth predictors and image-to-3D generators. It also preserves 2D-to-3D correspondence, enabling text-driven 3D scene editing, geometry-conditioned novel-view video synthesis, and training-free integration with textured-mesh generators.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.