WorldSculpt: Generating Compositional Worlds from Grounded Videos
AuthorsMuyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
Resources
WorldSculpt turns partial multi-view videos into large, editable 3D worlds made of individually generated object meshes, even in heavily cluttered scenes.
Key results
Photorealistic Unreal Engine scenes used for dense compositional evaluation.
Total individually annotated objects across the benchmark.
Total posed rendered views supplied across all benchmark scenes.
Largest scene scale tested by WorldSculpt.
Improvement from learned IBR fusion over arithmetic mean fusion.
What the paper found
WorldSculpt addresses a central limitation of 3D world generation: systems such as Marble and HY-World 2.0 produce explorable scenes, but usually as a fused mesh or Gaussian representation rather than independently editable objects. Given posed RGB images, instance masks, and coarse 3D boxes, WorldSculpt canonicalizes each object around an anchor view, lifts DINOv3 image features into a shared voxel volume, and fuses them with a permutation-invariant IBR-style aggregator. That conditioning is injected into the frozen Pixal3D prior, built on TRELLIS.2, using zero-initialized projections and LoRA adapters; a degradation curriculum trains robustness to occlusion, pose noise, mask errors, and low resolution. Although finetuned only on canonical single-object data, the system generates separate meshes and transforms them back into a common world frame without scene-level training. Its new UE-MeshyScene benchmark contains 6 photorealistic environments, 2299 objects, and 5964 views, with scenes ranging from 93 to 701 objects. WorldSculpt achieves F-scores of 0.995 on HouseCat6D, 0.981 on Toys4k-Scene, and 0.951 on UE-MeshyScene, outperforming prior compositional methods as clutter and occlusion increase. On UE-MeshyScene, learned IBR fusion reduces CD-ℓ2 by 12% versus arithmetic averaging. The approach also converts generated 3DGS worlds from Marble and HY-World 2.0 into object-level mesh scenes, though it currently generates geometry only and assumes static objects plus reliable upstream masks, poses, and localizations.
Original abstract
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.