NTH

WorldSculpt: Generating Compositional Worlds from Grounded Videos

AuthorsMuyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang

September 7, 2026 3 min read
Watch on YouTube
The one-line take

WorldSculpt turns partial multi-view videos into large, editable 3D worlds made of individually generated object meshes, even in heavily cluttered scenes.

Key results

6
UE-MeshyScene environments

Photorealistic Unreal Engine scenes used for dense compositional evaluation.

2299
UE-MeshyScene objects

Total individually annotated objects across the benchmark.

5964
UE-MeshyScene views

Total posed rendered views supplied across all benchmark scenes.

701
UE-MeshyScene maximum objects per scene

Largest scene scale tested by WorldSculpt.

12%
UE-MeshyScene IBR CD-ℓ2 reduction

Improvement from learned IBR fusion over arithmetic mean fusion.

What the paper found

WorldSculpt addresses a central limitation of 3D world generation: systems such as Marble and HY-World 2.0 produce explorable scenes, but usually as a fused mesh or Gaussian representation rather than independently editable objects. Given posed RGB images, instance masks, and coarse 3D boxes, WorldSculpt canonicalizes each object around an anchor view, lifts DINOv3 image features into a shared voxel volume, and fuses them with a permutation-invariant IBR-style aggregator. That conditioning is injected into the frozen Pixal3D prior, built on TRELLIS.2, using zero-initialized projections and LoRA adapters; a degradation curriculum trains robustness to occlusion, pose noise, mask errors, and low resolution. Although finetuned only on canonical single-object data, the system generates separate meshes and transforms them back into a common world frame without scene-level training. Its new UE-MeshyScene benchmark contains 6 photorealistic environments, 2299 objects, and 5964 views, with scenes ranging from 93 to 701 objects. WorldSculpt achieves F-scores of 0.995 on HouseCat6D, 0.981 on Toys4k-Scene, and 0.951 on UE-MeshyScene, outperforming prior compositional methods as clutter and occlusion increase. On UE-MeshyScene, learned IBR fusion reduces CD-ℓ2 by 12% versus arithmetic averaging. The approach also converts generated 3DGS worlds from Marble and HY-World 2.0 into object-level mesh scenes, though it currently generates geometry only and assumes static objects plus reliable upstream masks, poses, and localizations.

Original abstract

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis