Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views
AuthorsMijin Yoo, In Cho, Subin Jeon, Jiwoo Lee, Eunbyung Park, Seon Joo Kim
Resources
This paper turns 3D scenes into object-level tokens instead of raw point clouds or Gaussians, enabling reconstruction, segmentation, editing, and retrieval from unposed images.
Key results
feature lifting with 2 context views
feature lifting with 2 context views
novel-view reconstruction with 2 context views
novel-view reconstruction with 2 context views
novel-view reconstruction with 2 context views
class-agnostic instance segmentation with 8 context views
What the paper found
Scenes as Objects, Not Primitives proposes a feed-forward 3D reconstruction framework from unposed multi-view images that represents a scene as instance-structured 3D token groups rather than dense per-Gaussian primitives. Built on the pretrained VGGT geometry foundation model and evaluated on ScanNet, the method uses an image-anchor transformer to decode anchor tokens that spawn 3D Gaussians, then an anchor-grouping transformer to produce instance tokens that compete for anchor ownership via softmax assignment. Training relies only on 2D supervision: RGB rendering loss for reconstruction and Hungarian-matched instance-mask supervision for grouping, with no 3D annotations. The key novelty is a two-level semantic factorization: each instance stores a shared 512-dimensional group embedding, while 8-dimensional anchor residuals capture local variation, reducing feature storage from 8.4M scalars in Uni3R to 59.4K. On ScanNet with 2 context views, the model reaches 0.661 source-view mIoU and 0.657 target-view mIoU for feature lifting, while keeping reconstruction competitive at 25.28 PSNR, 0.771 SSIM, and 0.238 LPIPS. With 8 context views, it achieves the best class-agnostic instance segmentation, reaching 0.235 AP, 0.438 AP50, and 0.564 AP25, outperforming per-scene optimization baselines such as Gaussian Grouping and ObjectGS. The same token groups also enable direct instance-level editing and open-vocabulary retrieval, and the paper reports zero-shot transfer to MipNeRF360 and robustness on RealEstate10K using SAM2 pseudo-labels.
Original abstract
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that decomposes a scene into instance-structured 3D token groups directly from unposed multi-view images -- compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group pairs an instance token capturing entity-level identity with anchor tokens that encode local geometry and appearance, which are decoded into a set of 3D Gaussians. This two-level factorization decouples object identity from local appearance, making object instances a native interface of the representation rather than a derived product. The token groups are learned through differentiable rendering with joint reconstruction and segmentation supervision, requiring no 3D annotations. Our feed-forward model surpasses per-scene optimization baselines in class-agnostic instance segmentation while remaining competitive in novel view synthesis. Beyond these metrics, the same token groups directly unlock instance-level scene editing -- removing, translating, or inserting objects by operating on their groups -- as well as efficient open-vocabulary 3D instance retrieval, where retrieval complexity scales with the number of instances rather than primitives.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.