NTH

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

AuthorsMijin Yoo, In Cho, Subin Jeon, Jiwoo Lee, Eunbyung Park, Seon Joo Kim

July 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns 3D scenes into object-level tokens instead of raw point clouds or Gaussians, enabling reconstruction, segmentation, editing, and retrieval from unposed images.

Key results

0.661
ScanNet source-view mIoU

feature lifting with 2 context views

0.657
ScanNet target-view mIoU

feature lifting with 2 context views

25.28
ScanNet PSNR

novel-view reconstruction with 2 context views

0.771
ScanNet SSIM

novel-view reconstruction with 2 context views

0.238
ScanNet LPIPS

novel-view reconstruction with 2 context views

0.235
AP

class-agnostic instance segmentation with 8 context views

What the paper found

Scenes as Objects, Not Primitives proposes a feed-forward 3D reconstruction framework from unposed multi-view images that represents a scene as instance-structured 3D token groups rather than dense per-Gaussian primitives. Built on the pretrained VGGT geometry foundation model and evaluated on ScanNet, the method uses an image-anchor transformer to decode anchor tokens that spawn 3D Gaussians, then an anchor-grouping transformer to produce instance tokens that compete for anchor ownership via softmax assignment. Training relies only on 2D supervision: RGB rendering loss for reconstruction and Hungarian-matched instance-mask supervision for grouping, with no 3D annotations. The key novelty is a two-level semantic factorization: each instance stores a shared 512-dimensional group embedding, while 8-dimensional anchor residuals capture local variation, reducing feature storage from 8.4M scalars in Uni3R to 59.4K. On ScanNet with 2 context views, the model reaches 0.661 source-view mIoU and 0.657 target-view mIoU for feature lifting, while keeping reconstruction competitive at 25.28 PSNR, 0.771 SSIM, and 0.238 LPIPS. With 8 context views, it achieves the best class-agnostic instance segmentation, reaching 0.235 AP, 0.438 AP50, and 0.564 AP25, outperforming per-scene optimization baselines such as Gaussian Grouping and ObjectGS. The same token groups also enable direct instance-level editing and open-vocabulary retrieval, and the paper reports zero-shot transfer to MipNeRF360 and robustness on RealEstate10K using SAM2 pseudo-labels.

Original abstract

A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that decomposes a scene into instance-structured 3D token groups directly from unposed multi-view images -- compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group pairs an instance token capturing entity-level identity with anchor tokens that encode local geometry and appearance, which are decoded into a set of 3D Gaussians. This two-level factorization decouples object identity from local appearance, making object instances a native interface of the representation rather than a derived product. The token groups are learned through differentiable rendering with joint reconstruction and segmentation supervision, requiring no 3D annotations. Our feed-forward model surpasses per-scene optimization baselines in class-agnostic instance segmentation while remaining competitive in novel view synthesis. Beyond these metrics, the same token groups directly unlock instance-level scene editing -- removing, translating, or inserting objects by operating on their groups -- as well as efficient open-vocabulary 3D instance retrieval, where retrieval complexity scales with the number of instances rather than primitives.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis