NTH

GraphWrit3R: End-to-End 3D Scene Graph Writing

AuthorsLuka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

AffiliationsINSAIT, Sofia University “St. Kliment Ohridski” · Stanford University · Ulm University

October 5, 2026 3 min read
Watch on YouTube
The one-line take

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Key results

0.60
3DSSG object recall

Object R@5 on the in-domain 3DSSG benchmark.

0.83
3DSSG predicate recall

Predicate R@5 on the in-domain 3DSSG benchmark.

0.71
3DSSG triplet recall

Triplet R@100 on the in-domain 3DSSG benchmark.

0.92
ScanNet predicate recall

Zero-shot predicate R@5 after training exclusively on 3DSSG.

736.6M
Model parameter footprint

Approximate total inference-time parameters for GraphWrit3R.

What the paper found

GraphWrit3R reframes 3D scene-graph generation as direct language decoding: instead of detecting objects, constructing explicit intermediate graphs, and then classifying relations, it converts a point cloud, 3D Gaussian Splats, or both into voxel-aligned latent tokens and autoregressively writes a complete JSON graph containing object labels, 3D geometry, and subject–predicate–object relations. Sonata encodes point clouds, Chorus encodes Gaussian Splats, and a per-voxel alignment objective combines InfoNCE, cosine, and feature-level MSE losses before a Qwen2.5-0.5B decoder produces the graph. The system is fully local, needs no ground-truth object annotations at inference, and avoids proprietary API calls, unlike comparison setups involving OpenAI GPT-5.4. On 3DSSG, GraphWrit3R reaches 0.60 object R@5, 0.83 predicate R@5, and 0.71 triplet R@100 while jointly detecting objects and relations; on zero-shot ScanNet, it achieves 0.92 predicate R@5 after training only on 3DSSG. Its approximately 736.6M-parameter footprint supports second-scale inference, and the generated graph enables open-vocabulary queries and downstream spatial reasoning. The model parses valid JSON in 100% of cases when outputs fit its 8k-token context window, but scenes larger than roughly 40 objects and 1,600 relationships can be truncated. Ablations show asymmetric alignment toward Sonata and simple average token fusion are strongest, while degraded Gaussian-Splat reconstructions make point-cloud-only inference perform best.

Original abstract

3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.

Read the original paper

More in Graph Learning

Browse all 32 papers →