GraphWrit3R: End-to-End 3D Scene Graph Writing
AuthorsLuka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
AffiliationsINSAIT, Sofia University “St. Kliment Ohridski” · Stanford University · Ulm University
Resources
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Key results
Object R@5 on the in-domain 3DSSG benchmark.
Predicate R@5 on the in-domain 3DSSG benchmark.
Triplet R@100 on the in-domain 3DSSG benchmark.
Zero-shot predicate R@5 after training exclusively on 3DSSG.
Approximate total inference-time parameters for GraphWrit3R.
What the paper found
GraphWrit3R reframes 3D scene-graph generation as direct language decoding: instead of detecting objects, constructing explicit intermediate graphs, and then classifying relations, it converts a point cloud, 3D Gaussian Splats, or both into voxel-aligned latent tokens and autoregressively writes a complete JSON graph containing object labels, 3D geometry, and subject–predicate–object relations. Sonata encodes point clouds, Chorus encodes Gaussian Splats, and a per-voxel alignment objective combines InfoNCE, cosine, and feature-level MSE losses before a Qwen2.5-0.5B decoder produces the graph. The system is fully local, needs no ground-truth object annotations at inference, and avoids proprietary API calls, unlike comparison setups involving OpenAI GPT-5.4. On 3DSSG, GraphWrit3R reaches 0.60 object R@5, 0.83 predicate R@5, and 0.71 triplet R@100 while jointly detecting objects and relations; on zero-shot ScanNet, it achieves 0.92 predicate R@5 after training only on 3DSSG. Its approximately 736.6M-parameter footprint supports second-scale inference, and the generated graph enables open-vocabulary queries and downstream spatial reasoning. The model parses valid JSON in 100% of cases when outputs fit its 8k-token context window, but scenes larger than roughly 40 objects and 1,600 relationships can be truncated. Ablations show asymmetric alignment toward Sonata and simple average token fusion are strongest, while degraded Gaussian-Splat reconstructions make point-cloud-only inference perform best.
Original abstract
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.
Efficient Swing Computation for Retrieval in Large-Scale Recommender Systems
Runhao Jiang, Renchi Yang
This paper makes massive-scale item recommendations much faster by approximating graph-based similarity scores without sacrificing retrieval quality.