RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
AuthorsMaëlic Neau
AffiliationsSep 2026 still trained and evaluated on the 50 or 56 predicates of a single annotation style, with the relation head conditioned on object class labels and therefore tied to one detector and one label
Resources
RelateAnything aims to make scene-graph relation prediction as flexible as open-vocabulary detection by predicting arbitrary text-specified relations in real time.
Key results
RelateAnything model size
Images in the verified relation corpus
Relations in the training corpus
RelateAnything score versus 11.8 for OvSGTR
End-to-end pipeline latency per frame on NVIDIA A40
What the paper found
RelateAnything addresses the fixed-label bottleneck in visual scene-graph generation with a 53M-parameter model that takes pixels and regions from any detector or segmenter, discards object class labels, and scores relations against a predicate list supplied at inference. Its DINOv3 visual backbone and Relation Transformer produce pair embeddings compared directly with a frozen, distilled text-embedding bank, so changing the vocabulary is a matrix substitution rather than retraining. Training uses positive-unlabeled InfoNCE with synonym groups and discounted negatives, plus antonym repulsion to correct CLIP-derived text geometry, where relations such as above and below can otherwise reach cosine similarity 0.95. The accompanying RA-4M corpus contains 474k images, 4.3M machine-generated but geometrically verified relations, and 10,102 free-text predicates; annotations were produced with Gemma 4 and filtered by deterministic box checks. Across cross-dataset benchmarks, the model achieves 2.3× to 3.5× higher mean recall than comparable open-vocabulary systems, including ROBIN-3B, while using under 2% of its parameters, and reaches an OV-SGG-Bench composite score of 40.1 versus 11.8 for OvSGTR. With YOLO-World supplying regions, the gains persist, and the complete pipeline runs in 20 ms per frame on an NVIDIA A40. The paper also introduces OV-SGG-Bench, a six-axis protocol showing that conventional recall can reward object-category priors and annotation-style matching rather than visual relation understanding; in-domain improvements overstate transfer gains by approximately 5×. Qwen3-VL-8B is used as an independent graph-quality judge, while CLIP provides the tokenizer vocabulary for the distilled text encoder.
Original abstract
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class names: the taxonomy has left the model and become an input. Relation prediction has not. Scene-graph models are still trained and evaluated on the 50 or 56 predicates of one annotation style, their relation head conditioned on object labels and so tied to one detector. Three obstacles explain this, none primarily modelling: no relation corpus is both free-text and verified, a label-conditioned architecture cannot accept a vocabulary it was not trained on, and the standard metric rewards agreement with the training corpus, so a larger vocabulary scores as a regression. We present RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings. Object labels are never an input, so the region source can change without retraining, and the vocabulary is a bank of text embeddings, not a learned classifier. It runs at 20 ms/frame. Training over 19,103 predicates requires positive-unlabeled supervision and a text encoder that separates antonyms, which contrastive encoders embed at cosine 0.95. To supply the supervision we build RA-4M, 474k images and 4.3M relations over 10,102 free-text predicates, generated against numbered box markers and geometrically verified. To measure it we build OV-SGG-Bench, six axes scored across datasets that the priors standard recall rewards cannot satisfy. On three cross-dataset benchmarks and a fourth zero-shot, RelateAnything has 2.3-3.5x the mean recall of the strongest open-vocabulary method of comparable scale, margins that survive a real detector, and leads a 3B-VLM scene-graph model on both metrics at under 2% of its parameters. In-domain measurement overstates transfer gains ~5x. Model, corpus and benchmark are public.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.