NTH

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

AuthorsXinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu

September 6, 2026 3 min read
Watch on YouTube
The one-line take

This survey explores how AI can combine the structure of graphs with the visual intuition humans use to understand them.

Key results

8
VisionGraph task types

VisionGraph evaluates connectivity, shortest path, cycle detection, and other graph-theoretic tasks.

22
VGCure task count

VGCure covers node-, edge-, and graph-level visual reasoning tasks.

5
ChemDFM-X modalities

ChemDFM-X integrates 2D graphs, 3D conformations, images, mass spectra, and infrared spectra.

What the paper found

This survey defines “vision meets graphs” as treating rendered graph images as first-class inputs rather than relying only on adjacency matrices, GNN message passing, or LLM graph serialization. Its central diagnostic is the Rendering–Perception–Inference framework: rendering determines which structural cues survive, perception maps pixels to nodes, edges, labels, and endpoints, and inference performs prediction or multi-step reasoning. Across graph reasoning, graph learning, and scientific graphs, the recurring bottleneck is perception: models often miss nodes or misbind edge endpoints before reasoning begins. Benchmarks such as VisionGraph cover 8 graph-theoretic task types, while VGCure spans 22 node-, edge-, and graph-level tasks; methods including Description-Program-Reasoning, GITA, and MCDGraph improve reliability by combining visual and textual inputs, explicit structural descriptions, augmentation, masked graph infilling, and contrastive discrimination. For learning, GVN and DEL augment GNNs with visual features, while GraphAbstract shows that natural-image-pretrained vision encoders can outperform GNNs on graph-level global-structure tasks, despite limitations on node- and edge-level prediction. Scientific graphs are especially promising because conventions reduce visual ambiguity: ChemDFM-X combines 5 chemical modalities, GIT-Mol aligns molecular images, graphs, and text, and domain models such as ChemVLM outperform general-purpose VLMs on chemical reasoning. ImageMol pretrains on millions of molecular images, whereas protein systems remain less mature: Geneverse uses LLaVA with AlphaFold-derived structure images, but LiveProteinBench finds that structural images can fail to improve, or even degrade, sequence-only prediction. The survey’s roadmap emphasizes graph-native visual pretraining, controllable rendering, active visual reasoning, and tool-based structural verification.

Original abstract

Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision-language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis