When Vision Meets Graphs: A Survey on Graph Reasoning and Learning
AuthorsXinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu
Resources
This survey explores how AI can combine the structure of graphs with the visual intuition humans use to understand them.
Key results
VisionGraph evaluates connectivity, shortest path, cycle detection, and other graph-theoretic tasks.
VGCure covers node-, edge-, and graph-level visual reasoning tasks.
ChemDFM-X integrates 2D graphs, 3D conformations, images, mass spectra, and infrared spectra.
What the paper found
This survey defines “vision meets graphs” as treating rendered graph images as first-class inputs rather than relying only on adjacency matrices, GNN message passing, or LLM graph serialization. Its central diagnostic is the Rendering–Perception–Inference framework: rendering determines which structural cues survive, perception maps pixels to nodes, edges, labels, and endpoints, and inference performs prediction or multi-step reasoning. Across graph reasoning, graph learning, and scientific graphs, the recurring bottleneck is perception: models often miss nodes or misbind edge endpoints before reasoning begins. Benchmarks such as VisionGraph cover 8 graph-theoretic task types, while VGCure spans 22 node-, edge-, and graph-level tasks; methods including Description-Program-Reasoning, GITA, and MCDGraph improve reliability by combining visual and textual inputs, explicit structural descriptions, augmentation, masked graph infilling, and contrastive discrimination. For learning, GVN and DEL augment GNNs with visual features, while GraphAbstract shows that natural-image-pretrained vision encoders can outperform GNNs on graph-level global-structure tasks, despite limitations on node- and edge-level prediction. Scientific graphs are especially promising because conventions reduce visual ambiguity: ChemDFM-X combines 5 chemical modalities, GIT-Mol aligns molecular images, graphs, and text, and domain models such as ChemVLM outperform general-purpose VLMs on chemical reasoning. ImageMol pretrains on millions of molecular images, whereas protein systems remain less mature: Geneverse uses LLaVA with AlphaFold-derived structure images, but LiveProteinBench finds that structural images can fail to improve, or even degrade, sequence-only prediction. The survey’s roadmap emphasizes graph-native visual pretraining, controllable rendering, active visual reasoning, and tool-based structural verification.
Original abstract
Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision-language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.