Graph Evidence Is Not Enough: Diagnosing Native Decoder Use in Graph-Augmented LLMs
AuthorsXiaoyu Guo, Pengcheng Chen, Jiong Yu, Yi Lu, Yaohua Wang, Ziyang Li
Resources
The paper shows that feeding graphs to an LLM is not enough—the graph must be presented in a form the decoder can actually use.
Key results
Native-generation performance across the four Core-HopQA domains.
Native HopQA score on the DBLP citation-style graph.
Native HopQA score on the Biomedical relation graph.
Native HopQA score on the GoodReads recommendation graph.
Native HopQA score on the PubMed citation graph.
Percentage-point improvement reported across the four domains.
What the paper found
This paper argues that exposing graph evidence to an LLM does not mean the decoder can use it. It introduces five-choice HopQA, a bounded diagnostic requiring native generation of the exact shortest-hop distance between two nodes, and separates graph-signal existence from evidence inclusion, structural readability, adjacency preservation, and decoder usability. On DBLP, Biomedical, GoodReads, and PubMed, the native-generation baseline G-Retriever achieves 0.0% strict exact match, while LLaGA remains near zero, despite graph-only classifiers and retrieval-execution controls recovering hop information. The proposed Sampling-First Structured Graph Encoding, or S2GE, addresses the interface with query-aware sampling, source-target and proximity roles, ordered serialization, and an adjacency-alignment loss; it is evaluated with LLaMA-3-8B-Instruct using DeepSpeed and NVIDIA GPUs. S2GE reaches 36.5% strict exact match on DBLP, 57.8% on Biomedical, 76.6% on GoodReads, and 52.0% on PubMed, outperforming the strongest native-generation baseline by 53.5 percentage points on average. A readable-versus-shuffled-versus-no-graph intervention triangle reveals three regimes: harmful shuffling, shuffle robustness, and no-graph saturation. The central finding is that decoder scale cannot recover topology distinctions discarded by sampling or projection; graph-augmented systems therefore need interfaces that preserve query-relevant evidence, endpoint roles, and local adjacency before generation begins.
Original abstract
Graph-augmented large language models often assume that graph evidence produced by external computation and placed in the input can be used by the native decoder. We test this assumption with HopQA, a deliberately bounded diagnostic that asks for the shortest-hop distance between two query nodes. Because the answer is a small integer and the target is purely topological, failure cannot be dismissed as open-ended generation or ambiguous evaluation. Yet existing graph-augmented baselines still fail on this setting, showing that providing graph evidence is not the same as making it usable. We introduce an intervention triangle with three matched conditions: readable graph evidence, shuffled graph evidence, and no-graph input. This separates evidence inclusion, structural readability, and decoder-usable topology. Guided by this diagnosis, we present S$^2$GE as an instance showing that diagnosis-driven interface design can improve native decoder usability. S$^2$GE uses query-aware sampling, endpoint and proximity-based ordering, and structure-preserving alignment. Across DBLP, Biomedical, GoodReads, and PubMed, S$^2$GE achieves strict exact-match scores of $36.5\%$, $57.8\%$, $76.6\%$, and $52.0\%$, improving over the strongest native-generation baseline by $53.5$ points on average. The interventions further reveal harmful-shuffle, shuffle-robust, and no-graph-saturated regimes.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.