NTH

GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation

AuthorsMaya Arseven, Anette Frank, Beni Egressy, Johann Higl, Moritz Plenz

August 11, 2026 2 min read
Watch on YouTube
The one-line take

GLM-RAG shows that graph language models can make knowledge-graph retrieval more transferable for multi-hop question answering than conventional alternatives.

Key results

273,830
Training question-document pairs

Total training queries used across HotpotQA, MuSiQue, and 2WikiMultihopQA.

46.4
MuSiQue Recall@2

In-domain document retrieval performance.

60.0
MultiHopRAG GLM-RAG Recall@2

Zero-shot multi-hop retrieval performance.

76.9
G-Bench Medical accuracy

Zero-shot accuracy on the medical multi-hop benchmark.

76.6
G-Bench Computer Science accuracy

Zero-shot accuracy on the computer-science multi-hop benchmark.

692.3
GLM-RAG retrieval latency

Average retrieval latency in milliseconds, compared with 18.6 milliseconds for GFM-RAG*.

What the paper found

GLM-RAG replaces the GNN retriever in graph-based retrieval-augmented generation with a Graph Language Model initialized from T5-large, allowing node labels, relation labels, queries, and graph structure to interact at the token level. Its graph encoder combines language-model self-attention with graph-aware relative positional encoding, then ranks entities in a 2-hop subgraph capped at 600 triplets before mapping them back to source documents. Across HotpotQA, MuSiQue, and 2WikiMultihopQA, training uses 273,830 question-document pairs; GLM-RAG remains competitive in-domain, reaching 46.4 Recall@2 on MuSiQue, while vanilla RAG is generally sufficient for single-hop questions. Its main advantage appears in zero-shot multi-hop transfer: on MultiHopRAG, GLM-RAG achieves 60.0 Recall@2 versus 39.0 for GFM-RAG+, and it reaches 76.9 accuracy on G-Bench Medical and 76.6 on G-Bench Computer Science, surpassing GFM-RAG, G-Reasoner, and the Microsoft GraphRAG comparison. Ablations show scaling benefits from T5-small through T5-large, with the larger encoder reaching 336M parameters, but the semantic gains carry a substantial systems cost: retrieval latency is 692.3 milliseconds versus 18.6 milliseconds for GFM-RAG*. Answers are generated with OpenAI’s gpt-4o-mini, and the results suggest GLM-RAG is best suited to compositional, out-of-domain reasoning where graph traversal and deep text semantics must be combined.

Original abstract

Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model graph topology in multi-hop reasoning tasks. In parallel, graph language models (GLMs) have emerged as a promising paradigm that integrates graph reasoning and the semantic capabilities of language models. In this work, we introduce a GLM-based retriever and investigate the comparative strengths of GLM-based, GNN-based, and traditional vector-search-based retrievers in single- and multi-hop RAG settings, and with a particular focus on transferability to unseen domains. Our findings suggest that finetuned GLM retrievers generalize better out of domain, achieving SOTA on two multi-hop benchmarks. On in-domain multi-hop QA datasets they remain comparable to prior work, with promising scaling as parameters and subgraph coverage increase. GNN-based retrievers achieve higher graph coverage with an efficient training setup, whereas the vector-search baseline excels at single-hop datasets.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis