NTH

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation

AuthorsBao Long Nguyen Huu, Atsushi Hashimoto

July 14, 2026 2 min read
Watch on YouTube
The one-line take

AGE improves GraphRAG by learning which graph nodes to mask during self-supervised training, making graph embeddings better aligned with frozen LLMs for graph question answering.

Key results

0.5595
ExplaGraphs baseline accuracy

G-Retriever on ExplaGraphs with Llama3.2 1B

0.8267
ExplaGraphs AGE accuracy

AGE G-Retriever on ExplaGraphs with Llama3.2 1B

26.72%
ExplaGraphs improvement

Full AGE stack over G-Retriever baseline

15.46%
Ablation gain random mask

JEPA with random mask over baseline

80.3
WebQSP Hit@1

AGE G-Retriever with Llama3.1 8B LoRA

0.3
Best sampling rate

Chosen node sampling rate for experiments

What the paper found

AGE, from OMRON Corporation and OMRON SINIC X Corporation, targets a core failure mode in GraphRAG: frozen LLMs cannot reliably consume graph embeddings because graph latent features drift away from text-encoder spaces. The method replaces random node masking with an RL-trained node sampler that identifies key nodes and masks auxiliary nodes, then trains a Transformer-based concept encoder-decoder with a JEPA objective to reconstruct auxiliary-node representations from key nodes. This design is paired with a graph-structure aggregator and a frozen Llama backbone, including Llama3.2 1B, 3B, Llama3.1 8B, and Llama2 7B/13B. On ExplaGraphs, AGE lifts G-Retriever from 0.5595 to 0.8267 with Llama3.2 1B, and ablations show the full JEPA plus node sampler stack beats GA and random masking, reaching a 26.72% improvement over baseline versus 9.37% and 15.46% for weaker variants. On WebQSP, AGE G-Retriever reaches 80.3 Hit@1 with Llama3.1 8B LoRA, outperforming G-Retriever’s 70.2 and narrowing the gap to parametric retrievers such as DualR at 82.8. The paper also reports that a sampling rate of 0.3 is the best overall trade-off, and that AGE can preserve or improve accuracy while keeping training at 2.0 to 7.3 minutes per epoch and inference as high as 148.5 tokens per second.

Original abstract

GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referring to graph-structured data as external knowledge. While this technique ideally captures intricate relationships, it often struggles with graph representations for LLMs, particularly for frozen LLMs, due to the misalignment between graph-based and text-based latent features. We tackle this issue by introducing the {\it Adaptive-masking for Graph Embedding (AGE)}. AGE employs a Transformer in a mask-based self-supervised learning (SSL) approach. We designed the architecture similar to text embedding encoders, addressing the latent feature misalignment. In contrast to natural language texts, graphs are concise representations, and there exist {\it key nodes} that hold dominant contextual information, which are challenging to predict from their surroundings. Masking such key nodes leads to inefficiency in the SSL process. Therefore, AGE focuses on predicting nodes apart from key nodes, utilizing a learnable node sampler. Our experimental results indicate that AGE significantly improves approaches using non-parametric search component in GraphQA tasks, achieving superior accuracy across four benchmark datasets with distinct characteristics.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis