RAD: Rule-Augmented Relational Anomaly Detection
AuthorsNoah Dahle, Anne Tumlin, Ngoc Tran, Xenofon Koutsoukos, Tyler Derr
RAD detects unusual behavior in multi-table databases by combining graph-based learning with human-readable behavioral rules.
Key results
Scale of the LANL cybersecurity relational database.
Scale of the Amazon review-churn relational database.
Scale of the H&M purchase-churn relational database.
AUROC achieved by RAD without edge reconstruction.
AUPRC achieved by RAD without edge reconstruction.
RAD’s average rank across the three benchmark datasets; lower is better.
What the paper found
RAD, or Rule-Augmented Relational Anomaly Detection, targets anomalies whose meaning depends on linked entities, typed relationships, temporal history, and multi-hop context—signals often lost when relational databases are flattened into one feature matrix. It converts tables into a heterogeneous graph, mines candidate behavioral rules from flattened summaries using random-forest paths, refines and compacts selected predicates with an LLM, then injects binary rule features into target nodes before two-layer GraphSAGE message passing. A heterogeneous graph autoencoder scores anomalies through attribute reconstruction and pairwise ranking supervision, while edge reconstruction is optional. The evaluation benchmark covers LANL cybersecurity authentication events with 1.65B records, Amazon review-churn with 15.0M records, and H&M purchase-churn with 16.7M records. Across these tasks, RAD achieves the strongest average AUPRC rank of 1.2, emphasizing retrieval of rare anomalies under severe class imbalance; on LANL, its no-edge-reconstruction variant reaches 0.996 AUROC and 0.659 AUPRC, substantially exceeding flattened detectors and relational baselines. Ablations show that direct rule injection and ranking supervision are central, whereas edge reconstruction can hurt when sampled graph topology is noisy. LLM refinement contributes mainly through rule filtering, tightening, deduplication, and compaction rather than independent anomaly labeling or scoring.
Original abstract
Anomaly detection is often applied to data stored in relational databases, yet most existing methods require flattening multiple tables into a single feature matrix. This flattening can obscure entity identity, schema structure, and multi-hop dependencies, limiting the detection of anomalies that depend on relational context rather than isolated feature values. Beyond preserving relational structure, relational anomaly detection raises an additional challenge: how to incorporate symbolic behavioral evidence into learned relational representations. To address these challenges, we study relational anomaly detection, where the goal is to identify anomalous entities or events in a multi-table database. We propose RAD, a rule-augmented relational anomaly detector that combines heterogeneous graph representation learning with refined symbolic rule signals. RAD derives candidate rules from random-forest paths over flattened summaries of the entities or events being scored, refines them into compact interpretable predicates, injects the resulting rule features into the graph model, and learns anomaly scores using reconstruction-based and pairwise-ranking supervision. To evaluate this setting, we introduce a relational anomaly detection benchmark spanning three settings: LANL cybersecurity event detection and two unexpected user-churn anomaly tasks derived from Amazon and H&M relational databases. Experiments show that RAD improves anomaly ranking over flattened tabular detectors and relational baselines under natural class imbalance, achieving the best average rank on AUROC and AUPRC across the benchmark. Ablations show that direct rule injection and ranking-based supervision are key contributors to performance, while edge reconstruction is not uniformly beneficial. Our code and data are available at: https://github.com/noahd15/RAD_RelationalAnomalyDetection.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.