NTH

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

AuthorsJinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo, Minki Kang, Patara Trirat, Heejun Lee, Sung Ju Hwang

June 10, 2026 2 min read
Watch on YouTube
The one-line take

OmniRetrieval aims to be a universal translator for search, sending each query to the right mix of text, tables, and graph databases instead of forcing everything into one format.

Key results

13
datasets

benchmark datasets spanning all four backends

309
knowledge_bases

distinct knowledge bases in the benchmark

300
questions_per_dataset

questions sampled per dataset for evaluation

65.71
source_selection_accuracy

OmniRetrieval macro-averaged source selection accuracy

44.34
retrieval_accuracy

OmniRetrieval macro-averaged retrieval accuracy

65.88
judge_accuracy

OmniRetrieval macro-averaged LLM-as-a-Judge accuracy

What the paper found

OmniRetrieval, from KAIST and DeepAuto.ai, argues that retrieval should not collapse text, SQL, SPARQL, and Cypher into one shared embedding space; instead, it uses a three-stage LLM controller that first selects relevant knowledge sources from their native descriptors, then generates source-specific executable queries, and finally consolidates heterogeneous outputs into one evidence set. The system is evaluated on 13 datasets spanning 309 knowledge bases across unstructured corpora, relational databases, RDF knowledge graphs, and labeled property graphs, with 300 questions sampled per dataset. Using GPT-5.4, Gemini-3.1 (Pro), Sonnet-4.6, Qwen-3.5 (27B), and Gemma-4 (31B), OmniRetrieval outperforms KB Routing on the macro-averaged main metrics: source selection rises from 61.65 to 65.71, retrieval accuracy from 39.98 to 44.34, and LLM-as-a-Judge from 57.99 to 65.88. The strongest gains come from deferring commitment until evidence selection, because the gold source is often included among candidates even when the top-1 route is wrong; in multi-candidate questions, evidence-selection accuracy reaches 67.5% at k=3 but falls to 62.8% at k=10, showing the trade-off between broader exploration and selector quality. The paper also reports that a constrained unified-representation baseline remains well below OmniRetrieval, supporting the claim that structural operators such as joins, graph traversals, and property paths must be preserved rather than flattened.

Original abstract

Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification would collapse these sources into a shared space, but this erases the structural affordances (such as schemas, ontologies, compositional operators) that give each source its expressive power. Effective retrieval over diverse knowledge, therefore, requires not homogenization but an overarching layer that meets each source on its own terms. To achieve this, we present OmniRetrieval, a framework that takes any natural-language query, identifies appropriate knowledge sources, and dispatches source-native queries to their native execution engines. Across an extensive benchmark spanning 13 datasets and 309 distinct knowledge bases over text, relational, and graph-structured sources, OmniRetrieval exceeds single-source baselines, demonstrating that it can serve as a general-purpose interface to the heterogeneous sources while preserving the structural distinctions that make each source valuable.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →