OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
AuthorsJinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo, Minki Kang, Patara Trirat, Heejun Lee, Sung Ju Hwang
Resources
OmniRetrieval aims to be a universal translator for search, sending each query to the right mix of text, tables, and graph databases instead of forcing everything into one format.
Key results
benchmark datasets spanning all four backends
distinct knowledge bases in the benchmark
questions sampled per dataset for evaluation
OmniRetrieval macro-averaged source selection accuracy
OmniRetrieval macro-averaged retrieval accuracy
OmniRetrieval macro-averaged LLM-as-a-Judge accuracy
What the paper found
OmniRetrieval, from KAIST and DeepAuto.ai, argues that retrieval should not collapse text, SQL, SPARQL, and Cypher into one shared embedding space; instead, it uses a three-stage LLM controller that first selects relevant knowledge sources from their native descriptors, then generates source-specific executable queries, and finally consolidates heterogeneous outputs into one evidence set. The system is evaluated on 13 datasets spanning 309 knowledge bases across unstructured corpora, relational databases, RDF knowledge graphs, and labeled property graphs, with 300 questions sampled per dataset. Using GPT-5.4, Gemini-3.1 (Pro), Sonnet-4.6, Qwen-3.5 (27B), and Gemma-4 (31B), OmniRetrieval outperforms KB Routing on the macro-averaged main metrics: source selection rises from 61.65 to 65.71, retrieval accuracy from 39.98 to 44.34, and LLM-as-a-Judge from 57.99 to 65.88. The strongest gains come from deferring commitment until evidence selection, because the gold source is often included among candidates even when the top-1 route is wrong; in multi-candidate questions, evidence-selection accuracy reaches 67.5% at k=3 but falls to 62.8% at k=10, showing the trade-off between broader exploration and selector quality. The paper also reports that a constrained unified-representation baseline remains well below OmniRetrieval, supporting the claim that structural operators such as joins, graph traversals, and property paths must be preserved rather than flattened.
Original abstract
Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification would collapse these sources into a shared space, but this erases the structural affordances (such as schemas, ontologies, compositional operators) that give each source its expressive power. Effective retrieval over diverse knowledge, therefore, requires not homogenization but an overarching layer that meets each source on its own terms. To achieve this, we present OmniRetrieval, a framework that takes any natural-language query, identifies appropriate knowledge sources, and dispatches source-native queries to their native execution engines. Across an extensive benchmark spanning 13 datasets and 309 distinct knowledge bases over text, relational, and graph-structured sources, OmniRetrieval exceeds single-source baselines, demonstrating that it can serve as a general-purpose interface to the heterogeneous sources while preserving the structural distinctions that make each source valuable.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.