NTH

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

AuthorsLuigi Sigillo, Matteo Silvestri, Francesco Tabaro, Rajat Bhatnagar, Syed Irtaza Mubashar, Matt Jeffryes, Daljit Nijjer, Vittorio Perera, Ola Spjuth, Julio Saez-Rodriguez, Melissa Harrison, Fabio Petroni

August 5, 2026 3 min read
Watch on YouTube
The one-line take

EMBL AI Librarian helps biology-focused AI agents turn natural-language questions into literature-backed answers and evidence.

Key results

7
Librarian subqueries

Maximum complementary queries generated per user question.

73.8
ScholarQA-Bench Citation F1

GLM-5 synthesis agent using Librarian on the Bio split.

0.80
ProClaim average agreement

Claim verification with Librarian and Claude Sonnet 4.6.

78.9
LitQA2 accuracy

GPT-5.4 with Librarian on the 91-question open-form subset.

54.6
LAB-Bench GPT-5.4 macro accuracy

Macro accuracy after adding Librarian across four biology task categories.

What the paper found

Researchers at EMBL Rome and EMBL-EBI introduce EMBL AI Librarian, a model-agnostic knowledge layer that turns natural-language questions into compact, citable evidence from Europe PMC, rather than returning whole papers. A single GLM-5 controller generates complementary keyword and fielded subqueries, validates them, searches the live Europe PMC index, ranks paragraphs with BM25, and uses an LLM to filter, rerank, and extract supporting sentences. The evaluation used 7 subqueries, up to 50 records per query, and 16 full-text paragraphs per paper. On ScholarQA-Bench, a GLM-5 synthesis agent reached Citation F1 of 73.8 with Librarian versus 67.0 using BM25 over the OpenScholar Data Store. In ProClaim-eval, replacing the original PubMed and Semantic Scholar retriever with Librarian raised agreement from 0.75 to 0.80; the claim-verification agents used Anthropic’s Claude Sonnet 4.6. On the 91-question open-form LitQA2 subset, an OpenAI GPT-5.4 agent achieved 78.9 accuracy with Librarian, compared with 70.3 using web search. For LAB-Bench, GPT-5.4 macro accuracy increased from 50.5 to 54.6, with SeqQA accuracy rising from 52.5 to 63.8. The system runs GLM-5 on an NVIDIA DGX B200 and avoids maintaining a costly dense embedding index, while remaining inspectable through its Europe PMC queries and source metadata. Its limitations include single-round retrieval and no direct support for figures, tables, or supplementary material.

Original abstract

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than $16$ points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about $8$ points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis