NTH

Hakken: Predicting future discoveries to fill the gaps in today's knowledge

AuthorsTarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi, Samy Badreddine, Frederick Gifford, Daniel Evans-Yamamoto, Sucheendra K. Palaniappan, Miquel Ferrer, Kae Nagano, Iris Rossell, Tom Joy, Hatem ElShazly, Chrysa Iliopoulou, Christoph Wehner, Thiviyan Thanapalasingam, Susana Nunes, Pedro G. Cotovio, Peter Wurman, Peter Stone, Hiroaki Kitano, Michael Spranger

September 14, 2026 2 min read
Watch on YouTube
The one-line take

Hakken uses evolving scientific knowledge graphs and language models to propose explainable future discoveries, including two newly validated biomedical interactions.

Key results

254806
Cleaned biomedical entities

Number of entities in the cleaned biomedical knowledge graph.

7127960
Cleaned biomedical triples

Number of temporal knowledge-graph triples used for evaluation.

60.72%
THiGERLLM macro recall

Macro recall under the 2020 temporal cutoff.

51.16%
THiGERLLM macro F1

Macro F1 under the 2020 temporal cutoff.

1543297
Aging hypotheses

Above-confidence-threshold hypotheses generated for aging-related entities.

What the paper found

Hakken is a domain-agnostic system for predicting scientific relationships that have not yet appeared in the literature, rather than merely retrieving or recombining known facts. Its core model, THiGERLLM, combines temporal knowledge-graph encoding with publication semantics from Mistral-7B-Instruct-v0.3, using GraphSAGE neighborhood aggregation, hierarchical temporal Transformers, label-history features, positive-unlabeled learning, and calibrated multi-label prediction. PHELInE then explains each hypothesis through influential multi-hop graph paths, using a GraphSAGE surrogate to estimate sufficiency and necessity without retraining the predictor. In biomedicine, the cleaned graph contains 254806 entities and 7127960 triples spanning 23 relation types. With a 2020 temporal cutoff, THiGERLLM achieved 60.72% macro recall and 51.16% macro F1, improving coverage of rare relation types relative to the graph-only THiGER model, while retaining predictive signal across ten-year horizons. For aging research, it generated 1543297 above-threshold hypotheses; three were selected for laboratory testing, and two were supported: TP53 affecting BAMBI expression and RAF1 regulating TNF expression. The approach is positioned as complementary to discovery systems such as AlphaFold and GNoME and to broad LLM systems such as Google Gemini 2.0, offering structured, interrogable predictions with lower computational demands than large multi-agent workflows; training used two machines equipped with NVIDIA H100 GPUs.

Original abstract

We present Hakken, a domain-agnostic prediction and explanation system performing knowledge prediction, i.e., growing scientific knowledge by establishing novel relationships, ones that are not limited to the deductive hull of previous knowledge. Hakken uses a transformer-based prediction model built on temporal sequences of knowledge graphs extracted from vast bodies of research publications, fused with an LLM's semantic knowledge, to predict the presence and define the type of as-yet undocumented relationships between scientific concepts. It then calls a model-agnostic explanation framework to provide accompanying information for each prediction that allows scientists to evaluate the suggested new relationship. While general purpose, we demonstrate Hakken's practical capabilities by applying it to the biomedical domain. There, Hakken's prediction model establishes a new benchmark for time-aware multi-label relation prediction, and we show that the model's output stays coherent and informative over extended time spans in historic data. In addition, we scored 1.5 million above-confidence-threshold hypotheses related to aging, qualitatively validated batches of these predictions with biologists and progressed three of them for empirical validation in wet-lab. Two predictions with potentially significant impact in the context of drug discovery and repurposing were confirmed, introducing previously undocumented interactions between TP53 and BAMBI, and between RAF1 and TNF, to biomedical science.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis