NTH
Research collection

Natural Language Processing research

Explore methods for understanding and generating human language. Read research on linguistic tasks, multilingual systems, and evaluation.

26 papers · Latest edition October 2, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Natural Language Processing papers

Newest editions first.

06Nlp

TF-IDF and BM25 Are Exact KL Divergences

Ivan Silajev

This paper shows that two classic search-ranking methods, TF-IDF and BM25, can be understood precisely as measuring KL divergence between probability models.

Read analysis
16Nlp

NSF-SciFy: Mining the NSF Awards Database for Scientific Claims

Delip Rao, Weiqiu You, Eric Wong, Chris Callison-Burch

This paper turns millions of NSF award abstracts into a new dataset for extracting scientific claims and research plans, aiming to power better claim verification and science analysis.

Read analysis
17Nlp

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

Jinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo, Minki Kang, Patara Trirat, Heejun Lee, Sung Ju Hwang

OmniRetrieval aims to be a universal translator for search, sending each query to the right mix of text, tables, and graph databases instead of forcing everything into one format.

Read analysis
18Nlp

Xetrieval: Mechanistically Explaining Dense Retrieval

Zhixin Cai, Jun Bai, Yang Liu, Jiaqi Li, Yichi Zhang, Taichuan Li, Zhuofan Chen, Zixia Jia, Zilong Zheng, Wenge Rong

Xetrieval explains why dense retrievers rank documents highly by turning opaque embeddings into sparse human-readable features and using them to trace retrieval decisions.

Read analysis
20Nlp

KletterMix: Climbing Toward High-Quality German Pretraining Data

Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia, Abbas Goher Khan, Mehdi Ali, Michael Fromm, Kristian Kersting

This paper builds a large, carefully translated German pretraining corpus and shows that better curated data can noticeably improve German language models.

Read analysis
23Nlp

ICICLE: Expanding Retrieval with In-Context Documents

Yu-Chen Den, Yung-Yu Shih, Zhi Rui Tam, Kuan-Yu Chen, Pu-Jen Cheng, Yun-Nung Chen, Eugene Yang

ICICLE lets generative retrievers add new documents on the fly by using in-context document IDs instead of retraining the model each time.

Read analysis
26Nlp

Tokenisation via Convex Relaxations

Jan Tempus, Philip Whittington, Craig W. Schmidt, Dennis Komm, Tiago Pimentel

This paper turns tokenization from a greedy heuristic into a convex optimization problem, producing a tokenizer that can be near-optimal and often slightly better for language models.

Read analysis