NTH

TF-IDF and BM25 Are Exact KL Divergences

AuthorsIvan Silajev

September 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that two classic search-ranking methods, TF-IDF and BM25, can be understood precisely as measuring KL divergence between probability models.

Key results

4
Random variables in unified model

The construction models document index, token, key-token status, and common-token status.

1
BM25 IDF correction

The practical BM25 formulation includes the +1 correction in its IDF logarithm.

What the paper found

This paper gives TF-IDF and BM25 a unified probabilistic foundation by proving that their term-relevance scores arise from Kullback-Leibler divergences between explicitly constructed probability models. The derivation uses 4 random variables: document index, token, key-token status, and common-token status. For TF-IDF, the divergence assigns token probability d_i(t)/|d_i| and corpus rarity probability DF(t)/N, yielding the standard product of normalized term frequency and inverse document frequency exactly. For the practical BM25 variant used in production search software such as Milvus, the paper replaces these probabilities with a surrogate model incorporating term-frequency saturation through k1 and document-length normalization through b; the resulting divergence equals BM25 after the known scalar rescaling 1/(k1+1). Its IDF term includes the standard +1 correction, while the original BM25 formula without that correction is represented as a difference between divergences involving an adversarial probability model. In both cases, a query score is the query-token-weighted average of term divergences, measuring how strongly uncommon query keywords in a document depart from a naive commonness assumption. The result is theoretical rather than empirical: it supplies an exact information-theoretic interpretation and makes TF-IDF and BM25 directly comparable with other divergence-based retrieval methods.

Original abstract

TF-IDF and BM25 are two of the most widely used methods for scoring query-document relevance, yet neither has a standard probabilistic derivation that justifies it as a statistical method within a unified framework. We address this gap by showing that both scoring methods admit an exact interpretation as Kullback-Leibler divergences between two probability models. We treat the BM25 variant that includes the plus 1 correction in the IDF term, which is the one used in practice, and also discuss the original BM25 formulation without that correction. The resulting framework provides a common theoretical basis for TF-IDF and BM25, clarifies what they measure, and allows them to be compared theoretically with other information retrieval methods rather than only experimentally.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →