Concept
BM25 (also called Okapi BM25) is a ranking function that scores how relevant a
document is to a keyword query. For each query term found in a document, it adds a
score that rewards terms that are rare across the collection, gives diminishing credit
for repeated occurrences, and adjusts for document length. It is the default relevance
function in many full-text search engines.
Learning objectives
After reading this article you will be able to:
- Explain what BM25 scores and why many search engines default to it
- Describe how IDF, saturation (
k1), and length normalization (b) shape a score - Work through a BM25 calculation and explain why the documents rank as they do
- Compare BM25 with TF-IDF and recognize when it fits better than vector search
What is the BM25 formula?
For a queryQ and a document D, a widely used form of the BM25 score is:
f(q, D)is how many times termqappears in documentD.|D|is the length ofD, in terms, andavgdlis the average document length across the collection.Nis the number of documents in the collection, andn(q)is the number of documents that containq.k1andbare tuning parameters for saturation and length normalization.
How does BM25 score a document?
BM25 adds up a score for each query term the document contains. Each term’s score is its IDF weight times a term frequency that saturates and is adjusted for document length. Three ideas drive the result.Inverse document frequency: rare terms count more
IDF measures how much a term tells you. A term that appears in almost every document, such as “issue” in a ticket system, has an IDF near zero. A term that appears in a handful of documents, such as “timeout,” has a high IDF. All else being equal, a document that matches the rare terms in a query tends to rank above one that matches only the common ones.Term-frequency saturation: the k1 parameter
Mentioning a term more often makes a document more relevant, but only up to a point. The term-frequency part of the formula rises quickly and then flattens towardk1 + 1. With k1 = 1.2, for a document of average length:
A low
k1 saturates almost immediately, so the first occurrence is nearly all that
counts. A high k1 keeps rewarding repetition for longer. Saturation limits how much
a document can gain by repeating a keyword: no number of repetitions can push a term’s
factor past k1 + 1.
Length normalization: the b parameter
A long document contains more words, so it matches more terms by chance. Theb
parameter controls how much BM25 corrects for that. With b = 0 length is ignored.
With b = 1 term frequency is fully scaled by |D| / avgdl. With k1 = 1.2 and
b = 0.75, one occurrence of a term contributes a factor of about 1.26 in a document
half the average length, 1.00 at average length, and 0.71 at twice the average length.
Common parameter values
Commonly cited defaults arek1 between about 1.2 and 2.0 and b around 0.75. These
are conventions from the research literature and from many search engines, not
universal rules. Change them only when you can measure the effect on a set of labeled
queries.
How do you calculate a BM25 score?
Compute each query term’s IDF, compute its term-frequency factor in each document, multiply the two, and sum over the query terms. For example, take three short documents and the query “export timeout,” withk1 = 1.2, b = 0.75, and no
stemming or stop words:
Three effects are visible. D1 wins by a wide margin because it is the only document
with the rare term “timeout.” D3 mentions “export” twice, but saturation and its
above-average length mean it barely edges out D2. And because “export” is in every
document, it contributes little to any score.
How is BM25 different from TF-IDF?
BM25 differs from classic TF-IDF in three main ways: term frequency saturates instead of growing without limit, document length normalization is explicit and tunable, and its IDF is a smoothed, probabilistic form. It also has two tuning parameters,k1 and
b, where classic TF-IDF usually has none.
BM25 comes from the probabilistic relevance framework rather than the vector space
model, but it uses the same two signals as TF-IDF, term frequency and inverse document
frequency, and is often used in its place.
In practice BM25 is often found to produce better rankings than classic TF-IDF out of
the box, which is one reason many engines use it as their default.
When is BM25 a better fit than vector search?
BM25 fits best when the query contains exact terms that must match, such as names, IDs, error codes, and rare tokens; it needs no embedding model, and each term’s contribution to a score is visible. Vector search is the reverse: it matches meaning through embeddings, so it handles paraphrases and synonyms but can miss exact identifiers. Because the two fail in different places, hybrid search runs both and fuses the results.How does HelixDB use BM25?
HelixDB text indexes are BM25-ranked. Each index covers a string or string-array property of a node or edge label and uses one of three analyzers:standard,
standard_stem_en, or whitespace_lowercase. An index can optionally store term
positions. Search results come back best match first, ordered by BM25 score and then by
ID, so ties are broken deterministically.
In prefiltered search, where BM25 ranks
only the nodes or edges a traversal reaches, the BM25 statistics come from the full
tenant partition rather than only from the candidates. A document’s score therefore
does not depend on which records are in the candidate set. HelixDB does not fuse BM25
and vector results itself; the application combines them. See
Text indexes to create an index and
run a search.
Frequently asked questions
What does the 25 in BM25 mean?
BM stands for “best matching.” The number identifies one variant in a numbered series of weighting functions developed during research on the Okapi retrieval system. BM25 is the variant that became the standard.Are BM25 scores comparable across queries?
No. A BM25 score has no fixed upper bound and depends on the query’s terms and the collection’s statistics. Use scores to order results for one query, not as an absolute relevance threshold across queries.Does BM25 consider word order?
No. BM25 treats a document as a bag of words. Phrase matching and proximity scoring typically use term positions stored in the index and are applied as additional features by engines that support them.What is BM25F?
BM25F is an extension for documents with several fields, such as a title and a body. It weights each field’s term frequency before applying saturation, so a match in the title can count for more than the same match in the body.Related topics
What is full-text search?
Inverted indexes, analyzers, and relevance ranking.
What is hybrid search?
Combining keyword and vector results with rank fusion.
What is filtered vector search?
Pre-filtering, post-filtering, and returning the right top k.
Prefiltered search guide
Rank only the records a traversal reaches in HelixDB.