Skip to main content
Concept
BM25 (also called Okapi BM25) is a ranking function that scores how relevant a document is to a keyword query. For each query term found in a document, it adds a score that rewards terms that are rare across the collection, gives diminishing credit for repeated occurrences, and adjusts for document length. It is the default relevance function in many full-text search engines.

Learning objectives

After reading this article you will be able to:
  • Explain what BM25 scores and why many search engines default to it
  • Describe how IDF, saturation (k1), and length normalization (b) shape a score
  • Work through a BM25 calculation and explain why the documents rank as they do
  • Compare BM25 with TF-IDF and recognize when it fits better than vector search

What is the BM25 formula?

For a query Q and a document D, a widely used form of the BM25 score is:
  • f(q, D) is how many times term q appears in document D.
  • |D| is the length of D, in terms, and avgdl is the average document length across the collection.
  • N is the number of documents in the collection, and n(q) is the number of documents that contain q.
  • k1 and b are tuning parameters for saturation and length normalization.
The IDF shown adds 1 inside the logarithm, a widely used variant that keeps every weight positive. The classic form omits it and can assign negative weights to terms that appear in more than half of the documents. Variants differ in small details, but all share the three ideas below.

How does BM25 score a document?

BM25 adds up a score for each query term the document contains. Each term’s score is its IDF weight times a term frequency that saturates and is adjusted for document length. Three ideas drive the result.

Inverse document frequency: rare terms count more

IDF measures how much a term tells you. A term that appears in almost every document, such as “issue” in a ticket system, has an IDF near zero. A term that appears in a handful of documents, such as “timeout,” has a high IDF. All else being equal, a document that matches the rare terms in a query tends to rank above one that matches only the common ones.

Term-frequency saturation: the k1 parameter

Mentioning a term more often makes a document more relevant, but only up to a point. The term-frequency part of the formula rises quickly and then flattens toward k1 + 1. With k1 = 1.2, for a document of average length: A low k1 saturates almost immediately, so the first occurrence is nearly all that counts. A high k1 keeps rewarding repetition for longer. Saturation limits how much a document can gain by repeating a keyword: no number of repetitions can push a term’s factor past k1 + 1.

Length normalization: the b parameter

A long document contains more words, so it matches more terms by chance. The b parameter controls how much BM25 corrects for that. With b = 0 length is ignored. With b = 1 term frequency is fully scaled by |D| / avgdl. With k1 = 1.2 and b = 0.75, one occurrence of a term contributes a factor of about 1.26 in a document half the average length, 1.00 at average length, and 0.71 at twice the average length.

Common parameter values

Commonly cited defaults are k1 between about 1.2 and 2.0 and b around 0.75. These are conventions from the research literature and from many search engines, not universal rules. Change them only when you can measure the effect on a set of labeled queries.

How do you calculate a BM25 score?

Compute each query term’s IDF, compute its term-frequency factor in each document, multiply the two, and sum over the query terms. For example, take three short documents and the query “export timeout,” with k1 = 1.2, b = 0.75, and no stemming or stop words:
“export” appears in all three documents, and “timeout” appears only in D1:
Each cell below is a term’s IDF multiplied by its term-frequency factor, and the score is their sum: Three effects are visible. D1 wins by a wide margin because it is the only document with the rare term “timeout.” D3 mentions “export” twice, but saturation and its above-average length mean it barely edges out D2. And because “export” is in every document, it contributes little to any score.

How is BM25 different from TF-IDF?

BM25 differs from classic TF-IDF in three main ways: term frequency saturates instead of growing without limit, document length normalization is explicit and tunable, and its IDF is a smoothed, probabilistic form. It also has two tuning parameters, k1 and b, where classic TF-IDF usually has none. BM25 comes from the probabilistic relevance framework rather than the vector space model, but it uses the same two signals as TF-IDF, term frequency and inverse document frequency, and is often used in its place. In practice BM25 is often found to produce better rankings than classic TF-IDF out of the box, which is one reason many engines use it as their default. BM25 fits best when the query contains exact terms that must match, such as names, IDs, error codes, and rare tokens; it needs no embedding model, and each term’s contribution to a score is visible. Vector search is the reverse: it matches meaning through embeddings, so it handles paraphrases and synonyms but can miss exact identifiers. Because the two fail in different places, hybrid search runs both and fuses the results.

How does HelixDB use BM25?

HelixDB text indexes are BM25-ranked. Each index covers a string or string-array property of a node or edge label and uses one of three analyzers: standard, standard_stem_en, or whitespace_lowercase. An index can optionally store term positions. Search results come back best match first, ordered by BM25 score and then by ID, so ties are broken deterministically. In prefiltered search, where BM25 ranks only the nodes or edges a traversal reaches, the BM25 statistics come from the full tenant partition rather than only from the candidates. A document’s score therefore does not depend on which records are in the candidate set. HelixDB does not fuse BM25 and vector results itself; the application combines them. See Text indexes to create an index and run a search.

Frequently asked questions

What does the 25 in BM25 mean?

BM stands for “best matching.” The number identifies one variant in a numbered series of weighting functions developed during research on the Okapi retrieval system. BM25 is the variant that became the standard.

Are BM25 scores comparable across queries?

No. A BM25 score has no fixed upper bound and depends on the query’s terms and the collection’s statistics. Use scores to order results for one query, not as an absolute relevance threshold across queries.

Does BM25 consider word order?

No. BM25 treats a document as a bag of words. Phrase matching and proximity scoring typically use term positions stored in the index and are applied as additional features by engines that support them.

What is BM25F?

BM25F is an extension for documents with several fields, such as a title and a body. It weights each field’s term frequency before applying saturation, so a match in the title can count for more than the same match in the body.

What is full-text search?

Inverted indexes, analyzers, and relevance ranking.

What is hybrid search?

Combining keyword and vector results with rank fusion.

What is filtered vector search?

Pre-filtering, post-filtering, and returning the right top k.

Prefiltered search guide

Rank only the records a traversal reaches in HelixDB.