Skip to main content
Concept
Hybrid search is a retrieval approach that runs a keyword search, typically BM25, and a vector search for the same query, then merges the two ranked lists into one. Keyword search catches exact names, IDs, and error codes; vector search catches paraphrases and related concepts. Fusing them, often with reciprocal rank fusion, gives results that hold up for both kinds of query.

Learning objectives

After reading this article you will be able to:
  • Explain why BM25 and vector search fail in complementary ways
  • Describe the hybrid pipeline: retrieve from both indexes, fuse, optionally rerank
  • Calculate reciprocal rank fusion scores and compare them with weighted blending
  • Recognize when hybrid search is worth its extra cost

How do keyword search and vector search differ?

Keyword search matches the words in the query, while vector search matches meaning, so each one succeeds where the other fails. That complementary failure is the whole case for combining them. A support-ticket search shows the gap. A user typing “E1042” needs the tickets that contain that exact code, which an embedding may blur into generic “error” text. A user typing “I can’t log in after resetting my password” needs tickets about authentication failures, even when they share few words with the query. For each method on its own, see What is BM25? and What is vector search?

How does hybrid search work?

Hybrid search sends one query to both a BM25 index and a vector index, fuses the two ranked candidate lists into a single ranking, and optionally reranks the top of that list before returning results.
1

Retrieve from both indexes

Run BM25 and vector search for the same query, usually in parallel. Embed the query with the same model that embedded the documents. Retrieve more candidates from each search than you plan to return, so good results from one list are not cut off before fusion.
2

Fuse the lists

Combine the two rankings into one with a fusion method such as reciprocal rank fusion.
3

Rerank, optionally

Score the fused top results with a slower, more accurate model, then return the final top results.
Apply the same filters, such as tenant or access control, to both searches, ideally before ranking. Filtering after ranking can leave a list with fewer than k results, or none, which skews fusion toward the other list. Skipping the filter on either list can return records the user should not see. See What is filtered vector search?

How do you combine BM25 and vector search results?

Most systems use rank-based fusion, usually reciprocal rank fusion (RRF), or blend normalized scores with a weight. Either can be followed by a reranker that rescores the fused top candidates.

Reciprocal rank fusion

Reciprocal rank fusion ignores raw scores and uses only each document’s rank in each list:
The constant k dampens the advantage of the very top ranks. The value 60 comes from the paper that introduced RRF and is a common convention, not a requirement. Suppose BM25 returns A, B, C and vector search returns D, C, A: In this example, documents that appear in both lists rise to the top. In general, RRF rewards agreement between lists, but a high rank in one list can still outscore low ranks in both. RRF needs no score normalization and no training data, which makes it a robust default. Its limitation is that it discards score magnitude: a document that is far better than the rest in one list gets no extra credit. A weighted variant multiplies each list’s terms by a weight to favor one method.

Weighted score blending

The other common approach normalizes each list’s scores and blends them:
A document that appears in only one list has no score in the other. Its missing normalized score is typically treated as 0 or as that list’s lowest normalized score. Raw scores cannot be added directly because they are not on the same scale:
  • BM25 scores have no fixed range and shift with the query’s terms and the collection’s statistics.
  • Vector scores depend on the embedding model and the metric. Some are distances, where lower is better, and some are similarities, where higher is better.
  • Both distributions change from query to query, so a fixed threshold on either is unreliable.
Min-max or z-score normalization per query puts the scores on a common scale, and the weight w is then tuned on labeled queries. This can beat RRF when tuned well, but it is sensitive to outliers and to how many candidates each search returns.

Reranking

A reranker, often a cross-encoder model or an LLM, reads the query and each candidate together and produces a new relevance score. It is usually more accurate than either first-stage search but far more expensive per document, so it runs only on the fused top candidates.

Why does hybrid search matter for RAG?

Hybrid search matters for retrieval-augmented generation (RAG) because many questions mix a concept with an identifier, and vector retrieval alone can miss the one chunk that names it; adding BM25 reduces that risk. In “what changed in the refund policy for plan B-7,” BM25 finds the chunks that contain “B-7” even when their embeddings are not the closest to the question, and a RAG answer is only as good as the context it retrieves. For the rest of the pipeline, see What is RAG?, and for following relationships from retrieved hits, see GraphRAG.

When is hybrid search worth it?

Hybrid search pays off when queries mix natural language with exact terms: support tickets, product catalogs, code and logs, legal and policy documents, and agent memory. It adds less when every query is conversational and the corpus has no identifiers, or when every lookup is an exact key. The cost is a second index to build and keep current, a query embedding per request, and a fusion step to tune. Measure whether it pays for itself: compare BM25 alone, vector alone, and the fused result on a set of real queries with known relevant documents. In HelixDB, vector and BM25 text indexes are optional access paths over a label and a top-level property, on nodes or edges. When you define both on the same label, both searches rank records of that label:
  • One request is one ACID transaction over a committed snapshot. A single request can build a candidate set with a graph traversal, then run a prefiltered vector search and a prefiltered BM25 search over that same candidate set.
  • Prefiltering guarantees that neither search returns a result outside the candidate set, such as documents the current user cannot read. Each prefiltered search returns at most 800 results. Prefiltered BM25 takes its term statistics from the full tenant partition, not only from the candidate set.
  • Vector results are ordered by distance, closest first, then ID. BM25 results are ordered by score, then ID. These orders give the ranks that rank-based fusion uses.
  • The application computes embeddings; HelixDB stores and indexes the vectors.
  • HelixDB has no built-in rank-fusion operator. The application fuses the two result lists, for example with RRF, and applies any reranking.
See Prefiltered search, Vector indexes, and Text indexes.

Frequently asked questions

On workloads that mix exact terms with natural language, it often improves recall of exact terms while keeping most semantic matches, but it is not automatically better. Fusion reorders the combined list, so some hits that only vector search found can fall below the cutoff.

Can hybrid search combine more than two searches?

Yes. RRF sums over any number of ranked lists, so one query can fuse, for example, BM25 over titles, BM25 over body text, and vector search. Weighted blending also extends to more lists, but each list needs its own normalization and weight, which means more tuning.

How many results should each search return before fusion?

More than the final number you need, because a document ranked modestly in both lists can outrank one ranked highly in only one. The right depth depends on your corpus and latency budget, so tune it along with the fusion method.

Does hybrid search need two databases?

No. It needs a keyword index and a vector index over the same content. When both live in one database, the two searches can read the same snapshot of the data and apply the same filters, for example within one transaction. See Do you need separate graph, vector, and text databases?.

What is BM25?

The ranking function behind most keyword search, explained.

What is vector search?

Nearest neighbors, approximate search, and recall.

What are vector embeddings?

How models turn text and other inputs into comparable vectors.

What is retrieval-augmented generation (RAG)?

Grounding model answers in retrieved context.

What is filtered vector search?

Pre-filtering, post-filtering, and returning the right top k.

Prefiltered search guide

Rank only the records a traversal reaches in HelixDB.