Concept
Hybrid search is a retrieval approach that runs a keyword search, typically BM25, and a
vector search for the same query, then merges the two ranked lists into one. Keyword
search catches exact names, IDs, and error codes; vector search catches paraphrases
and related concepts. Fusing them, often with reciprocal rank fusion, gives results
that hold up for both kinds of query.
Learning objectives
After reading this article you will be able to:
- Explain why BM25 and vector search fail in complementary ways
- Describe the hybrid pipeline: retrieve from both indexes, fuse, optionally rerank
- Calculate reciprocal rank fusion scores and compare them with weighted blending
- Recognize when hybrid search is worth its extra cost
How do keyword search and vector search differ?
Keyword search matches the words in the query, while vector search matches meaning, so each one succeeds where the other fails. That complementary failure is the whole case for combining them.
A support-ticket search shows the gap. A user typing “E1042” needs the tickets that
contain that exact code, which an embedding may blur into generic “error” text. A user
typing “I can’t log in after resetting my password” needs tickets about authentication
failures, even when they share few words with the query. For each method on its own,
see What is BM25? and
What is vector search?
How does hybrid search work?
Hybrid search sends one query to both a BM25 index and a vector index, fuses the two ranked candidate lists into a single ranking, and optionally reranks the top of that list before returning results.1
Retrieve from both indexes
Run BM25 and vector search for the same query, usually in parallel. Embed the query
with the same model that embedded the documents. Retrieve more candidates from each
search than you plan to return, so good results from one list are not cut off
before fusion.
2
Fuse the lists
Combine the two rankings into one with a fusion method such as reciprocal rank
fusion.
3
Rerank, optionally
Score the fused top results with a slower, more accurate model, then return the
final top results.
How do you combine BM25 and vector search results?
Most systems use rank-based fusion, usually reciprocal rank fusion (RRF), or blend normalized scores with a weight. Either can be followed by a reranker that rescores the fused top candidates.Reciprocal rank fusion
Reciprocal rank fusion ignores raw scores and uses only each document’s rank in each list:k dampens the advantage of the very top ranks. The value 60 comes from
the paper that introduced RRF and is a common convention, not a requirement. Suppose
BM25 returns A, B, C and vector search returns D, C, A:
In this example, documents that appear in both lists rise to the top. In general, RRF
rewards agreement between lists, but a high rank in one list can still outscore low
ranks in both. RRF needs no score normalization and no training data, which makes it a
robust default. Its limitation is that it discards score magnitude: a document that is
far better than the rest in one list gets no extra credit. A weighted variant
multiplies each list’s terms by a weight to favor one method.
Weighted score blending
The other common approach normalizes each list’s scores and blends them:- BM25 scores have no fixed range and shift with the query’s terms and the collection’s statistics.
- Vector scores depend on the embedding model and the metric. Some are distances, where lower is better, and some are similarities, where higher is better.
- Both distributions change from query to query, so a fixed threshold on either is unreliable.
w is then tuned on labeled queries. This can beat RRF when tuned well, but it
is sensitive to outliers and to how many candidates each search returns.
Reranking
A reranker, often a cross-encoder model or an LLM, reads the query and each candidate together and produces a new relevance score. It is usually more accurate than either first-stage search but far more expensive per document, so it runs only on the fused top candidates.Why does hybrid search matter for RAG?
Hybrid search matters for retrieval-augmented generation (RAG) because many questions mix a concept with an identifier, and vector retrieval alone can miss the one chunk that names it; adding BM25 reduces that risk. In “what changed in the refund policy for plan B-7,” BM25 finds the chunks that contain “B-7” even when their embeddings are not the closest to the question, and a RAG answer is only as good as the context it retrieves. For the rest of the pipeline, see What is RAG?, and for following relationships from retrieved hits, see GraphRAG.When is hybrid search worth it?
Hybrid search pays off when queries mix natural language with exact terms: support tickets, product catalogs, code and logs, legal and policy documents, and agent memory. It adds less when every query is conversational and the corpus has no identifiers, or when every lookup is an exact key. The cost is a second index to build and keep current, a query embedding per request, and a fusion step to tune. Measure whether it pays for itself: compare BM25 alone, vector alone, and the fused result on a set of real queries with known relevant documents.How does HelixDB support hybrid search?
In HelixDB, vector and BM25 text indexes are optional access paths over a label and a top-level property, on nodes or edges. When you define both on the same label, both searches rank records of that label:- One request is one ACID transaction over a committed snapshot. A single request can build a candidate set with a graph traversal, then run a prefiltered vector search and a prefiltered BM25 search over that same candidate set.
- Prefiltering guarantees that neither search returns a result outside the candidate set, such as documents the current user cannot read. Each prefiltered search returns at most 800 results. Prefiltered BM25 takes its term statistics from the full tenant partition, not only from the candidate set.
- Vector results are ordered by distance, closest first, then ID. BM25 results are ordered by score, then ID. These orders give the ranks that rank-based fusion uses.
- The application computes embeddings; HelixDB stores and indexes the vectors.
- HelixDB has no built-in rank-fusion operator. The application fuses the two result lists, for example with RRF, and applies any reranking.
Frequently asked questions
Is hybrid search better than vector search?
On workloads that mix exact terms with natural language, it often improves recall of exact terms while keeping most semantic matches, but it is not automatically better. Fusion reorders the combined list, so some hits that only vector search found can fall below the cutoff.Can hybrid search combine more than two searches?
Yes. RRF sums over any number of ranked lists, so one query can fuse, for example, BM25 over titles, BM25 over body text, and vector search. Weighted blending also extends to more lists, but each list needs its own normalization and weight, which means more tuning.How many results should each search return before fusion?
More than the final number you need, because a document ranked modestly in both lists can outrank one ranked highly in only one. The right depth depends on your corpus and latency budget, so tune it along with the fusion method.Does hybrid search need two databases?
No. It needs a keyword index and a vector index over the same content. When both live in one database, the two searches can read the same snapshot of the data and apply the same filters, for example within one transaction. See Do you need separate graph, vector, and text databases?.Related topics
What is BM25?
The ranking function behind most keyword search, explained.
What is vector search?
Nearest neighbors, approximate search, and recall.
What are vector embeddings?
How models turn text and other inputs into comparable vectors.
What is retrieval-augmented generation (RAG)?
Grounding model answers in retrieved context.
What is filtered vector search?
Pre-filtering, post-filtering, and returning the right top k.
Prefiltered search guide
Rank only the records a traversal reaches in HelixDB.