Hybrid Search: When Pure Vector Search Isn’t Enough
The Kitebase support bot from article 04 handles “How do I reset my password if I never set a recovery email?” well. Then a customer pastes the message from a failed invite: “What does error KB-4012 mean?” The help center answers it in one sentence, in a section called “If an invite fails”. Vector search ranks that section 9th out of 18, behind every chunk about something going wrong. The model gets three excerpts about lockouts and blank pages, and says it doesn’t know.
Nothing is broken. Vector search matches meaning, and “KB-4012” has no meaning to match. The fix is old: run a plain keyword search next to it and merge the two result lists. That’s hybrid search, and it’s the default retrieval setup in production RAG systems.
What you’ll build: the Kitebase search from article 04 with keyword search added next to vector search and the two rankings merged, plus three test questions where each search alone misses one and the hybrid gets all three into the top 3. It all runs offline.
Why vector search misses exact strings
An embedding (article 04) turns text into a vector so that similar meanings get similar vectors. That’s what lets “I’m locked out and have no backup address” find a chunk that says “never set a recovery email”: different words, same need.
The flip side: an embedding keeps the gist and loses the details. The gist of “What does error KB-4012 mean?” is “something went wrong”, so that’s what vector search finds.
Article 04’s offline stand-in embedder hashed words, which is really keyword matching and would have found KB-4012 easily. This article needs a stand-in that behaves like a real model, so stand_in.py maps words to twelve hand-written concepts and ignores words it has no concept for:
CONCEPTS = {
"access": "access locked lock locks login log sign signing",
"trouble": "error fail fails failed can't won't didn't isn't stopped stuck blank ...",
"recovery": "recovery backup reset forgot",
# ... nine more
}
“Locked”, “can’t” and “error” all count toward trouble. “KB” and “4012” count toward nothing. That’s an exaggeration of real models, which do see an identifier: they split it into sub-word tokens (article 01), short pieces like “KB” and “40” with little meaning of their own. So KB-4012 tends to land close to KB-4015 and to error text in general. Same effect, less extreme.
Here’s each search failing on a different question:
The left panel is the vector problem, and it hits every exact string: error codes, product and plan names, SKUs, API paths, version numbers, people’s names. Users paste these all the time.
Keyword search with BM25
The obvious keyword search counts how many question words each chunk contains. That goes wrong in two ways. Common words swamp rare ones: “Kitebase” is in 12 of the 18 chunks, so matching it says almost nothing. And long chunks win just by containing more words.
BM25 (“Best Match 25”, a ranking formula from the 1990s) fixes both. For each question word found in a chunk, it adds points:
- Rare words are worth more. This weight is the IDF (inverse document frequency): a word in 1 of 18 chunks scores 2.54, “email” (in 7 of 18) scores 0.93, “Kitebase” 0.42.
- Repeats are worth less each time. “Backup” twice scores 3.48, not twice 2.53.
- Long chunks are penalised a little, so a match in a short chunk counts more.
The whole scorer is small:
def idf(self, term: str) -> float:
n, df = len(self.docs), self.doc_freq[term] # df: how many chunks contain the term
return math.log(1 + (n - df + 0.5) / (df + 0.5))
def scores(self, query: str) -> list[float]:
terms = words(query) # lowercase, split, drop filler words like "what" and "does"
result = []
for doc, length in zip(self.docs, self.lengths):
score = 0.0
for term in terms:
tf = doc[term] # times the term appears in this chunk
score += self.idf(term) * tf * (K1 + 1) / (tf + K1 * (1 - B + B * length / self.avg_length))
result.append(score)
return result
K1 = 1.2 controls how fast repeats stop helping and B = 0.75 how much length matters. They’re the defaults in Elasticsearch and Lucene, the search library under it; leave them alone.
Only one chunk contains any of those words, so BM25 returns exactly one result, and it’s the right one.
Two gotchas. First, the tokenizer (the code that splits text into words) decides what can match. This one splits “KB-4012” into “kb” and “4012”. One that dropped numbers would make every error code unfindable. Test yours with your real codes and SKUs.
Second, keyword search matches words, not phrases. Ask about “KB-4015”, a code the help center doesn’t mention, and BM25 still returns the KB-4012 section at 6.55, because “error”, “kb” and “mean” match. And it has no idea of meaning: in the right panel of the first diagram, “backup address” matches the two-factor chunk’s “backup codes”.
Run both and merge the rankings
Hybrid search runs both searches on every question. Each returns its best candidates (10 here, 20 to 50 at scale), and you merge the two lists into one ranking. The top 3 of that go into the prompt, as in article 04.
The obvious merge is adding the scores. It doesn’t work, because they’re on different scales. Cosine similarity, the vector score from article 04, runs from 0 to 1. BM25 has no upper limit: 8.74 here, and it grows with every matching word. The companion code’s score_sum does exactly that, and for the paraphrase question it puts the right chunk 5th, where BM25 alone put it. The vector side barely counts.
You could rescale both to 0 to 1 first, and some databases offer that. But the best BM25 score is 4.11 for one question and 8.74 for the next, so any rescaling is a guess. The usual default skips scores altogether.
Reciprocal rank fusion
Reciprocal rank fusion (RRF) ignores the scores and uses only the positions. Each list gives each chunk a vote worth 1 / (k + rank), and you add up the votes. A chunk near the top of both lists collects two big votes. A chunk found by only one search gets one.
def rrf(*rankings: Hits, k: int = 60) -> Hits:
fused = defaultdict(float)
by_id = {}
for ranking in rankings:
for rank, (_, chunk) in enumerate(ranking, start=1): # only the rank is used, never the score
fused[chunk.id] += 1 / (k + rank)
by_id[chunk.id] = chunk
return sorted(((score, by_id[cid]) for cid, score in fused.items()),
key=lambda pair: pair[0], reverse=True)
Here it is on the paraphrase question, and on the KB-4012 one:
k is there to flatten the votes. With k = 60, 1st place is worth 0.0164 and 5th 0.0154, so no single list can dominate and agreement between the two is what counts. With k = 0, 1st would be worth 1.0 and 5th 0.2, and whichever search ranked a chunk first would decide on its own. 60 comes from the original RRF paper and is the usual default. Leave it there.
The ranked helper that feeds rrf also drops chunks that scored zero. That matters: for the KB-4012 question BM25 matched one chunk. Without the filter, the other 17 would get BM25 votes in arbitrary file order.
Run the three test questions through all three searches:
Where the right chunk lands (top 3 go to the model):
question vector BM25 hybrid
How do I reset my password if I never set a recovery email? 3 1 2
I'm locked out and have no backup address 1 5 2
What does error KB-4012 mean? 9 1 1
Vector search alone gets two of three into the top 3. BM25 alone gets two of three. The hybrid gets all three.
RRF isn’t magic. It rewards agreement, which is why the paraphrase’s right answer comes 2nd, behind a chunk both searches liked. And it only reorders what the searches found: if neither returns the right chunk, the hybrid won’t either.
Metadata filters
A member who isn’t an admin asks about invoices and gets the admin-only billing pages. Or a question about one plan retrieves the page for another.
Metadata is the extra fields you store next to each chunk: product area, plan, language, who may see it. A filter narrows the chunks before either search scores anything. For Kitebase, tag each chunk with an area (sign-in, team, billing, troubleshooting), and a question asked from the billing screen only searches area = billing.
The most reliable filters come from your app, not the question: the page the user is on, their plan, their role. Extracting them from the question text is a fallback that guesses.
Apply the same filter to both searches. If only the vector search is filtered, BM25 brings the other chunks straight back into the fused list. The same mechanism handles permissions, as in article 04: filter by what the user may see, never ask the model to hide it.
Hybrid search in a real database
You won’t run the Python BM25 in production; it’s there so you can see every number.
- Elasticsearch and OpenSearch use BM25 as their default ranking and have vector search built in, so hybrid is mostly configuration.
- Qdrant, Weaviate, Pinecone and Vespa support hybrid search. Check which fusion they use.
- Postgres can do it with the
pgvectorextension for vectors plus its built-in full-text search, fused with RRF in a SQL query. Two catches: Postgres’sts_rankisn’t BM25 (it doesn’t weigh rare words higher), and its query parsers require every word to match by default, so you need to OR the words together. Extensions such as ParadeDB’spg_searchadd real BM25, but for catching exact codes plain full-text search is often enough.
The full pipeline is: apply filters, run both searches in parallel (20 to 50 candidates each), fuse with RRF, optionally re-rank (the re-ranker from article 04), and put the best 3 to 5 in the prompt. In parallel, hybrid adds little latency; the real cost is a second index to keep in sync. It’s worth it: in Anthropic’s contextual retrieval experiments, adding BM25 to embedding search (both with their contextual trick) cut the share of questions whose right chunk was missing from the top 20 from 3.7% to 2.9%, and a re-ranker on top brought it to 1.9%.
Try it yourself
The companion example is the Kitebase help center from article 04 plus two new articles (inviting teammates, troubleshooting): 8 files, 18 chunks. It runs BM25, the stand-in vector search and RRF, and prints the three rankings side by side.
Download the runnable example (zip)
cd 06-hybrid-search
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Question: What does error KB-4012 mean?
vector (stand-in) keyword (BM25) hybrid (RRF, k=60)
1 account-recovery#3 0.69 invite-teammates#2 8.74 invite-teammates#2 0.0309
2 troubleshooting#1 0.62 - account-recovery#3 0.0164
3 troubleshooting#2 0.36 - troubleshooting#1 0.0161
Then try these:
python main.py "I'm locked out and have no backup address". Findaccount-recovery#1in each column: 1st, 5th, and 2nd after fusion.- In
Searcher.hybrid, swaprrfforscore_sumand run step 1 again. The right chunk drops to 5th, because BM25’s bigger numbers drown out the vector scores. python main.py "What does error KB-4015 mean?". That code isn’t in the help center, and BM25 still returns the KB-4012 section, because three of the four words match.
pip install pytest && pytest -q runs the offline tests. Set OPENAI_API_KEY to use text-embedding-3-small for the vector side and see where a real model ranks KB-4012.
Common beginner mistakes
- Testing only with nicely worded questions. Vector search looks great on “how do I reset my password”. Put real questions from your logs in your test set: users paste error codes, order numbers and product names.
- Adding raw scores. BM25 and cosine live on different scales, so the bigger one wins. Use RRF.
- Filtering only one search. The unfiltered side sneaks the excluded chunks back into the fused list.
- Letting zero scores vote. A chunk with no matching words isn’t a BM25 result. Drop it before fusion.
- Tuning weights and
kwithout an eval set. Without 30 to 50 real questions with known answers, you can’t tell whether a change helped. Article 08 shows how to build one.
Questions you will face in production
“My users write full sentences. Do I still need keyword search?” Probably. Full sentences often contain an exact string: a plan name, a setting, an error message. Check your logs for questions containing codes, IDs or names. If there are any, add BM25.
“How many candidates should each search return?” Start with 20 each, fuse, and send the top 3 to 5 to the model. Go up to 50 if you add a re-ranker, since it can only promote what it’s given.
“Should one search count more than the other?” Some databases let you weight each list in the fusion. Start unweighted and change it only when your evals show one search is consistently better on your questions.
Check your understanding
A customer asks "Does the Relay integration work with Jira?" and vector search returns general integrations pages. Which search finds the Relay page, and why?
BM25. “Relay” is a product name that appears in very few chunks, so it has a high IDF and the chunk that contains it scores well above the rest. To an embedding model it’s just an ordinary English word, so it adds little to the vector.
You merge by adding scores: a chunk has cosine 0.9 and no BM25 match; another has cosine 0.2 and BM25 12.0. Which wins, and what should you do?
The second, 12.2 to 0.9, because BM25’s scale is so much bigger. That’s the vector side being ignored, not a real judgment. Merge by rank with RRF instead.
With k = 60, chunk A is 1st in BM25 and missing from vector search. Chunk B is 3rd in both. Which ranks higher after fusion?
B: 1/63 + 1/63 = 0.0317, against A’s 1/61 = 0.0164. RRF rewards agreement. If A was the exact-code answer, you’d catch it in evals and fix it with more candidates per search or a re-ranker, not by abandoning fusion.
You add an area = billing filter to the vector search, and non-billing chunks still reach the prompt. Why?
The BM25 search isn’t filtered, so it returns chunks from every area and RRF merges them in. Apply the same filter to both searches before fusion.
What to remember
- Vector search matches meaning and misses exact strings: error codes, product names, SKUs, IDs.
- BM25 matches exact words, weighs rare words higher, and misses paraphrases.
- Hybrid search runs both on every question and merges the two rankings.
- Merge with reciprocal rank fusion:
1 / (60 + rank)from each list, added up. Never add raw scores. - Apply metadata filters to both searches, and take filters from your app’s context where you can.
- Check every change against real questions with known answers.
What to study next
Hybrid search makes the keyword side exact. The vector side is only as good as the embedding model behind it, so article 07: Embeddings Deep Dive covers how embedding models work, how to choose one and how to benchmark it on your own data. After that, article 08: LLM Evaluation Pipelines shows how to measure whether changes like this one actually helped.
Further reading
- Robertson and Zaragoza (2009): The Probabilistic Relevance Framework: BM25 and Beyond. The standard BM25 reference, including where the formula comes from.
- Cormack, Clarke and Büttcher (2009): Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods. The RRF paper, and the source of k = 60. Short and readable.
- Anthropic: Introducing Contextual Retrieval. Embeddings plus BM25 plus a re-ranker on real benchmarks. The source of the failure-rate numbers above.
- Elastic: Practical BM25, Part 2: The BM25 Algorithm and its Variables. A friendly walk through IDF,
k1andbwith examples.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.