Learn / RAG in 7 lessons / Retrieval quality and hybrid search

Lesson 4 of 7 7 min

Retrieval quality and hybrid search

Why pure semantic search misses obvious matches, what hybrid search fixes, and how re-ranking tightens the final top-k.

Where pure vector search falls short

Semantic embeddings are trained to capture meaning, which is exactly why they’re good at matching a question phrased differently from the source text. It’s also exactly why they’re weak at:

  • Exact identifiers - error codes, SKUs, ticket numbers, acronyms. A rare token like ERR_504_TIMEOUT doesn’t move an embedding much; the vector still mostly represents the surrounding prose.
  • Negation and specificity - “return policy” and “no return policy” can embed close together, because they’re topically related even though they mean opposite things.
  • Very short, keyword-style queries - a two-word query has little semantic content for the model to work with.

A plain keyword method like BM25 (term-frequency ranking with document-length normalization - the same family of algorithm this site’s own chatbot uses, in src/lib/chat/bm25.ts) is the mirror image: excellent at exact-token matches, poor at matching different phrasing for the same idea.

Hybrid search: use both, merge the results

Hybrid search runs a query through both a keyword index and a vector index, then merges the two ranked lists into one. The standard merge technique is Reciprocal Rank Fusion (RRF), which sidesteps the problem that BM25 scores and cosine-similarity scores aren’t on comparable scales:

rrf_score(doc) = sum over each ranker of  1 / (k + rank_in_that_ranker)

k is a small constant (commonly 60) that dampens the effect of rank 1 vs rank 2 while still rewarding documents that rank highly in either list. A document that’s #1 in the keyword list and unranked in the vector list still surfaces; so does one that’s #2 in both. No score normalization required - only rank position matters.

Re-ranking: a second, more expensive pass

Retrieval (keyword or vector) has to be fast, because it searches the whole index. That speed comes from comparing lightweight representations. A re-ranker - usually a cross-encoder model that reads the query and one candidate chunk together - is slower per comparison but far more precise, because it can reason about the actual relationship between the two texts instead of comparing pre-computed vectors.

The pattern: retrieve a wider candidate set cheaply (say, top 20-50 from hybrid search), then re-rank only that shortlist down to the top 3-8 that actually go in the prompt. This keeps the expensive step bounded to a small, fixed number of comparisons per query instead of scaling with the size of the index.

What to actually do

  • Start with vector search alone; it’s simpler and covers most natural-language questions.
  • Add BM25 + RRF hybrid search once you notice exact-match queries (IDs, codes, acronyms) failing - a good sign is a support or search log full of short, specific queries.
  • Add a re-ranker last, only if retrieval precision is still the bottleneck after hybrid search - it’s the most expensive lever, so pull it last, not first.

Key takeaways

  • Pure vector search can miss exact matches - an error code, a product SKU, an acronym - because embeddings favor meaning over exact tokens.
  • Hybrid search combines a keyword method (BM25) with vector search and merges the two ranked lists, catching both kinds of query.
  • Reciprocal Rank Fusion (RRF) is a simple, effective way to merge two ranked lists without needing comparable raw scores.
  • A re-ranker (a small cross-encoder model) can improve precision on the merged top-k before it goes into the prompt, at extra latency cost.

Quick check

3 questions - see how much stuck.

1. Why can pure embedding-based search miss a query for an exact error code like 'ERR_504_TIMEOUT'?
2. What is the main idea behind hybrid search?
3. What does a re-ranker do that the initial retrieval step doesn't?