Hybrid search — fuse dense and sparse with RRF

Dense vector search captures meaning but misses exact tokens like SKUs and error codes that sparse keyword search nails, so hybrid search runs both and fuses the two ranked lists by RANK (Reciprocal Rank Fusion), which is scale-free and needs no score-calibration.

Previously

Filters constrain WHICH items are eligible; they don't fix similarity's own blind spot. Dense vectors understand meaning but miss exact tokens, while keyword search is the mirror image — so hybrid search runs both and fuses by rank with RRF, getting paraphrase and exact-token matches at once. We now have a feature-complete search on one machine. The last wall is the one from scene 1: one machine can't hold a billion vectors.

Scene 13

Hybrid search — fuse dense and sparse with RRF

  1. Watch
  2. Try it
  3. Predict
  4. Capture
QUERYcybersport deskReciprocal Rank Fusion · k = 60 · score = Σ 1/(k + rank)DENSE · semantic ANNcosine over embeddings1Gaming desk2Esports table3Office deskSPARSE · BM25 keywordterm-overlap score1Standing desk2Office desk3Desk lamp4Gaming deskFUSED · RRFranked by Σ 1/(k+rank)1Gaming deskBOTH0.03202Office deskBOTH0.03203Standing desk0.01644Esports table0.01615Desk lamp0.0159RRF: each list contributes 1/(k+rank); documents both retrievers liked rise to the top.
dense = ranks by meaning →
← sparse = ranks by exact tokens
each list has a blind spot
What to watch for

One search box has to handle two kinds of query: plain English ('comfy gaming chair') and exact strings ('SKU-44871'). Watch what happens to the query 'cybersport desk' under two different retrievers. On the LEFT, dense search turns the query into a vector and ranks by MEANING — it returns 'Gaming desk' and 'Esports table' because they mean the same thing, even though neither contains the word 'cybersport'. On the RIGHT, sparse search ranks by the literal TOKENS — it returns 'Standing desk', 'Office desk', 'Desk lamp' because they all contain the word 'desk'. Notice the blind spots: dense never returns the plain literal-token row, and sparse never returns the paraphrase 'Esports table'. Neither list alone is right.

Implementation

Highlighted lines are the ones running in the diagram right now.

rrf_fuse
merge ranked lists by reciprocal rank — scale-free
def rrf_fuse(ranked_lists, k=60):
score = defaultdict(float)
for ranking in ranked_lists: # dense, sparse, ...
for rank, doc in enumerate(ranking, start=1):
score[doc] += 1.0 / (k + rank) # rank, not score
return sorted(score, key=score.get, reverse=True)
naive_score_add
the broken fix — adds incomparable scales
def naive_score_add(dense, sparse):
score = defaultdict(float)
for doc, s in dense.items(): # cosine ~0..1
score[doc] += s
for doc, s in sparse.items(): # BM25 ~0..30 <-- swamps cosine
score[doc] += s
return sorted(score, key=score.get, reverse=True)

Where this sits in Build a vector database (Pinecone / Weaviate / pgvector style)

Scene 13 of 15, in the Production act — Filters, the graph-disconnection trap, hybrid RRF, and sharded scatter-gather.. Dense vectors capture meaning but miss exact tokens like SKUs and error codes; sparse keyword search nails them. Run both and fuse by rank with RRF — scale-free, no score calibration.

Up next. Everything so far lived on one machine. A billion vectors and their index don't fit on one box — so we split them across machines, ask every machine for its local best, and merge. And once we've built this whole engine, it's worth seeing where it actually sits in 2026: right next to the database, powering every LLM app.

All 15 scenes in Build a vector database (Pinecone / Weaviate / pgvector style) · Every curriculum

Built with Arqly
Every scene in Build a vector database (Pinecone / Weaviate / pgvector style) builds on the one before it.All 15 Build a vector database (Pinecone / Weaviate / pgvector style) scenes