Hybrid search — fuse dense and sparse with RRF
Dense vector search captures meaning but misses exact tokens like SKUs and error codes that sparse keyword search nails, so hybrid search runs both and fuses the two ranked lists by RANK (Reciprocal Rank Fusion), which is scale-free and needs no score-calibration.
Filters constrain WHICH items are eligible; they don't fix similarity's own blind spot. Dense vectors understand meaning but miss exact tokens, while keyword search is the mirror image — so hybrid search runs both and fuses by rank with RRF, getting paraphrase and exact-token matches at once. We now have a feature-complete search on one machine. The last wall is the one from scene 1: one machine can't hold a billion vectors.
Scene 13
Hybrid search — fuse dense and sparse with RRF
- Watch
- Try it
- Predict
- Capture
One search box has to handle two kinds of query: plain English ('comfy gaming chair') and exact strings ('SKU-44871'). Watch what happens to the query 'cybersport desk' under two different retrievers. On the LEFT, dense search turns the query into a vector and ranks by MEANING — it returns 'Gaming desk' and 'Esports table' because they mean the same thing, even though neither contains the word 'cybersport'. On the RIGHT, sparse search ranks by the literal TOKENS — it returns 'Standing desk', 'Office desk', 'Desk lamp' because they all contain the word 'desk'. Notice the blind spots: dense never returns the plain literal-token row, and sparse never returns the paraphrase 'Esports table'. Neither list alone is right.
Highlighted lines are the ones running in the diagram right now.
def rrf_fuse(ranked_lists, k=60):score = defaultdict(float)for ranking in ranked_lists: # dense, sparse, ...for rank, doc in enumerate(ranking, start=1):score[doc] += 1.0 / (k + rank) # rank, not scorereturn sorted(score, key=score.get, reverse=True)
def naive_score_add(dense, sparse):score = defaultdict(float)for doc, s in dense.items(): # cosine ~0..1score[doc] += sfor doc, s in sparse.items(): # BM25 ~0..30 <-- swamps cosinescore[doc] += sreturn sorted(score, key=score.get, reverse=True)
Where this sits in Build a vector database (Pinecone / Weaviate / pgvector style)
Scene 13 of 15, in the Production act — Filters, the graph-disconnection trap, hybrid RRF, and sharded scatter-gather.. Dense vectors capture meaning but miss exact tokens like SKUs and error codes; sparse keyword search nails them. Run both and fuse by rank with RRF — scale-free, no score calibration.
Up next. Everything so far lived on one machine. A billion vectors and their index don't fit on one box — so we split them across machines, ask every machine for its local best, and merge. And once we've built this whole engine, it's worth seeing where it actually sits in 2026: right next to the database, powering every LLM app.
All 15 scenes in Build a vector database (Pinecone / Weaviate / pgvector style) · Every curriculum