Design your vector database

Every vector-DB deployment is a deliberate point on the recall/latency/memory frontier plus filter, hybrid, and sharding choices — and the right configuration for RAG is wrong for recommendation, wrong for a semantic cache, and wrong for billion-on-a-budget, even though the primitives are identical.

Previously

The same primitives — vector, metric, index, filter, hybrid, shards — configure radically different systems. The capstone is choosing them deliberately for a real workload, with every knob traceable to the scene that justified it, and the trilemma as the compass.

Scene 15

Design your vector database

  1. Watch
  2. Try it
  3. Predict
  4. Capture
ACTIVERAG over docs50M docs · recall matters · fit…ragRecommendation500M items · latency-critical ·…recommendationSemantic cache2M Q&As · sub-2 ms · recall for…semantic-cacheBillion-on-a-budget1B vectors · 128 GB RAM hard ca…billion-budgetDESIGN PALETTEIndexFlat (exact)IVF + nprobeHNSW + M + ef_searchIVFPQ (compressed)Metriccosine, normalized at …dot product, un-normal…Hybriddense + sparse, fused …Shardingclustered shards + ove…single node (no shardi…Re-rankre-rank top-N with exa…VERIFIER✓FITSHNSW + M + ef_search↳ 50M fits in RAM and recall matters →…✓FITScosine, normalized at ingest↳ Normalize once at ingest, then the c…✓FITSsingle node (no sharding)↳ 50M fits one node, so skip the distr…Configuring: RAG over docs — 50M docs · recall matters · fits in RAM · moderate latency
four workloads, one toolkit
every knob traces back to a scene →
What to watch for

Here is everything the arc built, in one place. Four workload cards are docked across the top — RAG over docs, Recommendation, Semantic cache, Billion-on-a-budget — and each card states its real constraints: how many vectors, how tight the latency budget, how much RAM, and how much being wrong actually costs. Down the side is the palette: the index choices (Flat, IVF, HNSW, IVFPQ), the metric choice, hybrid (dense + sparse fused by RRF), sharding, and re-rank. Every one of those knobs was earned in an earlier scene. The job now is not to learn anything new — it's to read each workload's constraints off the trilemma (recall, latency, memory) and pick the honest configuration. The same toolkit; four different right answers.

Where this sits in Build a vector database (Pinecone / Weaviate / pgvector style)

Scene 15 of 15, in the Design canvas act — Pick every knob for RAG, recommendation, semantic cache, or billion-on-a-budget.. Capstone: pick the index, filter, hybrid, and sharding for RAG vs recommendation vs semantic cache vs billion-on-a-budget — each knob traceable to the scene that justified it.

All 15 scenes in Build a vector database (Pinecone / Weaviate / pgvector style) · Every curriculum

Built with Arqly
Every scene in Build a vector database (Pinecone / Weaviate / pgvector style) builds on the one before it.All 15 Build a vector database (Pinecone / Weaviate / pgvector style) scenes