Filtered search — pre vs post

Real queries combine a metadata predicate with similarity, and the two naive orders both break: post-filter (search then drop non-matches) can return fewer than k results, while pre-filter (restrict first, then search) sets up a subtler failure the next scene reveals.

Previously

We finished the clean index-choice trilemma; now production reality: queries carry metadata predicates. Bolting a filter onto vector search has two obvious orders — search-then-filter and filter-then-search — and the first one starves below k. Pre-filtering looked like the safe answer. It isn't, once the index underneath is an HNSW graph.

Scene 11

Filtered search — pre vs post

  1. Watch
  2. Try it
  3. Predict
  4. Capture
acoustic → electroniccalm → energeticLo-fi RainAcoustic Suns…Campfire FolkCoffeehouseIndie DriveSynth DawnNeon CityClub PulseRave PeakBass DropMidnight DriveGarage BeatStudy BeatsLonely SynthNow PlayingPost-filter: searched all 14, dropped non-matches → delivered 2 of 3.RECALL vs LATENCYrecallslower →FlatIVFPQHNSW
What to watch for

Up to now a query was just 'find the closest songs'. Real queries carry an extra condition — a metadata predicate — a plain attribute test like 'genre = electronic', 'in stock', or 'price < $50' that each item either passes or fails. Watch the songs paint in: solid ones match the filter, gray ones don't. Notice the closest songs to 'Now Playing' are a MIX — some match, some are gray. So 'closest' and 'matches the filter' are two different sets, and a real answer has to satisfy BOTH at once.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

PostFilter.search
search everything by closeness, then drop non-matches
def post_filter(query, predicate, k):
# one ANN pass over the whole index
cand = ann_search(query, fetch=FETCH)
out = []
for id in cand: # closest first
if predicate(id): # keep matches
out.append(id)
if len(out) == k:
return out
return out # may be < k
FilteredSearch.route
the ordering choice — and the over-fetch knob nobody can size
def filtered_search(query, predicate, k):
if PRE_FILTER:
return pre_filter(query, predicate, k)
# post-filter over-fetches to survive the drop:
FETCH = k * over_fetch_multiplier
return post_filter(query, predicate, k)

Where this sits in Build a vector database (Pinecone / Weaviate / pgvector style)

Scene 11 of 15, in the Production act — Filters, the graph-disconnection trap, hybrid RRF, and sharded scatter-gather.. Real queries combine a metadata filter with similarity, and both naive orders break: post-filter can return fewer than k results; pre-filter sets up a subtler failure.

Up next. Pre-filter dodged starvation by searching only matching points. But remember HNSW finds neighbors by HOPPING along edges — and if you delete most nodes from the graph, the path to a valid neighbor can run straight through a deleted node. The greedy walk hits a dead end. Let's freeze that picture: the filter that disconnects the graph.

All 15 scenes in Build a vector database (Pinecone / Weaviate / pgvector style) · Every curriculum

Built with Arqly
Every scene in Build a vector database (Pinecone / Weaviate / pgvector style) builds on the one before it.All 15 Build a vector database (Pinecone / Weaviate / pgvector style) scenes