Design your search cluster

Every Elasticsearch deployment is a deliberate trade across shard count, replica count, refresh interval, scoring strategy, and ILM tiering — and the right answer for an e-commerce catalog is wrong for a security-analytics archive even though the primitives are identical.

Previously

The same primitives — shards, replicas, refresh, scoring, tiers — configure radically different deployments. The capstone is choosing them deliberately, with each choice traceable to the scene that justifies it.

Scene 12

Design your search cluster

  1. Watch
  2. Try it
  3. Predict
  4. Capture
ACTIVEE-commerce search · 5M produc…Low write rate, high read rate, BM25 critica…ecommerceApplication logs · 10TB/dayWrite firehose, time-series, cheap historica…logsSecurity analytics · 1B event…Aggregations dominate, top-N over source IPs…securityDESIGN PALETTEprimary shards2primary_shardsreplicas1replicasrefresh_interval1srefresh_intervaldfs_query_then_fetchONdfs_query_then_fetchILM tieringnoneilm_tieringVERIFIER✓2 primary shards — fits a 5M catalog↳ scene 5 · small corpus; few sh…✓replicas=1 — read fan-out for catalog…↳ scene 6 · replicas scale reads…✓refresh=1s — interactive product brow…↳ scene 4✓dfs_query_then_fetch ON — rankings st…↳ scene 9 · global IDF; one extr…✓ILM off — product catalog is small an…↳ scene 11✓FITS THE WORKLOADhonest fit for ecommerceevery ✓ / ✗ cites the scene that justifies it
What to watch for

Three workloads, three honest configurations. The canvas snaps to the defaults for each archetype; read the verifier rows — every ✓ cites the scene that justifies it.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

archetype.ecommerce
small corpus, BM25 critical, dfs ON, ILM off
# 5M product catalog · interactive browse · ranking is part of SLO
# primary_shards = 2 (scene 5: small corpus, few shards)
# replicas = 1 (scene 6: HA + read fan-out)
# refresh = 1s (scene 4: live catalog UI)
# search_type = dfs_query_then_fetch (scene 9: stable rankings
# across reindex)
# terms_agg = default shard_size (scene 10: not the bottleneck)
# ILM = off (scene 11: products don't age out)

Where this sits in Build a distributed search engine (Elasticsearch / OpenSearch style)

Scene 12 of 12. Capstone: pick e-commerce, logs, or security analytics and configure shards, replicas, refresh, scoring, and ILM — the verifier traces every ✓/✗ back to the scene that earned it.

All 12 scenes in Build a distributed search engine (Elasticsearch / OpenSearch style) · Every curriculum

Built with Arqly
Every scene in Build a distributed search engine (Elasticsearch / OpenSearch style) builds on the one before it.All 12 Build a distributed search engine (Elasticsearch / OpenSearch style) scenes