Series overview
Part 18 of 18100% complete
2026-06-01•12 min read

From lab to production: rollout phases and checklists

The lab now has everything this series set out to build: a PostgreSQL source of truth, an outbox relay keeping a versioned, aliased Elasticsearch index within a second of it, a Search API with designed queries, facets, cursors, and caller-scoped authorisation, and a test suite. This chapter turns that into a plan for production at 300 million profiles. It sets out five phases, each with exit criteria that can be checked. Two of the checks are tools you run here: a relevance gate built on the _rank_eval API, and a readiness script that automates part of the production checklist. The chapter ends with the questions that must have answers before launch, and exercises to take the system further.

You need the lab and the user-search project from chapter 17. The chapter takes about 30 minutes.

Phase 1 — Local proof of concept with a small dataset

Goal. Show that the defined workload, chapter 01’s query shapes Q1–Q7, can be answered from a projection, on data you can inspect.

Scope. Chapters 01 to 07: the lab, the capability matrix, the first mapping, and analysis.

Exit criteria.

  • Every query shape Q1–Q7 has a working request, and its result count matches the equivalent PostgreSQL query on the synthetic data, as chapters 05 to 07 checked.
  • The capability matrix is agreed with the people who own the product and the data, including which fields are not in the projection.
  • The team can explain, with _analyze, why each text field produces the terms it does.

Phase 2 — Mapping and query validation

Goal. Settle the mapping and query design before anything depends on them, because changing them later costs a reindex.

Scope. Chapters 05 to 07 and 11 to 13: the reviewed index definition, the query builder, facets, and cursors.

A relevance gate. Unit tests prove that the query has a structure; they cannot say whether it finds the right profiles. The ranking evaluation API can. You give it search requests together with judgements, documents rated as relevant for each request, and it computes a metric such as recall or precision. Compare two versions of the name clause on a mistyped search, judging three Prashant Kumars as relevant:

GET user-profile-read/_rank_eval?filter_path=metric_score,details.*.metric_score
{
"requests": [
{ "id": "typo_plain",
"request": { "query": { "match": { "fullName": { "query": "prashnat kumr", "operator": "and" } } } },
"ratings": [ { "_index": "user-profile-v2", "_id": "600", "rating": 1 },
{ "_index": "user-profile-v2", "_id": "1200", "rating": 1 },
{ "_index": "user-profile-v2", "_id": "1800", "rating": 1 } ] },
{ "id": "typo_fuzzy",
"request": { "query": { "match": { "fullName": { "query": "prashnat kumr", "operator": "and",
"fuzziness": "AUTO", "prefix_length": 1 } } } },
"ratings": [ { "_index": "user-profile-v2", "_id": "600", "rating": 1 },
{ "_index": "user-profile-v2", "_id": "1200", "rating": 1 },
{ "_index": "user-profile-v2", "_id": "1800", "rating": 1 } ] }
],
"metric": { "recall": { "k": 2000, "relevant_rating_threshold": 1 } }
}
{"metric_score":0.5,"details":{"typo_fuzzy":{"metric_score":1.0},"typo_plain":{"metric_score":0.0}}}

The plain match finds none of the judged profiles; the fuzzy clause from chapter 11 finds all three. A real judgement set holds a few hundred recorded searches with their relevant profiles, judged by people who handle profile searches every day. The response’s unrated_docs lists results that have no judgement yet, which is how the set grows. Run it on every change to the mapping, the analysers, the synonyms, or the query builder, and fail the change when the score drops.

Exit criteria.

  • The index definition file for the first production version is reviewed and merged; no field is added without a query shape that needs it.
  • The query builder’s unit tests pass, including the scope regression test (chapter 17).
  • A judged relevance set exists, and its _rank_eval score is recorded as the baseline.
  • The Search API’s contract, including cursors, totals with their exactness, and error responses, is agreed with its consumers.

Phase 3 — Bulk backfill and a replayable sync pipeline

Goal. Prove that production-sized data can be loaded, kept in sync, and rebuilt, and that failures lose nothing.

Scope. Chapters 04 and 14: the outbox, the relay, the backfill, and zero-downtime reindexing. Run this phase in a staging environment with production-scale synthetic data.

Exit criteria.

  • A full backfill of production-sized data completes within the rebuild window you set in chapter 08’s worksheet, and reports zero dead letters.
  • The relay keeps up with the peak change rate, and the outbox backlog returns to zero after bursts.
  • Kill tests pass. Stop the relay in the middle of a batch; stop Elasticsearch during a backfill; restart both. After recovery, the index matches PostgreSQL in counts, and a sample of changed and deleted profiles is correct.
  • The resurrection test passes. Delete profiles while a backfill runs, as chapter 14 did, and confirm none of them reappears.
  • A complete v1-to-v2 reindex, including validation and the atomic swap, has been rehearsed, and so has the rollback.

Phase 4 — Load testing and capacity planning

Goal. Replace the hypothetical inputs in chapter 08’s worksheet with measurements, and confirm the shard plan.

Scope. Chapters 08 and 15, on production-like hardware.

Exit criteria.

  • The Rally track uses realistic data and varied queries drawn from recorded or generated values, not the fixed bodies of the lab track.
  • At the forecast peak search rate, while the relay writes at its peak rate, p99 latency is within the target, measured both in Rally and in the Search API.
  • The same holds with one data node stopped.
  • The worksheet is updated with measured values, and the shard, replica, and node plan is revised from them.
  • Every performance setting in use, such as refresh and bulk sizes, has a measurement behind it, not only a recommendation.

Phase 5 — Production hardening, observability, and a controlled rollout

Goal. Put the system in front of users gradually, with the ability to go back at every step.

Scope. Chapters 16 and 17, and the rollout itself.

The rollout.

  1. Shadow. The Search API serves results from the existing PostgreSQL path, and also runs the Elasticsearch query in the background and logs differences in counts and top results. Users see nothing new.
  2. Internal users. A feature flag sends a small group of internal support agents to the Elasticsearch path. Collect their judgements; they grow the relevance set.
  3. Percentage rollout. Increase the share of traffic in steps, watching latency, errors, and the outbox backlog after each.
  4. Full traffic. Keep the old path available until the rollback window closes.

Rollback at every step is the feature flag, not a redeploy. A problem in the index itself is rolled back with the alias swap from chapter 14.

Exit criteria.

  • Every service uses its own least-privilege API key, with rotation scheduled; the superuser is used only for break-glass access.
  • Monitoring and alerts cover chapter 16’s signal table, including the outbox backlog.
  • Snapshots run on a schedule, a restore has been tested, and snapshot retention fits the erasure deadline.
  • Erasure has been tested end to end: a profile deleted in PostgreSQL disappears from the live index, from retained index versions, and eventually from snapshots.
  • The service-level objectives for latency and freshness have been met at full traffic for the agreed period.

The automated part of the readiness checklist

Some checks are simple enough to run on every deployment. Create es/readiness-check.sh in the lab directory:

es/readiness-check.sh
#!/usr/bin/env bash
# Automated part of the production-readiness checklist (chapter 18), run against the lab.
# In production, run it with a monitoring API key instead of the elastic user.
set -uo pipefail
set -a; source .env; set +a
ES="https://localhost:9200"
es() { curl -s --cacert certs/ca/ca.crt -u "elastic:${ELASTIC_PASSWORD}" "$ES$1"; }
sql() { docker compose exec -T postgres psql -U profiles -d profiles -At -c "$1"; }
failures=0
check() { # name, condition result (0 = pass), detail
if [[ "$2" == 0 ]]; then printf 'PASS %-34s %s\n' "$1" "$3"; else printf 'FAIL %-34s %s\n' "$1" "$3"; failures=$((failures + 1)); fi
}
status=$(es "/_cluster/health?filter_path=status" | grep -o '"status":"[a-z]*"' | cut -d'"' -f4)
[[ "$status" == green ]]; check "cluster health is green" $? "$status"
read_idx=$(es "/_alias/user-profile-read" | grep -o '"user-profile-v[0-9]*"' | tr -d '"')
write_idx=$(es "/_alias/user-profile-write" | grep -o '"user-profile-v[0-9]*"' | tr -d '"')
[[ -n "$read_idx" && "$read_idx" == "$write_idx" && $(wc -w <<< "$read_idx") -eq 1 ]]
check "read and write aliases agree" $? "read=$read_idx write=$write_idx"
dynamic=$(es "/$read_idx/_mapping?filter_path=*.mappings.dynamic" | grep -o '"dynamic":"[a-z]*"' | cut -d'"' -f4)
[[ "$dynamic" == strict ]]; check "mapping is strict" $? "dynamic=$dynamic"
refresh=$(es "/$read_idx/_settings?include_defaults=true&filter_path=**.refresh_interval" | grep -o '"refresh_interval":"[^"]*"' | head -1 | cut -d'"' -f4)
[[ "$refresh" != "-1" ]]; check "refresh is enabled" $? "refresh_interval=$refresh"
fielddata=$(es "/$read_idx/_mapping" | grep -c '"fielddata":true')
[[ "$fielddata" == 0 ]]; check "no fielddata on text fields" $? "fielddata fields=$fielddata"
disk=$(es "/_cat/allocation?h=disk.percent" | awk 'NF {print $1; exit}')
[[ "$disk" -lt 85 ]]; check "disk below low watermark (85%)" $? "disk=${disk}%"
backlog=$(sql "SELECT count(*) FROM profile_search_outbox")
[[ "$backlog" -lt 1000 ]]; check "outbox backlog is draining" $? "pending=$backlog"
dead=$(sql "SELECT count(*) FROM profile_search_dead_letter")
[[ "$dead" == 0 ]]; check "no dead-lettered changes" $? "dead letters=$dead"
pg=$(sql "SELECT count(*) FROM user_profile")
idx=$(es "/user-profile-read/_count?filter_path=count" | grep -o '[0-9]*')
[[ $((pg - idx)) -le $backlog && $((idx - pg)) -le $backlog ]]
check "index count matches PostgreSQL" $? "postgres=$pg index=$idx (tolerance=$backlog)"
last=$(es "/_slm/policy/nightly-profiles?filter_path=*.last_success.time" | grep -o '"time":[0-9]*' | cut -d: -f2)
age_h=$(( ( $(date +%s) * 1000 - ${last:-0} ) / 3600000 ))
[[ -n "$last" && "$age_h" -lt 26 ]]; check "snapshot succeeded in last 26h" $? "age=${age_h}h"
echo "failures: $failures"
exit $(( failures > 0 ))
Terminal window
./es/readiness-check.sh
PASS cluster health is green green
PASS read and write aliases agree read=user-profile-v2 write=user-profile-v2
PASS mapping is strict dynamic=strict
PASS refresh is enabled refresh_interval=1s
PASS no fielddata on text fields fielddata fields=0
PASS disk below low watermark (85%) disk=20%
PASS outbox backlog is draining pending=0
PASS no dead-lettered changes dead letters=0
PASS index count matches PostgreSQL postgres=999997 index=999997 (tolerance=0)
PASS snapshot succeeded in last 26h age=0h
failures: 0

Check that it can fail: set refresh_interval on user-profile-v2 to -1 and run it again. It reports FAIL refresh is enabled refresh_interval=-1 and exits with status 1, which stops a deployment pipeline. Set the interval back to 1s afterwards.

The count check allows a difference up to the outbox backlog, because changes in the queue have not reached the index yet. On a busy system, run it when the backlog is small, or compare counts per updatedAt range as chapter 12’s composite aggregation allows.

The production-readiness checklist

Data and mapping

  • The projection holds only the fields in the capability matrix; each personal-data field has a named purpose.
  • Mappings are strict; the index definition is versioned in the repository.
  • Reads and writes go through user-profile-read and user-profile-write, never an index name.
  • Synonyms are search-time and updateable; the request-cache clear is part of the change procedure.

Sync and recovery

  • Every change reaches the index through the outbox (or CDC); nothing writes to Elasticsearch directly.
  • External versions come from a database counter, not a timestamp.
  • The backfill restores index settings in all cases, and raises gc_deletes while it runs.
  • Rebuild from PostgreSQL and restore from a snapshot have both been rehearsed, with their durations recorded.

Search API

  • Callers are authenticated; roles decide access to contact details; scope is a base filter.
  • Direct identifiers travel only in request bodies.
  • Results use source filtering; highlights cover only fields the caller may see.
  • Pagination uses cursors; totals report whether they are exact.
  • Timeouts are set from measured latencies; cluster failures return 502 or 503.

Operations

  • Least-privilege API keys per service, with expiry and a rotation procedure.
  • TLS on HTTP and, for multi-node clusters, on transport.
  • Alerts on health, disk, heap, rejections, latency, and outbox backlog.
  • Slow log thresholds set; slow logs treated as personal data.
  • Snapshot policy with retention aligned to the erasure deadline.
  • readiness-check.sh, or its equivalent, runs in the deployment pipeline.

Quality

  • Unit and Testcontainers suites pass against the production Elasticsearch version.
  • The _rank_eval relevance score is at or above the baseline.
  • Rally results at peak load, with a node down, meet the latency target.

Questions to answer before production

Each of these changes a decision somewhere in this series. If one has no answer, the decision it feeds is a guess.

QuestionDecides
What is the peak search rate, per query shape?Replicas and node count (chapter 08)
What are the p95 and p99 latency targets, per endpoint?Timeouts, shard size, whether index sorting is worth it (chapters 10, 15)
What is the average and maximum document size, measured on real data?Storage, primary shard count (chapter 08)
How many profiles change per second at peak, and in bursts?Relay capacity, refresh interval, write thread pool headroom (chapters 14, 15)
How stale may search results be?Refresh interval, relay poll interval, where exact reads go to PostgreSQL (chapter 04)
How fast must the profile count grow before the plan is revisited?Primary shard count and the reindex cadence (chapter 08)
What availability is required, and how many node or zone failures must be tolerated?Replicas, zones, node count (chapters 08, 16)
What are the recovery point and recovery time objectives?Snapshot frequency, rebuild versus restore (chapter 16)
What is the erasure deadline, and which copies does it cover?Snapshot retention, rollback window, log retention (chapters 14, 16)
Which roles may see contact details, and what scopes exist?The Search API’s security rules (chapter 16)
Which scripts and languages do names use?Analysers; whether transliteration is in scope (chapter 06)
Who is on call for the cluster, the relay, and the Search API?Whether to run the cluster yourself or use a managed service

What to build next

Each exercise extends something the series built, and each has a way to know it works.

  1. Automate the reindex. Turn chapter 14’s six steps into one pipeline job: create, backfill with the relay dual-writing, wait for the outbox to drain, run validate and the _rank_eval gate, swap, and schedule the old version’s deletion. Done when a mapping change goes from pull request to live index without manual steps, and a failed validation stops it before the swap.
  2. Add a cursor-based export. An endpoint that streams every profile matching a search, as NDJSON, using a point in time and search_after in pages of 1,000, closing the PIT when done or on failure. Done when an export of all active profiles in Bihar contains exactly PostgreSQL’s count, once each, while the relay is changing profiles.
  3. Parallelise the backfill. Split user_id into ranges, run one keyset reader per range, and share one BulkIngester. Done when Rally-measured throughput on the target cluster improves and the report still shows zero dead letters.
  4. Replace the polling outbox with Debezium. Run Kafka Connect with the Debezium PostgreSQL connector and the outbox event router, partitioned by user_id, and consume it with the same re-read-and-write logic as OutboxRelay. Done when chapter 14’s kill tests pass with the new source, and the relay’s polling query is gone.
  5. Benchmark realistic traffic and revise the shard plan. Replace the Rally track’s fixed query bodies with parameter sources that draw names, states, and cities from production-like distributions; add a parallel indexing task at the peak change rate; run the throughput ramp from chapter 08 with and without a node. Done when the worksheet’s inputs are all measured and the shard plan has been revised from them, or confirmed.
  6. Test the indexer. Add a PostgreSQLContainer next to SharedElasticsearch, load chapter 04’s SQL, and test the relay: an update, a delete, a stale event after a delete, and a mapping-rejected document that must land in the dead-letter table. Done when the resurrection scenario from chapter 03 is a failing test before the re-read logic and a passing one after it.

What this series built

Across eighteen chapters, a one-million-profile lab grew into a system with the shape of one that serves 300 million:

  • PostgreSQL remains the source of truth.
  • Elasticsearch holds a projection derived from it, which it can always be rebuilt from.
  • Every design decision was checked against the lab, and every count against PostgreSQL.

Along the way the series recorded where the usual advice did not hold on this setup: refresh during loads, filter caching, index sorting. It also recorded where the first version of the code was wrong: the poller that skipped events, the facets that leaked counts outside an agent’s scope, and the bootstrapper that created the wrong index version.

The lab’s numbers do not transfer to production. The method does: define the workload, measure before deciding, keep the source of truth authoritative, and let tests and benchmarks decide what stays.

ElasticsearchSpring Boot

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind