The production checklist
By the end of this chapter — and of the series — you have the complete method as a checklist, the common ways results get invalidated, and a hands-on exercise that runs the whole loop on a bottleneck you induce deliberately so its signature is unmistakable.
The workflow, condensed
The loop is the product. Every artifact in the series — the run sheet, the PromQL set, the troubleshooting table — exists to make one trip around it cheap and reliable.
The final checklist
Before calling a service performance-tested, each line should have an answer you can point at:
- Objectives: per-endpoint p95/p99 SLOs, error-rate bound, and a stated sustained-traffic target — written before the first test, not reverse-engineered from the first results.
- Production-like data: dataset size and skew large enough to make indexes and caches honest (chapter 05); reset strategy chosen.
- Instrumentation:
http_server_requestshistogram buckets, JVM/GC/CPU series, HikariCP and Tomcat pools, plus at least one business timer — all scraped and visible in Grafana. - Repeatability: run sheet filled for every meaningful run; identical re-run establishes the noise floor.
- Test scenarios: at minimum a read-heavy profile, a mixed read/write profile, and a stepped stress run; spike and soak where their failure modes apply.
- Database observability:
pg_stat_statementsenabled;EXPLAINevidence for the top queries under the test workload. - JVM diagnostics: heap, GC pause frequency and overhead, allocation rate — and the ability to take thread/heap dumps when needed.
- Regression thresholds: Gatling assertions encode the SLOs so a regression fails a build, not just a report.
- Capacity assumptions: measured knee under the stated profile; scaling limit identified (app vs database vs infrastructure); autoscaling lag measured, not assumed.
- Documented results: for every accepted change — before/after metrics, the hypothesis, and the run sheets.
Common mistakes
The short version of every failure this series is designed to prevent:
- Reporting mean latency or RPS as the result. Averages hide the tail; RPS without assertions is a description of what was offered, not what the service survived.
- Testing with the wrong workload model. Closed-model tests on open-model traffic commit coordinated omission and underreport tail latency — chapter 07.
- Equal traffic to every endpoint. Real distributions are skewed; the endpoint that breaks first under realistic mix is never the one equal weighting would predict.
- No warm-up. Cold-JVM numbers are real numbers — of a system that will never exist again after the first minute.
- Changing several variables between runs. The result is un-attributable; chapter 07’s five-step loop is the only version that produces an answer.
- Benchmarking with 1,000 rows. Indexes unused, everything cached, N+1 invisible — you measured a different system.
- Tuning without a signature. “Raise the pool/heap/threads” applied before naming the metric that says which queue filled is guessing with extra steps.
- Trusting numbers from a saturated generator. Check the load machine’s CPU first; a client-side-only gap is the generator, not the service.
- CPU limits in benchmarks without recording them. CFS throttling turns “code latency” into “quota latency” — chapter 12.
- Treating a soak trend as noise. A monotonically rising heap floor or pool count is a leak the short tests are structurally unable to see.
- Stress-testing shared or production environments without authorization. The technical findings do not justify the blast radius.
Capstone exercise: induce, prove, fix, compare
The deliberate-fault exercise that proves the method is in your hands, not just in these chapters:
- Baseline. Fresh database (template reset), run
MixedWorkloadSimulation, record p50/p95/p99 per endpoint and the saturation metrics into a run sheet. Repeat once — you now have the noise floor. - Induce a bottleneck. Remove the
@EntityGraphfromfindByIdWithItemsand letitemsstay lazy — or, for a database-layer fault,DROP INDEX idx_orders_customer_status. One change only. - Predict before you run. Write down what you expect to see: which metric should move first, what the signature should look like. The prediction is what makes the exercise a test of your model rather than a fishing trip.
- Prove it. Re-run the identical profile. Collect the evidence: query-per-request ratio (
spring_data_repository_invocationsvshttp_server_requestsrates),hikaricp_connections_pending,pg_stat_statements, orEXPLAIN. Confirm the signature matches the prediction — or find why it did not. - Fix and compare. Restore the entity graph or re-create the index, reset the data, re-run. Compare against the baseline and the degraded run at the same load levels.
- Write it up. One paragraph: hypothesis, evidence, fix, delta. This is the format real performance reviews want.
Pass/fail for the exercise: you can state, with a graph, which queue the induced defect filled — not merely that “it got slower”.
What to build next
Three directions, each with a way to know it worked:
- Add a second read path with caching (chapter 10 §6). Run ReadHeavy before and after; the evidence you want is
spring_data_repository_invocationsdecoupling from request rate — and then break invalidation on purpose to see the stale-read signature. - Push the search endpoint into deep pagination. Extend
SearchSimulationto walk to page 2000 against the million-row table, captureEXPLAINfor the deep offset, then implement keyset pagination and compare p99 by page depth. Done when you can state the crossover page where keyset wins. - Run the stepped stress profile in Kubernetes under the chapter-12 manifest, with and without CPU limits. The deliverable is the
container_cpu_cfs_throttledoverlay on the p99 graph — the single most convincing way to internalise what limits do to tail latency.
Recap
Measure the vocabulary before the system: SLI/SLO, percentiles, arrival rate versus concurrency. Instrument before loading: the metrics exist so every “why” has a named series. Control the environment so runs compare. Generate load with the open model and real data volume. Read results by finding the knee and asking which queue filled first. Change one variable, prove the hypothesis, and record it — then repeat at the boundaries with stress, spike, and soak. Everything else in the series is detail in service of that loop.
Interview-style discussion questions
Five questions you can now answer in depth — the kind a staff-level interview uses to separate “ran a load test once” from “owns the method”:
- “Your p50 is 20 ms but p99 is 2 s. Is the service healthy?” — What the tail is made of, why percentiles rather than averages, and which saturation metrics you would pull before answering.
- “A load test shows throughput rising but latency flat. Is that success?” — Offered vs achieved rate, error ratio, and why the assertions define the answer.
- “How do you size a HikariCP pool?” — Little’s law applied to connection hold time, the
max_connectionsceiling across replicas, and why a bigger pool can move the queue rather than remove it. - “The same code runs slower in Kubernetes than on your laptop. Where do you look first?” — CPU throttling and CFS quanta,
container_cpu_cfs_throttled_seconds_total, per-pod metrics during autoscale events, and the client/server latency gap. - “Why might your load test be underreporting latency?” — Coordinated omission in closed models, generator saturation, missing warm-up — and the checks that rule each out.
What this series built
Across thirteen chapters a minimal Spring Boot service grew into a measured system: schema and indexes owned by SQL, data volume chosen so caches cannot lie, Gatling profiles matched to real question shapes, Prometheus holding the server’s own account of each run, and a one-variable-at-a-time loop that turns “it got slower” into a named queue, a hypothesis, and a re-test. The numbers in it are example assumptions; the method — define, measure, isolate, verify — is not.