Running a load test correctly
By the end of this chapter you can run a load test whose numbers you would defend: warmed-up, measured in a steady-state window, repeated once to prove it is stable, and changed one variable at a time. You can also explain the two silent ways a load test lies — coordinated omission and generator saturation — and check for both.
The anatomy of a run
Every run in this series has three phases, and only the middle one is measured:
- Warm-up exists because the JVM starts cold: interpreted and JIT-profiling tiers compile hot paths over the first thousands of requests, the HikariCP pool lazily opens its connections, and PostgreSQL pulls working-set pages into
shared_buffers. Measuring the first minute measures the JVM warming up, not your code. One minute at ramping load is a reasonable lab value; the honest test is whether results stop changing between runs — if run 2 differs wildly from run 1, warm-up was too short. - Steady state is the window assertions should describe: flat arrival rate, long enough that percentiles are stable. Three minutes is the lab minimum; for anything you will quote, prefer ten or more.
- Ramp-down / cool-down answers a different question: does the service recover? Latency should collapse back to warm-up levels within seconds of load stopping. If it does not, you have found state that persists — a full pool, a long GC queue, a leak — which is a finding, not noise to discard.
Why duration matters: a 30-second test cannot see GC pressure, connection-pool drift, or a memory leak’s slope. Length is set by what you are hunting: minutes for latency percentiles, tens of minutes for load tests, hours for soak (chapter 11).
Commands
Local, against the app started with ./gradlew bootRun:
# full suite (every simulation class) — rare; prefer one at a time./gradlew gatlingRun --all
# a single simulation — the usual form./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation
# override the target when the app runs elsewhere — the plugin forwards# caller JVM system properties to the simulation process./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation \ -DbaseUrl=http://staging.internal:8080In CI the same Gradle invocation runs inside a job; the only additions are POSTGRES_PASSWORD from CI secrets and --no-daemon if your CI image does not already disable it. Chapter 12 builds the full pipeline step, including why CI perf jobs need generous timeouts and must not run in parallel with other jobs on shared runners.
The discipline: one variable per run
The whole experiment protocol fits in five steps:
- Baseline run. Record the run sheet, run the simulation, save the report.
- Repeat the identical run. Not to double-check Gatling — to measure the noise floor of your lab. If two identical runs differ by 3%, any “improvement” under ~10% is within noise until proven otherwise.
- Change exactly one variable. One index, one pool size, one JVM flag. One.
- Re-run identically — same simulation, same injection profile, same data (or the same reset strategy).
- Compare like with like: p95/p99 per endpoint, error rate, achieved throughput — and the saturation metrics that explain the difference.
The failure this prevents: change the heap size, add an index, and raise the HikariCP pool in one go, and latency improves 20%. Which change did it? You now know something helped and must re-run three more times to find out anyway — so nothing was saved. Worse is the correlated case: the index fix made queries cheap enough that the pool pressure you “fixed” would never have appeared. One variable at a time is not slower; it is the only version that produces an answer.
The same discipline applies to the environment. A laptop that picked up a video call mid-run, a Docker host under memory pressure, a docker pull in progress — all are variable changes you did not make. The run sheet’s “other significant load on the host” line is there because these are not hypothetical.
Coordinated omission: how naive tests underreport tail latency
This is the subtlest and most important idea in the chapter.
A naive load test thinks in requests per second achieved: it starts a request, waits for the response, then starts the next one. When the server slows, the test slows with it — the requests that would have arrived during the stall are simply never sent. The report then shows the server’s response times for the requests that got through, omitting the backlog that formed. That is coordinated omission: the client coordinated its request schedule with the server’s availability and thereby omitted precisely the observations that would have shown the queue.
Concrete numbers. Suppose the configured rate is 100 req/s and a 5-second stall hits the server. In a closed model with 10 virtual users, at most 10 requests are in flight during the stall — the 490 that should have arrived were never issued, and the report shows a modest p99 instead of the 5-second wall every real user hit. An open-model test keeps injecting: those 490 requests arrive into a queue, take seconds each, and the p99 tells the truth.
Gatling’s injectOpen does not commit coordinated omission at the injection level — it schedules users by arrival rate. The remaining trap is on your side: pause() between requests inside a scenario does not throttle arrivals in the open model (each virtual user is independent), but building “throughput” assertions around a closed-model mental picture does. Where coordinated omission still bites even open-model users is in reporting: a system that accepts connections but never responds can look “up” while its p99 explodes — which is why assertions are on percentiles and failure ratio, never on request count.
If you evaluate k6 for teams that prefer it: k6’s constant-arrival-rate executor is its open model and the right default; the default vus executors are closed-model. Same principle, same trap.
Client-side latency ≠ server-side latency
Gatling’s numbers are what the client saw: DNS + TCP + TLS + request + server time + response + generator-side queueing. Micrometer’s http_server_requests is what the server saw: from the servlet container onward. The gap between the two p99s is a diagnostic signal in itself:
- Gap ≈ constant and small (loopback: sub-millisecond): healthy; client numbers valid.
- Gap grows with load: the generator is becoming the bottleneck — its CPU is saturated or its event loop is queuing. Check the load generator’s own CPU; Gatling’s log also reports when it cannot keep up.
- Server-side high, client-side higher with correlation: real queueing inside the app (Tomcat, HikariCP) — chapter 08’s job.
- Server-side flat, client-side rising: almost always the network or the generator, not the service.
That last row is the mistake to avoid: a saturated generator produces beautiful “the API is slow” reports. Before any diagnosis, glance at the machine Gatling runs on — if its CPUs are pegged, the run is invalid regardless of what it measured.
A run, end to end
# 1. known-clean database (write-mix runs: restore the template copy)docker compose down -v && docker compose up -d # or point at orders_test
# 2. start the app fresh; note JVM flags./gradlew bootRun
# 3. warm-up run, discarded./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation # phase-1 ramp only matters here
# 4. the measured run — fill the run sheet./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation
# 5. sanity check: generator CPU was never pegged; server p99 tracks client p99# (compare report vs Grafana panel for the same window)Milestone check
Before moving on, you should be able to say of your last run: how long warm-up was and why; which injection profile produced the measurement window; what the noise floor between identical runs is; and whether the generator ever approached saturation. Chapter 08 is about reading what the run produced.