A vocabulary for measuring
By the end of this chapter you can name the seven kinds of performance test, say which question each one answers, and write acceptance criteria for an API that a Gatling report can actually pass or fail. You also get the queueing model the rest of the series uses to explain why response times degrade non-linearly — and why a high requests-per-second number on its own proves nothing.
This series is for developers who already build and run Spring Boot services. You should be comfortable with Java, REST controllers, Spring Data, and basic unit or integration testing; none of those are explained here. What the series does explain is the measurement method: how to instrument a service, generate honest load, read the result, and change exactly one variable at a time. The running example is an Order Management API — create an order, fetch it, search orders by customer and status, update an order’s status — built on PostgreSQL and tested with Gatling.
Versions and labels used throughout the series
| Component | Version | Check with |
|---|---|---|
| Java | 21 | java -version |
| Gradle | 9.x, via the wrapper | ./gradlew --version |
| Spring Boot | 4.1.1 | build.gradle.kts |
| PostgreSQL | 18 | SELECT version(); |
| Gatling | 3.15.1 via io.gatling.gradle 3.15.1.3 | ./gradlew dependencies |
| Prometheus | v3.7.0 | docker inspect on the container |
| Grafana | 12.2.0 | docker inspect on the container |
| Docker Engine with Compose v2 | any current release | docker compose version |
Numbers in this series carry one of three labels, because a configuration that is right for one workload is wrong for another:
- Principle: true of the JVM, Spring Boot, or queueing in general, with a link to the primary source where it matters.
- Example assumption: a choice made for the Order Management API. Change it when your context differs.
- Needs validation: a decision only your own measurements can settle. The series gives the method, not the answer.
Seven tests, seven questions
“Performance testing” is the umbrella term, not a test. Each variant answers a different question, generates a different traffic pattern, and fails in a different way. Conflating them is the most common reason a load-test report gets filed and ignored: the team ran a stress test, reported a throughput number, and nobody asked the question the number answers.
| Test | Question it answers | Traffic pattern | Expected outcome | Common failure signal | Example use case |
|---|---|---|---|---|---|
| Performance test | Does the service meet its latency and throughput targets under expected load? | Representative mix at target rate, held steady | SLOs met for the full duration | p95 or p99 above target while throughput looks fine | Release check: “p99 < 300 ms at 200 req/s” |
| Load test | Does the service hold up at the expected peak, for a sustained period? | Peak business load, ramped then held for tens of minutes | Stable latency, no error growth, no resource exhaustion | Latency creeping upward across the steady phase | Pre-launch validation of the Order API at projected Black Friday peak |
| Stress test | Where is the breaking point, and does the service recover? | Load increased stepwise past capacity until SLOs fail | Degradation is gradual and recoverable, not a cliff | Errors spike abruptly; recovery does not happen after load stops | Find the ceiling before marketing announces a campaign |
| Spike test | What happens when traffic arrives far faster than the service or autoscaler can react? | Sudden step change — e.g. 10x in under a minute — then back down | Queueing and brief elevated latency, then return to SLO | Errors during the spike; latency that never settles back | Flash sale start, push-notification-driven burst |
| Soak (endurance) test | Does anything degrade over hours at moderate load? | Moderate, steady load for 4–24+ hours | Flat memory, connection counts, and latency | Heap or connection pool climbing monotonically; GC pauses lengthening | Catching a connection leak that only matters over a weekend |
| Capacity test | How much load can one instance (or the cluster) serve within SLO? | Incremental levels with a measurement at each | A defensible “max sustained req/s” number | The knee: the level after which latency accelerates | Sizing replicas; answering “how many pods for the launch?” |
| Scalability test | Does adding resources add proportional capacity? | Same workload, repeated at 1, 2, 4 instances | Near-linear throughput gain until a shared bottleneck appears | Throughput flat or worse with more replicas | Proving the database, not the app, is the scaling limit |
Two things in that table do more work than the rest. The traffic pattern column is the test design — a soak test with a spike profile is just a badly-run stress test. And the failure signal column is the exit criterion: you run a stress test specifically to observe the failure signal, so “the test produced errors” is the expected outcome, not a reason to stop early.
Example assumption: this series treats a “performance test” as the SLO-conformance check, “load test” as the sustained-peak variant, and keeps stress, spike, soak, capacity, and scalability as distinct profiles. Your organisation’s names may differ; the patterns and questions are what carry over.
Why high RPS alone is not success
A Gatling report’s headline number is throughput: requests per second the system completed. It is the least informative number on the page, for three reasons.
- Throughput is decided by the client, not the server. In an open-model test — the default mental model, and the right one for user-facing traffic — the arrival rate is what you configured. The server does not “achieve” 500 req/s; 500 req/s arrived, and the question is what happened to them.
- Completed requests can all be errors. A service that returns HTTP 500 in 2 ms will show spectacular throughput and latency. The assertions that matter are on response time percentiles and failure ratio, never on the count.
- RPS says nothing about tail latency. At high utilisation, queueing delay concentrates in the tail. Two runs can show identical mean latency and RPS while one has a p99 ten times worse — a difference that decides whether users in that tail abandon the checkout.
The success criterion for every test in this series is therefore an assertion bundle, not a throughput figure: latency percentiles per endpoint, a failed-request ratio, and — when relevant — a throughput floor. Chapter 06 shows these as Gatling assertions(...), so a run fails loudly instead of producing a report nobody reads.
The measurement vocabulary: SLI, SLO, SLA
Before you can say whether a test passed, you need the three terms that define “passed”:
- SLI (service level indicator) — the measured quantity. For this API: request latency by endpoint, successful-response ratio, and throughput. An SLI must be measurable from instrumentation, not vibes.
- SLO (service level objective) — the internal target on an SLI, over a window. “p99 latency on
GET /api/orders/{id}stays below 300 ms over any 5-minute window” is an SLO. SLOs are what your load tests assert. - SLA (service level agreement) — the contractual promise to a customer, typically looser than the internal SLO and attached to penalties. The gap between SLA and SLO is your error budget — room to be imperfect without breaching the contract.
Two more terms do quiet but essential work:
- Saturation — how full a resource is: CPU utilisation, HikariCP connections in use, Tomcat worker threads busy. Saturation is the cause you correlate latency against; chapter 08 is built around this.
- Availability — the fraction of time (or requests) the service is usable, e.g. “99.9% of requests get a non-5xx response”. In load-test terms it maps to the failed-request ratio.
Percentiles, and why the average lies
Latency distributions under load are right-skewed: most requests are fast, a few are slow, and the slow ones are very slow. Reporting the mean of a skewed distribution describes a request nobody experienced.
- p50 (median) — half of requests were faster. Use it for “typical” experience.
- p95 — the request one user in twenty experiences. This is usually the first place overload shows.
- p99 — one in a hundred. On an endpoint hit 500 times a second, p99 is exceeded five times every second — hundreds of real users per minute.
- max — dominated by outliers: a single GC pause, a TCP retransmit, a container throttle slice. Useful as a tripwire (“max must stay under 2 s”), useless as a target.
The numerical example that makes this concrete: a service where p50 is 20 ms and p99 is 1.4 s has a mean somewhere around 40 ms. “Average response time 40 ms” reads as healthy and describes nothing real. Assert on p95 and p99, keep max as a guardrail, and treat mean as a smoke signal only.
Arrival rate, concurrency, and the queue
Three quantities control every load test, and they are related by a constraint you cannot configure away.
Arrival rate is how fast new requests arrive: “200 orders per second”. Concurrency is how many requests are in flight at once. Response time is how long each takes. Little’s law — a principle, not an approximation — ties them together:
concurrency = arrival rate × response timeAt 200 req/s with a 100 ms response time, roughly 20 requests are in flight at any instant. Double the response time under load and concurrency doubles to hold the same arrival rate. This is why the relationship between load and latency is a knee, not a line: a system has finite servers (Tomcat worker threads, HikariCP connections, CPU cores). While arrival rate stays below service capacity, concurrency stays low and response time is flat. Once arrivals exceed what the servers can drain, requests queue, and every additional arrival adds queueing delay to everyone behind it.
A request’s total latency is the sum of time spent waiting at each stage plus the service time: accept queue → worker thread → connection pool → database → serialisation. Server-side latency is the portion inside the application and its dependencies; client-side latency — what Gatling measures — adds network transit and its own queueing. Chapters 07 and 08 keep these two strictly separate, because confusing them is how teams conclude “the API is slow” when the load generator’s own thread pool was saturated.
Principle: above the knee, latency is dominated by queueing, not by the work itself. That is why “add resources” is sometimes the fix and sometimes useless — if the queue is behind the database connection pool, doubling CPU does nothing. Diagnosing which queue filled is what chapter 09 is for.
Acceptance criteria for the Order Management API
Example assumption: the following targets are invented for the running example — plausible for a mid-traffic internal API, not derived from a real business requirement. Replace every number before reusing them.
| Endpoint / property | SLO |
|---|---|
GET /api/orders/{id} | p99 < 150 ms, p50 < 25 ms |
GET /api/orders?customerId=&status= | p99 < 300 ms over any 5-minute window |
POST /api/orders | p99 < 400 ms, p50 < 60 ms |
PATCH /api/orders/{id}/status | p99 < 250 ms |
| Error rate (non-2xx, excluding 404s on bad IDs) | < 0.1% of requests |
| Sustained target | the above hold at 150 req/s mixed traffic for 30 minutes |
| Soak | no upward trend in heap, GC pause time, or pool utilisation over 4 hours at 60% of target |
Three properties make these testable rather than aspirational: every criterion names a metric, a threshold, and a window; the traffic target is attached to the latency targets, not separate; and the soak criterion is about trends, not values.
What chapter 02 builds
With the vocabulary settled, the next chapter creates the lab: a repository layout, a Docker Compose file for PostgreSQL, Prometheus, and Grafana, and — most important — the list of variables that must be controlled before any two test runs can be compared.