Stress, spike, and soak testing
By the end of this chapter you have dedicated profiles for the three tests that probe the edges of the system — stress (where is the knee?), spike (does it survive a burst?), soak (does it survive Tuesday?) — plus a definition of graceful overload behaviour and a post-test recovery checklist.
Safety first, once: these tests are designed to push a system past its limits. Run them against the lab or a dedicated environment you own; never against shared staging without coordination, and never against production without explicit authorization, a rollback plan, and a rate limit you can pull. A stress test is a controlled failure — controlled is the operative word.
How good overload behaves
Before hunting for breaking points, define what “broke well” means. Under load beyond capacity, a healthy service:
- Fails fast rather than slowly. A request that cannot get a connection in 30 s has already destroyed its user’s patience; HikariCP’s
connection-timeoutis a deliberate bound — the queue must end somewhere. Overload that returns a 503 in 50 ms beats overload that returns a 200 in 60 s. - Rate-limits deliberately. A 429 (or a dropped request at the gateway) sheds load at the edge instead of queueing it to collapse the middle. If your API lacks one, “unbounded queue under overload” is a finding.
- Applies backpressure. Queues of work — Tomcat’s accept queue, HikariCP’s pending list — are finite by configuration, so excess demand is rejected at the boundary rather than absorbed until the JVM dies.
- Times out downstream calls so a slow dependency cannot hold request threads hostage.
- Recovers. When load drops to normal, latency and error rate return to baseline within a bounded time — seconds to a couple of minutes. A service that stays degraded after load is removed is broken in a way the load test just proved.
Each of these is observable in the metrics from chapter 03: hikaricp_connections_timeout_total, tomcat_threads_busy versus max, response-code histograms, and the shape of the recovery after ramp-down.
Profile 1 — stress: find the knee
StressSimulation from chapter 06 already implements the shape — incrementUsersPerSec(25).times(6).eachLevelLasting(2m).separatedByRampsLasting(30s) — a staircase of increasing arrival rates. What this chapter adds is the reading protocol:
- Identify the last step where all SLO assertions held. That is measured capacity under this profile and environment — both qualifiers are part of the finding.
- Identify the first step where they broke, and which assertion broke first — latency or errors. The first casualty is almost always the bottleneck’s name: p99 bending while errors are still flat is queueing; errors first is a hard limit (pool timeout, connection refusal).
- Watch the knee step in Grafana: which saturation metric reached its ceiling at or before the bend. That is the component to work on first; everything after it is symptom.
- Confirm recovery: after the staircase ends, does p99 return to warm-up levels? Non-recovery is a separate finding — leaks, stuck threads, a brokered queue that never drains — and chapter 08’s table covers it.
To find the knee when the initial range doesn’t reach it, widen times()/startingFrom — the test exists to bracket the breaking point, not to confirm a guess about where it is.
Profile 2 — spike: survive the burst
A spike is not a bigger load test; it is a rate-of-change test. Autoscalers, caches, and connection pools all have reaction times, and a step change defeats all of them at once.
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;import java.time.Duration;
public class SpikeSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http .baseUrl(System.getProperty("baseUrl", "http://localhost:8080")) .acceptHeader("application/json") .shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular();
ScenarioBuilder spike = scenario("spike") .feed(orders) .exec(http("GET /api/orders/{id}") .get("/api/orders/#{order_id}").check(status().is(200)));
{ setUp(spike.injectOpen( constantUsersPerSec(20).during(Duration.ofMinutes(2)), // baseline rampUsersPerSec(20).to(300).during(Duration.ofSeconds(20)), // the spike constantUsersPerSec(300).during(Duration.ofMinutes(2)), // held burst rampUsersPerSec(300).to(20).during(Duration.ofSeconds(30)), constantUsersPerSec(20).during(Duration.ofMinutes(2)) // recovery window ).protocols(httpProtocol)) .assertions( // a spike may legitimately degrade — assert on the recovery, // not on the burst itself global().failedRequests().percent().lt(2.0) ); }}The questions a spike answers are different from a stress test’s: did errors cluster only in the transition (acceptable — the system shed load) or persist through the held burst (the system genuinely can’t do 300/s)? Did p99 settle back after the step down? A service that survives a spike is allowed to look ugly during it — degradation is fine, non-recovery is not.
Profile 3 — soak: find what accumulates
Soak tests exist for the failures with accumulating state: slow leaks invisible at any timescale shorter than their fill time. Four hours is the lab minimum; a weekend is the honest version. The load is deliberately modest — 60% of the measured target — because the point is duration, not pressure.
// inside a SoakSimulation — same chains as ReadHeavy, different injectionsetUp(scn.injectOpen( rampUsersPerSec(1).to(20).during(Duration.ofMinutes(5)), constantUsersPerSec(20).during(Duration.ofHours(4))).protocols(httpProtocol));What you are watching over those four hours — and what each trend means:
| Metric trend | Suspect |
|---|---|
jvm_memory_used_bytes{area="heap"} post-GC floor climbs steadily | Heap leak — objects retained: cache without bound, static collection, listener accumulation |
hikaricp_connections total or active ratchets up and never returns | Connection leak — a path checks out a connection without returning it (or a transaction that never closes) |
hikaricp_connections_pending grows over hours | Connection pool draining, not leaking — usually cumulative slow queries |
jvm_gc_pause_seconds frequency rises with flat allocation rate | Heap fragmentation pressure / shrinking effective heap — same fix path as a leak |
process_cpu_usage drifts up at flat load | Growing per-request work — caches filling, metrics cardinality, a growing in-memory structure |
Confirmation when a trend appears: take a heap dump (temporarily expose /actuator/heapdump, chapter 09’s security note applies), or watch whether down -v + fresh app clears the slope — a leak that survives restart is in the database or a queue, not the JVM.
The difference between soak findings and stress findings: stress finds limits, soak finds slopes. A pool that is 3% fuller each hour looks healthy for the first 90 minutes; a soak run is the only test patient enough to see it.
Finding the breaking point without breaking anything else
- Step, don’t slam.
incrementUsersPerSecexists precisely so the first failing step is visible; a test that goes straight to 10× tells you nothing about where the knee was. - Kill criteria in advance. Decide before the run what makes you stop it: error rate > 25% sustained, or the database host’s CPU pegged for 5 minutes — write it on the run sheet.
- Isolate the blast radius. Dedicated environment, dedicated database (the
orders_testtemplate from chapter 05), no shared state with anything a colleague is using. - Rate limits are the exit ramp. The API’s own limits — or a gateway’s — should be what sheds load once you are past the knee; if nothing sheds, you are testing the container’s OOM-killer.
Post-test recovery checklist
After any boundary test, verify the environment returned to baseline — the next run’s validity depends on it:
- Gatling fully stopped;
up{job="order-api"}still UP. - Error rate back to ~0 and p99 back to warm-up levels at idle-rate traffic.
-
hikaricp_connections_activereturned to idle levels;pending= 0. -
jvm_memory_used_bytespost-GC floor back near its pre-run level. -
pg_stat_activityshows no stuck sessions, lock waits, or abandoned transactions. - No OOM-restart events (
process_uptime_secondsnot reset). - Database reset performed if write traffic ran (template swap or reseed — chapter 05).
- Run sheet filed with knee level, first-failing assertion, and recovery time.
Milestone check
You can now name: the load level where the Order API’s SLO first breaks, which resource queued first, how it behaves while overloaded (sheds, queues, or collapses), and whether it fully recovers. Chapter 12 moves this same discipline into CI — where tests run on every change — and into Kubernetes, where CPU limits change what latency numbers mean.