Reading the results
By the end of this chapter you can open a Gatling report, decide in under a minute whether the run is worth analysing, and walk from a bent latency curve to a ranked list of suspects using Prometheus. The output artifact is a troubleshooting table and a decision tree you will reuse in chapters 09–11.
Anatomy of a Gatling report
Each run writes a self-contained HTML report under build/reports/gatling/<simulation>-<timestamp>/index.html. In reading order:
- Assertions banner —
OK/KOagainst theassertions(...)you set. A KO is a finding, not a build artifact to shrug at: it names the exact request group that crossed its threshold. - Global statistics table — per request name: count, min, max, mean, stdDev, percentiles, req/s, and error percentage. Read the error percentage first — a slow run is diagnosable, a failed run needs the failure understood before latency means anything.
- Response time percentiles over time — the single most useful chart. A flat line is health; a knee is capacity; a sawtooth correlated with GC is a JVM story; a staircase that tracks
incrementUsersPerSecsteps tells you exactly which load level broke the SLO. - Response time distribution — the histogram shape. A tight low-latency mode plus a long tail is normal; a second hump at high latency means two populations of requests — usually “the ones that hit the cache/index” and “the ones that did not”.
- Active users along the run — in the open model this climbs during queueing (users arrive faster than they drain). A growing active-user curve is queued load rendered visible.
- Requests/responses per second — achieved throughput. Compare the offered rate (your injection profile) to the completed rate: divergence after some timestamp is the moment the service stopped keeping up.
Reading order that works: assertions → errors → percentile-over-time → the timestamp where it bent → that same timestamp in Grafana.
The PromQL query set
These are the queries behind the dashboard checklist from chapter 03. All assume the application="order-api" label from chapter 02’s scrape config.
Latency percentiles per endpoint (the server-side counterpart of Gatling’s client-side numbers):
histogram_quantile(0.99, sum by (le, uri) ( rate(http_server_requests_seconds_bucket{application="order-api"}[1m])))Error ratio (5xx share of all requests):
sum(rate(http_server_requests_seconds_count{application="order-api", status=~"5.."}[1m]))/sum(rate(http_server_requests_seconds_count{application="order-api"}[1m]))Throughput as the server saw it:
sum by (uri) (rate(http_server_requests_seconds_count{application="order-api"}[1m]))Heap utilisation and trend:
sum(jvm_memory_used_bytes{application="order-api", area="heap"})/sum(jvm_memory_max_bytes{application="order-api", area="heap"})GC pause pressure — fraction of the last minute spent stopped, plus frequency:
sum(rate(jvm_gc_pause_seconds_sum{application="order-api"}[1m])) -- overhead ratiosum(rate(jvm_gc_pause_seconds_count{application="order-api"}[1m])) -- pauses per secondAllocation rate (what drives GC):
rate(jvm_gc_memory_allocated_bytes_total{application="order-api"}[1m])HikariCP pool utilisation and queueing:
hikaricp_connections_active{application="order-api"}/ hikaricp_connections_max{application="order-api"} -- utilisation
hikaricp_connections_pending{application="order-api"} -- queued threads; should be 0
histogram_quantile(0.99, sum by (le) (rate(hikaricp_connections_acquire_seconds_bucket{application="order-api"}[1m]))) -- wait time for a connectionTomcat worker saturation:
tomcat_threads_busy_threads{application="order-api"}/ tomcat_threads_config_max_threads{application="order-api"}Process CPU:
process_cpu_usage{application="order-api"} -- 1.0 ≈ one core saturatedTwo PromQL notes that bite in practice: always rate the _bucket/_count/_sum series, never the raw counters, and pick a window ([1m]) shorter than your load-test phases so the ramp-up does not smear into the steady state. If a query returns nothing, check up{job="order-api"} first — the target may be DOWN, and Grafana shows silence, not an error.
The correlation method
Numbers become a diagnosis when two views of the same second disagree or agree:
- In the Gatling report, find the timestamp where p99 first bends (or errors begin).
- In Grafana, pin the same timestamp. Which saturation metric moved first — CPU, Tomcat threads, HikariCP pending, GC?
- The metric that moved at or just before the bend is the primary suspect; metrics that moved after are symptoms of the pile-up behind it.
- Form one hypothesis (“connection pool exhausted at 40 req/s”), one experiment (chapter 09/10), one re-run.
The pairs that matter most:
| Gatling shows | Grafana shows | Reading |
|---|---|---|
| p99 climbs | hikaricp_connections_pending > 0 at the same time | DB-pool queue — requests wait for connections |
| p99 climbs | tomcat_threads_busy / max ≈ 1 | Tomcat queue — work is waiting for a worker thread |
| p99 climbs | process_cpu_usage at the granted-core ceiling | CPU-bound |
| p99 climbs | GC overhead rising into double digits | GC-bound — usually allocation rate, not “bad GC” |
| p99 climbs | all of the above flat | Look downstream — the database, or the client itself |
| p99 flat | errors rising | Failure path is fast — check status codes before latency |
| client p99 climbs | server p99 flat | Generator or network — the measurement, not the service |
Troubleshooting table
The working artifact of this chapter. “First safe experiment” means the cheapest change that would disprove the hypothesis — falsifiable, reversible.
| Symptom | Likely bottleneck | Evidence to collect | First safe experiment | Common incorrect conclusion |
|---|---|---|---|---|
p99 rises at a load level; process_cpu_usage at ceiling | CPU saturation | GC overhead, per-endpoint split, allocation rate | Halve per-request work (smaller page or payload) and re-run — if the knee moves proportionally, CPU confirmed | “Add more instances” — correct only if the bottleneck is per-instance CPU, wrong if the shared DB is already hot |
p99 rises; hikaricp_connections_pending climbs from 0 | Connection pool exhaustion | acquire_seconds p99, query count per request, pg_stat_statements top queries | Fix the query count first (N+1 check, chapter 10); only then resize the pool | “Raise maximum-pool-size” — moves the queue into PostgreSQL’s max_connections or worse, onto a CPU-bound DB |
| p99 rises; Tomcat busy threads pinned at max, CPU and DB calm | Worker-thread starvation — something inside requests blocks | Thread dump or threaddump actuator; look for BLOCKED/TIMED_WAITING on one frame | Find and bound the blocking call (timeout, async boundary); re-run | “Raise server.tomcat.threads.max” — adds queue depth, not throughput; the blocker still caps you |
| Long tail of very slow requests while p50 is fine | Periodic stall: GC pause, disk flush, lock | jvm_gc_pause max vs p99 gap, pg_stat_activity waits | Correlate the tail timestamps with GC events | “Average latency is fine, ship it” — the tail is the user experience at this percentile |
| Latency and errors climb over minutes, never recover during run | Resource leak: pool, memory, file handles | Heap floor after GC, hikaricp_connections total, open sockets | Soak test with heap dumps at intervals (chapter 11) | “Raise the heap/limits” — postpones the failure off the end of the test, not out of production |
| Everything slow from the first request, all saturation flat | Cold caches / wrong environment | Run sheet variables, first-minute vs steady-state split | Extend warm-up; re-run | “The code got slower” — you measured a cold JVM |
| Client p99 ≫ server p99, gap grows with rate | Load generator saturated | Generator CPU, its own event-loop metrics | Re-run from a bigger box or two generators | “API regression” — the most expensive wrong conclusion available; it sends everyone hunting a defect that does not exist |
| Error rate climbs while latency stays low | Fast failure path — timeouts, 429s, pool timeout_total | Response code histogram, hikaricp_connections_timeout_total | Read the actual status codes; a 500 in 3 ms is a different problem than a 30 s timeout | “Overloaded” — fast errors usually mean a limit was hit, not exceeded gradually |
The decision tree
The tree encodes one rule: queues are found by looking at what is waiting, not at what is busy. A system can be 100% busy on CPU and healthy; pending > 0 on a connection pool is never healthy under target load.
Milestone check
You can now take any Gatling report, locate the moment the SLO broke, and produce — from Prometheus alone — a one-sentence hypothesis with named evidence. Chapter 09 is the catalog those hypotheses come from: the recurring bottleneck shapes, each with its metric signature and the experiment that proves it.