Series overview
Part 6 of 1346% complete
2026-07-29•6 min read

Gatling simulations

By the end of this chapter Gatling is wired into the Gradle build and you have four complete simulations covering the four workloads this series needs: read-heavy, mixed read/write, paginated search, and a stepped stress test that pushes past the SLOs from chapter 01. Along the way: feeders, session state, request correlation, the open-versus-closed model decision, and assertions that make a run fail loudly.

Add the Gatling plugin

Add one line to the plugins block of build.gradle.kts (chapter 03):

build.gradle.kts
plugins {
java
id("org.springframework.boot") version "4.1.1"
id("io.spring.dependency-management") version "1.1.7"
id("io.gatling.gradle") version "3.15.1.3" // brings Gatling 3.15.1
}

The plugin adds a gatling source set — simulations live in src/gatling/java, test resources (feeder CSVs) in src/gatling/resources — and a gatlingRun task. With several simulation classes present, pick one by fully-qualified name (gatlingRun alone is interactive, and fails outright when CI=true):

Terminal window
./gradlew gatlingClasses # compile check only
./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation
./gradlew gatlingRun --all # every simulation, sequentially

Gatling does not need the application as a dependency — it is a pure HTTP client. That separation is deliberate: the load generator must never share a JVM, a connection pool, or a class path with the system under test.

Open model, closed model — pick deliberately

Gatling’s injection API splits in two, and the choice determines what your numbers mean:

  • Open model (injectOpen): you control the arrival rate — new virtual users start at the configured rate regardless of whether earlier requests finished. If the server slows down, in-flight requests pile up, which is exactly what real traffic does.
  • Closed model (injectClosed): you control concurrency — a fixed number of virtual users, each starting a new iteration only after finishing the last. If the server slows, the arrival rate drops automatically.

Principle: for a user-facing HTTP API, the open model is the honest one. Real visitors do not politely wait for earlier visitors to finish. The closed model hides saturation: as response times grow, the offered load silently shrinks to what the server can handle — this is coordinated omission, covered fully in chapter 07, and it is why a closed-model test can report “stable throughput” while every user is queued. Use closed models only when the real callers genuinely are bounded — a fixed fleet of batch workers, a thread pool of internal consumers — which is rare for REST APIs.

All four simulations below use the open model.

A note on feeders, headers, and auth

  • Feeders inject per-request data into the virtual user’s session — the CSV exports from chapter 05. circular() re-reads from the top when exhausted (right for read IDs); queue() consumes each row once and fails the run if exhausted (right for placed_orders.csv — a PATCHed order is consumed, and reusing it would generate 409s).
  • Session attributes are referenced in URLs and bodies with #{name} — the CSV column name.
  • Headers belong on HttpProtocolBuilder when they are universal (Accept, Content-Type), on the request otherwise.
  • Authentication: this lab API is unauthenticated. When yours is not, the pattern is a one-shot exec that posts credentials and .check(jsonPath("$.token").saveAs("accessToken")), followed by .header("Authorization", "Bearer #{accessToken}") on each request — fetch tokens in a before hook or once per user, never inside a measured request chain unless token issuance is itself in scope:
src/gatling/java/in/o612/eng/orders/load/AuthPattern.java
// Illustrative — not used by the lab API. Pattern for token-bearing APIs:
// .exec(http("authenticate").post("/auth/token")
// .body(StringBody("{\"client\":\"load-test\"}"))
// .check(status().is(200), jsonPath("$.token").saveAs("accessToken")))
// then on each request:
// .header("Authorization", "Bearer #{accessToken}")

Store real credentials in environment variables (System.getenv("LOADTEST_TOKEN")), never in the repository — a simulation with a committed token is a credential leak.

Simulation 1 — read-heavy

The dominant real-world shape: mostly fetches, some searches, no writes. Traffic distribution is 70% GET /api/orders/{id}, 30% search — deliberately unequal, because equal-weight endpoints are a lab fiction.

src/gatling/java/in/o612/eng/orders/load/ReadHeavySimulation.java
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import java.time.Duration;
public class ReadHeavySimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http
.baseUrl(System.getProperty("baseUrl", "http://localhost:8080"))
.acceptHeader("application/json")
.shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular();
FeederBuilder<String> customers = csv("data/customers.csv").circular();
ChainBuilder getOrder = feed(orders)
.exec(http("GET /api/orders/{id}")
.get("/api/orders/#{order_id}")
.check(status().is(200)));
ChainBuilder searchOrders = feed(customers)
.exec(http("GET /api/orders?customer")
.get("/api/orders?customerId=#{customer_id}&size=20")
.check(status().is(200)));
ScenarioBuilder readHeavy = scenario("read-heavy")
.randomSwitch().on(
percent(70.0).then(getOrder),
percent(30.0).then(searchOrders));
{
setUp(readHeavy.injectOpen(
// phase 1: warm-up — arrival rate ramps, results discarded
rampUsersPerSec(1).to(30).during(Duration.ofMinutes(1)),
// phase 2: steady state — the measurement window
constantUsersPerSec(30).during(Duration.ofMinutes(3)),
// phase 3: ramp-down — watches recovery, kept out of assertions
rampUsersPerSec(30).to(0).during(Duration.ofSeconds(30))
).protocols(httpProtocol))
.assertions(
details("GET /api/orders/{id}").responseTime().percentile(99.0).lt(150),
details("GET /api/orders?customer").responseTime().percentile(99.0).lt(300),
global().failedRequests().percent().lt(0.1)
);
}
}

shareConnections() makes virtual users share the underlying HTTP connection pool like real browsers and API clients behind keep-alive — without it, each virtual user gets a private connection and you benchmark TCP setup instead of your API. The assertions are the SLOs from chapter 01, expressed per request name: this run fails (non-zero exit, KO in the report) if p99 on fetches exceeds 150 ms.

Simulation 2 — mixed read/write

Adds create and status-update traffic in a realistic ratio. The chain that matters is createThenRead: the POST response’s id is saved into the session and immediately used in a follow-up GET — request correlation, the pattern for “create then fetch the thing you created”.

src/gatling/java/in/o612/eng/orders/load/MixedWorkloadSimulation.java
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import java.time.Duration;
public class MixedWorkloadSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http
.baseUrl(System.getProperty("baseUrl", "http://localhost:8080"))
.acceptHeader("application/json")
.contentTypeHeader("application/json")
.shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular();
FeederBuilder<String> customers = csv("data/customers.csv").circular();
FeederBuilder<String> placed = csv("data/placed_orders.csv").queue();
ChainBuilder getOrder = feed(orders)
.exec(http("GET /api/orders/{id}")
.get("/api/orders/#{order_id}").check(status().is(200)));
ChainBuilder search = feed(customers)
.exec(http("GET /api/orders?customer")
.get("/api/orders?customerId=#{customer_id}&size=20")
.check(status().is(200)));
ChainBuilder createThenRead = feed(customers)
.exec(http("POST /api/orders")
.post("/api/orders")
.body(StringBody("""
{"customerId":#{customer_id},
"items":[{"sku":"SKU-LT-1","quantity":2,"unitPrice":19.99},
{"sku":"SKU-LT-7","quantity":1,"unitPrice":4.50}]}
""")).asJson()
.check(status().is(201), jsonPath("$.id").saveAs("createdOrderId")))
.pause(Duration.ofMillis(300))
.exec(http("GET created order")
.get("/api/orders/#{createdOrderId}").check(status().is(200)));
ChainBuilder advanceStatus = feed(placed)
.exec(http("PATCH /api/orders/{id}/status")
.patch("/api/orders/#{order_id}/status")
.body(StringBody("{\"status\":\"PAID\"}")).asJson()
.check(status().is(200)));
ScenarioBuilder mixed = scenario("mixed").randomSwitch().on(
percent(55.0).then(getOrder),
percent(25.0).then(search),
percent(12.0).then(createThenRead),
percent(8.0).then(advanceStatus));
{
setUp(mixed.injectOpen(
rampUsersPerSec(1).to(20).during(Duration.ofMinutes(1)),
constantUsersPerSec(20).during(Duration.ofMinutes(3))
).protocols(httpProtocol))
.assertions(
global().responseTime().percentile(95.0).lt(250),
global().failedRequests().percent().lt(0.5)
);
}
}

Two details are load-bearing. First, placed uses queue() — each PLACED order can be transitioned to PAID exactly once, and if the feeder exhausts, the run fails rather than silently generating 409s. 12% of 20 users/s is 2.4 creations per second; the status-update pool (~120,000 PLACED orders) is far deeper than any single run consumes, which is why chapter 05’s template-reset strategy exists for repeated write runs. Second, every write is checked — a silent 500 on 8% of traffic would contaminate every latency percentile while looking like “the server held up”.

Why these weights? Example assumption: reads dominate e-commerce order traffic by roughly 5:1 in this fictional system. Your real ratio comes from production access logs — pull a day of traffic, count by route shape, and weight randomSwitch accordingly. The method is the deliverable, not the numbers.

Search with pagination is where deep OFFSET and index quality show up. Each virtual user pages through three pages of results for one customer.

src/gatling/java/in/o612/eng/orders/load/SearchSimulation.java
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import java.time.Duration;
public class SearchSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http
.baseUrl(System.getProperty("baseUrl", "http://localhost:8080"))
.acceptHeader("application/json")
.shareConnections();
FeederBuilder<String> customers = csv("data/customers.csv").circular();
ScenarioBuilder search = scenario("paginated search")
.feed(customers)
.exec(session -> session.set("page", 0))
.repeat(3).on(
exec(http("GET /api/orders search page #{page}")
.get(s -> "/api/orders?customerId=" + s.getString("customer_id")
+ "&page=" + s.getInt("page") + "&size=20")
.check(status().is(200)))
.pause(Duration.ofMillis(800))
.exec(session -> session.set("page", session.getInt("page") + 1))
);
{
setUp(search.injectOpen(
rampUsersPerSec(1).to(15).during(Duration.ofMinutes(1)),
constantUsersPerSec(15).during(Duration.ofMinutes(3))
).protocols(httpProtocol))
.assertions(
global().responseTime().percentile(99.0).lt(300),
global().failedRequests().percent().lt(0.1)
);
}
}

exec(session -> ...) is the escape hatch for session arithmetic — here, a page counter incremented per iteration. pause(800ms) models a human reading a page; in the open model it does not reduce the arrival rate, it only changes how long each virtual user stays active — which raises the concurrent-in-flight count by Little’s law. Both effects are intended.

Simulation 4 — stepped stress

This one’s job is to break the SLO on purpose, in controlled steps, so you can see where the knee is and what saturates first.

src/gatling/java/in/o612/eng/orders/load/StressSimulation.java
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import java.time.Duration;
public class StressSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http
.baseUrl(System.getProperty("baseUrl", "http://localhost:8080"))
.acceptHeader("application/json")
.contentTypeHeader("application/json")
.shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular();
FeederBuilder<String> customers = csv("data/customers.csv").circular();
ScenarioBuilder probe = scenario("stepped stress").randomSwitch().on(
percent(60.0).then(feed(orders).exec(
http("GET /api/orders/{id}")
.get("/api/orders/#{order_id}").check(status().is(200)))),
percent(30.0).then(feed(customers).exec(
http("GET /api/orders?customer")
.get("/api/orders?customerId=#{customer_id}&size=20")
.check(status().is(200)))),
percent(10.0).then(feed(customers).exec(
http("POST /api/orders")
.post("/api/orders")
.body(StringBody("""
{"customerId":#{customer_id},
"items":[{"sku":"SKU-ST","quantity":1,"unitPrice":9.99}]}
""")).asJson()
.check(status().is(201)))));
{
setUp(probe.injectOpen(
// 25 → 150 req/s in six 2-minute steps, 30 s ramps between.
// Widen the range until the SLO assertions fail; that level
// IS the finding.
incrementUsersPerSec(25.0)
.times(6)
.eachLevelLasting(Duration.ofMinutes(2))
.separatedByRampsLasting(Duration.ofSeconds(30))
.startingFrom(25.0)
).protocols(httpProtocol))
.assertions(
global().responseTime().percentile(99.0).lt(300),
global().failedRequests().percent().lt(0.5)
);
}
}

The assertions here are not pass/fail gates — they are the tripwire that tells you which step crossed the line. Read the report’s “response time over time” graph against the injection profile: the step where p99 bends sharply upward is the knee, and the level just below it is the service’s honest capacity under this profile. Chapter 08 shows how to correlate that knee with the saturation metrics from chapter 03.

Injection profiles, briefly

DSLShapeUse for
rampUsersPerSec(a).to(b)arrival rate linear a→bwarm-up, ramp-down
constantUsersPerSec(r).during(t)flat arrival ratesteady-state measurement
incrementUsersPerSec(step).times(n).eachLevelLasting(t)staircasestress / capacity tests
nothingFor(t)silencecool-down
constantConcurrentUsers(n) / rampConcurrentUsersfixed concurrencyclosed model only — bounded caller pools
atOnceUsers(n)all at oncespike tests (chapter 11)

Common failure: forgetting .protocols(httpProtocol) — the scenario runs against no HTTP config and fails cryptically. And naming requests ("GET /api/orders/{id}") is not cosmetic: the report and the details(...) assertions are keyed by those names, so use the route template, not a friendly sentence.

Milestone

Terminal window
./gradlew gatlingClasses # compiles
./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation

A full run takes ~4.5 minutes including warm-up and ramp-down. The console ends with a path to an HTML report under build/reports/gatling/. Open it — chapter 07 is about making sure the numbers in it are trustworthy.

Spring BootJavaPerformanceTesting

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind