Test performance and resilience: measuring the framework, not the model
Checkpoint tag: chapter-14-performance — performance-tests runs reproducible Gatling scenarios, every reported number comes from a documented local run, and initial SLOs are drafted from measurements.
What will be built
A performance-tests module with Gatling scenarios for the REST message path, SSE concurrency, and read-tool throughput; a stub-model latency mode so application overhead is measurable without a real model; fault runs (latency, flaky, outage) quantifying how degradation propagates; a JVM-level look at virtual threads, pool sizing, and JFR; and docs/perf/results-template.md so every number in the repo is reproducible.
Why it matters
The single most common lie in AI-benchmark content is reporting “agent latency” that is 95% hosted-model time measured over hotel Wi-Fi. This chapter’s core discipline: measure the application’s overhead against a deterministic stub; measure model latency separately, labeled, and never conflate them. The stub isn’t a toy here — it’s the control that makes the framework’s cost visible.
Concepts explained
Gatling over k6 (ADR-009): JVM-native DSL, same toolchain, feeds results straight into the same observability stack. The scenarios are small on purpose — this is a smoke envelope, not a capacity plan.
Virtual threads under blocking MCP calls. The MCP client is synchronous by design; on virtual threads, blocking is cheap. The risk isn’t blocking — it’s pinning (synchronized blocks holding a carrier thread) and unbounded downstream concurrency. JFR’s jdk.VirtualThreadPinned events are how you check the first; the Chapter 11 bulkhead is the second.
Pool arithmetic. Tomcat threads, the HTTP client’s connection pool to MCP, the MCP server’s pool to the simulator, the JDBC pool — the slowest link’s concurrency is the system’s real limit. The scenario matrix exists to find that link rather than assume it.
Scenario matrix
| Scenario | Target | Assertion |
|---|---|---|
messageStubLoad | 50 rps for 60s, MODEL_PROVIDER=stub | p95 app overhead within your measured budget; errors < 1% |
sseFanout | 100 concurrent SSE subscribers | no emitter leaks (memory flat after close) |
readToolThroughput | get_service_status through the loop | per-tool p95 recorded; MCP→simulator hop cost isolated |
degradedTool | same + simulator latency 2000 | tool timeout fires, loop completes or degrades, error shape correct |
bulkheadSaturation | 20 concurrent vs. permit=8 | fast SATURATED, no queue blowup |
realModelSmoke (opt-in) | Ollama, 5 rps, 60s | model latency reported separately |
The Gatling simulation (Java DSL) for the headline scenario:
package in.o612.eng.opsagent.perf;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;
import java.time.Duration;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
public class MessageStubSimulation extends Simulation {
HttpProtocolBuilder http = http.baseUrl(System.getenv().getOrDefault( "AGENT_URL", "http://localhost:8080")) .header("Authorization", "Bearer " + System.getenv("PERF_TOKEN"));
ScenarioBuilder scn = scenario("message-stub") .exec(http("create").post("/api/v1/conversations") .check(jsonPath("$.conversationId").saveAs("convId"))) .exec(http("message").post("/api/v1/conversations/#{convId}/messages") .body(StringBody("{\"message\":\"status check\"}")) .check(status().is(200)));
{ setUp(scn.injectOpen(constantUsersPerSec(50).during(Duration.ofSeconds(60)))) .protocols(http) .assertions(global().responseTime().percentile(95.0).lt(2000), global().failedRequests().percent().lt(1.0)); }}Running and reporting
MODEL_PROVIDER=stub ./gradlew :agent-api:bootRun &./gradlew :performance-tests:gatlingRun# report -> performance-tests/build/reports/gatling/*/index.htmlEvery table in docs/perf/ records: machine (CPU/RAM), JVM flags, provider mode, dataset size (chunks in document_chunks — vector search latency means nothing without a corpus size), and the raw Gatling output path. If a number can’t name its run, it doesn’t ship.
JFR walkthrough: -XX:StartFlightRecording on agent-api during readToolThroughput; inspect jdk.VirtualThreadPinned (expect none — the path is lock-free by construction; find some and you found a bug worth a chapter footnote), jdk.SocketRead for MCP wait time, and allocation rate under the bulkhead.
SLO seeds from measurements
Drafted after the runs, not before: e.g., “p95 framework overhead per message < X ms at 50 rps on this hardware; tool-call hop adds < Y ms” — X and Y filled in by your run, kept honest by the template.
Failure-injection lab
degradedTool is the lab: compare readToolThroughput baseline vs. latency 2000 vs. flaky 0.5. The interesting observation is retry amplification — flaky at the simulator doubles call volume in the client (2 attempts); at 50 rps input that’s 100 rps downstream. Do the math before enabling retries anywhere bigger.
Checkpoint verification checklist
-
gatlingRungreen on the stub profile with assertions enabled. - Real-model numbers labeled separately — no conflation in the report.
- Bulkhead saturation tested, not assumed.
-
results-template.mdfilled for at least one run.
Commit message and Git tag
test(perf): Gatling scenarios, stub-model baselines, fault-mode runs, results templategit tag chapter-14-performance
What comes next
Chapter 15 packages everything — container images, the full Compose stack, and Kubernetes manifests that reflect what the previous fourteen chapters actually built.
Project State Ledger — chapter-14-performance
- Module:
performance-tests(Gatling Java DSL); scenarios: stub load, SSE fanout, tool throughput, degraded, saturation, opt-in real-model - Rules: stub isolates framework cost; model latency reported separately; every number cites its run
- Diagnostics: JFR for VT pinning + socket waits; pool arithmetic documented
- Next:
chapter-15-deployment