Series overview
Part 17 of 17100% complete
2026-09-10•5 min read

Production hardening and the final exercise: prove it, then ship it

Checkpoint tag: v1.0.0 — production-readiness checklist complete, known limitations documented, final scenario green end to end.

What will be built

The last chapter is a review, not a feature: the architecture re-inspected against the running system, the threat model updated from skeleton to as-built, four failure drills executed and recorded, secret rotation and data retention answered concretely, the upgrade strategy written down, and the final incident scenario — the one from Chapter 0 — run end to end as the acceptance test for the whole series.

Why it matters

Every prior chapter built a mechanism; this chapter answers “does the mechanism survive contact.” Readiness reviews are where the gap between implemented and operable shows up: a runbook nobody rehearsed, a dashboard nobody loads during the drill, a rotation procedure that requires downtime nobody noticed. You ship what you rehearsed.

The architecture review — as built

Walk the original boundaries and check they survived implementation:

Boundary (Ch 0/1 promise)As-built enforcement
domain-contracts is framework-lightArchUnit rule, still green; zero Spring AI imports
Model behind a portModelGateway; stub/Ollama/hosted interchangeable
Retrieval tenant-safeWHERE tenant_id + supersede filter; eval case sec-001 guards it
Tool calls policy-gatedToolPolicyRegistry in dispatch, not in prompts
Writes approval-gatedagent.approvals + arg-hash + single-use transition
Audit separated from logsagent.audit_events append-only + AUDIT marker
Telemetry redactedTelemetryRedactionTest + collector transform

Any row you can’t point at in code is documentation debt — fix the code or fix the doc, never leave a phantom control.

Threat model — as-built update

Review the Chapter 0 table against shipped reality:

  • Closed: cross-tenant retrieval/tool access (claim-vs-arg enforcement + eval cases), approval replay/TOCTOU (arg hash + conditional update), unbounded orchestration (three budget dimensions), telemetry leakage (double redaction).
  • Residual, accepted: indirect injection can still produce a wrong answer — mitigated by citations, abstention, and the fact that answers can’t mutate anything; local-dev profiles relax controls (labeled); claims-based identity propagation trusts agent-api — token exchange remains an optional lab.
  • Documented non-goals: multi-agent delegation, internet-facing deployment, real ITSM integration — the simulator stands in; the seams for the real thing are the contract types and SimulatorClient.

Operational drills

Run each, record output in docs/runbooks/drill-results.md:

  1. Dependency death: kill Ollama → MODEL_UNAVAILABLE fast-fails; trace shows the failure boundary; recovery on restart is clean (no stuck loops).
  2. IdP outage: Keycloak down → 401s at both boundaries; deny metrics spike; recovery requires no code.
  3. Bad deploy rollback: deploy a deliberately broken agent-api image tag → readiness fails → rollout halts (maxUnavailable: 0 did its job) → rollout undo.
  4. Secret rotation: rotate AGENT_MCP_CLIENT_SECRET in Keycloak, update Secret, rolling restart → zero failed auth on the old secret before expiry — because you rotate the Keycloak credential first and the Pod env second; the order is the lesson.

Retention, cost, and upgrade strategy

  • Retention: agent.messages and audit_events get stated policies (90d operational, per-org audit) with a purge job sketched; runbook chunks are self-cleaning via supersede.
  • Cost: token usage metrics from Chapter 6/12 roll up into a per-conversation estimate; local-debug prompt capture stays file-only.
  • Upgrades: Spring AI and the MCP SDK move fast — the upgrade procedure is bump catalog → ./gradlew check → ./gradlew :evaluation-suite:evaluate → drill 1. The eval suite is the upgrade gate; that is why it exists.

The final exercise

One scenario, every mechanism:

priya (operator, acme) asks: “payment-gateway is returning 5xx since the last deploy — assess it, and open an incident.”

Expected sequence — each step verifiable in telemetry:

  1. Retrieval finds rb-payment-gateway-degraded chunks (acme-scoped, current version).
  2. The loop calls get_service_status + list_recent_incidents under ops:read — ToolProposed events on SSE.
  3. IncidentAssessment returns schema-valid, evidence refs all real chunk IDs.
  4. create_incident proposed → policy: HIGH risk, ops:incident:write present, approval required → ApprovalRequired event with args JSON and 5-min expiry.
  5. priya approves → arg-hash verified → idempotent execution → incident inc-… created.
  6. One Tempo trace contains the whole path; agent.audit_events holds APPROVAL_REQUESTED, APPROVED, TOOL_EXECUTED with matching hashes.
  7. The same question from sam dies at step 4 with INSUFFICIENT_SCOPE — and from rin at retrieval, where globex can’t see acme’s runbook.

If all seven hold on a clean checkout, the series delivered what the blueprint promised.

Known limitations — v1.0.0

Single-node event bus and approval queue; claims-based identity propagation (not RFC 8693); simulator instead of real ITSM; eval judge coverage is thin by design; Compose is not HA; the K8s overlay assumes ingress+storage; model quality depends on the deployed model, which the eval suite measures but cannot guarantee.

Checkpoint verification checklist

  • All four drills executed and recorded.
  • Threat-model residual risks written down, not vibes.
  • Final scenario green, including the two negative paths.
  • git tag v1.0.0 on a tree where ./gradlew clean check + evaluate + gatlingRun all pass.

What comes next

The optional labs — Java 27 structured concurrency, A2A, Kotlin ADK, reranking, GraalVM, local model benchmarks — each isolated, each safe to skip. What you have now is the thing most “AI tutorials” never produce: a system whose limits you can name because you built the boundaries yourself.

Project State Ledger — v1.0.0

  • Complete: 9 modules, 5 MCP tools, 4 scopes, 2 tenants, full observability + eval + perf + deploy
  • Verification suite: check + evaluate + gatlingRun + drill script
  • Series complete. Carry this ledger into any maintenance or extension session.
Spring BootKotlinAIArchitecture

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind