What we are building: an enterprise operations agent
Checkpoint tag: chapter-00-architecture — this chapter adds documentation only; there is nothing to compile yet.
Assumptions for this chapter and the rest of the series: Java 25, Spring Boot 4.1.1, Spring AI 2.0.1, Kotlin 2.3.21, Gradle 9.7.1, PostgreSQL 17 with pgvector 0.8.6, Keycloak 26.7.x, and Docker Compose. Versions were verified against official release pages when this blueprint was written; each chapter re-states what it depends on. The fictional domain — an internal operations platform for two tenants, acme and globex — recurs across the series, so the boundaries drawn here are the seams the code cuts along later.
What will be built
By the last chapter you will have ops-agent-platform, a multi-module Gradle repository containing four deployable services and a verification suite. An ops engineer sends a question to agent-api; the agent retrieves relevant runbook sections with citations, optionally calls read-only MCP tools for live service status and incident data, proposes mutating tools only when policy allows, and executes writes only after a human approves a single-use token bound to the exact arguments. Every step lands in one distributed trace.
This chapter lays out the whole map first: what the system does, what it deliberately refuses to do, where the trust boundaries sit, and which versions and ADRs everything downstream depends on. Skipping this chapter to reach the code sooner is the usual reason agent projects end up as a chatbot with a YAML file glued on.
Why it matters
Most JVM-side AI material stops at “call the model, print the answer.” Production stops much later: after you have bounded the agent loop, decided who may invoke which tool, proven that a mutated argument cannot slip past an approval, and can show an auditor a trace for the incident ticket the agent filed at 03:00. The hard parts are the same engineering disciplines as any distributed system — isolation, authorization, idempotency, observability — applied to a component that is nondeterministic by design. This series treats the LLM as one unreliable dependency inside a normal Spring system, not as the system.
Prerequisites and starting point
There is no starting Git tag — this chapter creates the documentation skeleton. You need nothing installed. To follow the series as a whole you should be comfortable with Java, Spring Boot, REST, SQL, Docker, and Gradle; you do not need prior agent, MCP, or Kotlin experience.
Chatbot, RAG app, tool-using agent, MCP server, multi-agent system
These terms get used interchangeably in marketing copy; in an architecture document they are different systems with different failure modes.
- Chatbot: a model plus a prompt template. Stateless-ish text in, text out. Failure mode: it says something wrong confidently.
- RAG application: a chatbot plus retrieval. Answers are grounded in documents you control, which adds a new failure mode — retrieving the wrong document, or a document a tenant was never meant to see.
- Tool-using agent: a model that can propose structured calls against real systems. New failure modes: wrong tool, wrong arguments, unbounded loops, and — the dangerous one — a correct-looking call nobody authorized.
- MCP server: not an agent at all. It is a capability boundary — a service that exposes tools, resources, and prompts over a protocol so that any compliant client can discover and invoke them. Putting operational capability behind MCP instead of local
@Toolbeans is what gives us an independent authorization boundary. - Multi-agent system: agents delegating to agents. This series deliberately stops short of that (there is an optional A2A lab) because every additional delegating hop multiplies the audit and authorization surface.
ops-agent-platform is a tool-using agent talking to a remote MCP server, with RAG on one side and an approval gate on the other.
System context and architecture
The simulator stands in for the real systems of record — a service catalog, deployment registry, and incident tracker — so the series is self-contained and can inject failures on demand. In a real deployment, mcp-operations-server would wrap ServiceNow, PagerDuty, or your internal equivalents; the agent code would not change.
Module boundaries
The rule that matters most: domain-contracts — identifiers, commands, results, the IncidentAssessment type, error envelopes — depends on nothing from Spring AI, JPA, HTTP, or a model provider. AI types may flow inward through ports; they may not leak into contracts. ArchUnit enforces this from Chapter 1.
The two request shapes
A read-only question (RAG) never touches a tool:
A read-only tool call adds the MCP hop:
And the write path, which is where the series spends its paranoia:
Security note: the model proposes every tool call, but the decision is never the model’s. Scope check, risk classification, approval requirement, and argument integrity all happen in deterministic code. If the model changes the arguments after approval — or a prompt injection asks it to — the argument-hash check fails and the call is rejected.
Trust boundaries
Prompts, model output, retrieved runbook text, and MCP metadata are all treated as untrusted input. The delimiters in the system prompt are a parsing aid, not a security boundary — the security boundary is the deterministic policy code.
Ingestion and retrieval
Ingestion is a separate command-line module, not a side effect of an HTTP request: document parsing, section-aware chunking, deterministic chunk IDs, per-document checksums so only changed runbooks are re-embedded.
The advisor chain and agent loop
Every loop iteration consumes budget — model calls, tool calls, wall-clock time, tokens. The loop terminates on an answer, a refusal, an approval wait, or budget exhaustion; there is no unbounded “the model decides when to stop.”
Deployment
Compose is the runnable reference environment:
Kubernetes (Chapter 15) mirrors the same shapes with Deployments, a one-shot ingestion Job, ConfigMaps, Secrets, NetworkPolicies, and PodDisruptionBudgets — the Compose file is the honest local environment, not a mini-production.
Threat model skeleton
docs/threat-model/ grows through the series; the skeleton committed at this checkpoint names assets, actors, entry points, and the threats that shape the design.
Assets: tenant-scoped runbooks; live operational data (service health, incidents); the ability to create incidents and notes; credentials and tokens; model spend; audit integrity.
Actors: ops viewer, ops operator, service accounts, a compromised or malicious document author, a prompt-injection attacker, a compromised MCP client.
Threats with their primary controls:
| Threat | Control |
|---|---|
| Direct prompt injection | Prompts never decide authorization; policy code does |
| Indirect injection via runbook content | Retrieved text delimited and treated as data; ingestion review gate |
| Cross-tenant retrieval or tool access | Tenant filter in retrieval query and tenant policy on every tool call |
| Confused deputy (agent acts with ambient privileges) | Per-tool scopes; service account cannot approve its own writes |
| Approval replay / TOCTOU | Single-use tokens bound to user, tenant, tool, normalized args hash, expiry |
| Duplicate incident on retry | Idempotency-Key required on all writes, enforced by simulator |
| Token theft / wrong issuer / wrong audience | JWT issuer + audience + scope validation on both HTTP boundaries |
| SSRF via model-generated URLs/identifiers | Typed identifiers and allowlists; model never supplies raw URLs |
| Secrets/prompts in logs, traces, evals | Redaction filters; audit sink separate from diagnostics |
| DoS via long prompts, recursive calls, retry storms | Size limits, step budgets, deadlines, retry classification, bulkheads |
| Supply chain (deps, models) | Locked versions, pinned image tags, pinned evaluator prompts |
Residual risks (accepted and documented): a clever injection may still produce a wrong answer — which is why writes are gated but answers are advisory; local Ollama profiles relax some controls for developer convenience, and that relaxation is labeled, not hidden.
Technology choices and alternatives
| Decision | Choice | Alternative considered |
|---|---|---|
| Runtime | Java 25 LTS | Java 27 — Boot 4.1 supports through 26; 27 appears only in an optional lab |
| Framework | Spring Boot 4.1.1 / Framework 7.x | Boot 3.5 — the 4.x/AI-2.0 line is the current pairing |
| AI integration | Spring AI 2.0.1 (spring-ai-bom) | LangChain4j — fine library; Spring AI keeps the BOM-managed stack coherent |
| MCP transport | Streamable HTTP, stateless tools | SSE transport — deprecated in MCP SDK 2.0; stdio — no network auth boundary |
| Tool server language | Kotlin 2.3.21 | Java — Kotlin earns its place here (data classes, when); one module, not the whole repo |
| Vector store | PostgreSQL 17 + pgvector 0.8.6 | Qdrant/pgvector-managed — extra infra for runbook-scale data |
| Identity | Keycloak 26.7.x, OIDC | Hand-rolled JWT — never in a security chapter |
| Resilience | Framework 7 retry/concurrency primitives | Resilience4j — added only if Chapter 11 hits its limits |
| Load test | Gatling | k6 — JVM DSL keeps everything in one toolchain |
Expected resource requirements
A laptop with 8 GB free RAM runs the whole Compose stack: Ollama with qwen3:4b plus nomic-embed-text fits in roughly 4 GB; a llama3.2:3b path is documented for tighter machines. Ports used: 3000, 3100, 3200, 4317/4318, 5432, 8080–8082, 8085, 9090, 11434. The hosted-provider profile is optional and excluded from default builds; no paid account is required anywhere in the series.
Repository roadmap and the final demo
The tree from the blueprint is committed empty in Chapter 1 and filled in chapter order: simulator before agent (Chapter 2 before 3), knowledge before tools (4–5 before 7–8), security before writes (9 before 10). Each chapter ends at a Git tag.
The final exercise in Chapter 16 replays one scenario end to end: a user asks why payment-gateway is degraded; the agent cites the runbook section, calls get_service_status and list_recent_incidents, produces a schema-valid IncidentAssessment, proposes create_incident, waits for human approval, executes exactly once under an idempotency key, and the whole path is visible as one trace in Tempo with an audit record for the write.
What comes next
Chapter 1 turns this map into a buildable skeleton: the Gradle wrapper, version catalog, convention plugins, and the first ArchUnit rules — so that domain-contracts physically cannot drift toward Spring AI no matter what a later chapter adds.
Project State Ledger — chapter-00-architecture
- Modules declared:
build-logic,domain-contracts,agent-api,knowledge-ingestion,mcp-operations-server,operations-simulator,evaluation-suite,architecture-tests,test-support - Root package:
in.o612.eng.opsagent.<module>; Kotlin escapesinwith backticks - Ports: agent-api 8080, mcp 8081, simulator 8082, keycloak 8085, postgres 5432, ollama 11434, otel-collector 4317/4318, prometheus 9090, tempo 3200, loki 3100, grafana 3000
- Databases:
simdb(schemasim),opsdb(schemasknowledge,agent) - Scopes:
agent:invoke;ops:read;ops:incident:write;ops:note:write - MCP tools (planned):
get_service_status,list_recent_incidents,get_incident,create_incident,append_incident_note - Models: Ollama
qwen3:4bchat,nomic-embed-textembeddings (vector(768)); hosted profile viaspring-ai-starter-model-openai - Tenants:
acme,globex; userspriya(operator),sam(viewer) - Next:
chapter-01-build-foundation— Gradle scaffold, convention plugins, CI, first architecture tests