Add typed, self-correcting outputs: the IncidentAssessment schema
Checkpoint tag: chapter-06-structured-output — IncidentAssessment responses deserialize safely, invalid output fails predictably after a bounded number of correction attempts, and every retry’s token cost is visible.
What will be built
A second agent capability: POST /api/v1/conversations/{id}/assess produces a typed IncidentAssessment — severity, confidence, evidence references, recommended actions — from the retrieved context. The flow validates syntax (JSON schema from the record itself, provider-native constraints where supported), validates semantics (evidence IDs must name chunks that were actually retrieved), and retries malformed output a bounded number of times before returning a typed failure.
Why it matters
Chapter 5’s citation parsing is the limit of what prose output can give you: regex, hope, and a counter for when it goes wrong. The assessment is different — it feeds downstream automation, so “the model usually returns JSON” is not a plan. Schema validity buys you shape; it does not buy correctness — a response can be perfectly valid JSON and still cite evidence that was never retrieved. Both layers get their own check, and the chapter measures what self-correction costs instead of assuming it’s free.
Concepts explained
entity() in Spring AI 2.0. .call().entity(Type.class, spec -> …) generates a JSON schema from the Java type, appends it to the prompt, parses the response, and — with the two reliability switches — can push the schema to the provider as an API-level constraint (useProviderStructuredOutput()) and auto-retry schema-invalid responses (validateSchema()). entity() is call()-only: typed parsing needs the whole response, so assessments are a synchronous endpoint by design.
Schema-valid ≠ true. The schema proves the response parses. EvidenceValidator proves the response is grounded — every evidence entry names a chunk_id from this request’s retrieval set. Confusing those two layers is how hallucinated citations ship.
Bounded correction. Retry-with-error-feedback is the honest version of “the model will get it right”: at most 3 attempts, each augmented with the validation error, cumulative token usage tracked. Unbounded correction is an infinite loop that bills you.
Files added or changed
domain-contracts: + assessment/{IncidentAssessment, Confidence}.javaagent-api/src/main/java/in/o612/eng/opsagent/agent/ assessment/AssessmentService.java, assessment/EvidenceValidator.java api/AssessmentController.javaagent-api/src/main/resources/prompts/assessment.stComplete code
The contract — deliberately in domain-contracts, so the MCP server could produce or consume the same shape without importing Spring AI:
package in.o612.eng.opsagent.contracts.assessment;
import in.o612.eng.opsagent.contracts.incident.Severity;import java.util.List;
public record IncidentAssessment( String summary, Severity severity, Confidence confidence, String serviceId, List<String> evidence, // chunk_ids backing the assessment List<String> recommendedActions) {
public enum Confidence { LOW, MEDIUM, HIGH }}assessment.st — a separate prompt file; assessment is a different task than chat and deserves its own versioned prompt:
You are producing a structured incident assessment for an ops engineer.Base every claim on the numbered runbook excerpts and tool results provided.evidence must list only the excerpt numbers used, as "chk-" ids from theprovided metadata. If evidence is thin, lower confidence — do not inflate it.AssessmentService:
package in.o612.eng.opsagent.agent.assessment;
import in.o612.eng.opsagent.agent.retrieval.RunbookRetriever;import in.o612.eng.opsagent.contracts.assessment.IncidentAssessment;import io.micrometer.core.instrument.MeterRegistry;import org.springframework.ai.chat.client.ChatClient;import org.springframework.stereotype.Service;
import java.util.Set;import java.util.stream.Collectors;
@Servicepublic class AssessmentService {
public sealed interface AssessmentOutcome { record Assessed(IncidentAssessment assessment, int attempts) implements AssessmentOutcome {} record Rejected(String reason) implements AssessmentOutcome {} }
private final ChatClient chatClient; private final RunbookRetriever retriever; private final EvidenceValidator evidence; private final MeterRegistry metrics; private final String prompt;
public AssessmentOutcome assess(String tenantId, String serviceId, String question) { var retrieval = retriever.retrieve(tenantId, serviceId, question); if (!retrieval.sufficient()) { return new AssessmentOutcome.Rejected("insufficient evidence"); } Set<String> validEvidence = retrieval.chunks().stream() .map(RunbookRetriever.RetrievedChunk::chunkId) .collect(Collectors.toSet());
IncidentAssessment raw; try { raw = chatClient.prompt() .system(prompt + "\n\nValid evidence ids: " + validEvidence) .user(buildUserMessage(question, retrieval)) .call() .entity(IncidentAssessment.class, spec -> spec.validateSchema()) // provider-native structured output is enabled only on the // hosted profile where the provider supports it ; } catch (RuntimeException e) { // schema validation exhausted its retries metrics.counter("agent.schema.corrections_exhausted").increment(); return new AssessmentOutcome.Rejected("model produced no schema-valid assessment"); }
var violations = evidence.check(raw, validEvidence); if (!violations.isEmpty()) { metrics.counter("agent.assessment.semantic_rejections").increment(); return new AssessmentOutcome.Rejected("assessment cited non-retrieved evidence: " + violations); } return new AssessmentOutcome.Assessed(raw, attemptsUsed()); }
// attemptsUsed() reads the counter delta of agent.schema.corrections around the call; // buildUserMessage() assembles the numbered context exactly as ConversationService does}EvidenceValidator — pure, deterministic, and trivially unit-testable:
package in.o612.eng.opsagent.agent.assessment;
import in.o612.eng.opsagent.contracts.assessment.IncidentAssessment;import org.springframework.stereotype.Component;
import java.util.List;import java.util.Set;
@Componentpublic class EvidenceValidator {
public List<String> check(IncidentAssessment a, Set<String> validEvidence) { var problems = new java.util.ArrayList<String>(); if (a.evidence() == null || a.evidence().isEmpty()) { problems.add("no evidence cited"); } else { a.evidence().stream() .filter(id -> !validEvidence.contains(id)) .forEach(id -> problems.add("unknown evidence id: " + id)); } if (a.confidence() == IncidentAssessment.Confidence.HIGH && (a.evidence() == null || a.evidence().size() < 2)) { problems.add("HIGH confidence requires at least two evidence refs"); } if (a.recommendedActions() == null || a.recommendedActions().isEmpty()) { problems.add("no recommended actions"); } return problems; }}AssessmentController maps Assessed → 200, Rejected → 422 with the reason in the ApiError envelope — an unprocessable assessment is a normal outcome, not a crash.
Automated tests
EvidenceValidatorTest(pure): unknown chunk ID rejected; HIGH confidence with one evidence ref rejected; happy path passes.AssessmentServiceTest: stubChatClient— actually a stubbedModelGateway-level fake is too low here, so the test uses a scriptedChatModelbean that returns (a) malformed JSON →Rejectedafter the bounded retries, withagent.schema.correctionsincremented; (b) valid JSON citingchk-deadbeefnot in the retrieval set → semanticRejected; (c) valid + grounded →Assessed.AssessmentSchemaTest: the generated schema’s required fields match the record components — a guard against someone renamingevidenceand silently widening what the model may omit.RetryBudgetTest: schema-invalid stub makes exactlymaxAttemptscalls, not more — verified via the scripted model’s invocation counter.
Failure-injection lab
Scripted ChatModel stubs make failure deterministic:
| Stub returns | Outcome |
|---|---|
| prose around JSON | retried; correction count +1 |
{"summary": 42} (type violation) | retries, then Rejected |
| valid JSON, invented evidence | semantic rejection, 422 |
| valid + grounded | Assessed |
Measure what self-correction costs: log token usage per attempt (Spring AI’s ChatResponse.getMetadata().getUsage() is cumulative across validation retries — surface it as agent.assessment.tokens per outcome). Do not claim a number you didn’t run; record yours in docs/runbooks/ results template.
Security considerations
Assessments are advisory — nothing in this flow mutates state — but serviceId still comes from the request path, never from model output, and evidence checking prevents a crafted retrieval corpus from laundering fake support into a high-confidence assessment.
Observability checks
agent.assessment.attempts histogram, agent.schema.corrections, agent.assessment.semantic_rejections, agent.assessment.tokens — the four numbers that tell you whether corrections are a safety net or a money leak.
Checkpoint verification checklist
- Malformed output retries ≤ the bound, then fails as
Rejected(422), never as a parse exception to the caller. - Schema-valid but ungrounded output is rejected semantically.
- Token usage is cumulative across retries and visible in metrics.
-
entity()is never invoked on the streaming path.
Commit message and Git tag
feat(agent-api): schema-validated IncidentAssessment with semantic evidence checksgit tag chapter-06-structured-output
What comes next
Chapter 7 builds the other half of the platform — the Kotlin MCP server that will carry the operational tools the agent calls in Chapter 8.
Project State Ledger — chapter-06-structured-output
- Contracts added:
IncidentAssessment(summary, severity, confidence, serviceId, evidence[], recommendedActions[]),Confidenceenum - Endpoint:
POST /api/v1/conversations/{id}/assess→ 200IncidentAssessment| 422ApiError - Validation layers: schema (entity/validateSchema, bounded) → semantic (
EvidenceValidator) → outcome - Metrics:
agent.assessment.attempts,agent.schema.corrections,agent.assessment.semantic_rejections,agent.assessment.tokens - Rule:
entity()on.call()only; streaming stays text - Next:
chapter-07-mcp-server