Series overview
Part 4 of 1724% complete
2026-08-19•5 min read

Create the agent API: ChatClient, provider profiles, and a deterministic stub model

Checkpoint tag: chapter-03-basic-agent-api — agent-api answers messages through a real model locally, streams over SSE, and the entire default test suite runs offline against a stub.

What will be built

agent-api gains its request surface: create a conversation, post a message, get an answer, and stream events over SSE. Underneath sits the seam the whole series protects — a ConversationService that talks to a ChatModel port, with three interchangeable bindings: Ollama (local default), a hosted OpenAI-compatible profile, and a deterministic StubChatModel for tests and CI.

Why it matters

The first architectural decision in an AI application is where the model lives in the dependency graph. If ChatClient types leak into your application service, every later concern — retrieval, tools, approval, evaluation — gets welded to one provider’s SDK. We put the model behind a port on day one, which is why Chapter 6 can swap in structured output and Chapter 8 can add tool calls without touching HTTP code.

The second decision is how you stream. We use Spring MVC plus SseEmitter, not WebFlux: SSE is a server-push channel with trivial backpressure requirements (the client reads or it doesn’t), and a reactive stack would buy us nothing here while taxing every subsequent chapter.

Prerequisites and starting tag

Starting tag: chapter-02-operations-simulator. For the real-model path: Ollama running (ollama pull qwen3:4b) or the Compose service from the updated stack file below. Everything testable runs without either.

Concepts explained

ChatClient vs. ChatModel. ChatModel is the low-level SPI (call(Prompt) → ChatResponse, stream(Prompt) → Flux<ChatResponse>). ChatClient is the fluent facade — prompt assembly, advisors, tool callbacks. Our application code depends on a ModelGateway port that wraps ChatClient; the stub implements the port directly so no Spring AI machinery sits between a test and a canned answer.

Profiles as provider selection. agent.model.provider selects the bean wiring, not just config values: ollama, openai, or stub. Credentials never live in the repo; the hosted profile reads OPENAI_API_KEY from the environment and the app fails fast if it is absent, rather than silently falling back to a local model with different behavior.

SSE over MVC. SseEmitter gives each connected client a queue; the conversation’s event bus fans events out to its emitters. When the emitter times out or the client disconnects, we remove it — the same path Chapter 10 reuses for approval-request events.

Files added or changed

domain-contracts: + conversation types (ConversationId, ChatRequest, ChatReply, AgentEvent)
agent-api/src/main/resources/application.yml (+ application-openai.yml)
agent-api/src/main/resources/prompts/system.st
agent-api/src/main/java/in/o612/eng/opsagent/agent/
api/ConversationController.java, api/ApiExceptionHandler.java
conversation/ConversationService.java, conversation/ConversationStore.java, conversation/ConversationEventBus.java
model/ModelGateway.java, model/ChatClientGateway.java, model/StubModelGateway.java, model/ModelConfig.java
config/AgentProperties.java
agent-api/src/main/resources/db/migration/V1__agent_init.sql
agent-api/src/test/java/... (stub-driven tests)
infra/compose/docker-compose.yml (+ ollama)
api-requests/agent.http

Complete code — contracts

domain-contracts/src/main/java/in/o612/eng/opsagent/contracts/conversation/AgentEvent.java
package in.o612.eng.opsagent.contracts.conversation;
import java.time.Instant;
public sealed interface AgentEvent {
String conversationId();
Instant at();
record Token(String conversationId, String text, Instant at) implements AgentEvent {}
record ToolProposed(String conversationId, String tool, String argsJson, Instant at) implements AgentEvent {}
record ApprovalRequired(String conversationId, String approvalId, String tool,
String argsJson, Instant expiresAt, Instant at) implements AgentEvent {}
record Completed(String conversationId, String answer, Instant at) implements AgentEvent {}
record Failed(String conversationId, String code, String message, Instant at) implements AgentEvent {}
}

Plus ChatRequest(String message), ChatReply(String answer, List<Citation> citations, List<String> evidence), and a Citation record (documentId, headingPath, chunkId) that Chapter 5 fills in — until then both lists are empty.

Complete code — agent-api

application.yml

agent-api/src/main/resources/application.yml
server:
port: 8080
spring:
application:
name: agent-api
datasource:
url: ${OPS_DB_URL:jdbc:postgresql://localhost:5432/opsdb}
username: ${OPS_DB_USER:agent}
password: ${OPS_DB_PASSWORD:agent-dev-password}
flyway:
schemas: agent
default-schema: agent
ai:
ollama:
base-url: ${OLLAMA_BASE_URL:http://localhost:11434}
chat:
options:
model: qwen3:4b
temperature: 0.1
agent:
model:
provider: ${MODEL_PROVIDER:ollama} # ollama | openai | stub
timeout: 60s
max-prompt-chars: 16000
system-prompt: classpath:prompts/system.st

application-openai.yml — the hosted profile. It only activates alongside agent.model.provider=openai; the missing-key failure is loud:

agent-api/src/main/resources/application-openai.yml
spring:
ai:
openai:
api-key: ${OPENAI_API_KEY:?Set OPENAI_API_KEY for the hosted profile}
chat:
options:
model: ${OPENAI_MODEL:gpt-4o-mini}
temperature: 0.1

Dependencies added to agent-api/build.gradle.kts: spring-ai-starter-model-ollama, spring-ai-starter-model-openai, spring-boot-starter-jdbc, spring-boot-starter-flyway, runtimeOnly postgresql.

ModelGateway — the port:

agent-api/src/main/java/in/o612/eng/opsagent/agent/model/ModelGateway.java
package in.o612.eng.opsagent.agent.model;
import java.util.List;
import java.util.function.Consumer;
public interface ModelGateway {
/** Blocking call; returns the model's complete text answer. */
String complete(String systemPrompt, List<Turn> history, String userMessage);
/** Streaming call; invokes onToken per chunk, onDone when finished. */
void stream(String systemPrompt, List<Turn> history, String userMessage,
Consumer<String> onToken, Runnable onDone, Consumer<Throwable> onError);
record Turn(String role, String content) {}
}

ChatClientGateway — the Spring AI binding:

agent-api/src/main/java/in/o612/eng/opsagent/agent/model/ChatClientGateway.java
package in.o612.eng.opsagent.agent.model;
import org.springframework.ai.chat.client.ChatClient;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.messages.Message;
import org.springframework.ai.chat.messages.UserMessage;
import java.util.ArrayList;
import java.util.List;
import java.util.function.Consumer;
public class ChatClientGateway implements ModelGateway {
private final ChatClient chatClient;
public ChatClientGateway(ChatClient chatClient) { this.chatClient = chatClient; }
@Override
public String complete(String systemPrompt, List<Turn> history, String userMessage) {
return chatClient.prompt()
.system(systemPrompt)
.messages(toMessages(history))
.user(userMessage)
.call()
.content();
}
@Override
public void stream(String systemPrompt, List<Turn> history, String userMessage,
Consumer<String> onToken, Runnable onDone, Consumer<Throwable> onError) {
chatClient.prompt()
.system(systemPrompt)
.messages(toMessages(history))
.user(userMessage)
.stream()
.content()
.doOnNext(onToken)
.doOnError(onError)
.doOnComplete(onDone)
.subscribe();
}
private List<Message> toMessages(List<Turn> history) {
var messages = new ArrayList<Message>();
for (Turn t : history) {
messages.add(t.role().equals("assistant")
? new AssistantMessage(t.content())
: new UserMessage(t.content()));
}
return messages;
}
}

StubModelGateway — deterministic, offline, and the reason default CI needs no GPU, no container, and no API key:

agent-api/src/main/java/in/o612/eng/opsagent/agent/model/StubModelGateway.java
package in.o612.eng.opsagent.agent.model;
import java.util.List;
import java.util.function.Consumer;
public class StubModelGateway implements ModelGateway {
@Override
public String complete(String systemPrompt, List<Turn> history, String userMessage) {
return "stub-answer: " + userMessage;
}
@Override
public void stream(String systemPrompt, List<Turn> history, String userMessage,
Consumer<String> onToken, Runnable onDone, Consumer<Throwable> onError) {
String answer = complete(systemPrompt, history, userMessage);
for (String word : answer.split(" ")) {
onToken.accept(word + " ");
}
onDone.run();
}
}

ModelConfig — explicit bean selection, no classpath roulette:

agent-api/src/main/java/in/o612/eng/opsagent/agent/model/ModelConfig.java
package in.o612.eng.opsagent.agent.model;
import org.springframework.ai.chat.client.ChatClient;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.boot.autoconfigure.condition.ConditionalOnProperty;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
@Configuration
public class ModelConfig {
@Bean
@ConditionalOnProperty(name = "agent.model.provider", havingValue = "stub")
ModelGateway stubGateway() {
return new StubModelGateway();
}
@Bean
@ConditionalOnProperty(name = "agent.model.provider", havingValue = "ollama", matchIfMissing = true)
ChatClient ollamaChatClient(ChatModel ollamaChatModel) {
return ChatClient.builder(ollamaChatModel).build();
}
@Bean
@ConditionalOnProperty(name = "agent.model.provider", havingValue = "openai")
ChatClient openAiChatClient(ChatModel openAiChatModel) {
return ChatClient.builder(openAiChatModel).build();
}
@Bean
@ConditionalOnProperty(name = "agent.model.provider", havingValue = "ollama", matchIfMissing = true)
ModelGateway chatClientGateway(ChatClient ollamaChatClient) { return new ChatClientGateway(ollamaChatClient); }
}

ConversationService — provider-neutral orchestration. History is capped (last 20 turns), prompt size is enforced before the model is ever invoked, and every message is persisted with its tenant and correlation IDs:

agent-api/src/main/java/in/o612/eng/opsagent/agent/conversation/ConversationService.java
package in.o612.eng.opsagent.agent.conversation;
import in.o612.eng.opsagent.agent.config.AgentProperties;
import in.o612.eng.opsagent.agent.model.ModelGateway;
import in.o612.eng.opsagent.contracts.conversation.AgentEvent;
import org.springframework.stereotype.Service;
import java.time.Instant;
import java.util.List;
@Service
public class ConversationService {
private final ConversationStore store;
private final ConversationEventBus events;
private final ModelGateway model;
private final AgentProperties props;
private final String systemPrompt;
public ConversationService(ConversationStore store, ConversationEventBus events,
ModelGateway model, AgentProperties props,
org.springframework.core.io.ResourceLoader loader) {
this.store = store;
this.events = events;
this.model = model;
this.props = props;
this.systemPrompt = loadPrompt(loader, props.systemPrompt());
}
public String createConversation(String tenantId, String userId) {
return store.create(tenantId, userId);
}
public void handleMessage(String conversationId, String tenantId, String message) {
if (message.length() > props.maxPromptChars()) {
events.publish(conversationId,
new AgentEvent.Failed(conversationId, "PROMPT_TOO_LARGE",
"message exceeds " + props.maxPromptChars() + " chars", Instant.now()));
return;
}
store.append(conversationId, "user", message);
List<ModelGateway.Turn> history = store.recentTurns(conversationId, 20);
StringBuilder answer = new StringBuilder();
model.stream(systemPrompt, history, message,
token -> {
answer.append(token);
events.publish(conversationId,
new AgentEvent.Token(conversationId, token, Instant.now()));
},
() -> {
store.append(conversationId, "assistant", answer.toString());
events.publish(conversationId,
new AgentEvent.Completed(conversationId, answer.toString(), Instant.now()));
},
err -> events.publish(conversationId,
new AgentEvent.Failed(conversationId, "MODEL_ERROR",
"model call failed", Instant.now())));
}
private static String loadPrompt(org.springframework.core.io.ResourceLoader l, String path) {
try (var in = l.getResource(path).getInputStream()) {
return new String(in.readAllBytes(), java.nio.charset.StandardCharsets.UTF_8);
} catch (java.io.IOException e) {
throw new IllegalStateException("system prompt not readable: " + path, e);
}
}
}

ConversationEventBus — a ConcurrentHashMap<String, CopyOnWriteArrayList<SseEmitter>>; subscribe(conversationId) registers an emitter with onTimeout/onCompletion removal, publish iterates and prunes dead emitters. Single-node by design — the multi-node variant (a real broker) is a stated limitation, not hidden complexity.

ConversationStore — JdbcClient over the agent schema below. recentTurns returns the last N messages oldest-first.

ConversationController:

agent-api/src/main/java/in/o612/eng/opsagent/agent/api/ConversationController.java
package in.o612.eng.opsagent.agent.api;
import in.o612.eng.opsagent.agent.conversation.ConversationEventBus;
import in.o612.eng.opsagent.agent.conversation.ConversationService;
import in.o612.eng.opsagent.contracts.conversation.ChatReply;
import in.o612.eng.opsagent.contracts.conversation.ChatRequest;
import org.springframework.http.MediaType;
import org.springframework.web.bind.annotation.*;
import org.springframework.web.servlet.mvc.method.annotation.SseEmitter;
import java.util.List;
import java.util.Map;
@RestController
@RequestMapping("/api/v1")
public class ConversationController {
private final ConversationService conversations;
private final ConversationEventBus events;
public ConversationController(ConversationService c, ConversationEventBus e) {
this.conversations = c;
this.events = e;
}
@PostMapping("/conversations")
public Map<String, String> create(@RequestHeader("X-Tenant-Id") String tenant,
@RequestHeader("X-User-Id") String user) {
return Map.of("conversationId", conversations.createConversation(tenant, user));
}
@PostMapping("/conversations/{id}/messages")
public ChatReply message(@PathVariable String id,
@RequestHeader("X-Tenant-Id") String tenant,
@RequestBody ChatRequest request) {
// synchronous convenience path: same pipeline, non-streamed
conversations.handleMessage(id, tenant, request.message());
return new ChatReply("streamed via SSE; see /events", List.of(), List.of());
}
@GetMapping(path = "/conversations/{id}/events", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public SseEmitter events(@PathVariable String id) {
return events.subscribe(id);
}
}

V1__agent_init.sql

agent-api/src/main/resources/db/migration/V1__agent_init.sql
CREATE SCHEMA IF NOT EXISTS agent;
CREATE TABLE agent.conversations (
id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
user_id TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE agent.messages (
id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
conversation_id TEXT NOT NULL REFERENCES agent.conversations(id),
role TEXT NOT NULL CHECK (role IN ('user','assistant','system','tool')),
content TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX messages_conversation_idx ON agent.messages (conversation_id, id);

prompts/system.st — the baseline system prompt, versioned as a file, not a string literal in a service:

agent-api/src/main/resources/prompts/system.st
You are an operations assistant for an internal platform team.
Answer only from the provided runbook excerpts and tool results.
If the evidence does not support an answer, say so explicitly.
Never invent incidents, statuses, commands, or runbook content.
Cite the document id and heading for every factual claim.

Finally, docker-compose.yml gains an ollama service (ollama/ollama image, port 11434, a small entrypoint script that pulls qwen3:4b and nomic-embed-text on first start).

Commands to build and run

terminal
docker compose -f infra/compose/docker-compose.yml up -d postgres ollama
MODEL_PROVIDER=stub ./gradlew :agent-api:bootRun # fully offline
./gradlew :agent-api:bootRun # real model via Ollama
MODEL_PROVIDER=openai OPENAI_API_KEY=sk-... ./gradlew :agent-api:bootRun # hosted

API requests and expected responses

api-requests/agent.http
POST http://localhost:8080/api/v1/conversations
X-Tenant-Id: acme
X-User-Id: priya
# -> {"conversationId": "conv-..."}
GET http://localhost:8080/api/v1/conversations/conv-.../events
Accept: text/event-stream
POST http://localhost:8080/api/v1/conversations/conv-.../messages
X-Tenant-Id: acme
Content-Type: application/json
{"message": "how do I restart a degraded service"}

With MODEL_PROVIDER=stub, the SSE channel emits data: {"text":"stub-answer: "}-style Token events followed by Completed. With Ollama, you get real tokens; latency is the model’s, not the framework’s — keep that distinction, Chapter 14 depends on it.

Automated tests

  • ConversationServiceTest (unit, stub gateway): oversized message produces a Failed event and no model call; tokens arrive before Completed; assistant turn persisted.
  • ConversationControllerIT (MODEL_PROVIDER=stub): create → message → SSE receives Token then Completed.
  • ModelConfigTest: provider=stub yields StubModelGateway; provider=openai without OPENAI_API_KEY fails context startup with a clear missing-env error.
  • No test touches the network.

Failure-injection lab

  1. Kill Ollama mid-stream: SSE emits Failed(code=MODEL_ERROR); the partially persisted assistant turn is not written (only completed answers persist — verify in agent.messages).
  2. Send a 20 KB message: Failed(code=PROMPT_TOO_LARGE); the model was never invoked — check the stub’s call counter.
  3. Set MODEL_PROVIDER=openai with no key: startup fails fast. Silent provider fallback would be a production incident masquerading as convenience.

Security considerations

Headers X-Tenant-Id/X-User-Id are placeholders — trusted-header auth is a lie and Chapter 9 replaces them with JWT validation; they exist now only so tenant scoping is structural, not bolted on. Message size caps already bound the cheapest DoS vector. The stub provider means CI artifacts never contain prompts or API keys.

Observability checks

agent.conversations.created and agent.messages.handled counters; agent.model.latency timer tagged by provider (stub|ollama|openai) — the label that will later separate our latency from the model’s.

Troubleshooting

  • ChatModel bean ambiguity (both Ollama and OpenAI on the classpath): the @ConditionalOnProperty wiring above avoids it; if you see it anyway, check for a stray auto-configured ChatClient.Builder injection.
  • SSE connects but no events: events publish to the conversation’s emitter list — confirm the GET /events subscription happened before POST /messages.
  • Ollama 404 on chat: model not pulled; docker exec ollama ollama pull qwen3:4b.

Checkpoint verification checklist

  • MODEL_PROVIDER=stub ./gradlew :agent-api:bootRun works with zero external services beyond Postgres.
  • SSE stream delivers Token events then Completed.
  • Default ./gradlew check is fully offline.
  • Conversations persist per tenant; oversized prompts are rejected pre-model.

Commit message and Git tag

feat(agent-api): conversation API with provider-neutral model gateway and SSE

git tag chapter-03-basic-agent-api

What comes next

Chapter 4 feeds the agent something worth saying: the runbook ingestion pipeline that turns versioned Markdown into pgvector chunks.

Project State Ledger — chapter-03-basic-agent-api

  • New endpoints: POST /api/v1/conversations, POST /conversations/{id}/messages, GET /conversations/{id}/events (SSE)
  • New tables: agent.conversations, agent.messages (V1__agent_init.sql)
  • Ports: ModelGateway (complete, stream); impls ChatClientGateway, StubModelGateway
  • Provider switch: agent.model.provider = ollama|openai|stub; env OPENAI_API_KEY, OPENAI_MODEL, OLLAMA_BASE_URL, MODEL_PROVIDER
  • Prompts: classpath:prompts/system.st, versioned
  • Limits: agent.model.max-prompt-chars=16000, history capped at 20 turns
  • Known limitation: event bus is single-node; auth headers are placeholders until Ch 9
  • Next: chapter-04-knowledge-ingestion
Spring BootJavaAI

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind