Add guardrails and bounded orchestration: budgets, timeouts, and degraded modes
Checkpoint tag: chapter-11-guardrails-resilience — runaway loops, slow models, injected instructions, and retry storms all terminate safely and observably.
What will be built
The reliability layer: an explicit timeout hierarchy (HTTP → orchestration → per-tool → downstream), retry classification (retryable in the policy registry now means something), a concurrency bulkhead around model calls, prompt/context size enforcement end to end, safe handling of model-generated identifiers, and a DEGRADED response mode so a dead vector store or MCP server produces a bounded answer instead of a stack trace.
Why it matters
Every unbounded thing in an agent system is a billing vector or an outage vector. The model is slow; tools are slower; a retry storm multiplies both. The disciplines here are the same ones you’d apply to any fan-out service — budgets, deadlines, bulkheads — but the nondeterministic component makes them mandatory rather than nice: a model that loops is not a bug you fix, it’s a behavior you bound.
Concepts explained
Retry classification, not retry enthusiasm. The registry’s retryable flag exists because “503 from a status lookup” (retry, once, with jitter) and “403 tenant mismatch” (never retry) are different animals. Rules: retry only on transport errors and 5xx/429; never on 4xx, timeouts past deadline, or schema-invalid tool output; at most 2 attempts; jittered backoff (100–400 ms) so parallel callers don’t synchronize into a thundering herd.
The timeout hierarchy must compose inward. A 45s loop deadline containing 12s tool calls containing a 10s downstream read timeout: each layer’s budget must fit inside its parent’s, or outer timeouts silently truncate inner retries. The invariant is checked by a unit test, not by faith.
Bulkheads on the model. Semaphore around ModelGateway calls caps concurrent model invocations (default 8). Excess requests get Failed(code=SATURATED) immediately rather than queueing into memory. The model is the most expensive dependency in the system — protect its callers from each other first.
Degraded mode. RetrievalResult failure → answer without RAG (labeled “no runbook context available”); MCP unavailable → read tools return UPSTREAM_* results the model can summarize honestly; model itself down → Failed(code=MODEL_UNAVAILABLE) + 503. Each degradation is a product decision written in code, not an accident.
Files added or changed
agent-api/…/resilience/{RetryExecutor, DeadlineContext, ModelBulkhead, DegradedModeAdvisor}.javaagent-api/…/tools/McpToolExecutor.java (retry classification)agent-api/…/config/ResilienceProperties.javaagent-api/src/test/... (resilience labs)Complete code (the load-bearing parts)
package in.o612.eng.opsagent.agent.resilience;
import java.util.concurrent.ThreadLocalRandom;import java.util.function.Supplier;
public class RetryExecutor {
public static <T> T execute(Supplier<T> call, boolean retryable) { int attempts = retryable ? 2 : 1; RuntimeException last = null; for (int i = 0; i < attempts; i++) { try { return call.get(); } catch (RuntimeException e) { if (!isRetryable(e) || i == attempts - 1) throw e; last = e; sleep(100 + ThreadLocalRandom.current().nextInt(300)); // jitter } } throw last; }
static boolean isRetryable(RuntimeException e) { // transport failures and 5xx/429 only; 4xx, auth, schema errors -> never return switch (e) { case org.springframework.web.client.ResourceAccessException rae -> true; case org.springframework.web.client.HttpServerErrorException hse -> true; default -> false; }; }
private static void sleep(long ms) { try { Thread.sleep(ms); } catch (InterruptedException ie) { Thread.currentThread().interrupt(); } }}package in.o612.eng.opsagent.agent.resilience;
import org.springframework.stereotype.Component;import java.util.concurrent.Semaphore;import java.util.function.Supplier;
@Componentpublic class ModelBulkhead {
private final Semaphore permits = new Semaphore(8);
public <T> T guard(Supplier<T> call) { if (!permits.tryAcquire()) { throw new SaturatedException("model concurrency limit reached"); } try { return call.get(); } finally { permits.release(); } }
public static class SaturatedException extends RuntimeException { public SaturatedException(String m) { super(m); } }}The orchestrator wraps each chatClient…call() in bulkhead.guard(...); AgentOrchestrator.run passes its remaining deadline into executor.execute so a late-stage tool call gets its remaining budget, not a fresh 12 seconds — deadline propagation, not parallel clocks. McpToolExecutor wraps each callTool in RetryExecutor.execute(…, policy.retryable()).
Framework 7’s @Retryable annotation was the alternative; we chose the explicit executor because the retry decision needs the policy’s flag and our error classification, and Chapter 14 measures retry overhead you can’t see inside an annotation. Resilience4j stays on the bench — nothing here exceeds what two JDK classes express.
Failure-injection lab
The whole lab, mapped to observable outcomes:
| Injection | Expected observable |
|---|---|
| model stub: infinite tool proposals | TOOL_BUDGET_EXCEEDED at cap; agent.loop.steps capped |
simulator latency 13000 | TOOL_TIMEOUT at 12s; loop continues or degrades |
simulator flaky 1.0 | 2 attempts, jittered, then UPSTREAM_* — count calls in the simulator log: exactly 2, not 10 |
| 20 concurrent requests with stub latency | 9th+ request gets SATURATED 503; no queue growth |
| 50 KB prompt | PROMPT_TOO_LARGE pre-model; token spend = 0 |
| Ollama down | MODEL_UNAVAILABLE; retrieval still worked (degraded, not dead) |
Security considerations
Bulkheads double as DoS containment; prompt caps double as cost containment. Tool results now pass through a size cap too (10 KB) — a compromised downstream can’t fill the context window with garbage. Model-proposed identifiers were already typed/bound; nothing in this chapter relaxes Chapter 9–10 gates — degraded mode narrows capability, never widens it.
Observability checks
agent.resilience.retries (tool × outcome), agent.resilience.saturated, agent.model.bulkhead.wait — plus the existing loop metrics. A retry-rate alert (rate(agent.resilience.retries[5m]) > threshold) is the early-warning for downstream sickness.
Checkpoint verification checklist
- Timeout hierarchy composes (inner < outer) and is unit-tested as an invariant.
- Retries: transport/5xx only, ≤2, jittered; 4xx never retried.
- Saturation returns 503 fast, not slow.
- Each dependency failure has a named degraded outcome.
Commit message and Git tag
feat(agent-api): bounded orchestration, retry classification, bulkheads, degraded modesgit tag chapter-11-guardrails-resilience
What comes next
Chapter 12 makes all of this visible: OpenTelemetry end to end, Prometheus metrics, Tempo traces, Loki logs, Grafana dashboards — with redaction, because the interesting payloads are exactly the ones you must not export.
Project State Ledger — chapter-11-guardrails-resilience
- New:
RetryExecutor(retryable={5xx,transport}, max 2, jitter),ModelBulkhead(8 permits,SATURATED), deadline propagation into per-tool calls, tool-result cap 10 KB - Degraded modes: RAG-down→labeled ungrounded answer; MCP-down→structured tool errors; model-down→503
- Budgets: 8 model calls / 10 tool calls / 45s deadline (Ch 8), prompt ≤16k, context ≤12k chars
- Decision: Framework-level retry annotation rejected — policy needs the flag + error classes; Resilience4j still unnecessary
- Next:
chapter-12-observability