Series overview
Part 13 of 1776% complete
2026-09-03•4 min read

End-to-end observability: one trace from HTTP to tool call

Checkpoint tag: chapter-12-observability — one Tempo trace links HTTP → agent loop → retrieval → model → MCP → simulator, with zero sensitive payload content exported.

What will be built

The observability stack joins Compose: OTel Collector receiving OTLP from all four modules, Prometheus scraping, Tempo storing traces, Loki aggregating logs, Grafana wired to all three. On the app side: explicit Observation-based instrumentation on the agent loop, retrieval, tool dispatch, and approval decisions; a correlation_id carried in baggage; audit events as a separate sink; and redaction tests that prove prompts and tool arguments never reach telemetry.

Why it matters

“Why did the agent do that” is the question this system exists to answer, and it is only answerable with a trace that crosses process boundaries: HTTP request → authz → retrieval → model → policy decision → MCP call → simulator row insert. Logging alone gives you five disconnected stories; metrics alone give you aggregates. The trace is the connective tissue — and the discipline is in what it doesn’t carry, because a span attribute containing the user prompt is a data leak with a query UI.

Concepts explained

Signals, separated. Metrics are aggregates (mcp.tool.calls{tool,outcome}) — low cardinality or Prometheus becomes a bill. Traces are per-request causality — spans carry timing and references (tool.name=get_service_status), never payloads. Logs are structured context with correlation_id/tenant_id in MDC. Audit events are the fourth signal — compliance facts (APPROVED by priya, args_hash=…) going to the append-only table + a dedicated logger, because “who approved what” must survive a log-retention purge.

OTLP through a Collector. Apps export to the Collector; the Collector batches, redacts (a transform processor drops any attribute matching prompt|args|content), and fans out to Tempo/Prometheus. Redaction in both places: at the app (don’t emit) and at the collector (drop if emitted anyway) — defense in depth for telemetry.

Baggage for correlation. correlation_id propagates via W3C baggage — present in MDC for logs and as a span baggage entry, but not as a metric label, because UUIDs are the cardinality reaper.

Files added or changed

infra/observability/{otel-collector.yml, prometheus.yml, tempo.yml, loki.yml, grafana-datasources.yml, dashboards/agent.json}
infra/compose/docker-compose.yml (full stack now)
agent-api/…/observability/{AgentObservability, RedactionFilter}.java
mcp-operations-server + operations-simulator: OTLP config
agent-api/src/test/.../TelemetryRedactionTest.java

Complete code (the load-bearing parts)

infra/observability/otel-collector.yml
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
batch: {}
redact-payloads:
transform:
trace_statements:
- context: span
statements:
- delete_key(attributes, "http.request.body")
- delete_key(attributes, "tool.args")
- delete_key(attributes, "prompt.content")
exporters:
otlp/tempo:
endpoint: tempo:4317
tls: { insecure: true }
prometheus:
endpoint: 0.0.0.0:9464
loki:
endpoint: http://loki:3100/loki/api/v1/push
service:
pipelines:
traces: { receivers: [otlp], processors: [redact-payloads, batch], exporters: [otlp/tempo] }
metrics: { receivers: [otlp], processors: [batch], exporters: [prometheus] }
logs: { receivers: [otlp], processors: [redact-payloads, batch], exporters: [loki] }
agent-api/src/main/resources/application.yml (additions)
management:
otlp:
tracing: { endpoint: ${OTEL_ENDPOINT:http://localhost:4318}/v1/traces }
metrics:
export: { endpoint: ${OTEL_ENDPOINT:http://localhost:4318}/v1/metrics }
logging: { endpoint: ${OTEL_ENDPOINT:http://localhost:4318}/v1/logs }
tracing:
sampling: { probability: 1.0 } # dev; production example uses 0.1 + tail-sampling note
observations:
key-values:
application: ${spring.application.name}

AgentObservability — the spans worth naming are the decisions, not the framework plumbing:

agent-api/src/main/java/in/o612/eng/opsagent/agent/observability/AgentObservability.java
// Wrapped around orchestrator steps:
// Observation.start("agent.model.call", registry) -> low-cardinality tags only
// Observation.start("agent.tool.dispatch") -> tag tool + outcome, never args
// Observation.start("agent.approval.decision") -> tag outcome
// correlation_id enters MDC on request entry; baggage propagates it into
// MCP calls; every emitted event carries it in structured log output.

TelemetryRedactionTest — the test that makes the promise checkable:

agent-api/src/test/java/in/o612/eng/opsagent/agent/observability/TelemetryRedactionTest.java
// Asserts on captured observations/meter registry:
// - no span attribute key matches /prompt|args|content|token|secret/i
// - metric tag sets are a subset of the allowlist {tool,outcome,provider,risk}
// - audit logger output contains args_hash but never args_json
// - a deliberately poisoned tool arg ("Bearer xyz") never appears in any
// captured telemetry string

Grafana dashboards and SLO seeds

dashboards/agent.json ships four panels: request rate + error rate per boundary; agent.loop.steps histogram heatmap; tool-call latency by tool; approval outcome funnel. SLO seeds, marked as starting points: availability: 99.5% of POST /messages return non-5xx; approval latency p95 < 60s (human-in-the-loop, so generous); tool error rate < 5%.

Failure-injection lab

  1. Simulator outage → follow one trace: agent.tool.dispatch span shows TOOL_FAILED, simulator span absent (connection refused never created a server span — learn to read that absence).
  2. Wrong-audience token → 401 span with security.denial event attribute bad_audience — security failures are first-class trace content.
  3. Sampling check: flip probability to 0.0, confirm dashboards still update (metrics ≠ traces — losing traces must not lose the counters).

Security considerations

The redaction contract: prompts, retrieved content, tool args/results, tokens, and secrets never leave the app in telemetry; tenant_id/user_id appear in logs and audit but not in metric labels or span attributes (they’re join keys for a human querying, not labels for aggregation). The local-debug profile that does capture prompts prints a banner, writes to a file-not-OTLP sink, and is excluded from the container image’s default env.

Checkpoint verification checklist

  • One Tempo trace spans agent→MCP→simulator with correlation_id baggage visible.
  • TelemetryRedactionTest green; collector transform verified by an injected marker attribute.
  • Metric labels confined to the allowlist.
  • Audit rows and logs both carry correlation IDs.

Commit message and Git tag

feat(obs): OTLP pipeline, decision spans, redaction contract, Grafana dashboards

git tag chapter-12-observability

What comes next

Chapter 13 builds the evaluation suite — the versioned dataset and graders that tell you whether a prompt, model, or retrieval change actually made things better or just different.

Project State Ledger — chapter-12-observability

  • Stack: OTel Collector :4317/:4318 → Tempo/Prometheus/Loki → Grafana :3000
  • Spans: agent.model.call, agent.retrieval, agent.tool.dispatch, agent.approval.decision, MCP + simulator server spans
  • Redaction: app-side (never emit) + collector transform (drop on receipt); tested
  • Sampling: 1.0 dev / 0.1 prod example; metrics survive trace sampling
  • Config: OTEL_ENDPOINT; local-debug prompt-capture profile exists but is file-sink-only
  • Next: chapter-13-evaluation
ObservabilityPrometheusSpring BootTesting

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind