Secure both boundaries: Keycloak, JWT validation, and tenant authorization
Checkpoint tag: chapter-09-oauth-security — every boundary validates issuer, audience, and scope; the negative security matrix is green; X-Tenant-Id headers are gone from the request path.
What will be built
Keycloak joins the Compose stack with a committed realm import (ops): users priya (operator, tenant acme) and sam (viewer, tenant acme), a second tenant’s operator, client registrations for user login and service-to-service calls, and the four scopes from the blueprint. agent-api and mcp-operations-server become OAuth2 resource servers with explicit audience validation. agent-api acquires client-credentials tokens for MCP calls. Tenant claims replace header trust end to end.
Why it matters
Until now, “tenant” was a header anyone could send — acceptable scaffolding while nothing real was at stake, unacceptable the moment Chapter 10 adds write tools. This chapter is also where the identity design gets decided honestly: the user’s token authenticates the agent boundary; the agent calls MCP with its own service token carrying the resolved tenant and subject as claims the MCP server independently verifies. We do not forward the user JWT — audience mismatch alone would (correctly) reject it, and forwarding user tokens to internal services erodes the boundary between “who asked” and “who is allowed to call.”
Concepts explained
The three independent checks. Issuer (iss = our Keycloak realm) proves where the token came from. Audience (aud contains this service) proves it was minted for us — a token for agent-api is not valid currency at the MCP server. Scope/authorities prove what the bearer may do. OIDC libraries get the first almost free; audience and scope are where real systems under-check.
Delegation vs. impersonation. agent-api calls MCP as itself (client_credentials grant on the agent-api client), carrying tenant_id and sub propagation claims inside the tool-call metadata rather than presenting the user’s token. Trade-off, stated plainly: the MCP server trusts agent-api to have already authenticated the user — acceptable inside one trust domain; a zero-trust variant would use token exchange (RFC 8693) to carry user context cryptographically. That upgrade is an optional lab; the claims-based propagation is the documented baseline.
Deny by construction. Missing scope → 403 at the boundary; tenant mismatch between claim and tool argument → INSUFFICIENT_SCOPE tool result. Neither reaches business logic.
Files added or changed
infra/keycloak/realm-ops.jsoninfra/compose/docker-compose.yml (+ keycloak)agent-api/…/security/SecurityConfig.java, AudienceValidator.java, CallerContext.javamcp-operations-server/…/security/SecurityConfig.kt, TenantPolicy.ktagent-api/…/mcp/McpAuthCustomizer.javatest-support/…/TestTokens.java (jwt fixtures)*/src/test — security matrix testsComplete code
realm-ops.json (excerpted to the material parts — the repo carries the complete import):
{ "realm": "ops", "clients": [ { "clientId": "ops-cli", "publicClient": true, "directAccessGrantsEnabled": true }, { "clientId": "agent-api", "serviceAccountsEnabled": true, "secret": "dev-only-agent-secret", "defaultClientScopes": ["ops:read", "ops:incident:write", "ops:note:write"] }, { "clientId": "mcp-operations-server", "bearerOnly": true, "protocolMappers": [{ "name": "audience", "protocol": "openid-connect", "protocolMapper": "oidc-audience-mapper", "config": {"included.client.audience": "mcp-operations-server"} }] } ], "users": [ {"username": "priya", "enabled": true, "credentials": [{"type": "password", "value": "dev-password"}], "attributes": {"tenant_id": ["acme"]}, "clientRoles": {"ops-cli": ["operator"]}}, {"username": "sam", "enabled": true, "credentials": [{"type": "password", "value": "dev-password"}], "attributes": {"tenant_id": ["acme"]}, "clientRoles": {"ops-cli": ["viewer"]}}, {"username": "rin", "enabled": true, "credentials": [{"type": "password", "value": "dev-password"}], "attributes": {"tenant_id": ["globex"]}, "clientRoles": {"ops-cli": ["operator"]}} ]}The realm also defines client scopes agent:invoke, ops:read, ops:incident:write, ops:note:write, and a tenant_id attribute mapper so every user token carries "tenant_id": "acme" as a claim. Operator role gets all three ops:* scopes; viewer gets ops:read only.
agent-api security:
package in.o612.eng.opsagent.agent.security;
import org.springframework.context.annotation.Bean;import org.springframework.context.annotation.Configuration;import org.springframework.security.config.annotation.web.builders.HttpSecurity;import org.springframework.security.oauth2.core.DelegatingOAuth2TokenValidator;import org.springframework.security.oauth2.jwt.*;import org.springframework.security.web.SecurityFilterChain;
@Configurationpublic class SecurityConfig {
@Bean SecurityFilterChain api(HttpSecurity http) throws Exception { http.authorizeHttpRequests(auth -> auth .requestMatchers("/actuator/health/**").permitAll() .requestMatchers("/api/v1/**").hasAuthority("SCOPE_agent:invoke") .anyRequest().denyAll()) .oauth2ResourceServer(oauth -> oauth.jwt(jwt -> jwt.decoder(jwtDecoder()))); return http.build(); }
@Bean JwtDecoder jwtDecoder( @org.springframework.beans.factory.annotation.Value( "${spring.security.oauth2.resourceserver.jwt.issuer-uri}") String issuer) { var decoder = JwtDecoders.fromIssuerLocation(issuer); var validator = new DelegatingOAuth2TokenValidator<Jwt>( JwtValidators.createDefaultWithIssuer(issuer), new AudienceValidator("agent-api")); ((NimbusJwtDecoder) decoder).setJwtValidator(validator); return decoder; }}package in.o612.eng.opsagent.agent.security;
import org.springframework.security.oauth2.core.OAuth2Error;import org.springframework.security.oauth2.core.OAuth2TokenValidator;import org.springframework.security.oauth2.core.OAuth2TokenValidatorResult;import org.springframework.security.oauth2.jwt.Jwt;
public class AudienceValidator implements OAuth2TokenValidator<Jwt> {
private final String expected;
public AudienceValidator(String expected) { this.expected = expected; }
@Override public OAuth2TokenValidatorResult validate(Jwt jwt) { return jwt.getAudience().contains(expected) ? OAuth2TokenValidatorResult.success() : OAuth2TokenValidatorResult.failure( new OAuth2Error("invalid_token", "missing required audience " + expected, null)); }}CallerContext replaces the header plumbing: a record CallerContext(String tenantId, String userId, Set<String> scopes) extracted once per request from JwtAuthenticationToken — tenantId from the tenant_id claim, scopes from scope. ConversationService and AgentOrchestrator.dispatch now consume CallerContext; the executor injects ctx.tenantId().
On the MCP side (Kotlin): same decoder pattern, plus method-level checks inside the tool methods — the boundary check and the business check are different lines of defense:
package `in`.o612.eng.opsagent.mcp.security
import org.springframework.security.oauth2.jwt.Jwtimport org.springframework.stereotype.Component
@Componentclass TenantPolicy {
fun enforce(jwt: Jwt, requestedTenant: String, requiredScope: String) { val scopes = jwt.getClaimAsString("scope")?.split(" ").orEmpty() require(requiredScope in scopes) { "missing scope $requiredScope" } val tenant = jwt.getClaimAsString("tenant_id") require(tenant == requestedTenant) { "tenant mismatch" } }}Each @McpTool method resolves the current Jwt (via SecurityContextHolder) and calls tenantPolicy.enforce(jwt, tenantArg, "ops:read") — so even a model-crafted tenant argument fails against the caller’s claim. And mcp-operations-server’s decoder uses AudienceValidator("mcp-operations-server"): an aud=agent-api token is rejected even if every other check would pass.
McpAuthCustomizer — agent-api’s outbound identity:
// OAuth2 client-credentials for the agent -> MCP hop.// Registration "agent-mcp": client-id agent-api, grant client_credentials,// token-uri http://keycloak:8085/realms/ops/protocol/openid-connect/token.// A request customizer bean adds "Authorization: Bearer <service token>"// to every MCP transport request via OAuth2AuthorizedClientManager.// Propagation claims (tenant_id, sub) ride as tool-call metadata; the MCP// server validates them against its own JWT's claims — never a forwarded// user token.Verification note: the exact customizer interface name differs across MCP SDK versions (McpSyncHttpClientRequestCustomizer in current 2.0 docs). Confirm against the Spring AI MCP client docs pinned to 2.0.1 before copying — the contract is “a bean that may add headers to outbound MCP HTTP requests,” whatever the interface is called this month.
The security matrix — automated
| Token presented | Expected |
|---|---|
valid user JWT, agent:invoke, tenant acme | 200, tools work |
valid JWT missing agent:invoke | 403 |
| expired JWT | 401 |
| JWT from a different realm/issuer | 401 |
user token presented directly to /mcp | 401 (audience mismatch) |
service token, ops:read, tenant arg globex while claim acme | tool result tenant mismatch |
sam (viewer) proposing create_incident | INSUFFICIENT_SCOPE before any MCP call |
Agent-side tests use spring-security-test’s jwt() post-processor for the controller layer plus TestTokens (test-support) minting real JWTs with a test keypair for decoder-level tests — wrong-issuer/audience/expiry cases need a real JwtDecoder path, not a mocked Authentication.
Failure-injection lab
- Stop Keycloak: boundary requests 401 (decoder can’t fetch JWKS — fail closed, not open).
- Replay a user token at
/mcp: 401 viaAudienceValidator. rin(globex) asks forpayment-gatewaystatus:tenant mismatch— the arg reaches the tool but dies inTenantPolicy.- Strip
ops:readfrom the service client’s scope set and restart the realm: every tool call returnsmissing scope— scope is checked twice (agent dispatch + MCP method) and fails at the first wall now.
Observability checks
security.auth.denials counter labeled boundary + reason (bad_issuer|bad_audience|missing_scope|tenant_mismatch) — the matrix above, expressed as telemetry so a misconfigured realm looks like a denial spike, not silence.
Checkpoint verification checklist
- All seven matrix rows behave as specified in automated tests.
- No
X-Tenant-Id/X-User-Idheaders remain on/api/v1/**. - User JWT is never forwarded to MCP; service token + propagation claims only.
- MCP endpoint rejects anything without
aud=mcp-operations-server.
Commit message and Git tag
feat(security): OIDC resource servers, audience validation, tenant policy on both boundariesgit tag chapter-09-oauth-security
What comes next
Chapter 10 finally arms the write path: create_incident and append_incident_note behind idempotency keys and the human-approval state machine.
Project State Ledger — chapter-09-oauth-security
- IdP: Keycloak realm
opsat:8085; clientsops-cli(public),agent-api(confidential svc acct),mcp-operations-server(bearer-only); userspriya/sam/rin; tenant claimtenant_id - Scopes:
agent:invoke(agent-api),ops:read,ops:incident:write,ops:note:write(mcp) - Identity model: user JWT terminates at agent-api; service token to MCP + propagated
tenant_id/subclaims;TenantPolicyenforces claim==arg per tool call - Env:
KEYCLOAK_URL,AGENT_MCP_CLIENT_SECRET; realm import committed underinfra/keycloak/ - Known limitation: user-context is claims-propagated, not cryptographically delegated (token exchange = optional lab); Keycloak dev-mode secrets are placeholders
- Next:
chapter-10-human-approval