Series overview
Part 27 of 2896% complete
2026-07-17•18 min read

Multi-tenancy patterns

Every chapter so far has assumed Northwind Commerce is the only business running on this platform. This chapter changes that assumption: Northwind is launching “Northwind Platform,” white-labeling the same order/inventory/payment system for other retailers — and every pattern this series has built needs a second look through that lens.

1. Problem the Pattern Solves

Northwind signs its first white-label customer, a smaller retailer called “Fenwick Goods,” who will run their storefront on the exact same order-service, inventory-service, and payment-service Northwind’s own store uses — different product catalog, different customers, different Keycloak realm even, but the same underlying platform. The team’s first instinct — give Fenwick their own complete copy of every service and every database — would technically work, but means every deploy, every schema migration, and every bug fix now has to happen twice, and a third tenant means three times, scaling operational cost linearly with tenant count for infrastructure that’s otherwise identical.

The opposite instinct — one shared database, one set of tables, with a tenant_id column added everywhere — is cheaper to operate but introduces a new, severe risk this series has never had to consider: a single missing WHERE tenant_id = ? clause in a query, or a bug in RefundAuthorizationPolicy (Chapter 26) that doesn’t check tenant boundaries, could let Fenwick’s customer-service rep see or refund a Northwind customer’s order — a data breach between two competing businesses sharing the same infrastructure, not just an internal bug.

Forces in tension:

  • Isolation vs. operational cost. Full isolation (separate deployments, separate databases per tenant) makes a cross-tenant data leak structurally impossible, at the cost of N times the infrastructure and operational overhead for N tenants. Shared infrastructure is dramatically cheaper to operate but makes tenant isolation a code-level responsibility, present in every single query and every authorization check, with real consequences if it’s ever missed.
  • Noisy-neighbor risk vs. resource efficiency. Sharing compute and database resources across tenants is efficient when tenants have complementary usage patterns, but one tenant’s flash sale (a scenario this series has built extensively around for Northwind’s own traffic) can degrade every other tenant’s service if resource isolation isn’t deliberately built in.
  • Per-tenant customization vs. platform consistency. Different tenants may want different business rules (Fenwick’s own refund-approval threshold, say) — supporting this without forking the codebase per tenant requires the platform’s configuration and policy mechanisms (Chapters 5, 26) to become tenant-aware, not just environment-aware.
  • Compliance and data residency vs. shared infrastructure. A future tenant in a different regulatory jurisdiction might require its data to reside in a specific region or even a fully separate infrastructure stack — a constraint the isolation model chosen now needs to be able to accommodate without a full re-architecture later.

2. Core Idea

Three tenancy isolation models, each trading operational cost against isolation guarantees differently:

  • Silo (isolated deployment per tenant): each tenant gets fully separate services, databases, and infrastructure — maximum isolation, maximum cost, scaling linearly with tenant count.
  • Pool (shared deployment, shared or partitioned data): tenants share the same running services, with data separated either by a tenant_id column in shared tables (the riskiest variant) or by separate schemas/databases per tenant behind the same application code (a middle ground) — much better cost efficiency, with isolation enforced in code and configuration rather than by physical separation.
  • Bridge (a mix, chosen per capability): some capabilities run siloed (data with the highest sensitivity or compliance requirement), others pooled (lower-risk, high-volume operations) — the pragmatic middle Northwind actually adopts in this chapter.

Pool: shared, for order-service and inventory-service

order-service, inventory-service

(one shared deployment)

order_schema

WHERE tenant_id = 'northwind'

order_schema

WHERE tenant_id = 'fenwick'

Silo: per-tenant, for payment-service

payment-service

(Northwind instance)

payment_db

Northwind

payment-service

(Fenwick instance)

payment_db

Fenwick

Pool: shared, for order-service and inventory-service

order-service, inventory-service

(one shared deployment)

order_schema

WHERE tenant_id = 'northwind'

order_schema

WHERE tenant_id = 'fenwick'

Silo: per-tenant, for payment-service

payment-service

(Northwind instance)

payment_db

Northwind

payment-service

(Fenwick instance)

payment_db

Fenwick

Northwind’s actual choice, a bridge: payment-service — handling money, the highest compliance and blast-radius sensitivity — runs siloed, one full deployment and database per tenant, accepting the operational cost for the strongest isolation where it matters most. order-service and inventory-service — lower individual risk, higher shared-infrastructure value — run pooled, with tenant isolation enforced at the code and query level (Section 5).

Participants:

  • Tenant identifier — propagated through every request from Chapter 26’s JWT (a tenant_id claim, verified by Keycloak per-realm or per-organization configuration) through every downstream call and every database query.
  • Tenant-aware data access layer — for pooled services, a mechanism (Section 5’s Hibernate filter) that makes it structurally difficult, not just a matter of remembering, to query across tenant boundaries.
  • Per-tenant configuration — Chapter 5’s centralized configuration, extended with a tenant dimension, so Fenwick’s refund-approval threshold (Chapter 26) can differ from Northwind’s own without forking code.

Commonly confused with:

  • Feature flags (Chapter 24). A feature flag controls whether a behavior is active, often per-user or per-cohort; multi-tenancy controls whose data a request can access and which tenant’s configuration applies — related mechanisms (both make runtime decisions based on context), but answering fundamentally different questions. A tenant boundary is a hard isolation requirement; a feature flag is a soft rollout control.
  • Environment separation (staging vs. production). Environments isolate deployment stages of the same tenant’s data; multi-tenancy isolates different tenants’ data within the same environment — Northwind’s staging environment might itself need to simulate multiple tenants to test this chapter’s isolation correctly, a distinct concern from environment separation itself.
  • The anti-corruption layer (Chapter 14). An ACL translates between different domain models at a system boundary; tenant isolation partitions the same domain model’s data by owner. Fenwick and Northwind use the identical Order domain model — the isolation is about data ownership, not conflicting models.

3. When to Use It

Strong indicators:

  • The business itself is expanding to serve multiple distinct customer organizations on shared infrastructure — Northwind Platform’s white-label expansion is exactly this, not a hypothetical future need.
  • A demonstrated or clearly anticipated need to balance per-tenant customization (business rules, thresholds) against maintaining one shared codebase rather than forking per tenant.
  • Cost sensitivity that makes full silo isolation for every capability impractical, requiring a deliberate choice about where isolation is worth its cost and where it isn’t.

Concrete use cases:

  • White-label platforms, as here: e-commerce, logistics, or any B2B2C platform where the platform operator’s own business and its customers’ businesses run on shared infrastructure.
  • B2B SaaS generally: nearly every multi-customer SaaS product is a multi-tenancy problem, whether or not the team frames it that way — customer A’s data must never be visible to customer B, and per-customer configuration (plan tier, feature access) is a near-universal requirement.
  • Healthcare platforms serving multiple provider organizations: tenant isolation here often has direct regulatory weight (HIPAA-adjacent data segregation requirements), making the silo model’s stronger guarantees worth their cost for at least the most sensitive data categories.
  • Government platforms serving multiple agencies or jurisdictions: similar to healthcare, isolation requirements are often externally mandated, not just an internal engineering preference.

Prerequisites:

  • A tenant identifier that’s verifiably present and trustworthy on every request — Chapter 26’s JWT claims are the natural carrier, provided the identity provider (Keycloak) is configured to issue them correctly and they can’t be forged or spoofed by the client.
  • An explicit, reviewed decision about which capabilities need silo-level isolation versus which can safely pool — made deliberately, per capability, based on data sensitivity and blast radius (Section 5), not defaulted uniformly in either direction.
  • Tenant-aware testing (Section 7) — a test suite that never verifies cross-tenant isolation is a test suite that can’t actually catch the exact class of bug Section 1 warns about.

4. When Not to Use It

  • A genuinely single-tenant system with no plan to serve multiple independent customer organizations. Building multi-tenancy infrastructure speculatively, before a second tenant is a real, committed business plan, adds real complexity (tenant-aware queries, tenant-scoped configuration) for a need that may never materialize.
  • Every capability defaulting to full silo isolation “to be safe.” This is Section 1’s first instinct, and it’s not wrong exactly — it’s expensive at a scale that may not be justified for every capability; evaluate per capability, as Section 5’s bridge model does, rather than applying maximum isolation uniformly regardless of actual sensitivity.
  • Every capability defaulting to shared, single-table pooling “to be simple.” The opposite mistake — treating multi-tenancy as just another WHERE clause everywhere, including for the most sensitive data (payment records) — accepts more risk than the operational savings justify for that specific data.
  • Overengineering signal: building a fully generic, configurable-per-capability tenancy framework before Northwind has even a second real tenant to validate the design against. Fenwick Goods is real, concrete evidence to design against — building for a hypothetical fifth or tenth tenant’s unknown requirements risks over-engineering for needs that haven’t been validated.

5. Implementation Example

Tenant identifier propagation, extending Chapter 26’s JWT claims:

Example JWT claims, Keycloak-issued
{
"sub": "customer-abc123",
"tenant_id": "fenwick-goods",
"realm_access": { "roles": ["customer"] }
}
order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/TenantContext.kt
package `in`.o612.eng.northwind.order.internal
import org.springframework.security.core.context.SecurityContextHolder
import org.springframework.security.oauth2.jwt.Jwt
object TenantContext {
fun currentTenantId(): String {
val jwt = SecurityContextHolder.getContext().authentication.principal as Jwt
return jwt.getClaimAsString("tenant_id") ?: error("Request missing required tenant_id claim")
}
}

Pooled tenancy for order-service, using Hibernate’s @Filter to make tenant isolation structural rather than a manually-remembered WHERE clause on every query — the direct fix for Section 1’s “missing WHERE clause” risk:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/Order.kt
package `in`.o612.eng.northwind.order.internal
import org.hibernate.annotations.Filter
import org.hibernate.annotations.FilterDef
import org.hibernate.annotations.ParamDef
import jakarta.persistence.*
@Entity
@Table(name = "orders")
@FilterDef(name = "tenantFilter", parameters = [ParamDef(name = "tenantId", type = String::class)])
@Filter(name = "tenantFilter", condition = "tenant_id = :tenantId")
class Order(
@Id val id: java.util.UUID,
@Column(name = "tenant_id") val tenantId: String,
var status: String,
// ... other fields, unchanged since Chapter 1
)
order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/TenantFilterAspect.kt
package `in`.o612.eng.northwind.order.internal
import jakarta.persistence.EntityManager
import org.aspectj.lang.annotation.Aspect
import org.aspectj.lang.annotation.Before
import org.hibernate.Session
import org.springframework.stereotype.Component
/** Enables the Hibernate tenant filter on EVERY repository call,
* automatically — no individual query author can forget it, because
* it's applied at the session level, not per-query. */
@Aspect
@Component
class TenantFilterAspect(private val entityManager: EntityManager) {
@Before("execution(* in.o612.eng.northwind.order.internal.*Repository.*(..))")
fun enableTenantFilter() {
entityManager.unwrap(Session::class.java)
.enableFilter("tenantFilter")
.setParameter("tenantId", TenantContext.currentTenantId())
}
}

This is the structural difference that matters: instead of every developer remembering to add AND tenant_id = ? to every query they write (Section 1’s actual, realistic failure mode), the filter is enabled automatically for every repository call, at the session level — a developer would have to actively work around it to accidentally query across tenants, rather than simply forgetting one line.

Silo tenancy for payment-service, one full deployment per tenant, using Chapter 5’s service discovery to route correctly:

k8s/payment-service-fenwick.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-service-fenwick
labels: { app: payment-service, tenant: fenwick-goods }
spec:
template:
spec:
containers:
- name: payment-service
image: northwind/payment-service:3.2.0 # identical image — same code, different deployment and data
env:
- name: SPRING_DATASOURCE_URL
value: "jdbc:postgresql://payment-db-fenwick.default.svc.cluster.local/payment_db"
api-gateway/src/main/kotlin/in/o612/eng/northwind/gateway/TenantAwareRoutingFilter.kt
package `in`.o612.eng.northwind.gateway
import org.springframework.cloud.gateway.filter.GatewayFilter
import org.springframework.stereotype.Component
@Component
class TenantAwareRoutingFilter : GatewayFilter {
override fun filter(exchange: org.springframework.web.server.ServerWebExchange, chain: org.springframework.cloud.gateway.filter.GatewayFilterChain) =
// Route payment-service calls to the tenant-specific deployment
// based on the JWT's tenant_id claim — order-service and
// inventory-service calls route to the one shared pooled deployment.
chain.filter(exchange) // routing logic resolves target host by tenant_id for silo'd services
}

payment-service’s own code needs zero tenant-awareness at all — it’s an identical image, deployed once per tenant, exactly as Chapter 3’s database-per-service pattern already established for a different reason (operational isolation rather than tenant isolation, but the same underlying mechanism serves both purposes here).

Per-tenant configuration, extending Chapter 5’s Spring Cloud Config and Chapter 26’s ABAC policy with a tenant dimension:

platform-config/payment-service-fenwick-goods.yml
northwind:
refund-policy:
rep-approval-limit: 50.00 # Fenwick's own, lower threshold — different business, different risk tolerance
payment-service/src/main/kotlin/in/o612/eng/northwind/payment/internal/RefundAuthorizationPolicy.kt (revised)
package `in`.o612.eng.northwind.payment.internal
import org.springframework.boot.context.properties.ConfigurationProperties
import org.springframework.cloud.context.config.annotation.RefreshScope
import org.springframework.stereotype.Component
import java.math.BigDecimal
@RefreshScope
@Component
@ConfigurationProperties(prefix = "northwind.refund-policy")
class RefundPolicyProperties(var repApprovalLimit: BigDecimal = BigDecimal("100.00"))

Since payment-service is siloed per tenant (each with its own deployment and its own application.yml fetched from its own tenant-specific config profile), Fenwick’s threshold simply is different — no runtime tenant-branching logic needed inside the code at all, another direct benefit of choosing silo isolation for this specific, high-sensitivity capability.

6. Step-by-Step Flow

payment-service-fenwick (siloed)order-service (pooled)API GatewayFenwick customerpayment-service-fenwick (siloed)order-service (pooled)API GatewayFenwick customerPOST /orders (JWT: tenant_id=fenwick-goods)forward (pooled deployment, shared with Northwind)TenantFilterAspect enables Hibernate filter: tenant_id='fenwick-goods'query orders table -- automatically scoped, cannot see Northwind's rowscall payment-service (routed to Fenwick's OWN deployment/database)process using Fenwick's own config (rep-approval-limit: 50.00)response202 Accepted
payment-service-fenwick (siloed)order-service (pooled)API GatewayFenwick customerpayment-service-fenwick (siloed)order-service (pooled)API GatewayFenwick customerPOST /orders (JWT: tenant_id=fenwick-goods)forward (pooled deployment, shared with Northwind)TenantFilterAspect enables Hibernate filter: tenant_id='fenwick-goods'query orders table -- automatically scoped, cannot see Northwind's rowscall payment-service (routed to Fenwick's OWN deployment/database)process using Fenwick's own config (rep-approval-limit: 50.00)response202 Accepted
  1. Client action. A Fenwick Goods customer places an order — from their point of view, indistinguishable from placing an order on any other platform.
  2. API request. The gateway routes based on the JWT’s tenant_id claim — to the shared pooled deployment for order-service, or to Fenwick’s own siloed payment-service deployment.
  3. Service behavior. order-service’s TenantFilterAspect automatically scopes every database query to Fenwick’s tenant ID, with no per-query code needed to remember this.
  4. Database interaction. The Hibernate filter makes Fenwick’s rows in the shared orders table structurally invisible to any query running under a different tenant context — the same table, same schema, same instance as Northwind’s own orders, safely partitioned.
  5. Inter-service communication. The call to payment-service is routed, at the gateway, to Fenwick’s own siloed deployment and database — a fundamentally different, stronger isolation guarantee for the highest-sensitivity data.
  6. Error or failure handling. A bug in order-service’s query logic that somehow bypassed the tenant filter would still be contained to the pooled, lower-sensitivity data — it could never expose payment data, which never shares infrastructure across tenants at all, by design.
  7. Observability signals. Track and alert on any query executed without an active tenant filter (a defensive check TenantFilterAspect can itself log) — this is the specific signal that would catch a bug attempting to bypass tenant isolation before it becomes a real incident.
  8. Final response/outcome. Fenwick’s order is processed correctly, using Fenwick’s own business rules (the $50 refund threshold) where relevant, with data isolation enforced structurally for the pooled services and physically for the siloed one — the bridge model’s actual payoff.

7. Production Concerns

  • Timeouts, retries, idempotency. Unaffected directly by tenancy model, though note that a siloed service’s per-tenant deployment means per-tenant resilience configuration (Chapter 16) is also possible — Fenwick’s payment-service could have different circuit-breaker thresholds tuned to their specific traffic pattern, independent of Northwind’s own.
  • Data consistency. Cross-tenant data consistency is never a requirement by definition — tenants are, by design, isolated from each other’s transactions and consistency boundaries entirely, simplifying this specific concern even as it complicates data access.
  • API versioning. A pooled service’s API is shared across all tenants — a breaking change affects every tenant simultaneously, which argues for even more conservative versioning discipline (Chapter 2) than a single-tenant service would need, since a mistake has a wider, harder-to-coordinate blast radius.
  • Authentication and service-to-service trust. The tenant_id claim must come from a source the platform trusts completely — verify Keycloak issues it correctly per tenant/realm and that it cannot be client-supplied or spoofed; a forged tenant ID would defeat every isolation mechanism this chapter builds.
  • Logging, metrics, tracing, audit trails. Tag every log line and every trace span with tenant_id — beyond debugging, this is often a direct compliance and billing requirement (usage-based billing per tenant needs exactly this data), and it’s the dimension along which a security investigation into a suspected isolation breach would need to search.
  • Kubernetes deployment, health probes, autoscaling. Siloed services scale their resource requests (Chapter 25) per tenant deployment independently — Fenwick’s smaller transaction volume means their payment-service deployment can run with a much smaller resource allocation than Northwind’s own, a direct cost benefit of the silo model for a lower-volume tenant.
  • Testing strategy. Write explicit cross-tenant isolation tests — attempt to query or access tenant A’s data while authenticated as tenant B, and assert it fails or returns nothing — as a required, standing part of the test suite, not an afterthought; this is the test that actually verifies Section 1’s core risk is mitigated, not assumed away.
  • Migration strategy. Onboard each new tenant by explicitly deciding, capability by capability, whether it needs its own siloed deployment or can join the pooled tier — Northwind’s own bridge model (payment siloed, order/inventory pooled) is itself a decision that should be revisited as tenant count and diversity grow, not treated as permanently fixed.

8. Common Mistakes

  1. Relying on manually-added WHERE tenant_id = ? clauses instead of a structural enforcement mechanism. Trusting every developer to remember this on every query, forever, is exactly Section 1’s realistic failure mode — one missed clause is a cross-tenant data breach. Fix: enforce tenant scoping structurally (Hibernate’s @Filter plus an aspect, as Section 5 shows), so it can’t be silently omitted.
  2. Choosing one isolation model uniformly for every capability, in either direction. Full silo for everything is prohibitively expensive at scale; full pooling for everything (including payment data) accepts more risk than the savings justify. Fix: evaluate isolation model per capability based on data sensitivity and blast radius, as Northwind’s bridge model does.
  3. Trusting a client-supplied tenant identifier instead of one verified by the identity provider. Accepting a tenant_id from a request header or query parameter the client controls, rather than a JWT claim issued and signed by Keycloak, lets any client claim to be any tenant. Fix: the tenant identifier must come from a cryptographically verified source, exactly like every other identity claim since Chapter 26.
  4. No explicit cross-tenant isolation testing. Testing only “tenant A’s own data is correctly returned for tenant A” without ever testing “tenant A absolutely cannot see tenant B’s data” leaves the actual risk this chapter addresses unverified. Fix: write and maintain explicit negative tests attempting cross-tenant access, as Section 7 requires.
  5. Building a generic, fully configurable tenancy framework before a second real tenant exists to validate it against. Speculatively designing for hypothetical future tenants’ unknown requirements, rather than Fenwick’s actual, concrete needs, risks over-engineering for the wrong things. Fix: design against the real tenant in front of you; generalize only once a second and third real tenant’s requirements reveal what actually needs to be configurable.
  6. Forgetting that per-tenant configuration changes (like a refund threshold) are business decisions, not just technical configuration. Treating RefundPolicyProperties’ per-tenant values as purely an engineering concern, without the business rigor a financial threshold deserves. Fix: route per-tenant business-rule configuration through the same review process as any other business-critical decision, not just a config file merge.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Platform is expanding to serve multiple distinct customer organizations on shared infrastructureYesThis is the core problem multi-tenancy patterns exist to solve—
A specific capability handles highly sensitive data (payments, health records)Silo that capabilityStrongest isolation guarantee justifies its operational cost for the highest-risk data—
A specific capability is lower-risk, high-volume, benefits from shared infrastructure economicsPool that capability, with structural tenant enforcementCost-efficient, provided isolation is enforced structurally, not by convention—
Single-tenant system, no committed plan for a second tenantNoBuilding this speculatively adds real complexity for an unvalidated needRevisit when a second tenant is a real, committed plan
Tenant identifier would come from a client-controlled, unverified sourceFix that firstAny isolation mechanism is defeated if the tenant identifier itself can be spoofedVerify tenant identity through the trusted identity provider (Chapter 26)

10. Hands-On Exercise

Extend it: onboard a third tenant, “Ridgeline Outdoors,” and decide explicitly, with justification, whether inventory-service should remain pooled for them or move to silo — what evidence about Ridgeline’s specific business (transaction volume, compliance requirements) would change your answer from what Northwind chose for Fenwick?

Simulate a failure: temporarily disable TenantFilterAspect in a test environment and confirm that a query, run under Fenwick’s tenant context, can now see Northwind’s order data — then re-enable it and confirm the isolation holds. This concretely demonstrates why structural enforcement, not convention, is the actual protection.

Decision question, with justification required: Fenwick Goods requests a custom checkout field (a loyalty-program number) that Northwind’s own storefront doesn’t use and doesn’t want cluttering its own schema. Should this be a nullable, generally-available column on the pooled orders table, a separate per-tenant extension table, or a reason to move Fenwick’s order-service usage to a siloed deployment instead? Weigh the trade-offs from Section 1 and Section 4.

11. Key Takeaways

  • Multi-tenancy isolation exists on a spectrum from silo (fully separate deployments and data, strongest isolation, highest cost) to pool (shared infrastructure, code-enforced isolation, lowest cost) — most real platforms, like Northwind’s, land on a deliberate bridge, choosing per capability based on data sensitivity.
  • Enforce tenant isolation structurally, not by convention — a manually-remembered WHERE tenant_id = ? clause is exactly the kind of single point of failure that turns into a cross-tenant data breach the moment one developer forgets it once.
  • The tenant identifier must come from a cryptographically verified source (Chapter 26’s JWT claims, correctly configured per tenant in the identity provider), never from anything the client itself controls.
  • Reserve silo isolation for the highest-sensitivity, highest-blast-radius capabilities (payment data, here) where its operational cost is clearly justified; pool lower-risk, high-volume capabilities where shared infrastructure’s cost efficiency matters more.
  • Explicit, standing cross-tenant isolation tests — actively attempting and expecting to fail at accessing another tenant’s data — are the only way to verify this pattern’s core promise is actually being kept, not just assumed.
  • Per-tenant configuration (business rules, thresholds) should route through the same review rigor as any other business-critical decision, not be treated as a purely technical configuration change.
  • Design against your real, current tenants’ actual needs rather than a hypothetical future tenant’s unknown requirements — Fenwick Goods is concrete evidence to build against; a fifth or tenth imagined tenant is not.
Spring BootKotlinMicroservices

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind