Database per service
Chapter 2 split Northwind’s monolith into three deployables — order-service, inventory-service, payment-service — but left all three still connecting to order_schema, inventory_schema, and payment_schema on the one shared Postgres instance from Chapter 1. This chapter finishes the decomposition by giving each service its own database instance.
1. Problem the Pattern Solves
Three months after the Chapter 2 split, Northwind’s platform team gets paged at 2 a.m.: the shared Postgres instance hit its connection limit. inventory-service’s nightly reconciliation job opened enough connections to starve order-service’s connection pool during a flash sale. Nobody on the order team could have prevented this — they don’t own the instance, don’t control its max_connections, and didn’t even know the reconciliation job existed until the incident review.
A week later, a different problem: the payment team wants to upgrade to Postgres 16 to use a feature that helps with their ledger queries, but can’t, because order and inventory are still validating their code against Postgres 14 and a shared major-version upgrade requires all three teams to coordinate a migration window. The database, not the code, has become the last remaining coupling point between three otherwise-independent services.
Forces in tension:
- Isolation vs. operational cost. One shared instance is cheaper to run and back up than three, but every shared resource — connection limits, disk I/O, a major-version upgrade, a maintenance window — becomes a cross-team negotiation.
- Consistency vs. autonomy. A shared instance makes an ad-hoc cross-schema SQL join possible, even when it’s against policy — a temptation that database-per-service removes structurally rather than by convention.
- Blast radius. A shared instance means a runaway query or a full disk in one service’s schema can degrade or take down every service on that instance. Separate instances contain the failure to the service that caused it.
- Cost and complexity. Three managed Postgres instances (or three self-hosted clusters) cost more, in both money and operational surface — backup policies, monitoring, patching — than one. This is a real, recurring cost that only pays for itself once the coupling problems above are actually biting.
2. Core Idea
Database-per-service means each service owns a database (or database cluster) that no other service accesses directly, ever — not through a shared connection pool, not through a read replica another team queries, not through a “just this once” cross-schema join. All access to a service’s data goes through that service’s API.
Intent: make data ownership match deployable ownership exactly, so a service’s team has full control over its schema, its indexing strategy, its scaling profile, and its upgrade cadence, with zero coordination required from other teams for any of it.
Participants: unchanged from Chapter 2 — the same three services, the same public REST contracts. What moves is only the connection string and the deployment topology behind each service.
Commonly confused with:
- Schema-per-service on a shared instance (Chapter 2’s state). This is a legitimate, often sufficient intermediate step — it gets you logical isolation and prevents accidental cross-schema queries via application-level discipline, but it does not give you independent scaling, independent upgrade cadence, or blast-radius isolation. Don’t skip straight past it if your team size and load don’t yet justify full database-per-service; it’s a real destination, not just a waypoint.
- A shared database with per-service database users and row-level security. This narrows who can touch what data but doesn’t solve the shared-resource contention problem (connections, I/O, upgrade windows) that motivated this chapter’s 2 a.m. page.
- Data replication for read scaling (e.g., a read replica another service queries directly). A read replica queried directly by another service is not database-per-service — it’s a second undocumented consumer of your schema, coupled to your internal table structure exactly as tightly as a direct write would be. If another service needs your data, expose it through your API or through an explicit, versioned event (Publish-Subscribe, a later chapter) — never through a replica connection string handed to another team.
3. When to Use It
Strong indicators:
- Two or more services on a shared instance have measurably competed for connections, I/O, or maintenance windows — a real incident, not a hypothetical one.
- Services have genuinely different operational requirements:
payment-service’s data needs encryption-at-rest and a stricter backup/retention policy thaninventory-service’s. - A service’s data volume or query pattern needs different infrastructure —
inventory-service’s reconciliation batch job needs more I/O bandwidth thanorder-service’s OLTP traffic pattern.
Concrete use cases:
- Payments platforms: regulatory requirements (PCI DSS) often mandate that cardholder-data-adjacent systems live in isolated infrastructure with tighter access control than the rest of the platform — database-per-service is close to a hard requirement here, not just good practice.
- Healthcare: patient records typically require audit logging, encryption, and retention policies distinct from, say, an appointment-reminder service — sharing an instance would force the stricter policy onto data that doesn’t need it, or worse, the laxer policy onto data that does.
- SaaS multi-tenant platforms: a
billingservice handling financial transactions often needs stronger consistency guarantees and slower, more careful migrations than ausage-analyticsservice ingesting high-volume event data — different velocity, different risk tolerance. - Logistics: a
route-optimizationservice running heavy batch computations benefits from an instance sized and tuned for throughput, entirely separate from ashipment-trackingservice tuned for low-latency point reads.
Prerequisites:
- The service boundary itself must already be clean (Chapter 2’s precondition) — migrating a database out from under a service whose queries still secretly depend on another service’s tables will surface every one of those dependencies as an outage.
- A migration plan that doesn’t require downtime for the whole platform — the technique in Section 5 (dual-write with a cutover) is the standard approach.
- Budget and operational capacity for N databases instead of one — someone has to patch, back up, and monitor each one.
4. When Not to Use It
- Team or platform is too small to operate multiple database instances well. A two-person platform team keeping one Postgres instance healthy is a reasonable job; keeping five healthy, each with its own backup verification and patching schedule, is a different, larger job that doesn’t pay for itself below a certain service count.
- No evidence of actual resource contention or coupling pain. Splitting databases pre-emptively, “because microservices best practices say so,” multiplies infrastructure cost without a problem to solve — schema-per-service on one instance may be the right permanent state for a smaller deployment.
- Cost and risk: each additional database instance is another single point of failure to monitor, another backup to verify restorable (not just “backup completed,” but “we’ve tested restoring from it”), and another source of connection-string secrets to manage. A five-service platform with five under-monitored databases is often less reliable than three well-monitored schemas on one carefully managed instance.
- Overengineering signal: provisioning a dedicated database instance for a service that stores less than a few gigabytes and takes single-digit requests per second. The isolation benefit is real but the fixed operational cost per instance doesn’t scale down — right-size the pattern to the service’s actual footprint.
5. Implementation Example
This chapter’s code is deliberately about the migration, not new business logic — InventoryService’s domain code from Chapter 1 and Chapter 2 doesn’t change at all. What changes is the connection configuration and the cutover procedure, which is the part most tutorials skip and most incidents happen in.
Docker Compose, three instances instead of one:
services: order-db: image: postgres:16 environment: { POSTGRES_DB: order_db, POSTGRES_USER: order_service, POSTGRES_PASSWORD: ${ORDER_DB_PASSWORD} } ports: ["5433:5432"] inventory-db: image: postgres:16 environment: { POSTGRES_DB: inventory_db, POSTGRES_USER: inventory_service, POSTGRES_PASSWORD: ${INVENTORY_DB_PASSWORD} } ports: ["5434:5432"] volumes: ["inventory-data:/var/lib/postgresql/data"] # larger disk profile for reconciliation batches payment-db: image: postgres:16 environment: { POSTGRES_DB: payment_db, POSTGRES_USER: payment_service, POSTGRES_PASSWORD: ${PAYMENT_DB_PASSWORD} } ports: ["5435:5432"]
volumes: inventory-data:The cutover, without downtime, for one service (inventory-service, chosen first because — as in Chapter 2’s migration order — it’s under the least cross-service query pressure):
package `in`.o612.eng.northwind.inventory.internal
import org.springframework.stereotype.Repository
/** Temporary migration shim, active only during the Phase 1 dual-write * window described in the runbook. Removed entirely once cutover completes * — this class should never outlive the migration it exists for. */@Repositoryinternal class DualWriteStockRepository( private val primary: StockRepository, // still points at the shared instance private val shadow: StockRepository, // points at the new inventory-db private val migrationMetrics: MigrationMetrics,) : StockRepository by primary {
override fun decrement(sku: String, quantity: Int) { primary.decrement(sku, quantity) runCatching { shadow.decrement(sku, quantity) } .onFailure { migrationMetrics.recordShadowWriteFailure(sku, it) } }}The shadow write is deliberately best-effort and non-blocking on failure — a shadow-write failure must never fail the real request, since the shadow database isn’t authoritative yet. migrationMetrics.recordShadowWriteFailure feeds a dashboard the team watches during the dual-write window; a rising failure rate blocks the cutover in Phase 3 until it’s understood.
Configuration change at cutover — this is the entire “migration” from the application’s point of view once dual-writing has been running cleanly:
spring: datasource: url: jdbc:postgresql://shared-db.internal:5432/northwind?currentSchema=inventory_schema username: inventory_servicespring: datasource: url: jdbc:postgresql://inventory-db.internal:5432/inventory_db username: inventory_serviceNo repository code changes at cutover — StockRepository’s JPA queries never referenced the schema name or any table outside inventory_schema, a discipline that goes all the way back to Chapter 1’s enforced module boundary. That discipline is what makes this cutover a configuration change instead of a code change.
Verification test, run continuously during the dual-write window, comparing row counts and checksums between old and new:
package `in`.o612.eng.northwind.inventory.internal
import org.junit.jupiter.api.Testimport org.assertj.core.api.Assertions.assertThat
class MigrationReconciliationTest {
@Test fun `shadow database stock levels match primary for every known SKU`() { val primarySnapshot = primaryJdbc.queryForList("SELECT sku, quantity FROM stock ORDER BY sku") val shadowSnapshot = shadowJdbc.queryForList("SELECT sku, quantity FROM stock ORDER BY sku")
assertThat(shadowSnapshot).isEqualTo(primarySnapshot) }}Production note. This reconciliation check should run as a scheduled job against real production data during the dual-write window, not only as a one-off test — drift here is exactly the signal that tells you whether it’s safe to cut over.
6. Step-by-Step Flow
Walking the migration itself as the “event” for this chapter, rather than a runtime request:
- Client action (of the migration): the platform team declares
inventory-serviceready for extraction, based on Section 3’s indicators (its reconciliation job’s I/O contention withorder-service). - API request equivalent:
DualWriteStockRepositoryis deployed, writing to both the shared instance and the newinventory-dbon every request, with shadow-write failures logged but not surfaced to the caller. - Service behavior: the reconciliation job in Section 5 runs nightly during the dual-write window, comparing snapshots.
- Database interaction: historical data is backfilled once via
pg_dump/pg_restore(or a logical replication slot, for larger datasets where a full dump-and-restore window isn’t acceptable), then kept current by the dual write. - Inter-service communication: unchanged —
order-servicestill callsinventory-service’s same REST contract from Chapter 2 throughout the entire migration. This is the payoff of database-per-service being a strictly internal concern: no other service, and no client, observes the migration happening. - Error or failure handling: a shadow-write failure rate above a threshold (say, 0.1%) pauses the migration and pages the on-call engineer — cutting over on top of unexplained drift is how migrations turn into incidents.
- Observability signals:
migration.shadow_write.failures,migration.reconciliation.drift_rows, and connection-pool utilization on both instances, watched side by side for the length of the dual-write window (Northwind ran this for two weeks before cutting over). - Final response/outcome: once reconciliation shows zero drift for a full billing cycle, the configuration change in Section 5 ships, the old schema is set to read-only for a rollback window, then dropped.
7. Production Concerns
- Timeouts, retries, idempotency. The dual-write shim must be idempotent-safe on retry (a retried
decrementcall shouldn’t double-decrement the shadow database) — reuse the same idempotency key discipline this series’ Outbox chapter formalizes. - Data consistency and transaction boundaries. The shadow write is not transactional with the primary write — accept that as a defined, monitored gap during migration (that’s what the reconciliation job is for), not a defect to eliminate before starting.
- Database-per-service and schema ownership. Once cut over,
inventory-service’s team can now choose its own Postgres version, connection pool size, backup schedule, and vacuum tuning without asking anyone — the entire point of this chapter. - API versioning and backward compatibility. Not directly affected — the REST contracts between services are untouched by a database migration, which is exactly the isolation this pattern is supposed to provide. If a database migration ever requires an API change, that’s a sign the service boundary and the data boundary weren’t as separate as assumed.
- Authentication, authorization, and service-to-service trust. Each new database instance needs its own credentials, ideally issued per-service from a secrets manager (Kubernetes Secrets backed by Vault or a cloud KMS) rather than a shared connection string checked into any config file.
- Logging, metrics, tracing, audit trails. Track the migration itself as an auditable event — who approved the cutover, when, based on what reconciliation numbers — separately from the application’s normal business audit trail.
- Kubernetes deployment, health probes, secrets, configuration. The readiness probe should verify connectivity to whichever database is currently primary, and the cutover should be a config/secret change (a new
ConfigMap/Secretversion) rather than a code deploy, so it can be rolled back independently of any application release. - Testing strategy. Beyond the reconciliation test shown above, run a full disaster-recovery drill against the new instance — restore its backup into a fresh instance and run the integration test suite against it — before considering the migration complete. An unverified backup is not a backup.
- Migration strategy. Dual-write-then-cutover, one service at a time, in the order established by cross-service query pressure (Chapter 2), never all at once — the same incremental discipline this series has applied to every extraction so far.
8. Common Mistakes
- Cutting over before verifying the backfill. Flipping the connection string before confirming the historical data actually landed correctly in the new instance risks serving stale or missing data on day one. Fix: the reconciliation job (Section 5) must show zero drift for a defined window before cutover, not just “looks right” from a spot check.
- Making the shadow write blocking. If
shadow.decrement()can throw and fail the primary request, the migration itself becomes a source of production incidents. Fix: wrap shadow writes inrunCatching(or equivalent) and treat failures as metrics, never as request failures. - Migrating all three services’ databases in one release. This recreates the “big bang” risk Chapter 2 warned against, at the database layer instead of the deployable layer. Fix: one service, one dual-write window, one cutover, fully verified, before starting the next.
- No rollback window. Dropping the old schema immediately after cutover removes the ability to revert if a problem surfaces a day later. Fix: keep the old schema read-only and intact for a defined window (Northwind used one week) before decommissioning it.
- Treating database-per-service as “one instance per service, forever, regardless of load.” Provisioning a dedicated multi-node cluster for a service handling a handful of requests per minute is the overengineering flagged in Section 4. Fix: size the instance to the service’s actual, measured footprint, and revisit if it changes.
- Letting another service keep a direct connection to the old schema “temporarily.” The most common way this pattern fails silently: a forgotten batch job or reporting tool still points at the old shared instance months after cutover. Fix: revoke the old database credentials for every consumer except the owning service as part of the cutover checklist, not as an afterthought.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Measured resource contention (connections, I/O) between services on a shared instance | Yes | Isolating instances removes the shared-resource coupling directly | — |
| Regulatory requirement for isolated infrastructure (PCI, health data) | Yes | Often close to mandatory; also simplifies audit scope | — |
| Small team, low per-service data volume, no contention observed | No | Fixed per-instance operational cost isn’t justified yet | Schema-per-service on one well-monitored instance |
| Service boundary is still unclean (cross-schema queries exist) | No | Migrating the database won’t fix a boundary problem — it will turn it into an outage | Fix the service boundary first (Chapter 2), then reassess |
| One service has a distinct backup/retention/version requirement | Yes | Independent instance lets that service set its own policy without affecting others | — |
10. Hands-On Exercise
Extend it: design (and if you like, implement against local Docker Compose instances) the same dual-write migration for payment-service, which — unlike inventory-service — needs zero data loss with much lower tolerance for drift, given it holds financial records. What does that tighter tolerance change about the dual-write window length or the reconciliation frequency?
Simulate a failure: during a simulated dual-write window, kill the shadow database (inventory-db) mid-migration. Confirm that order-service’s calls to inventory-service are completely unaffected, and that the shadow-write failure metric captures the outage. If either assumption is wrong in your implementation, that’s a bug in the migration shim, not an acceptable trade-off.
Decision question, with justification required: order-service and payment-service both need to answer “what’s the total value of orders placed by customer X in the last 30 days” — a query that, pre-decomposition, was a single SQL join. Post database-per-service, should this be solved by payment-service calling order-service’s API at request time, by order-service publishing events that payment-service folds into its own read model, or by a dedicated reporting service that reads from both via their APIs? Name the trade-offs from Section 1 that apply, and identify which later chapter in this series directly addresses the pattern you’d reach for.
11. Key Takeaways
- Splitting deployables (Chapter 2) and splitting databases (this chapter) are two separate decisions with two separate sets of evidence required — don’t do the second just because you did the first.
- Schema-per-service on a shared instance is a legitimate destination, not just a migration waypoint, for teams and workloads that haven’t outgrown it.
- The migration is safest as dual-write-then-verify-then-cutover, never a direct cutover, and it should be invisible to every other service’s API contract throughout.
- A clean pre-existing service boundary (going back to Chapter 1’s enforced module discipline) is what makes this cutover a configuration change instead of a code rewrite.
- Never let a shadow-write failure during migration fail the primary request — the migration path must be strictly additive risk, not a new failure mode for live traffic.
- Revoking old database access for every consumer at cutover is not optional — a forgotten direct connection to a decommissioned schema is a silent, delayed outage waiting to happen.
- More database instances is real, ongoing operational cost — right-size the pattern to actual, measured contention or compliance need, not to a checklist of “what microservices are supposed to look like.”