Microservices decomposition by business capability
This chapter picks up Northwind Commerce exactly where Chapter 1 left it: a modular monolith with order, inventory, and payment modules, enforced Gradle boundaries, and event-driven cross-module communication. The domain model, the event names, and the schema-per-module discipline all carry forward unchanged — only the deployment topology changes in this chapter.
1. Problem the Pattern Solves
Eighteen months after Chapter 1, Northwind has grown to 45 engineers across six feature teams, all committing to the same northwind-platform repository and shipping one deployable JAR. The symptoms the modular monolith was supposed to defer have arrived on schedule: every deploy requires all six teams’ changes to pass CI together, so a flaky test in payment blocks a order-only bug fix from shipping. The inventory module’s nightly stock-reconciliation job needs four times the CPU of the rest of the application, but scaling the whole monolith to satisfy one module’s batch job wastes money on the other five days it isn’t running. The payment team wants to move to Kotlin coroutines and a different Postgres connection pool tuning profile than the rest of the app, but can’t, because it’s one JVM with one configuration.
These are exactly the leading indicators Chapter 1 named as the signal to extract services: team-to-module ratio, deploy contention, and a module with a distinct scaling profile. The question this chapter answers is not whether to split — that decision is already made — but along which lines, because getting this wrong is far more expensive than the monolith ever was.
Forces in tension:
- Coupling vs. autonomy. The wrong split (say, cutting along
OrderController/OrderService/OrderRepositorylayers, one “service” per layer) turns every business operation into a chatty, synchronous network conversation, trading in-process coupling for the strictly worse network-coupling. - Consistency vs. team ownership. A capability boundary that splits a single business invariant (e.g., “stock can’t go negative”) across two teams’ services means neither team can enforce that invariant alone — it becomes a cross-team contract, negotiated instead of compiled.
- Latency vs. scalability. Every module boundary that becomes a service boundary adds a network hop to any workflow that crosses it. Order placement, which was a few method calls in Chapter 1, now involves at least two HTTP round trips before this chapter is done — that cost only pays for itself if the resulting services can scale, deploy, and fail independently in ways that matter.
- Operational complexity vs. cost. Three services means three deployables, three sets of dashboards, three on-call rotations to define — real, recurring cost that the modular monolith didn’t have. This chapter only makes sense because Northwind’s team size has already outgrown one deployable’s coordination cost.
2. Core Idea
Decomposition by business capability means drawing service boundaries around cohesive units of business functionality — each capable of making decisions and owning data independently — rather than around technical layers, database tables, or organizational convenience. In Domain-Driven Design terms, each service should correspond to (at most) one bounded context: a boundary within which a domain model, its terminology, and its invariants are internally consistent, and outside of which the same word can mean something different.
Intent: each service should be independently understandable (a new engineer can hold its whole model in their head), independently deployable (its release doesn’t require coordinating with another team), and independently ownable of its invariants (it never needs another service’s cooperation, in the same transaction, to keep its own data correct).
Participants:
- Business capability — something the business does, stated as a verb phrase independent of any current implementation (“manage stock levels,” not “the
Stocktable”). - Bounded context — the DDD boundary inside which a capability’s domain model and ubiquitous language hold. Northwind’s
inventorycontext has its own concept of a “reservation”; that word means something different inpayment’s context (an authorization hold on a card), and that’s fine — they’re different models, on purpose. - Context map — the diagram of how bounded contexts relate (which is a customer of which, which translates the other’s language via an anti-corruption layer — covered in a later chapter).
Northwind’s split, and why each cut lands where it does:
Each box became a bounded context in Chapter 1 before it became a network boundary here — that ordering is the entire point of the pattern. The Chapter 1 InventoryApi interface becomes, almost mechanically, the inventory-service REST contract in this chapter, because the boundary was already correct.
Commonly confused with:
- Decomposition by technical layer. Splitting into a
presentation-service, abusiness-logic-service, and adata-serviceproduces three services that must all be called, in sequence, for any single business operation — maximum coupling, zero autonomy, none of the benefits this pattern promises. - Decomposition by entity/table (a “database-first” split). A
customer-service,order-service, andproduct-servicedrawn straight from an ER diagram often cuts through a single business capability — “placing an order” touches all three — recreating distributed-transaction problems that a capability-first cut avoids. Capability boundaries usually don’t align one-to-one with tables; a capability can own several tables, and a table can conceptually belong to exactly one capability even if a naive ER diagram suggests otherwise. - Decomposition by organizational chart alone. Conway’s Law means org structure influences good boundaries, but “we have six teams, so six services” without checking capability cohesion produces services that are really just teams with a network cable between them — the “distributed monolith” from Chapter 1’s confusion list.
3. When to Use It
Strong indicators:
- You have concrete, lived evidence of where a modular monolith’s module boundaries hold under real usage (Chapter 1’s prerequisite) — this pattern formalizes boundaries you’ve already validated, it doesn’t discover new ones from scratch.
- Deploy coordination cost, measured as “how often does team A’s change block or get blocked by team B’s,” is a recurring, named pain point in retrospectives.
- At least one capability has a genuinely distinct non-functional profile — scaling, latency tolerance, technology needs, compliance boundary — that the shared deployable can’t satisfy for all capabilities at once.
Concrete use cases:
- E-commerce, as here:
order,inventory,paymenteach have different scaling curves (flash sales spike order and payment traffic without touching inventory reconciliation) and different compliance surfaces (payment touches PCI scope; the others don’t). - Payments platforms: a
ledgercapability (immutable, audited, rarely changed) is usually a poor fit to co-deploy with afraud-scoringcapability (frequently retrained, needs fast iteration) — different change cadences are themselves a decomposition signal. - Logistics:
route-planning(CPU-heavy, batch-oriented) andshipment-tracking(high-volume, low-latency reads) have opposite scaling shapes and are natural separate capabilities even before team size forces the question. - Healthcare:
patient-records(strict access control, audit trail, rarely changes shape) andappointment-scheduling(frequent feature iteration) benefit from separate deployment and separate blast radius for compliance reasons alone.
Prerequisites:
- A context map, even an informal one, reviewed by more than one team — bounded-context boundaries that only one person believes in tend to be technical convenience, not business reality.
- An existing modular monolith or equivalent proof that the proposed boundary holds without constant cross-boundary transactions (Chapter 1’s boundary-erosion warning applies doubly once the boundary is a network call).
- A plan for what happens to cross-capability consistency — Northwind’s “don’t reserve stock that isn’t there and then fail payment” concern doesn’t disappear when it becomes two services; it becomes the subject of the saga pattern later in this series.
4. When Not to Use It
- The organization is smaller than the capability count implies. Three business capabilities and four engineers means one team will context-switch across three service codebases, three CI pipelines, and three on-call surfaces — strictly worse than one team owning one well-modularized monolith.
- The capabilities aren’t actually independent yet. If
ordercannot process a checkout without a synchronous, in-the-same-request answer frominventoryfor every single field it needs, splitting them turns a fast method call into a network dependency that fails independently — measure this before splitting, using exactly the kind of event-driven boundary Chapter 1 already put in place. - Overengineering pattern to watch for: decomposing “for scale” a system that has not demonstrated it needs to scale non-uniformly. A single Postgres instance and a Spring MVC monolith comfortably serves tens of millions of requests per day for most CRUD-shaped domains; splitting earlier than the team-size or scaling evidence demands multiplies operational surface for a benefit that may never materialize.
- Risk: a bounded context boundary that looked right on a whiteboard but wasn’t tested against real cross-module event traffic (Chapter 1’s integration tests) will produce a service boundary that requires constant renegotiation of its contract — effectively distributed refactoring, the most expensive kind.
5. Implementation Example
What actually changes from Chapter 1, and what doesn’t. The internal implementation classes inside each module — InventoryService, StockRepository, the domain logic — move with almost no changes into their own Gradle project and their own Spring Boot application. What changes is the transport: InventoryApi, an in-process Kotlin interface, becomes an HTTP contract; ApplicationEventPublisher calls become an HTTP client call. This chapter deliberately uses a plain, undecorated RestClient call with no retries, circuit breakers, or timeouts configured beyond a sane default — those concerns get their own, dedicated treatment in the “Synchronous communication” and “Circuit breaker, retry, timeout” chapters later in this series. Introducing them here would bury the decomposition lesson under resilience-engineering detail that belongs elsewhere.
Repository layout after extraction (three independently deployable Gradle projects, three independent CI pipelines, three Postgres databases):
northwind-order-service/ ← was the `order` module + `app` composition rootnorthwind-inventory-service/ ← was the `inventory` modulenorthwind-payment-service/ ← was the `payment` moduleThe inventory service keeps its bounded-context model completely intact — only the entry point changes, from a Spring bean implementing InventoryApi to a @RestController exposing the same operation:
package `in`.o612.eng.northwind.inventory.api
import org.springframework.http.ResponseEntityimport org.springframework.web.bind.annotation.*
@RestController@RequestMapping("/api/v1/reservations")class InventoryController(private val inventoryService: InventoryService) {
@PostMapping fun reserveStock(@RequestBody request: ReservationRequest): ResponseEntity<ReservationResponse> { val result = inventoryService.reserveStock(request.orderId, request.items) return when (result) { is ReservationResult.Reserved -> ResponseEntity.ok(ReservationResponse.reserved(request.orderId)) is ReservationResult.Unavailable -> ResponseEntity.status(409).body(ReservationResponse.unavailable(request.orderId, result.skus)) } }}
data class ReservationRequest(val orderId: java.util.UUID, val items: List<ReservationItem>)data class ReservationItem(val sku: String, val quantity: Int)
data class ReservationResponse(val orderId: java.util.UUID, val status: String, val unavailableSkus: List<String> = emptyList()) { companion object { fun reserved(orderId: java.util.UUID) = ReservationResponse(orderId, "RESERVED") fun unavailable(orderId: java.util.UUID, skus: List<String>) = ReservationResponse(orderId, "UNAVAILABLE", skus) }}InventoryService.reserveStock is the exact same class from Chapter 1’s internal package, unchanged — it still returns a sealed result, still writes only to its own schema, still knows nothing about HTTP. That’s the payoff of enforcing the public-API/internal split at the module level before extraction: the domain code didn’t need a rewrite, only a new front door.
order-service replaces its event listener with a synchronous HTTP client call to the new contract:
package `in`.o612.eng.northwind.order.internal
import org.springframework.web.client.RestClientimport java.util.UUID
class InventoryClient(private val restClient: RestClient) {
fun reserveStock(orderId: UUID, items: List<ReservationItemDto>): ReservationOutcome = restClient.post() .uri("/api/v1/reservations") .body(ReservationRequestDto(orderId, items)) .exchange { _, response -> when (response.statusCode.value()) { 200 -> ReservationOutcome.Reserved 409 -> ReservationOutcome.Unavailable(response.bodyTo(ReservationResponseDto::class.java)?.unavailableSkus.orEmpty()) else -> throw InventoryServiceUnavailable(orderId) } }}
sealed interface ReservationOutcome { data object Reserved : ReservationOutcome data class Unavailable(val skus: List<String>) : ReservationOutcome}
data class ReservationRequestDto(val orderId: UUID, val items: List<ReservationItemDto>)data class ReservationItemDto(val sku: String, val quantity: Int)data class ReservationResponseDto(val orderId: UUID, val status: String, val unavailableSkus: List<String> = emptyList())class InventoryServiceUnavailable(orderId: UUID) : RuntimeException("Inventory service unavailable for order $orderId")Trade-off. Notice what was lost in this transport swap: the Chapter 1 version was eventually consistent by design (an
AFTER_COMMITevent, no caller blocked waiting). This version blocks the order-placement request on a live network call to another service. That is a real regression in latency and availability — the order request now fails ifinventory-serviceis down, which it never could in the monolith. Section 7 and the later Saga chapter both come back to this; for now, treat it as the honest cost of this particular boundary, not a bug to silently fix.
Consumer-driven contract test, proving the two independently-deployed services still agree on the shape of the boundary — this is a lightweight hand-rolled version; the dedicated contract-testing chapter later in this series replaces it with Pact:
package `in`.o612.eng.northwind.order.internal
import org.junit.jupiter.api.Testimport org.springframework.boot.test.web.client.TestRestTemplateimport org.springframework.web.client.RestClientimport com.github.tomakehurst.wiremock.junit5.WireMockExtensionimport com.github.tomakehurst.wiremock.client.WireMock.*import org.junit.jupiter.api.extension.RegisterExtensionimport org.assertj.core.api.Assertions.assertThatimport java.util.UUID
class InventoryClientContractTest {
@RegisterExtension val wiremock = WireMockExtension.newInstance().build()
@Test fun `treats HTTP 409 as an Unavailable outcome, not an exception`() { wiremock.stubFor( post(urlEqualTo("/api/v1/reservations")) .willReturn(aResponse().withStatus(409).withHeader("Content-Type", "application/json") .withBody("""{"orderId":"${'$'}{java.util.UUID.randomUUID()}","status":"UNAVAILABLE","unavailableSkus":["WIDGET-1"]}""")) ) val client = InventoryClient(RestClient.create(wiremock.baseUrl()))
val outcome = client.reserveStock(UUID.randomUUID(), listOf(ReservationItemDto("WIDGET-1", 2)))
assertThat(outcome).isInstanceOf(ReservationOutcome.Unavailable::class.java) }}6. Step-by-Step Flow
- Client action. Storefront sends
POST /orderstoorder-servicealone — it never talks toinventory-serviceorpayment-servicedirectly, preserving the same external contract as Chapter 1. - API request.
order-servicevalidates and persists the order inorder_schema— still the same shared Postgres instance from Chapter 1, since decomposing the deployables and decomposing the databases are two separate decisions; Chapter 3 covers moving each schema to its own instance. - Service behavior.
order-service’sInventoryClientissues a synchronous REST call toinventory-service. - Database interaction.
inventory-servicedecrements stock ininventory_schema— still on the shared instance, butorder-service’s code has no repository, connection, or credential that can reach it; the only door in is the HTTP contract. - Inter-service communication. On success,
order-servicecallspayment-servicenext, sequentially — this chapter keeps the call chain simple and synchronous on purpose; whether this should be choreographed via events instead is exactly the question the Saga chapter answers. - Error or failure handling. A
409frominventory-serviceshort-circuits the flow before any payment is attempted — the same business rule as Chapter 1, now expressed as an HTTP status code contract instead of a Kotlin sealed type, which is the real cost of crossing a network boundary: the contract has to be serialized and versioned. - Observability signals. Each service emits its own
orders.placed/reservations.created/charges.capturedmetrics; without a shared process, correlating “this order’s” three service calls into one story requires a correlation ID threaded through HTTP headers — the Observability chapter builds this out fully. - Final response. The client still gets
202 Acceptedimmediately after step 2, matching Chapter 1’s contract — the internal transport changed completely; the client-facing behavior didn’t have to.
7. Production Concerns
- Timeouts, retries, idempotency. The synchronous call chain in step 5 means a slow
payment-servicenow directly threatensorder-service’s own request latency and thread pool — a concern that didn’t exist in Chapter 1. Configure an explicit connect/read timeout on everyRestClient(shown here without one, deliberately, to keep focus on decomposition — never ship it without one). - Data consistency and transaction boundaries. The single Chapter 1 transaction boundary is gone; “reserve stock, then charge payment, then mark paid” is now three independently-committing operations across a network, with a real, reachable failure window between each step (e.g., stock reserved,
payment-servicecrashes before responding). This is the exact problem the Saga pattern exists to name and solve — this chapter intentionally leaves it partially unsolved as the motivation for that chapter. - Database-per-service and schema ownership. Each service still connects to its own schema on the one shared Postgres instance from Chapter 1 — the deployables split before the data did. The next chapter walks through moving each schema to its own instance and what that buys you beyond schema separation alone.
- API versioning and backward compatibility.
POST /api/v1/reservations’s request/response shape is now a cross-team contract that can’t be changed in one atomic commit — a field rename requires a deprecation window, unlike the Chapter 1InventoryApiinterface which could be refactored compiler-wide in a single PR. - Authentication and service-to-service trust.
order-servicecallinginventory-serviceover the network introduces a real trust boundary for the first time in this series — for now, assume a private network with no cross-service auth (label this explicitly as a temporary simplification); the OAuth2/token-propagation chapter later replaces this assumption. - Logging, metrics, tracing, correlation IDs. Generate a correlation ID at
order-service’s ingress and propagate it as an HTTP header (X-Correlation-Id) on every outbound call — without it, reconstructing “what happened to order X” means grepping three separate log streams by timestamp and guessing. - Kubernetes deployment. Three Deployments, three Services, independent replica counts and resource requests —
inventory-servicecan now finally get the CPU allocation its nightly reconciliation job needs without over-provisioning the other two. - Testing strategy. Unit and slice tests stay inside each service. The full-stack Testcontainers test from Chapter 1 (booting every module together) no longer exists as a single-process test — it’s replaced by the WireMock-based consumer contract test shown above, plus, ideally, a small number of true end-to-end tests running all three real services together in CI.
- Migration strategy. Extract one service at a time, starting with the module under the least cross-module event traffic (here,
payment, which only reacts to one event type), and keep the others as monolith modules until each extraction is proven — a “big bang” three-way split in one release multiplies the number of things that can go wrong simultaneously.
8. Common Mistakes
- Splitting along technical layers instead of capabilities. A
northwind-api-gateway-service,northwind-business-logic-service,northwind-data-servicetriad forces every request through all three network hops for zero autonomy gained. Fix: every service should be able to answer at least one complete business question without calling another service synchronously in the common case. - Ignoring the bounded-context language shift. Reusing the word “reservation” identically in
inventory-service(a stock hold) and, later, in a hypotheticalshipping-service(a delivery slot) without noticing they’re different concepts creates silent contract bugs. Fix: maintain an explicit glossary per bounded context and treat a shared word with different meanings as a signal to rename, not to unify. - Extracting a service before its data-ownership boundary is provably clean. If
inventory-servicestill needs a live SQL join againstorder_schemato answer a common query, the extraction happened before the boundary was ready. Fix: verify — as Chapter 1’s Testcontainers tests did — that a module’s queries never cross into another module’s schema, before the network split, not after. - No plan for partial failure. Treating the new synchronous call chain (
order→inventory→payment) as if it fails atomically, the way the old in-process transaction did. Fix: name the failure window explicitly (as Section 7 does here) and treat it as a tracked, temporary gap, closed by the Saga and Outbox chapters — not something to hope doesn’t happen. - Decomposing every module at once. A single release that turns three monolith modules into three services simultaneously means any integration failure could originate from any of nine new network paths. Fix: extract incrementally, one service, one release, with the others still in the monolith, exactly as Section 7’s migration strategy describes.
- Treating the new REST contract as internal and casual. Because the team that owns
order-servicealso wroteinventory-service’s controller last quarter, changes get made to the contract without version negotiation “since we all know about it.” Fix: the instant a boundary crosses a deployable, it needs the same contract discipline as a public API, regardless of team overlap — that discipline gets tooling support in the contract-testing chapter.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Modular monolith has proven, stable module boundaries; deploy coordination is now the bottleneck | Yes | Boundaries are validated; the remaining cost is coordination, which extraction solves | — |
| One module has a measurably distinct scaling or compliance profile | Yes | Independent scaling/isolation is worth the network cost for that module specifically | Extract just that one module; leave the rest as a monolith |
| Boundaries are still shifting frequently during normal development | No | Extracting an unstable boundary means renegotiating a network contract repeatedly | Stay in the modular monolith; keep refining module boundaries |
| Team is small (fewer engineers than proposed services) | No | Operational overhead per service exceeds the coordination benefit | Modular monolith, possibly with one service extracted for a specific hard constraint |
| A cut point requires a live cross-module SQL join to work today | No | The data-ownership boundary isn’t clean enough yet | Fix the module boundary in the monolith first (Chapter 1 discipline), then re-evaluate |
10. Hands-On Exercise
Extend it: extract the notification module (from Chapter 1’s exercise) into its own northwind-notification-service, exposing a POST /api/v1/notifications endpoint, and have order-service call it after a terminal order state — synchronously or fire-and-forget, your choice, but justify the choice against Section 1’s latency-vs-coupling trade-off.
Simulate a failure: stop inventory-service entirely and send an order through order-service. Observe exactly what the client receives and how long it takes to receive it. Compare that behavior to Chapter 1’s monolith, where inventory could never be “down” independently of the whole app. What does this tell you about the true cost of this decomposition, separate from its benefits?
Decision question, with justification required: Northwind’s payment capability now needs to add refunds, which requires reading order-service’s order history to validate a refund request against the original order. Should payment-service call order-service synchronously for this, maintain its own read-only copy of the relevant order data, or should order-service push the data it needs? Name which forces from Section 1 and Section 7 drive your answer, and note which later chapter in this series (there are two strong candidates) would let you revisit the decision with more tools available.
11. Key Takeaways
- Decompose along business capabilities and bounded contexts, not technical layers, database tables, or team headcount alone — the modular monolith’s proven module boundaries are the input to this decision, not a starting guess.
- A clean extraction is mostly mechanical when the module boundary was already enforced pre-network (Chapter 1’s discipline pays off directly here); a messy one signals the boundary wasn’t ready.
- Crossing a network boundary always costs something real — latency, a new failure mode, contract versioning — that a compile-time module boundary didn’t have. Never present decomposition as free.
- The single transaction that guaranteed consistency in the monolith is gone the moment a boundary becomes a network call; name that gap explicitly rather than assuming it away.
- Extract one service at a time, starting with the lowest-cross-traffic module, and validate each extraction before starting the next.
- A shared word used differently across two contexts (like “reservation”) is a decomposition signal, not a naming coincidence to paper over.
- This pattern trades a real, ongoing operational cost (more deployables, more on-call surfaces, more contracts to version) for real, ongoing autonomy benefits — it only pays off once team size and scaling needs actually demand that autonomy.