Event-driven architecture
Since Chapter 2, placing an order has meant order-service calling inventory-service then payment-service synchronously, one after another, over REST. This chapter replaces that chain with events on Kafka — the foundation the next several chapters (pub-sub, sagas, outbox) all build on.
1. Problem the Pattern Solves
Chapter 2 named the cost of its own design honestly: order-service’s checkout request now blocks on two sequential network calls, and if inventory-service is down, checkout fails entirely, even though nothing about placing an order requires an immediate answer from inventory — a customer is used to seeing “processing” for a few seconds. During Northwind’s last flash sale, a brief payment-service latency spike (its own downstream card processor was slow) cascaded backward: order-service’s threads piled up waiting on payment-service, which piled up waiting on the processor, and the whole checkout path degraded together, even though inventory-service was perfectly healthy the entire time.
There’s a second, quieter problem: every time Northwind adds a new capability that cares about order placement — a loyalty-points service, an analytics pipeline, a fraud-check service — order-service’s code has to change to call it. The set of things interested in “an order was placed” keeps growing, and order-service has become a hub that must know about, and call, every one of them.
Forces in tension:
- Coupling in time vs. coupling in interface. A synchronous call couples two services in both time (both must be up simultaneously) and interface (the caller must know the callee’s contract). Event-driven communication removes the time coupling — the publisher doesn’t wait for, or even know whether, anyone is currently listening — while still coupling on the event’s schema.
- Consistency vs. availability. A synchronous chain can (with effort) approximate strong consistency by failing the whole operation if any step fails. An event-driven flow is eventually consistent by construction — the order exists in
order_schemabeforeinventory-servicehas necessarily processed anything about it. - Simplicity of a linear chain vs. flexibility of a fan-out. A synchronous
order → inventory → paymentchain is easy to trace by reading the code top to bottom. An event-driven fan-out — one event, multiple independent consumers — is easier to extend (a new consumer doesn’t touch the publisher) but harder to trace end to end without dedicated tooling (the correlation-ID discipline from Chapters 4–6, and the Observability chapter later). - Operational cost. A message broker (Kafka, here) is new, stateful infrastructure to run, monitor, and reason about — partitions, consumer groups, retention — a real cost that only pays off once the coupling and availability problems above are actually being felt.
2. Core Idea
Event-driven architecture means services communicate by publishing facts about things that already happened (events — OrderPlaced, not commands like PlaceOrder) to a broker, and other services subscribe to the events they care about and react independently, without the publisher knowing or caring who’s listening or whether they’ve finished reacting yet.
Intent: decouple producers from consumers in time (publisher doesn’t block on consumers), in cardinality (any number of consumers, including zero, without publisher changes), and in failure (a consumer being down doesn’t prevent the event from being published, and — with the right broker semantics — doesn’t lose the event either).
Compare this to Chapter 2’s diagram: there, order-service explicitly called inventory-service, which then called payment-service — a chain the publisher had to know about end to end. Here, order-service publishes one event and knows nothing about who consumes it or how many consumers exist.
Commonly confused with:
- Message queues used for point-to-point work distribution (the Competing Consumers chapter, later). Event-driven publish-subscribe means every interested subscriber gets its own copy of the event; a work queue means exactly one worker picks up each message. Kafka can express either pattern depending on consumer-group configuration — this chapter uses the pub-sub shape (each service its own consumer group, so each gets every event); the Competing Consumers chapter uses the queue shape (multiple instances of one service sharing a consumer group, competing for messages).
- Request-reply over a broker (RPC-over-messaging). Publishing an event and later correlating a reply message is possible but reintroduces the synchronous chain’s temporal coupling with extra steps — this chapter’s events are fire-and-forget facts, not disguised remote procedure calls.
- Event sourcing (a later chapter). Event-driven architecture is about how services communicate; event sourcing is about how a service stores its own state — as a sequence of events rather than current-state rows.
order-servicein this chapter still stores anOrderrow in Postgres and merely also publishes an event about the change; it hasn’t adopted event sourcing.
3. When to Use It
Strong indicators:
- A synchronous chain has demonstrated cascading-failure risk (Northwind’s flash-sale incident) — one slow downstream service degrading callers that don’t strictly need an immediate answer from it.
- The number of things interested in “X happened” keeps growing, and each new interested party currently requires a code change in the publisher.
- The business operation genuinely tolerates eventual consistency — the customer experience already expects “order placed, confirmation to follow,” not an instant, fully-settled outcome.
Concrete use cases:
- E-commerce, as here: order placement fanning out to inventory, payment, loyalty, analytics, and notifications — a growing, independently-evolving set of interested consumers.
- Logistics: a “package scanned at facility” event consumed independently by customer notification, route optimization, and delay-prediction systems, none of which should block the scan operation itself.
- SaaS: a “subscription upgraded” event triggering billing recalculation, feature-flag updates, and a customer-success team notification — three independent reactions to one fact, none of which the billing system should have to know about explicitly.
- Healthcare: a “lab result finalized” event notifying the ordering physician, updating the patient’s chart, and triggering a billing code — each a separate concern that shouldn’t require the lab system to call three other systems synchronously.
Prerequisites:
- A message broker operated to production standards (this series uses Kafka throughout) — under-provisioned or unmonitored broker infrastructure becomes a new single point of failure worse than the synchronous chain it replaced.
- Explicit tolerance, agreed with the business, for eventual consistency on this specific operation — not every operation should become eventually consistent by default (see Section 4).
- A plan for the two hard problems this pattern introduces and defers: what happens when a sequence of dependent steps needs to happen reliably across services (the Saga chapter), and how a service guarantees it published an event for every state change it made (the Transactional Outbox chapter). This chapter’s example has a real gap on both counts, named honestly in Section 7.
4. When Not to Use It
- Operations that genuinely need strong, immediate consistency. A funds transfer between two accounts inside the same service’s transaction boundary should stay a single ACID transaction — event-driven eventual consistency solves a coordination problem you don’t have yet, and adds real complexity for no benefit, if the operation never needed to cross a service boundary in the first place.
- Low-volume, simple request/response interactions with a single consumer. If exactly one service ever needs to react to an event, and it needs to react synchronously (the caller needs to know the outcome before proceeding), a direct REST or gRPC call (Chapter 6) is simpler and doesn’t need broker infrastructure at all.
- Teams without the operational maturity to run and monitor a broker. Kafka (or any broker) run without dashboards for consumer lag, without alerting on under-replicated partitions, and without a retention/replay strategy becomes a black box that silently drops or delays messages — worse than a synchronous system whose failures are at least immediately visible.
- Risk: debugging complexity. A synchronous chain’s failure shows up as a stack trace at the point of failure. An event-driven flow’s failure might show up as “nothing happened” three services away, hours later, discoverable only through correlation IDs and distributed tracing — a real cost that Section 7 and the Observability chapter address, but never fully eliminate.
5. Implementation Example
What changes from Chapter 2, precisely. order-service’s HTTP-based InventoryClient and its synchronous call chain are replaced by a Kafka producer. inventory-service and payment-service stop exposing (for this flow specifically) a synchronous endpoint for order-service to call, and instead become Kafka consumers. Their internal domain logic — InventoryService.reserveStock, unchanged since Chapter 1 — does not change at all.
dependencies { implementation("org.springframework.kafka:spring-kafka")}package `in`.o612.eng.northwind.order.internal
import `in`.o612.eng.northwind.order.api.OrderPlacedimport org.springframework.kafka.core.KafkaTemplateimport org.springframework.stereotype.Component
@Componentinternal class OrderEventPublisher( private val kafkaTemplate: KafkaTemplate<String, OrderPlaced>,) { fun publish(event: OrderPlaced) { // Keyed by orderId so all events for one order land on the same // partition, preserving per-order ordering — see Section 7. kafkaTemplate.send("orders.events", event.orderId.toString(), event) }}package `in`.o612.eng.northwind.order.internal
import `in`.o612.eng.northwind.order.api.OrderPlacedimport org.springframework.stereotype.Serviceimport org.springframework.transaction.annotation.Transactionalimport org.springframework.transaction.event.TransactionPhaseimport org.springframework.transaction.event.TransactionalEventListenerimport org.springframework.context.ApplicationEventPublisher
@Serviceinternal class OrderService( private val orderRepository: OrderRepository, private val localEvents: ApplicationEventPublisher,) { @Transactional fun placeOrder(request: PlaceOrderCommand): Order { val order = orderRepository.save(Order.pending(request)) // Publish to Kafka only after this transaction commits — never // before, or a rolled-back order could still emit OrderPlaced. localEvents.publishEvent(OrderPlaced(order.id, order.items)) return order }}
@org.springframework.stereotype.Componentinternal class KafkaBridge(private val publisher: OrderEventPublisher) { @TransactionalEventListener(phase = TransactionPhase.AFTER_COMMIT) fun onOrderPlaced(event: `in`.o612.eng.northwind.order.api.OrderPlaced) = publisher.publish(event)}The AFTER_COMMIT-gated bridge is the same discipline Chapter 1 established for in-process events, now applied across a real broker — publish only what’s actually true and committed. Section 7 flags exactly where this still isn’t airtight (the gap between “transaction committed” and “Kafka send succeeded” is the Transactional Outbox chapter’s entire subject).
inventory-service as a consumer, replacing its Chapter 2 REST endpoint for this flow:
package `in`.o612.eng.northwind.inventory.internal
import `in`.o612.eng.northwind.order.api.OrderPlacedimport org.springframework.kafka.annotation.KafkaListenerimport org.springframework.stereotype.Component
@Componentinternal class OrderPlacedListener( private val inventoryService: InventoryService,) { @KafkaListener(topics = ["orders.events"], groupId = "inventory-service") fun onOrderPlaced(event: OrderPlaced) { inventoryService.reserveStock(event.orderId, event.items.toReservationRequests()) // reserveStock still publishes StockReserved / StockUnavailable — // now onto Kafka instead of the in-process bus from Chapter 1. }}spring: kafka: bootstrap-servers: kafka.default.svc.cluster.local:9092 consumer: group-id: inventory-service auto-offset-reset: earliest properties: spring.json.trusted.packages: "in.o612.eng.northwind.order.api,in.o612.eng.northwind.inventory.api"groupId = "inventory-service" is what makes this publish-subscribe rather than a work queue: every replica of inventory-service shares this group and Kafka load-balances partitions across them (so exactly one replica processes each order — correct, since reserving stock twice for the same order would be wrong), while payment-service, in its own separate consumer group, receives its own independent copy of every event.
Test using an embedded Kafka broker, verifying the publish-and-consume round trip without a real cluster:
package `in`.o612.eng.northwind.order.internal
import org.junit.jupiter.api.Testimport org.springframework.boot.test.context.SpringBootTestimport org.springframework.kafka.test.context.EmbeddedKafkaimport org.springframework.kafka.test.utils.KafkaTestUtilsimport org.assertj.core.api.Assertions.assertThatimport java.time.Duration
@EmbeddedKafka(partitions = 1, topics = ["orders.events"])@SpringBootTestclass OrderEventPublishingTest {
@Test fun `placing an order publishes OrderPlaced after the transaction commits`() { val consumer = KafkaTestUtils.getSingleRecord(testConsumer(), "orders.events", Duration.ofSeconds(5)) assertThat(consumer.value()).contains(sampleOrderId.toString()) }}6. Step-by-Step Flow
- Client action.
POST /orders, identical request shape to every previous chapter — the client-facing contract hasn’t changed since Chapter 1. - API request.
order-servicevalidates and returns as soon as its own write commits — it no longer waits oninventory-serviceorpayment-serviceat all. - Service behavior. The
AFTER_COMMITbridge publishesOrderPlacedto Kafka. - Database interaction. Only
order-service’s own database is touched synchronously;inventory_schemaandpayment_schemaare updated later, independently, by their own consumers. - Inter-service communication.
inventory-serviceandpayment-serviceboth receiveOrderPlacedindependently and in parallel — neither knows or cares about the other, a genuine change from Chapter 2’s explicit chain. - Error or failure handling. Notice what’s not handled yet: nothing currently stops
payment-servicefrom charging a customer for an order whose stock reservation later fails. This chapter deliberately leaves that gap visible rather than papering over it — it’s exactly what the Saga pattern (Chapter 11) exists to close. - Observability signals. Kafka consumer lag per service (
inventory-service,payment-service) becomes a first-class health signal — a growing lag means a consumer is falling behind or stuck, invisible in a synchronous system where a stuck call at least ties up a visible thread. - Final response. The client gets
202 Acceptedimmediately after step 2 — measurably faster and more available than Chapter 2’s chain, sinceorder-serviceno longer depends oninventory-serviceorpayment-servicebeing up to accept the order.
7. Production Concerns
- Timeouts, retries, idempotency. A consumer that crashes after processing a message but before committing its Kafka offset will reprocess that message on restart —
InventoryService.reserveStockmust be idempotent for a givenorderId(flagged as a design requirement all the way back in Chapter 1; this is where it becomes load-bearing rather than hypothetical). - Data consistency and transaction boundaries. The gap between “
order-service’s transaction committed” and “the Kafka send actually succeeded” is real and unaddressed by this chapter’s code — a crash in that narrow window means an order exists with no event ever published for it. The Transactional Outbox chapter closes exactly this gap; until then, treat it as a known, documented risk, not a solved problem. - Ordering guarantees. Kafka guarantees order only within a partition — keying every event by
orderId(as shown in Section 5) ensures all events for one order are processed in order relative to each other, but gives no ordering guarantee across different orders, which is usually fine and should be verified against your actual requirement, not assumed. - API versioning and backward compatibility.
OrderPlaced’s schema is now a contract every current and future consumer depends on — adding a field is safe if consumers ignore unknown fields (verify your deserializer does); removing or renaming one breaks every consumer silently unless you have schema validation (a schema registry, covered as a specific gotcha in the Kafka-adjacent chapters). - Authentication and service-to-service trust. Kafka should be configured with SASL/TLS and ACLs restricting which services can produce to or consume from which topics — an unauthenticated broker lets any service on the network publish or read any event, undermining the ownership model this whole series has built up.
- Logging, metrics, tracing, correlation IDs. Propagate the correlation ID as a Kafka message header, not just an HTTP header — without it, tracing “what happened to order X” across a publish-and-three-independent-consumers flow is materially harder than tracing Chapter 2’s linear chain.
- Kubernetes deployment, health probes, autoscaling. Consumers should scale based on consumer lag, not just CPU — a growing backlog is the clearest signal that a consumer group needs more replicas (up to its partition count; beyond that, more partitions are needed too).
- Testing strategy. Embedded Kafka (as shown above) for fast, hermetic tests of publish/consume logic; a small number of true end-to-end tests against a real (Testcontainers-managed) Kafka broker to catch configuration issues embedded Kafka can mask.
- Migration strategy. Northwind moved one flow (order placement) from synchronous to event-driven, keeping the REST reservation endpoint from Chapter 2 available for other callers that still need a synchronous answer (the warehouse dashboard’s stock check, for instance) — migrate flow by flow, based on which ones actually show the cascading-failure or extensibility pain from Section 1, not wholesale.
8. Common Mistakes
- Publishing before the transaction commits. Calling
kafkaTemplate.send()directly inside the@Transactionalmethod, before commit, risks publishing an event for a change that later rolls back. Fix: always gate the publish onAFTER_COMMIT, as shown in Section 5. - Treating events as disguised remote procedure calls. Publishing
ReserveStockCommand(an instruction) instead ofOrderPlaced(a fact) recreates synchronous coupling in async clothing — the publisher now has to know what it wants the consumer to do, not just what happened. Fix: publish facts about the past (past-tense event names), and let each consumer decide independently what to do about them. - Ignoring consumer idempotency. Assuming “Kafka delivers exactly once” without designing for at-least-once semantics (which is what Kafka actually provides by default, and what most production configurations use) leads to duplicate stock decrements on redelivery. Fix: design every consumer to be safely re-runnable for the same event, using the event’s own identifier as the idempotency key.
- No plan for cross-service consistency. Leaving the payment-before-stock-confirmed gap from Section 6 unaddressed in a real production system, rather than as a deliberately named, tracked risk pending the Saga chapter. Fix: name the gap explicitly, decide on an interim mitigation (e.g., a compensating refund path) if the Saga pattern isn’t implemented yet, and don’t let “eventually consistent” become “sometimes silently wrong.”
- Under-provisioning or under-monitoring the broker. Running Kafka without dashboards for consumer lag, disk usage, and under-replicated partitions turns the broker into a silent failure point nobody notices until a backlog has grown for hours. Fix: treat broker health as a first-class production concern with the same rigor as database health, from day one of adopting this pattern.
- Migrating every service-to-service call to events at once. Converting the warehouse dashboard’s stock-check call (Chapter 6, a synchronous read that genuinely needs an immediate answer) to an asynchronous event-driven flow “for consistency with order placement” adds latency and complexity to a call path that never needed it. Fix: keep synchronous calls where an immediate answer is actually required; use events only where the operation tolerates and benefits from eventual consistency.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Synchronous chain shows cascading-failure risk from downstream latency | Yes | Removes temporal coupling; a slow consumer no longer blocks the producer | — |
| Growing number of independent consumers interested in one fact | Yes | New consumers subscribe without changing the publisher | — |
| Operation needs an immediate, synchronous answer to proceed | No | Eventual consistency doesn’t fit a request that can’t continue without the result | REST or gRPC (Chapter 6) |
| Team lacks broker operational maturity (no lag monitoring, no alerting) | No, not yet | An unmonitored broker is a worse failure mode than a visible synchronous failure | Invest in broker operations first, or stay synchronous |
| Strong, immediate consistency is a hard business requirement for this specific operation | No | Eventual consistency introduces exactly the risk this requirement forbids | Keep it inside one service’s transaction, or use a saga with explicit compensation |
10. Hands-On Exercise
Extend it: add a new loyalty-service consumer to orders.events, in its own consumer group, that increments a customer’s loyalty points on every OrderPlaced — without touching a single line of order-service’s code. This is the extensibility benefit from Section 1, made concrete.
Simulate a failure: stop inventory-service entirely, place several orders, then restart it. Confirm the queued OrderPlaced events are processed once inventory-service comes back — and confirm nothing was lost. Then compare this recovery behavior to what Chapter 2’s synchronous version would have done in the same scenario (checkout requests failing outright while inventory-service was down).
Decision question, with justification required: payment-service’s consumer of OrderPlaced currently charges the customer immediately, without waiting to know whether inventory-service successfully reserved stock — a real risk, named in Section 6. Should payment-service be changed to wait for a StockReserved event before charging (introducing an ordering dependency between two independent consumers), or should the system accept the current risk and handle failed reservations with a refund? Name the trade-offs from Section 1 and Section 7 this decision turns on, and note which upcoming chapter is built specifically to answer this question properly.
11. Key Takeaways
- Event-driven architecture decouples services in time and cardinality: publishers don’t wait for consumers, and new consumers can subscribe without any change to the publisher.
- Publish facts about what already happened, not commands telling another service what to do — that distinction is what keeps the coupling loose rather than just relocated.
- This pattern trades immediate consistency for availability and extensibility — it’s the right trade only for operations that genuinely tolerate eventually-consistent outcomes.
- The gap between “committed a local transaction” and “successfully published the resulting event” is real, unaddressed by this chapter alone, and closed properly by the Transactional Outbox pattern later in this series.
- Consumers must be idempotent — Kafka’s at-least-once delivery means reprocessing will happen, and the system must tolerate it by design, not by hope.
- A message broker is real, stateful infrastructure requiring the same operational rigor (monitoring, alerting, capacity planning) as a database — don’t adopt this pattern without that investment.
- Migrate flow by flow, based on evidence of the specific pain (cascading failure, consumer sprawl) this pattern solves — not every service-to-service call benefits from becoming asynchronous.