Circuit breaker, retry, timeout, fallback, and bulkhead
Chapter 7 named the risk directly: “a slow payment-service… cascaded backward” during a flash sale. This chapter gives order-service’s remaining synchronous calls (the warehouse dashboard’s gRPC path from Chapter 6, and the legacy ambassador from Chapter 15) the specific defenses that prevent exactly that cascade.
1. Problem the Pattern Solves
During Northwind’s next flash sale, inventory-service’s database starts running slow — a bad query plan, not an outage — so every reservation call now takes eight seconds instead of eighty milliseconds, but still eventually succeeds. order-service has no timeout configured on that call (a gap flagged but not fixed since Chapter 2’s implementation example), so its threads pile up waiting. Within two minutes, order-service’s entire request-handling thread pool is occupied by requests waiting on inventory-service, and every incoming order — even ones that don’t need inventory at all, like a customer checking their own order history — starts timing out, because there are no threads left to serve them.
The failure never should have spread that far. inventory-service being slow is a real, specific problem; order-service’s order-history endpoint failing because of it is a design flaw — the two operations share nothing except an accidentally shared resource pool.
Forces in tension:
- Fast failure vs. tolerance for transient blips. A short timeout fails fast, protecting the caller’s resources, but may reject a request that would have succeeded one second later, on a call that was merely momentarily slow.
- Automatic recovery vs. hammering a struggling service. Blind, immediate retries can turn a struggling downstream service’s temporary slowness into a full outage, by multiplying the very load that’s causing it to be slow in the first place.
- Isolation vs. resource efficiency. Dedicating separate thread pools per downstream dependency (bulkheads) prevents one slow dependency from starving unrelated request paths, at the cost of provisioning and tuning multiple pools instead of one shared, simpler one.
- User experience vs. protecting the system. A circuit breaker that opens (stops trying) protects the caller and the struggling downstream service, but means the caller now has to decide what to tell the user instead of the real answer — a fallback response, which is itself a product decision, not just an engineering one.
2. Core Idea
Five distinct, complementary techniques, each addressing a different failure mode of a synchronous call:
- Timeout: bound how long a caller will wait for a response, so a slow call fails predictably instead of holding resources indefinitely.
- Retry: automatically re-attempt a failed call, with backoff, for failures likely to be transient — bounded in count, and only for operations confirmed idempotent (the discipline this series has required since Chapter 1).
- Circuit breaker: after enough consecutive or recent failures, stop attempting calls to a struggling dependency entirely for a cool-down period, failing fast instead of continuing to pile on load — protecting both the caller’s own resources and the struggling dependency’s chance to recover.
- Fallback: when a call fails (or the circuit is open), return a reasonable default or degraded response instead of propagating the failure to the end user.
- Bulkhead: isolate the resources (typically a thread pool or a concurrent-call limit) used for one dependency from those used for others, so one dependency’s slowness can’t exhaust resources needed to serve unrelated requests.
Commonly confused with:
- The Saga pattern (Chapter 11). A saga coordinates multi-step business operations across services with compensation for already-completed steps. These resilience patterns operate at the level of a single call — they don’t know or care about business semantics, only about that one call’s success, failure, and timing. A saga step’s individual calls should each be wrapped in these resilience patterns; the two operate at different layers and compose naturally.
- The Transactional Outbox (Chapter 12). The outbox guarantees an event is eventually published despite a crash; these patterns handle a live, synchronous call’s failure modes in real time. Different failure windows, different mechanisms — a service can and typically does need both.
- Load balancing. Spreading requests across multiple healthy instances (handled by Kubernetes’
Servicediscovery, Chapter 5) is a different concern from what happens when the instances you’re routing to are all slow or failing — resilience patterns handle the latter.
3. When to Use It
Strong indicators:
- Any synchronous call to another service or external system (Chapters 2, 6, 13, 15’s legacy ambassador) — which describes every remaining blocking call in Northwind’s system.
- A demonstrated or plausible cascading-failure scenario, exactly as Section 1 describes — the risk named since Chapter 7 becoming concrete.
- Multiple, independent request paths sharing a resource pool that a slow dependency could exhaust, starving unrelated functionality.
Concrete use cases:
- E-commerce, as here: any checkout-adjacent call to inventory, payment, or a legacy system needs the full set — timeout, bounded retry, circuit breaking, and a bulkhead separating it from unrelated request paths.
- Payments: circuit-breaking calls to an external card processor prevents a processor outage from also taking down the parts of a payment service that don’t need the processor (fetching payment history, for instance).
- Logistics: a carrier API integration (rate quotes, tracking lookups) benefits from a bulkhead separating it from core shipment-processing logic, so a slow third-party carrier can’t stall internal operations.
- Any platform with third-party dependencies of varying reliability: the whole point of these patterns is that you don’t control the reliability of what you call, only how you respond to its failures.
Prerequisites:
- Confirmed idempotency for any call configured with automatic retries — retrying a non-idempotent call (as this series has repeated since Chapter 1) risks duplicate side effects.
- A defined, product-approved fallback behavior for each critical call path — “what do we tell the customer if inventory is unreachable” is a business decision this pattern needs an answer to, not something to improvise mid-incident.
- Monitoring for circuit-breaker state transitions and bulkhead saturation — these patterns are only useful if their activation is visible, not silently absorbing failures nobody notices.
4. When Not to Use It
- Calls with no plausible failure mode worth protecting against. An in-process method call (Chapter 1’s monolith-era design) has no network to fail — applying circuit-breaker machinery there is meaningless ceremony.
- A call whose failure should always immediately and visibly fail the whole request, with no fallback possible. For some operations, there genuinely is no reasonable degraded response — attempting to synthesize a “fallback” for a payment-capture failure, for instance, would be actively wrong (Section 8’s mistake #3). Not every call needs, or should have, a fallback.
- Overengineering signal: wrapping every internal call, including ones already protected by a larger, coarser mechanism (a saga’s own retry/compensation logic, for instance), with a full independent circuit-breaker-plus-bulkhead-plus-fallback stack, multiplying configuration surface for marginal additional protection.
- Risk of misconfigured retries: setting a retry count and backoff without considering the combined worst-case latency (retries × per-attempt timeout) can make a slow dependency’s failure take even longer to surface to the caller than doing nothing at all would have — always calculate and bound the total worst-case time, not just each individual setting.
5. Implementation Example
Resilience4j configuration, applied to order-service’s call to inventory-service — the exact call path that caused Section 1’s incident:
dependencies { implementation("io.github.resilience4j:resilience4j-spring-boot3:2.2.0") implementation("org.springframework.boot:spring-boot-starter-aop")}resilience4j: timelimiter: instances: inventoryReservation: timeout-duration: 2s retry: instances: inventoryReservation: max-attempts: 3 wait-duration: 200ms exponential-backoff-multiplier: 2 retry-exceptions: - in.o612.eng.northwind.order.internal.TransientInventoryException ignore-exceptions: - in.o612.eng.northwind.order.internal.StockUnavailableException # business outcome, never retry circuitbreaker: instances: inventoryReservation: sliding-window-size: 20 failure-rate-threshold: 50 wait-duration-in-open-state: 30s permitted-number-of-calls-in-half-open-state: 5 bulkhead: instances: inventoryReservation: max-concurrent-calls: 20 max-wait-duration: 100ms orderHistoryLookup: max-concurrent-calls: 20 max-wait-duration: 100msTwo separate bulkheads — inventoryReservation and orderHistoryLookup — is the single most important line in this configuration: it’s the direct fix for Section 1’s incident, guaranteeing the order-history endpoint has its own 20 threads that inventory slowness can never touch, no matter how badly inventoryReservation’s pool saturates.
package `in`.o612.eng.northwind.order.internal
import io.github.resilience4j.circuitbreaker.annotation.CircuitBreakerimport io.github.resilience4j.bulkhead.annotation.Bulkheadimport io.github.resilience4j.retry.annotation.Retryimport io.github.resilience4j.timelimiter.annotation.TimeLimiterimport org.springframework.stereotype.Componentimport java.util.UUIDimport java.util.concurrent.CompletableFuture
@Componentclass InventoryClient(private val restClient: org.springframework.web.client.RestClient) {
@Bulkhead(name = "inventoryReservation") @CircuitBreaker(name = "inventoryReservation", fallbackMethod = "reserveStockFallback") @Retry(name = "inventoryReservation") @TimeLimiter(name = "inventoryReservation") fun reserveStock(orderId: UUID, items: List<ReservationItemDto>): CompletableFuture<ReservationOutcome> = CompletableFuture.supplyAsync { restClient.post().uri("/api/v1/reservations").body(ReservationRequestDto(orderId, items)) .exchange { _, response -> when (response.statusCode.value()) { 200 -> ReservationOutcome.Reserved 409 -> throw StockUnavailableException(orderId) // business outcome, not a transient failure else -> throw TransientInventoryException(orderId) } } }
// Invoked when retries are exhausted, the circuit is open, or the // bulkhead is full — a defined, product-approved degraded response, // never a silent swallow. fun reserveStockFallback(orderId: UUID, items: List<ReservationItemDto>, ex: Throwable): CompletableFuture<ReservationOutcome> = CompletableFuture.completedFuture(ReservationOutcome.PendingRetryAsync(orderId))}
class TransientInventoryException(orderId: UUID) : RuntimeException("Transient failure reserving stock for order $orderId")class StockUnavailableException(orderId: UUID) : RuntimeException("Stock genuinely unavailable for order $orderId")The distinction between StockUnavailableException (a real, final business outcome — retrying won’t help, the stock simply isn’t there) and TransientInventoryException (a network or infrastructure hiccup — retrying might help) is what makes the retry configuration’s ignore-exceptions/retry-exceptions split correct. Retrying a genuine “out of stock” response would be pointless work and, worse, could give a misleading impression that the system is “trying harder” when the actual answer will never change.
ReservationOutcome.PendingRetryAsync is the fallback’s chosen degraded behavior — a deliberate product decision (made with the checkout team, not invented by an engineer mid-incident) that when inventory can’t be reached, the order is accepted provisionally and reconciled asynchronously once inventory recovers, rather than failing the checkout outright. This connects directly back to the outbox pattern (Chapter 12): the provisional order state and its eventual reconciliation are exactly the kind of “guarantee eventual correctness despite a live failure” problem that pattern’s machinery already solves.
Test verifying the circuit breaker actually opens under sustained failure, using Resilience4j’s test utilities:
package `in`.o612.eng.northwind.order.internal
import io.github.resilience4j.circuitbreaker.CircuitBreakerimport org.junit.jupiter.api.Testimport org.assertj.core.api.Assertions.assertThat
class InventoryClientResilienceTest {
@Test fun `circuit opens after failure threshold and fallback is invoked`() { val registry = io.github.resilience4j.circuitbreaker.CircuitBreakerRegistry.ofDefaults() val breaker = registry.circuitBreaker("inventoryReservation") val failingClient = InventoryClient(alwaysFailingRestClient())
repeat(20) { runCatching { failingClient.reserveStock(sampleOrderId, sampleItems).get() } }
assertThat(breaker.state).isEqualTo(CircuitBreaker.State.OPEN) }}6. Step-by-Step Flow
- Client action.
POST /orders, unchanged. - API request.
order-servicecallsreserveStock, entering theinventoryReservationbulkhead — capped at 20 concurrent calls, isolated from theorderHistoryLookupbulkhead. - Service behavior. The call times out at 2 seconds (not 8), retries once with backoff, times out again.
- Database interaction. Never reached in this failure scenario — the timeout fires before
inventory-service’s slow query even returns. - Inter-service communication. After enough recent failures cross the configured threshold, the circuit opens — subsequent calls fail immediately, without even attempting the network call, for the 30-second cool-down.
- Error or failure handling. The fallback method returns a defined, product-approved degraded outcome (
PendingRetryAsync) rather than propagating a raw exception to the client. - Observability signals. Circuit-breaker state transitions (
CLOSED→OPEN→HALF_OPEN→CLOSED) and bulkhead rejection counts are first-class metrics, watched on their own dashboard — the exact signal that would have made Section 1’s incident visible and diagnosable in seconds instead of minutes. - Final response. The client gets
202 Acceptedwith a provisional order status — degraded, but honest and fast — while, critically, order-history requests on the separate bulkhead continue succeeding throughout, the specific failure this chapter exists to prevent.
7. Production Concerns
- Timeouts, retries, idempotency. Never configure a retry without confirming the underlying operation is idempotent — Section 5’s split between
StockUnavailableException(never retried) andTransientInventoryException(retried) is the concrete mechanism for this; apply the same discipline to every retry-configured call in the system. - Calculating worst-case latency. With a 2s timeout and 3 max attempts with exponential backoff (200ms, 400ms), the absolute worst case before the fallback fires is roughly
2s + 0.2s + 2s + 0.4s + 2s ≈ 6.6s— a number that must be explicitly acceptable to whatever’s callingorder-service(its own timeout budget, per Chapter 4’s cross-hop timeout composition) or the retry configuration needs tightening. - Data consistency and transaction boundaries. A fallback that returns a degraded, provisional result (as
PendingRetryAsyncdoes) must have a real, tested reconciliation path — this is not optional cleanup; an order stuck in “provisional” forever because nobody built the reconciliation job is a worse outcome than a clean failure would have been. - API versioning. Resilience configuration (timeouts, retry counts) should be treated as operationally tunable (via the centralized configuration from Chapter 5, ideally with
@RefreshScopefor rapid incident response) rather than requiring a full redeploy to adjust during an active incident. - Authentication and service-to-service trust. Unaffected directly by these patterns, but note that a circuit breaker opening due to authentication failures (a misconfigured or expired service credential) looks identical, from the caller’s metrics, to a genuine downstream outage — alert on the underlying exception type, not just the circuit state, to distinguish these during triage.
- Logging, metrics, tracing, correlation IDs. Log every fallback invocation with the triggering exception and the correlation ID — a fallback silently returning a degraded response with no trace of why is exactly the “invisible failure” this series has warned against since Chapter 7’s introduction of asynchronous flows.
- Kubernetes deployment, autoscaling. Bulkhead sizing should account for the pod’s actual thread pool and expected concurrency, not an arbitrary number — undersized bulkheads reject legitimate load even when the downstream dependency is healthy, while oversized ones fail to actually isolate anything.
- Testing strategy. Test each resilience behavior explicitly and in isolation — a test that forces the circuit open (as Section 5 shows), a test that verifies retries stop for a business-outcome exception, and a test that verifies the bulkhead actually caps concurrency, are all different, necessary tests, not one general “resilience works” test.
- Migration strategy. Apply these patterns to the highest-risk call path first — Northwind started with
inventoryReservation, the exact path that caused a real incident — rather than instrumenting every call in the system uniformly and speculatively on day one.
8. Common Mistakes
- No timeout at all, relying only on retries and circuit breaking. A retry policy around a call with no timeout can still let each individual attempt hang indefinitely, defeating the entire point of bounding worst-case latency. Fix: always set an explicit timeout as the foundation; retries and circuit breaking build on top of it, not instead of it.
- Retrying a call without confirming idempotency. Applying a blanket retry policy to every outbound call, including ones with side effects that aren’t safe to repeat (a legacy system call that isn’t confirmed idempotent, for instance), risks duplicate effects under retry. Fix: explicitly classify every call as safe or unsafe to retry, as Section 5’s exception-type split does, before configuring any retry policy.
- Building a fallback for an operation with no reasonable degraded behavior. Synthesizing a fake “success” fallback for a payment-capture failure, just to keep the checkout flow moving, actively lies to the business about whether money changed hands. Fix: for calls with no honest degraded response, let the failure propagate and fail the request clearly — not every call needs, or should have, a fallback (Section 4).
- Sharing one bulkhead across unrelated call types. Giving
inventoryReservationandorderHistoryLookupthe same thread pool “for simplicity” reproduces exactly the incident from Section 1 — one slow dependency exhausting resources needed for an unrelated, healthy one. Fix: a separate bulkhead per logically distinct dependency or call type, as Section 5 configures. - Not calculating combined worst-case latency across timeout × retries. Configuring a generous timeout and a generous retry count independently, without multiplying them together, can produce a worst-case latency the caller’s own timeout budget (Chapter 4) never anticipated. Fix: always compute and document the worst-case total latency for any timeout-plus-retry configuration, as shown in Section 7.
- No monitoring on circuit-breaker state or bulkhead rejection rate. Letting these mechanisms silently absorb failures with no dashboard or alert means an ongoing, significant degradation (a circuit stuck open for hours) can go unnoticed by the team while customers experience it directly. Fix: treat circuit state transitions and bulkhead saturation as first-class, alerted metrics from day one of adopting these patterns.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Synchronous call to another service or external system with any failure risk | Yes (timeout + retry, at minimum) | Bounds worst-case latency and handles transient failures without manual intervention | — |
| A dependency’s slowness could exhaust resources needed for unrelated request paths | Yes (bulkhead) | Isolates the blast radius, the exact fix for Section 1’s incident | — |
| Sustained downstream failure risks piling load onto an already-struggling dependency | Yes (circuit breaker) | Fails fast, protecting both caller resources and the struggling dependency’s recovery | — |
| The call has a genuine, honest degraded response available | Yes (fallback) | Improves user experience during partial failure, if the degraded response is truthful | — |
| No reasonable degraded response exists for this specific operation | No (fallback) | A fabricated success or misleading fallback is worse than a clear failure | Let the failure propagate; handle it explicitly upstream |
| In-process call, no network involved | No | Nothing to protect against — these patterns address network/dependency failure modes | — |
10. Hands-On Exercise
Extend it: apply the same resilience configuration pattern (timeout, retry, circuit breaker, dedicated bulkhead) to order-service’s call to the legacy ambassador from Chapter 15, using a more generous timeout and retry budget appropriate for the legacy system’s known lower reliability — justify your specific numbers.
Simulate a failure: using a tool like Toxiproxy or a simple sleep injected into a test double for inventory-service, simulate exactly Section 1’s scenario (slow but not down) and verify: the circuit opens after the configured threshold, the fallback fires, and — critically — a concurrent request to order-service’s order-history endpoint succeeds normally throughout, proving the bulkhead isolation actually works.
Decision question, with justification required: Northwind’s PendingRetryAsync fallback for stock reservation currently has no automated reconciliation job actually implemented yet — orders can sit “provisional” indefinitely if inventory-service doesn’t recover. What component should own reconciling these, how would you detect an order that’s been provisional too long, and what should happen to it? Connect your answer back to the Saga (Chapter 11) and Outbox (Chapter 12) patterns.
11. Key Takeaways
- Timeout, retry, circuit breaker, fallback, and bulkhead are five distinct, complementary techniques — each addresses a different failure mode of a synchronous call, and a well-protected call path typically needs several of them together, as Section 5’s configuration shows.
- Always set an explicit timeout as the foundation; it’s the one piece none of the others can substitute for.
- Never configure automatic retries without first confirming the operation is idempotent — a distinction this series has required since Chapter 1, made concrete here through explicit exception-type classification.
- Bulkheads isolate one dependency’s failure from unrelated request paths sharing the same process — the direct, specific fix for the cascading-failure risk this series named all the way back in Chapter 7.
- Not every call needs a fallback — a fabricated or misleading degraded response (especially around money, as in payment capture) is worse than letting a failure propagate honestly.
- Always calculate worst-case latency as timeout × retry attempts, not each setting in isolation — an unconsidered combination can silently violate a caller’s own timeout budget.
- Monitor circuit-breaker state and bulkhead saturation as first-class, alerted metrics — these mechanisms are only protective if their activation is visible to the team, not silently absorbing an ongoing problem.