Strangler Fig migration pattern
Chapter 1 mentioned, in passing, that Northwind’s new platform was “replacing a fifteen-year-old PHP monolith.” Every chapter since has built the replacement — but never addressed how a live, revenue-generating legacy system actually gets replaced without a shutdown weekend and a prayer. This chapter answers that.
1. Problem the Pattern Solves
The legacy PHP monolith still runs Northwind’s checkout, catalog browsing, and customer account management — real, paying-customer traffic, today. The new Kotlin platform (Chapters 1–12) has working order, inventory, and payment capabilities, fully tested, running in production for exactly zero real customers, because nothing yet routes any live traffic to them.
The obvious plan — finish every capability, then cut over all traffic in one release — is the plan every postmortem about a failed migration describes in hindsight. A big-bang cutover means the first time the new system sees real production load, real customer data edge cases, and real integration failures is the same moment the old system is switched off, with no fallback if something’s wrong. Northwind’s leadership, having read enough of those postmortems, wants proof the new order-placement flow works under real traffic before the PHP checkout code is deleted — not after.
Forces in tension:
- Risk reduction vs. running two systems simultaneously. Migrating incrementally means, for a period, both the legacy system and pieces of the new system are live and must interoperate — a genuine, temporary increase in system complexity, in exchange for a large reduction in cutover risk.
- Speed of migration vs. safety. A slow, capability-by-capability migration takes calendar time and sustained team focus; a fast, big-bang migration is quicker to attempt but has essentially no ability to partially roll back if something’s wrong post-cutover.
- Data consistency during the transition. If both the legacy system and a new service can touch overlapping data (e.g., a customer’s order history split across old and new order records) during the migration window, keeping that data consistent and queryable as one coherent whole is a real, temporary engineering problem the pattern must address, not ignore.
- Team morale and stakeholder confidence. A years-long migration with no visible progress milestones risks losing organizational support; each strangled capability needs to be a demonstrable, working, in-production win, not just progress toward a distant finish line.
2. Core Idea
The Strangler Fig pattern migrates a legacy system incrementally by placing a routing facade in front of it, then moving individual capabilities behind that facade from the legacy system to a new implementation, one at a time, until nothing remains that still needs to be routed to the legacy system — at which point it can be safely decommissioned. The name comes from the strangler fig vine, which grows around a host tree, gradually taking over its structural role, until the original tree is no longer needed.
Northwind’s facade is the API gateway already built in Chapter 4 — no new component needed, just new routing rules added to the existing spring-cloud-gateway configuration:
Participants:
- Facade — the API gateway, making the routing decision per request path, invisible to clients who only ever see one platform.
- Legacy system — the PHP monolith, untouched by this pattern except for having traffic gradually redirected away from it.
- New implementation — the Kotlin services this series has built, taking over one routed capability at a time.
- Shared data bridge (Section 5) — the temporary mechanism keeping legacy and new data consistent during the overlap window, for capabilities where both systems’ data must remain coherent as a whole.
Commonly confused with:
- A simple reverse-proxy migration (“lift and shift”). Moving the same legacy code behind a new proxy without actually replacing any of its functionality isn’t strangling anything — the pattern specifically means capabilities are being reimplemented, not just re-hosted, behind the facade.
- The API gateway pattern itself (Chapter 4). The gateway is the mechanism this chapter reuses, not a different pattern — Strangler Fig is the migration strategy (route by capability, incrementally, toward full replacement); the gateway is simply the routing tool already available for the job.
- Feature flags / canary releases (a later chapter). A canary release routes a percentage of traffic to a new version of the same service to de-risk a single deployment; Strangler Fig routes traffic between two structurally different systems (legacy monolith vs. new microservices) based on which capability a request needs, typically over a much longer migration timeline. The two techniques can combine — canary-releasing a newly-strangled capability’s rollout — but they answer different questions.
3. When to Use It
Strong indicators:
- A live, business-critical legacy system needs replacing, and a full shutdown-and-cutover window is unacceptable (true for essentially any revenue-generating system, and definitely true for Northwind’s checkout).
- The legacy system’s capabilities can be identified and separated well enough to route by URL path, feature, or another observable request attribute — if the legacy system is so tangled that no capability boundary is externally visible, that untangling has to happen first (which may itself use techniques from Chapter 2’s decomposition-by-capability analysis, applied to the legacy codebase).
- The organization needs, or benefits from, visible incremental progress and the ability to pause or roll back individual migrated capabilities independently.
Concrete use cases:
- E-commerce, as here: replacing a legacy monolith’s checkout, catalog, and account capabilities one at a time behind a routing gateway.
- Government platforms: migrating a decades-old benefits-processing system to modern infrastructure while it continues serving live claims, where downtime or a failed cutover has direct, serious consequences for real people’s benefits.
- Banking: replacing a legacy core-banking module (account management, say) while it processes live transactions, with regulatory requirements that likely mandate an auditable, gradual, reversible migration over a risky big-bang one.
- Any system where the legacy codebase’s owners/experts are limited or leaving. Incremental migration lets a small number of people with legacy-system knowledge validate each newly-strangled capability against their understanding of the old system’s actual behavior, before that knowledge is lost.
Prerequisites:
- A facade capable of routing by the capability boundary you’ve chosen — Northwind’s path-based routing is the simplest case; some legacy systems need content-based or header-based routing if capabilities aren’t cleanly separated by URL.
- A data-consistency strategy for any period where both systems touch related data (Section 5) — this is usually the hardest part of a real Strangler Fig migration, harder than writing the new capability itself.
- Monitoring that can compare legacy and new implementations’ behavior for the same capability, ideally before fully cutting traffic over (a shadow-traffic or dual-write verification step, similar in spirit to Chapter 3’s database migration dual-write).
4. When Not to Use It
- The legacy system is small enough, or low-risk enough, that a full rewrite-and-cutover is genuinely simpler. A small internal tool with few users and an easy rollback (redeploy the old version) may not need the incremental machinery this pattern requires — match the migration strategy’s weight to the system’s actual blast radius.
- No clean way to route by capability, and untangling that would itself be a larger project than the migration. If capabilities are so deeply intertwined in the legacy codebase that no observable request attribute distinguishes them, forcing a Strangler Fig approach before that untangling work is done can produce a facade with hundreds of special-case routing rules that’s harder to reason about than either system alone.
- Overengineering signal: building an elaborate, general-purpose “migration facade framework” for a one-time migration effort, when the existing API gateway (as Northwind reuses here) already does the job with a handful of routing rules.
- Risk: treating the migration as indefinitely paused after strangling the “easy” capabilities first, leaving the hardest, most tangled parts of the legacy system running forever “for now.” A Strangler Fig migration without a tracked plan and a target end-state can stall in a permanently half-migrated state, which is arguably the worst outcome — running two systems’ worth of complexity forever, with neither fully retired.
5. Implementation Example
Routing rule addition to the existing gateway (Chapter 4), migrating order placement first — the highest-value, most-tested new capability:
spring: cloud: gateway: routes: - id: order-placement-migrated uri: http://order-service.default.svc.cluster.local predicates: - Path=/api/orders/** # New implementation — Chapters 1-12 - id: legacy-checkout uri: http://legacy-php-monolith.default.svc.cluster.local predicates: - Path=/checkout/**, /catalog/**, /account/** # Not yet migrated — still routes to the legacy systemThe harder part: data consistency during the overlap window. Northwind’s legacy monolith and the new order-service both need a consistent view of “does this customer account exist and what’s their loyalty tier” during the migration, since order placement (migrated) needs customer data that account management (not yet migrated) still owns in the legacy database:
package `in`.o612.eng.northwind.order.internal.legacy
import org.springframework.web.client.RestClientimport java.util.UUID
/** Temporary bridge, active only during the migration window. Calls the * legacy monolith's internal API for customer data not yet owned by any * new service. Deleted once account management is itself strangled * (tracked as a follow-up migration step, not an indefinite dependency). */class LegacyCustomerBridge(private val legacyClient: RestClient) { fun getCustomerTier(customerId: UUID): String = legacyClient.get().uri("/internal/api/customers/{id}/tier", customerId) .retrieve().body(CustomerTierResponse::class.java)?.tier ?: "STANDARD"}
data class CustomerTierResponse(val tier: String)This is exactly the anti-corruption layer pattern (the next chapter) in embryonic form — order-service doesn’t want the legacy monolith’s data model leaking into its own domain, so it translates at the boundary. The next chapter formalizes this technique; here it appears as the natural, minimal thing to do at a Strangler Fig seam.
Shadow-traffic verification, run before fully cutting a capability over — sending real production requests to both implementations and comparing results, without the new implementation’s response actually being returned to the client yet:
package `in`.o612.eng.northwind.gateway
import org.springframework.cloud.gateway.filter.GatewayFilterimport org.springframework.stereotype.Componentimport reactor.core.publisher.Mono
@Componentclass ShadowTrafficFilter(private val shadowClient: ShadowHttpClient, private val comparisonMetrics: ComparisonMetrics) : GatewayFilter {
override fun filter(exchange: org.springframework.web.server.ServerWebExchange, chain: org.springframework.cloud.gateway.filter.GatewayFilterChain): Mono<Void> { // Legacy response is what's actually returned to the client; // the new service's response is fetched and compared, but discarded. shadowClient.mirrorTo("order-service", exchange.request) .subscribe { shadowResponse -> comparisonMetrics.compare(exchange.request.path.value(), shadowResponse) } return chain.filter(exchange) }}Production note. Shadow traffic must never cause a side effect twice — mirroring a
POST /checkoutto a shadoworder-servicethat actually reserves stock and charges a card would double-charge real customers. Shadow-traffic comparison is safe for read paths, or requires the shadow target to run in a mode that simulates the write without committing it — a real, non-trivial design constraint worth naming explicitly rather than glossing over.
6. Step-by-Step Flow
Tracing the migration of order placement, from “still legacy” to “fully strangled”:
- Client action (of the migration): the platform team decides order placement is the highest-confidence capability to migrate first, based on Chapters 1–12’s test coverage.
- API request equivalent: the gateway is configured to shadow-route order-placement traffic to
order-servicewithout affecting real responses. - Service behavior:
order-serviceprocesses shadow requests exactly as it would real ones, its results captured for comparison only. - Database interaction: during the shadow phase, only the legacy database is authoritative;
order-service’s writes (if shadow mode permits any) are treated as throwaway. - Inter-service communication:
order-servicecalls theLegacyCustomerBridgefor data account management still owns — the temporary cross-system dependency this pattern accepts during the overlap window. - Error or failure handling: a mismatch between legacy and shadow responses is logged and investigated before cutover, not discovered as a customer-facing bug after — the entire value of the shadow phase.
- Observability signals: a comparison-mismatch rate, tracked per capability being migrated, is the specific metric that gates the go/no-go cutover decision — not a deadline on a calendar.
- Final response/outcome: once mismatch rate is acceptably low for a sustained period, the gateway’s routing rule flips real traffic to
order-service, and the legacy checkout code becomes dead code, safely removable after a rollback-safety window.
7. Production Concerns
- Timeouts, retries, idempotency. Shadow traffic must not affect the real request’s latency or reliability — fire-and-forget the shadow call (as shown, via
.subscribe()rather than blocking on it) so a slow or failing shadow target never impacts the client-facing legacy response. - Data consistency during the transition. The
LegacyCustomerBridgeis a real, if temporary, cross-system coupling — track every such bridge explicitly as migration debt with an owner and a removal plan, not as a permanent architectural feature. - API versioning and backward compatibility. The gateway’s routing rules are, in effect, the migration’s version-control mechanism — each capability’s cutover is a single, reversible configuration change (route back to legacy) rather than a code rollback, a genuine operational advantage of centralizing the routing decision at the gateway.
- Authentication and service-to-service trust. Both legacy and new systems need to accept the same authentication scheme during the overlap — if the legacy system uses session cookies and the new services use JWTs (Chapter 4’s assumption), the gateway may need to bridge between the two auth schemes for the duration of the migration, another piece of temporary complexity to track and remove.
- Logging, metrics, tracing, correlation IDs. Tag every request with which implementation (legacy or new) actually served it, in addition to the usual correlation ID — this makes it possible to correlate a customer-reported issue back to “was this the old or new system” during the ambiguous overlap period.
- Kubernetes deployment. The legacy monolith, however old, needs to be containerized and deployed into the same cluster as the new services (or reachable from the gateway) for this pattern to route to it at all — if it isn’t already, that containerization is itself a prerequisite migration step, smaller and lower-risk than rewriting its logic.
- Testing strategy. Shadow-traffic comparison (Section 5) is this pattern’s signature testing technique — it validates the new implementation against real production inputs and legacy outputs, catching edge cases no test suite anticipated, before any real customer is exposed to the new code path.
- Migration strategy (the pattern’s own subject): migrate capabilities in order of confidence and value — Northwind chose order placement first because it had the most test coverage (Chapters 1–12) and because a positive outcome there builds organizational confidence for tackling the harder-to-untangle capabilities (like account management) later.
8. Common Mistakes
- No end-state plan, migrating only the “easy” capabilities. Strangling order placement and catalog browsing, then leaving account management on the legacy system indefinitely because it’s tangled, means running two systems’ operational overhead forever. Fix: track every capability’s migration status explicitly, with owners and target dates for the hard ones too, even if those dates are revised.
- Shadow-traffic side effects causing real-world harm. Mirroring a checkout request to a shadow implementation that actually charges a card, as warned in Section 5, is a serious, customer-facing failure mode disguised as a safety measure. Fix: verify shadow targets are genuinely side-effect-free (read-only comparison, or a simulate-without-commit mode) before enabling shadow traffic on any write path.
- Cutting over based on a deadline instead of a measured comparison result. Flipping the routing rule because “the migration was scheduled to finish this sprint,” regardless of what the shadow-traffic comparison showed, reintroduces the exact big-bang risk this pattern exists to avoid. Fix: gate cutover decisions on the comparison-mismatch metric from Section 6, not the calendar.
- Leaving temporary bridges (like
LegacyCustomerBridge) undocumented as temporary. Six months later, nobody remembersLegacyCustomerBridgewas meant to be deleted once account management was strangled, and it becomes permanent, undocumented technical debt. Fix: track every migration-era bridge as an explicit, dated item in the migration plan, not just a comment in the code. - No rollback path once traffic is cut over. Deleting the legacy checkout code immediately after cutover, with no safety window, removes the ability to route back if a production issue surfaces post-cutover that shadow traffic didn’t catch. Fix: keep the legacy implementation deployable and the routing rule reversible for a defined safety window after each cutover.
- Treating the gateway’s routing rules as the whole migration. Assuming the migration is “done” once traffic is routed to the new service, without verifying the new service’s actual behavior matches what customers depended on, skips the entire value of the shadow-traffic phase. Fix: always validate via shadow traffic (or an equivalent comparison mechanism) before cutover, not just build the new capability and switch a route.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Live, business-critical legacy system needs replacing without downtime | Yes | Incremental, reversible migration avoids the big-bang cutover risk | — |
| Legacy system’s capabilities are cleanly separable by route/feature | Yes | Facade routing is straightforward to implement and reason about | — |
| Small, low-risk legacy system with easy rollback via redeploy | No | Full rewrite-and-cutover may be genuinely simpler for low blast-radius systems | Direct rewrite and cutover |
| Legacy capabilities are deeply entangled with no observable routing boundary | Not yet | Forcing routing-based migration before untangling produces a fragile facade | Untangle the legacy codebase’s capabilities first, then migrate |
| Migration has stalled indefinitely on the “easy” capabilities | No longer following the pattern correctly | An indefinite half-migrated state runs two systems’ overhead forever | Recommit to a tracked plan covering every remaining capability |
10. Hands-On Exercise
Extend it: design the routing and shadow-traffic plan for migrating /catalog/** next, after order placement’s successful cutover. What makes catalog browsing an easier or harder shadow-traffic candidate than order placement, given it’s a read-heavy path?
Simulate a failure: in a test environment, introduce a deliberate behavioral difference between order-service and a stubbed “legacy” response for the same request (e.g., a different tax-calculation rounding rule), and confirm your shadow-traffic comparison mechanism actually surfaces the mismatch rather than silently ignoring it.
Decision question, with justification required: account management (customer profile, addresses, saved payment methods) is the last capability left on the legacy monolith, and it’s also the most deeply tangled with the legacy database’s schema. Should Northwind invest in untangling and migrating it via Strangler Fig, or is this a case (per Section 4) where a different approach — a full, scheduled rewrite with a maintenance window, given its lower relative traffic compared to checkout — might actually be justified? Name the trade-offs from Section 1 that inform your answer.
11. Key Takeaways
- The Strangler Fig pattern migrates a legacy system incrementally, capability by capability, behind a routing facade — avoiding the all-or-nothing risk of a big-bang cutover.
- The facade doesn’t need to be new infrastructure — Northwind reused the API gateway already built in Chapter 4, since routing-by-capability was already its job.
- The hardest part is rarely writing the new capability — it’s managing data consistency and cross-system dependencies during the overlap window, and tracking those dependencies as temporary, not permanent.
- Shadow-traffic comparison validates a new implementation against real production inputs before any customer is exposed to it — but must be designed carefully to avoid duplicating real side effects on write paths.
- Gate every cutover decision on a measured comparison result, not a deadline — the entire point of incremental migration is making the decision reversible and evidence-based.
- Track every migration-era bridge and every remaining capability with an explicit owner and plan — an indefinitely stalled, half-migrated state is a worse outcome than either fully legacy or fully migrated.
- This pattern and the anti-corruption layer (next chapter) are close cousins — the boundary between old and new systems needs exactly the same translation discipline as the boundary between any two differently-modeled systems.