API gateway and backend for frontend
Northwind now runs order-service, inventory-service, and payment-service (Chapters 1–3), each with its own database. This chapter addresses what happens once real clients — a web storefront and a mobile app — start calling them directly.
1. Problem the Pattern Solves
Northwind ships a mobile app alongside the web storefront. Both need an “order confirmation” screen: order status from order-service, a payment receipt from payment-service, and — on mobile only, for a simplified view — nothing from inventory-service. Today, both clients call all the services they need directly, over the public internet, each with its own TLS certificate, its own CORS configuration, its own rate limiting (or, more often, none at all).
Three problems surface within a month. First, every service now needs to implement authentication, TLS termination, and CORS itself — three copies of the same cross-cutting concern, drifting out of sync as each team patches its copy differently. Second, the mobile team, on a slow cellular connection, complains that three separate round trips (order, payment, and an unnecessary inventory call the app never asked to drop) make the confirmation screen feel sluggish. Third, when payment-service’s URL changes because it moved to a new cluster, both the web and mobile teams have to ship a client update — the internal topology has leaked into public client contracts.
Forces in tension:
- Simplicity vs. duplication. Handling auth, rate limiting, and TLS once, centrally, is simpler to operate correctly than three times, independently — but a central component is also a new single point of failure and a new deployable to own.
- Client fit vs. one-size-fits-all. Mobile and web often need different payload shapes, different levels of aggregation, and different latency budgets; a single generic API punishes one client type to satisfy the other.
- Latency vs. coupling. Aggregating calls server-side (fewer round trips for the client) trades client-side latency for a new internal coupling point that has to call multiple services and handle their partial failures.
- Security surface vs. flexibility. Exposing every internal service directly to the internet maximizes the attack surface; funneling all traffic through one edge component narrows it, at the cost of that component becoming a high-value target and a potential bottleneck.
2. Core Idea
API gateway: a single edge component that all external traffic passes through, responsible for cross-cutting concerns — routing to the correct backend service, TLS termination, authentication, rate limiting, and request/response logging — applied uniformly, once, instead of duplicated per service.
Backend for frontend (BFF): a thin, client-specific service sitting behind (or as a specialized route on) the gateway, that shapes and aggregates backend calls into exactly the contract one client type needs — a web-bff for the storefront, a mobile-bff for the app — rather than forcing every client to consume the same generic, lowest-common-denominator API.
These are related but distinct decisions: a gateway is about cross-cutting infrastructure concerns applied uniformly; a BFF is about tailoring the contract per client. You can have one without the other — a gateway with no BFF (all clients get the same downstream contract, just routed and secured centrally) or a BFF with no shared gateway (unusual, but possible in a very small system).
Commonly confused with:
- A reverse proxy / load balancer. A plain reverse proxy routes and load-balances but is typically unaware of business-level concerns like per-route auth policy or request aggregation — an API gateway is a reverse proxy with application-aware routing and cross-cutting policy layered on top.
- API composition / aggregator services (a later chapter in this series). A BFF may aggregate calls internally, but its defining trait is client-ownership of the contract, not the aggregation mechanic itself. API composition is the general technique — timeout budgets, partial-failure handling, fan-out/fan-in — that a BFF, a gateway, or a standalone aggregator can all use. This chapter treats aggregation lightly; the composition chapter covers doing it robustly.
- Service mesh (a later chapter). A service mesh handles service-to-service traffic inside the cluster (mTLS, retries, traffic shaping between internal services); an API gateway handles client-to-cluster traffic at the edge. They solve adjacent but different problems and are commonly used together, not as alternatives.
3. When to Use It
Strong indicators for a gateway:
- More than one service is directly reachable from outside the cluster, and each has independently implemented (or forgotten to implement) auth, rate limiting, or TLS.
- You need a single place to enforce a security policy change (e.g., a new required header, a WAF rule) without touching every service.
Strong indicators for a BFF:
- Two or more client types (web, mobile, a partner API) have genuinely different data shape, aggregation, or latency needs from the same underlying services — not just cosmetic differences that a single flexible response could satisfy.
- A client-facing contract is churning quickly to satisfy one client’s UI changes, and forcing every other client to absorb those changes (or version around them) is causing friction.
Concrete use cases:
- E-commerce: exactly Northwind’s scenario — web and mobile confirmation screens with different aggregation and payload needs.
- Banking/fintech: a partner-facing API (for a third-party integrator) typically needs a far more conservative, versioned, and rate-limited contract than the bank’s own mobile app — a natural BFF split by consumer trust level, not just device type.
- SaaS with an admin console and an end-user app: the admin console often needs broader, more detailed data (audit fields, internal IDs) than the end-user app should ever see — a BFF per audience prevents over-exposure by construction rather than by careful field filtering in a shared endpoint.
- Government platforms: a public-facing citizen portal and an internal caseworker tool against the same case-management services usually warrant separate BFFs, since their authorization models and data visibility rules differ substantially.
Prerequisites:
- A stable set of backend service contracts to gateway/aggregate (Chapters 2–3) — building a BFF against services whose contracts are still churning means maintaining the BFF becomes a second, redundant churn point.
- Clarity on which cross-cutting concerns belong at the gateway (auth, rate limiting) versus which belong in each service (business validation) — pushing business logic into the gateway is a common failure mode covered in Section 8.
4. When Not to Use It
- A single client type, or clients with near-identical needs. Building separate BFFs for web and mobile when both consume exactly the same data in the same shape is pure duplication — one shared API (possibly still behind a gateway for cross-cutting concerns) is simpler.
- Very small systems (one or two services). A gateway in front of one service adds a network hop and an operational component for no routing benefit — direct access, secured at the service itself, is simpler until there’s more than one thing to route between.
- Risk: the gateway becomes a monolith by another name. If business logic, orchestration, or per-client conditional behavior accumulates in the gateway’s routing rules, it becomes an undocumented, hard-to-test, single point of coupling for the whole system — the opposite of what decomposition (Chapter 2) was for.
- Overengineering signal: standing up a BFF per client “for future flexibility” before a second client type actually exists. Build the shared API first; split into BFFs when a second, meaningfully different consumer actually shows up.
5. Implementation Example
Stack choice. Spring Cloud Gateway (reactive, built on WebFlux/Reactor) for the edge gateway — it’s designed specifically for this role: non-blocking I/O suits a component whose entire job is proxying many concurrent connections with low per-request overhead. The BFFs themselves use plain Spring MVC, since they perform a small, bounded number of downstream calls per request rather than needing WebFlux’s high-concurrency streaming model.
Gateway routing and cross-cutting config:
dependencies { implementation("org.springframework.cloud:spring-cloud-starter-gateway") implementation("org.springframework.boot:spring-boot-starter-oauth2-resource-server") implementation("org.springframework.boot:spring-boot-starter-actuator")}spring: cloud: gateway: routes: - id: web-bff uri: http://web-bff:8080 predicates: - Path=/web/** filters: - StripPrefix=1 - name: RequestRateLimiter args: redis-rate-limiter.replenishRate: 50 redis-rate-limiter.burstCapacity: 100 - id: mobile-bff uri: http://mobile-bff:8080 predicates: - Path=/mobile/** filters: - StripPrefix=1 - name: RequestRateLimiter args: redis-rate-limiter.replenishRate: 100 redis-rate-limiter.burstCapacity: 200Rate limits differ per client type deliberately — mobile traffic patterns (bursty, on reconnect) differ from web traffic, and the gateway is the single place that policy is expressed, instead of being duplicated (and inevitably drifting) inside each BFF.
Authentication is enforced once, at the gateway, before any request reaches a BFF:
package `in`.o612.eng.northwind.gateway
import org.springframework.context.annotation.Beanimport org.springframework.context.annotation.Configurationimport org.springframework.security.config.annotation.web.reactive.EnableWebFluxSecurityimport org.springframework.security.config.web.server.ServerHttpSecurityimport org.springframework.security.web.server.SecurityWebFilterChain
@Configuration@EnableWebFluxSecurityclass SecurityConfig { @Bean fun filterChain(http: ServerHttpSecurity): SecurityWebFilterChain = http .authorizeExchange { it.anyExchange().authenticated() } .oauth2ResourceServer { it.jwt {} } .build()}Security note. The full OAuth2/JWT validation setup (issuer configuration, claim mapping, token propagation to downstream services) is covered in depth in this series’ security chapter — this config is deliberately minimal to keep the gateway/BFF distinction in focus. Don’t copy this snippet into production without that chapter’s additions.
The mobile-bff, aggregating exactly what the mobile confirmation screen needs and nothing else:
package `in`.o612.eng.northwind.mobilebff
import org.springframework.web.bind.annotation.*import java.util.UUID
@RestController@RequestMapping("/orders")class OrderConfirmationController( private val orderClient: OrderServiceClient, private val paymentClient: PaymentServiceClient,) { @GetMapping("/{orderId}/confirmation") fun confirmation(@PathVariable orderId: UUID): MobileOrderConfirmation { val order = orderClient.getOrder(orderId) val receipt = paymentClient.getReceipt(orderId) return MobileOrderConfirmation( orderId = order.id, status = order.status, totalCharged = receipt?.amount, lastFourDigits = receipt?.cardLastFour, ) // Deliberately no inventory data — the mobile confirmation screen // never shows stock detail. web-bff's equivalent endpoint below does. }}
data class MobileOrderConfirmation( val orderId: UUID, val status: String, val totalCharged: java.math.BigDecimal?, val lastFourDigits: String?,)The web-bff’s equivalent endpoint calls the same three services but returns a richer payload including line-item and stock-availability detail the web confirmation page displays and mobile doesn’t — two genuinely different contracts serving two genuinely different clients, each independently versionable without affecting the other.
package `in`.o612.eng.northwind.mobilebff
import org.springframework.web.client.RestClientimport java.util.UUID
class OrderServiceClient(private val restClient: RestClient) { fun getOrder(orderId: UUID): OrderView = restClient.get().uri("/api/v1/orders/{id}", orderId).retrieve().body(OrderView::class.java) ?: throw OrderNotFound(orderId)}
data class OrderView(val id: UUID, val status: String)class OrderNotFound(orderId: UUID) : RuntimeException("Order $orderId not found")6. Step-by-Step Flow
- Client action. Mobile app requests
/mobile/orders/{id}/confirmationwith a bearer token — it has never heard oforder-serviceorpayment-serviceby name or address. - API request. The gateway validates the JWT and applies the mobile-specific rate limit before forwarding.
- Service behavior.
mobile-bfffans out two concurrent calls toorder-serviceandpayment-service. - Database interaction. Handled entirely inside each downstream service, unchanged from Chapter 3 — the BFF holds no database of its own in this example.
- Inter-service communication. The two calls in step 3 run in parallel, not sequentially, since they’re independent reads with no ordering dependency — unlike Chapter 2’s order-placement chain, where each step depended on the previous one’s outcome.
- Error or failure handling. If
payment-serviceis slow or down,mobile-bffmust decide whether to fail the whole confirmation or return a partial response (order status without a receipt) — this chapter’s example fails closed for simplicity; the API composition chapter builds out a principled partial-failure policy. - Observability signals. The gateway logs the route, status code, and latency for every request; the BFF and both downstream services each emit their own metrics, correlated by the request ID the gateway assigns at ingress.
- Final response. The client receives one aggregated JSON payload from one round trip, shaped exactly for the mobile screen — the two internal calls, and the two services behind them, are invisible to it.
7. Production Concerns
- Timeouts, retries, idempotency. The gateway and each BFF need their own timeout budgets, and those budgets must compose correctly — if the gateway times out at 5s but a BFF’s downstream call can take 8s, the client sees a generic gateway timeout instead of a meaningful error from the BFF.
- Data consistency and transaction boundaries. BFFs should be read-heavy aggregators; pushing writes that span multiple services into a BFF turns it into an undocumented orchestrator — write orchestration belongs in a service that owns that responsibility explicitly (see the Saga chapter).
- API versioning and backward compatibility. Because a BFF’s contract is owned by (and versioned with) its one client type, it can evolve quickly without a deprecation negotiation across multiple consumer teams — one of the strongest arguments for splitting BFFs in the first place.
- Authentication, authorization, and service-to-service trust. The gateway terminates the client’s auth; calls from the BFF to downstream services need their own service-to-service credential (a client-credentials token, not the end user’s raw token, generally) — covered fully in the security chapter.
- Logging, metrics, tracing, correlation IDs. Assign the correlation ID at the gateway, the earliest point in the request’s life, and propagate it through every downstream hop — this is the first point in the series where a single request genuinely spans four processes (gateway, BFF, two services), making this discipline non-optional.
- Kubernetes deployment, health probes, autoscaling. The gateway is a horizontally-scaled, stateless Deployment fronted by a Kubernetes
Service/Ingress; scale it on request rate, since its job is proxying, not computation. Rate-limiter state (if using Redis-backed limits, as above) must live outside the gateway’s own pods so limits hold correctly across replicas. - Testing strategy. Test each BFF’s aggregation logic against WireMock stubs of its downstream services (as in Chapter 2’s contract test), and test the gateway’s routing rules with a small number of true end-to-end tests hitting real, deployed BFFs — don’t try to unit test declarative YAML routing rules; verify them by exercising the routes.
- Migration strategy. Introduce the gateway first, routing directly to existing services with no BFF (a pure reverse-proxy step), to get cross-cutting concerns centralized with minimal risk; add BFFs afterward, once a second client type’s divergent needs actually justify the split.
8. Common Mistakes
- Putting business logic in the gateway. A gateway route filter that inspects request bodies and makes business decisions (e.g., rejecting orders over a certain value) creates undocumented, hard-to-test logic outside any service’s ownership. Fix: the gateway enforces generic, infrastructure-level policy (auth, rate limits, routing); business rules live in the owning service.
- One BFF trying to serve every client “to avoid duplication.” A single “universal” BFF with conditional logic branching on a client-type header reproduces the original one-size-fits-all problem one layer down. Fix: if clients’ needs genuinely diverge, give them genuinely separate BFFs; if they don’t diverge, you don’t need more than one BFF (or any BFF) in the first place.
- The gateway calling multiple services itself instead of delegating to a BFF. Cramming aggregation logic into gateway filters (rather than a dedicated BFF service) makes that logic hard to test, deploy, and reason about independently of routing. Fix: keep the gateway’s job to routing and cross-cutting policy; aggregation belongs in a BFF (or aggregator service).
- No independent timeout budget per hop. Letting the gateway’s timeout be shorter than the sum of a BFF’s downstream call timeouts guarantees confusing, generic failures under load. Fix: explicitly budget and document each hop’s timeout, and make sure they compose (gateway timeout > BFF’s own worst-case downstream timeout, with margin).
- Skipping the gateway for internal service-to-service calls. Routing
order-service’s call toinventory-servicethrough the public-facing API gateway (rather than direct internal networking or a service mesh) adds unnecessary latency and conflates two different traffic classes with two different concerns. Fix: the gateway handles client-to-cluster traffic; internal service-to-service traffic uses direct networking or a service mesh (a later chapter), not the same edge component. - Building BFFs before there are two different clients. Standing up
web-bffandmobile-bffwhen only the web client exists yet, “for when mobile ships,” is speculative infrastructure with no client to validate it against. Fix: build a shared API (behind a gateway) first; split into BFFs when the second client’s genuinely different needs are real and known, not anticipated.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Multiple services directly exposed, each duplicating auth/TLS/rate limiting | Yes (gateway) | Centralizes cross-cutting concerns applied uniformly | — |
| Two+ client types with genuinely different data/aggregation needs | Yes (BFF) | Avoids a lowest-common-denominator API punishing every client | — |
| Single client type, single shared contract works fine | No (BFF) | Splitting contracts you don’t need is pure overhead | Shared API behind a gateway |
| Only one backend service exists | No (gateway) | Nothing to route between yet | Secure the single service directly |
| Aggregation logic is business-critical and multi-step (writes, not reads) | No (BFF alone) | A read-aggregating BFF isn’t the right place for write orchestration | Saga or a dedicated orchestrator service |
10. Hands-On Exercise
Extend it: add a third BFF, partner-bff, for a hypothetical third-party integration partner that should only see order status (not payment details at all) and is subject to a much stricter rate limit than either web or mobile. Configure its gateway route accordingly.
Simulate a failure: make payment-service return HTTP 500 for every request, and observe what mobile-bff’s confirmation endpoint returns today (per Section 6, it fails closed). Then modify it to return a partial MobileOrderConfirmation with totalCharged = null instead. Which behavior is more correct for this specific screen, and why might the answer differ for a different endpoint?
Decision question, with justification required: Northwind is now asked to support a third client — a smart-speaker voice assistant integration — that needs only a single field (status) from the confirmation flow. Do you build voice-bff, route it through mobile-bff with a stripped-down response, or add a query parameter to an existing BFF that trims the response? Justify against Section 4’s overengineering warning and Section 1’s simplicity-vs-duplication trade-off.
11. Key Takeaways
- An API gateway centralizes cross-cutting, infrastructure-level concerns (auth, TLS, rate limiting, routing) applied uniformly to all client traffic — it is not the place for business logic or multi-service orchestration.
- A backend for frontend exists to give genuinely different clients genuinely different contracts, aggregated and shaped for their specific needs — it’s justified by real client divergence, not anticipated future clients.
- Gateway and BFF are separate decisions: you can centralize cross-cutting concerns without splitting contracts, and you can split contracts (in a very small system) without a shared gateway.
- Keep aggregation in a BFF read-oriented and simple; multi-service write orchestration belongs to a pattern that owns that responsibility explicitly, not to a BFF by accident.
- Every hop added (gateway → BFF → service) needs its own timeout budget that composes correctly with the others, or failures become confusing and generic under load.
- Correlation IDs must originate at the gateway — the earliest point in the request’s life — and propagate through every hop, or debugging a four-process request becomes guesswork.
- Don’t build BFFs speculatively; build the shared API first and split only when a second client’s real, divergent needs justify the operational cost of another deployable.