Series overview
Part 5 of 2818% complete
2026-06-10•17 min read

Service discovery and centralized configuration

Northwind now runs order-service, inventory-service, payment-service, an api-gateway, and two BFFs (Chapter 4) — six deployables, each currently configured with hardcoded URLs and its own application.yml checked into its own repository.

1. Problem the Pattern Solves

Northwind’s platform team starts autoscaling inventory-service for flash sales — sometimes two replicas, sometimes eight. order-service’s InventoryClient (Chapter 2) points at a single hardcoded hostname, which worked when there was exactly one inventory-service instance behind a fixed address, but now needs to reach whichever replicas are currently healthy, without order-service being redeployed every time the replica count changes.

Separately, a security audit requires rotating the shared Redis password used by the gateway’s rate limiter (Chapter 4) across all six services within 24 hours. Today that means six separate pull requests, six separate deploys, and six chances to make a typo in a connection string — for a value that isn’t even application logic, just configuration that happens to be duplicated everywhere it’s needed.

Forces in tension:

  • Static addressing vs. dynamic infrastructure. Hardcoded hostnames assume a service’s location is fixed; container orchestration (Kubernetes scaling, rolling deploys, spot-instance replacement) means it never is.
  • Configuration duplication vs. central control. Config in each service’s own repo is simple to reason about locally but means any cross-cutting change (a rotated secret, a feature flag, a shared timeout value) requires N coordinated deploys.
  • Availability vs. an extra moving part. A discovery service or config server is itself infrastructure that must be highly available — if it goes down, does the rest of the platform go down with it? This pattern’s design has to answer that question explicitly, not by accident.
  • Security. Centralizing configuration also centralizes secrets — a config server holding every service’s database password is a high-value target that needs its own access control, distinct from the services that consume it.

2. Core Idea

Service discovery lets a service find the current network location of another service by name, without a hardcoded address, through either a client-side registry (each client queries a registry and load-balances itself — e.g., Netflix Eureka) or server-side discovery (the platform’s own infrastructure resolves the name — e.g., Kubernetes’ built-in DNS-based Service objects, which route to whichever pods are currently healthy).

Centralized configuration stores configuration — non-secret application properties and, via a secrets backend, sensitive values — in one place that every service fetches from at startup (and optionally refreshes at runtime), instead of each service carrying its own static config file as the sole source of truth.

Centralized configuration

Service discovery (Kubernetes DNS, this chapter's choice)

resolve by name

fetch config at startup

fetch secrets

Kubernetes Service objects

inventory-service.default.svc.cluster.local

Spring Cloud Config Server

backed by a Git repo

Secrets: HashiCorp Vault /

Kubernetes Secrets

order-service

inventory-service pod 1

inventory-service pod 2

inventory-service pod 3

Centralized configuration

Service discovery (Kubernetes DNS, this chapter's choice)

resolve by name

fetch config at startup

fetch secrets

Kubernetes Service objects

inventory-service.default.svc.cluster.local

Spring Cloud Config Server

backed by a Git repo

Secrets: HashiCorp Vault /

Kubernetes Secrets

order-service

inventory-service pod 1

inventory-service pod 2

inventory-service pod 3

Stack decision, stated up front. Northwind runs on Kubernetes (established from Chapter 1’s production-concerns sections onward), so this chapter uses Kubernetes’ built-in DNS-based service discovery rather than adding a separate discovery registry like Eureka. This is a deliberate choice worth explaining: Eureka and similar client-side registries solve discovery for platforms without a built-in answer (VM-based deployments, for instance); on Kubernetes, the platform already provides server-side discovery for free, and adding Eureka on top would be duplicate infrastructure solving an already-solved problem. For configuration, this chapter uses Spring Cloud Config Server backed by Git, because — unlike discovery — Kubernetes’ ConfigMap/Secret primitives handle static, per-deployment config well but don’t natively provide the dynamic refresh-without-restart behavior Spring Cloud Config offers for values that change between deploys (feature flags, tunable timeouts).

Commonly confused with:

  • Load balancing. Discovery answers “where are the current instances of X”; load balancing answers “which one do I send this request to.” Kubernetes’ Service object actually does both (DNS resolution plus round-robin/iptables-based balancing across pod IPs) — a client-side registry like Eureka historically separated the two, requiring a client-side load balancer (Ribbon, or Spring Cloud LoadBalancer) alongside it.
  • Feature flags. Centralized configuration can carry feature-flag values, but a feature flag also needs runtime toggling, targeting rules, and often a UI — a dedicated flag system (touched on in this series’ deployment chapter) usually outgrows a plain config server.
  • A service mesh’s control plane. A mesh’s control plane also distributes configuration (routing rules, mTLS certificates) to its sidecars, but that configuration is about traffic behavior between services, not application-level business configuration — different scope, covered in the service mesh chapter.

3. When to Use It

Strong indicators:

  • More than a handful of service instances, with counts that change dynamically (autoscaling, rolling deploys) — static addressing breaks down exactly at this point.
  • Configuration values (timeouts, feature flags, shared connection details) are duplicated across multiple services’ config files and have drifted or needed synchronized rotation before.
  • You’re running on infrastructure (VMs, bare metal, or an orchestrator without built-in service discovery) that doesn’t already solve name resolution for you.

Concrete use cases:

  • E-commerce during peak events: Northwind’s flash-sale autoscaling of inventory-service is the canonical case — instance counts change by the hour, and every caller must resolve current, healthy instances rather than a fixed list.
  • SaaS platforms with per-environment configuration: the same service code running in staging, sandbox, and production needs environment-specific config (different rate limits, different downstream URLs) fetched centrally rather than baked into separate build artifacts per environment.
  • Logistics platforms with regional deployments: services deployed per region need region-specific configuration (currency, tax rules, carrier integrations) that a central config server can serve per profile, without regional code forks.
  • Multi-team platforms rotating shared secrets: any organization with a compliance requirement to rotate credentials on a schedule benefits from one place to update a secret that every consuming service picks up, rather than N separate deploys per rotation.

Prerequisites:

  • A deployment platform’s discovery mechanism understood and chosen deliberately (Kubernetes DNS here) rather than defaulting to a library because a tutorial used it — the wrong choice adds an unnecessary component, as explained above for Eureka-on-Kubernetes.
  • A secrets management strategy separate from general configuration — never store credentials in the same Git-backed config repo as non-sensitive properties, even if the tooling makes it technically possible.

4. When Not to Use It

  • A handful of instances at fixed, rarely-changing addresses. Two or three services deployed as single, long-lived instances (no autoscaling, infrequent redeploys) don’t need dynamic discovery — a static configuration file naming their addresses is simpler and has one less moving part to operate.
  • Kubernetes already solves discovery — don’t add a second registry on top “for consistency” with a non-Kubernetes past project. This is one of the most common overengineering cases in this pattern: teams migrating from VM-based Eureka deployments to Kubernetes often keep Eureka running out of habit, maintaining a redundant, less-integrated discovery layer next to the platform’s native one.
  • Configuration that’s genuinely static per deployment. Values that never change without a full redeploy anyway (the JDBC driver class name, for instance) don’t need to live in a dynamic config server — plain application.yml, baked into the artifact or supplied via a Kubernetes ConfigMap at deploy time, is simpler and avoids a runtime dependency on the config server being available.
  • Risk: the config server as a new single point of failure. If every service fetches its configuration from one config server at startup with no fallback, an outage of that one component prevents every other service from starting — Section 7 addresses mitigating this explicitly.

5. Implementation Example

Discovery: Kubernetes-native, no extra library. order-service reaches inventory-service purely by DNS name — no discovery client dependency needed in the application at all:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/InventoryClientConfig.kt
package `in`.o612.eng.northwind.order.internal
import org.springframework.boot.context.properties.ConfigurationProperties
import org.springframework.context.annotation.Bean
import org.springframework.context.annotation.Configuration
import org.springframework.web.client.RestClient
@ConfigurationProperties(prefix = "northwind.inventory-service")
data class InventoryServiceProperties(val baseUrl: String)
@Configuration
class InventoryClientConfig(private val props: InventoryServiceProperties) {
@Bean
fun inventoryClient(): InventoryClient =
InventoryClient(RestClient.builder().baseUrl(props.baseUrl).build())
}
order-service — config fetched from Spring Cloud Config, not a local file
northwind:
inventory-service:
base-url: http://inventory-service.default.svc.cluster.local

inventory-service.default.svc.cluster.local is a Kubernetes Service DNS name — it resolves to whichever pods are currently Ready, load-balanced automatically, regardless of whether there are two replicas or eight. order-service’s code and configuration never mention an instance count or an IP address.

k8s/inventory-service.yaml
apiVersion: v1
kind: Service
metadata:
name: inventory-service
spec:
selector:
app: inventory-service
ports:
- port: 80
targetPort: 8080
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: inventory-service
spec:
replicas: 3
selector:
matchLabels: { app: inventory-service }
template:
metadata:
labels: { app: inventory-service }
spec:
containers:
- name: inventory-service
image: northwind/inventory-service:1.4.0
readinessProbe:
httpGet: { path: /actuator/health/readiness, port: 8080 }
initialDelaySeconds: 10

The readinessProbe is what makes discovery correct, not just present — Kubernetes only routes traffic to pods that pass it, so a pod still warming up its connection pool is invisible to order-service until it’s actually ready to serve.

Centralized configuration: Spring Cloud Config Server, Git-backed:

config-server/build.gradle.kts
dependencies {
implementation("org.springframework.cloud:spring-cloud-config-server")
}
config-server/src/main/kotlin/in/o612/eng/northwind/config/ConfigServerApplication.kt
package `in`.o612.eng.northwind.config
import org.springframework.boot.autoconfigure.SpringBootApplication
import org.springframework.boot.runApplication
import org.springframework.cloud.config.server.EnableConfigServer
@EnableConfigServer
@SpringBootApplication
class ConfigServerApplication
fun main(args: Array<String>) = runApplication<ConfigServerApplication>(*args) {}
config-server/src/main/resources/application.yml
spring:
cloud:
config:
server:
git:
uri: https://github.in/o612/eng/northwind/platform-config
default-label: main

Each service fetches its own config at startup by name and active profile — order-service, profile production, resolves platform-config/order-service-production.yml from the Git repo:

order-service/src/main/resources/bootstrap.yml
spring:
application:
name: order-service
config:
import: "configserver:http://config-server.default.svc.cluster.local"
platform-config/order-service-production.yml (in the Git repo, not the service repo)
northwind:
inventory-service:
base-url: http://inventory-service.default.svc.cluster.local
payment-service:
base-url: http://payment-service.default.svc.cluster.local
resilience4j:
ratelimiter:
instances:
inventoryReservation:
limitForPeriod: 200

The Redis password rotation from Section 1 now means one commit to platform-config, not six pull requests — every service picks up the new value the next time it refreshes (see the actuator refresh endpoint below), with no code change and no rebuild.

Runtime refresh without a restart, for values that should apply immediately rather than only at next deploy:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/RateLimitProperties.kt
package `in`.o612.eng.northwind.order.internal
import org.springframework.boot.context.properties.ConfigurationProperties
import org.springframework.cloud.context.config.annotation.RefreshScope
import org.springframework.stereotype.Component
@RefreshScope
@Component
@ConfigurationProperties(prefix = "northwind.rate-limit")
data class RateLimitProperties(var requestsPerSecond: Int = 50)

A POST /actuator/refresh to a running instance re-reads its config and rebinds any @RefreshScope bean — used sparingly, for values that genuinely need to change without a redeploy (a temporary rate-limit override during an incident, for instance), not as a substitute for a proper deploy pipeline.

6. Step-by-Step Flow

Tracing “platform team rotates the shared Redis password used by the gateway’s rate limiter”:

api-gatewayConfig Serverplatform-config repoPlatform engineerapi-gatewayConfig Serverplatform-config repoPlatform engineerNext Redis connection uses new password;no pod restart, no redeploycommit new Redis passwordPOST /actuator/refreshGET /api-gateway/productionfetch latest platform-configupdated YAMLupdated propertiesrebind @RefreshScope beans
api-gatewayConfig Serverplatform-config repoPlatform engineerapi-gatewayConfig Serverplatform-config repoPlatform engineerNext Redis connection uses new password;no pod restart, no redeploycommit new Redis passwordPOST /actuator/refreshGET /api-gateway/productionfetch latest platform-configupdated YAMLupdated propertiesrebind @RefreshScope beans
  1. Client action (of this operational flow): a platform engineer merges a config change to platform-config.
  2. API request equivalent: the engineer calls each affected service’s /actuator/refresh endpoint (or triggers it via a fleet-wide script, or a webhook from the Git repo).
  3. Service behavior: each service’s Config Server client re-fetches its configuration.
  4. Database interaction: none in this flow — this is a pure configuration-propagation path.
  5. Inter-service communication: the config server itself pulls from Git; consuming services pull from the config server — a strict one-directional fan-out, never the reverse.
  6. Error or failure handling: if the config server is unreachable during a refresh call, the service keeps running on its last-known-good configuration (Spring Cloud Config’s client fails the refresh call but does not crash the running application) — this fallback behavior is exactly why Section 7 treats config-server availability as a startup concern, not a runtime one.
  7. Observability signals: the config server logs every fetch by service name and profile; each consuming service logs a refresh event with the properties that changed (values redacted for anything secret-shaped).
  8. Final response/outcome: every replica of api-gateway is using the rotated password within seconds of the refresh call, with zero pod restarts and zero client-visible disruption.

7. Production Concerns

  • Availability of the config server itself. Because services fetch config at startup, a config-server outage during a rolling deploy or a scale-up event prevents new pods from starting correctly. Mitigate with Spring Cloud Config’s local file-based fallback cache (spring.cloud.config.fail-fast=false plus a cached copy) so a new pod can start with its last-successfully-fetched config if the server is briefly unreachable — never let the config server’s uptime become the platform’s uptime ceiling.
  • Secrets vs. configuration. Never put credentials directly in the Git-backed config repo, even a private one — Git history is forever. Use Spring Cloud Config’s backend integration with Vault, or better, keep secrets in Kubernetes Secret objects (backed by a KMS) injected as environment variables, and reserve the config server for non-sensitive values.
  • Data consistency. Configuration changes are not transactional across services — a rolling refresh means, briefly, some replicas are on the old value and some on the new. Design any config value that must be consistent across all replicas simultaneously (a routing-critical flag, for instance) with that propagation delay in mind, or use a feature-flag system with an explicit, atomic activation model instead.
  • API versioning. Not directly relevant here, but service discovery does interact with deployment strategy — Kubernetes’ rolling updates mean both old and new versions of a service can be resolvable simultaneously during a deploy, which the blue-green/canary chapter addresses directly.
  • Authentication and service-to-service trust. The config server itself needs access control — not every service should be able to fetch every other service’s configuration profile. Scope config-server access per service identity, and audit fetches the same way you’d audit any secrets access.
  • Logging, metrics, tracing. Track config-fetch latency and failure rate as a first-class metric — a slowly degrading config server manifests as slow service startups platform-wide, which is easy to misdiagnose as “something’s wrong with the services” instead of the shared dependency.
  • Kubernetes deployment, health probes, autoscaling. Kubernetes’ own Service/Endpoints objects are the actual discovery mechanism in production, as shown in Section 5; the config server is a separate Deployment that itself needs a readiness probe and, ideally, more than one replica for its own availability.
  • Testing strategy. Test service startup against a local, static config file in unit and integration tests — don’t make every test suite depend on a running config server. Reserve a small number of true integration tests that verify a service actually starts correctly against a real config server instance.
  • Migration strategy. Introduce the config server first for genuinely shared, frequently-rotated values (the Redis password case), leaving stable, per-service values in local application.yml — migrating every property to central config on day one adds risk for values that never needed it.

8. Common Mistakes

  1. Running a separate discovery registry on Kubernetes out of habit. Deploying Eureka alongside Kubernetes’ own DNS-based discovery, because a previous project used it, adds a redundant, less-integrated component. Fix: use the platform’s native discovery mechanism; only add a library-based registry when the deployment platform doesn’t already provide one.
  2. Storing secrets in the Git-backed config repo. Even a private repository’s history retains every value ever committed — a “temporary” password in a YAML file is a permanent leak. Fix: route secrets through Vault or Kubernetes Secrets, never through the plain config server backend.
  3. No fallback when the config server is unreachable at startup. A hard dependency on the config server with no cached fallback means a config-server blip cascades into every other service failing to start during a redeploy. Fix: enable a local fail-safe cache and treat config-server availability as a monitored SLO, not an assumption.
  4. Treating /actuator/refresh as a substitute for a real deploy pipeline. Using runtime config refresh to change values that should really go through code review and a proper rollout (business rule thresholds, for instance) bypasses the safety net a deploy pipeline provides. Fix: reserve dynamic refresh for genuinely operational values (rate limits, timeouts, temporary overrides); route everything else through normal deploys.
  5. Hardcoding a service’s address anywhere after discovery is in place. A stray RestClient.create("http://10.2.3.4:8080") left in test fixtures or a debug script silently breaks the moment that pod is rescheduled. Fix: always resolve service addresses through the same DNS name the rest of the platform uses, even in scripts and local tooling.
  6. Migrating every configuration value to the central server at once. This makes an unrelated config-server hiccup capable of affecting every service’s startup simultaneously, for values that mostly never change. Fix: migrate incrementally, starting from the values that actually motivated the change (shared, rotatable secrets and cross-cutting settings).

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Instance counts change dynamically (autoscaling, rolling deploys)Yes (discovery)Static addresses break the moment topology changes—
Deployed on Kubernetes alreadyYes, but use built-in DNSThe platform already solves discovery; a separate registry duplicates itAdd a registry only if the platform lacks one
Shared config/secrets duplicated and drifting across servicesYes (central config)One place to change, one place to audit—
Two or three long-lived, fixed instances, no autoscalingNoStatic config is simpler with fewer moving partsLocal application.yml per service
Value never changes without a full redeploy anywayNo (central config)No benefit to dynamic fetch for a value that’s effectively staticBake into the deploy artifact or a Kubernetes ConfigMap

10. Hands-On Exercise

Extend it: add a new configuration value — a feature flag, northwind.features.express-checkout-enabled — to the config server, marked @RefreshScope in order-service, and verify you can toggle it via /actuator/refresh without a restart.

Simulate a failure: stop the config server entirely, then start a fresh order-service pod. Does it start successfully using a cached fallback, or does it fail? If it fails, that’s the gap Section 7 warns about — go implement the local fallback cache and re-run the test.

Decision question, with justification required: Northwind is evaluating whether to move from Kubernetes’ built-in DNS discovery to a full service mesh (a later chapter) partly for more sophisticated traffic-aware discovery (e.g., routing based on service version, not just health). Given the operational cost a mesh adds (a later chapter details this), what specific traffic pattern or failure mode would need to show up in production before that trade-off is worth making? Name the forces from Section 1 you’re weighing.

11. Key Takeaways

  • Service discovery answers “where is a healthy instance of X right now” — on Kubernetes, the platform already answers this via DNS-based Service objects; don’t add a separate registry unless your platform lacks one.
  • Centralized configuration solves duplicated, drifting, hard-to-rotate config and secrets across services — it’s justified by real duplication pain, not by “best practice” alone.
  • Keep secrets out of the config server’s Git backend entirely; route them through a dedicated secrets manager with its own access control and audit trail.
  • A config server is a new dependency every other service’s startup relies on — give it its own availability plan (replicas, a local fallback cache) rather than letting it become a silent single point of failure.
  • Reserve runtime config refresh for genuinely operational values; route anything that should go through review and rollout through a normal deploy instead.
  • Readiness probes are what make discovery correct, not just present — a pod that’s technically running but not ready to serve must be invisible to the platform’s routing.
  • Migrate configuration to the central server incrementally, starting from the values that actually motivated the move, not all at once.
Spring BootKotlinMicroservices

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind