Workflow orchestration with BPMN versus distributed sagas
Chapter 11 built a plain Kotlin saga orchestrator for order placement and named BPMN engines as “heavier tooling” for a different class of problem, deferred to this chapter. This closing chapter is where that comparison gets made concrete — and where this series’ Northwind Platform story reaches the process complex enough to actually need it.
1. Problem the Pattern Solves
Northwind Platform (Chapter 27’s white-label expansion) needs a proper returns-and-refunds process, and it looks nothing like Chapter 11’s order-placement saga. A return request triggers: an automated eligibility check (is the order within the 30-day window), then — only if the refund exceeds Chapter 26’s approval threshold — a human customer-service manager must review and approve it, which might take anywhere from minutes to several days depending on their workload. Once approved, the customer has up to two weeks to ship the item back, tracked via a carrier webhook; if it doesn’t arrive within that window, the process needs to escalate to a different manual review queue instead of just failing. Only once the item is confirmed received does the actual refund (Chapter 26’s RefundAuthorizationPolicy, now invoked as one step of a much longer process) get issued.
Chapter 11’s saga orchestrator was designed for a short-lived, fully-automated sequence measured in seconds — its SagaState enum and command-dispatch loop have no natural way to represent “waiting up to two weeks for a package,” no built-in way to model a human task with an assignable queue and a UI for a manager to act on, and no visual representation a non-engineer (a business analyst documenting the return policy, an auditor reviewing the process) could review without reading Kotlin source code.
Forces in tension:
- Process duration and human involvement vs. a code-based orchestrator’s assumptions. Chapter 11’s orchestrator assumes each step resolves quickly and automatically; a process spanning days or weeks with human decision points needs durable, resumable state and a task-assignment model that plain application code doesn’t provide for free.
- Developer ownership vs. business-analyst visibility. A saga expressed in Kotlin is fully within engineering’s control and version-controlled alongside the rest of the codebase — but it’s opaque to a business stakeholder who needs to review or even adjust the process (the 30-day window, the two-week return period) without necessarily going through a full engineering change cycle.
- Tooling weight vs. process complexity. A BPMN engine brings real capabilities (durable timers, human task lists, a graphical process definition) that a hand-rolled orchestrator would have to build from scratch — but it’s also a new, heavier piece of infrastructure with its own operational and learning cost, not justified for every multi-step flow.
- Auditability of a long-running, human-involved process vs. implicit code logic. A regulator or an internal auditor asking “show me exactly what happened to this specific return, including who approved it and when” is far better served by a process engine’s built-in execution history than by reconstructing the story from application logs (even with Chapter 22’s tracing, which wasn’t designed for multi-day-spanning workflows).
2. Core Idea
Distributed sagas (Chapter 11) suit short-lived, fully-automated, developer-owned sequences of service calls with programmatic compensation — order placement’s reserve-then-charge-then-confirm, resolved in seconds, is the right shape for this.
BPMN-based workflow orchestration suits longer-running, potentially human-involved, business-visible processes — a graphical process definition (BPMN, Business Process Model and Notation, a standardized diagram format) that a workflow engine executes, persisting the process’s state durably between steps, providing built-in human task assignment, timers, and boundary events (like “the two-week return window expired”) as first-class modeling constructs rather than something a developer builds by hand.
Notice the diagram itself is the point: this is a BPMN-style process definition, reviewable by a business analyst without reading any Kotlin — and the IssueRefund box is not a reimplementation of Chapter 11’s saga logic, it’s a single step in this larger process that invokes it. BPMN orchestration and sagas aren’t competitors here; the process engine handles the long-running, human-involved shape, and delegates the short automated sub-sequences to exactly the saga pattern Chapter 11 already built.
Participants:
- Process engine — Camunda 8 (Zeebe), this chapter’s choice, a horizontally-scalable, cloud-native BPMN engine with first-class Spring Boot integration — chosen over older BPMN engines specifically for its cloud-native, non-monolithic architecture, consistent with every infrastructure choice this series has made since Chapter 1.
- Process definition — the BPMN diagram itself (an XML file, typically authored in a visual modeler), the actual source of truth for the process’s shape, distinct from and more visible than a hand-written orchestrator’s control flow.
- Human task — a first-class BPMN construct representing work assigned to a person or a queue, with the engine tracking its pending/claimed/completed state durably, something Chapter 11’s orchestrator has no native concept of at all.
- Service tasks — steps where the engine calls out to actual services (invoking
order-service’s eligibility check, or Chapter 11’s saga-based refund issuance) — the connective tissue between the process engine and this series’ existing microservices.
Commonly confused with:
- The Saga pattern (Chapter 11), addressed directly above — sagas are the right tool for short, automated, code-owned sequences; BPMN orchestration is the right tool for longer, human-involved, business-visible ones. This chapter’s returns process uses both: BPMN for the overall shape, a saga (invoked as one service task) for the automated refund-issuance sub-sequence.
- A generic workflow/task queue (like the competing-consumers work queue from Chapter 18). A task queue distributes independent units of work for parallel processing; a BPMN process models a single entity’s (one return request’s) stateful, multi-step, potentially branching journey over time — a fundamentally different shape of problem.
- A simple state machine enum (like Chapter 11’s
SagaStep). A hand-rolled enum-based state machine works well for a small, fixed set of automated transitions; it has no native support for durable timers, human task assignment, or a business-reviewable visual definition — exactly the capabilities that justify the heavier BPMN engine once a process actually needs them.
3. When to Use It
Strong indicators:
- The process spans a duration where “just keep a thread or a saga orchestrator instance waiting” isn’t practical — hours, days, or weeks, as Northwind’s two-week return window requires.
- The process includes genuine human decision points that need task assignment, a work queue, and potentially reassignment or escalation — not just an automated service call that might occasionally be slow.
- Business stakeholders (product, compliance, operations) need to review, audit, or even participate in defining the process’s shape — a visual BPMN diagram serves this need in a way source code cannot.
- Regulatory or audit requirements demand a clear, engine-maintained execution history of exactly what happened to a specific process instance, including who made which human decision and when.
Concrete use cases:
- E-commerce, as here: returns and refunds, dispute resolution, and fraud investigation workflows all combine automation with human judgment over realistic timeframes.
- Insurance claims processing: a claim’s lifecycle (filed, reviewed, possibly escalated to a specialist, approved or denied, paid) is a textbook long-running, human-involved BPMN process.
- Loan or credit approval workflows: multiple automated checks combined with human underwriting decisions, often with regulatory requirements for auditable process history.
- Employee or vendor onboarding: a genuinely cross-departmental, multi-day process with sequential and parallel human approval steps, historically one of BPMN’s most common application domains outside pure software.
Prerequisites:
- A process genuinely complex or long-running enough to justify the engine — Section 4’s overengineering warning is especially important here, since BPMN tooling is easy to reach for reflexively once introduced.
- Integration points (service tasks) connecting the process engine to existing services — this chapter’s
IssueRefundstep calling into Chapter 26’sRefundAuthorizationPolicyand Chapter 11’s saga machinery is exactly this connective work, and it’s real implementation effort, not just drawing a diagram. - Dedicated operational capacity for the process engine itself — Camunda 8, like the service mesh (Chapter 21) and Keycloak (Chapter 26) before it, is another significant piece of shared infrastructure requiring its own availability and operational ownership.
4. When Not to Use It
- A short, fully-automated sequence with no human involvement and no multi-day waiting. Chapter 11’s order-placement saga remains exactly the right tool for exactly that problem — introducing a BPMN engine for it would add substantial tooling weight to solve a problem a plain Kotlin orchestrator (or even choreography) already solves well.
- A process whose “steps” are really just a few sequential method calls with no real branching, waiting, or human decision. Reaching for BPMN because it’s now available in the platform, for a flow that’s genuinely simple, reproduces the same overengineering risk this series has flagged for every heavyweight pattern — evaluate against the process’s actual shape, not the tool’s availability.
- A team with no capacity to operate and maintain a BPMN engine. Camunda (or any process engine) is real infrastructure with its own upgrade path, monitoring needs, and learning curve — introducing it without that investment risks the same under-resourced-adoption problem Chapter 21 warned about for service meshes.
- Overengineering signal: modeling every multi-step business operation in BPMN “for consistency,” including ones that are genuinely simple sagas, diluting the value of having BPMN specifically for the processes that actually need its human-task and long-duration capabilities.
5. Implementation Example
The BPMN process definition (conceptually — normally authored visually in Camunda Modeler, shown here as the underlying XML structure a modeler produces):
<bpmn:process id="returns-process" isExecutable="true"> <bpmn:startEvent id="ReturnRequested" /> <bpmn:serviceTask id="EligibilityCheck" zeebe:type="check-eligibility" /> <bpmn:exclusiveGateway id="ApprovalNeeded" /> <bpmn:userTask id="ManagerApproval" name="Review return over threshold"> <bpmn:extensionElements> <zeebe:assignmentDefinition candidateGroups="customer-service-managers" /> </bpmn:extensionElements> </bpmn:userTask> <bpmn:boundaryEvent id="ReturnWindowTimer" attachedToRef="WaitForItem"> <bpmn:timerEventDefinition><bpmn:timeDuration>P14D</bpmn:timeDuration></bpmn:timerEventDefinition> </bpmn:boundaryEvent> <bpmn:serviceTask id="IssueRefund" zeebe:type="issue-refund-saga" /> <bpmn:userTask id="EscalationReview" name="Item not received in time"> <bpmn:extensionElements> <zeebe:assignmentDefinition candidateGroups="returns-escalation-team" /> </bpmn:extensionElements> </bpmn:userTask></bpmn:process>P14D (ISO 8601 duration) is the two-week return window expressed as a durable timer the engine itself manages — Camunda persists this process instance’s state and simply wakes it up when the timer fires or the awaited event (item received) arrives, whichever comes first, with zero application code needed to implement the waiting itself.
The service task worker, connecting the BPMN process to Northwind’s existing services — this is where the process engine’s abstract steps become real calls into the platform:
package `in`.o612.eng.northwind.returns
import io.camunda.zeebe.client.api.worker.JobHandlerimport io.camunda.zeebe.spring.client.annotation.JobWorkerimport org.springframework.stereotype.Component
@Componentclass EligibilityCheckWorker(private val orderServiceClient: OrderServiceClient) {
@JobWorker(type = "check-eligibility") fun checkEligibility(job: io.camunda.zeebe.client.api.response.ActivatedJob): Map<String, Any> { val orderId = job.variablesAsMap["orderId"] as String val order = orderServiceClient.getOrder(java.util.UUID.fromString(orderId)) val daysSinceOrder = java.time.Duration.between(order.placedAt, java.time.Instant.now()).toDays()
return mapOf( "eligible" to (daysSinceOrder <= 30), "refundAmount" to order.totalAmount, "requiresApproval" to (order.totalAmount > java.math.BigDecimal("100.00")), ) }}The refund-issuance service task, deliberately delegating to Chapter 11’s saga machinery rather than reimplementing it — the concrete demonstration that BPMN orchestration and sagas compose rather than compete:
package `in`.o612.eng.northwind.returns
import `in`.o612.eng.northwind.order.internal.saga.OrderSagaOrchestrator // reused from Chapter 11import io.camunda.zeebe.spring.client.annotation.JobWorkerimport org.springframework.stereotype.Component
@Componentclass IssueRefundWorker( private val refundAuthorizationPolicy: RefundAuthorizationPolicy, // reused from Chapter 26 private val paymentServiceClient: PaymentServiceClient,) { @JobWorker(type = "issue-refund-saga") fun issueRefund(job: io.camunda.zeebe.client.api.response.ActivatedJob): Map<String, Any> { val orderId = job.variablesAsMap["orderId"] as String val amount = java.math.BigDecimal(job.variablesAsMap["refundAmount"].toString())
// The BPMN process handled the LONG-RUNNING, human-involved shape // (approval, waiting for the item). This step is a short, fully // automated sequence — exactly the shape Chapter 11's saga pattern // was built for — invoked here as one BPMN service task. val outcome = paymentServiceClient.issueRefund(java.util.UUID.fromString(orderId), amount) return mapOf("refundStatus" to outcome.status) }}This is the chapter’s central technical point made concrete in code: the BPMN process owns the two-week, human-approval-involving shape; a short, automated saga-style sequence is just one step within it, reusing Chapter 11’s and Chapter 26’s existing components without modification.
6. Step-by-Step Flow
- Client action. A customer initiates a return request, starting a new process instance in the engine.
- API request equivalent. The engine immediately dispatches the
check-eligibilityservice task toEligibilityCheckWorker— a fast, synchronous call, no different in shape from any service call this series has built since Chapter 2. - Service behavior. Because the refund amount exceeds Chapter 26’s threshold, the process creates a human task rather than proceeding automatically — a decision the process definition itself encodes, visible in the diagram.
- Database interaction. The engine persists the process instance’s state (which step it’s on, its variables) to its own durable store — this is what allows the process to survive for two weeks without any application server needing to keep it in memory.
- Inter-service communication. The manager’s approval, arriving through whatever UI Camunda’s Tasklist (or a custom frontend) provides, resumes the persisted process instance — a fundamentally different interaction model than any synchronous or event-driven call this series has used, since it’s mediated by a human’s action on their own schedule.
- Error or failure handling. If the 14-day timer fires before the item-received event arrives, the boundary event redirects the process to the escalation human task instead of failing outright — a branch the process definition itself specifies, not exception-handling code buried in an orchestrator’s failure path.
- Observability signals. Camunda’s own operational tooling (Operate) shows every currently-running process instance’s exact position, its variable values, and its full history — a genuinely different and richer observability surface than Chapter 22’s request tracing, purpose-built for long-running process visibility rather than request-scoped latency debugging.
- Final response/outcome. The refund is issued via the reused saga-style service task, and the entire process — including the manager’s specific approval decision and timestamp — is durably recorded in the engine’s execution history, directly satisfying the audit requirement Section 1 named.
7. Production Concerns
- Timeouts, retries, idempotency. Service tasks (like
IssueRefundWorker) need the same idempotency discipline this series has required since Chapter 1 — the engine may redeliver a job if a worker doesn’t complete it within its configured time, exactly analogous to Kafka’s at-least-once redelivery (Chapter 7) requiring idempotent consumers. - Data consistency and transaction boundaries. The process engine’s own state (which step a process instance is on) is a separate concern from the business data each service task manipulates —
IssueRefundWorkershould ensure its own operation is transactionally sound withinpayment-service, independent of the engine’s own persistence guarantees for the process state itself. - API versioning and backward compatibility. BPMN process definitions themselves need a versioning strategy — a running process instance started under version 1 of the returns process should either complete under that version or have an explicit, deliberate migration path to version 2, not be silently affected by a mid-flight process definition change.
- Authentication and service-to-service trust. Human task assignment (Section 5’s
candidateGroups) should map to the same role structure Chapter 26 established (Keycloak groups/roles), so a manager’s ability to claim and approve a task is governed by the same identity and authorization system as every other authenticated action in the platform. - Logging, metrics, tracing, audit trails. The process engine’s own execution history is often the primary audit trail for these long-running processes — richer and more directly reviewable for this purpose than reconstructing a story from Chapter 22’s distributed traces, which are better suited to request-scoped, seconds-long flows.
- Kubernetes deployment, health probes, autoscaling. Camunda 8’s Zeebe engine is itself a distributed, cloud-native system deployed on Kubernetes, following the same lifecycle and health-probe discipline as Chapter 25 — it needs its own dedicated operational attention, not an afterthought bolted onto existing service deployments.
- Testing strategy. Test service task workers as ordinary unit and integration tests (they’re just Kotlin classes calling existing services); test the process definition’s actual branching and timer behavior using Camunda’s own process-testing tools, which can simulate time passing without actually waiting fourteen real days in a test run.
- Migration strategy. Introduce BPMN orchestration for the first process that genuinely demonstrates the need (Northwind’s returns process, with its real human-approval and multi-week-waiting requirements) rather than retrofitting Chapter 11’s already-working, appropriately-scoped order-placement saga into BPMN for uniformity’s sake.
8. Common Mistakes
- Migrating Chapter 11’s order-placement saga to BPMN “for consistency.” Order placement is short, fully automated, and has no human involvement — exactly the shape a plain saga orchestrator handles well; forcing it into a BPMN engine adds tooling weight with no corresponding benefit. Fix: match the tool to the process’s actual shape (duration, human involvement, audit needs), as Section 4 insists, not to a desire for platform-wide uniformity.
- Reimplementing saga-style automated logic inside BPMN service tasks instead of reusing existing sagas. Duplicating Chapter 11’s and Chapter 26’s refund logic inside a new, BPMN-specific implementation, rather than calling the existing components as this chapter’s
IssueRefundWorkerdoes. Fix: treat BPMN service tasks as integration points into existing services and sagas, not a reason to rewrite logic that already works. - No idempotency on service task workers. Assuming the engine calls a service task exactly once, when in fact job redelivery (analogous to Kafka’s at-least-once semantics) can occur. Fix: design every service task worker to be safely re-runnable, exactly as this series has required for every other message- or job-driven consumer since Chapter 7.
- Changing a process definition while instances are actively running under the old version. Deploying a new version of the returns process without a deliberate migration strategy can leave in-flight process instances in an inconsistent or undefined state relative to the new definition. Fix: version process definitions explicitly and decide, per change, whether in-flight instances complete under their original version or need an explicit migration.
- Adopting a BPMN engine without dedicated operational capacity. Treating Camunda as a drop-in tool rather than significant new infrastructure requiring its own monitoring, upgrade planning, and on-call familiarity. Fix: invest in operating the engine properly, exactly as Chapter 21 insisted for service mesh adoption — the same discipline applies to any significant new platform component.
- Using BPMN’s human task assignment without integrating it into the platform’s actual identity and role system. Managing task-queue membership as a separate, disconnected concept from Chapter 26’s Keycloak roles creates two parallel, potentially inconsistent authorization systems. Fix: map BPMN candidate groups directly onto the platform’s existing role structure, as Section 7 specifies.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Process is short, fully automated, seconds-to-complete | No (BPMN) | A plain saga orchestrator or choreography (Chapter 11) already fits this shape well | Distributed saga |
| Process spans hours to weeks, or includes genuine human decision points | Yes (BPMN) | Durable timers and human task assignment are first-class, not something to build by hand | — |
| Business stakeholders need to review or help define the process shape | Yes (BPMN) | A visual process definition is reviewable without reading source code | — |
| Regulatory/audit requirements need detailed, engine-maintained execution history of a long process | Yes (BPMN) | Purpose-built execution history beats reconstructing a story from request-scoped traces | — |
| No dedicated capacity to operate a new, significant piece of infrastructure | No, not yet | An under-resourced BPMN engine adoption risks becoming a reliability and maintenance burden | Defer until the operational investment is realistic, or handle the process manually in the interim |
10. Hands-On Exercise
Extend it: add a second human task to the returns process for cases where the eligibility check itself is ambiguous (an order exactly at the 30-day boundary, say), routing to a review queue before the approval-threshold branch is even evaluated.
Simulate a failure: using Camunda’s process-testing tools, simulate the 14-day timer expiring before an item-received event arrives, and verify the process correctly routes to the escalation human task rather than hanging indefinitely or failing silently.
Decision question, with justification required: Northwind’s fraud-investigation process (mentioned in Section 3) has a similar shape to returns but involves coordinating with an external law-enforcement reporting system with unpredictable, sometimes multi-month response times. Does BPMN’s timer and human-task model still fit at that duration, or does a months-long external dependency call for a different pattern entirely? Justify using the forces from Section 1.
11. Key Takeaways
- Sagas (Chapter 11) and BPMN orchestration solve different shapes of multi-step problem: sagas fit short, fully-automated, developer-owned sequences; BPMN fits longer-running, human-involved, business-visible processes — they compose, with BPMN service tasks invoking existing sagas rather than replacing them.
- A BPMN process definition’s visual, reviewable nature is a genuine capability a hand-written orchestrator can’t match — valuable specifically when business stakeholders need visibility into or input on the process shape.
- Durable timers and first-class human task assignment are what a process engine provides that a plain Kotlin state machine has to build from scratch — the concrete, specific reason to reach for the heavier tool once a process genuinely needs them.
- Never force a short, fully-automated process into BPMN for platform consistency — match the tool to the process’s actual duration, human involvement, and audit needs, exactly as every pattern in this series has insisted on evidence-driven adoption over uniformity.
- A BPMN engine is real, significant infrastructure requiring dedicated operational investment — the same lesson this series drew for service meshes (Chapter 21) and identity providers (Chapter 26) applies here without exception.
- Service task workers need the same idempotency discipline as every other message- or job-driven consumer in this series, since job redelivery is a real, expected occurrence, not an edge case.
- This chapter closes the series’ pattern catalog, but the underlying discipline threading through every chapter — draw boundaries from real evidence, add complexity only when a concrete, demonstrated need justifies it, and keep every pattern’s cost honestly weighed against its benefit — is the actual, transferable lesson, more durable than any single technology choice made along the way.