Workflows and orchestration
Durable execution for multi-step processes: orchestration vs choreography, sagas, timers, retries and idempotent activities.
When a checkout, review or migration can run for days, the hard part is not the first call. It is remembering what already happened after a crash, retry or human delay.
Read this if your last attempt…
- Your answer says "put it on a queue" but the process has several dependent steps.
- You cannot explain what resumes after a worker crashes halfway through a payment flow.
- You model long waits with cron scans and many nullable columns.
- You mention sagas without naming the coordinator, history or compensation policy.
The concept
The model: workflow code plus a durable history
A workflow engine runs a long-lived business process as code while storing process history outside the worker. Temporal describes a Workflow Execution as durable and recoverable after failure in the Workflow Execution overview, where replay checks Commands against Event History so a worker can rebuild local state.
The engine persists the workflow history. Workers replay workflow code from that history, then call activities, timers or signal handlers for side effects and waits.
Pick the smallest coordination model that preserves the business invariant.
| Approach | Use when | Watch out for |
|---|---|---|
| Queue plus state table | One or two async steps, short waits, simple retry state and no compensation chain. | You own leases, retries, timeouts, poison rows, schema changes and operator visibility. |
| Choreographed events | Services should stay independent and each event has local meaning even if no central flow exists. | The full user journey can become implicit. Add correlation IDs, audit views and idempotent consumers. |
| Temporal-style workflow engine | You want code workflows, replay, durable history, Activities, timers and messages; Temporal documents these in [Workflow Execution](https://docs.temporal.io/workflow-execution), [Activity Definition](https://docs.temporal.io/activity-definition) and [Sending messages](https://docs.temporal.io/sending-messages). | Workflow code must stay deterministic, and running workflows need a versioning plan. |
| AWS Step Functions | You want a managed AWS state machine. AWS says Standard workflows can run for up to one year, while Express workflows run for up to five minutes in [Choosing workflow type](https://docs.aws.amazon.com/step-functions/latest/dg/choosing-workflow-type.html). | The workflow type is immutable after creation. Standard uses exactly-once workflow execution and state-transition billing. Asynchronous Express is at-least-once; Synchronous Express is at-most-once. Express billing uses executions, duration and memory [Choosing workflow type](https://docs.aws.amazon.com/step-functions/latest/dg/choosing-workflow-type.html). |
| Cadence | You are in the Cadence ecosystem or comparing the durable-function family. Cadence describes workflows as durable execution functions that survive process restarts, infrastructure failures and arbitrary pauses in [Cadence Workflows](https://cadenceworkflow.io/docs/concepts/workflows). | The same Cadence page says workflow state recovery uses event sourcing and workflow code must be deterministic. |
- Workflow engines reduce coordination code. They do not remove the need for idempotent Activities, careful compensation and operational alerts.
- Step Functions Standard and Express are different products for different durations and execution semantics; do not treat them as one generic box.
How interviewers grade this
- You draw a durable workflow owner, not only a queue and workers.
- You say what lives in workflow code and what lives in Activities.
- You explain replay and determinism before promising crash recovery.
- You choose orchestration when one place must own retries, timers, signals and compensation.
- You link the duplicate-safe parts to Idempotency.
- You can say when Long-running tasks are enough without a workflow engine.
Variants
Durable code workflow
Write the business process as code; the engine records enough history to replay it after failures.
This is the Temporal and Cadence family. It fits when the business process has branches, waits, side effects and messages that are clearer as code than as a diagram-only state machine. The workflow function stays deterministic. Activities call the outside world. Timers and messages are part of the durable history rather than ad hoc rows in a database.
Pros
- +Natural fit for branching business logic.
- +A single history gives auditability and recovery.
- +Long waits do not tie up worker processes.
Cons
- −You must learn replay and determinism rules.
- −Activity boundaries and idempotency are design work.
- −Versioning mistakes can break running executions.
Choose this variant when
- Checkout, onboarding, subscription billing, fulfilment, infrastructure provisioning and manual review flows.
Declarative state machine
Model the flow as states and transitions in a managed service.
AWS Step Functions is the common interview example. It is strongest when your flow is mostly service integrations, waits, choices, maps and retries. The state-machine definition is explicit and operationally visible. The trade-off is that complex business branching can become definition-heavy, and Standard versus Express matters because AWS documents different duration and execution semantics for them.
Pros
- +Clear visual state machine.
- +Managed integrations and retries.
- +Strong fit inside AWS-centric systems.
Cons
- −Provider limits and semantics shape the design.
- −Complex code-like branching can be awkward.
- −Type choice is not something to hand-wave.
Choose this variant when
- AWS-native workflows, data pipelines, approval flows and integration-heavy orchestration.
Choreographed events
Let services listen for facts and publish the next facts.
Orchestration gives one workflow owner the next-step decision. Choreography lets services react to events and publish the next event.
Choreography works when each event is useful on its own and no central owner needs to decide the next step. It is often a good fit for fan-out: analytics, search indexing, notification and audit consumers. It is weaker when the product needs one owner for a customer-visible process, because retries and compensation can scatter across listeners.
Pros
- +Loose coupling between services.
- +Easy to add subscribers.
- +A log can support replay and audit consumers.
Cons
- −Harder to inspect the whole journey.
- −Compensation can become implicit.
- −Event schema changes need discipline.
Choose this variant when
- Event fan-out, eventually consistent projections and independent subscribers.
Queue plus state table
A small job table and a queue are enough for simple asynchronous work.
A job table holds status and attempts. A worker leases one row through a queue message and updates the table. This is enough when the state machine is small.
This is the humble default for simple jobs. Store job_id, status, attempts, next_run_at, locked_until and result pointer. Push job_id to a queue. Workers lease, run and update the table. Once you add many timers, human signals, multiple compensations and versioned long-running branches, you are reimplementing a workflow engine in pieces.
Pros
- +Few moving parts.
- +Easy to explain and operate for simple jobs.
- +No replay constraints.
Cons
- −State machine logic spreads across workers and SQL.
- −Long waits become polling or delayed queue hacks.
- −Audit history is only as good as the rows you remember to write.
Choose this variant when
- Image processing, report generation, import jobs and one-step webhooks.
Worked example
Numbers in this section are illustrative.
Scenario: ticket checkout for an illustrative concert platform. Payment capture can succeed while booking fails because the hold expired, a service is unavailable, or a worker crashes. The product requirement is simple: charge once, book once, and refund if the charge cannot end in a ticket.
1. Start one workflow per checkout. Use a workflow id such as checkout:{hold_id}. If the client retries the start request, the same logical checkout is reused or rejected by policy. The API boundary still uses Idempotency.
2. Charge through an idempotent Activity. The workflow calls chargePayment with an idempotency key derived from the workflow id and activity id. If the payment provider times out, the Activity retry policy backs off. If the worker crashes after the provider charged but before completion is recorded, the retry uses the same key and observes the prior charge instead of charging again.
3. Book the ticket. The workflow calls bookTicket after the charge is known. Booking is idempotent on hold_id or ticket_id. If it succeeds, the workflow sends confirmation and completes.
4. Compensate on failure. If booking reaches a terminal failure after payment succeeded, the workflow calls refundPayment. The compensation is not magic rollback; it is another business action that can retry and can itself fail temporarily. History records charge, booking failure, refund request and refund completion.
A workflow records which steps completed. If booking fails after payment succeeds, it runs a refund compensation and records the terminal state.
5. Wait for people and clocks. For example, the workflow waits an illustrative 15 minutes for a payment callback, one day for manual fraud review, or three days before closing an unresolved refund task. These numbers are product policy, not platform claims. Durable timers mean the worker does not need to sit in memory during those waits.
When not to use a workflow engine here: if checkout were just "enqueue an email receipt after a local transaction", a queue plus email_jobs table would be enough. The workflow engine becomes useful because payment, booking, refunds, waits, retries and human review are one customer-visible process.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“You have an order flow: charge card, reserve inventory, book shipment, then email the customer. How do you make it survive crashes and partial failure?”
Weak answer
"I would put each step on a queue. If something fails, the worker retries, and if it still fails we put it in a DLQ."
Strong answer
"I would make the order a durable workflow. The workflow history records which step has completed, so after a worker crash a new worker replays the history and continues from the next decision. Side effects are Activities: charge payment, reserve inventory, book shipment, send email. Each Activity is idempotent because retries can execute it more than once. If shipment fails after payment and inventory succeeded, the workflow runs compensations in reverse order, such as release inventory and refund payment. For simple one-step jobs I would not use this machinery; a queue plus state table is enough."
Why it wins: It names the durable owner of the flow, separates workflow code from side-effecting Activities, handles crash recovery through history and replay, includes compensation, and still knows when a simpler queue design is enough.
When it comes up
- A user action crosses services and can partially succeed.
- A process waits for hours or days, such as review, payment settlement or expiry.
- The interviewer asks what happens after a worker crashes mid-flow.
- You need retries, compensation, audit history and human approval in one place.
Order of reveal
- 1Name the process owner. "I need one durable owner for the business process, not just independent queue messages."
- 2Separate workflow code from Activities. "The workflow decides what happens next. Activities do side effects and are idempotent because retries can re-run them."
- 3Explain crash recovery. "The engine persists an event history. A new worker replays that history and reaches the same decision point before continuing."
- 4Add compensation and timers. "If a later step fails, I run compensating Activities for the steps that already committed. Long waits are durable timers, not sleeping threads."
- 5Close with the simpler fallback. "If it stays one async step, I will keep the queue and table."
Signature phrases
- ““A queue remembers work; a workflow remembers progress.”” — Separates delivery from process state.
- ““Workflow code is deterministic; Activities touch the outside world.”” — Shows you understand replay.
- ““Compensation is a business action, not rollback.”” — Avoids overpromising ACID across services.
- ““If it is only one async step, I will keep the jobs table.”” — Prevents over-engineering.
Likely follow-ups
?“How is this different from a saga?”Reveal
A saga is the business pattern: local transactions plus compensations. A workflow engine is one way to implement an orchestrated saga with durable history, timers and retries. If the interviewer wants the pattern in depth, point to Event-driven / saga.
?“What do you store in the workflow history?”Reveal
Store decisions and results needed to replay: scheduled Activities, completions, failures, timers, signals and terminal state. Do not treat it as a large blob store. Put big payloads in object storage and keep references in history.
?“How do you handle a human approval that arrives late?”Reveal
The workflow waits on a durable timer and a signal or update. If the approval signal arrives before the deadline, continue. If the timer fires first, follow the timeout branch. Late signals are either ignored, rejected or treated as a new workflow depending on product policy.
?“What happens when you deploy a new workflow version?”Reveal
New executions can use the new path. Existing executions either stay on old workers, take a patched branch, or continue as a new workflow at a safe boundary. The unsafe move is to reorder command-producing calls underneath open histories.
Code examples
import { condition, proxyActivities } from '@temporalio/workflow';
import type * as activities from './activities';
const { chargePayment, bookTicket, refundPayment, sendReceipt } =
proxyActivities<typeof activities>({
startToCloseTimeout: '30 seconds',
retry: { maximumAttempts: 5, initialInterval: '2 seconds' },
});
export async function checkoutWorkflow(input: CheckoutInput): Promise<void> {
let approved = false;
setApprovalSignalHandler(() => { approved = true; });
const charge = await chargePayment({
paymentId: input.paymentId,
idempotencyKey: 'checkout:' + input.checkoutId + ':charge',
});
const approvedInTime = await condition(() => approved, '1 day');
if (!approvedInTime) {
await refundPayment({ chargeId: charge.id });
return;
}
try {
await bookTicket({ holdId: input.holdId });
await sendReceipt({ checkoutId: input.checkoutId });
} catch (err) {
await refundPayment({ chargeId: charge.id });
throw err;
}
}CREATE TABLE jobs (
job_id TEXT PRIMARY KEY,
status TEXT NOT NULL CHECK (status IN ('pending', 'running', 'succeeded', 'failed')),
attempt INTEGER NOT NULL DEFAULT 0,
next_run_at TIMESTAMPTZ NOT NULL DEFAULT now(),
locked_until TIMESTAMPTZ,
result_uri TEXT,
last_error TEXT
);
-- Worker leases a due job, runs it once, then updates status/result.
-- If this table grows compensation branches, human signals and long timers,
-- you are approaching a workflow engine design.Common mistakes
A queue delivers work. A workflow engine owns process state, timers, messages, retry policy and event history. If you need to answer "which business step already committed?", use durable state that can answer that question.
A replayed workflow must issue the same commands. Local random numbers, wall-clock checks, direct HTTP calls and database reads can change on replay. Put those in Activities or workflow-safe APIs, then record their results through the engine.
The engine can retry an Activity, but it cannot make a payment provider forget a duplicate request. External calls still need idempotency keys, unique business ids or dedupe tables.
A refund is a new business action, not time travel. It can fail, be delayed, be partial, or require support review. Strong answers model compensations as first-class Activities with their own retries and alerts.
Running histories were produced by old code. Reordering Activities, adding timers in the middle, or removing command-producing calls can make replay fail. Use worker versioning, patching or a new workflow type for incompatible changes.
A one-step import job with status and retry count does not need replay, signals or compensation. A simple table is easier to debug and operate.
Practice drills
Numbers in this section are illustrative.
Why does workflow code need to be deterministic?Reveal
The worker may replay the code from stored history after a crash. If the code emits a different sequence of timers, Activity schedules or completions during replay, the engine cannot safely know what state it is in. Determinism keeps replay aligned with history.
A payment Activity succeeds at the provider, then the worker crashes before reporting completion. What protects the customer?Reveal
The Activity can be retried, so the provider call needs an idempotency key based on the logical operation. The retry uses the same key and observes or returns the first charge result instead of charging again.
When is choreography better than orchestration?Reveal
Use choreography when services publish facts that are useful independently and no central owner needs to decide the next step. Analytics, search indexing and notifications often fit. Use orchestration when the customer-visible process needs one place for retries, timers, compensation and audit.
What makes a queue plus jobs table enough?Reveal
The flow has a small number of states, short waits, simple retries, no human signals and no compensation chain. The table can store status, attempts, lease and result pointer without turning into an implicit workflow engine.
Deep dives
History, replay and the Activity boundary
The mental model that prevents most mistakes is: workflow history is the source of truth for progress, not worker memory. A workflow worker is a stateless executor for durable state. It pulls a workflow task, replays history, rebuilds local variables and then issues the next command. If the worker dies, another worker can replay the same history. This is why the design feels like event sourcing but is presented to the author as ordinary code.
The boundary is important. Workflow code should be pure decision logic over already-recorded facts. It may schedule an Activity, start a timer, wait for a signal, or complete. It should not call the payment API directly, read a database for a fresh random branch, or use local time to decide whether to skip a step. Those values can change between original execution and replay. Activities are where external uncertainty lives.
The Activity boundary also explains idempotency. Completed Activities are recorded in history, so replay does not re-run them. But a started Activity can finish its external side effect and fail to report back. The engine may schedule another attempt because it has no completion event. The external call therefore needs an idempotency key or business unique id. The workflow engine gives you durable orchestration; it does not guarantee external systems are duplicate-safe.
Keep history lean. Large payloads, uploaded files and generated reports belong in blob storage or a database. Store references and small decision results in the workflow. If history becomes enormous, many engines offer continue-as-new or a similar rollover technique, but the interview-safe answer is to keep history as process evidence, not as your data lake.
Temporal, Step Functions and Cadence without hype
Temporal is the reference mental model for this lesson because its docs expose the core ideas plainly: Workflow Executions, replay, Event History, Activities, Signals, Timers and versioning. Use it when you want to write the workflow as application code and are willing to follow deterministic replay rules.
AWS Step Functions is the comparison when the architecture is AWS-native and the flow fits a managed state machine. The key interview detail is not the logo. It is the workflow type. AWS documents Standard workflows as long-running, durable and auditable, with exactly-once execution. Express workflows are high-volume and short-lived; asynchronous Express is at-least-once, while synchronous Express is at-most-once AWS Step Functions. Choose deliberately.
Cadence is another durable-function workflow engine. Its docs describe workflows as fault-oblivious stateful functions with durable timers and event handlers, and they describe event-sourced recovery plus deterministic workflow restrictions. If a candidate has used Cadence, the transferable idea is the same: durable orchestration through replayed history and side-effecting Activities.
The practical comparison question is operational fit. Are you already running the engine? Do you need code workflows or a declarative state machine? Are the Activities idempotent? How will you debug stuck executions? How will you version code for open workflows? Those answers matter more than naming a tool.
Cheat sheet
- •A queue remembers work; a workflow remembers progress.
- •Workflow code decides; Activities touch external systems.
- •Replay requires deterministic workflow code.
- •Activities need idempotency because retries can re-execute them.
- •Orchestration has one owner; choreography spreads ownership across event listeners.
- •A saga is local steps plus compensations; the workflow engine can orchestrate it.
- •Durable timers model days-long waits without sleeping workers.
- •Signals and Updates model humans and external events.
- •Version running workflows with worker versions, patches or new workflow types.
- •Use a queue plus state table when the state machine is small.
Practice this skill
These problems exercise Workflows and orchestration. Try one now to apply what you just learned.
Read this if