Temporal
A durable workflow engine that stores event history, lets workers replay workflow code, and runs retried Activities from task queues.
Also worth naming: Temporal Server · Temporal Service · Temporal Cloud · Cadence-related durable workflow model
Temporal is useful when a business process lasts longer than one request and cannot forget which side effects already happened. Treat it as a workflow engine with strict replay rules, not as a nicer queue.
What it is
Temporal is a durable workflow engine. You write Workflow code for long-running decision logic, run Worker Processes that poll for work, and let the Temporal Service persist progress. Temporal defines a Workflow Execution as durable, reliable and recoverable after failure in its Workflow Execution overview. Read Workflow orchestration first for the general pattern; this page is the concrete Temporal shape.
The Service is not where your business code runs. Temporal's Worker docs say Worker Processes poll Task Queues, execute Workflow or Activity code, and run outside the Temporal Service, while the Service orchestrates state transitions and dispatches tasks Workers. That split is why you scale your workers for CPU, I/O and credentials, while the Service owns history, task routing and timers.
Use Temporal when the workflow has multiple dependent steps, long waits, human messages, compensation, or a need to answer "what already happened?" from one history. Use Long-running tasks or a queue plus a state table when the job is small: one or two steps, short retries, simple status, no human wait and no compensation chain.
When to reach for it
Reach for this when…
- A customer-visible process spans several services and can partially succeed.
- You need durable timers, external messages, retries and compensation in one inspectable history.
- A worker crash must resume from recorded progress rather than rerun the whole flow.
- A queue plus a jobs table is growing many states, cron repairs and manual recovery paths.
Not really this pattern when…
- The work is one background job with simple retry and status; use a queue, a state table, or Temporal Standalone Activities.
- You only need publish-subscribe fan-out or a retained event stream; use a message broker or Kafka-like log.
- The process needs low-latency request and response only; keep the hot path synchronous and enqueue follow-up work.
- The team cannot operate deterministic workflow code, Activity idempotency and versioning yet; start smaller.
How it works
1. The Service stores progress; workers execute code
Temporal's server has four independently scalable services: Frontend for API routing and authorization, History for Workflow mutable state, queues and timers, Matching for user-facing Task Queues, and the internal Worker Service for background system work Temporal Server. Persistence stores Tasks, Workflow mutable state, the History table, Namespace metadata and Visibility data. Temporal documents Cassandra, PostgreSQL and MySQL as tested persistence databases, with SQLite meant for development and testing rather than production Persistence.
Clients call Frontend. History persists workflow state and timers. Matching hosts Task Queues. Your Worker Processes poll those queues and run Workflow and Activity code outside the Service.
A Namespace is the isolation boundary for Workflow IDs, Task Queues, retention and access configuration. Temporal says Task Queues and Workflow Executions belong to a Namespace, and a Workflow ID is unique inside that Namespace Namespaces. A Task Queue is a lightweight queue that one or more workers poll over synchronous RPC; Temporal notes that workers poll only when they have spare capacity and that Workflow and Activity Tasks persist in a Task Queue if a worker goes down Task Queues.
2. Event History is the source of truth
Temporal tracks each Workflow Execution by appending Events such as Activity scheduled, Timer fired, Signal received and Workflow completed. Its Event History docs call this an append-only, durably persisted log that lets the application state survive crashes and also acts as an audit log Events and Event History. When a worker resumes a Workflow, it replays the history and compares generated Commands with existing Events Workflow Execution.
A Workflow Worker receives a task, replays prior Events, emits the next Commands, and the Service appends the resulting Events. Completed Activities are read from history rather than run again.
That replay model creates the main rule: Workflow code must be deterministic. Temporal states that Workflow code must make the same Workflow API calls in the same sequence for the same input, and that non-deterministic work such as API calls, database queries and LLM calls should go in Activities Workflow Definition. Use workflow-safe time and random APIs. Do not branch on local wall-clock time inside workflow code.
3. Activities are the side-effect boundary
An Activity is a normal function for one well-defined action, such as calling a payment API or writing a derived file Activity Definition. Temporal recommends idempotent Activities because completed Activities do not rerun during replay, but a worker can crash after the external side effect and before the completion is recorded, causing a retry Activity idempotency. Link this directly to Idempotency: the Activity needs a business key that makes a repeated call safe.
A Workflow schedules an Activity. The Activity Worker reports completion, failure or heartbeats. Retry policy and timeouts decide whether the Service schedules another attempt.
Activities retry by default. Temporal documents the default Retry Policy as an initial interval of 1 second, a 2.0 backoff coefficient, a maximum interval of 100 seconds, unlimited maximum attempts and no non-retryable errors Retry Policies. Activity timeouts decide how failures are detected: Start-To-Close bounds a single attempt, Schedule-To-Close bounds the whole Activity Execution including retries, and Heartbeat Timeout detects a long-running Activity that stopped heartbeating Detecting Activity failures.
4. Messages, timers and code changes are first-class
Temporal has three workflow message types. Signals are asynchronous writes, Queries read current state without recording an Event, and Updates are synchronous writes that can return a result; Temporal's feature guide shows the differences in a comparison table Workflow message passing. These are the clean way to model approval, cancellation, progress reads and admin intervention without inventing a side channel.
Timers are durable. The TypeScript timer docs say a Workflow can sleep for months, timers are persisted, and the worker process is not tied up during the wait Timers. Running workflows also need a change plan. Temporal's TypeScript versioning docs warn that adding, removing or reordering await calls on command-producing APIs such as Activities and timers can cause replay problems, and recommend Worker Versioning or Patching Versioning.
Performance envelope
Temporal envelope: limits and defaults from Temporal docs; throughput and latency depend on workflow shape, payload size, persistence and worker capacity.
| Dimension | Documented anchor | Design consequence |
|---|---|---|
| Workflow duration | [No time constraint](https://docs.temporal.io/workflow-execution/event#time-constraints) on how long a Workflow Execution can run | Long waits are fine, but history growth and code versioning still need design. |
| Event History count | [Warns after 10,240 Events; limited to 51,200 Events](https://docs.temporal.io/workflow-execution/limits) | Looping workflows and chatty messages should Continue-As-New before the limit. |
| Event History size | [Warns after 10 MB; limited to 50 MB](https://docs.temporal.io/workflow-execution/limits) | Keep payloads small and store large data in external storage with references. |
| Activity retry defaults | [1s initial interval, 2.0 backoff, 100s maximum interval and unlimited attempts](https://docs.temporal.io/encyclopedia/retry-policies#default-values-for-retry-policy) | Set explicit policies for permanent failures, user-visible deadlines and noisy dependencies. |
| Activity timeouts | [Every Activity needs Start-To-Close or Schedule-To-Close](https://docs.temporal.io/evaluate/features/timeouts-and-retries#activity-timeouts) | Start-To-Close is the usual crash detector; long Activities also heartbeat. |
| Pending operations | [2,000 incomplete Activities, Child Workflows, Signals or cancellations per Workflow Execution by default](https://docs.temporal.io/workflow-execution/limits) | Chunk large fan-outs or use child workflows rather than scheduling huge batches at once. |
Capabilities in interviews
Durable code workflows
Write the process as code while Temporal persists its Event History and resumes it after failures.
Use this for order fulfilment, onboarding, provisioning and subscription flows where the process state matters more than any single worker. The Workflow code decides the next step; the History Service records Events; a Worker replays history and issues the next Command. The general pattern is covered in Workflow orchestration; the Temporal-specific constraint is deterministic replay.
Choose this variant when
- The business process has branches, waits and partial success.
- You need one place to inspect progress and recovery.
- You want code workflows rather than a hand-rolled state machine.
Activities with retry policies and timeouts
Put external side effects in Activities, then set retry, timeout and heartbeat policy per operation.
Activities are where payment calls, database writes, emails and file work happen. Temporal retries Activities by policy, but the external system may see a repeated attempt, so use Idempotency keys. Use Start-To-Close for one attempt, Schedule-To-Close for the total budget across retries, and Heartbeats when a long Activity can report progress.
Choose this variant when
- The call can fail transiently and should be retried.
- The external operation can be made duplicate-safe.
- You need progress or cancellation for a long-running Activity.
Signals, Queries and Updates
Interact with a running Workflow without reaching into worker memory or a side table.
Signals are asynchronous writes, Queries read current state, and Updates write state while returning a result. Use them for approval, cancellation, progress reads, cart-style mutations and operator fixes. This keeps the Workflow as the state owner instead of splitting truth across a database row, a queue message and a worker process.
Choose this variant when
- A human or external system must affect an open workflow.
- The caller needs a progress read without mutating state.
- The caller needs a validated write and a synchronous result.
Durable timers and safe workflow changes
Model waits as timers and deploy workflow changes with versioning or patches.
Temporal timers are durable, so a workflow can wait through worker restarts without sleeping a process. The cost is that open histories were produced by old code. When a change adds, removes or reorders command-producing calls, use Worker Versioning, patching or a new Workflow Type so old histories still replay safely.
Choose this variant when
- The process waits for deadlines, callbacks, expiry or review.
- You deploy while old workflows are still open.
- A change affects Activities, timers or child workflow sequencing.
Operating knobs
Namespace boundaries
A Namespace gives Workflow ID uniqueness, resource isolation and configuration boundaries such as retention Namespaces. Use separate Namespaces for environments, regulated domains or teams that need independent access controls and limits. Do not rely on a Namespace alone for multi-tenant product data unless naming, auth and visibility are designed around it.
Task Queue and worker split
Task Queues route work to workers that registered the right Workflow and Activity types. Split queues when operations need different worker fleets, credentials, regions or throttles. Keep all workers on one queue compatible, because Temporal says a worker that accepts an unknown Workflow or Activity Task will fail that Task Task Queues.
Retry and timeout policy per Activity
The documented default retry policy is useful for transient failure, but it is too open-ended for many product promises. Set Start-To-Close based on one attempt, Schedule-To-Close based on the user-visible budget, Heartbeat Timeout for long work, and non-retryable errors for permanent failures. The policy should match the external system and the compensation path.
History budget and Continue-As-New
A Workflow can run for a long time, but the Event History has count and size limits. Continue-As-New closes the current run, passes latest relevant state to a new run with the same Workflow ID, and starts a fresh history Continue-As-New. Use it for entity workflows, long loops, high message counts and periodic version cutovers.
Self-hosted vs Temporal Cloud
Self-hosting means operating the Temporal Service, persistence, visibility, upgrades, security and monitoring yourself; the self-hosted guide covers deployment, namespaces, security, monitoring and visibility Self-hosted guide. Temporal Cloud runs the Service for you while your workers remain in your environment; the Cloud overview says Temporal operates persistence, replication, version upgrades and availability, while workers execute your code Temporal Cloud overview.
Versus the alternatives
Temporal against nearby choices.
| Dimension | Temporal | Queue plus state table | Event broker / saga | Managed state machine |
|---|---|---|---|---|
| State owner | Workflow Execution with Event History | Your jobs table and worker code | Distributed across event listeners | Provider state-machine service |
| Best fit | Long-lived orchestration with retries, timers, messages and compensation | Small async jobs with simple status | Independent services reacting to facts | Cloud-native integration flows and explicit state diagrams |
| Replay model | Deterministic workflow code replays from history | No replay unless you build it | Consumers replay from a log if available | Provider-specific execution semantics |
| Side effects | Activities with retry and timeout policy | Worker handler code | Event consumers and local transactions | Tasks, service integrations or functions |
| Main risk | Non-deterministic code, history growth and Activity idempotency mistakes | Reimplementing workflow features in SQL and cron | Implicit full journey and scattered compensation | Definition sprawl and provider-specific limits |
- For the general orchestration decision, use [Workflow orchestration](/learn/workflow-orchestration).
- For choreographed compensation, use [Event-driven / saga](/learn/patterns/event-driven-saga).
Failure modes & gotchas
A queue delivers work; Temporal owns workflow progress. If the process is one email send or one report build, a queue plus a small table is easier. Temporal earns its place when the process has durable state, timers, messages, compensation or a history that operators must inspect.
Local time, random branching, direct HTTP calls or database reads inside Workflow code can produce different Commands during replay. Temporal documents deterministic constraints and says non-deterministic operations belong in Activities Workflow Definition. Keep Workflow code as decision logic over recorded facts.
An Activity can perform its external side effect, then the worker can crash before reporting completion. Temporal may retry because no completion Event exists. For payment, booking, emails and writes, use an idempotency key or unique business ID so a repeated attempt observes the prior result rather than causing a duplicate.
Without a useful Start-To-Close Timeout, a worker crash can leave detection until a broad timeout. Without Heartbeats, a long Activity cannot report progress or receive cancellation promptly. Temporal recommends Start-To-Close and Heartbeats for long-running Activities Detecting Activity failures.
High-frequency signals, tight loops and large payloads can push a run towards the documented Event History limits: warning at 10,240 Events or 10 MB, hard limit at 51,200 Events or 50 MB Workflow Execution limits. Continue-As-New at safe checkpoints and store large payloads elsewhere.
Running histories expect the command sequence produced by older code. Reordering Activities, adding a timer in the middle, or removing a command-producing call can fail replay. Use Worker Versioning, patching, or a new Workflow Type when the command sequence changes Versioning.
In production
Ticket checkout platform (illustrative)
Charge, book, review and refund as one workflow
A checkout service starts a Workflow when a seat hold is created. Payment capture, ticket booking, refund and receipt email are Activities with idempotency keys. Review approval is a Signal or Update, and status reads are Queries. The useful property is not higher throughput than a queue; it is that one Event History records the exact branch the checkout took.
Infrastructure provisioning control plane (illustrative)
Provision, poll, repair and deprovision over time
A provisioning API starts a Workflow per resource. Activities call cloud APIs to create, tag, verify and deprovision resources. Timers schedule periodic checks, Updates request configuration changes, and Continue-As-New rolls history forward after repeated maintenance cycles. A queue could run the first create call, but it would not naturally own the whole lifecycle.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“A checkout charges a card, books a ticket, waits for possible manual review and may need a refund. Would you use Temporal?”
Weak answer
"Yes. I would put each step on a Temporal queue and let retries handle failures. If something fails, I put it in a dead-letter queue and an operator can fix it."
Strong answer
"I would use one Workflow Execution per checkout only if the process really needs a durable owner. The Workflow decides the sequence: charge, book, wait for review or cancellation, then confirm or compensate. Payment, booking, refund and email are Activities with idempotency keys, because retries can execute an Activity more than once. Review approval is a Signal or Update, progress reads are Queries, and deadlines are durable timers. I would watch history size and Continue-As-New only at a safe boundary. If this were just one receipt email after a local transaction, I would keep a queue plus a small jobs table."
Why it wins: It separates Workflow code from Activities, names the message types, handles duplicate-safe side effects, includes timers and compensation, and still rejects Temporal when a simple queue is enough.
Interview playbook
When it comes up
- The prompt has a multi-step business process that can partially succeed.
- The process waits for a callback, human approval, deadline or cancellation.
- A worker crash must not lose progress or repeat side effects unsafely.
- A queue-and-table design is starting to need compensation, signals and audit history.
Order of reveal
- 11. Name the durable owner. I would model the process as a Temporal Workflow Execution with one durable Event History for progress.
- 22. Split decisions from side effects. Workflow code decides what happens next; Activities touch payment, databases, email and other external systems.
- 33. Explain replay. Workers replay history after failures, so workflow code must be deterministic and old histories need a versioning plan.
- 44. Add retry and timeout policy. Each Activity gets explicit Start-To-Close, Schedule-To-Close, Heartbeat and retry settings that match the user-facing deadline.
- 55. State the fallback. If the flow stays a small one-step job, I would use a queue plus a state table instead.
Signature phrases
- ““Temporal remembers progress, not just work to do.”” — Separates a workflow engine from a queue.
- ““Workflow code replays; Activities do side effects.”” — Captures the main implementation boundary.
- ““Retries require idempotent Activity design.”” — Prevents overpromising around external systems.
- ““Continue-As-New before history becomes the limit.”” — Shows you know the scaling limit on a single run.
Likely follow-ups
?“How do workers scale?”Reveal
Workers poll Task Queues. Add more compatible Worker Processes to a Task Queue for more capacity, or split queues when work needs different credentials, hardware, regions or throttling. The Temporal Service does not execute your business code; your worker fleet does.
?“How do you keep a workflow open for months without huge memory use?”Reveal
Use durable timers and recorded messages. A sleeping workflow does not pin a worker process. The state that matters is in Event History, so the worker can be evicted and later replay the history when a timer fires or a message arrives.
?“How do you change workflow code safely?”Reveal
Compatible changes are fine, but adding, removing or reordering command-producing calls needs Worker Versioning, patching, or a new Workflow Type. I also use replay tests against real histories before rolling out changes.
Worked example
Numbers in this section are illustrative.
Setup. Suppose a ticket checkout holds a seat, charges a customer, waits for possible fraud review, books the ticket and sends a receipt. The failure to reason about is payment succeeded but booking or review failed.
Workflow ID. Start one Workflow per checkout, using a stable business ID such as the hold ID. Client retries at the API boundary still need an idempotency key, but the Workflow gives the process one durable owner.
Activities. Charge payment, book ticket, refund payment and send receipt are Activities. Each one receives a business idempotency key. The payment Activity can be retried after a timeout without charging again, because the provider sees the same key.
Messages and timers. Manual review arrives as a Signal or Update. A progress endpoint uses a Query. A deadline is a durable timer, so no worker process sleeps in memory while the customer waits.
Compensation. If booking reaches a terminal failure after payment succeeds, the Workflow calls refundPayment. A refund is a new business action with its own retries and alerts. This is the concrete Temporal implementation of the saga shape from Event-driven / saga.
History management. If the checkout run is short, history stays small. For entity-style workflows that receive many messages over time, carry latest state into Continue-As-New before hitting the Event History warning thresholds.
When not to use it. If the task were only "send receipt later", a queue plus a jobs table from Long-running tasks would be enough. Temporal becomes useful when payment, booking, review, cancellation, refunds and audit are one customer-visible process.
Cheat sheet
- •Temporal = durable Workflow Execution + Event History + Worker Processes polling Task Queues.
- •Service pieces: Frontend, History, Matching and internal Worker Service; persistence stores history and mutable state.
- •Workflow code must be deterministic because workers replay it from history.
- •Activities do side effects and need idempotency because retries can execute them more than once.
- •Retry defaults are [1s initial, 2.0 backoff, 100s max interval and unlimited attempts](https://docs.temporal.io/encyclopedia/retry-policies#default-values-for-retry-policy).
- •Set Activity Start-To-Close, Schedule-To-Close and Heartbeat Timeout deliberately.
- •Signals write async, Queries read, Updates write and return a result.
- •History limits [warn at 10,240 Events or 10 MB; limit at 51,200 Events or 50 MB](https://docs.temporal.io/workflow-execution/limits).
- •Use Continue-As-New for long loops, entity workflows and version cutovers.
- •Use a queue plus state table when the state machine is small.
Drills
Numbers in this section are illustrative.
Why must Workflow code be deterministic?Reveal
A worker may replay Workflow code from Event History after a crash. If the code emits a different sequence of command-producing calls, the Service cannot match Commands to Events safely and replay can fail. Keep non-deterministic work in Activities or workflow-safe APIs.
An Activity charged a card, then the worker crashed before reporting completion. What prevents a duplicate charge?Reveal
Temporal can retry the Activity because no completion Event was recorded. The Activity must call the payment provider with the same idempotency key, usually derived from the Workflow ID and Activity identity, so the provider returns or observes the original charge instead of making a second one.
When do you use Continue-As-New?Reveal
Use it when a run is accumulating too much Event History, receives many messages, loops for a long time, or should move to newer code at a safe checkpoint. Continue-As-New passes the latest relevant state to a new run with a fresh history.
When is a queue plus state table enough?Reveal
It is enough when the work has one or two states, short retries, simple status, no human messages, no durable timers across long waits, and no compensation chain. Temporal is useful when those concerns become the core of the process rather than incidental plumbing.
What it is