Developer Workflows
Durable Workflow Checkpoints: Resume Long Jobs Without Starting Over
Design workflow checkpoints with explicit stages, verified outputs, leases, and recovery tests so long-running jobs can resume safely after interruption.
In this article
A long-running job can fail after doing most of its useful work. The process disappears, its temporary directory is gone, and the next attempt starts from zero because nobody recorded what had actually finished. Adding another retry doesn't solve that missing memory.
Durable workflow checkpoints preserve progress outside the worker that performs the task. They let a new worker distinguish completed work, uncertain work, and work that has never started. This matters for report generation, media processing, data imports, and scheduled publishing pipelines alike.
Model stages instead of one giant function
Break the workflow into observable stages with explicit inputs and outputs. A fictional report job might collect authorised source data, normalise records, generate a draft, validate it, store the final artifact, and verify the stored artifact. Each stage needs a clear completion condition.
Avoid a single done: true flag. It can't explain whether validation failed, storage was interrupted, or the final result was created but not confirmed. Use states that support recovery, such as pending, running, verified, and failed, with a reason and an update time.
Temporal's workflow-execution documentation describes persistence and recovery as core workflow properties. You don't need to adopt that product to use the design principle: progress must survive the individual process, and recovery must use recorded state rather than guesswork.
Store checkpoints in a durable location
Choose storage with durability appropriate to the job: a database, a durable workflow service, or a reliable object store plus metadata. A worker's local scratch folder is convenient for intermediate work but isn't a dependable recovery record when the worker can be replaced.
A simplified checkpoint could look like this:
{
"run_id": "report-example-001",
"stage": "validation",
"input_version": "source-set-7",
"draft_object_key": "reports/example/draft.json",
"validated": false,
"attempt": 2
}The values are illustrative. Keep authentication tokens and private content out of checkpoint metadata where possible. Use JSON Formatter to inspect a sanitised state record, and remember that readable structure doesn't prove the referenced object still exists.
Record evidence of completion
A stage should become verified only when its output passes the checks needed by the next stage. For a file, that may mean successful parsing, required records, a content digest, and a confirmed storage identifier. For a data import, it may mean a committed transaction and checked invariants.
Store the input version too. If the source data changes after a checkpoint, decide whether to resume the old snapshot or begin a new run. Silently mixing an old draft with a new source set makes the result difficult to explain and can invalidate earlier checks.
Don't confuse a successful network response with a verified final outcome. A storage API can acknowledge a request while your application loses the returned identifier. Recovery should inspect the expected destination before submitting another final artifact.
Keep only one worker responsible for a run
Use a lease, lock, or workflow-engine guarantee to prevent two workers from independently advancing the same run. Leases should expire so a crashed worker doesn't block progress forever. Renew them while work continues, and reject writes from a worker whose lease is no longer valid.
A fencing token or version check can protect against a paused worker returning after another worker takes over. Without it, the old worker may overwrite a newer checkpoint. Record state transitions atomically when possible rather than updating several unrelated flags separately.
Cancellation needs its own state. A cancelled job should not be revived by a generic recovery sweep. Decide which temporary artifacts can be cleaned up, which final artifacts remain useful, and how the user can request a deliberate restart.
Handle uncertain external outcomes explicitly
Some failures happen between an external action and your checkpoint update. Perhaps a final upload succeeded, but the worker crashed before recording it. Mark that outcome as uncertain and reconcile it against the destination using a stable run identifier or expected artifact name.
The important question is not “did this worker remember success?” but “what does the destination now contain?” Confirm the content and identity before resuming. Don't delete or overwrite an unfamiliar artifact merely because it resembles an expected output.
Keep final publication separate from drafting. Validation should happen before an artifact becomes visible as finished work. A useful related pattern appears in safe database migrations: intermediate progress and a verified release are different events.
Test recovery at awkward boundaries
In staging, stop a worker after every important transition. Stop it after creating an output but before recording that output. Let a lease expire, launch a replacement worker, and check that the old worker cannot commit stale state when it returns.
Compare sanitised checkpoint snapshots with JSON Diff. Verify that resumed jobs preserve their input version, skip verified stages, and repeat only the work needed. Also test missing intermediate objects and corrupted drafts; checkpoints shouldn't blindly trust references that no longer work.
Track recovery age and repeated failures. A job that remains uncertain for hours needs an operational alert, not endless silent retries. Set a bounded retry policy and give the user a clear partial-result or failure message when recovery cannot proceed safely.
Conclusion
Durable checkpoints turn a long job into a series of verifiable transitions. Preserve state outside the worker, record output evidence, coordinate ownership, and test failures between steps. The result is a workflow that resumes deliberately instead of repeatedly hoping a fresh start will work.