Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 07 / Durable Execution & Integration

Accepted is a promise.

Durable, Replayable Execution

Work accepted by an interface should survive downstream failures, resume from a verified checkpoint, and finish in an explicit state.

Maturity
Reference design / Specified
Critical uncertainty
Accepting work does not prove that an external destination completed it. Completion requires reconciliation or an explicitly owned exception.
Human boundary
Resolve exceptions, correct inputs, authorize guarded replay or compensation
Applications
Web forms · CRM synchronization · Document processing · Batch imports
Four accepted workloads converge on a durable execution record, pass through eight checkpointed layers, and end as completed, owned exception, or recoverable failure.

Problem shape

The problem beneath them

Persist accepted work before acknowledging it, advance it through explicit replay-safe states, reconcile external effects, and leave every attempt in a state with a named owner.

Distributed work crosses failure boundaries. A process can crash after committing locally but before notifying a queue. A worker can time out after a destination committed an effect. A retry can run against changed source data. A batch can contain completed, rejected, and unknown items at the same time. “Try again” is therefore not a complete recovery policy.

The pattern separates intent, attempt, and effect. Intent records what the system accepted. Attempts record executions and their evidence. Effects identify externally observable changes. State transitions connect them without pretending they form one atomic transaction. This separation makes uncertain outcomes representable: an attempt may be unknown even though the accepted intent remains recoverable.

Durability also requires ownership. Exhausting retries is not a final state unless a person or service owns the resulting exception and has enough evidence to act. Likewise, “sent” is not “completed” when the destination is authoritative. Completion should follow a destination receipt, read-after-write check, or another explicit reconciliation rule.

Accepting work does not prove that an external destination completed it. Completion requires reconciliation or an explicitly owned exception.

This pattern governs the lifecycle of accepted work. It does not make a non-idempotent destination transactional, guarantee that a third party retains data, or decide the business validity of an input. Those concerns need destination-specific controls and policy.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Web forms

The user saw success. Did the work survive?

An interface confirms receipt before downstream work finishes.

02 / CRM integration

The destination was down. Where is the accepted record?

A remote write fails after the source has acknowledged the request.

03 / Document processing

Page 18 failed. Must pages 1–17 run again?

A multi-stage job retains no trustworthy progress boundary.

04 / Batch imports

The job stopped halfway. What is safe to replay?

Partial effects exist, but their relationship to the input is unclear.

All four questions converge into Durable, Replayable Execution.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Acceptance boundary

    Question
    At what point may the interface promise that work has been accepted?
    Responsibility
    Validate the minimum envelope, assign a stable work identifier, and durably commit intent before acknowledging receipt.
    Input
    Request payload, caller identity, tenancy, policy context, and any client idempotency key.
    Output
    Durable acceptance record and acknowledgement correlated to its identifier.
    Stops when
    The record is committed or the caller receives an explicit rejection with no implied acceptance.
  2. 02

    Immutable input

    Question
    Can a replay reconstruct the exact material input originally accepted?
    Responsibility
    Preserve the canonical payload, source references, version identifiers, and integrity metadata without later mutation.
    Input
    Accepted envelope and retrievable source artifacts.
    Output
    Versioned input snapshot or content-addressed references with provenance.
    Stops when
    Material inputs are frozen, integrity-checkable, and subject to an explicit retention policy.
  3. 03

    Explicit state machine

    Question
    Which transitions are legal, and who may perform them?
    Responsibility
    Represent pending, active, waiting, reconciling, failed, and final states with guarded transitions.
    Input
    Acceptance record, current state, event, actor, and transition preconditions.
    Output
    Atomic state transition plus timestamped transition evidence.
    Stops when
    The event is applied once, rejected as illegal, or retained for investigation.
  4. 04

    Durable checkpoint

    Question
    What verified progress can a later attempt reuse safely?
    Responsibility
    Commit stage outputs and their input/version fingerprints at meaningful restart boundaries.
    Input
    Stage result, input identity, code or policy version, and validation evidence.
    Output
    Checkpoint that is reusable only under matching preconditions.
    Stops when
    The checkpoint and state transition are durably associated.
  5. 05

    Retry policy

    Question
    Is another attempt safe and likely to improve the outcome?
    Responsibility
    Classify failures, schedule bounded retries with backoff, and prevent permanent errors from cycling.
    Input
    Failure class, attempt history, destination health, deadline, and retry budget.
    Output
    Retry schedule, deferred state, or exception route.
    Stops when
    Work succeeds, reaches its retry/deadline bound, or a non-retryable condition is established.
  6. 06

    Idempotency protection

    Question
    Could replay duplicate an internal or external effect?
    Responsibility
    Give each effect stable identity and suppress, deduplicate, or reconcile repeated attempts.
    Input
    Work identifier, effect key, destination scope, and prior effect evidence.
    Output
    One authorized effect, an existing result, or an ambiguous-effect state.
    Stops when
    Uniqueness is enforced or uncertainty is routed without repeating the effect blindly.
  7. 07

    Destination reconciliation

    Question
    What does the authoritative destination say actually happened?
    Responsibility
    Compare intended and observed state using receipts, stable destination identifiers, or safe read-back.
    Input
    Intended effect, attempt evidence, destination response, and observable destination state.
    Output
    Confirmed completion, confirmed absence, conflict, or unresolved status.
    Stops when
    Destination state is confirmed or an owned exception records why confirmation is unavailable.
  8. 08

    Explicit final state

    Question
    How does the lifecycle end without abandoned work?
    Responsibility
    Record a controlled outcome, owner, evidence, retention, and permitted recovery action.
    Input
    State history, reconciliation result, policy, and unresolved conditions.
    Output
    Completed, owned exception, or recoverable failure.
    Stops when
    No work remains in an ownerless or semantically ambiguous terminal state.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
No durable acceptance record Reject or ask caller to retry The system cannot honestly promise recovery The caller reports success but no record can be located
Matching completed effect key Return the prior result Repeating the effect adds risk without new intent Destination state conflicts with stored evidence
Transient destination outage Defer and retry with bounds Availability may recover without changing intent Deadline or retry budget is exhausted
Permanent validation rejection Stop automatic retry Identical input will reproduce the rejection Correction requires business authority
Timeout after dispatch Reconcile before retry The destination may have committed despite no response The effect cannot be queried safely
Checkpoint fingerprint mismatch Recompute from the last valid boundary Prior output belongs to different material inputs Recalculation could invalidate downstream effects
Unknown final destination state Create owned exception Neither success nor absence is proven Consequence or age exceeds policy tolerance

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Accept, snapshot, transition through verified checkpoints, apply one effect, reconcile, and finalize.

Final state
Completed.
Owner
Execution service until reconciliation confirms completion.
Evidence
Acceptance ID, input fingerprint, transitions, checkpoint records, effect key, and destination receipt.
Recovery
Replay from the last compatible checkpoint if later verification invalidates completion.
Ambiguous path

Preserve intent, stop duplicate dispatch, inspect destination state, and route unresolved evidence.

Final state
Owned exception.
Owner
Named operations or domain queue with authority to reconcile or correct.
Evidence
Attempt chronology, timeout or conflict, destination queries, and operator disposition.
Recovery
Confirm the existing effect, authorize a guarded replay, compensate, or close with reason.
Failure path

Classify the failure, retain valid checkpoints, exhaust only the permitted retry budget, and pause.

Final state
Recoverable failure.
Owner
Platform service for technical repair, then execution service for resumption.
Evidence
Failure class, stack or protocol detail, affected stage, checkpoint fingerprint, and retry history.
Recovery
Repair the dependency or code, validate compatibility, and resume from the verified boundary.

Authority map

Capability does not grant authority.

RULE

May decide
Validate transitions, retry classes, budgets, and uniqueness constraints
May not decide
Declare an unobserved external effect complete
Required evidence
Rule version, evaluated inputs, and result

MODEL

May decide
Classify an error or propose exception context when policy permits
May not decide
Cause replay, suppress an exception, or assert completion alone
Required evidence
Model/version, bounded input, confidence, and explanation

SYSTEM

May decide
Persist intent, schedule attempts, enforce transitions, and reconcile
May not decide
Invent corrected business data or exceed retry authority
Required evidence
Event log, effect keys, attempts, and destination observations

HUMAN

May decide
Resolve exceptions, correct inputs, authorize guarded replay or compensation
May not decide
Rewrite immutable history or erase contrary evidence
Required evidence
Identity, decision, rationale, and referenced evidence

EXCEPTION

May decide
Hold uncertain work with explicit owner and service expectation
May not decide
Masquerade as success or remain ownerless
Required evidence
Trigger, current state, owner, age, and available recovery actions

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Acknowledgement precedes durable commit Accepted ID is absent after response Stop further acknowledgement on unhealthy storage Re-submit only with caller evidence and deduplication Interface owner
Worker crashes after effect commit Attempt lease expires without final transition Block duplicate effect by key Read destination and finalize or route uncertainty Execution service
Corrupt or incomplete checkpoint Integrity or schema validation fails Mark checkpoint unusable; retain earlier valid state Recompute from the previous verified checkpoint Pipeline owner
Retry storm during outage Attempt rate or shared failure class exceeds policy Open circuit and defer work Probe recovery, release gradually, preserve budgets Platform owner
Destination changed independently Reconciliation observes conflicting state Suspend automatic overwrite Apply domain conflict policy with human authority Domain owner
Exception queue loses ownership Missing owner, stale age, or breached policy Prevent exception from being treated as final success Assign, escalate, or return to recoverable failure Operations lead

Invariants and guarantees

Properties the structure is designed to preserve.

  • A successful acceptance response corresponds to a durable, addressable intent record.
  • Material input and provenance remain reconstructable for the required retention period.
  • State changes are legal, atomic, attributable, and append evidence rather than rewrite history.
  • Replay never knowingly repeats an effect already confirmed under the same effect identity.
  • A timeout is represented as uncertainty, not silently converted to success or failure.
  • Every non-completed outcome has an explicit owner and a permitted recovery route.
  • Completion is supported by destination evidence when the destination is authoritative.

What changes between implementations

The constraints determine the final mechanism.

Checkpoint granularity follows the cost and determinism of recomputation. Retry limits follow deadlines, dependency behavior, and harm from repeated calls. Effect identity may be per request, item, document page, or business operation. Some destinations offer native idempotency keys; others require a local ledger and reconciliation query. High-volume systems may partition state and use leases, while low-volume systems can favor simpler transactional records.

Retention, encryption, tenancy isolation, and operator access follow the sensitivity of accepted inputs. Human-review capacity determines escalation thresholds and exception service levels. None of these choices removes the need for a durable acceptance boundary, explicit transitions, and truthful final states.

Evidence chain

Follow the pattern into systems and software.

Architecture and controls specified; no production results published.

The lead-operations design study is a relevant reference-design relationship because accepted lead events may cross validation, enrichment, identity, and CRM-delivery boundaries. The invoice-intake design study is related because document and page processing can preserve verified progress while downstream extraction or review pauses. These links show where the pattern can compose into a larger design; they do not claim that this exact reference design was implemented or measured in either case.

Entity resolution can consume replayed records only when original identity evidence remains stable. Risk-aware model cascades can run as checkpointed stages, but route selection must not change replay semantics or bypass an authority gate.

Status is REFERENCE DESIGN and maturity is specified. The architecture, states, controls, boundaries, and recovery responsibilities are described here. No production results are published.

No open-source implementation is linked. No evaluation fixture, test, or benchmark is linked. The related design studies are design relationships, not evidence that this pattern has been implemented, tested, or measured as a standalone artifact.

Known boundaries

Limitations and non-fit

This pattern cannot provide atomicity across an external destination that exposes neither idempotency nor observable state. It can make uncertainty explicit, but a human or compensating process may still be required. It also does not guarantee deterministic model output; checkpoints must capture outputs and versions when exact replay matters.

For synchronous, single-transaction work entirely inside one reliable database, a full durable executor may add needless machinery. For fire-and-forget telemetry where loss is explicitly acceptable, the promise is intentionally weaker. The pattern is a poor fit when nobody can own exceptions or when source retention rules prevent enough evidence from being preserved; those constraints must be resolved before claiming recoverability.

Related patterns

Continue through the adjacent decision structures.

The central design move is to make acceptance a durable obligation rather than a successful moment at an interface. Once intent, attempts, effects, and ownership are separate records, interruption becomes a state the system can reason about. Replay then means continuing from verified evidence—not hoping that running everything again produces the right world.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real