Problem shape
The problem beneath them
Persist accepted work before acknowledging it, advance it through explicit replay-safe states, reconcile external effects, and leave every attempt in a state with a named owner.
Distributed work crosses failure boundaries. A process can crash after committing locally but before notifying a queue. A worker can time out after a destination committed an effect. A retry can run against changed source data. A batch can contain completed, rejected, and unknown items at the same time. “Try again” is therefore not a complete recovery policy.
The pattern separates intent, attempt, and effect. Intent records what the system accepted. Attempts record executions and their evidence. Effects identify externally observable changes. State transitions connect them without pretending they form one atomic transaction. This separation makes uncertain outcomes representable: an attempt may be unknown even though the accepted intent remains recoverable.
Durability also requires ownership. Exhausting retries is not a final state unless a person or service owns the resulting exception and has enough evidence to act. Likewise, “sent” is not “completed” when the destination is authoritative. Completion should follow a destination receipt, read-after-write check, or another explicit reconciliation rule.
Accepting work does not prove that an external destination completed it. Completion requires reconciliation or an explicitly owned exception.
This pattern governs the lifecycle of accepted work. It does not make a non-idempotent destination transactional, guarantee that a third party retains data, or decide the business validity of an input. Those concerns need destination-specific controls and policy.
Different problems, same shape
The surface changes. The decision structure persists.
The user saw success. Did the work survive?
An interface confirms receipt before downstream work finishes.
The destination was down. Where is the accepted record?
A remote write fails after the source has acknowledged the request.
Page 18 failed. Must pages 1–17 run again?
A multi-stage job retains no trustworthy progress boundary.
The job stopped halfway. What is safe to replay?
Partial effects exist, but their relationship to the input is unclear.
All four questions converge into Durable, Replayable Execution.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Acceptance boundary
- Question
- At what point may the interface promise that work has been accepted?
- Responsibility
- Validate the minimum envelope, assign a stable work identifier, and durably commit intent before acknowledging receipt.
- Input
- Request payload, caller identity, tenancy, policy context, and any client idempotency key.
- Output
- Durable acceptance record and acknowledgement correlated to its identifier.
- Stops when
- The record is committed or the caller receives an explicit rejection with no implied acceptance.
-
02
Immutable input
- Question
- Can a replay reconstruct the exact material input originally accepted?
- Responsibility
- Preserve the canonical payload, source references, version identifiers, and integrity metadata without later mutation.
- Input
- Accepted envelope and retrievable source artifacts.
- Output
- Versioned input snapshot or content-addressed references with provenance.
- Stops when
- Material inputs are frozen, integrity-checkable, and subject to an explicit retention policy.
-
03
Explicit state machine
- Question
- Which transitions are legal, and who may perform them?
- Responsibility
- Represent pending, active, waiting, reconciling, failed, and final states with guarded transitions.
- Input
- Acceptance record, current state, event, actor, and transition preconditions.
- Output
- Atomic state transition plus timestamped transition evidence.
- Stops when
- The event is applied once, rejected as illegal, or retained for investigation.
-
04
Durable checkpoint
- Question
- What verified progress can a later attempt reuse safely?
- Responsibility
- Commit stage outputs and their input/version fingerprints at meaningful restart boundaries.
- Input
- Stage result, input identity, code or policy version, and validation evidence.
- Output
- Checkpoint that is reusable only under matching preconditions.
- Stops when
- The checkpoint and state transition are durably associated.
-
05
Retry policy
- Question
- Is another attempt safe and likely to improve the outcome?
- Responsibility
- Classify failures, schedule bounded retries with backoff, and prevent permanent errors from cycling.
- Input
- Failure class, attempt history, destination health, deadline, and retry budget.
- Output
- Retry schedule, deferred state, or exception route.
- Stops when
- Work succeeds, reaches its retry/deadline bound, or a non-retryable condition is established.
-
06
Idempotency protection
- Question
- Could replay duplicate an internal or external effect?
- Responsibility
- Give each effect stable identity and suppress, deduplicate, or reconcile repeated attempts.
- Input
- Work identifier, effect key, destination scope, and prior effect evidence.
- Output
- One authorized effect, an existing result, or an ambiguous-effect state.
- Stops when
- Uniqueness is enforced or uncertainty is routed without repeating the effect blindly.
-
07
Destination reconciliation
- Question
- What does the authoritative destination say actually happened?
- Responsibility
- Compare intended and observed state using receipts, stable destination identifiers, or safe read-back.
- Input
- Intended effect, attempt evidence, destination response, and observable destination state.
- Output
- Confirmed completion, confirmed absence, conflict, or unresolved status.
- Stops when
- Destination state is confirmed or an owned exception records why confirmation is unavailable.
-
08
Explicit final state
- Question
- How does the lifecycle end without abandoned work?
- Responsibility
- Record a controlled outcome, owner, evidence, retention, and permitted recovery action.
- Input
- State history, reconciliation result, policy, and unresolved conditions.
- Output
- Completed, owned exception, or recoverable failure.
- Stops when
- No work remains in an ownerless or semantically ambiguous terminal state.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| No durable acceptance record | Reject or ask caller to retry | The system cannot honestly promise recovery | The caller reports success but no record can be located |
| Matching completed effect key | Return the prior result | Repeating the effect adds risk without new intent | Destination state conflicts with stored evidence |
| Transient destination outage | Defer and retry with bounds | Availability may recover without changing intent | Deadline or retry budget is exhausted |
| Permanent validation rejection | Stop automatic retry | Identical input will reproduce the rejection | Correction requires business authority |
| Timeout after dispatch | Reconcile before retry | The destination may have committed despite no response | The effect cannot be queried safely |
| Checkpoint fingerprint mismatch | Recompute from the last valid boundary | Prior output belongs to different material inputs | Recalculation could invalidate downstream effects |
| Unknown final destination state | Create owned exception | Neither success nor absence is proven | Consequence or age exceeds policy tolerance |
Operating paths