The duplicate that looks like success

A form provider sends a webhook. The workflow creates a lead, but the network drops before it records success. The provider retries and a second execution creates another lead. Both runs look reasonable; the CRM is wrong.

Replace “lead” with invoice, onboarding project or refund and the consequence grows. AI is incidental here. The moment an AI-enriched workflow performs a side effect, ordinary distributed-systems problems apply.

Three identities, three jobs

A correlation ID follows one business event across services. An execution ID distinguishes attempts. An idempotency key identifies the business operation that must happen at most once from the organisation’s point of view.

Provider event IDs are often good keys. Otherwise derive a canonical key from stable fields—tenant plus source lead ID, or vendor plus invoice number—after careful normalisation. Do not use model-generated summaries.

Claim before the side effect

Create an idempotency record with a unique constraint and a state such as started, effect_pending, completed or needs_reconciliation. Claim it atomically. If the key already exists, return the completed result or resume according to state.

A read-then-write check without a unique constraint can race: two workers see no record and both proceed. The database must participate in the guarantee.

The unavoidable uncertainty window

Suppose an external refund API accepts the request, then the worker crashes before storing the response. Your database says pending; the provider may have acted. Retrying blindly risks a double refund.

Use the provider’s idempotency key when available. Otherwise query by stable business reference or move the item to reconciliation. Exactly-once delivery across arbitrary systems is not a switch an orchestrator can provide.

Design replay explicitly

A replay tool should show prior attempts, known side effects, current source state and why retry is safe. It should preserve the business key and create a new execution ID. Dangerous states may require human confirmation.

Dead-letter queues without safe replay are archives of regret. Test replay using synthetic duplicate, timeout and crash fixtures before launch.

Idempotency and AI steps

Model output may vary across attempts. Persist the accepted structured result or version the model and prompt. A retry should not silently reclassify the same lead and choose a different deterministic branch unless that behaviour is intended.

Keep side-effect decisions explainable: record the inputs, validated AI fields, rule version and action key that led to the write.

The operational rule

Assume every trigger can arrive twice and every response can be lost. Ask which business action must not repeat, choose its key, store its state atomically and define reconciliation for ambiguous outcomes.

This small piece of backend discipline is often the difference between a convincing workflow demo and a production system.

The practical next step

Map one real execution and one failure.

That will reveal more about the right architecture than a tool comparison or model demo.

Let's build something real