Approval is an architecture boundary
Refunds, external emails, price changes, contracts, vendor selection and account deletion carry consequences that justify named authority. The workflow can assemble facts, run policy checks and draft the action. A person decides whether that proposed action should cross the boundary.
This is not a fallback screen attached after the “real” automation. It is a control plane with identity, policy, state and audit requirements.
Prepare the decision
A reviewer should see the request, source evidence, account context, applicable rule, failed checks, proposed action and alternatives. If they must open five systems again, the automation saved little.
Make uncertainty visible at the field or rule level. “Confidence: 82%” is less actionable than “purchase order is missing; total arithmetic passed; vendor is known.”
Authenticate the decision
Resolve approvers from explicit policy and authoritative groups. Bind a decision to an immutable action request. Use a short-lived signed token or authenticated callback, prevent replay and capture the approver identity and reason.
For higher-risk actions, require two approvers or separation of duties. A chat message is a useful interface; it is not itself an authorisation model.
Model the waiting state
Approval can expire, be rejected, be cancelled by the requester or outlive the facts that produced it. Store a state machine: pending, approved, rejected, expired, cancelled and completed. Revalidate time-sensitive conditions before execution.
Escalation should be deliberate. An expired refund request may return to the queue; it should not become approved because nobody responded.
Resume safely
The originating workflow must verify the continuation token, confirm the action request and ensure it has not already completed. The downstream write should have its own idempotency key. Record the final provider result against the decision.
If the callback fails after approval, the system should retry without asking the person to approve again—and without performing the action twice.
Measure the human loop
Track approval time, rejection rate, change rate, reason distribution, expired requests and time saved per review. A high edit or reject rate points to weak preparation, unclear policy or excessive automation coverage.
Approved corrections are valuable evaluation data. They should improve schemas, retrieval, rules and model prompts without turning individual reviewers into invisible labellers.
Earn narrower review
Over time, evidence may justify automatic execution for a low-risk, tightly defined class. Keep the approval service for everything outside that class. The boundary should narrow through measured performance, not enthusiasm.
Good human-in-the-loop design creates faster decisions, clearer responsibility and safer automation. That is a product capability.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real