Start with the process, not the model

An operator does not wake up wanting an agent. They want qualified leads to reach the right owner, invoices to enter a review queue, or support agents to stop searching four systems before every reply. “Agent” describes one possible implementation pattern; it does not describe the job.

I map the trigger, inputs, systems, decisions, actions, exceptions and completion condition before deciding whether AI belongs anywhere. That map often reveals that most steps already have crisp rules: validate an email, look up a customer, calculate a refund window, check a purchase order, write a record and notify an owner.

A useful 80/20 shape

Consider inbound lead operations. A language model may be useful for reading a company website and returning a structured business category. It should not usually decide whether the lead can enter the CRM, which territory owns it or whether consent rules permit outreach. Those are explicit policies.

The resulting system can be 80% deterministic orchestration and 20% probabilistic interpretation. That is not a compromise. It is a design that puts flexibility exactly where fixed rules struggle and predictability everywhere else.

  • Deterministic: validation, deduplication, arithmetic, thresholds, permissions, routing and database writes.

  • AI-assisted: extraction, classification, synthesis, drafting and fuzzy matching.

  • Human-owned: exceptions, relationship judgment and decisions with material consequences.

Why over-agentic designs fail operationally

A general agent can choose tools and paths dynamically. That can be valuable when the task is genuinely open-ended. It also expands the state space you must test. The same input may produce a different tool sequence, a different number of model calls and a different intermediate assumption.

In routine operations that variability appears as difficult debugging, inconsistent latency, unpredictable cost and unclear accountability. A workflow graph makes expected states and error paths visible. A bounded AI node inside the graph can still use powerful language understanding without owning the whole process.

A decision test for every AI step

For each proposed model call, I ask four questions. Can rules solve this accurately? What happens if the answer is wrong? Can the output be validated? What evidence would show that the model is better than the current human step? If rules work, use rules. If failure is consequential and validation is weak, keep a person in the path.

The model should receive the minimum context needed, return a small explicit schema and be allowed to abstain. Downstream code should reject unknown enum values, missing required fields and impossible combinations before they touch a business system.

Autonomy is a rollout decision

Support automation provides a useful ladder: classification, context assembly, response drafting, human-approved sending, then automatic sending for narrow eligible cases. Each level can be measured before moving to the next. The organisation does not have to bet the whole support queue on a model demo.

The same principle applies to finance, RevOps and onboarding. Start in shadow mode, compare against human outcomes, expose failure reasons, then grant narrowly scoped write permission only when the evidence and consequence justify it.

The boring final system is the point

When this is done well, operators stop thinking about AI. A valid event enters, known systems are queried, ambiguity is interpreted, rules are enforced, exceptions reach an owner and every run can be explained. The system looks boring because its uncertainty has boundaries.

That is the outcome I optimise for: fewer manual steps, faster cycle time, cleaner records and a process that fails visibly—not an impressive autonomy diagram.

The practical next step

Map one real execution and one failure.

That will reveal more about the right architecture than a tool comparison or model demo.

Let's build something real