Pick the decision boundary
For a lead workflow, the AI step may return industry, company type, intent and a short context summary. The workflow then applies explicit ICP and territory rules. Evaluate those model-produced fields, not the final owner as though the model chose it.
This separation makes regressions diagnosable. A wrong owner can come from classification, a routing table, stale CRM data or a code bug. One end-to-end success score hides the cause.
Build a representative evaluation set
Use permissioned historical examples where appropriate, supplemented with synthetic edge cases. Sample across sources, languages, company types, document quality and rare categories. Include ambiguous inputs where abstention is correct.
Keep a locked test set away from prompt iteration. Label guidelines matter: if two competent reviewers disagree, the category may be underspecified rather than the model weak.
Measure the contract
Track schema success, per-field accuracy, critical-class precision/recall, abstention, unsupported claims, latency and cost. For summaries, use a rubric tied to source fidelity and required facts instead of a generic similarity score.
Slice results. A 94% average can hide a 60% recall for the category that controls an expensive route. Set release thresholds by consequence.
Test the surrounding workflow
Feed invalid JSON, missing fields, unknown enum values, low confidence, timeout and provider errors. Verify that the workflow repairs only format, falls back or enters review as designed. Ensure no forbidden side effect occurs.
Then run a smaller end-to-end set to confirm model fields and deterministic rules compose correctly. This is where versioned routing tables and fixtures earn their keep.
Compare changes as releases
Record model identifier, provider settings, prompt version, schema version and retrieval snapshot. Compare candidates on the same data. A prompt that improves accuracy but doubles latency may be unacceptable for synchronous support.
n8n’s current evaluation features provide ecosystem support for test datasets and comparing AI workflow outputs. Whether using those features or an external harness, keep evaluation in the release path—not in a notebook nobody reruns.
Research note: External ecosystem reference
Monitor production correction signals
Shadow results, human overrides, draft edits, exception reasons and downstream reconciliation reveal drift. Do not treat every human change as model error; distinguish preference, new policy and source-data problems.
Sample production outcomes for review and alert on shifts in category mix, abstention, schema failure, latency and cost. Evaluation is a loop, not a launch gate.
Keep the model’s job small enough to test
The more open-ended the assignment, the harder it is to label, measure and control. A narrow structured task can still remove substantial manual work when the rest of the process is automated.
A company does not need an agent evaluation programme to use AI responsibly. It needs a clear contract, representative data, consequence-aware metrics and tested failure paths.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real