Problem shape
The problem beneath them
Apply deterministic eligibility and authority floors first, use the least costly qualified model for routine work, and escalate novelty, uncertainty, or consequence through controlled routes.
Sending every task to the strongest available model wastes latency and compute, but optimizing only for cost creates a more serious error: the cheap route can become an authority bypass. Model confidence is not calibrated uniformly across task types, distributions, versions, or consequences. A routine-looking input can contain a novel edge case; a stronger model can still be wrong; and neither should override mandatory review.
The pattern separates routing from decision authority. Deterministic controls decide whether automation is eligible at all. A risk floor establishes the minimum route. Task and novelty signals select among eligible models. Output checks decide whether the selected route can stop. Human authority remains a distinct gate for consequential actions.
A cascade must also be evaluated as a policy, not a collection of model scores. Route choice changes which errors are observed, and human outcomes can be delayed or selective. Capturing input strata, route reasons, overrides, and final dispositions is necessary to detect systematic under-escalation.
No model may downgrade a policy-required human review or gain additional authority because it produces a confident answer.
This pattern allocates inference and review effort. It does not prove that a small route is accurate enough, define organizational risk appetite, or make a larger model an oracle. Those require task-specific evidence and accountable policy.
Different problems, same shape
The surface changes. The decision structure persists.
Does a typo need the same model as a database migration?
Change complexity and blast radius vary substantially.
Is this known policy language or a novel evasion?
Routine violations resemble history while adversarial behavior shifts.
Can the routine route answer this, or does the evidence conflict?
A familiar request contains incompatible sources or unusual consequence.
Is this ordinary extraction or a high-risk exception?
The same document task can feed either clerical review or a consequential action.
All four questions converge into Risk- and Novelty-Aware Model Cascades.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Deterministic eligibility
- Question
- Is model automation permitted for this task and data at all?
- Responsibility
- Enforce consent, data class, jurisdiction, feature availability, task allowlists, and required prerequisites.
- Input
- Request, actor, data classification, policy, and system health.
- Output
- Eligible route set or immediate non-model disposition.
- Stops when
- Prohibited work is rejected, deferred, or sent to an authorized human path.
-
02
Non-negotiable risk floor
- Question
- What minimum scrutiny and authority does policy require?
- Responsibility
- Encode actions and consequences that mandate stronger analysis, specialist review, or human approval.
- Input
- Proposed action, affected subject, permissions, reversibility, and regulatory or internal policy.
- Output
- Minimum route and mandatory gates.
- Stops when
- No later component can select a route below the floor.
-
03
Task and consequence classification
- Question
- What kind of work is this, and what happens if it is wrong?
- Responsibility
- Classify task family, complexity signals, blast radius, and error asymmetry using bounded features.
- Input
- Eligible request, metadata, artifacts, and intended downstream use.
- Output
- Task class, consequence class, and routing features.
- Stops when
- Unknown or conflicting classification is itself routed safely.
-
04
Novelty detection
- Question
- Is this input represented by cases the routine route is known to handle?
- Responsibility
- Compare to evaluated distributions, detect unfamiliar patterns and conflicts, and avoid equating distance with certainty.
- Input
- Task features, reference data, policy signatures, and historical route strata.
- Output
- Known, novel, or indeterminate novelty disposition with reasons.
- Stops when
- Novelty is bounded enough to select or escalate a route.
-
05
Small-model route
- Question
- Can the least costly eligible model produce a checkable answer?
- Responsibility
- Handle well-scoped routine tasks under strict output, tool, evidence, and action limits.
- Input
- Bounded task, permitted context, route contract, and required output schema.
- Output
- Candidate result, confidence signals, evidence, and stop/escalate recommendation.
- Stops when
- Validation passes within authority or any escalation trigger fires.
-
06
Stronger-model or specialist route
- Question
- Does additional capability or specialization resolve the identified difficulty?
- Responsibility
- Reassess novel, complex, or conflicting inputs without relaxing policy controls.
- Input
- Original task, lower-route trace, escalation reason, and expanded permitted context.
- Output
- Deeper candidate result or human-review requirement.
- Stops when
- Checks pass within route authority or uncertainty remains material.
-
07
Confidence calibration
- Question
- Do route signals predict acceptable performance for this task stratum?
- Responsibility
- Apply versioned calibration and output validation; treat unsupported confidence as untrusted.
- Input
- Candidate result, model/version, task stratum, evidence, and evaluation contract.
- Output
- Acceptable-to-stop, escalate, or invalid result.
- Stops when
- The route satisfies the stopping contract or cannot do so.
-
08
Human authority gate
- Question
- Which conclusion or action requires accountable human judgment?
- Responsibility
- Present evidence, route history, uncertainty, and permitted decisions to an authorized reviewer.
- Input
- Candidate output, policy floor, consequence, evidence, and unresolved issues.
- Output
- Approved, corrected, rejected, deferred, or escalated disposition.
- Stops when
- An authorized decision is recorded or the case remains explicitly owned.
-
09
Outcome capture and route evaluation
- Question
- Is the cascade sending the right cases to the right routes over time?
- Responsibility
- Record route reasons, versions, outputs, overrides, delayed outcomes, and evaluable cohorts.
- Input
- Full route trace, human disposition, downstream feedback, and policy version.
- Output
- Auditable decision record and data for controlled route evaluation.
- Stops when
- Outcome status and limitations are recorded without inventing ground truth.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Policy prohibits model processing | Use approved non-model or human path | Cost cannot override eligibility | Policy applicability is unclear |
| Mandatory human gate applies | Preserve gate for every model route | Authority is fixed before inference | Reviewer specialization is required |
| Known low-consequence task with valid inputs | Try small-model route | Routine work may not need deeper capacity | Output or evidence validation fails |
| Novel or out-of-distribution input | Use stronger or specialist route | Routine-route evidence does not cover novelty | Stronger route also lacks support |
| Conflicting evidence | Escalate depth or human review | Confidence cannot settle source conflict | Conflict changes a consequential action |
| High confidence but invalid schema/citation | Reject route output | Confidence cannot waive contract failure | A corrected rerun still fails |
| Route service unavailable | Defer or use approved equal/higher route | Availability must not lower controls | Capacity deadline creates human-operational risk |
Operating paths