Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 11 / Operations, Privacy & Security

Private means the fallback stays private too.

Private Inference with No-Egress Fallback

Keep the complete inference path—not only the model endpoint—inside an approved boundary, and fail safely when local capacity or components are unavailable.

Maturity
Reference design / Specified
Critical uncertainty
Private inference includes preprocessing, embeddings, retrieval, logs, traces, evaluation artifacts and fallback behaviour—not only the language-model server.
Human boundary
Approve scoped interpretation, correct output, or authorize permitted specialist handling
Applications
Sales-call analysis · Contract review · Source-code assistance · Capacity-constrained local inference
Sensitive sales calls, contracts, source code, and capacity requests remain inside a bounded nine-layer inference path and end as private result, capacity deferred, or manual review.

Problem shape

The problem beneath them

Authorize the workload before intake, keep every data-bearing stage inside an enforced boundary, and defer or route to an approved human process instead of failing over to unapproved infrastructure.

Inference systems create many representations: uploaded bytes, decoded text, temporary files, features, embeddings, retrieved passages, prompts, model caches, outputs, logs, traces, reviewer packets, and evaluation examples. Any one can leave through network calls, shared observability, package downloads, support tooling, backups, or permissive failover.

“On-prem model” therefore describes deployment, not a complete control. The system needs an explicit approved boundary, data-flow inventory, identity and purpose checks, egress enforcement independent of application intent, and failure routes that preserve the boundary under pressure. The no-egress requirement should be tested at network and dependency layers, not merely documented in code.

Capacity deserves special treatment. A full queue is an availability condition, not authorization to send data elsewhere. Safe outcomes can include deferral, reduced approved processing, or manual/specialist review. The selected route must disclose its state and owner rather than quietly producing a less-private result.

Private inference includes preprocessing, embeddings, retrieval, logs, traces, evaluation artifacts and fallback behaviour—not only the language-model server.

This pattern controls data location and route eligibility. It does not by itself establish lawful purpose, eliminate insider risk, prevent every model memorization concern, or certify a deployment against a regulatory framework. Those claims require governance and evidence beyond architecture.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Sales calls

Can this recording leave the network?

Raw audio and transcript contain customer and employee information.

02 / Contracts

Can document embeddings use an external service?

Derived representations may disclose protected document content.

03 / Source code

Can prompts or traces expose proprietary code?

Operational observability can copy sensitive context outside the inference host.

04 / Capacity

If the local model is full, may the request leave?

A generic failover policy points toward an unapproved provider.

All four questions converge into Private Inference with No-Egress Fallback.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Eligibility, consent and jurisdiction

    Question
    May this data be processed for this purpose in this boundary?
    Responsibility
    Verify actor, consent or lawful basis, purpose, jurisdiction, retention class, and permitted operations before intake.
    Input
    Request identity, subject/data classification, policy, consent record, and intended use.
    Output
    Authorized processing envelope or explicit rejection/manual route.
    Stops when
    Scope, boundary, and allowed outputs are established without inference.
  2. 02

    Encrypted intake

    Question
    Can protected input enter without exposure or substitution?
    Responsibility
    Authenticate sender, encrypt transport, validate integrity and type, assign stable identity, and avoid unsafe temporary handling.
    Input
    Authorized envelope and source artifact.
    Output
    Encrypted, integrity-checked intake record inside the boundary.
    Stops when
    Durable custody is confirmed or the transfer is rejected cleanly.
  3. 03

    Private source storage

    Question
    Where do originals, derivatives, and keys live, and for how long?
    Responsibility
    Apply encryption, tenancy separation, least privilege, retention, deletion, backup, and access audit controls.
    Input
    Intake record, data class, identity, retention policy, and key context.
    Output
    Controlled source object and access-scoped references.
    Stops when
    Storage and lifecycle controls match the authorized envelope.
  4. 04

    Local preprocessing

    Question
    Do decoding, transcription, parsing, redaction, and feature extraction stay private?
    Responsibility
    Run approved local components with constrained dependencies, temporary storage, and telemetry.
    Input
    Controlled source references and processing contract.
    Output
    Versioned local derivatives with provenance.
    Stops when
    Required derivatives are produced or an internal failure route is entered.
  5. 05

    Local embeddings and retrieval

    Question
    Can indexing and context selection occur without external data or metadata leakage?
    Responsibility
    Compute embeddings locally, isolate indexes, enforce retrieval permissions, and retain source/version links.
    Input
    Local derivatives, authorized corpus, identity, and query.
    Output
    Permission-filtered local context and trace.
    Stops when
    Context is bounded to permitted sources or evidence is insufficient.
  6. 06

    Local model inference

    Question
    Can an approved model produce the bounded output within local resource limits?
    Responsibility
    Run pinned models and serving components with controlled prompts, tools, caches, and output validation.
    Input
    Authorized task, local context, model/version, and resource allocation.
    Output
    Candidate private result, uncertainty, and internal execution evidence.
    Stops when
    Output passes its contract, escalates, or reaches a capacity/technical bound.
  7. 07

    Egress enforcement

    Question
    What prevents any component from sending protected material outside?
    Responsibility
    Deny unapproved network paths, proxies, DNS, telemetry, remote tools, and fallback endpoints at independent control layers.
    Input
    Workload identities, allowlists, network policy, dependency inventory, and runtime events.
    Output
    Enforced communication boundary and auditable deny/allow events.
    Stops when
    Every required flow is explicitly authorized and unexpected egress is blocked.
  8. 08

    Capacity-safe failure route

    Question
    What happens when private infrastructure cannot process the request?
    Responsibility
    Apply queue bounds, deadlines, prioritization, approved degradation, and no-egress fallback policy.
    Input
    Capacity state, workload priority, deadline, privacy class, and eligible local routes.
    Output
    Capacity deferred or manual/specialist review state.
    Stops when
    The request has a truthful owner and no unapproved route is attempted.
  9. 09

    Evidence packet and human decision

    Question
    What may an authorized reviewer see and decide without widening exposure?
    Responsibility
    Assemble minimal local evidence, enforce reviewer authorization, and record scoped disposition.
    Input
    Candidate result or failure, source locators, policy, and reviewer role.
    Output
    Private result, correction, deferral, or specialist disposition with audit evidence.
    Stops when
    Decision, authority, evidence, and permitted downstream release are recorded.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Purpose, consent, or jurisdiction fails Reject or use approved governance route Infrastructure location cannot cure unlawful processing Applicability requires privacy/legal authority
Workload is authorized and local capacity exists Run complete private path All stages can remain within approved controls Evidence or output contract fails
Retrieval source is outside user permission Exclude source Local storage does not imply universal access Required evidence cannot be lawfully retrieved
Component attempts unapproved egress Block, alert, and contain Boundary enforcement must be independent of component intent Data exposure may have occurred
Local queue reaches policy bound Defer or prioritize locally Capacity does not authorize public failover Deadline requires authorized manual handling
Approved specialist environment exists Transfer only under explicit policy Another boundary may be eligible if separately authorized Data, purpose, or recipient exceeds authorization
No private route can satisfy task Return no result or human route Privacy requirement outranks convenience Consequence requires accountable disposition

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Authorize, ingest encrypted, process and retrieve locally, enforce egress, validate, and release only the permitted output.

Final state
Private result.
Owner
Private inference service until authorized release; consuming role owns subsequent action.
Evidence
Authorization envelope, artifact IDs, component/model versions, access and retrieval trace, egress events, and output disposition.
Recovery
Reprocess from retained approved inputs when a component or policy version changes.
Ambiguous path

Stop automated release when evidence, authorization, or interpretation is unresolved and present minimal local evidence to an authorized reviewer.

Final state
Manual or specialist review.
Owner
Named privacy, domain, or specialist reviewer.
Evidence
Eligibility result, evidence gaps, local processing trace, candidate result, and reviewer decision.
Recovery
Narrow purpose, obtain authorization or evidence, approve a scoped result, or reject processing.
Failure path

Block unapproved egress, place recoverable requests in bounded local storage, and expose capacity or component failure explicitly.

Final state
Capacity deferred.
Owner
Platform operations for restoration and service owner for deadline/queue decisions.
Evidence
Capacity state, failed component, request age and priority, denied fallback, and containment actions.
Recovery
Restore validated local capacity, resume from a safe checkpoint, or move to an independently approved human route.

Authority map

Capability does not grant authority.

RULE

May decide
Enforce eligibility, retention, route allowlists, egress, and mandatory review
May not decide
Invent consent or expand purpose because processing is local
Required evidence
Policy/version, identity, data class, conditions, and result

MODEL

May decide
Analyze permitted local context and report bounded uncertainty
May not decide
Choose an external fallback, change retention, or authorize disclosure
Required evidence
Model/version, local inputs, output, uncertainty, and tool trace

SYSTEM

May decide
Store, process, queue, deny egress, and release approved outputs
May not decide
Bypass network controls or silently downgrade privacy under load
Required evidence
State history, access, component versions, network events, and release record

HUMAN

May decide
Approve scoped interpretation, correct output, or authorize permitted specialist handling
May not decide
Export data outside delegated authority or erase audit evidence
Required evidence
Identity, role, evidence reviewed, purpose, decision, and destination

EXCEPTION

May decide
Hold capacity, authorization, evidence, and integrity failures inside the boundary
May not decide
Become public fallback or indefinite ownerless storage
Required evidence
Trigger, protected artifacts, owner, retention deadline, and recovery options

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Preprocessor calls cloud API Denied network event or dependency audit Terminate workload and isolate component Replace with approved local component; review attempted payload Security owner
Shared tracing exports prompts Data-flow test or telemetry inspection Disable exporter and rotate affected credentials Deploy local redacted tracing and assess exposure Observability owner
Cross-tenant vector retrieval Permission test or result audit finds mismatch Block index and affected outputs Repair partition/filter controls, rebuild, and review access Data platform
Local capacity triggers generic failover Route trace shows external endpoint selection Deny at network and mark requests deferred Correct fallback policy and replay privately Inference platform
Temporary files outlive retention Storage scan finds expired derivatives Quarantine worker and prevent new intake Securely remove, repair lifecycle, and verify Platform owner
Reviewer exports evidence improperly Audit event or control alert Revoke access and contain destination Incident response, disposition review, and control correction Privacy/security owner

Invariants and guarantees

Properties the structure is designed to preserve.

  • Eligibility, purpose, and boundary are established before protected content is processed.
  • Originals, derivatives, embeddings, prompts, outputs, logs, and evaluation artifacts follow explicit lifecycle controls.
  • Retrieval authorization is enforced at query time and preserves source provenance.
  • Egress is denied by infrastructure controls, not application convention alone.
  • Capacity or component failure never selects an unapproved external route.
  • Every deferred or manual state has an owner and retention deadline.
  • Human review occurs inside an approved boundary and does not silently widen permitted use.

What changes between implementations

The constraints determine the final mechanism.

The approved boundary may be a disconnected host, private datacenter, controlled cloud tenancy, confidential-computing environment, or another architecture accepted by the governing policy. Threat model, jurisdiction, data classes, and audit obligations determine isolation strength. Workload size and latency determine model serving, batching, queueing, and prioritization.

Some systems can redact before deeper processing; others must treat redaction itself as sensitive local inference. Retrieval may use per-tenant indexes or policy filters over shared infrastructure where authorized. Observability can range from aggregate health signals to locally retained detailed traces. Retention and deletion need to cover backups, caches, evaluation corpora, and reviewer exports—not just primary storage.

Evidence chain

Follow the pattern into systems and software.

Architecture and controls specified; no production results published.

on-prem-sales-call-compliance is a relevant reference-design relationship because audio preprocessing, transcription, retrieval, analysis, evidence packets, and review can all fall inside the privacy boundary. This relationship does not claim that the complete pattern described here was implemented, tested, certified, or measured in that case.

Evidence-bounded inference ensures local conclusions remain grounded. Risk-aware cascades can choose only among privacy-eligible routes. Durable execution lets capacity-deferred work resume without moving outside the boundary.

This is a REFERENCE DESIGN with specified architecture and controls; no production results are published. No open-source implementation is linked. No evaluation fixture, test, or benchmark is linked. No certification, penetration-test result, privacy assessment, or measured no-egress guarantee is claimed.

Known boundaries

Limitations and non-fit

No-egress controls reduce one class of exposure but do not eliminate compromised administrators, malicious dependencies, hardware attacks, inference leakage to authorized users, or improper downstream use. Encryption at rest does not protect data during authorized processing. Local operation can also delay security updates or reduce model capability; those trade-offs need explicit ownership.

The pattern is not justified merely as branding where data is already public and policy permits managed services. It is non-fit when the organization cannot patch, monitor, capacity-plan, and incident-manage the private stack. If a hard deadline cannot tolerate deferral and no approved alternate boundary exists, the service contract must reject that workload rather than promise unavailable private processing.

Related patterns

Continue through the adjacent decision structures.

Private inference is a data-flow property that must remain true during normal operation, overload, component failure, review, and evaluation. The decisive control is often the refusal to fall back: when capacity disappears, the request waits or reaches an approved owner, but the privacy boundary does not move.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real