Problem shape
The problem beneath them
Authorize the workload before intake, keep every data-bearing stage inside an enforced boundary, and defer or route to an approved human process instead of failing over to unapproved infrastructure.
Inference systems create many representations: uploaded bytes, decoded text, temporary files, features, embeddings, retrieved passages, prompts, model caches, outputs, logs, traces, reviewer packets, and evaluation examples. Any one can leave through network calls, shared observability, package downloads, support tooling, backups, or permissive failover.
“On-prem model” therefore describes deployment, not a complete control. The system needs an explicit approved boundary, data-flow inventory, identity and purpose checks, egress enforcement independent of application intent, and failure routes that preserve the boundary under pressure. The no-egress requirement should be tested at network and dependency layers, not merely documented in code.
Capacity deserves special treatment. A full queue is an availability condition, not authorization to send data elsewhere. Safe outcomes can include deferral, reduced approved processing, or manual/specialist review. The selected route must disclose its state and owner rather than quietly producing a less-private result.
Private inference includes preprocessing, embeddings, retrieval, logs, traces, evaluation artifacts and fallback behaviour—not only the language-model server.
This pattern controls data location and route eligibility. It does not by itself establish lawful purpose, eliminate insider risk, prevent every model memorization concern, or certify a deployment against a regulatory framework. Those claims require governance and evidence beyond architecture.
Different problems, same shape
The surface changes. The decision structure persists.
Can this recording leave the network?
Raw audio and transcript contain customer and employee information.
Can document embeddings use an external service?
Derived representations may disclose protected document content.
Can prompts or traces expose proprietary code?
Operational observability can copy sensitive context outside the inference host.
If the local model is full, may the request leave?
A generic failover policy points toward an unapproved provider.
All four questions converge into Private Inference with No-Egress Fallback.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Eligibility, consent and jurisdiction
- Question
- May this data be processed for this purpose in this boundary?
- Responsibility
- Verify actor, consent or lawful basis, purpose, jurisdiction, retention class, and permitted operations before intake.
- Input
- Request identity, subject/data classification, policy, consent record, and intended use.
- Output
- Authorized processing envelope or explicit rejection/manual route.
- Stops when
- Scope, boundary, and allowed outputs are established without inference.
-
02
Encrypted intake
- Question
- Can protected input enter without exposure or substitution?
- Responsibility
- Authenticate sender, encrypt transport, validate integrity and type, assign stable identity, and avoid unsafe temporary handling.
- Input
- Authorized envelope and source artifact.
- Output
- Encrypted, integrity-checked intake record inside the boundary.
- Stops when
- Durable custody is confirmed or the transfer is rejected cleanly.
-
03
Private source storage
- Question
- Where do originals, derivatives, and keys live, and for how long?
- Responsibility
- Apply encryption, tenancy separation, least privilege, retention, deletion, backup, and access audit controls.
- Input
- Intake record, data class, identity, retention policy, and key context.
- Output
- Controlled source object and access-scoped references.
- Stops when
- Storage and lifecycle controls match the authorized envelope.
-
04
Local preprocessing
- Question
- Do decoding, transcription, parsing, redaction, and feature extraction stay private?
- Responsibility
- Run approved local components with constrained dependencies, temporary storage, and telemetry.
- Input
- Controlled source references and processing contract.
- Output
- Versioned local derivatives with provenance.
- Stops when
- Required derivatives are produced or an internal failure route is entered.
-
05
Local embeddings and retrieval
- Question
- Can indexing and context selection occur without external data or metadata leakage?
- Responsibility
- Compute embeddings locally, isolate indexes, enforce retrieval permissions, and retain source/version links.
- Input
- Local derivatives, authorized corpus, identity, and query.
- Output
- Permission-filtered local context and trace.
- Stops when
- Context is bounded to permitted sources or evidence is insufficient.
-
06
Local model inference
- Question
- Can an approved model produce the bounded output within local resource limits?
- Responsibility
- Run pinned models and serving components with controlled prompts, tools, caches, and output validation.
- Input
- Authorized task, local context, model/version, and resource allocation.
- Output
- Candidate private result, uncertainty, and internal execution evidence.
- Stops when
- Output passes its contract, escalates, or reaches a capacity/technical bound.
-
07
Egress enforcement
- Question
- What prevents any component from sending protected material outside?
- Responsibility
- Deny unapproved network paths, proxies, DNS, telemetry, remote tools, and fallback endpoints at independent control layers.
- Input
- Workload identities, allowlists, network policy, dependency inventory, and runtime events.
- Output
- Enforced communication boundary and auditable deny/allow events.
- Stops when
- Every required flow is explicitly authorized and unexpected egress is blocked.
-
08
Capacity-safe failure route
- Question
- What happens when private infrastructure cannot process the request?
- Responsibility
- Apply queue bounds, deadlines, prioritization, approved degradation, and no-egress fallback policy.
- Input
- Capacity state, workload priority, deadline, privacy class, and eligible local routes.
- Output
- Capacity deferred or manual/specialist review state.
- Stops when
- The request has a truthful owner and no unapproved route is attempted.
-
09
Evidence packet and human decision
- Question
- What may an authorized reviewer see and decide without widening exposure?
- Responsibility
- Assemble minimal local evidence, enforce reviewer authorization, and record scoped disposition.
- Input
- Candidate result or failure, source locators, policy, and reviewer role.
- Output
- Private result, correction, deferral, or specialist disposition with audit evidence.
- Stops when
- Decision, authority, evidence, and permitted downstream release are recorded.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Purpose, consent, or jurisdiction fails | Reject or use approved governance route | Infrastructure location cannot cure unlawful processing | Applicability requires privacy/legal authority |
| Workload is authorized and local capacity exists | Run complete private path | All stages can remain within approved controls | Evidence or output contract fails |
| Retrieval source is outside user permission | Exclude source | Local storage does not imply universal access | Required evidence cannot be lawfully retrieved |
| Component attempts unapproved egress | Block, alert, and contain | Boundary enforcement must be independent of component intent | Data exposure may have occurred |
| Local queue reaches policy bound | Defer or prioritize locally | Capacity does not authorize public failover | Deadline requires authorized manual handling |
| Approved specialist environment exists | Transfer only under explicit policy | Another boundary may be eligible if separately authorized | Data, purpose, or recipient exceeds authorization |
| No private route can satisfy task | Return no result or human route | Privacy requirement outranks convenience | Consequence requires accountable disposition |
Operating paths