Technical design
Call review inside a private network
The recording stayed home. One search query didn't.
Company Aurum is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed. The story explains the problem and follows one item through the system; this page is the engineering detail behind it.
Scope and assumptions
The system gives every recorded call a first-pass compliance check and a review queue of timestamped findings. It runs entirely inside the firm's private network. It does not score advisers, trigger disciplinary action, or contact customers.
Illustrative figures assume: 6,000 calls × 18 minutes = 1,800 hours of audio a week; transcription at a real-time factor (processing time ÷ audio time) of 0.1 on the chosen GPU needs about 180 GPU-hours a week; 10% of calls produce at least one finding needing review, about 1.5 segments each (≈900 segments); 1% need a full listen. Real-time factor must be measured on the actual hardware and model; the 0.1 is a planning placeholder.
Architecture
telephony recordings (on-prem storage)
│ pull, read-only
▼
job queue (Postgres) ── admission control: queue age vs deadline
│
▼
GPU node(s), no internet egress
├─ speech recognition: transcript, word timestamps, word confidence
├─ diarization / channel mapping
├─ local LLM: checklist evaluation → findings with timestamps
└─ local embeddings: search in the internal policy library
│
▼
findings store ──► review UI (plays audio at timestamp) ──► reviewer decision
│
audit log: model versions, inputs hashes, decisions
- Network. The processing subnet has default-deny egress. Allowed connections: the recording store (read), the database, the internal artifact registry (for approved model imports), monitoring. Denied attempts are logged with source process and destination.
- Speech recognition. A Whisper-family model run locally (for example with faster-whisper), producing word timestamps and per-word probabilities. The stereo telephony channels are used for speaker separation when available.
- Checklist evaluation. A local instruction-tuned model (a 7B to 14B open-weight model, quantized to fit the GPU memory budget) reads the transcript in windows and returns, per checklist item, a finding with start and end timestamps and quoted words. Quoted words must match the transcript at those timestamps, which code verifies.
- Policy search. Local embeddings over the internal policy library, so the model can cite which rule a finding relates to. This is the step the leaky design sent outside; here it runs next to the model.
- Review UI. An internal web page with the audio player, transcript around the finding, both readings for low-confidence words, and decision buttons (breach, no breach, needs full review).
For the model-calling layer I would use hone-models, which I built for exactly this kind of local setup: one interface over local servers, a registry that refuses a request that won't fit the model's context window instead of silently truncating it, GPU leases so transcription and the language model take turns on one card, and a local record of every call. Any equivalent layer works; the requirements are refusal on overflow and a record of which model version produced each finding.
Data inventory inside the boundary
| Artifact | Sensitivity | Retention |
|---|---|---|
| Original recording | High | Existing policy |
| Normalized audio | High | Deleted after transcription |
| Transcript with timestamps | High | Same as recording |
| Embeddings of transcript windows | High (can leak content) | Same as transcript |
| Prompts and model outputs | High | Same as transcript |
| Findings and reviewer decisions | High | Compliance record retention |
| Application logs | Must not contain content | Standard |
Logs are the easiest leak to miss. The logging configuration drops prompt and transcript fields, and a weekly check greps logs for transcript fragments.
Failure handling
| Failure | Detection | Response |
|---|---|---|
| Outbound connection attempt | Egress deny log | Alert platform owner; identify process |
| Queue age beyond review deadline | Admission control | Oldest calls routed to manual review; decision logged |
| Low word confidence on a material phrase | Per-word probability below threshold | Finding shows alternatives; reviewer listens first |
| Model output quote doesn't match transcript | Code check | Finding dropped and logged |
| GPU node down | Health check | Work drains to remaining node; backlog visible |
| New model version regresses on the fixed test set | Pre-activation test | Import rejected; previous version stays |
Capacity planning
Inputs to collect before buying hardware: calls per day and their length distribution, peak hour arrival rate, the review deadline, measured real-time factor for the chosen speech model on the candidate GPU, model memory footprint and concurrency, and headroom for one node being down. The illustrative 180 GPU-hours a week spread over a working week is about 4.5 GPUs busy full time for transcription alone, before the language model; overnight batch processing can halve the needed hardware if the review deadline allows next-day findings.
Evaluation
- Boundary: controlled outbound test (must be denied), egress log review, log content check.
- Transcription: word error rate on 50 reviewed calls, with separate attention to negations and numbers.
- Findings: recall of known breaches on 200 reviewed calls, per checklist item; precision measured as the share of findings reviewers accept.
- Operations: queue age distribution under a replayed peak week; restore test of the findings store.
Trade-offs I considered
- Hosted transcription and LLM with a data processing agreement. Cheaper and better models, and ruled out by the firm's constraints. If the constraints change, this design's review workflow still applies.
- Keyword spotting on transcripts only. Cheap, and catches exact phrases. Misses paraphrases and can't tell whether a risk disclosure was given. Useful as a fast first filter.
- Real-time checking during the call. Attractive for coaching, but doubles the infrastructure and changes the purpose. Out of scope.
Stack
Postgres for jobs and findings, GPU servers with faster-whisper and a quantized open-weight instruction model served by llama.cpp or vLLM, local embeddings, hone-models (or equivalent) for model calls and GPU leases, an internal web app for review, and the firm's existing monitoring. Model artifacts come from an internal registry with checksums.
Security and operations
Change control covers model files like code. Access to findings follows compliance roles. The boundary test is rerun after every infrastructure change, and its result is part of the monthly compliance report, stated as what it is: evidence that the tested path is blocked, not proof of overall compliance.