Sales compliance · Reference design
Call review inside a private network
The recording stayed home. One search query didn't.
Transcribe and check every call on the firm's own hardware, enforce the boundary with the network, and send reviewers to timestamped moments instead of whole calls.
Company Aurum is a hypothetical company. Every figure is illustrative.
About Company Aurum
Company Aurum is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.
- Industry
- Distributor of regulated insurance and investment products by phone
- Size
- About 250 advisers; a compliance team of nine
- Systems
- Telephony recording to on-premises storage, a CRM with advice records, a private data centre with a small GPU budget
- Volume
- About 6,000 calls a week, 18 minutes on average
- Constraints
- Recordings and anything derived from them stay inside the private network; no hosted AI; change control for model updates
In brief
How the work ran before
Company Aurum sells regulated insurance and investment products by phone. About 250 advisers make roughly 6,000 calls a week, averaging 18 minutes each. Every call is recorded to storage inside Aurum's own data centre.
A compliance team of nine checks the calls. They pick a random 2% sample, listen at one and a half times speed, and tick a checklist: did the adviser explain the risks, disclose the costs, ask the suitability questions, and avoid promising returns? A reviewer gets through perhaps ten calls a day, and each one takes full concentration.
Where it broke
A 2% sample is a lottery. An adviser who tells customers "you can't lose money on this" might be sampled twice a year. An external review eventually listened to calls nobody had sampled and found exactly that phrase, repeatedly, from one adviser, over several months.
The obvious answer is to transcribe every call and let a model check it. Aurum's contracts and its regulator's expectations rule out the obvious way to do that. Recordings may not leave the private network, and neither may anything made from them: transcripts, summaries, even the numerical embeddings a search system builds. A cloud transcription API is out. So is a cloud language model.
There's a subtler trap. A team can build a "local" pipeline that transcribes on-premises and still leaks. If one step calls a hosted embedding service to search the policy library with a sentence from the transcript, the recording stayed home but its content left.
What I would build
A pipeline where every step that touches call content runs inside the boundary, and the boundary is enforced by the network, not by good intentions.
- Transcribe locally. A speech-recognition model on Aurum's own GPU servers produces a transcript with word-level timestamps and confidence, and separates the speakers.
- Check against the checklist. A local language model reads the transcript and, for each checklist item, returns either the timestamped passage that satisfies it, the passage that may breach it, or "not found".
- Send reviewers to the moment, not the call. The reviewer's queue shows each finding with a link that plays the audio from ten seconds before the passage. Reviewers listen to the words that matter.
- Deny everything else. The servers that process calls can't reach the internet at all. Model updates come in through a controlled import, like any other software change.
- Show what's waiting. If a busy week produces more calls than the servers can process, the backlog is visible as a backlog. An unprocessed call is never shown as "no findings".
Following one call through
A constructed call, sample data. Call C-2291, 17 minutes, adviser and customer on separate channels.
| Checklist item | Finding | Timestamp | Transcript confidence | Route |
|---|---|---|---|---|
| Risk disclosure | Found: "the value can go down as well as up" | 06:41 | High | None |
| Cost disclosure | Found: ongoing charge stated | 09:12 | High | None |
| Suitability questions | Partly found: income asked, investment horizon not found | none | none | Reviewer |
| No guaranteed returns | Possible breach: "you [can / can't] lose on this" | 13:05 | Low on one word | Reviewer, listen first |
Constructed sample.
The last row is the interesting one. Speech recognition often confuses "can" and "can't", because the difference is a short sound at the end of a word. The transcript's confidence on that word is low. The system doesn't guess. It sends the reviewer to 13:05 with both readings shown, and the reviewer spends 20 seconds listening instead of 17 minutes.
context before segment after context
Checklist item: no guaranteed returns · possible breach
Customer: "So what happens if the market drops?"
Adviser: "Honestly, you can / can't lose on this one."
low confidence on one word · both readings shown
the reviewer listens from 12:55 before deciding
When things go wrong
A step tries to call out. Say a developer adds a library that phones home, or someone wires in a hosted embedding API for convenience. The network policy blocks the connection, and the block is logged with the process and destination. The test for this is deliberate: before launch, a controlled outbound request is sent from the processing servers to a test destination, and it must be denied.
The GPUs are saturated. Month-end, a campaign, a server down for maintenance. Calls queue. The dashboard shows the queue age and which calls are waiting. If the backlog passes the review deadline, the oldest calls go to reviewers the old way, and that decision is recorded.
Speakers are mixed up. Diarization separates voices; it doesn't know who is who. The adviser channel from the telephony system is used when available. When it isn't, a finding that depends on who said something goes to review.
A model update changes behaviour. Model files are imported through change control and tested on a fixed set of reviewed calls before they're switched on. The previous version stays available for rollback.
What changes for the team
| Question | Before | After |
|---|---|---|
| Share of calls checked at all | 2%, by listening | Every call, first pass by the pipeline |
| What a reviewer listens to | Whole calls at 1.5× | Timestamped segments, whole calls when needed |
| Advisers with a repeated pattern | Found by luck | Visible across calls, per adviser |
| Where call data goes | Recording storage | Recording storage and processing servers, nowhere else |
Illustrative figures from assumptions in the technical design. Not measured.
| Illustrative week (6,000 calls, 18 min average) | Manual sample | Local first pass + targeted review |
|---|---|---|
| Calls with any automated check | 0 | 6,000 |
| Calls a reviewer listens to in full | about 120 | about 60 (low-confidence or complex) |
| Segments reviewed | none | about 900 |
| GPU hours needed (at real-time factor 0.1 for transcription) | 0 | about 180 plus model checks |
The GPU line is the price of the boundary. Hosted services would be cheaper and easier. Aurum has decided they aren't allowed, so the cost belongs in the design from day one.
What this does not solve
It doesn't decide whether a call breached a rule; reviewers do. Transcription errors can hide a breach as well as invent one, so "no findings" means "no findings in the transcript", not "compliant". A private deployment is slower to update than a cloud service and needs people to patch, monitor and restore it. And the network boundary proves only what it was tested against.
How I would prove it
Two separate proofs. For the boundary: run the outbound-request test, and review the egress logs of a full week for any denied attempts. For usefulness: take 200 calls the compliance team already reviewed, including known breaches, run them through the pipeline, and check whether each known breach produced a finding with a timestamp pointing at it. Report misses by checklist item, because missing a guaranteed-return phrase is worse than missing a cost disclosure the adviser gave in an email.
The technical design
The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 5 minutes.
Data that can't leave the building?
Bring your boundary and your volumes.
Call volume, review deadlines and what may not leave your network decide the architecture. No recordings needed for a first conversation.