Customer support · Reference design
Support drafts from current policy
The citation was real. The policy had changed.
Check eligibility twice, before retrieval and before sending, so an AI draft can't quote superseded policy or answer a question the customer has since changed.
Company Ledgerline is a hypothetical company. Every figure is illustrative.
About Company Ledgerline
Company Ledgerline is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.
- Industry
- Subscription billing and invoicing software for small businesses
- Size
- About 300 people; 45 support agents across three shifts
- Systems
- Hosted help desk with saved replies, an internal policy wiki owned by billing operations, a public help center
- Volume
- About 14,000 tickets a month, roughly 30% about billing
- Change rate
- Refund and credit policy changes two to four times a month
- Constraints
- Agents send every reply; ticket data stays in EU hosting; no automatic credits
In brief
How the work ran before
Company Ledgerline sells billing software to small businesses. Its support team is 45 agents across three shifts, answering about 14,000 tickets a month. Nearly a third are about money: refunds, duplicate charges, plan changes.
The rules for those answers live in an internal policy wiki owned by billing operations. The refund window, what counts as a duplicate charge, when a credit is allowed. The pages change two to four times a month, usually after a pricing change. Agents also rely on saved replies in the help desk, and those drift out of date quietly.
To speed things up, Ledgerline's support team trialled an AI assistant that drafts replies. It searched the help center for the most similar articles and wrote an answer from them. Agents liked it on easy tickets.
Where it broke
A customer asks for a refund on an annual plan bought 20 days ago. The assistant finds a well-written help-center article about refunds and drafts a friendly reply offering a full refund within 30 days. The agent, busy, sends it.
The article was accurate once. Ledgerline's billing operations team had shortened the window to 14 days two weeks earlier and updated the wiki, but the older help-center article was still online and matched the question better than the new page. The draft even had a citation. The citation proved where the text came from. It did not prove the text was still the policy.
A second failure is harder to see. An agent opens a draft, gets pulled into a call, and comes back 20 minutes later. In the meantime the customer has replied: "Actually, I was charged twice." The draft on screen answers the old question.
Both failures come from the same gap. Relevance is not eligibility. A document can be the best match and still be the wrong source, and a draft can be well grounded and still be out of date.
What I would build
The fix is to check eligibility twice: once before retrieval, and once before sending.
- Before retrieval, filter. Only approved, current policy versions that apply to this customer's plan and region can enter the search. Superseded articles are excluded before ranking, however well they match.
- Bind the draft. Every draft records which policy versions it used and which revision of the ticket it answered.
- Check every claim. The model links each material statement ("you are eligible for a refund") to the passage that supports it. Code rejects links to sources that aren't eligible. Statements without support are removed, and the draft says what it couldn't answer.
- Before sending, re-check. When the agent clicks send, the system compares the draft's bound versions with the current ones. If the policy changed or the customer wrote again, the draft is marked stale and rebuilt.
The agent still reads, edits and sends every reply. The system's job is to make sure what they read is current.
Following one ticket through
Sample data throughout. The policy set is fixed for this example: billing policy v4 §3.2 covers ordinary refund windows; v4 §3.4 covers duplicate charges. The older help-center article refunds-help-v3 was superseded when v4 was approved at 09:02.
Ticket SUP-48213, revision 3, asks about a refund on a Growth monthly plan, EU customer.
| Candidate source | Status | Eligible? | Why |
|---|---|---|---|
| billing-policy-v4 §3.2 | Approved 09:02 | Yes | Current, applies to Growth plans in the EU |
| billing-policy-v4 §3.4 | Approved 09:02 | Yes, but not relevant yet | Duplicate charges; nothing in the ticket mentions one |
| refunds-help-v3 | Superseded 09:02 | No | Excluded before ranking |
Sample data.
The draft answers from §3.2 and binds to v4 and ticket revision 3. At 09:13 the customer adds, "this was a duplicate charge". That creates revision 4. When the agent clicks send at 09:15, the pre-send check sees revision 4 and marks the draft stale. The system rebuilds it against §3.4, and the agent sends a reply about the duplicate-charge process instead of a refund window.
Sample customer · revision 3
Asks for a refund on a monthly plan.
Sample customer · revision 4 · 09:13
"This was a duplicate charge."
Suggested reply · built on revision 3 · not sent
Draft answers the refund-window question from billing policy v4 §3.2.
When things go wrong
Two eligible sources disagree. Sometimes a regional addendum and the main policy both apply and say different things. The draft drops the affected statement and shows both passages to the agent. A confidence score does not settle a policy conflict; billing operations does.
Nothing eligible supports an answer. The draft says so: "I couldn't find a current policy for refunds on legacy plans." That's a useful draft. It tells the agent where to look and tells billing operations where the policy has a gap.
The support check is wrong. Linking a claim to a passage is itself a model judgement and can be mistaken. The design treats it as a filter that reduces risk, not a guarantee. The evaluation measures how often unsupported claims slip through.
What changes for the team
| Question | Before | After |
|---|---|---|
| Which policy did this answer use? | Unknown | Version and section stored on the ticket |
| Can an agent send a draft built on old policy? | Yes | Blocked at send and rebuilt |
| What happens after a policy change? | Macros and articles drift | Waiting drafts are invalidated automatically |
| How does QA audit billing answers? | Read and compare by hand | Filter by policy version |
Illustrative figures, computed from assumptions stated in the technical design. Not measured.
| Illustrative month (4,200 billing tickets) | Similarity-only drafts | Version-aware drafts |
|---|---|---|
| Drafts built on a superseded source | about 170 | 0 at retrieval, if lifecycle metadata is correct |
| Drafts stale at send (ticket changed) | not detected | about 250 caught and rebuilt |
| Drafts that abstain on at least one claim | 0 | about 300 |
Abstaining drafts look worse in a demo and are better in production. Each one marks a place where a confident wrong answer would otherwise have gone out.
What this does not solve
It depends on policy lifecycle data being right. If billing operations forgets to mark an article superseded, the filter will trust it. It doesn't write policy, and it can't make an unclear policy clear. And it only helps with questions that policy answers; an angry customer still needs a person.
How I would prove it
Replay 500 closed billing tickets from the last two months against the policy versions that were current at the time. Label each draft: correct and current, correct but cited a superseded source, contained an unsupported claim, or abstained. Then replay the same tickets against today's policy to see how many old answers would now be wrong. That second number shows billing operations how much a policy change costs in support.
The technical design
The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 4 minutes.
Drafts citing old policy?
Bring a few tickets and the policies behind them.
A handful of closed tickets and the policy versions agents used is enough to see where versioning creates risk. Remove customer details first.