Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Technical design

Invoice checks before ERP entry

Every field matched. The bank account didn't.

← Back to the story

Company Brenner is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed. The story explains the problem and follows one item through the system; this page is the engineering detail behind it.

Scope and assumptions

The system starts when a document lands in the AP mailbox or scanner inbox and ends when a parked ERP invoice exists or an exception sits in an owner's queue. Payment, approval workflows and vendor-master maintenance are outside it.

Illustrative figures assume: 3,500 invoices a month; 20% of invoices fail at least one check (unknown supplier, PO mismatch, arithmetic, duplicate, bank); a clerk needs about 6 minutes to key and match a routine invoice and about 1.5 minutes to review a pre-filled, pre-checked one. Monthly routine clerk time: 3,500 × 6 = 21,000 minutes before; 2,800 × 1.5 + 700 × 6 = 8,400 minutes after. These are assumptions to be replaced by the pilot's timings.

Architecture

AP mailbox / scanner
      │
      ▼
 intake: hash, dedupe file, store original (immutable)
      │
      ▼
 page preparation: native text or OCR per page, page images
      │
      ▼
 extraction model ──► candidate record with page/region per field
      │
      ▼
 check engine (code): arithmetic · supplier · duplicate · PO/GR match · bank
      │
  ┌───┴──────────────────────────┐
  ▼                              ▼
 ERP adapter: parked invoice   exception queues: AP · procurement · treasury
  • Intake. Stores the original file under its SHA-256 hash, so a resent PDF is recognized before any processing. Splits multi-invoice PDFs by page ranges when a separator is detected, and otherwise treats the file as one package.
  • Page preparation. Uses embedded text when the PDF has it and OCR only for image pages. OCR confidence is stored per page, separately from model output, so a bad scan is reported as a bad scan.
  • Extraction. A vision-capable model or a document-extraction service returns JSON matching one schema. Every field has value, page, bbox and raw_text. Fields the model can't find are null, never guessed.
  • Check engine. Plain Python. Each check is a small function with a name, a version, inputs, a result and an owner for failures. Checks never modify candidates.
  • ERP adapter. Creates a parked invoice with the package ID in a reference field, then reads it back. The adapter is the only code with ERP credentials, and those credentials have no posting or payment rights.

Data model

Record Contents
package file hash, received_at, sender, page count, original file reference
candidate package, schema version, model and prompt version, extracted JSON with page/region per field
check_result candidate, check name and version, pass / fail / not applicable, detail, owner on failure
erp_draft candidate, ERP document ID, created_at, read-back status
exception package, failed checks, owner queue, due date, resolution and resolving user

Checks and their owners

Check Rule Owner on failure
Line arithmetic qty × unit price = line total, within 0.01 per line AP
Totals sum of lines = net; net + tax = gross AP
Tax plausibility tax rate in the allowed set for the supplier's country AP
Supplier identity tax ID or registered name matches exactly one vendor Procurement
Duplicate same supplier + invoice number, or same supplier + amount + date within 7 days AP
PO / goods receipt lines match open PO lines within configured price and quantity tolerance AP, then buyer
Bank details remittance account equals the vendor master Treasury, always

The bank check has no tolerance and no override in the pipeline. The only way past it is treasury changing the vendor master through its own procedure.

Failure handling

Failure What happens
Unreadable scan Package goes to AP as "manual entry", with the OCR confidence shown
Model returns invalid JSON One retry with the validation error; then manual entry
Model returns a field without a page reference Field treated as missing
ERP unavailable Draft creation waits and retries; the package stays "checked, not drafted"
ERP response lost after create Search by package reference before retrying
Same file sent twice Second copy linked to the first by hash; no second candidate
Exception not handled in 3 working days Escalates to the AP lead

Evaluation

The pilot set is 300 already-entered invoices with the clerk's entry as reference, plus 30 constructed hard cases (continuation pages without headers, credit notes, multi-currency, a changed bank token). Metrics:

  • field accuracy per field after normalization (dates, number formats), reported separately for header fields, line items and remittance data;
  • false-draft rate: invoices the system would have parked that the clerks corrected afterwards (the key number);
  • exception precision: share of held invoices that a clerk agrees needed attention;
  • time per invoice for review versus manual entry, measured on a sample with a stopwatch, not estimated.

I would rerun the same set whenever the model, prompt, OCR engine or schema changes, and block the change if the false-draft rate rises.

Trade-offs I considered

  • Buy an AP automation product. Often the right answer. The design above is also a checklist for evaluating one: does it keep page evidence per field, can the bank check be made unskippable, and can you see why an invoice was held?
  • Template-based OCR per supplier. Accurate for the top suppliers, brittle for the long tail of 600. A model handles layout variety better; checks catch its mistakes.
  • Auto-post when everything passes. Tempting for small, PO-matched invoices. I would not start there. Parked drafts keep the four-eyes rule intact and cost only seconds per invoice.

Stack

Python services, Postgres for records, object storage for originals and page images, an OCR engine for scans (Tesseract or a cloud OCR within the EU region), a vision-capable model with JSON-schema output for extraction, and the ERP's standard API for parked invoices. A small internal review page shows the page image with highlighted regions next to each failed check.

Security and operations

Invoices contain bank details and personal data of sole traders. Originals stay in EU storage with access limited to AP, treasury and audit. The model provider must not retain inputs. Monitoring watches four numbers daily: packages received, drafts created, exceptions opened and exceptions older than three days.