Classify the document problem
Invoices, purchase orders, emails, scans and contract sections fail differently. Before selecting a model, inventory document classes, languages, page counts, image quality, table density, handwriting, field variation and downstream consequences. Keep a representative sample with permission and a synthetic set for routine tests.
A single “document accuracy” number hides the difference between missing an optional reference and inventing a bank account. Define field criticality and allowed null behaviour.
Normalise before interpretation
Validate file type and size, scan or quarantine as required, hash for duplicate detection and preserve an immutable source reference. Extract native text when available. Use OCR for image-only pages and record OCR quality separately from model quality.
Page boundaries, tables and reading order matter. Provide the model only the necessary content and stable page/source identifiers so a reviewer can trace each field.
Design the schema with operations
Start from the destination record and the exception process. Mark fields required, optional, derived or forbidden. Use bounded enums where the business already has a taxonomy. Include source spans or page references for critical fields.
Do not ask the model to “complete” missing values. A missing purchase-order number should become null with a review reason, not a plausible-looking guess.
Validate in layers
First validate JSON and types. Then run cross-field checks: line items plus tax equal total; currency matches purchase order; invoice date is sensible; vendor exists; duplicate key is absent; contract clause actually occurs in the cited page.
A repair pass can fix formatting when the source supports the value. It must not invent a value to satisfy the schema. Persist both the original response and repair reason when policy permits.
Evaluate fields, not vibes
The demonstration extraction benchmark reports exact-match or normalised field accuracy, schema validity, missed fields, hallucinated fields, latency and provider cost. Critical fields are weighted separately, and documents are sliced by class and quality.
A useful test set includes clean normals, messy normals, rare layouts and adversarial ambiguities. Track abstention: a system that safely refuses difficult records can create more value than one with a higher headline coverage and hidden false values.
-
Schema validity: can downstream code parse the record?
-
Field accuracy: is each value correct after legitimate normalisation?
-
Hallucination rate: are unsupported values introduced?
-
Coverage: what share can progress without transcription?
-
Review time: does the exception path actually save effort?
Route by reason, not one confidence score
Separate unreadable source, unknown vendor, missing required field, arithmetic mismatch and policy exception. Each has a different owner and remedy. Show field-level evidence in the review interface.
For an invoice, successful extraction can create an ERP draft after deterministic checks. It should not approve payment. This boundary removes transcription while preserving finance control.
Operate the pipeline
Monitor document mix, OCR failures, invalid schema, field errors, review volume, cost and latency. Re-run the benchmark when models, prompts, OCR or schema change. Reconcile intake documents with destination records every day.
Extraction is reliable when operators can see uncertainty, correct it efficiently and prove where each consequential field came from.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real