Technical design
Contract review across linked clauses
The cap matched the playbook. The definition three documents away didn't.
Company Halden is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed. The story explains the problem and follows one item through the system; this page is the engineering detail behind it.
Scope and assumptions
The system prepares review packets for incoming agreement packages against the approved playbook. It starts when a package is uploaded to the contract lifecycle tool and ends when a reviewer records a decision per topic. It gives no legal advice and approves nothing.
Illustrative figures assume: 120 packages a month; 8% arrive with a referenced document missing (≈10); a standard package takes about 90 minutes to read end to end and about 45 minutes packet-first; 120 × 1.5 h = 180 h versus 120 × 0.75 h = 90 h. The pilot replaces these with measured times.
Architecture
package upload (CLM)
│
▼
document preparation: text + layout per page, OCR confidence, clause segmentation
│
▼
inventory: document types, versions, dates, referenced-but-missing documents
│
▼
reference graph: cross-references · defined terms · amendment targets (code)
│ + suggested implicit links (model, marked)
▼
topic assembly: start clause per playbook topic → expand along links (depth-limited)
│
▼
playbook comparison (model, quotes required)
│
▼
reviewer packet ──► decision per topic (reviewer) ──► stored with packet version
Clause segmentation and the graph
Clauses are segmented using numbering patterns ("14.2", "§14.2", "Clause 14(2)") and headings, with each clause keeping document ID, page range and text. Edges:
| Edge | Found by | Example |
|---|---|---|
| cross_reference | regex over clause text, resolved against segmented clause IDs | "subject to §14.4" |
| uses_term → defines_term | defined-term extraction: capitalized terms with a definition pattern | "Confidential Information" |
| amends | amendment clauses naming a target document and clause | "Clause 3.1 of the DPA is replaced by…" |
| precedence | clauses containing precedence language | "in case of conflict, this Order Form prevails" |
| suggested (model) | model proposal per amendment, with quoted evidence | "obligations regarding data are extended" |
Suggested edges are stored with the model version and are always rendered as suggestions.
Topic assembly
The playbook is converted into topics, each with an approved position, fallback positions, and one or more start patterns (for "liability cap": clauses whose heading or text match limitation-of-liability language). Assembly starts there and walks the graph breadth-first to depth 3, following cross_reference, uses_term and amends edges. Precedence clauses are always included for every topic, because they can change the meaning of everything else.
The depth limit is a trade-off. Deeper walks pull in more of the contract and make the packet longer; shallower ones risk missing a chain. The evaluation tunes it.
Comparison output
For each topic the model returns: summary of the assembled language, the playbook position it compares against, deviation (none / possible / likely), the quoted spans for each statement, and open questions. Code verifies each quote appears verbatim in the cited clause. Statements with unverifiable quotes are removed.
Failure handling
| Failure | Response |
|---|---|
| Referenced document missing | Package marked incomplete; review can proceed only with the gap acknowledged |
| Cross-reference to a clause not found | Edge marked unresolved and shown |
| OCR confidence low on a page | Page flagged; clauses on it listed as uncertain |
| Conflicting precedence clauses | Both shown; no resolution attempted |
| Quote not found in cited clause | Statement removed from the comparison |
| Playbook topic has no start clause match | Packet says the topic wasn't located |
Evaluation
Gold set: 30 reviewed packages with the lawyer's notes marking which clauses mattered for each topic. Metrics:
- clause recall per topic (did the packet include every clause the lawyer relied on), split by explicit and implicit links;
- packet length (clauses per topic), since a packet that includes everything is useless;
- incompleteness detection (packages with a missing referenced document);
- deviation agreement with the lawyer's position, reported as agreement, not accuracy, because legal judgements can differ;
- reviewer time, measured.
Trade-offs I considered
- Plain retrieval over chunks. Works for self-contained questions. It is structurally blind to "subject to §14.4", which is the whole problem.
- Long-context model over the full package. Increasingly feasible and worth testing as a baseline. Its weakness is traceability: it may reach the right conclusion without showing the chain, and it won't flag a missing document it never saw.
- A contract lifecycle tool's built-in AI. Worth evaluating first if the team already has one. The questions to ask it are the ones above: does it follow references, and does it show them?
Stack
Document parsing with a layout-aware PDF library and OCR for scans, Python for segmentation and graph building, Postgres for clauses and edges (the graph is small enough that a graph database adds nothing), an EU-hosted model with zero retention for suggestions and comparisons, and an integration with the CLM tool to attach the packet and record decisions.
Security and operations
Agreements are confidential and sometimes privileged. Processing stays in the EU, access follows the CLM's permissions, and model inputs are limited to the clauses in the packet. The playbook owner gets a monthly report of topics where reviewers disagreed with the comparison most often; those are usually playbook wording problems.