Problem shape
The problem beneath them
Preserve source records, compare plausible candidates with consequence-aware evidence, and make every canonical identity decision attributable and reversible.
Identifiers are rarely complete across system boundaries. Names vary, addresses age, phone numbers are reassigned, and records are copied. Exact equality is therefore too strict, while fuzzy similarity alone is too permissive. The cost is asymmetric: a false merge can contaminate history, permissions, balances, or communication; a missed match may create duplicate work but remain easier to repair.
Entity resolution should preserve three distinctions. A source record is what a system observed. A resolved entity is a current decision about which observations belong together. A canonical representation is a selected or synthesized view used downstream. Collapsing these into one mutable row destroys the ability to explain or reverse a decision.
The mechanism combines deterministic exclusions, exact identifiers, candidate generation, multi-signal comparison, consequence-aware thresholds, and explicit human authority. It also treats “needs review” as a valid controlled outcome. A forced binary choice is not more decisive when the available evidence cannot support it.
Similarity is evidence, not identity. Consequential merges require sufficient evidence and a reversible decision trail.
This pattern can organize identity evidence; it cannot prove real-world identity where authoritative evidence is unavailable. Canonicalization must not erase source records, transfer permissions by implication, or turn a probabilistic score into legal identity.
Different problems, same shape
The surface changes. The decision structure persists.
Are these listings the same product?
Names overlap while sellers, variants, and identifiers differ.
Are these records the same company?
Domains, legal names, addresses, and contacts only partially agree.
Are these tickets about the same incident?
Symptoms and timing align, but affected accounts differ.
Are these articles copies of the same source?
Text is similar while publication dates and attribution conflict.
All four questions converge into Entity Resolution with Reversible Canonicalization.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Preserve original records
- Question
- What exactly did each source assert before resolution?
- Responsibility
- Store source payload, source identity, observation time, and ingestion provenance without destructive merging.
- Input
- Raw records and trustworthy source metadata.
- Output
- Immutable source observations with stable record identifiers.
- Stops when
- Every comparison can refer back to intact source evidence.
-
02
Normalize without destroying provenance
- Question
- Which superficial differences can be made comparable?
- Responsibility
- Produce normalized names, domains, addresses, dates, and identifiers alongside—not over—the originals.
- Input
- Preserved source fields and versioned normalization rules.
- Output
- Comparable derived fields linked to their source values.
- Stops when
- Transformations are reproducible and original meaning remains available.
-
03
Generate plausible candidates
- Question
- Which record pairs deserve comparison without scanning everything?
- Responsibility
- Use blocking keys, search indexes, or locality methods to retrieve plausible matches with acceptable recall.
- Input
- Normalized attributes, entity scope, and candidate policy.
- Output
- Candidate pairs with retrieval reasons.
- Stops when
- The bounded candidate set and excluded scopes are recorded.
-
04
Apply exact rules
- Question
- Do authoritative matches or contradictions settle the pair?
- Responsibility
- Apply hard exclusions, scoped unique identifiers, and deterministic equivalence rules before probabilistic comparison.
- Input
- Candidate pair and verified identifiers.
- Output
- Same, different, or unresolved rule disposition.
- Stops when
- A sufficient rule settles the pair or passes it onward.
-
05
Measure multi-signal similarity
- Question
- How strongly do independent attributes support or oppose one entity?
- Responsibility
- Compare names, identifiers, location, chronology, relationships, and domain-specific signals without hiding contradictions.
- Input
- Candidate records and comparable features.
- Output
- Per-signal evidence, conflicts, missingness, and an aggregate assessment.
- Stops when
- Material positive and negative evidence is exposed for decision.
-
06
Consider confidence and consequence
- Question
- Is the evidence sufficient for the harm profile of this decision?
- Responsibility
- Apply calibrated thresholds and stricter requirements for irreversible, financial, access, or compliance effects.
- Input
- Evidence assessment, proposed action, reversibility, and false-decision costs.
- Output
- Auto-link eligibility, keep-separate decision, or review requirement.
- Stops when
- The route reflects both confidence and consequence.
-
07
Route uncertainty
- Question
- Who can resolve ambiguity, and what must they see?
- Responsibility
- Create a review packet with source records, signals, contradictions, and permitted actions.
- Input
- Unresolved pair, evidence, policy, and review capacity.
- Output
- Prioritized review item or deliberate unresolved state.
- Stops when
- An authorized owner accepts the item or policy permits deferral.
-
08
Canonicalize or keep separate
- Question
- What grouping and representative view should downstream systems use?
- Responsibility
- Link records to an entity, retain separate entities, and select canonical attributes under field-level provenance rules.
- Input
- Authorized disposition and source observations.
- Output
- Entity membership and canonical view without deleting sources.
- Stops when
- The disposition is applied consistently and downstream scope is explicit.
-
09
Record and reverse
- Question
- Can the decision be reconstructed and safely undone?
- Responsibility
- Version membership, rationale, actor, evidence, affected dependents, and reversal operations.
- Input
- Prior entity graph, new disposition, and decision evidence.
- Output
- Append-only decision event and compensating reversal path.
- Stops when
- Both forward impact and reversal requirements are known.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Same authoritative identifier in valid scope | Treat as same entity | Scoped identifier outweighs cosmetic differences | Identifier reuse or source authority is disputed |
| Mutually exclusive verified identifiers | Keep separate | Contradictory identity evidence blocks merging | A source is known to issue duplicates |
| Strong similarity with no material conflict | Auto-link when consequence policy allows | Multiple independent signals support the same explanation | Merge would transfer consequential state |
| Sparse records with moderate similarity | Needs review or remain unresolved | Missing evidence is not positive evidence | Delay itself creates material harm |
| High text similarity but incompatible variant | Keep separate | Product or content variant is identity-relevant | Variant taxonomy is incomplete |
| Existing canonical entity gains contradictory data | Re-open resolution | New evidence can invalidate an earlier decision | Reversal affects many downstream records |
| Reviewer lacks authority or evidence | Defer, do not force | A review click cannot manufacture certainty | Service expectations or consequence require specialist action |
Operating paths