Compare Two Records and Return a Probability You Can Threshold
String similarity solves the easy pairs and then stops. What is left needs judgement about which evidence is decisive: a shared registration number outweighs a changed address; a similar name outweighs nothing at all.
These decisions are destructive downstream — a merge fuses two histories — so the useful output is a probability plus a reason, not a verdict.
Pages
Data Matching in Detail
Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.
The shape
What These Decisions Have in Common
They are asymmetric. A missed match is a duplicate someone fixes later; a wrong merge is a billing dispute. That asymmetry becomes two thresholds and a review queue, which needs a calibrated number to sit on.
They run on candidate sets. Blocking still leaves tens of thousands of pairs on a mid-sized catalogue, so unit cost decides whether this runs on every import or gets deferred forever.
They must be auditable. Asking which evidence was decisive gives a data steward something to spot-check, and sometimes a rule worth promoting into deterministic code.
Ready to run
Scenarios You Can Load in One Click
These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.
Résumé screening
Compare an application against the stated requirements of a role and return a match score, a next step and whether the hard requirements are met. Judges the written requirements only — it is a first-pass sort, not a hiring decision.
requirements_matchnext_stepmeets_domain_requirement
Questions
Data Matching FAQ
Do I still need blocking?
Yes. All-pairs comparison is quadratic and no per-pair price makes it sensible. Use cheap keys to generate candidates, then spend a decision on each candidate.
How do I pick the auto-merge threshold?
Label a few hundred pairs, run them, and pick the point where your false-merge rate hits what the business tolerates. Calibrated probabilities are what let a threshold set on a sample keep its meaning on the rest of the data.
Can I encode my own domain rules?
Write them into the instructions — "a shared company number is decisive, an address change is not evidence against" — and the judgement follows your data’s rules rather than generic similarity.
Try It on Your Own Text
Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.
