Entity Matching When the Strings Do Not Match
Deduplication is easy until the easy cases are gone. What is left is the tail: a company that moved office, a supplier who files under a different legal suffix, a product listed by its SKU on one side and its marketing name on the other, a person whose name was transliterated twice.
Those pairs are decisions, not distance calculations. A decision model reads both records, weighs the identifiers against the surface text, and returns a probability you can threshold — plus a reason you can audit.
- Calibrated match probability
- Evidence label per pair
- Merge-safety score
- Cheap enough for blocked candidate sets
The problem
The Job: Decide Which Pairs Are Safe to Merge Automatically
String similarity is a good first filter and a bad decision procedure. Levenshtein distance says "Northwind Logistics Ltd" and "Northwind Logistics Limited (UK)" are close, and also says "Acme Corp" and "Acme Corps" are close — but one pair shares a company registration number and the other is two different businesses. The evidence that settles it is structured and the scoring function does not know that.
Deterministic rules on identifiers solve the pairs that have identifiers. The reason record linkage is still a job is that the interesting pairs have partial ones: a VAT number on one side, a domain on the other, an address that changed between the two imports.
Merging is destructive, so the cost of the two error types is wildly asymmetric. A missed match leaves a duplicate that someone fixes later. A bad merge fuses two customers’ histories, and people find out through a billing dispute. A system that returns only "match" or "no match" gives you nowhere to put that asymmetry; a probability plus a merge-safety score gives you two thresholds and a review queue between them.
Why a decision model
Why a Decision Model Beats Asking an LLM Whether Two Rows Are the Same
Record linkage runs on candidate sets, not on interesting examples. That changes what matters.
A probability is what a merge threshold needs
Auto-merge above one threshold, auto-reject below another, and send the band between them to a reviewer. That policy requires a calibrated number. "Yes, these look like the same company" cannot be tuned.
The evidence label makes the decision auditable
Asking which evidence was decisive — a shared registration number, a name-and-domain agreement, nothing but string similarity — turns a black box into something a data steward can spot-check and a rule you might promote into deterministic code.
Blocking produces a lot of pairs, so unit cost decides feasibility
Even after blocking, a mid-sized catalogue produces tens of thousands of candidate pairs. At the price measured below that is a job you run on every import; at frontier-model prices it is a proposal that gets deferred.
You can ask about a domain rule the model cannot guess
Write "a shared company registration number is decisive; an address change is not evidence against" into the instructions, and the judgement follows your data’s rules rather than a generic notion of similarity.
The same call gives you the review queue
Match probability, evidence, merge safety and a needs-review flag come back together, so the pipeline that decides also populates the queue for the pairs it declined to decide.
Two Records, One Company Number
Different name, different address, different domain, same registration number. Run it and see which evidence the model says was decisive.
Same company, or two companies? With a probability you can threshold
1Text
Two records518 / 100,000
2Questions
4 in this requestEditing is free. Sign in and you come straight back here with your text and questions — no card needed.
3Answers
Answers appear here
Each answer returns a probability for every option and a confidence score.
In code
What You Would Write Next
Pick your thresholds from your own labelled sample — the asymmetry between a missed match and a bad merge is a property of your business, not of the model.
async function scorePair(a: Record, b: Record) {
const { answers, usage } = await jev.decide({
state: `RECORD A\n${format(a)}\n\nRECORD B\n${format(b)}`,
questions: {
same_entity: {
type: "noul",
instructions:
"Do these two records describe the same real-world company? " +
"A shared company registration number is decisive evidence; " +
"an address change is not evidence against.",
},
evidence: {
type: "choice",
instructions: "Select the strongest evidence for or against a match.",
criteria: {
identifier_match: "A unique registered identifier is identical.",
identifier_conflict: "Two different unique identifiers are present.",
name_and_domain: "Names and domains line up allowing for suffixes.",
weak_or_none: "Nothing beyond a similar-looking name.",
},
},
},
});
return {
p: answers.same_entity.noul, // 0..1, calibrated
why: answers.evidence.choice,
tokens: usage.charged_tokens,
};
}
// Two thresholds and a queue, which is the whole policy.
const { p, why } = await scorePair(a, b);
if (p >= 0.97 && why === "identifier_match") return merge(a, b);
if (p <= 0.35) return keepSeparate(a, b);
return reviewQueue.push({ a, b, p, why });The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.
Real cost
One Run Costs $0.000032
Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the two records in the playground above, with all four questions in one call.
It checks out against the list rate of $0.042 per million input tokens: 761 ÷ 1,000,000 × 0.042 = $0.000032. The 114 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.
- Input tokens
- 761
- One run
- $0.000032
- 1,000 runs
- $0.032
- 1,000,000 runs
- $31.96
The only thing charged
Round trip 0.28s from a laptop, network included
Same questions, same length of input
At the model list rate, before our margin
Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.
Questions
Entity Matching: The Questions People Actually Ask
Do I still need blocking or a fuzzy-match first pass?
Yes. Comparing every row against every other row is quadratic and no per-pair price makes that sensible. Use cheap blocking keys to produce candidate pairs, then spend a decision on each candidate — that is where the ambiguity actually lives.
How do I choose the auto-merge threshold?
Label a few hundred pairs by hand, run them, and pick the threshold where your false-merge rate hits what the business can tolerate. Because the probabilities are calibrated, a threshold you set on that sample keeps roughly its meaning on the rest of the data.
Does this work for products and people, not just companies?
The shape is the same: two records in the state, your domain rules in the instructions, a probability out. What changes is which evidence you tell it to treat as decisive — a GTIN for products, a national identifier for people. Validate on your own data before trusting a threshold.
Can I send more than two records at once?
Keep one pair per call. A cluster judged in a single call gives you one answer where you need one per edge, and the answer stops being something you can threshold.
Is any of my data retained?
We pass your state to the upstream model to get the answer and store the request for your own history and billing. Read the privacy policy before sending regulated personal data, and consider sending identifiers rather than full records where your matching rules allow it.
What does a run cost at scale?
The measured figure for the example above is on this page, along with the price of a thousand and a million of them. Your pairs will differ in length; the charge is the input tokens the call reads, and output is free.
Related
Other Decisions in This Shape
LLM as a judge
Grade generated answers against a rubric and get sortable scores instead of a paragraph of praise.
RoutingLLM router
Decide which model, tool or workflow should handle a request, with a probability you can threshold.
RetrievalRAG evaluation
Check retrieval quality and answer grounding chunk by chunk, cheaply enough to run it on everything.
Run It on Your Own Data
Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.
