On this page

A classifier-weight resource evaluates records that an indexer placed in the same candidate group. The indexer determines which records meet; the classifier determines the resulting classification.

Core settings

This is sample-customer-classifier from the Aurelia Utilities sample, shortened to four of its seven mappings:

{
  "type": "classifier-weight",
  "_id": "sample-customer-classifier",
  "description": "Decides whether two records in a bucket are the same person",
  "dataset": "sample-customer",
  "nonMatchThreshold": 0.55,
  "matchThreshold": 0.85,
  "bucketSizeReviewThreshold": 10,
  "defaultWeight": 1.0,
  "defaultOptions": ["AGGRESSIVE"],
  "mappings": [
    { "id": "tax-id",    "keys": ["taxId"],    "weight": 3.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
    { "id": "email",     "keys": ["email"],    "weight": 2.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
    { "id": "full-name", "keys": ["fullName"], "weight": 2.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
    { "id": "phone",     "keys": ["phone"],    "weight": 1.5, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false }
  ]
}

The remaining three are street at 1.0, and postcode and city at 0.5 each, for a total weight of 10.5.

Comparisons are derived from dataset column semantics and mapping settings. There is no comparer property with values such as LEVENSHTEIN or JARO_WINKLER in this resource model.

Mapping fieldPurpose
idName shown with the comparison result
keysOne or more same-type dataset columns compared together
weightRelative influence compared with other mappings
objectCompare supported date, number, or coordinate values as their semantic type
nullMismatchDecide whether a missing-versus-present value counts as disagreement
thresholdAllowed semantic distance for supported object comparisons
optionsText normalization override for the mapping

Text options include NONE, IGNORE_CASE, IGNORE_SPACES, ONLY_ALPHANUMERIC, IGNORE_ACCENTS, and AGGRESSIVE. Apply them according to the field’s business meaning; identifiers often need stricter normalization than names or addresses.

Comparison scores

The reported score summarizes comparisons according to the configured mappings and weights. Read both the result and its field-level evidence; a high score is not a probability that the records belong to one subject.

GET /api/golden/{entity}/duplicates/clusters/{clusterId}/evaluate returns those similarities so a reviewer can inspect the evidence. Two clusters from the sample:

ColumnWeightAna’s pairTwo different people called Javier Torres Melgar
taxId3.01.0000.000
email2.01.0000.983
fullName2.00.9861.000
phone1.50.0000.000
street1.00.8630.969
postcode0.51.0001.000
city0.51.0001.000
Score10.58456

The Javier pair has matching names but different tax identifiers. That contradiction matters even when other fields resemble each other. Choose weights that reflect the business significance of each field and test known non-matches.

The phone similarity is 0.000 in both cases. 611000000 against +34 611 00 00 00 is one number written two ways, and an exact comparison sees two different strings. Normalize phone values in a transformation before comparing them, then evaluate the mapping’s contribution on representative data.

Mappings default to object: false, threshold: 50 for object comparison, and nullMismatch: false. A mapping without an explicit weight uses the classifier’s defaultWeight, initially 1.0.

Thresholds

SettingDefaultPurpose
nonMatchThreshold0.5Upper boundary for non-matches
matchThreshold0.8Lower boundary for matches
reviewScoreThreshold0.75Minimum representative score for the default candidates view; unset or invalid values use this default
bucketSizeReviewThreshold10Groups larger than this require review

reviewScoreThreshold accepts values above 0 and up to 1. It controls which cases appear in the default review queue, independently of their classification.

Configuration thresholds use a scale from 0 to 1. Reported cluster scores use integers from 0 to 100. A matchThreshold of 0.85 corresponds to a reported score of 85.

WhereScaleExample
matchThreshold, nonMatchThreshold in this resource0 to 10.85, 0.55
score on a cluster, and the interface0 to 10084, 56
A steward rule’s SCORE condition0 to 100min: 90

Multiply by 100 to go from one to the other, and state which one you mean when you write a runbook. A steward rule written with min: 0.9 does not select the clusters scoring 90 and above; it applies a threshold close to zero on the reported score scale. See Configure automatic stewardship.

Scores at or below nonMatchThreshold become NON_MATCH; scores at or above matchThreshold become MATCH; values between them become REVIEW.

bucketSizeReviewThreshold sends large groups to review as a safety control, whatever they scored.

The interval between the thresholds affects review volume. In the initial sample, thresholds of 0.55 and 0.85 yield 31 REVIEW groups and 214 MATCH groups. Classification alone does not apply a merge.

Tune these settings with labeled representative examples. Do not copy the illustrative values above into production without measuring false matches, missed duplicates, and review volume for your data.

Golden 3.0.0 · Published 2026-10-04