On this page
Classifier configuration
Configure a Trazadera Golden weighted classifier with dataset-based comparisons, weights, options, and decision thresholds.
A classifier-weight resource evaluates records that an
indexer placed in the same candidate group. The indexer determines
which records meet; the classifier determines the resulting classification.
Core settings
This is sample-customer-classifier from the
Aurelia Utilities sample, shortened to four
of its seven mappings:
{
"type": "classifier-weight",
"_id": "sample-customer-classifier",
"description": "Decides whether two records in a bucket are the same person",
"dataset": "sample-customer",
"nonMatchThreshold": 0.55,
"matchThreshold": 0.85,
"bucketSizeReviewThreshold": 10,
"defaultWeight": 1.0,
"defaultOptions": ["AGGRESSIVE"],
"mappings": [
{ "id": "tax-id", "keys": ["taxId"], "weight": 3.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
{ "id": "email", "keys": ["email"], "weight": 2.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
{ "id": "full-name", "keys": ["fullName"], "weight": 2.0, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false },
{ "id": "phone", "keys": ["phone"], "weight": 1.5, "options": ["AGGRESSIVE"], "object": false, "nullMismatch": false }
]
}
The remaining three are street at 1.0, and postcode and city at 0.5 each,
for a total weight of 10.5.
Comparisons are derived from dataset column semantics and mapping settings.
There is no comparer property with values such as LEVENSHTEIN or
JARO_WINKLER in this resource model.
| Mapping field | Purpose |
|---|---|
id | Name shown with the comparison result |
keys | One or more same-type dataset columns compared together |
weight | Relative influence compared with other mappings |
object | Compare supported date, number, or coordinate values as their semantic type |
nullMismatch | Decide whether a missing-versus-present value counts as disagreement |
threshold | Allowed semantic distance for supported object comparisons |
options | Text normalization override for the mapping |
Text options include NONE, IGNORE_CASE, IGNORE_SPACES,
ONLY_ALPHANUMERIC, IGNORE_ACCENTS, and AGGRESSIVE. Apply them according to
the field’s business meaning; identifiers often need stricter normalization
than names or addresses.
Comparison scores
The reported score summarizes comparisons according to the configured mappings and weights. Read both the result and its field-level evidence; a high score is not a probability that the records belong to one subject.
GET /api/golden/{entity}/duplicates/clusters/{clusterId}/evaluate returns
those similarities so a reviewer can inspect the evidence. Two
clusters from the sample:
| Column | Weight | Ana’s pair | Two different people called Javier Torres Melgar |
|---|---|---|---|
taxId | 3.0 | 1.000 | 0.000 |
email | 2.0 | 1.000 | 0.983 |
fullName | 2.0 | 0.986 | 1.000 |
phone | 1.5 | 0.000 | 0.000 |
street | 1.0 | 0.863 | 0.969 |
postcode | 0.5 | 1.000 | 1.000 |
city | 0.5 | 1.000 | 1.000 |
| Score | 10.5 | 84 | 56 |
The Javier pair has matching names but different tax identifiers. That contradiction matters even when other fields resemble each other. Choose weights that reflect the business significance of each field and test known non-matches.
The phone similarity is 0.000 in both cases. 611000000 against +34 611 00 00 00 is
one number written two ways, and an exact comparison sees two different
strings. Normalize phone values in a transformation before
comparing them, then evaluate the mapping’s contribution on representative data.
Mappings default to object: false, threshold: 50 for object comparison,
and nullMismatch: false. A mapping without an explicit weight uses the
classifier’s defaultWeight, initially 1.0.
Thresholds
| Setting | Default | Purpose |
|---|---|---|
nonMatchThreshold | 0.5 | Upper boundary for non-matches |
matchThreshold | 0.8 | Lower boundary for matches |
reviewScoreThreshold | 0.75 | Minimum representative score for the default candidates view; unset or invalid values use this default |
bucketSizeReviewThreshold | 10 | Groups larger than this require review |
reviewScoreThreshold accepts values above 0 and up to 1. It controls which
cases appear in the default review queue, independently of their classification.
Configuration thresholds use a scale from 0 to 1. Reported cluster scores
use integers from 0 to 100. A matchThreshold of 0.85 corresponds to a
reported score of 85.
| Where | Scale | Example |
|---|---|---|
matchThreshold, nonMatchThreshold in this resource | 0 to 1 | 0.85, 0.55 |
score on a cluster, and the interface | 0 to 100 | 84, 56 |
A steward rule’s SCORE condition | 0 to 100 | min: 90 |
Multiply by 100 to go from one to the other, and state which one you mean when
you write a runbook. A steward rule written with min: 0.9 does not select the
clusters scoring 90 and above; it applies a threshold close to zero on the reported score scale. See Configure automatic stewardship.
Scores at or below nonMatchThreshold become NON_MATCH; scores at or above
matchThreshold become MATCH; values between them become REVIEW.
bucketSizeReviewThreshold sends large groups to review as a safety control,
whatever they scored.
The interval between the thresholds affects review volume. In the initial
sample, thresholds of 0.55 and 0.85 yield 31 REVIEW groups and 214 MATCH
groups. Classification alone does not apply a merge.
Tune these settings with labeled representative examples. Do not copy the illustrative values above into production without measuring false matches, missed duplicates, and review volume for your data.
Related reading
- Configure an indexer
- Configure the merger
- Configure automatic stewardship
- Steward duplicate candidates
- Review duplicates manually, which walks these two clusters