On this page
Compare records and group candidates
Understand indexing, comparison, classification, buckets, and clusters before deciding whether records represent the same subject.
Golden first selects records worth comparing, then classifies those candidates. A stewardship decision determines whether to merge, separate, or retain them for review. Each stage answers a different question:
Indexer → Classifier → Stewardship decision → Merger
Candidate generation
The indexer creates named search and candidate keys from dataset fields. Records that share a relevant key can appear in the same bucket or logical cluster. Search-only mappings can opt out of duplicate candidate generation.
If a known duplicate never becomes a candidate, review mappings, column names, normalization, compound-key choice, and matching type before changing classifier thresholds.
Candidate classification
The classifier compares configured dataset values within a candidate and assigns an outcome:
| Classification | Customer-visible meaning |
|---|---|
MATCH | The configured evidence meets the match boundary |
NON_MATCH | The evidence meets the non-match boundary |
REVIEW | The evidence falls between the decision boundaries or requires review |
IGNORE | The candidate is set aside until restored |
Mappings choose the columns, relative weights, missing-value behavior, semantic comparison, and text normalization. The two thresholds establish the three main decision regions.
GET /api/golden/{entity}/duplicates/clusters/{clusterId}/evaluate returns
pairwise scores and field-level similarities. Inspect which fields agree and
which contradict the proposed identity. The final score is evidence under the
configured comparison policy, not a probability or a count of agreeing fields.
Candidates requiring review
REVIEW identifies candidates between the configured decision thresholds or
subject to another review condition, such as group size. Inspect their evidence
before deciding whether to merge or separate records.
In the initial Aurelia Utilities sample, 214 of 245
groups are MATCH and 31 are REVIEW. Validate threshold changes against
known matches and non-matches, as well as the resulting review volume.
Buckets and clusters
Two words appear in the interface and the API, and they are not synonyms:
| Group | Definition |
|---|---|
| Bucket | One index’s contribution: the records that share one key of one mapping |
| Cluster | The logical group a person decides on, made of one or more buckets |
A cluster reports the bucketIds it is made of. The API has a parallel
operation family for each. Use the identifier returned by the corresponding list operation.
Continue with Resolve groups and subsets.