On this page

Golden first selects records worth comparing, then classifies those candidates. A stewardship decision determines whether to merge, separate, or retain them for review. Each stage answers a different question:

Indexer → Classifier → Stewardship decision → Merger

Candidate generation

The indexer creates named search and candidate keys from dataset fields. Records that share a relevant key can appear in the same bucket or logical cluster. Search-only mappings can opt out of duplicate candidate generation.

If a known duplicate never becomes a candidate, review mappings, column names, normalization, compound-key choice, and matching type before changing classifier thresholds.

Candidate classification

The classifier compares configured dataset values within a candidate and assigns an outcome:

ClassificationCustomer-visible meaning
MATCHThe configured evidence meets the match boundary
NON_MATCHThe evidence meets the non-match boundary
REVIEWThe evidence falls between the decision boundaries or requires review
IGNOREThe candidate is set aside until restored

Mappings choose the columns, relative weights, missing-value behavior, semantic comparison, and text normalization. The two thresholds establish the three main decision regions.

GET /api/golden/{entity}/duplicates/clusters/{clusterId}/evaluate returns pairwise scores and field-level similarities. Inspect which fields agree and which contradict the proposed identity. The final score is evidence under the configured comparison policy, not a probability or a count of agreeing fields.

Candidates requiring review

REVIEW identifies candidates between the configured decision thresholds or subject to another review condition, such as group size. Inspect their evidence before deciding whether to merge or separate records.

In the initial Aurelia Utilities sample, 214 of 245 groups are MATCH and 31 are REVIEW. Validate threshold changes against known matches and non-matches, as well as the resulting review volume.

Buckets and clusters

Two words appear in the interface and the API, and they are not synonyms:

GroupDefinition
BucketOne index’s contribution: the records that share one key of one mapping
ClusterThe logical group a person decides on, made of one or more buckets

A cluster reports the bucketIds it is made of. The API has a parallel operation family for each. Use the identifier returned by the corresponding list operation.

Continue with Resolve groups and subsets.

Golden 3.0.0 · Published 2026-10-04