On this page
Aurelia configuration example
Connect the sample indexer, classifier, merger, and steward choices with their roles in the data flow.
How the matching is configured
The indexer: four ways of grouping
| Mapping | Matching | Keys | Groups found |
|---|---|---|---|
by-tax-id | Exact | taxId | 84 |
by-email | Exact, case and accents normalized | email | 76 |
by-name-postcode | Fuzzy, combined, up to 2 typos | fullName + postcode | 85 |
by-phone | Exact | phone | 0 |
The exact phone mapping finds no duplicate groups in the initial sample.
Phone values use different formats, such as 611000000 and
+34 611 00 00 00. Normalize comparable values before relying on an exact
phone mapping in your own configuration.
The classifier: weights and two thresholds
| Setting | Value |
|---|---|
| Match threshold | 0.85, equivalent to a reported score of 85 |
| Non-match threshold | 0.55, equivalent to a reported score of 55 |
| Review above bucket size | 10 records |
| Column | Weight |
|---|---|
taxId | 3.0 |
email | 2.0 |
fullName | 2.0 |
phone | 1.5 |
street | 1.0 |
postcode | 0.5 |
city | 0.5 |
The initial sample has 245 candidate groups: 214 classified MATCH and
31 classified REVIEW.
Ana’s two CRM records score 84. The tax identifier and email agree, while the
name, phone, and street comparisons reduce the score. It falls between the
non-match and match thresholds, so the group is classified REVIEW.
The merger: which record survives
sample-customer-merger sorts on contractDate and keeps the highest. In this
data the most recently contracted record is always the clean 2024 CRM original
rather than the 2022 re-registration or the 2021 billing row.
Inspect the contributing values before adopting a date-based precedence rule. A newer contract date does not generally establish that a record is more accurate.
The steward, installed but not attached
sample-customer-steward ships with the sample and is deliberately not
attached to the entity. The entity is DUPLICATES, not AUTO_DUPLICATES. This keeps the initial candidate groups available for the manual exercises.
The rules run in this order:
leave-large-buckets-to-a-humanignores groups with nine or more records.auto-merge-same-tax-idmerges groups found byby-tax-idand classifiedMATCH.queue-fuzzy-name-matches-for-reviewexcludesby-name-postcodegroups from automatic merging.
The first rule that selects a group acts on it. Place the size guard before the merge rule so that groups meeting the size limit are excluded from merging.
To attach it, see Configure a steward. Setting
steward on a DUPLICATES entity is accepted and persisted but creates no
scheduling; the entity has to be AUTO_DUPLICATES and in automatic mode.
Use the design review to consider a solution beyond the sample.