On this page
Automate duplicate decisions
Inspect Golden duplicate clusters through the API, read why a group scored what it did, and automate merge, disconnect, escalation, and undo safely.
Golden groups records into duplicate candidates. API callers with VIEWER can
inspect granted groups. Merge, disconnect, delete, and escalation require
STEWARD or ADMIN; ignore, unignore, and resolving an escalation require ADMIN.
Human stewards should use Steward duplicate candidates in the web application.
Every example on this page runs against the
Aurelia Utilities sample, whose entity is
sample-customer-entity. Install it in its initial state and configure the connection values. Mutating
examples change the state required by later calls; reset the sample between
independent cases.
Before you automate
- Use cluster identifiers, not bucket identifiers, unless you have a reason not to. A cluster is the logical group a person sees; a bucket is one index’s contribution to it. Both families exist and they do not mix.
- Test classification and action policy on representative non-production data.
- Use
VIEWERfor inspection and a separateSTEWARDtoken when an automated workflow is approved to change candidate state. - Preserve a decision comment, stable identifiers, and enough evidence to verify or escalate the result.
The entity must be enabled and its synchronization complete. If a queue appears stale, check the entity state and tasks before changing classifier or indexer configuration.
List the candidates
curl --fail-with-body --silent --show-error \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters?page=0&pageSize=100"
Each cluster reports the index that found it, the score, the classification, and its escalation state:
{
"clusterId": "by-email-damianarribasexamplecom",
"label": "damian.arribas@example.com",
"mappingId": "by-email",
"recordCount": 3,
"bucketIds": ["by-email-damianarribasexamplecom"],
"classification": "MATCH",
"score": 92,
"deviation": 7,
"reasoning": "Made 3 comparisons with score 95 and deviation 6.7: classification is match",
"ignore": false,
"escalated": false,
"resolved": false
}
This is a response excerpt from the initial sample. Its reported group score
is 92; the explanation contains a different comparison summary (95). Do
not extract a score from reasoning or replace the structured score by an
average calculated from the evaluation pairs. Use classification, score,
and mappingId for the corresponding API fields; reasoning is human-readable
text whose wording can change.
For a cross-entity total, GET /api/golden/duplicates/record-counts returns
the number of records carrying at least one possible duplicate, per entity.
Inspect comparison evidence
evaluate compares every pair of records in a cluster exhaustively and returns
the per-column similarities behind the score. Use this evidence to inspect agreements and differences between fields.
curl --fail-with-body --silent --show-error \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters/by-email-anagarciaexamplecom/evaluate"
{
"clusterId": "by-email-anagarciaexamplecom",
"recordIds": ["CRM-0001", "CRM-DUP-0001"],
"pairs": [
{
"recordA": "CRM-0001",
"recordB": "CRM-DUP-0001",
"score": 84,
"similarities": {
"taxId": 1.0,
"email": 1.0,
"fullName": 0.9857142857142858,
"phone": 0.0,
"street": 0.8625,
"postcode": 1.0,
"city": 1.0
}
}
]
}
Read per-field similarities as evidence. Ana’s pair reports score 84, below
the sample’s match threshold of 85. Its phone similarity is zero: the sample
compares differently formatted phone strings exactly. Normalization, selected
columns, weights, and thresholds all affect classification. The score is not
a probability that the two records identify one person.
The same operation on a pair that must not merge
curl --fail-with-body --silent --show-error \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters/by-name-postcode-javiertorresmelgar-37001/evaluate"
{
"score": 56,
"similarities": {
"taxId": 0.0,
"email": 0.9826086956521739,
"fullName": 1.0,
"phone": 0.0,
"street": 0.9692307692307692,
"postcode": 1.0,
"city": 1.0
}
}
Compare the two. This pair agrees more than Ana’s does on fullName
(1.0 against 0.986) and on street (0.969 against 0.863), and it still scores
28 points lower. The tax-identifier comparison contributes strongly to that
difference: it has weight 3.0 out of a total of 10.5, and its similarity is
0.0 rather than 1.0. Inspect the other field comparisons as well.
Test automation against both known matches and known non-matches. In this sample, matching names and addresses do not establish identity; the different tax identifiers are relevant evidence against a merge.
Choose an action
| Action | Path suffix | Use it when |
|---|---|---|
| Merge | /merge | Every record represents the same subject |
| Merge a subset | /merge/records | Only some of them do |
| Preview a subset merge | /merge/records/preview | Before the above, to see the record it would produce |
| Disconnect | /disconnect | The records must no longer be treated as related |
| Disconnect selected | /disconnect/records | Record separations between selected and remaining records |
| Ignore, unignore | /ignore, /unignore | Defer the candidate without deleting its records |
| Escalate, resolve | /escalate, /resolve-escalation | A person with more context should rule |
| Delete | /delete, /delete/records | The records themselves should be removed |
Delete is materially different from disconnect. Use it only when removal of the records is intended.
Merge a cluster
curl --fail-with-body --silent --show-error \
--request POST \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
--header "Content-Type: application/json" \
--data '{"comment":"Matching NIF and a recognizable surname transposition"}' \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters/by-email-anagarciaexamplecom/merge"
The body takes comment and crossSeparations. Leave crossSeparations
false unless the workflow explicitly permits overriding an existing separation.
Merge only part of a cluster
recordIds is required and must name at least two records, every one of
them in the cluster.
curl --fail-with-body --silent --show-error \
--request POST \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
--header "Content-Type: application/json" \
--data '{"recordIds":["CRM-0001","CRM-DUP-0001"],"comment":"The billing row stays out"}' \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters/by-email-anagarciaexamplecom/merge/records"
POST .../merge/records/preview takes the same body and computes the record
the merge would produce without merging. Use it in an automated workflow
that needs to check the outcome before committing to it.
Separations block a later merge
When a group holds two records somebody previously disconnected, the merge is refused:
{
"errors": [
"The records CRM-NEAR-0001 and CRM-NEAR-0002 were recorded as different, so they cannot be merged together. To merge them anyway, cross the separation explicitly and give a reason"
]
}
crossSeparations: true merges anyway and revokes that judgment.
Keep crossSeparations false unless the approved workflow permits an override.
When overriding a separation, provide a comment identifying the rule and the
reason for reversing the earlier decision.
Escalate a case
curl --fail-with-body --silent --show-error \
--request PUT \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
--header "Content-Type: application/json" \
--data '{"comment":"Phones disagree and the billing row has no NIF"}' \
"${GOLDEN_URL}/api/golden/sample-customer-entity/duplicates/clusters/by-email-anagarciaexamplecom/escalate"
The cluster’s escalated, escalatedBy, escalatedAt, and escalationNote
fields are then set, and it appears in the web application’s Escalated
cases queue. PUT .../resolve-escalation clears it and requires ADMIN.
Escalate when the available evidence does not meet the workflow’s decision criteria. The case remains available for an authorized reviewer.
Undo a merge
curl --fail-with-body --silent --show-error \
--request POST \
--header "Authorization: Bearer ${GOLDEN_TOKEN}" \
--header "Content-Type: application/json" \
--data '{"recordId":"<the surviving record>","comment":"Merged on a rule that has since been corrected"}' \
"${GOLDEN_URL}/api/golden/sample-customer-entity/merge/undo"
recordId is required and names the record a merge produced. The absorbed
records come back out of history, restored as they were before the merge plus
any separation recorded against the survivor while it existed; the survivor is
deleted and kept in history. Anything else edited on the survivor since the
merge is not carried over.
Four refusals, each with its own message:
| Condition | Message |
|---|---|
| The entity keeps no history | “the records a merge absorbed were not archived and the merge cannot be undone” |
| The record is not a merge survivor | “is not the result of a merge, so there is no merge to undo” |
| The survivor was merged again | “was itself merged into {other}. Undo that merge first, then this one” |
| A source is gone from history | “is no longer in history, so the merge that produced {record} cannot be undone” |
Undo is not a general rollback. It reverses one merge, on a table that kept history, provided nothing has merged the survivor since. Do not design a recovery procedure that assumes it will always be available.
Verify the outcome
Re-read the cluster and confirm:
- it has the expected state, or no longer appears under that filter;
- the resulting mastered or disconnected records are correct; and
- any task the action started reached a terminal success state.
If the result is unexpected, stop related decisions for the entity and follow the approved recovery or support process. Record the entity, cluster identifier, action, comment, and time.
Selected-record decisions
For a cluster, use POST .../merge/records/preview before applying
POST .../merge/records. Both take recordIds; a merge requires at least two
members of that cluster. POST .../disconnect/records separates the selected
records from the remaining members in both directions, without separating the
remaining members from each other. It preserves existing separations.
DELETE .../delete/records removes the selected records only. Ignore/unignore
are group operations, with PUT .../ignore and PUT .../unignore, and require
ADMIN. See the A/B/C exercise for a
reproducible partial decision and checks on unselected records.