On this page
Automatic steward configuration
Configure ordered Golden stewardship rules that select duplicate candidates and apply controlled actions automatically, and attach one to an entity.
A steward-bucket resource applies configured actions to the duplicate
candidates of one entity without per-group human confirmation. The resource states what to
do; Golden task scheduling controls when it runs.
Resource structure
| Property | Purpose |
|---|---|
_id | Steward resource identifier |
description | Human-readable intent. Display only |
entity | The entity whose duplicates these rules resolve |
coffeeBreak | Clusters resolved in one run before stopping. 0 means no break |
rules | Ordered selection and action rules. At least one is required |
With coffeeBreak: 0, a run has no cluster-count limit. Use a bounded value
while validating the policy or processing a large candidate set.
Rule order
Rules run in order. Each rule processes its selected clusters before the next rule starts. A cluster already processed is not offered to later rules.
Place exclusion rules before merge or delete rules that could select the same clusters. Clusters selected by no rule remain unchanged.
Rule properties
Every rule requires id, description, and action.
| Field | Notes |
|---|---|
id | Written into the audit comment of every cluster the rule touches |
description | Required explanation included verbatim in the audit comment |
action | MERGE, DISCONNECT, SPLIT, DELETE, or IGNORE |
filtering | index and classification. Limits the candidate clusters |
conditions | Extra tests, joined with OR |
sorting | The order clusters are taken in |
crossSeparations | MERGE only. See below |
Use a stable rule id and a description that explains the decision criteria.
Both appear in audit comments for the rule’s actions.
Actions
| Action | Effect |
|---|---|
MERGE | Collapses the cluster into a single golden record |
DISCONNECT | Records pairwise separation and preserves the records |
SPLIT | Removes the records that clash |
DELETE | Removes the cluster’s records |
IGNORE | Marks the cluster ignored; preserves its records |
IGNORE changes candidate state without merging or deleting records. It can
be undone by an administrator, but it still removes work from active review.
Prefer the resource test to inspect selection before applying any action.
Filtering
Use index to restrict a rule to candidates produced by an approved mapping.
For example, you can apply different rules to exact tax-identifier candidates
and approximate name candidates.
Set classification to MATCH when the rule should select only candidates
that meet the configured match threshold.
Without either filter, the rule can select every duplicate cluster in the entity. Review its conditions before combining that scope with a destructive action.
Conditions
Conditions use OR: a cluster is selected if it satisfies any condition. Adding a condition can broaden selection. With no conditions, all clusters returned by the filter are selected.
type | Unit of min and max |
|---|---|
RECORD_COUNT | Number of records in the cluster |
SCORE | Bucket score as a percentage, 0 to 100 — not the classifier’s 0-to-1 threshold scale |
DEVIATION | Present in the schema, but validation rejects this condition in the documented revision; do not use it |
interval says which side is bounded — MINIMUM, MAXIMUM, or INTERVAL —
and both ends are inclusive: a minimum of 90 on the score also takes the
clusters scoring exactly 90.
min and max accept decimal values. SCORE bounds use the reported 0–100
scale. For a single-sided score condition, its bound must be below 100; an
inverted interval is invalid. Test the exact condition before saving it.
Sorting
SIZE_ASC, SIZE_DESC, SCORE_ASC, SCORE_DESC, DEVIATION_ASC,
DEVIATION_DESC, NATURAL. Each criterion can appear only once.
Sorting decides what gets resolved first when the run stops at its break
rather than finishing. SCORE_DESC processes higher-scoring clusters first. A score does not
establish identity on its own.
crossSeparations
MERGE only, and off by default. Off, a cluster holding two records somebody
recorded as different is skipped and counted, never merged. On, this rule
merges it anyway and revokes that separation, using the rule’s own sentence as
the comment.
Enable this override only when policy permits reversing existing separations.
State the reason in the rule’s description, which supplies the action comment.
A worked example
This is sample-customer-steward from the
Aurelia Utilities sample. Three rules, and
the order is part of the logic:
{
"type": "steward-bucket",
"_id": "sample-customer-steward",
"entity": "sample-customer-entity",
"coffeeBreak": 100,
"rules": [
{
"id": "leave-large-buckets-to-a-human",
"description": "FIRST ON PURPOSE. Nine or more records usually means an over-permissive key, not nine copies of one person.",
"action": "IGNORE",
"conditions": [{ "type": "RECORD_COUNT", "interval": "MINIMUM", "min": 9 }],
"sorting": ["SIZE_DESC"]
},
{
"id": "auto-merge-same-tax-id",
"description": "An exact tax identifier the classifier also scored MATCH is merged without asking.",
"action": "MERGE",
"filtering": { "index": "by-tax-id", "classification": "MATCH" },
"sorting": ["SCORE_DESC"]
},
{
"id": "queue-fuzzy-name-matches-for-review",
"description": "Buckets found by the fuzzy name key are never merged automatically.",
"action": "IGNORE",
"filtering": { "index": "by-name-postcode" },
"sorting": ["SCORE_DESC"]
}
]
}
The example applies three restrictions:
- The size guard excludes large groups before the merge rule runs.
- The merge rule requires both the
by-tax-idmapping andMATCHclassification. - The final rule ignores name-and-postcode groups, including the two distinct customers named Javier Torres Melgar, without merging their records.
Rule 1 has no classification filter: the classifier already
forces REVIEW on any bucket of ten records or more, so a size guard filtered
on MATCH could only ever select buckets of exactly nine.
Attach it to an entity
Automatic stewardship requires an entity of type AUTO_DUPLICATES with
automatic mode enabled. A DUPLICATES entity can retain steward and
stewardCron settings, but those settings do not create a schedule for that type.
In the entity editor, preserve its sources, destinations, data views and other
resource references while changing its type to AUTO_DUPLICATES, choosing
sample-customer-steward, and setting the intended steward schedule. Save and
verify the configuration before enabling automatic mode. Run this exercise only
on disposable Aurelia data: automatic decisions can begin once enabled.
curl --fail-with-body --silent --show-error --request PUT \
--header "Authorization: Bearer ${GOLDEN_ADMIN_TOKEN}" \
"${GOLDEN_URL}/api/entities/sample-customer-entity/automatic/true"
The canonical format for a custom entity cron expression is Quartz, for
example 0 0 4 * * ? for 04:00 every day. Golden also accepts five-field input
such as 0 4 * * * and normalizes it to Quartz.
This is the entity scheduling setting, distinct from a job definition’s full
schedule representation.
Find the resulting steward definition in Tasks → Schedulings or with
GET /api/jobs/definitions. Set DEFINITION_ID to that returned identifier;
do not construct it from naming conventions. To run it immediately:
curl --fail-with-body --silent --show-error --request POST \
--header "Authorization: Bearer ${GOLDEN_ADMIN_TOKEN}" \
"${GOLDEN_URL}/api/jobs/definitions/${DEFINITION_ID}/run"
Follow the returned runId and inspect the records. Disable automatic mode
again with the same PUT route ending in /false when the exercise is done.
This stops future automatic work; it does not undo completed decisions or
cancel a run already active.
Bounded runs
With coffeeBreak: 100, a run resolves at most 100 clusters. Inspect the
outcome before starting another run; the candidate count alone does not
establish that the decisions were correct.
A run that decided nothing reports why it selected nothing, rather than reporting success and leaving you to guess. Read that message before changing a rule: often the answer is that an earlier rule consumed the clusters.
Roll out safely
- Test the resource and inspect the selected candidates.
- Confirm selection in the test report before applying rules.
IGNOREis a real candidate-state change, not a dry-run flag. - Start with a small
coffeeBreak. - Review the affected records and the task outcome.
- Expand only after measured results meet the acceptance policy.
- Disable automatic entity behavior and escalate if an action produces an unexpected outcome.