On this page

A steward-bucket resource applies configured actions to the duplicate candidates of one entity without per-group human confirmation. The resource states what to do; Golden task scheduling controls when it runs.

Resource structure

PropertyPurpose
_idSteward resource identifier
descriptionHuman-readable intent. Display only
entityThe entity whose duplicates these rules resolve
coffeeBreakClusters resolved in one run before stopping. 0 means no break
rulesOrdered selection and action rules. At least one is required

With coffeeBreak: 0, a run has no cluster-count limit. Use a bounded value while validating the policy or processing a large candidate set.

Rule order

Rules run in order. Each rule processes its selected clusters before the next rule starts. A cluster already processed is not offered to later rules.

Place exclusion rules before merge or delete rules that could select the same clusters. Clusters selected by no rule remain unchanged.

Rule properties

Every rule requires id, description, and action.

FieldNotes
idWritten into the audit comment of every cluster the rule touches
descriptionRequired explanation included verbatim in the audit comment
actionMERGE, DISCONNECT, SPLIT, DELETE, or IGNORE
filteringindex and classification. Limits the candidate clusters
conditionsExtra tests, joined with OR
sortingThe order clusters are taken in
crossSeparationsMERGE only. See below

Use a stable rule id and a description that explains the decision criteria. Both appear in audit comments for the rule’s actions.

Actions

ActionEffect
MERGECollapses the cluster into a single golden record
DISCONNECTRecords pairwise separation and preserves the records
SPLITRemoves the records that clash
DELETERemoves the cluster’s records
IGNOREMarks the cluster ignored; preserves its records

IGNORE changes candidate state without merging or deleting records. It can be undone by an administrator, but it still removes work from active review. Prefer the resource test to inspect selection before applying any action.

Filtering

Use index to restrict a rule to candidates produced by an approved mapping. For example, you can apply different rules to exact tax-identifier candidates and approximate name candidates.

Set classification to MATCH when the rule should select only candidates that meet the configured match threshold.

Without either filter, the rule can select every duplicate cluster in the entity. Review its conditions before combining that scope with a destructive action.

Conditions

Conditions use OR: a cluster is selected if it satisfies any condition. Adding a condition can broaden selection. With no conditions, all clusters returned by the filter are selected.

typeUnit of min and max
RECORD_COUNTNumber of records in the cluster
SCOREBucket score as a percentage, 0 to 100 — not the classifier’s 0-to-1 threshold scale
DEVIATIONPresent in the schema, but validation rejects this condition in the documented revision; do not use it

interval says which side is bounded — MINIMUM, MAXIMUM, or INTERVAL — and both ends are inclusive: a minimum of 90 on the score also takes the clusters scoring exactly 90.

min and max accept decimal values. SCORE bounds use the reported 0–100 scale. For a single-sided score condition, its bound must be below 100; an inverted interval is invalid. Test the exact condition before saving it.

Sorting

SIZE_ASC, SIZE_DESC, SCORE_ASC, SCORE_DESC, DEVIATION_ASC, DEVIATION_DESC, NATURAL. Each criterion can appear only once.

Sorting decides what gets resolved first when the run stops at its break rather than finishing. SCORE_DESC processes higher-scoring clusters first. A score does not establish identity on its own.

crossSeparations

MERGE only, and off by default. Off, a cluster holding two records somebody recorded as different is skipped and counted, never merged. On, this rule merges it anyway and revokes that separation, using the rule’s own sentence as the comment.

Enable this override only when policy permits reversing existing separations. State the reason in the rule’s description, which supplies the action comment.

A worked example

This is sample-customer-steward from the Aurelia Utilities sample. Three rules, and the order is part of the logic:

{
  "type": "steward-bucket",
  "_id": "sample-customer-steward",
  "entity": "sample-customer-entity",
  "coffeeBreak": 100,
  "rules": [
    {
      "id": "leave-large-buckets-to-a-human",
      "description": "FIRST ON PURPOSE. Nine or more records usually means an over-permissive key, not nine copies of one person.",
      "action": "IGNORE",
      "conditions": [{ "type": "RECORD_COUNT", "interval": "MINIMUM", "min": 9 }],
      "sorting": ["SIZE_DESC"]
    },
    {
      "id": "auto-merge-same-tax-id",
      "description": "An exact tax identifier the classifier also scored MATCH is merged without asking.",
      "action": "MERGE",
      "filtering": { "index": "by-tax-id", "classification": "MATCH" },
      "sorting": ["SCORE_DESC"]
    },
    {
      "id": "queue-fuzzy-name-matches-for-review",
      "description": "Buckets found by the fuzzy name key are never merged automatically.",
      "action": "IGNORE",
      "filtering": { "index": "by-name-postcode" },
      "sorting": ["SCORE_DESC"]
    }
  ]
}

The example applies three restrictions:

  1. The size guard excludes large groups before the merge rule runs.
  2. The merge rule requires both the by-tax-id mapping and MATCH classification.
  3. The final rule ignores name-and-postcode groups, including the two distinct customers named Javier Torres Melgar, without merging their records.

Rule 1 has no classification filter: the classifier already forces REVIEW on any bucket of ten records or more, so a size guard filtered on MATCH could only ever select buckets of exactly nine.

Attach it to an entity

Automatic stewardship requires an entity of type AUTO_DUPLICATES with automatic mode enabled. A DUPLICATES entity can retain steward and stewardCron settings, but those settings do not create a schedule for that type.

In the entity editor, preserve its sources, destinations, data views and other resource references while changing its type to AUTO_DUPLICATES, choosing sample-customer-steward, and setting the intended steward schedule. Save and verify the configuration before enabling automatic mode. Run this exercise only on disposable Aurelia data: automatic decisions can begin once enabled.

curl --fail-with-body --silent --show-error --request PUT \
  --header "Authorization: Bearer ${GOLDEN_ADMIN_TOKEN}" \
  "${GOLDEN_URL}/api/entities/sample-customer-entity/automatic/true"

The canonical format for a custom entity cron expression is Quartz, for example 0 0 4 * * ? for 04:00 every day. Golden also accepts five-field input such as 0 4 * * * and normalizes it to Quartz. This is the entity scheduling setting, distinct from a job definition’s full schedule representation.

Find the resulting steward definition in Tasks → Schedulings or with GET /api/jobs/definitions. Set DEFINITION_ID to that returned identifier; do not construct it from naming conventions. To run it immediately:

curl --fail-with-body --silent --show-error --request POST \
  --header "Authorization: Bearer ${GOLDEN_ADMIN_TOKEN}" \
  "${GOLDEN_URL}/api/jobs/definitions/${DEFINITION_ID}/run"

Follow the returned runId and inspect the records. Disable automatic mode again with the same PUT route ending in /false when the exercise is done. This stops future automatic work; it does not undo completed decisions or cancel a run already active.

Bounded runs

With coffeeBreak: 100, a run resolves at most 100 clusters. Inspect the outcome before starting another run; the candidate count alone does not establish that the decisions were correct.

A run that decided nothing reports why it selected nothing, rather than reporting success and leaving you to guess. Read that message before changing a rule: often the answer is that an earlier rule consumed the clusters.

Roll out safely

  1. Test the resource and inspect the selected candidates.
  2. Confirm selection in the test report before applying rules. IGNORE is a real candidate-state change, not a dry-run flag.
  3. Start with a small coffeeBreak.
  4. Review the affected records and the task outcome.
  5. Expand only after measured results meet the acceptance policy.
  6. Disable automatic entity behavior and escalate if an action produces an unexpected outcome.
Golden 3.0.0 · Published 2026-10-04