On this page
Indexer configuration
Configure how Trazadera Golden creates searchable keys and candidate duplicate groups from dataset columns or reviewed scripts.
An indexer resource controls which records can be found together for search
and duplicate comparison. It creates named mappings from dataset values; the
classifier evaluates the candidates those mappings produce.
Define mappings
These mappings follow sample-customer-indexer from the
Aurelia Utilities sample, with shortened descriptions:
{
"type": "indexer",
"_id": "sample-customer-indexer",
"description": "How customer records are grouped into buckets",
"dataset": "sample-customer",
"defaultKeyOptions": ["AGGRESSIVE"],
"mappings": [
{
"id": "by-tax-id",
"description": "Exact tax identifier. Exact comparison",
"duplicates": true, "matching": "EXACT", "type": "COLUMNS",
"combine": false, "keys": ["taxId"]
},
{
"id": "by-email",
"description": "Exact email, case and accents normalized",
"duplicates": true, "matching": "EXACT", "type": "COLUMNS",
"combine": false, "keys": ["email"]
},
{
"id": "by-phone",
"description": "Exact phone",
"duplicates": true, "matching": "EXACT", "type": "COLUMNS",
"combine": false, "keys": ["phone"]
},
{
"id": "by-name-postcode",
"description": "Fuzzy full name combined with postcode",
"duplicates": true, "matching": "FUZZY", "type": "COLUMNS",
"combine": true, "keys": ["fullName", "postcode"],
"fuzzyMaximumTypos": 2
}
]
}
Use description to explain the matching criterion. Golden displays it in
search and duplicate-review results.
Mapping controls
| Field | Purpose |
|---|---|
id, description | Name and explain the mapping in search and review results |
duplicates | Include the mapping in duplicate candidate generation; when false it remains available for search |
matching | Choose EXACT, PREFIX, SUFFIX, INFIX, FUZZY, FUZZY_LSH, or GEOGRAPHIC behavior |
type | Build keys from COLUMNS or a reviewed SCRIPT |
keys | Dataset columns used by a column mapping |
combine | Require the selected columns together in one key instead of producing independent keys |
keyOptions | Override the default text normalization for this mapping |
fuzzyMaximumTypos | Maximum typographical errors for fuzzy matching; default 2 |
geoPrecision | Geographic precision; default L6 |
geoSteps | Neighboring geographic steps; default 1 |
combine: true makes the selected fields one compound criterion. With it
false, each selected field can create a key independently. Use a compound
mapping to narrow a common field such as a name with another stable signal.
Candidate groups in Aurelia
Over the sample’s 378 records:
| Mapping | Groups found | What it reached |
|---|---|---|
by-name-postcode | 85 | The surname typos and the accent-stripped copies |
by-tax-id | 84 | Records sharing an exact tax identifier |
by-email | 76 | The records where the billing system lost the NIF |
by-phone | 0 | Nothing |
The exact phone mapping produces no candidate groups in the initial sample.
Values such as 611000000 and +34 611 00 00 00 require normalization to
match under an exact comparison. Test normalization in a
transformation before indexing.
The name-and-postcode mapping can group Ana’s billing record despite its missing tax identifier, missing phone, and incomplete email. It also groups the two different customers named Javier Torres Melgar. Review both cases when tuning the indexer; candidate membership does not establish identity.
When defaultKeyOptions is omitted, its default is ["AGGRESSIVE"]. The
example uses that default. Treat fuzzy and geographic
settings as recall-versus-volume controls and tune them with representative
examples, not as universal defaults.
Design for search and deduplication
- Set
duplicates: falseon mappings intended only for lookup. - Give each mapping a business-readable identifier; it appears in review evidence and filters.
- Normalize names and addresses deliberately, but keep durable identifiers as strict as the source contract requires.
- Avoid mappings on low-selectivity fields unless combined with another field.
- Test missed known duplicates and unrelated records that happen to share a common value.
Changing an indexer changes search and candidate behavior. Revalidate the entity and use its supported synchronization workflow before judging the new results.
Mapping defaults
| Property | Default |
|---|---|
duplicates | true |
matching | EXACT |
type | COLUMNS |
combine | false |
fuzzyMaximumTypos | 2 |
geoSteps | 1 |
Set matching and key selection explicitly when maintaining configuration across environments. See the API schemas for the complete mapping shape and geographical precision values.