On this page
Dataset configuration
Define Trazadera Golden record structure, semantic column types, identity, validation, nested data, lookups, and column definitions.
A dataset resource defines record structure. Tables and most source,
processing, destination, matching, and presentation resources refer to one.
Core properties
| Property | Purpose |
|---|---|
_id | Stable resource identifier |
description | Human-readable purpose |
dataType | RECORD, JSON, TEXT, BINARY, or XML |
identityType | DEFAULT, COLUMN, or SCRIPT identity behavior |
identityScript | Identity expression when identityType is SCRIPT |
merger | Optional nested-data merger reference |
lookup | Allows this dataset to be a lookup target |
columns | Column definitions for RECORD data |
provenance | Definitions for the source and reference values in record _source entries |
Define record columns
{
"type": "dataset",
"_id": "customer_dataset",
"description": "Customer record structure",
"dataType": "RECORD",
"identityType": "DEFAULT",
"columns": [
{
"key": "customer_id",
"description": "Source customer identifier",
"type": "TOKEN",
"token": "ID",
"identity": true,
"empty": false
},
{
"key": "email",
"description": "Primary email address",
"type": "TOKEN",
"token": "EMAIL",
"validation": "DEFAULT",
"error": "IGNORE"
},
{
"key": "addresses",
"description": "Known addresses",
"type": "DATASET",
"dataset": "address_dataset",
"array": true
}
]
}
empty: false requires a value. The lookup flag on a column means other
datasets can reference that column as a foreign key; it does not make the field
a duplicate-search index. Configure search and candidate grouping with an
indexer.
Reference another table with FOREIGN_ID
A column whose token is FOREIGN_ID holds an identifier resolved against a
lookup table, and it requires a foreignKey naming the table and column it
resolves against:
{
"key": "province",
"description": "Province",
"type": "TOKEN",
"token": "FOREIGN_ID",
"empty": true,
"validation": "DEFAULT",
"foreignKey": "sample_province.province"
}
A FOREIGN_ID column requires a foreignKey. Without it, dataset validation
fails with “Invalid null or empty foreign key”. Use TEXT when the value
does not require lookup resolution.
A lookup can point at another lookup
In the Aurelia Utilities sample, the chain is three deep:
sample_customer.province → sample_province.province
sample_province.region → sample_region.regionCode
sample_customer.supplyPoint → sample_supply_point.cups
Create lookup dependencies before the datasets that reference them. A dataset
with a FOREIGN_ID column requires its target table to exist. Create and
populate the region catalog before the province catalog, then the customer data.
The size of a lookup table changes how it is presented. A record form embeds a catalog of up to 500 values as a plain select; above that the picker becomes a typeahead that searches. The sample’s supply-point catalog holds 1200 rows deliberately, so that both behaviors are reachable.
GET /api/resources/lookup/{table} reads a page of a lookup table’s values for
an integration that has to render the same choice.
Choose semantic token types
Use the dataset token reference for the full semantic vocabulary, including identity, contact, name, address, geographic, text, and ignored values.
Choose the value’s meaning rather than approximating it as generic text. Token
choice affects validation, formatting, comparison, and presentation. Use the token reference and the published API schema for the supported values.
GET /api/resources/enums lists configured references and roles; it does not
list the static token vocabulary.
Identity
| Identity type | Behavior |
|---|---|
DEFAULT | Keeps the incoming record identifier without deriving one from business fields; repeated loads are not guaranteed to identify the same source row |
COLUMN | Combines columns marked identity; use stable source keys |
SCRIPT | Derives identity with the configured script |
Both COLUMN and SCRIPT require a merger referencing the same dataset.
To create this dependency through individual saves, create the dataset with
DEFAULT, create its merger, then update the dataset with the final identity
and merger reference before loading data. The manual tutorial
shows the complete sequence. The structure example above uses the initial
DEFAULT state; marking a column identity:true alone does not activate column identity.
An unstable identity turns an update into another record. Prefer durable source keys and validate repeated sample loads before using the dataset in production.
Aurelia uses SCRIPT with identityScript: "input['_id']": its source files
carry _id, and the load retains that identifier. _id is a reserved member,
not a business column to declare in columns.
Record shapes and provenance
A TOKEN column holds a simple value whose meaning is defined by token. A
DATASET column refers to a nested dataset through dataset. TOKEN columns can be single-valued or use array: true. DATASET columns
require array: true, including when a record contains only one nested item.
A single nested object with array: false is rejected. See
Records and their representation for the supported JSON
shapes and the difference between nesting and lookup references.
The dataset’s provenance.source and provenance.reference definitions
describe the vocabulary and validation of _src and _src_id inside each
_source entry. They do not create business columns or copy arbitrary source
fields into provenance. See Record metadata.
Validation and error policy
Validation values are NONE, DEFAULT, PARSER, REGEX, and SCRIPT.
When a cleaner rejects a value, it applies the column’s error policy:
| Policy | Outcome |
|---|---|
CLEAR | Remove the offending value |
REPLACE | Use errorReplacement |
FIX | Use a supported fix; remove the value if no fix is available |
IGNORE | Keep the submitted value |
REFUSE | Refuse this record when a cleaner rejects the value |
These policies require a cleaner in the processing path. REFUSE does not
refuse a record merely because a mandatory field is empty.
Test missing, malformed, nested, and multiple values. Confirm both the stored record and its quality information; validation policy need not reject the whole record.