On this page
Pipeline configuration
Apply ordered cleaner, transformer, and script processors to records in a Trazadera Golden pipeline.
A pipeline applies an ordered sequence of processors to records described by
one dataset.
Processor types
| Type | Public purpose |
|---|---|
CLEANER | Apply dataset formatting, validation, error policy, and quality handling |
TRANSFORMER | Apply a named transformation resource |
SCRIPT | Apply reviewed custom record logic |
Define the sequence
{
"type": "pipeline",
"_id": "customer_pipeline",
"description": "Clean customer records",
"dataset": "customer_dataset",
"processors": [
{
"processorType": "CLEANER",
"properties": []
}
]
}
Processors do not have an id field. Their supported fields are
processorType, transformation, script, and properties, with relevance
depending on the selected type.
Order is observable behavior: each processor receives the previous processor’s result. Choose the sequence from the data contract. For example, a script that expects already validated values belongs after the cleaner; a repair that must happen before validation belongs before it.
properties is an array of key and value pairs. Values are text and the
selected processor interprets them. Unknown names can be ignored, so use only
settings documented for that processor and release; leave the array empty when
no property is required.
CLEANER applies the dataset’s formatting, validation, error policy, and
quality handling, and removes fields the dataset does not declare.
TRANSFORMER requires a transformation whose source and target both equal the
pipeline dataset. A transformation that changes datasets, such as the CRM
mapping in Transformation configuration, runs outside
this pipeline stage. The example above needs only a cleaner.
SCRIPT requires the reviewed script text. At least one processor is required.
A cleaner can refuse a record under the dataset’s REFUSE policy. That record
does not continue through the pipeline, but the load can still complete
successfully. Verify accepted and refused counts as well as task status. See
Validation and cleaning.
Verify the pipeline
Test records that cover valid values, malformed values, missing optional and required fields, and nested or array data. Inspect both output and reported quality before using the pipeline in an entity or table flow.