On this page

A pipeline applies an ordered sequence of processors to records described by one dataset.

Processor types

TypePublic purpose
CLEANERApply dataset formatting, validation, error policy, and quality handling
TRANSFORMERApply a named transformation resource
SCRIPTApply reviewed custom record logic

Define the sequence

{
  "type": "pipeline",
  "_id": "customer_pipeline",
  "description": "Clean customer records",
  "dataset": "customer_dataset",
  "processors": [
    {
      "processorType": "CLEANER",
      "properties": []
    }
  ]
}

Processors do not have an id field. Their supported fields are processorType, transformation, script, and properties, with relevance depending on the selected type.

Order is observable behavior: each processor receives the previous processor’s result. Choose the sequence from the data contract. For example, a script that expects already validated values belongs after the cleaner; a repair that must happen before validation belongs before it.

properties is an array of key and value pairs. Values are text and the selected processor interprets them. Unknown names can be ignored, so use only settings documented for that processor and release; leave the array empty when no property is required.

CLEANER applies the dataset’s formatting, validation, error policy, and quality handling, and removes fields the dataset does not declare. TRANSFORMER requires a transformation whose source and target both equal the pipeline dataset. A transformation that changes datasets, such as the CRM mapping in Transformation configuration, runs outside this pipeline stage. The example above needs only a cleaner.

SCRIPT requires the reviewed script text. At least one processor is required.

A cleaner can refuse a record under the dataset’s REFUSE policy. That record does not continue through the pipeline, but the load can still complete successfully. Verify accepted and refused counts as well as task status. See Validation and cleaning.

Verify the pipeline

Test records that cover valid values, malformed values, missing optional and required fields, and nested or array data. Inspect both output and reported quality before using the pipeline in an entity or table flow.

Golden 3.0.0 · Published 2026-10-04