Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Pipeline YAML Structure

A Clinker pipeline is a single YAML file with three top-level sections: pipeline (metadata), nodes (the processing graph), and optionally error_handling.

Top-level shape

pipeline:
  name: my_pipeline            # Required — pipeline identifier
  memory:                      # Optional — see ops/memory.md
    limit: "256M"              # Optional (K/M/G suffixes), default 512M
    backpressure: pause        # Optional, default `pause`
  vars:                        # Optional typed static configuration
    threshold: { type: int, default: 500 }
    label: { type: string, default: "Monthly Report" }
  date_formats: ["%Y-%m-%d"]   # Optional — custom date parsing formats
  rules_path: "./rules/"       # Optional — CXL module search path
  concurrency:                 # Optional
    threads: 4
    chunk_size: 1000
  metrics:                     # Optional
    spool_dir: "./metrics/"

nodes:                         # Required — flat list of pipeline nodes
  - type: source
    name: raw_data
    config:
      name: raw_data
      type: csv
      path: "./data/input.csv"
      schema:
        - { name: id, type: int }
        - { name: value, type: string }

  - type: transform
    name: clean
    input: raw_data
    config:
      cxl: |
        emit id = id
        emit value = value.trim()

  - type: sink
    name: result
    input: clean
    config:
      name: result
      type: csv
      path: "./output/result.csv"

error_handling:                # Optional
  strategy: fail_fast

Pipeline metadata

The pipeline: block carries global settings that apply to the entire run.

FieldRequiredDescription
nameYesPipeline identifier. Used in logs and metrics.
memoryNoMemory-arbitrator tuning. Nested fields: limit (RSS budget, K/M/G suffixes, default 512M) and backpressure (spill/pause/both, default pause). See Memory Tuning.
varsNoTyped static configuration accessible in CXL via $vars.*. Each key declares type and an optional default; see Scoped Variables.
date_formatsNoList of strftime-style patterns for date parsing.
rules_pathNoDirectory for CXL use module resolution.
concurrencyNothreads and chunk_size for parallel chunk processing.
metricsNospool_dir for per-run JSON metric files.
date_localeNoUnsupported. Any explicit value is rejected with E119. Use explicit date_formats entries.
log_rulesNoUnsupported. Any explicit value is rejected with E124; configure runtime logging outside pipeline YAML.
include_provenanceNoUnsupported. Any explicit value is rejected with E125. Use write_meta: true on each intended Output.

These three names are admitted only far enough to produce precise, spanned diagnostics. Empty strings/maps and false are still explicit values and are rejected before execution; omission is the only accepted form.

Reserved metadata contract

The current status and locked owner are explicit:

FieldCurrent statusLocked target and owner
date_localeRejected (E119)Remove it and express supported parsing with date_formats:.
log_rulesRejected (E124)Remove it; runtime telemetry policy is not authored in pipeline YAML.
include_provenanceRejected (E125)Remove it and set write_meta: true on each Output that needs a provenance sidecar.

For provenance sidecars that work today, set write_meta: true on an Output node. D-24 keeps that spelling. See Approved exceptions and rejected placeholders.

The nodes list

Every pipeline has a flat nodes: list. Each entry is a node with a type: discriminator that determines its kind:

TypeRole
sourceReads data from a file
transformApplies CXL expressions to each record
aggregateGroups and summarizes records
routeSplits records into named branches by condition
mergeConcatenates multiple upstream branches that share a schema
combineJoins records across N inputs with where: predicates
reshapeMutates or synthesizes records within correlation groups
cullRemoves whole correlation groups to a side-output port
envelopeFrames body records with optional document header and trailer streams
outputWrites records to a file
compositionImports a reusable transform fragment

Node naming

Every node must have a name: field. Names must be unique within the pipeline and must not contain dots – the dot addresses something other than a node: a route branch (split.high, see below) and a node inside a composition call site (enrich.ref). The rule covers every node kind, including sources, outputs, and composition call sites, and it covers nodes declared inside a .comp.yaml body. Names are used for wiring, logging, and diagnostics.

A dotted name is refused at plan time with E010, which names the node and the name to use instead:

node name "enrich.ref" is invalid: '.' is reserved for branch references and
composition call-site paths; rename the node to "enrich_ref" (use underscores
or hyphens) and update every reference to it

Wiring by node kind

Input fields live at the node’s top level, alongside name: and type:. Their shape is specific to the node kind:

Node kindInput shape
sourceNo input field
transform, aggregate, route, reshape, cull, outputOne upstream reference in input:
mergeOrdered list of upstream references in inputs:
combineQualifier-to-upstream map in singular input:
envelopeRequired body: plus optional header: and trailer: upstream references
compositionRequired primary input: plus an inputs: map binding every required composition port; the map is authoritative for DAG wiring

Single upstream – used by ordinary one-input consumers:

- type: transform
  name: clean
  input: raw_data       # References the source node named "raw_data"
  config: ...

Port syntax – for consuming a specific branch from a route node, use node.port:

- type: sink
  name: high_value_out
  input: split.high     # Consumes the "high" branch of route node "split"
  config: ...

Multiple upstreams – merge nodes use inputs: (plural) instead of input::

- type: merge
  name: combined
  inputs:
    - east_processed
    - west_processed
  config: {}

Qualified inputs – Combine uses singular input: with a map whose keys become CXL qualifiers:

- type: combine
  name: enriched
  input:
    orders: clean_orders
    products: product_catalog
  config:
    where: "orders.product_id == products.product_id"
    match: first
    on_miss: null_fields
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.name
    propagate_ck: driver

Envelope and Composition have additional port semantics. See Envelope Nodes and Compositions before wiring those node kinds.

Source nodes have no input field. They are entry points – adding an input: field to a source is a parse error.

Using the wrong field or value shape for a node kind is caught at parse time by strict deserialization.

Optional fields on all nodes

Every node type supports these optional fields:

  • description: – human-readable text for documentation. Ignored by the engine.
  • _notes: – arbitrary metadata (JSON object). Ignored by the engine and available to external tooling.
- type: transform
  name: enrich
  description: "Add customer tier based on lifetime value"
  _notes:
    color: "#4a9eff"
    position: { x: 300, y: 200 }
  input: customers
  config:
    cxl: |
      emit tier = if lifetime_value >= 10000 then "gold" else "standard"

Strict parsing

All config structs use deny_unknown_fields. If you misspell a field name – for example, writing inputt: instead of input: or stratgy: instead of strategy: – the YAML parser rejects it immediately with a diagnostic pointing to the typo. This catches configuration errors before any data processing begins.

Environment variable: CLINKER_ENV

The CLINKER_ENV environment variable can be used for conditional logic outside of pipelines (e.g., selecting channel directories or controlling CLI behavior). It is not directly referenced within pipeline YAML but is available to the channel and workspace systems.