Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Clinker

Clinker is a pure-Rust, bounded-memory batch DAG executor for CSV, JSON, XML, and fixed-width data. It reads finite inputs, drives them through a directed acyclic graph of transformation nodes one record at a time, and exits when the inputs are drained. It ships as a single static binary with no interpreter, no runtime, and no install dependencies.

Pipelines are declared in YAML. Data transformation logic is written in CXL, a custom expression language purpose-built for ETL. Together they replace legacy tools like Informatica, SSIS, Talend, and NiFi with something deterministic, lightweight, and easy to reason about.

What Clinker is, plainly

A finite batch executor with per-record streaming evaluation, not a long-running stream processor. A pipeline run is a job: Sources read until EOF, the DAG drains, the process exits. Within a run, stateless operators (Transform, Route, most Combine probe-side work, Sink) evaluate records one at a time without accumulating per-record state. Every stage is charged against the configured RSS budget. Fused Source → Transform → Sink paths run streaming with no per-stage materialization; non-fused boundaries (Route fan-out, Merge fan-in, Composition bodies, diamond DAGs) materialize records into per-stage buffers that charge against the same envelope. The engine spills buffers to disk at 80% of the limit and fails fast with E310 MemoryBudgetExceeded at the hard limit, naming the offending producer. Blocking operators (Aggregate, sort, grace-hash Combine) accumulate state inside that same budget and spill to disk when soft and hard memory thresholds trip, rather than OOM-killing the process.

If you have used Flink, Kafka Streams, or Beam in unbounded mode: Clinker is not that. There are no watermarks against wall-clock time, no infinite-source semantics, no exactly-once delivery across restarts. The closest prior art is Pentaho Kettle / Apache Hop, Embulk, Singer, Benthos in batch mode, and Vector running file-to-file – finite ETL jobs with per-record evaluation and a hard memory ceiling.

Three pillars of what Clinker is:

  1. Finite inputs. Files (CSV, JSON, XML, fixed-width, EDIFACT, X12, HL7 v2, SWIFT MT) are the canonical shape. Finite-cursor network sources (paginated REST APIs with hard page/record caps) fit the same model – they exhaust their cursor and EOF. Unbounded sources (Kafka topics, Kinesis streams, Server-Sent Events, webhooks, tail -f-style file followers) are out of scope and will remain so.
  2. Finite jobs. A pipeline run begins when you invoke clinker run, drains the DAG, and exits with a status code. No long-running daemon, no service surface, no infinite event loop.
  3. Single process. One clinker binary invocation is one operating- system process. Parallelism happens inside the process via threads (std::thread, Rayon). Clinker does not spawn worker processes, does not coordinate a cluster, and does not shuffle data between machines. Scale by giving the host more cores, more RAM, and more disk – the DuckDB / Polars / Kettle model. If a single host genuinely can’t fit the work, partition the input by file or by key and run multiple clinker invocations from a shell script; that’s a five-line bash script, not an architectural addition.

Why Clinker?

Single binary, zero dependencies. Download it, run it. No JVM, no Python, no package manager. Runs on Linux, macOS, and Windows out of the box — CI builds and tests on all three, and the spill, staging, and RSS-sampling layers have platform-specific paths so behavior is consistent across them.

Good neighbor on busy servers. Clinker enforces a strict memory ceiling (default 512 MB) so it can run alongside JVM applications, databases, and other services without competing for RAM. Aggregation spills to disk when memory pressure rises.

Reproducible output. Given the same input and pipeline, Clinker produces byte-identical output across runs. No nondeterminism from thread scheduling, hash randomization, or floating-point reordering.

Operability-first design. Per-stage metrics, dead-letter queues for error records, explain plans for understanding execution, and structured exit codes for scripting. Built for production from day one.

Two binaries:

BinaryPurpose
clinkerRun pipelines against real data
cxlCheck, evaluate, and format CXL expressions interactively

A taste of Clinker

Here is a complete pipeline that reads a customer CSV, filters to active customers, classifies them into tiers, and writes the result:

pipeline:
  name: customer_etl

nodes:
  - type: source
    name: customers
    config:
      name: customers
      type: csv
      path: "./data/customers.csv"
      schema:
        - { name: customer_id, type: int }
        - { name: first_name, type: string }
        - { name: last_name, type: string }
        - { name: status, type: string }
        - { name: lifetime_value, type: float }

  - type: transform
    name: enrich
    input: customers
    config:
      cxl: |
        filter status == "active"
        emit customer_id = customer_id
        emit full_name = first_name.concat(" ", last_name)
        emit tier = if lifetime_value >= 10000 then "gold" else "standard"

  - type: sink
    name: result
    input: enrich
    config:
      name: result
      type: csv
      path: "./output/enriched_customers.csv"

Run it:

clinker run customer_etl.yaml

That is the entire workflow. No project scaffolding, no configuration files, no compile step. One YAML file, one command.

Next steps

Non-Goals

This page lists what Clinker is deliberately not. These are architectural commitments — design surfaces Clinker will not grow into, not just features that haven’t been built yet.

If you arrived here because you were considering Clinker for one of the scenarios below, the answer is “a different tool is the right fit.” Each non-goal is paired with the kind of tool that is the right fit.

Not an unbounded stream processor

Clinker reads sources that have an end. A pipeline run is a finite job: Sources read until EOF, the DAG drains, the process exits.

Out of scope:

  • Kafka topics, Kinesis streams, Pub/Sub subscriptions (long-running consumers without a natural end).
  • Server-Sent Events, WebSocket subscriptions, webhooks-as-input.
  • tail -f-style file followers.
  • Watermarking against wall-clock time.
  • Exactly-once delivery across process restarts.
  • Stateful infinite-stream windowing (tumbling / sliding / session windows over event time without a finite boundary).

Right fit instead: Apache Flink, Kafka Streams, Apache Beam in unbounded mode, Vector with streaming sources, Benthos with streaming inputs, Apache NiFi.

Not a multi-process or distributed engine

One clinker run invocation is one operating-system process. Clinker does not spawn worker processes, does not coordinate a cluster, and does not shuffle data between machines.

Out of scope:

  • Worker-process pools on a single machine.
  • Multi-machine sharded execution.
  • Network shuffle between executors.
  • Cluster managers (Kubernetes operators, YARN, Mesos integrations).
  • Distributed memory accounting.
  • Partial-failure recovery across worker boundaries.

Right fit instead: Apache Spark, Trino / Presto, Apache Flink in cluster mode, Apache Beam on Dataflow, Hadoop MapReduce.

Scaling Clinker: give the host more cores, more RAM, more disk — the DuckDB / Polars / Kettle / Hop model. If a single host genuinely can’t fit the work, partition the input by file or by key and run multiple clinker invocations from a shell script. That’s a five-line script, not an architectural addition.

Not a long-running service

Clinker is a CLI binary, not a server. There is no daemon mode, no HTTP control plane, no JDBC/ODBC listener, no UI server, no scheduled job runner inside Clinker itself.

Out of scope:

  • HTTP API exposing pipeline execution.
  • Built-in cron / scheduler / orchestrator.
  • Persistent connection pool living across pipeline runs.
  • A long-lived process accepting new pipeline submissions over a socket.

Right fit instead:

  • For scheduling: cron, systemd timers, Airflow, Dagster, Prefect, Temporal.
  • For HTTP-fronted ETL: any of the above orchestrators wrapping clinker run invocations.
  • For interactive queries against finite data: DuckDB, Polars, or any embedded query engine.

Orchestration by Temporal (and the others above) is supported via a shell-out contract — the orchestrator runs clinker run as a child process and reads its exit code, logs, and metrics. Clinker embeds no Temporal client or worker; that coupling is a decided non-goal (issue #622). See Running Under a Workflow Orchestrator for the exit-code, cancellation, and output-atomicity guarantees that contract depends on.

Not an OLAP / SQL query engine

Clinker is a per-record expression engine with explicit nodes: in a DAG. It does not parse SQL, does not optimize joins via cost-based optimization across the whole pipeline, and does not present a relational table model.

Out of scope:

  • SQL parsing (the CXL language is the surface; no SELECT ... FROM is accepted).
  • Cost-based join reordering across more than the local Combine node.
  • Materialized views or query caching.
  • Interactive query latencies under a second.
  • ANSI-SQL semantics for NULL, type coercion, or aggregate behavior.

Right fit instead: DuckDB, ClickHouse, DataFusion, Trino, Postgres, or any RDBMS. If you want SQL-driven transformation over files, DuckDB is the closest single-binary alternative to Clinker for the cases where SQL is the right surface.

Not a connector marketplace

Clinker ships with a deliberately small set of source and sink types: CSV, JSON, XML, fixed-width, EDIFACT, X12, HL7 v2, and SWIFT MT files, plus a finite-cursor REST source. Writing to a network endpoint is not supported: a REST Output sink (issue #224) and finite-cursor SQL sources and sinks (#225, #226) are tracked but unbuilt. There is no plugin registry, no third-party connector store, no SaaS-API catalog.

Out of scope:

  • Hundreds of pre-built SaaS integrations (Salesforce, HubSpot, Stripe, etc.).
  • A central registry of community-maintained connectors.
  • Schema discovery against arbitrary external APIs.
  • Change-data-capture (CDC) sources.

Right fit instead: Airbyte, Fivetran, Stitch, Singer with its tap ecosystem, dlt (data load tool).

Not a streaming-CDC engine

Clinker treats each pipeline run as a fresh, finite pass over the input. It does not maintain a persistent log of source changes, does not replicate row-level changes from a database, and does not produce an append-only stream of inserts / updates / deletes.

Out of scope:

  • Postgres logical replication subscriptions.
  • MySQL binlog tailing.
  • Debezium-style CDC stream production.
  • Maintaining a target database in continuous sync with a source.

Right fit instead: Debezium, Maxwell, AWS DMS, Striim, Estuary Flow, or vendor-native CDC like Snowflake Streams.

What Clinker is

For the positive framing, see the Introduction and Key Concepts. The short version:

  • A pure-Rust, single-binary, bounded-memory batch DAG executor for finite file and finite-cursor inputs.
  • Per-record evaluation through a directed acyclic graph of Source, Transform, Aggregate, Route, Merge, Combine, Output, and Composition nodes.
  • Pipelines declared in YAML, transformation logic written in CXL (a custom per-record expression language).
  • One process, finite job, EOF-then-exit. Disk spill under memory pressure rather than OOM.

Installation

Clinker is a single static binary with no runtime dependencies. Download it, put it on your PATH, and you are ready to go.

Binaries

Clinker ships two binaries:

  • clinker – the pipeline executor. This is the main tool you use to validate and run pipelines against data.
  • cxl – the CXL expression checker, evaluator, and formatter. Use it during development to test expressions interactively, check types, and format CXL blocks.

Verify installation

After placing the binaries on your PATH, confirm they work:

clinker --version
clinker 0.1.0
cxl --version
cxl 0.1.0

Both commands should print a version string and exit. If you see command not found, check that the directory containing the binaries is in your PATH.

Building from source

Clinker requires Rust 1.91+ (edition 2024). If you have a Rust toolchain installed, build and install both binaries directly from the repository:

# Clone the repository
git clone https://github.com/rustpunk/clinker.git
cd clinker

# Install the pipeline executor
cargo install --path crates/clinker

# Install the CXL expression tool
cargo install --path crates/cxl-cli

This compiles release-optimized binaries and places them in ~/.cargo/bin/, which is typically already on your PATH.

To verify the build:

cargo test --workspace

This runs the full test suite (approximately 1100 tests) and confirms everything is working correctly on your system.

Rust toolchain

The repository includes a rust-toolchain.toml that pins the exact Rust version. If you use rustup, it will automatically download the correct toolchain when you build.

RequirementValue
Rust edition2024
Minimum version1.91
C dependenciesNone

A working C compiler is not one of the requirements: nothing in the build graph runs one. TLS for the rest source and the OTLP exporter goes through rustls with the graviola provider, which ships as Rust and inline assembly rather than the C and per-architecture assembly that other providers build, and content hashing uses blake3’s pure-Rust SIMD paths. CI builds the workspace, including its test and benchmark targets, with every C-compiler environment variable pointed at a program that fails, so a dependency that starts needing one is caught rather than noticed by whoever first builds without a compiler installed. That job also builds a crate that deliberately compiles C and requires it to fail, so the check is known to still work and not merely to be passing.

Graviola supports x86_64 and aarch64. Those are the architectures Clinker is built and tested on; another one needs a different rustls provider, and the ones available today build C.

Optional capabilities

Three parts of Clinker are Cargo features of the clinker crate. All three are on by default — a downloaded binary, or one built with plain cargo install --path crates/clinker, has every one of them, and nothing in this documentation assumes otherwise.

FeatureWhat it adds
restThe rest source transport (transport: rest on a source node).
otlpOTLP/HTTP export of logs, metrics and spans to a collector.
lineageOpenLineage emission: --lineage and --lineage-events.

They exist for deployments that want a smaller binary or a narrower dependency graph — a build with no rest and no otlp links no HTTP client and no TLS stack at all:

# File sources and lineage only: no network transport is compiled in.
cargo install --path crates/clinker --no-default-features --features lineage

A build without a capability still parses every construct the full one does. What it will not do is run one silently: a pipeline that declares a rest source, a clinker.toml that sets observability.otlp.endpoint, or a --lineage flag is refused at validation — before any source is opened or any output written — with a diagnostic that names what was asked for and says the capability is not in this binary.

Turning otlp off also stops telemetry being recorded: the fixed arena is reserved for an exporter to drain, so with no exporter there is nothing to reserve it for, and the --machine ndjson-v1 terminal then carries no observability field at all.

Verify a Release

Verify both the archive checksum and its build attestation before running a downloaded Clinker binary. A checksum detects changed bytes; the attestation also binds those bytes to this repository and its release workflow.

Download one platform archive

Replace vX.Y.Z and the archive name with the release you intend to install:

gh release download vX.Y.Z \
  --repo rustpunk/clinker \
  --pattern 'clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz*'

Each native archive has a sibling .sha256 file. The release also contains a SHA256SUMS inventory covering all supported archives.

Check the SHA-256 digest

On Linux:

sha256sum --check clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz.sha256

On macOS:

shasum -a 256 -c clinker-vX.Y.Z-aarch64-apple-darwin.tar.gz.sha256

On Windows PowerShell:

$archive = "clinker-vX.Y.Z-x86_64-pc-windows-msvc.zip"
$expected = (Get-Content "$archive.sha256").Split()[0].ToLowerInvariant()
$actual = (Get-FileHash -Algorithm SHA256 $archive).Hash.ToLowerInvariant()
if ($actual -ne $expected) { throw "Clinker archive checksum mismatch" }

Stop if the checksum command fails. Do not extract or run the archive.

Verify build provenance

With a current GitHub CLI and network access:

gh attestation verify \
  clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz \
  --repo rustpunk/clinker

The verification must identify rustpunk/clinker as the source repository and the release archive itself as the attested subject. An attestation proves where and how the bytes were built; it is not a claim that the program is free of security defects.

If either checksum or provenance verification fails, keep the archive quarantined and report the release tag, archive name, and failing command.

Your First Pipeline

This walkthrough builds a pipeline from scratch, runs it, and explores the tools Clinker provides for validating and understanding pipelines before they touch real data.

1. Create sample data

Save the following as employees.csv:

id,name,department,salary
1,Alice Chen,Engineering,95000
2,Bob Martinez,Marketing,62000
3,Carol Johnson,Engineering,88000
4,Dave Williams,Sales,71000

2. Write the pipeline

Save the following as my_first_pipeline.yaml:

pipeline:
  name: salary_report

nodes:
  - type: source
    name: employees
    config:
      name: employees
      type: csv
      path: "./employees.csv"
      schema:
        - { name: id, type: int }
        - { name: name, type: string }
        - { name: department, type: string }
        - { name: salary, type: int }

  - type: transform
    name: classify
    input: employees
    config:
      cxl: |
        emit id = id
        emit name = name
        emit department = department
        emit salary = salary
        emit level = if salary >= 90000 then "senior" else "junior"

  - type: sink
    name: report
    input: classify
    config:
      name: report
      type: csv
      path: "./salary_report.csv"

This pipeline has three nodes:

  1. employees (source) – reads the CSV file and declares the schema.
  2. classify (transform) – passes all fields through and adds a level field based on salary.
  3. report (sink) – writes the result to a new CSV file.

The input: field on each consumer node wires the DAG together. Data flows from employees through classify to report.

3. Validate before running

Before processing any data, check that the pipeline is well-formed:

clinker run my_first_pipeline.yaml --dry-run

Dry-run parses the YAML, resolves the DAG, and type-checks all CXL expressions against the declared schemas. If there are errors – a typo in a field name, a type mismatch, a missing input: reference – Clinker reports them with source-location diagnostics and stops. No data is read.

4. Preview records

To see what the output will look like without writing files, preview a few records:

clinker run my_first_pipeline.yaml --dry-run -n 2

This reads at most 2 records from this source, runs them through the pipeline, and writes the preview to stdout without opening salary_report.csv. For a pipeline with several Sources, the limit applies separately to each Source; output counts can differ after filters, joins, or aggregates. Use --dry-run-output preview.csv to select an explicit preview destination.

At revision 3b343a4e, this example’s bounded preview can fail with an internal node-buffer cleanup error. Its ordinary run produces the output below. See the preview limitation.

5. Understand the execution plan

To see how Clinker will execute the pipeline:

clinker run my_first_pipeline.yaml --explain

The explain plan shows the DAG topology, the order nodes will execute, per-node parallelism strategy, and schema propagation through the pipeline. This is valuable for understanding complex pipelines with routes, merges, and aggregations.

6. Run it

clinker run my_first_pipeline.yaml

Clinker reads employees.csv, applies the transform, and writes salary_report.csv. The output:

id,name,department,salary,level
1,Alice Chen,Engineering,95000,senior
2,Bob Martinez,Marketing,62000,junior
3,Carol Johnson,Engineering,88000,junior
4,Dave Williams,Sales,71000,junior

Alice’s salary of 95,000 meets the threshold, so she is classified as senior. Everyone else is junior.

What just happened

The pipeline executed as a streaming process:

  1. The source node read employees.csv one record at a time.
  2. Each record flowed through the classify transform, which evaluated the CXL block to produce the output fields.
  3. The Sink node wrote each transformed record to salary_report.csv.

At no point was the entire dataset loaded into memory. This is how Clinker processes files of any size under its memory ceiling.

Next steps

Key Concepts

This page covers the mental model behind Clinker pipelines. If you have experience with other ETL tools, most of this will feel familiar – but pay attention to where Clinker diverges, especially around Clinker Expression Language (CXL), per-record evaluation, and the memory budget.

Batch jobs, not unbounded streams

A Clinker run is a finite batch job. Source nodes read their files until EOF, the DAG drains, and the process exits. There are no watermarks against wall-clock time, no infinite-source semantics, no exactly-once delivery across restarts. If you have used Flink, Kafka Streams, or Beam in unbounded mode: Clinker is not that.

The word “streaming” in Clinker’s documentation always refers to per-record evaluation within a single batch run – records flow through the graph one at a time rather than being materialized as a whole table – not to long-running stream-processor semantics. Internal identifiers in the codebase (function names like streaming_output_task, config fields like strategy: streaming, error messages, log lines) use the word in the same row-by-row sense; if you see it in a stack trace, it is not Flink leaking through.

Finite inputs only

Clinker reads sources that have an end. Files are the canonical shape, and finite-cursor network sources (paginated REST APIs with hard page/record caps) fit the same model – they exhaust their cursor and EOF. Unbounded sources (Kafka, Kinesis, Server-Sent Events, webhooks, tail -f-style file followers) are explicitly out of scope and will remain so.

Single process, ever

One clinker run invocation is one OS process. Parallelism happens inside that process via threads. Clinker does not spawn worker processes, does not coordinate a cluster, and does not shuffle data between machines. Scale by giving the host more cores, more RAM, more disk – the DuckDB / Polars / Kettle model. If a single host genuinely can’t fit the work, partition the input by file or by key and run multiple clinker invocations from a shell script; that’s a five-line script, not an architectural addition to Clinker.

For the full list of what Clinker deliberately does not do, see Non-Goals.

Pipelines are DAGs

A pipeline is a directed acyclic graph of nodes. Data flows from sources, through processing nodes, to outputs. There are no cycles – a node cannot consume its own output, directly or indirectly.

You define the graph with each node kind’s input fields. Most consumers use a single input:, Merge uses an inputs: list, Combine uses a named input: map, and Envelope names its body plus optional header and trailer streams. Clinker resolves these references, validates that the graph is acyclic, and determines execution order automatically.

From YAML to a run

Clinker first parses and validates the pipeline, binds schemas, compiles CXL, and produces a typed CompiledPlan. Execution entry points require that plan, so invalid pipelines are rejected before records are read. The current runtime then recompiles the configuration embedded in the supplied plan and executes the newly validated artifacts. It does not yet execute the stored DAG and other stored artifacts directly.

The locked D-01 through D-11 contract changes that lifecycle: the supplied CompiledPlan becomes authoritative, survives execution, and may be reused sequentially in one process while only a defined runtime envelope may refresh. Phase 5 owns that correction and any versioned persistent plan cache. Direct stored-plan execution and persistent caching are not current capabilities. See Stored-plan execution and cache identity for the current status and owner of each contract.

The nodes: list

Every pipeline has a single flat list of nodes. Each node has a type: discriminator that determines its behavior. The eleven node types are:

TypePurpose
sourceRead data from a file (CSV, JSON, XML, fixed-width)
transformApply CXL logic to reshape, filter, or enrich records
aggregateGroup records and compute summary values (sum, count, etc.)
routeSplit a stream into named ports based on conditions
mergeConcatenate multiple streams that share a schema
combineJoin records across N inputs with cross-input predicates
reshapeMutate or synthesize records within correlation groups
cullRemove whole correlation groups to a side-output port
envelopeFrame body records with document headers and trailers
outputWrite data to a file
compositionEmbed a reusable sub-pipeline

You can have as many nodes of each type as your pipeline requires. The only constraint is that the resulting graph must be a valid DAG.

CXL is per-record ETL

CXL is Clinker’s per-record ETL expression language. Each record flows through a CXL block independently. A program maps, filters, or enriches the current record in statement order; it is not a table query or join planner.

The core statements:

  • emit name = expr – produce a field in the output record, adding it or replacing the field of the same name. The record’s other fields are carried through unchanged, so there is no need to write emit id = id; a Sink can narrow what it writes (see Sink Nodes).
  • let name = expr – bind a local variable for use in later expressions. Local variables do not appear in the output.
  • filter condition – discard the record if the condition is false. A filtered record produces no output and is not counted as an error.
  • distinct / distinct by field – deduplicate records. distinct deduplicates on all output fields; distinct by field deduplicates on a specific field.

CXL uses and, or, and not for boolean logic – not && or ||. String concatenation uses +. Conditional expressions use if ... then ... else ... syntax.

System namespaces use a $ prefix: $pipeline.*, $source.*, $record.*, $window.*, and $vars.*. These provide access to pipeline and source context, per-record scoped state, window function state, and static configuration respectively.

Per-record evaluation and the memory budget

Within a run, stateless operators evaluate records one at a time: a CXL block sees exactly one record, with no table-level context and no per-record state carried across records. Clinker does not load an entire file into memory before processing it. This is what “streaming” means in Clinker – row-by-row evaluation inside a finite batch job, not Flink-style unbounded stream processing.

“One at a time” describes the evaluation model, not the transport. For efficiency, records move between stages in bounded batches (default 2048 events, tunable per stage or pipeline-wide via batch_size) rather than as individual messages – but each record is still evaluated independently against the CXL block. The batch is a handoff unit, not a window: an operator never needs the whole batch to process any single record. (Blocking operators – Aggregate, sort, grace-hash Combine – are the exception that do accumulate across records; see below.)

Per-record evaluation keeps per-row memory usage bounded for the stateless parts of the graph (Transform, Route, Merge, most Combine probe-side work, Output). Every stage is charged against the configured RSS budget. Fused Source → Transform → Sink paths run streaming, with no per-stage materialization, so a 100 GB CSV passes through with the same footprint as a 100 KB CSV. A stage that hands its output to a single downstream Sink also avoids a charged inter-stage buffer – single-branch Route, non-fused Merge, streaming Aggregate, and the Combine probe-side stream their result straight to the writer (see Streaming vs. Blocking Stages). The remaining boundaries – multi-branch Route fan-out, output that forks to several consumers, Composition bodies, diamond DAGs – materialize records into per-stage buffers that charge against the same budget envelope. Every materialized buffer can spill past the soft threshold, including buffers shared by several readers and Route/Cull output-port buffers. Readers run sequentially over the same immutable memory-or-spill backing; each opens one cursor, and the final reader takes the authoritative buffer regardless of declaration or dispatch order. A consumer that needs a full resident vector reserves that materialization first. If the overlap would exceed the hard limit, the engine fails before allocating with a structured E310 MemoryBudgetExceeded diagnostic that names the consumer.

Use clinker run --explain to see which nodes will materialize (buffer: materialized) versus which will stream (buffer: streaming) before runtime – that label is the canonical “which stages charge the budget” signal. See the --explain reference and the memory-tuning page.

Stateful operators must accumulate. Aggregate, sort, and grace-hash Combine cannot emit until they have seen enough input – sums need every addend, a full sort needs the last row, a hash join needs the build side complete. These operators run inside a configured RSS budget (default 512 MB) and degrade gracefully under pressure rather than OOM:

  • Aggregate uses hash aggregation by default and spills partitions to disk when soft/hard memory thresholds trip. When the input is already sorted by the group key, the planner picks streaming aggregation, which requires only constant memory.
  • Sort spills runs to disk and merges them.
  • Combine picks among in-memory hash join, grace hash join (spilled), and IEJoin / sort-merge depending on predicates and memory pressure. A pure-range Combine (band join with no equality key) runs the block-band IEJoin, which external-sorts each side and spills a matched-output sort to disk so both its input and its result stay inside the budget.

The memory ceiling is a first-class promise. Clinker is designed to share a server with JVM applications, databases, and other services without competing for RAM.

Input wiring

Consumer nodes reference their upstream via the input: field:

- type: transform
  name: enrich
  input: customers    # reads from the node named "customers"

Route nodes produce named output ports. Downstream nodes reference a specific port using dot notation:

- type: route
  name: split_by_region
  input: customers
  config:
    routes:
      us: region == "US"
      eu: region == "EU"
    default: other

- type: sink
  name: us_output
  input: split_by_region.us    # reads from the "us" port

Merge nodes accept multiple inputs using inputs: (plural):

- type: merge
  name: combined
  inputs:
    - us_transform
    - eu_transform

Combine, Envelope, and Composition have their own multi-input shapes. See Pipeline YAML Structure for the complete wiring table.

Schema declaration

Source nodes require an explicit schema: that declares every column’s name and type:

config:
  schema:
    - { name: customer_id, type: int }
    - { name: email, type: string }
    - { name: balance, type: float }
    - { name: created_at, type: date }

Clinker uses these declarations to type-check CXL expressions at compile time, before any data is read. If a CXL block references a field that does not exist in the upstream schema, or applies an operation to an incompatible type, the error is caught during validation – not at row 5 million of a production run.

Supported types include int, float, string, bool, date, and datetime.

Error handling

A pipeline picks one error handling strategy, in the top-level error_handling: block:

StrategyBehavior
fail_fastStop the pipeline on the first error (default)
continueRoute error records to a dead-letter queue file and continue

When using continue, Clinker writes rejected records to a DLQ file alongside the output. Each DLQ entry includes the original record, the error category, the error message, and the node that rejected it. This makes diagnosing production issues straightforward: check the DLQ, fix the data or the pipeline, and rerun. A run that dead-letters at least one record exits with code 2 rather than 0, so a scheduler can tell a clean run from a partial one.

See Error Handling & DLQ for the DLQ columns, the error categories, and the per-source options.

Interactive Guides

Some parts of Clinker are easier to understand by trying them than by reading about them. Each guide below is a single page that runs in your browser, on a desktop or a phone. You change a setting or tap a record, and the page shows what the engine does with it. The guides don’t run a pipeline; each one reproduces the rules described on the reference page it links back to.

Where does a null go?

How CXL works out an expression when a field is empty: and, or and not with null, why null == null is true, and why a filter drops a record whose condition comes out null. Reference: Null Handling.

Correlation keys

How one failing line takes the rest of its order to the DLQ, and how an Aggregate’s group_by decides between rejecting a whole group and recomputing totals without the failed line. Reference: Correlation Keys.

Document context

Where $doc.* values come from, how each file becomes its own document with its own Aggregate roll-up, and what dlq_granularity: document rejects. Reference: Document Envelope Context.

How many documents?

What an Envelope node’s preserve and concat do to the documents a Sink writes, how a synthesized footer differs between them, how an Aggregate after the Envelope rolls up, and which pipeline shapes E347 and E355 reject. Reference: Envelope Nodes.

Which files get read?

How a Source turns glob:, regex:, paths: and its filters into the list of files it reads, and in what order. Reference: Source Nodes → Choosing files.

Is my data still sorted?

Which stages keep a Source’s declared sort_order and which drop it, and when an Aggregate can stream instead of holding every group. Reference: Aggregate Nodes and Source Nodes.

Route and Merge

Where each record goes when a Route’s conditions are true, not true, or fail, in exclusive and inclusive mode, and how a Merge rejoins the branches. Reference: Route Nodes and Merge Nodes.

Combine playground

Which build rows where: matches for each driver row, and what match:, on_miss: and drive: do with them, including range joins. Reference: Combine Nodes.

Window functions

Which rows of a partition each $window.* function reads, and why $window.sum is a partition total while $window.cumulative_sum is the running total. Reference: Window Functions.

Where did my rows go?

What the end-of-run line N total, N ok, N written, N dlq counts, why the numbers often don’t add up, and which rows are in no number at all: filtered and duplicate rows, rows folded into Aggregate groups, Combine misses and Route branches nothing reads. Reference: Metrics & Monitoring.

Streaming vs. blocking

Which stages stream and which hold their output in a buffer, for several pipeline shapes, with the --explain lines each one produces. Reference: Streaming vs. Blocking Stages.

For engine developers

The Clinker Engine Internals book has three more:

  • a memory system explainer, linked from its Memory Arbitration & Scheduling chapter. It covers the memory budget, back-pressure, spilling to disk and the scheduler, with a simulator.
  • a range-join explainer, linked from its Combine Internals chapter. It runs the block-band IEJoin on small inputs: sorting and slicing into blocks, pruning block pairs, the kernel step by step, and the nested-loop fallback.
  • a retraction-loop explainer, linked from its Retraction Protocol chapter. It replays the commit loop that corrects an Aggregate whose group_by leaves out a correlation-key field, on three of the engine’s test pipelines.

Pipeline YAML Structure

A Clinker pipeline is a single YAML file with three top-level sections: pipeline (metadata), nodes (the processing graph), and optionally error_handling.

Top-level shape

pipeline:
  name: my_pipeline            # Required — pipeline identifier
  memory:                      # Optional — see ops/memory.md
    limit: "256M"              # Optional (K/M/G suffixes), default 512M
    backpressure: pause        # Optional, default `pause`
  vars:                        # Optional typed static configuration
    threshold: { type: int, default: 500 }
    label: { type: string, default: "Monthly Report" }
  date_formats: ["%Y-%m-%d"]   # Optional — custom date parsing formats
  rules_path: "./rules/"       # Optional — CXL module search path
  concurrency:                 # Optional
    threads: 4
    chunk_size: 1000
  metrics:                     # Optional
    spool_dir: "./metrics/"

nodes:                         # Required — flat list of pipeline nodes
  - type: source
    name: raw_data
    config:
      name: raw_data
      type: csv
      path: "./data/input.csv"
      schema:
        - { name: id, type: int }
        - { name: value, type: string }

  - type: transform
    name: clean
    input: raw_data
    config:
      cxl: |
        emit id = id
        emit value = value.trim()

  - type: sink
    name: result
    input: clean
    config:
      name: result
      type: csv
      path: "./output/result.csv"

error_handling:                # Optional
  strategy: fail_fast

Pipeline metadata

The pipeline: block carries global settings that apply to the entire run.

FieldRequiredDescription
nameYesPipeline identifier. Used in logs and metrics.
memoryNoMemory-arbitrator tuning. Nested fields: limit (RSS budget, K/M/G suffixes, default 512M) and backpressure (spill/pause/both, default pause). See Memory Tuning.
varsNoTyped static configuration accessible in CXL via $vars.*. Each key declares type and an optional default; see Scoped Variables.
date_formatsNoList of strftime-style patterns for date parsing.
rules_pathNoDirectory for CXL use module resolution.
concurrencyNothreads and chunk_size for parallel chunk processing.
metricsNospool_dir for per-run JSON metric files.
date_localeNoUnsupported. Any explicit value is rejected with E119. Use explicit date_formats entries.
log_rulesNoUnsupported. Any explicit value is rejected with E124; configure runtime logging outside pipeline YAML.
include_provenanceNoUnsupported. Any explicit value is rejected with E125. Use write_meta: true on each intended Output.

These three names are admitted only far enough to produce precise, spanned diagnostics. Empty strings/maps and false are still explicit values and are rejected before execution; omission is the only accepted form.

Reserved metadata contract

The current status and locked owner are explicit:

FieldCurrent statusLocked target and owner
date_localeRejected (E119)Remove it and express supported parsing with date_formats:.
log_rulesRejected (E124)Remove it; runtime telemetry policy is not authored in pipeline YAML.
include_provenanceRejected (E125)Remove it and set write_meta: true on each Output that needs a provenance sidecar.

For provenance sidecars that work today, set write_meta: true on an Output node. D-24 keeps that spelling. See Approved exceptions and rejected placeholders.

The nodes list

Every pipeline has a flat nodes: list. Each entry is a node with a type: discriminator that determines its kind:

TypeRole
sourceReads data from a file
transformApplies CXL expressions to each record
aggregateGroups and summarizes records
routeSplits records into named branches by condition
mergeConcatenates multiple upstream branches that share a schema
combineJoins records across N inputs with where: predicates
reshapeMutates or synthesizes records within correlation groups
cullRemoves whole correlation groups to a side-output port
envelopeFrames body records with optional document header and trailer streams
outputWrites records to a file
compositionImports a reusable transform fragment

Node naming

Every node must have a name: field. Names must be unique within the pipeline and must not contain dots – the dot addresses something other than a node: a route branch (split.high, see below) and a node inside a composition call site (enrich.ref). The rule covers every node kind, including sources, outputs, and composition call sites, and it covers nodes declared inside a .comp.yaml body. Names are used for wiring, logging, and diagnostics.

A dotted name is refused at plan time with E010, which names the node and the name to use instead:

node name "enrich.ref" is invalid: '.' is reserved for branch references and
composition call-site paths; rename the node to "enrich_ref" (use underscores
or hyphens) and update every reference to it

Wiring by node kind

Input fields live at the node’s top level, alongside name: and type:. Their shape is specific to the node kind:

Node kindInput shape
sourceNo input field
transform, aggregate, route, reshape, cull, outputOne upstream reference in input:
mergeOrdered list of upstream references in inputs:
combineQualifier-to-upstream map in singular input:
envelopeRequired body: plus optional header: and trailer: upstream references
compositionRequired primary input: plus an inputs: map binding every required composition port; the map is authoritative for DAG wiring

Single upstream – used by ordinary one-input consumers:

- type: transform
  name: clean
  input: raw_data       # References the source node named "raw_data"
  config: ...

Port syntax – for consuming a specific branch from a route node, use node.port:

- type: sink
  name: high_value_out
  input: split.high     # Consumes the "high" branch of route node "split"
  config: ...

Multiple upstreams – merge nodes use inputs: (plural) instead of input::

- type: merge
  name: combined
  inputs:
    - east_processed
    - west_processed
  config: {}

Qualified inputs – Combine uses singular input: with a map whose keys become CXL qualifiers:

- type: combine
  name: enriched
  input:
    orders: clean_orders
    products: product_catalog
  config:
    where: "orders.product_id == products.product_id"
    match: first
    on_miss: null_fields
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.name
    propagate_ck: driver

Envelope and Composition have additional port semantics. See Envelope Nodes and Compositions before wiring those node kinds.

Source nodes have no input field. They are entry points – adding an input: field to a source is a parse error.

Using the wrong field or value shape for a node kind is caught at parse time by strict deserialization.

Optional fields on all nodes

Every node type supports these optional fields:

  • description: – human-readable text for documentation. Ignored by the engine.
  • _notes: – arbitrary metadata (JSON object). Ignored by the engine and available to external tooling.
- type: transform
  name: enrich
  description: "Add customer tier based on lifetime value"
  _notes:
    color: "#4a9eff"
    position: { x: 300, y: 200 }
  input: customers
  config:
    cxl: |
      emit tier = if lifetime_value >= 10000 then "gold" else "standard"

Strict parsing

All config structs use deny_unknown_fields. If you misspell a field name – for example, writing inputt: instead of input: or stratgy: instead of strategy: – the YAML parser rejects it immediately with a diagnostic pointing to the typo. This catches configuration errors before any data processing begins.

Environment variable: CLINKER_ENV

The CLINKER_ENV environment variable can be used for conditional logic outside of pipelines (e.g., selecting channel directories or controlling CLI behavior). It is not directly referenced within pipeline YAML but is available to the channel and workspace systems.

Scoped Variables

Clinker’s scoped-variable system lets a pipeline read and write named values at three lifetimes: the pipeline run, the source, and the record. Each variable is declared on the Transform that writes it via that Transform’s declares: block (type, scope, optional default), written by the same Transform’s CXL with emit $<scope>.<name> = ..., and read inline from any downstream node via the $pipeline.*, $source.*, and $record.* namespaces.

The three scopes

ScopeLifetimeResetReader namespace
pipelineEntire pipeline runNever (per run)$pipeline.<key>
sourceOne per source file (Arc<str>-keyed)Per source-file$source.<key>
recordA single record as it flows through nodesPer record$record.<key>

Record-scope variables are the per-record private store: every transform along the row’s path can read them, but they never serialize as output columns unless explicitly re-emitted as a regular column. They are written with emit $record.<key> = ... from a transform that declares them.

Declaring variables

A scoped variable is declared on the Transform that writes it, in that Transform’s config.declares: list. Each entry is named, scoped, typed, and optionally given a default that satisfies reads firing before the writer has run:

- type: transform
  name: enrich
  input: orders
  config:
    declares:
      - { name: cutoff_date,  scope: pipeline, type: date,   default: "2024-01-01" }
      - { name: ingest_label, scope: source,   type: string, default: "prod" }
      - { name: fuzzy_score,  scope: record,   type: float }
    cxl: |
      emit id = id
      emit $pipeline.cutoff_date = "2024-01-01"
      emit $source.ingest_label = $source.file.file_stem()
      emit $record.fuzzy_score = fuzzy_match(name, $pipeline.canonical_name)

Allowed types: int, float, string, bool, date, date_time.

Each (scope, name) pair must be declared on exactly one Transform — the same pair declared on two Transforms is rejected at config-validation time, ahead of compilation. $pipeline, $source, and $record are flat shared namespaces; declare each name once and read it from every consumer.

The pipeline’s top-level vars: block is a separate, flat registry for static configuration read via $vars.<key> — it does not carry the nested pipeline: / source: / record: scopes:

pipeline:
  name: order_processing
  vars:
    fuzzy_threshold: { type: float, default: 0.85 }   # read as $vars.fuzzy_threshold

Built-in members of each scope ($source.file, $source.name, $source.row, $source.path, $source.count, $source.batch, $source.ingestion_timestamp; $pipeline.start_time, $pipeline.name, $pipeline.execution_id, $pipeline.batch_id, $pipeline.total_count, $pipeline.ok_count, $pipeline.dlq_count, $pipeline.filtered_count, $pipeline.distinct_count) are reserved — declaring a user variable with one of those names is rejected at parse time.

$source.count semantics

$source.count is the per-source record total for the Source that produced the current record. The total isn’t known until the source finishes, so you can’t use it during per-record evaluation: a read on a mid-stream record (in a Transform, Route, Window, or Merge) resolves to Null, while reads after the source has finished (such as a terminal aggregate emit) resolve to the per-source total.

This means a streaming denominator like value / $source.count yields Null on mid-stream records. If you need a running row counter, declare a scope: source variable on a Transform and increment it from that Transform’s CXL instead.

Reading variables

CXL access is identical for declared and built-in keys:

- type: transform
  name: filter_recent
  input: orders
  config:
    cxl: |
      emit id = id
      filter received_at > $pipeline.cutoff_date
      emit batch = $source.batch_id
      emit confidence = $record.fuzzy_score

Reads of undeclared keys are rejected with E203 (CXL name resolution failed) at compile time, with a “did you mean” suggestion that scans the declared registry.

Writing variables

A scoped variable is written by the Transform that declares it: list the variable in the Transform’s declares: block and assign it from the same Transform’s CXL with emit $<scope>.<name> = <expr>. The Transform still processes records normally — declaring and writing a scoped var is additive to its ordinary emit/filter logic.

- type: transform
  name: capture_header
  input: salesforce_in
  config:
    declares:
      - { name: batch_id,        scope: source, type: string }
      - { name: ingestion_label, scope: source, type: string }
    cxl: |
      emit id = id
      emit $source.batch_id = batch
      emit $source.ingestion_label = $source.file.file_stem()

- type: transform
  name: row_score
  input: enrich
  config:
    declares:
      - { name: fuzzy_score, scope: record, type: float }
    cxl: |
      emit id = id
      emit $record.fuzzy_score = fuzzy_match(name, $pipeline.canonical_name)

An emit $<scope>.<name> write to a variable the Transform does not declare is rejected at compile time. Requiring the declares: entry keeps the dependency between writers and readers visible at plan time.

Init phase: pre-runtime population

Set phase: init on a Transform to pre-compute a $pipeline.* or $source.* value from a config-file source before the main run starts:

- type: source
  name: config_src
  config:
    name: config_src
    type: csv
    path: config.csv
    schema:
      - { name: cutoff, type: int }

- type: aggregate
  name: max_agg
  input: config_src
  config:
    group_by: []
    cxl: |
      emit cap = max(cutoff)

- type: transform
  name: precompute_cutoff
  input: max_agg
  config:
    phase: init
    declares:
      - { name: cutoff_date, scope: pipeline, type: int }
    cxl: |
      emit cap = cap
      emit $pipeline.cutoff_date = cap

Init-phase nodes must be terminal — no runtime-phase node may consume from an init-phase Transform. (Init-phase nodes can chain through init-only descendants for compositions.) Use disjoint Sources for init vs runtime when you need both: a Source shared between an init and a runtime branch only feeds the init pass.

Compile-time validation

Scoped variables are checked before the run starts. Every reference and every writer is validated, and every flow from a writer to its readers is checked against the pipeline. Each code below tells you what to fix.

CodeWhat it catches
E109Channel targets a composition but carries vars: overrides.
E116Channel var changes an existing type, or any default mismatches its declared type.
E117Channel var name shadows a reserved system field for that scope.
E118Channel vars.source.<src> references an unknown source-node name.
E164An init-phase Transform has a runtime descendant.
E171A reader is not a transitive DAG descendant of its writer.
E172Bare $source.<custom> read downstream of a Merge or Combine.
E173Composition body reads a parent scoped var without opting in.
E174Composition _compose.scoped_vars declares a different type than the parent.
E175An init-phase node reads a runtime-only writer’s variable.
E203A reference to an undeclared scoped variable (resolver-level failure).

Cross-Transform duplicate declares: (the same (scope, name) declared on two Transforms) is rejected before the run starts. $pipeline, $source, and $record are flat shared namespaces; declare each name once and reference it from every consumer.

Each diagnostic points at the exact place you read or wrote the variable, plus the conflicting writer or parent declaration, so the report lands where you can act on it rather than in some unrelated configuration block.

Post-merge access: qualified $source.<input>.<key>

After a Merge or Combine, the bare $source.<custom> form is ambiguous: each record carries its own source’s value, but the reader’s intent is usually to compare across inputs. E172 rejects the unqualified form and the qualified form is the legal alternative:

- type: transform
  name: read_after_merge
  input: merged
  config:
    cxl: |
      emit id = id
      emit lt = $source.left_input.left_label
      emit rt = $source.right_input.right_label

The <input_name> segment matches the named input on the Combine (its IndexMap key) or the upstream node name on the Merge.

Composition opt-in

A composition body cannot see parent scoped variables by default — the seal is enforced by E173. To pass values across the boundary, the composition declares the schema of parent vars it consumes in its _compose.scoped_vars block:

# read_pipeline_var.comp.yaml
_compose:
  name: read_pipeline_var
  inputs:
    inp:
      schema:
        - { name: id, type: int }
  outputs:
    out: tap
  scoped_vars:
    pipeline:
      cutoff:
        type: int

nodes:
  - type: transform
    name: tap
    input: inp
    config:
      cxl: |
        emit id = id
        emit cutoff_seen = $pipeline.cutoff

The parent must declare cutoff with the matching type; mismatches raise E174.

What scoped variables are not

These are intentional non-features:

  • No persistence across runs. State is in-memory only. A pipeline run starts with declaration defaults; the writes don’t survive the process.
  • No undeclared writes. A Transform may only write a scoped variable it lists in declares:; an emit $pipeline.x to an undeclared name is a compile error. Requiring the declaration keeps every writer visible at plan time and the writer→reader dependency explicit in the DAG.
  • No dynamic var creation. The set of variables is closed at plan time, by design. This bounds memory and makes the validation matrix above tractable.

Channel overrides

A channel can both override a pipeline’s declaration defaults and add new entries across all four registries ($vars.*, $pipeline.*, $source.*, $record.*). Each registry has its own sub-block under vars: on a .channel.yaml, and each entry uses the same { type, default } shape that pipeline-side declarations use:

# Pipeline declarations
pipeline:
  name: orders
  vars:
    fuzzy_threshold: { type: float, default: 0.85 }   # $vars.*
nodes:
  - type: source
    name: orders_src
    config: { name: orders_src, type: csv, path: in.csv,
              schema: [{ name: id, type: int }] }
  - type: transform
    name: enrich
    input: orders_src
    config:
      declares:
        - { name: cutoff_date,  scope: pipeline, type: date,   default: "2024-01-01" }
        - { name: ingest_label, scope: source,   type: string, default: "prod" }
        - { name: tier,         scope: record,   type: string, default: "bronze" }
      cxl: |
        emit id = id

# channel/acme-prod/orders.channel.yaml
channel:
  target: ../../pipelines/orders.yaml
vars:
  static:
    fuzzy_threshold: { type: float, default: 0.95 }
  pipeline:
    cutoff_date: { type: date, default: "2026-01-01" }
  source:
    orders_src:
      ingest_label: { type: string, default: "acme-prod" }
  record:
    tier: { type: string, default: "platinum" }

The overlay lives in the tenant’s folder (channel/acme-prod/) and is applied with --channel acme-prod; the channel.target field is authoritative.

Override semantics (entry name already declared) require the channel’s type to match the declared type — mismatches produce E116. Add semantics (entry name not yet declared) extend the registry with a new declaration. In both cases, a default that does not match the entry’s type also produces E116. $source overrides are keyed by source-node name; an unknown source name produces E118. The reserved-name guard (E117) blocks channels from shadowing system fields like $pipeline.execution_id or $source.path. Channels that target a .comp.yaml may not carry vars: (E109 if they do).

See Channels for the full overlay rules and the channel manifest reference.

Channels

Channels make one pipeline serve many tenants. A single base pipeline is authored once; each tenant (a channel) layers its own configuration, variable defaults, and structural changes on top — without copying or editing the base YAML. The system is built for scale: thousands of per-tenant channels against one pipeline, with strict validation and per-value provenance.

A channel is a tenant. A group is a reusable overlay shared by many channels — selected automatically from a channel’s labels, or invoked by name. Groups and target files can contribute value clobber (config: / vars: / resources:) and an ordered op list (overrides:). A channel-wide manifest is narrower: it may contain only labels plus declared config and variables.

Workspace layout

Channels live in a channel-centric workspace. A clinker.toml at the workspace root declares the layout roots; the rest is folders of YAML:

workspace/
  clinker.toml                       # declares the [channel] and [group] roots
  pipeline/       *.yaml             # base pipelines  (the pipeline-default layer)
  composition/    *.comp.yaml        # reusable sub-pipelines
  schema/         *.schema.yaml      # shared schemas
  group/          *.group.yaml       # group overlays: selector, priority, overrides
  channel/<tenant>/                  # one cataloged channel resource folder
    channel.cfg.yaml                 # required manifest: identity, targets, labels, wide values
    orders.yaml                      # filename is descriptive; channel.target is authoritative

The channel id is the stable logical key in [catalog.channels]. A --channel tenant.globex invocation resolves that catalog entry directly and then selects the target file by its declared logical pipeline id. Neither a folder name, a file basename, nor the current working directory is an identity.

clinker.toml roots

[channel]
root = "channel"      # per-channel folders live under <root>/<channel-id>/
shard = "none"        # enumeration layout: none (default) | first-char | hash

[group]
root = "group"        # *.group.yaml definitions live here

Both tables are optional; omitting them defaults [channel].root to channel, [channel].shard to none, and [group].root to group. shard is an enumeration-ergonomics choice for very large channel trees (it splits the folder fan-out); a channel is always looked up by computed path regardless of shard scheme, so shard never changes resolution semantics.

Typed workspace catalog

The same clinker.toml declares stable logical identities in separate catalog namespaces for rules, schemas, compositions, pipelines, channels, and typed composition resources.

[catalog]
rules_root = "rules"

[catalog.rules]
"shared.dates" = "rules/shared/dates.cxl"

[catalog.schemas]
"shared.dates" = "schema/shared/dates.schema.yaml"

[catalog.compositions]
"etl.normalize" = "composition/normalize.comp.yaml"

[catalog.pipelines]
"daily.orders" = "pipeline/orders.yaml"

[catalog.channels]
"tenant.globex" = "channel/globex"

[catalog.resources.shared_orders]
kind = "file"
path = "data/orders.csv"
access = "read"

Logical identities are kind-scoped. The rule and schema named shared.dates above are distinct typed resources; asking for one kind never substitutes an entry from another kind. A missing identity or a reference through the wrong kind fails planning and names the catalog table that must contain it.

catalog.resources is additionally descriptor-typed. The current file kind accepts only path and access (read, write, or read-write), is admitted under fixed catalog entry and descriptor-byte caps, and contains no credential fields. Resource bindings use the logical key (shared_orders above), never the path.

Every catalog path and rules root is anchored to the selected workspace (from --base-dir or workspace discovery). Parent traversal, an absolute path outside the workspace, and a symlink whose canonical target escapes the workspace are rejected. Duplicate identities within one kind are rejected, as are two catalog identities—even across kinds—that alias the same canonical file. These checks happen before compilation, so neither lexical aliases nor symlink aliases can create a hidden second authority for one file.

For CXL modules, an explicit rule entry is used when present; otherwise the logical identity maps beneath one selected rules root. Root precedence is explicit CLI --rules-path, then pipeline.rules_path, then [catalog].rules_root, then the workspace-relative rules/ default. Selection chooses one root rather than searching several. Planning admits the bounded direct/transitive module closure into the compiled plan, after which execution does not reopen module source files. See Modules and use and the --rules-path reference.

The layer model

Every value and every op is attributed to exactly one layer. Layers apply in a fixed semantic order — never lexical or file order:

pipeline-default  <  group(s) by priority  <  channel-wide  <  channel-per-target
  • pipeline-default — the base pipeline’s own configuration.
  • group(s) by priority — every group applied to the run, ordered by priority (higher priority applies later and thus wins).
  • channel-wide — the channel manifest (channel.cfg.yaml): overlays that apply to every pipeline this channel runs.
  • channel-per-target — the per-target overlay file (<target>.channel.yaml): the highest-precedence layer.

Clobber, never deep-merge

A higher layer’s value replaces the lower layer’s value wholesale. There is no deep-merge and no list-append: overriding a list swaps the entire list. To override individual elements, model them as a keyed map (which the config: and overrides: surfaces already are), not a list — so each element is addressed and replaced by key. Every resolved value maps 1:1 back to the single layer that supplied it, and channels resolve / explain --field report that layer.

Structural ops (overrides:) apply in a total order — layer precedence first, then declaration order within a layer. Collisions are errors, never silent no-ops: adding a node whose name already exists, or targeting a missing or already-removed node, fails with a diagnostic anchored to the offending op.

Overlays are resolved before executable compilation. Structural op streams are concatenated in total order and folded over the base node list. Clinker then compiles that target once for typed candidate validation; only when every config: candidate passes name, type, ambiguity, and fixed-lock checks does the winning config map enter executable compilation. Scoped vars: are likewise validated before executor initialization. One invocation produces one validated effective plan.

Value clobber: config, vars, and resources

The value-clobber surface carries scalar overrides. It appears identically on a group, a channel manifest, and a per-target overlay.

config: overrides composition config knobs, keyed by node.param dotted paths (the composition node’s name, then the parameter name):

config:
  scorer.threshold: { value: 0.95 } # override the `threshold` knob of `scorer`

The override changes executed behavior, not just the rendered provenance: the composition body reads the knob as $config.<param>, which the planner constant-folds to the resolved value for that instantiation at compile time. The winning layer is still recorded in the provenance side-table, so channels resolve / explain --field continue to report which layer supplied the value.

A config: key that matches no parameter in the compiled plan is a hard error (E113) — a misspelled or stale key aborts the run rather than silently doing nothing.

Rebinding a composition resource

resources: changes which logical catalog resource supplies one declared composition slot. The key is composition-node.slot; the leaf uses the same { value, fixed } shape as other clobbers:

# channel.cfg.yaml, a group file, or a per-target file
resources:
  lookup.orders: { value: tenant_orders }

The base composition call must already declare orders under _compose.resources_schema, and tenant_orders must exist under [catalog.resources] with the required kind and capabilities. Group, channel-wide, and per-target candidates use the ordinary precedence order and retain every attempted layer plus the winner. fixed: true locks a lower binding against higher layers just as it does for config values.

Only a scalar logical identity is accepted as value. Inline descriptors and credential/profile/secret/token selectors are rejected at the strict YAML leaf. An overlay cannot introduce a slot, address an internal nested slot, or change ports, composition names, or config through this surface. Resource rebinding changes the semantic plan fingerprint but does not resolve credentials or open runtime handles.

Locking a value: fixed

fixed is metadata on the value it locks, never a sibling map. A config leaf uses { value, fixed }; a variable leaf uses { type, default, fixed }. fixed defaults to false. Unknown spellings and a misplaced top-level fixed: block fail at the authored line with the corrected leaf form.

# channel.cfg.yaml — the channel-wide manifest
channel:
  name: tenant.globex
  targets: [daily.orders]
config:
  scorer.threshold: { value: 0.9, fixed: true }
# order_fulfillment.channel.yaml — the per-target overlay (a higher layer)
channel: { target: daily.orders }
config:
  scorer.threshold: { value: 0.95 } # rejected: channel-wide locked this key

The per-target candidate is invalid because the channel-wide value is fixed; the diagnostic points to the per-target leaf and the run does not start. The resolved provenance remains 0.9, and channels resolve marks that winning layer (fixed). Invalid candidates are validated even when another layer would win, so a typo or type mismatch cannot hide behind precedence.

vars: overrides or adds scoped-variable defaults, using the same four scopes a pipeline’s own vars: block uses ($vars.* / $pipeline.* / $source.* / $record.*). Each leaf is the same { type, default } shape a pipeline declaration uses:

vars:
  static:                    # $vars.*
    currency: { type: string, default: "USD", fixed: true }
  pipeline:                  # $pipeline.*
    cutoff_date: { type: date, default: "2026-01-01" }
  source:                    # $source.<src>.*  — outer key is the source-node name
    orders:
      ingest_label: { type: string, default: "prod" }
  record:                    # $record.*
    tier: { type: string, default: "bronze" }

See Variables for the scoped-variable model these overlay.

Structural ops: overrides

The overrides: surface is an ordered list of discrete, name-addressed ops applied to the base pipeline’s node list before compilation. Each op is a mapping with an op: discriminant. Unknown keys, or keys that belong to a different op kind, are rejected at parse time.

The op vocabulary is add / remove / replace / set / bypass / patch_schema.

add — splice in a node

Insert a new node, either inline or as a composition reference. The splice anchor is exactly one of after: / before: / an explicit input:.

overrides:
  # Inline transform, spliced after `normalize` (its former consumers now read `stamp`):
  - op: add
    node:
      type: transform
      name: stamp
      input: normalize
      config:
        cxl: "emit order_id = order_id"
    after: normalize

  # A composition, named by `alias`, with a config knob for the injected node:
  - op: add
    composition: ../composition/fraud_check.comp.yaml
    alias: fraud_check
    after: normalize
    config:
      threshold: 0.8

after: X reads from X and repoints X’s former consumers onto the new node; before: X feeds X, taking over X’s former upstream. An inline node with no splice anchor keeps its own declared input:. Adding a node whose name already exists is an error.

remove — delete a node and rewire

Delete a node by name, repointing its named consumers through an explicit rewire: map so no dangling reference is left behind:

overrides:
  - op: remove
    target: legacy_audit
    rewire:
      route_priority.input: product_lookup   # <consumer>.input: <new upstream>

Each rewire: key is a <node>.input path; each value is the replacement upstream. Any consumer still referencing the removed node afterward is an error, as is removing a node that does not exist.

bypass — remove a linear node

Sugar for remove on a 1-in/1-out node: it auto-rewires the node’s sole consumer onto its sole upstream.

overrides:
  - op: bypass
    target: legacy_audit

bypass only applies to a single-input, single-consumer node; a fan-in/fan-out node must use the explicit remove op with a spelled-out rewire: map.

replace — swap a node’s definition

Replace a whole node by name, keeping its identity (and therefore every consumer edge) intact. The replacement node’s own name: must equal target:.

overrides:
  - op: replace
    target: normalize
    node:
      type: transform
      name: normalize
      input: orders
      config:
        cxl: "emit order_id = upper(order_id)"

set — set one field within a node

Set a single field within a named node by path. The currently addressable path is config.cxl — the primary CXL body of a transform / aggregate / combine node — so replacing a stage’s logic wholesale is a set, not a special case:

overrides:
  - op: set
    target: route_priority
    field: config.cxl
    value: >
      emit _route = if priority_level == "urgent"
        then "priority_report" else "fulfilled_orders"

Here _route is an ordinary audit field; it does not select an Output. Direct Outputs sharing route_priority each receive every record. To partition rows by destination, add a Route node with conditions that read the field (or express the conditions directly on the Route).

Any other field path is a hard error, never a silent no-op.

patch_schema — shape a source’s columns

Add / rename / modify / remove columns on a source node’s declared schema, via a column-name-keyed map (the map key is the column name). Each column carries exactly one op:

overrides:
  - op: patch_schema
    target: orders
    schema:
      amount:      { type: float, scale: 2 }       # modify: set any subset of attrs
      cust_id:     { rename: customer_id }         # rename (a physical->logical alias)
      order_notes: remove                          # drop an existing column (bare scalar)
      region:      { add: { type: string } }       # add a new column (map key = new name)

The modify leaf is a bare attribute map: it sets any subset of the column’s attributes (type, scale, precision, format, width, …), leaf-replace, keeping every attribute it does not name. A typo’d attribute is rejected rather than silently appended. The same grammar applies identically at every override layer (pipeline / group / channel).

The keyed-map shape (rather than a list) is deliberate: a column op is addressed and leaf-replaced by name, with first-class rename / remove / add, exactly matching the source-config schema patch grammar so the two surfaces resolve columns and their diagnostics identically.

rename is a source-column alias, not a bare relabel: the reader still binds the original physical column and re-labels its value under the new name, so downstream CXL and the output see the new name carrying the original column’s data. A missing column, an add that collides with an existing name, or a rename onto an existing name are all errors (E231–E233).

To see which layer set a given attribute on a patched column, trace it with clinker explain <pipeline> --field <source>.<column>.<attribute> (optionally --channel <name>); the output names the winning Base < Pipeline < Group < Channel layer and each shadowed one. See Field provenance.

Groups and selectors

A group (group/<name>.group.yaml) is a reusable overlay layer that sits between the pipeline default and the channel layers. It carries the same two surfaces every layer carries — config: / vars: value clobber and an overrides: op list:

group:
  name: enterprise
  targets:
    pipelines: [daily.orders]
    compositions: [etl.normalize]
  match: 'tier == "enterprise"'   # optional selector; higher priority wins
  priority: 20
config:
  scorer.threshold: { value: 0.8 }
overrides:
  - op: add
    node:
      type: transform
      name: fraud_stamp
      input: normalize
      config:
        cxl: "emit order_id = order_id"
    after: normalize

A group plays two roles under one concept:

  • Selector-derived — when match: is present, the group is applied automatically to every channel whose labels satisfy the CXL boolean. Multiple matching groups are ordered by priority (higher wins; the default priority is 0).
  • Standalone / explicit — when match: is absent, the group is never auto-selected; it applies only when invoked by name with --group. Groups are channel-agnostic — their overrides never read channel labels — so any group can run standalone against the base pipeline, with or without a channel.

Every group owns a non-empty explicit targets: set of catalog pipeline and/or composition ids. A selector only narrows that set: a matching label can never make the group global. Forced --group use is target-bounded by the same set.

Selectors are label-only CXL

match: is a bare CXL boolean expression evaluated in a restricted label-only context: the only names in scope are the channel’s labels. $record / $source / $pipeline / $vars / $doc, window and aggregate calls, now, and wildcards are all rejected, so a selector is a pure, deterministic predicate over labels.

match: 'region == "west" and tier == "enterprise"'

Labels are typed from their YAML/JSON scalar kind (string, bool, int, float), so the typechecker rejects label/literal type mismatches. A selector that references a label a channel does not declare is a hard error, never a silent false — a typo surfaces as an unresolved-identifier error rather than quietly excluding the channel.

The channel manifest

channel.cfg.yaml declares the channel identity, its non-empty pipeline target set, identity labels, and optional channel-wide values:

channel:
  name: tenant.globex
  targets: [daily.orders]
labels: { region: west, tier: enterprise }   # identity — drives group selectors
config:
  scorer.threshold: { value: 0.9, fixed: true }
vars:
  static:
    currency: { type: string, default: "USD", fixed: false }

Labels are identity, never a pipeline override. The manifest and its target set are required. Channel-wide overrides: and sources: are forbidden because they would apply graph/source/schema changes without a single admitted target; move those operations into the corresponding target file.

The per-target overlay

A target file overlays exactly one manifest-declared catalog pipeline and its admitted composition closure. The channel.target: logical id is authoritative; the filename has no identity semantics:

channel:
  target: daily.orders
config:
  scorer.threshold: { value: 0.95 }
overrides:
  - op: patch_schema
    target: orders
    schema:
      tax_exempt: { add: { type: bool } }

Complete admission and execution identity

Channel loading is fail closed. Clinker canonicalizes the workspace root and candidate path, rejects traversal and symlink escapes, opens each admitted file once, verifies its post-open identity, and reads it into one bounded byte buffer. UTF-8 validation, parsing, and content identity all use that exact buffer; the bytes cannot be swapped between validation and hashing.

Planning validates the whole admitted catalog before selecting a requested pipeline or target:

  • every manifest target must have exactly one target overlay;
  • every declared target is parsed and validated, including targets not selected for this run;
  • the complete reachable pipeline and composition closure is validated for every target; and
  • group discovery, channel discovery, or file I/O errors abort admission rather than silently skipping an entry.

The planned execution identity includes the selected pipeline bytes and every applied layer in precedence order: defaults, ordered groups, the channel-wide overlay, and the selected target overlay. Group priority, declaration sequence, and whether membership was derived or explicit are part of that identity. Changing applied bytes or their order changes the identity; changing an unapplied overlay does not.

CLI surface

Running with overlays

# Run as a tenant: resolves catalog identities and derives target-bounded groups.
clinker run pipeline/order_fulfillment.yaml --channel globex --base-dir .

# Force-include a group by name, with or without a channel.
clinker run pipeline/order_fulfillment.yaml --group enterprise --base-dir .

run resolves the overlay stack from the workspace (rooted at --base-dir, default the current directory) and folds the resolved overrides into the plan before execution. Overlay flags shared across run and explain:

FlagMeaning
--group <NAME>Force-include a group overlay by name (repeatable), provided its explicit target set admits the selected pipeline or composition closure.
--no-auto-groupsSuppress selector-derived group membership; only explicit --group overlays apply.
--channel <ID>Apply a logical id from [catalog.channels]; the selected [catalog.pipelines] id must appear in its manifest targets. Derives only target-admitted matching groups.

explain --field <node.param> --group <NAME> reports the same overlay stack for provenance lookups, mirroring run.

Inspecting overlays

channels resolve renders the effective post-overlay DAG for one target under a chosen channel and/or groups, with per-value provenance — which layer supplied each value and which group injected which node:

# Resolve the effective plan for the globex channel (derives matching groups from its labels)
clinker channels resolve pipeline/order_fulfillment.yaml --channel globex --base-dir .

# Preview a group overlay standalone (no channel)
clinker channels resolve pipeline/order_fulfillment.yaml --group enterprise --base-dir .

Here --channel is a logical id from [catalog.channels]. The selected pipeline must likewise appear in [catalog.pipelines]; resolve never guesses identity from the filename. Matching groups are considered only after their explicit target sets admit that pipeline or one of its composition dependencies.

channels lint compiles every cataloged channel target and reports every failure through the same resolver used by run and explain:

clinker channels lint --base-dir .

Membership and labels

# List the channels a group's selector currently matches
clinker channels group members enterprise --base-dir .

# Stamp/overwrite a label across one or more channels (idempotent)
clinker channels label set tier=enterprise globex initech --base-dir .

channels label set takes a key=value assignment; the value is typed by YAML scalar inference (true/false → bool, integers → int, decimals → float, else string) so numeric and boolean labels compare correctly against selectors. The channel manifest must already exist with its explicit channel.targets list; the command never creates a targetless manifest.

Renaming a base node

refactor rename-node renames a base node and propagates the rename to every overlay that references it (splice anchors, target:, rewire: keys) across the workspace:

# Preview every file that would change
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --dry-run

# Apply it, then re-lint
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --base-dir .
clinker channels lint --base-dir .

The new name must be letters, digits, and _ only.

Source config patches

Independent of the overlay op engine, a channel file can patch a source node’s parsed config directly through a sources: block, applied before validation and compile so the run behaves exactly as if the source YAML had been hand-edited. This is the same column-keyed schema grammar patch_schema reuses, plus multi-value and per-format option patches.

sources:
  transactions:                            # source-node name (unknown -> E230)
    options:
      record_path: batch_records           # set a scalar per-format option (bad key -> E235)
    split_to_rows:                          # keyed by field name
      items:      { mode: split, position_column: line_no }  # add-or-modify
      tags:       { position_column: ~ }    # clear one attribute
      line_items: remove                    # drop an entry (unknown field -> E234)
    max_output_rows_per_input: 10000        # replace the source ceiling; 0 disables it
    split_values:                           # keyed by field name
      codes:      { delimiter: "|" }        # add-or-modify an entry
      tags:       { delimiter: ~ }          # reset to the default delimiter
      notes:      remove                    # drop an entry (unknown field -> E234)
    schema:                                 # keyed by column name
      amount:      { type: float, scale: 2 }
      cust_id:     { rename: customer_id }
      order_notes: remove
      region:      { add: { type: string } }

All ops are keyed and leaf-replace — there is no deep-merge. On an existing split_to_rows / split_values entry a partial map is a modify: an omitted key keeps its current value, and a new entry takes the same defaults hand-written config would. Because an omitted key means “keep current”, clearing an attribute that is already set needs its own form — an explicit YAML null. On position_column that removes the attribute; on delimiter, which always holds some separator, it restores the ; default. options are merged onto the source’s current options and re-validated through the format’s option struct, so an unknown or mistyped key is rejected exactly as in hand-written config. A schema rename is a source-column alias — the same alias a base column can declare directly with source_name::

schema:
  # read the physical `cust_id` column, expose it downstream as `customer_id`
  - { name: customer_id, type: string, source_name: cust_id }

Format-structure patches (X12 / HL7 v2)

Beyond the format-agnostic ops above, a sources: patch can reshape the format-layer structures an X12 or HL7 source declares in its options: block — with keyed add/modify/remove grammar instead of blob-replacing the whole options map:

sources:
  interchange:                             # an X12 source
    group_section:                         # the GS functional-group declaration
      name: fg                             # rename the section (omit to keep)
      fields:
        e04: int                           # set/add a typed field
        e05: remove                        # drop a declared field
    set_section: remove                    # drop the whole ST declaration
  messages:                                # an HL7 v2 source
    split_fields:                          # keyed by positional field name
      f08: { components: 3 }               # add-or-modify a composite split
      f03: remove                          # drop a declared split

group_section / set_section patch the X12 nested-envelope declarations (the GS functional-group and ST transaction-set levels); split_fields patches the HL7 composite-field splits, keyed by positional field name and resolved by wire position (f8 and f08 address the same split). Each op applies only to a source of the matching format (anything else is E238). The set form is a partial modify on an existing declaration — an omitted name or axis width keeps its current value — and creates the declaration when absent, in which case name (X12) or components (HL7) is required (E240). Removing a declaration, field, or split the source does not carry is E239. These ops apply after the options merge, so they layer on top of an options value that replaces the same declaration in one patch.

Multi-record patches (discriminator-driven flat files)

A multi-record flat file interleaves several record layouts in one file, each identified by a discriminator tag. A sources: patch reshapes that layout with records: (keyed by record-type id) and a discriminator: merge, so a tenant’s record set can differ from the base without editing the pipeline:

sources:
  ledger:                                  # a multi-record source
    discriminator: { start: 2 }            # move the tag byte range (partial merge)
    records:
      detail:  { tag: X }                  # retag; a nested `columns:` reshapes fields
      trailer: remove                      # drop a record type
      header:                              # add a record type (map key = its id)
        add:
          tag: H
          columns:
            - { name: hdr_id, type: string, start: 1, width: 8 }

A records entry follows the same keyed grammar as schema: a bare remove drops the record type, { add: { tag, columns, ... } } declares a new one, and a bare attribute map modifies an existing one. A modify sets any subset of the record type’s tag / parent / join_key / description and carries a nested columns: map that runs the column-op grammar (modify / rename / add / remove) against that record type’s own fields. The discriminator: op merges field by field onto the current discriminator — a named field overwrites, an omitted one is kept — and the merged result must be a byte range (start + optional width) XOR a field.

These ops apply only to a multi-record schema (E241). Modifying or removing an unknown record-type id is E242, adding an id that already exists is E243, a merged discriminator that is neither pure byte-range nor pure field is E244, and a discriminator tag shared by two record types after the patch is E245.

Sources inside a composition body

A plain sources: key names a top-level source node. To patch a source declared inside a composition body, qualify the key with the composition call-site node name: <composition-node>.<source>. The composition body is expanded during compile, so the patch is applied to the body’s source when the body is bound — before the body typechecks — exactly as a top-level patch shapes a top-level source before it binds:

sources:
  enrich.lookups:                          # source `lookups` inside composition node `enrich`
    schema:
      code: { rename: lookup_code }

Resolution is one level deep: the qualifier must name a composition node in the pipeline (an unknown composition — or a nested a.b.c key naming a source inside a nested composition body — is E230), and the source half must name a source node declared in that composition’s body (an unknown one is E230, naming the body file). A plain unqualified key still targets a top-level source, and a name that matches no top-level source still fails with E230 — now hinting at the qualified form when the pipeline has compositions.

Note: an authored body Source must link to one declared composition resource slot with resource: <slot>; a direct path, glob, regex, or paths matcher is rejected. The source patch above changes schema/reader configuration but does not select the resource. Planning binds the slot and compiles a call-scoped Source instance. On a data run, a credential-free file binding opens through the CLI-prepared catalog factory and streams via the executor’s bounded Source path. Credential-bearing bindings remain unsupported and fail before opening because no profile-selection surface is available.

When a patch changes the effective source config, the run’s pipeline identity differs from the base and from other patched variants, so their outputs and lineage do not collide.

Diagnostics

CodeMeaning
E103A config: candidate has the wrong value type or attempts to override a lower fixed value, or a resource binding names an unknown/undeclared/incompatible slot or catalog identity. Every candidate is checked at its own leaf, including one a later layer would shadow.
E107A channel/group variable candidate disagrees with the pipeline declaration or its default does not match the declared type.
E110A variable candidate shadows a reserved scoped-variable name.
E111A vars.source candidate names no source in the selected pipeline.
E113A config: / override key matches no composition parameter in the compiled plan. A misspelled or stale key aborts the run instead of silently doing nothing.
E114An overlay op failed to apply (missing splice anchor, duplicate node name, missing/removed target, invalid set field, invalid bypass node). The diagnostic is anchored to the offending op’s source span, not the base pipeline.
E118A shorthand node.param candidate is ambiguous in the selected composition closure; use the exact target-specific node path.
E230A source patch (sources.<src> or patch_schema) targets a source that does not exist: an unknown top-level source, an unknown composition for a qualified <composition>.<source> key, a <composition>.<source> naming no source in that composition’s body, or a nested (a.b.c) key.
E231A schema rename / modify / remove of a column that does not exist.
E232A schema add of a column name that already exists.
E233A schema rename whose target name collides with an existing column.
E234A split_to_rows / split_values remove of a field with no matching entry.
E235An options patch sets an unknown or mistyped option key for the source’s format.
E236A renamed/aliased column’s exposed name collides with a real input field, which would mislocate that field. Raised at read time.
E237A schema patch on a multi-record / generated / external-file schema — column ops apply only to a single-record column list.
E238A group_section / set_section patch on a non-X12 source, or a split_fields patch on a non-HL7 source.
E239A remove of a nested-section declaration, declared section field, or field split the source does not carry.
E240A malformed format-structure patch: creating a nested section without a name, adding a split without components, a split key that is not a positional fNN name, or a zero axis width.
E241A records / discriminator patch on a single-record / generated / external-file schema — these ops apply only to a multi-record schema.
E242A records modify / remove of a record-type id the source does not declare.
E243A records add of a record-type id that already exists.
E244A merged discriminator that is neither a pure byte range (start + optional width) nor a pure field.
E245Two record types share a discriminator tag after the patch, which would make the reader’s discriminator dispatch ambiguous.

Compositions

Compositions are reusable pipeline fragments that can be imported into multiple pipelines. They encapsulate common transform patterns – date derivations, address normalization, currency conversion – into self-contained, testable units.

Using a composition

A composition node in your pipeline references an external .comp.yaml file:

- type: composition
  name: risk
  input: orders
  use: "./compositions/risk_score.comp.yaml"
  inputs:
    inp: orders
  config:
    threshold: 0.5

The use: field points to the composition definition file. The inputs: map binds each declared composition input port to an upstream node. The top-level input: is also required by the current node shape, but inputs: is the authoritative port wiring. The config: block passes parameters that customize this invocation.

Resolving the use: path

A use: value names a .comp.yaml in the workspace. It is resolved relative to the directory of the pipeline file being compiled, then against the set of .comp.yaml files discovered under the workspace root, finally falling back to a filename match. A use: that resolves to no .comp.yaml — a typo, a wrong relative prefix, or a file that does not exist — fails compilation with a spanned E103 diagnostic naming the composition node. The whole run aborts loudly; it does not silently drop the composition and write an empty output. The same holds for the other composition-binding errors (E102–E108): an ill-bound call site fails compile rather than producing a run that writes zero records. Run clinker explain --code E103 for details.

Composition definition file

A .comp.yaml file declares its interface in _compose: and its executable subgraph in nodes::

# compositions/risk_score.comp.yaml
_compose:
  name: risk_score
  inputs:
    inp:
      schema:
        - { name: order_id, type: string }
        - { name: amount, type: float }
  outputs:
    out: scored
  config_schema:
    threshold:
      type: float
      default: 0.5
      range: [0.0, 1.0]

nodes:
  - type: transform
    name: scored
    input: inp
    config:
      cxl: |
        emit order_id = order_id
        emit amount = amount
        emit high_value = amount >= $config.threshold * 2000.0

Composition fields

FieldRequiredDescription
_compose.nameYesComposition identifier
_compose.inputsYesNamed input ports and their minimum required schemas
_compose.outputsYesOutput port aliases pointing to body nodes or route ports
_compose.config_schemaNoTyped configuration parameters, defaults, and constraints
_compose.scoped_varsNoExplicit scoped-variable names the sealed body may read from its caller
_compose.resources_schemaNoTyped resource slots; see Resource bindings
nodesYesUnified node list for the sealed composition body

Reading config parameters in the body

A composition body reads its own config parameters as $config.<param>. The planner constant-folds each reference to the value resolved for that instantiation — the call site’s config: value, or a channel/group config: override, or the declared default — so the same composition used with different config: compiles to different bodies. Because the resolution happens per instantiation, a channel or group config: override changes what the body computes, not just the reported provenance.

Explaining config provenance

Every resolved composition parameter retains its base value and each attempted group, channel-wide, and per-target override. The winning layer, shadowed values, fixed locks, and source spans remain attached to the stable compiled node identity. Inspect a value with either a unique shorthand or its exact versioned address:

clinker explain pipeline.yaml --field 'risk.threshold'
clinker explain pipeline.yaml \
  --field '/v1/config/nodes/risk/fields/threshold'

The output includes the canonical exact address and lists layers in a stable order. Exact addresses include every enclosing composition call. For example, two sibling calls may each contain a local node named shared:

/v1/config/calls/left/nodes/shared/fields/threshold
/v1/config/calls/right/nodes/shared/fields/threshold

In that case shared.threshold fails with E118 instead of selecting one by insertion order. The diagnostic lists both exact addresses in deterministic order; copy the intended --field correction. An unknown query fails with E117 and lists only same-field candidates, never unrelated nodes or fields. An empty query fails with E116.

Address segments use RFC 6901 escaping: ~ becomes ~0 and / becomes ~1. Unicode remains unchanged. This makes rendering and parsing lossless, including after provenance serialization and repeated inspection of a compiled plan.

Source-schema provenance uses the parallel exact form /v1/schema/sources/<source>/columns/<column>/attributes/<attribute>; the three-part source.column.attribute shorthand remains available.

Body validation

Nodes inside a composition body are validated with the same node-scoped config checks as top-level pipeline nodes. A body node that would be rejected at the top level — an envelope wiring the not-yet-supported trailer: port, a transform declaring a reserved variable name or a default that does not match its declared type, an invalid log directive, or a batch_size: 0 — fails compilation with an E115 diagnostic naming the composition call site, the body file, and the violation. Run clinker explain --code E115 for details.

A body source or sink that sets a CSV delimiter or quote_char which is not exactly one ASCII byte is likewise rejected at compile time, not first at run, with the same one-byte rule top-level nodes get.

A body source or sink whose schema: names an external .schema.yaml file has that path resolved relative to the composition file’s own directory (not the invoking pipeline’s), and the file’s columns are inlined before the body binds. A body Sink therefore rounds decimal columns to their declared scale at the write boundary exactly as a top-level Sink does.

A body sink cannot be combined with a Source that declares dlq_granularity: document: document-level dead-lettering needs every Sink at pipeline level, and the pipeline fails compilation with E378. Run clinker explain --code E378 for how to move the Sink out of the body.

Executable example corpus

The five fragments under examples/pipelines/compositions/ are executable examples of the current authoring surface:

  • Clean Names
  • Fiscal Date Fields
  • Order Classification
  • Shipping Cost
  • Validate Email

Each fragment uses _compose.inputs, _compose.outputs, config_schema, and the unified nodes list. The two date-dependent examples require an explicit as_of_date configuration value so their results do not depend on the day the test runs.

The composition example test recursively inventories every .comp.yaml file in that directory. Its case manifest must name exactly that discovered set: empty or missing inventories, duplicate keys, missing or extra cases, and paths that escape the corpus directory fail with distinct diagnostics before any example runs. Each case is then loaded through the production composition loader, placed in a generated pipeline, and executed by the real clinker binary. The test checks the exit status, record counters, and output bytes.

clinker run --explain proves that a generated pipeline compiles. It does not prove that the example produces the documented result; the executable corpus’s byte comparison is the behavior check.

Advanced wiring

For compositions with multiple input ports, bind every declared port by name:

- type: composition
  name: enrich_address
  input: orders
  use: "./compositions/order_product_enrich.comp.yaml"
  inputs:
    orders: orders
    products: product_catalog

The primary input: must be present for the current YAML node shape. The planner builds composition edges from the named inputs: map, so that map must contain every required port declared by _compose.inputs.

Downstream nodes consume a composition’s declared output ports using the usual node.port syntax. If there is only one output port, the bare composition node name selects it.

Resource bindings

A composition declares the external capabilities it needs as typed slots. The currently admitted kind is file; it requires read capability and a run-local file opener:

_compose:
  name: order_lookup
  inputs:
    input: { schema: [{ name: order_id, type: string }] }
  outputs: { out: order_reference }
  config_schema: {}
  resources_schema:
    orders:
      kind: file
      required: true

nodes:
  - type: source
    name: order_reference
    config:
      name: order_reference
      type: csv
      resource: orders
      schema: [{ name: order_id, type: string }]

resource: is an explicit body-Source-to-slot link. The slot must be declared by the enclosing _compose.resources_schema; the Source name and its format do not select a resource implicitly. A resource-backed body Source must not also declare path, glob, regex, or paths, because the bound catalog resource is its only external target. Every authored body Source must declare resource:. Composition input ports are separate synthetic roots: to consume caller-provided rows, declare _compose.inputs.<port> and set a downstream node’s input: <port> instead of authoring a Source node for that port.

Top-level Sources are unchanged: they continue to require exactly one direct matcher. resource: on a top-level Source is rejected until a separate top-level binding surface is designed.

The workspace supplies a concrete, secret-free descriptor under a logical identity in clinker.toml:

[catalog.resources.shared_orders]
kind = "file"
path = "data/orders.csv"
access = "read"

The call site binds only the declared slot to that logical identity:

- type: composition
  name: lookup
  input: orders
  use: ./compositions/order_lookup.comp.yaml
  inputs: { input: orders }
  resources: { orders: shared_orders }

Resource descriptors are strict. Unknown kinds or fields, inline objects, unknown catalog identities, undeclared slots, missing required slots, and kind/capability mismatches fail planning. A call site cannot contain a path, credential profile, secret, token, or opened handle. File descriptors must remain inside the workspace and the catalog is admitted under fixed entry and descriptor-byte limits.

Planning retains the winning logical identity and every attempted overlay layer for each binding. That identity also participates in the semantic plan fingerprint. For each authored body Source, planning compiles a distinct call-site-scoped instance carrying only the slot, logical identity, resource kind, required capabilities, opener family, run lifetime, provenance, and stable logical dataset identity. It does not retain the catalog’s physical path. During clinker run, the CLI resolves that credential-free file requirement at the workspace edge and transfers an opaque single-use reader factory to the executor. The executor acquires the complete compiled group, opens all of its Sources before starting any of them, and streams their finite records through the ordinary bounded Source path. An open/read failure or interruption closes every opened session and releases the group. This surface still does not select credentials: a group requiring credentials fails before runtime effects because no credential-profile option exists yet.

Ordinary composition calls do not have outputs: or alias: fields. Either key fails with E377 at its authored location. Use _compose.outputs for the public output contract and the composition node’s name for its caller-visible namespace. add.alias remains valid only inside an overlay add operation, where it names the inserted node.

Call-site fields

FieldRequiredDescription
inputYesPrimary upstream required by the current node shape
useYesPath to the .comp.yaml definition
inputsYes for declared portsMap of composition input ports to upstream node references
configNoParameter overrides (key-value pairs)
resourcesNoDeclared slot to logical [catalog.resources] identity; scalar values only
outputsRejectedE377: declare ports under _compose.outputs and use node.port downstream
aliasRejectedE377: use this composition node’s name as the namespace

Contract status

The bounded catalog, typed file slot/binding, overlay provenance, stable logical dataset identity, and E377 call-surface rejection implement the planning half of D-12 through D-16. Credential references, credential resolution, and runtime handle activation are not implemented; a resource binding is therefore a validated planning contract, not permission to perform I/O.

See the canonical composition-resource and call-site contract for status, evidence, compatibility impact, and the AUTH-01 boundary.

Complete example

pipeline:
  name: order_pipeline

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./data/orders.csv"
      schema:
        - { name: order_id, type: string }
        - { name: amount, type: float }

  - type: composition
    name: risk
    input: orders
    use: "./compositions/risk_score.comp.yaml"
    inputs:
      inp: orders
    config:
      threshold: 0.5

  - type: sink
    name: result
    input: risk
    config:
      name: result
      type: csv
      path: "./output/scored_orders.csv"

Correlation Keys

A correlation key declares a set of records from a single source as an atomic group: if any record in the group fails validation or processing, the whole group is sent to the DLQ. This is the right shape for transactional data where partial processing is worse than total rejection – the canonical example is an order with multiple line items where one bad line should reject the entire order.

This page describes how to declare a correlation key and how it behaves through each node that can fan out, fan in, group, or join records.

Interactive companion: the correlation keys explainer lets you make lines fail and see what reaches the output and the DLQ, with and without an Aggregate.

Declaration

Correlation keys are declared per source. Each source’s config: block carries an optional correlation_key: field naming the column (or list of columns) whose value identifies a record’s correlation group within that source.

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: ./data/orders.csv
      correlation_key: order_id
      schema:
        - { name: order_id, type: string }
        - { name: amount, type: int }

  - type: source
    name: customers
    config:
      name: customers
      type: csv
      path: ./data/customers.csv
      correlation_key: [customer_id, region]   # multi-column key
      schema:
        - { name: customer_id, type: string }
        - { name: region, type: string }
        - { name: name, type: string }

  - type: source
    name: sensor_readings
    config:
      name: sensor_readings
      type: csv
      # No correlation_key: record-level errors land in the DLQ as
      # standalone entries with no group atomicity.
      schema:
        - { name: ts, type: date_time }
        - { name: value, type: float }

A record’s correlation group is identified by the tuple of values for that source’s listed fields. Records sharing the same tuple within the same source belong to the same group. There is no pipeline-level correlation key — declare it on each contributing source.

The group identity is captured at ingest, so rewriting the key column in a later Transform does not change a record’s group — anonymizing or transforming order_id downstream still keeps the original grouping intact.

A source whose declared correlation_key: field names a column not present in its own schema: block is rejected at compile time with diagnostic E153. The fix is to add the field to the schema or remove it from correlation_key:.

Row order

Declaring correlation_key on a Source changes the order its rows reach the rest of the pipeline. When the Source’s declared sort_order already begins with the correlation-key fields, in either direction, that declared order is kept as it is. Otherwise the planner inserts a sort after the Source, ascending on the correlation key and then on the remaining declared sort_order fields, so that each group’s rows are adjacent. Rows with equal keys keep their order in the file.

Every consumer that depends on row order sees this order: Sink output order, an Aggregate’s first-arriving value, Cull and Reshape ties, and which build record a Combine’s match: first picks. Adding a correlation key to an existing pipeline can therefore change those results even when no record fails.

DLQ semantics

When a record fails inside a correlation group:

  • The failing record produces a trigger DLQ entry. Its category reflects the actual failure (e.g. type_error, validation_failed).
  • Every other record from the same source in that group produces a collateral DLQ entry, carrying the category correlated.
  • Records belonging to other (clean) groups proceed normally.

A record with a null value for the correlation-key field is treated as its own group: it has no peers, so DLQ atomicity does not span multiple records.

The dlq_count counter sums triggers and collaterals. It counts one trigger per failure, so a row that fails twice in a group counts twice, as it does without a key.

Group buffering

The engine buffers records per correlation group until either the group completes or a failure triggers a flush. The max_group_buffer: field on the pipeline-level error_handling: block caps per-group buffering across every source’s groups:

error_handling:
  max_group_buffer: 100000     # Default: 100,000

A group that goes over the cap is dead-lettered whole when the run commits it. It is not a hard error: the run continues.

  • Each row of the group that failed on its own is written as its own trigger, with its own category and _cxl_dlq_id.
  • The group’s other rows are written under one group_size_exceeded trigger, the first of them; the rest are correlated rows that carry its id as their _cxl_dlq_trigger_id.
  • A group whose rows all failed writes only those failures, with no group_size_exceeded row.

Each failure is written once, with its Combine build row if it has one (see Combine interaction), each condemned row once, and every row counts toward dlq_count and the DLQ rate limits. The group_size_exceeded row’s _cxl_dlq_timestamp is when the group went over the cap, and its error detail states the cap and how many entries the group held.

The cap counts the entries a group holds, not its distinct rows: a row counts once for each Sink it reaches, and each failure counts once. A row that an inclusive Route sends to two Sinks counts twice, and so does a row that fails on one branch and reaches a Sink on another. A failing Combine match counts once, although it holds both the driver row and the matched build row.

Going over the cap does not stop a group from buffering. The group keeps buffering its rows until the run commits it, so today the cap decides how a large group is dead-lettered but does not bound the memory it uses.

Per-operator interactions

Route interaction (fan-out)

A correlation group can span multiple route branches. Group atomicity is preserved across branches: if any record in the group fails (in any branch’s transform, or in the route predicate itself), the entire group is rejected from every branch.

For an inclusive route where one record reaches both branches, a single failure DLQ’s that source row exactly once — not once per branch. A record that fails on both branches has two failures and is written twice, once as each branch’s trigger with its own stage, as it is without a key.

Merge interaction (fan-in)

Merge concatenates upstream branches that share a schema. Records keep their correlation identity through the merge, so rows from different sources that share the same key value become one correlation group downstream: a failure on any one of them DLQ’s the whole group across both sources.

Per-source rollback narrowing

When two sources contribute records to the same correlation group, a failure originating from one source does not collaterally DLQ records from the other source. The collateral fan-out is scoped to the failing source’s records only.

For example, with [src_a, src_b] → merge → transform → out where both declare correlation_key: id, an error that fires on a src_b row produces a trigger for that row while the src_a row sharing the same id is spared and reaches the output. Single-source pipelines behave exactly as a pipeline-wide collateral DLQ would, since every co-grouped record shares the one source.

Two cases stay group-wide rather than narrowing per source:

  • max_group_buffer overflow DLQ’s every record in the overflowing group — no single source is to blame for the overflow.
  • Combine output failures DLQ the synthesized output row, which has no single-source attribution. The exception concerns the output row only: the matched build record’s dead letter follows the failing driver’s group, as described under Combine interaction, and never widens the narrowing to the build record’s source.

Aggregate interaction

When an aggregate’s group_by covers every correlation-key field, the aggregate stays on the strict path: each emitted row inherits the correlation identity of its inputs, and any DLQ trigger in the group rolls back every record in the group, including the aggregate output row.

- type: aggregate
  name: order_totals
  input: orders                         # correlation_key: order_id
  config:
    group_by: [order_id]                # covers the key
    cxl: |
      emit total = sum(amount)

When an aggregate’s group_by omits a correlation-key field, the engine automatically retracts only the failing records and recomputes the affected groups, so the surviving contributions still produce a correct aggregate row. You do not configure this — the engine picks the path from the group_by content. (One restriction: this mode cannot be combined with strategy: streaming, which is rejected at compile time.)

- type: aggregate
  name: dept_totals
  input: orders                         # correlation_key: order_id
  config:
    group_by: [department]              # omits the key — surviving rows recomputed
    cxl: |
      emit total = sum(amount)

Combine interaction

Every combine declares propagate_ck: to select which correlation-key fields its output rows carry:

  • propagate_ck: driver — output inherits only the driver input’s correlation identity. The common case; today’s strict-correlation pipelines stay on this setting.
  • propagate_ck: all — output carries the union of correlation-key fields across every input. Use when the build side carries keys that downstream operators need to read.
  • propagate_ck: { named: [<field>, ...] } — output carries exactly the named subset. Use to project a multi-field key down after a join.
- type: combine
  name: enriched
  input:
    o: orders                          # driver (correlation_key: employee_id)
    d: departments                     # build side
  config:
    where: "o.employee_id == d.employee_id"
    match: first
    on_miss: skip
    cxl: |
      emit employee_id = o.employee_id
      emit amount = o.amount
      emit dept = d.dept
    propagate_ck: driver

How match mode fills the propagated key:

  • match: first — the single matched build’s key fills the slot.
  • match: all — one output row per matched build, each carrying its own build’s key.
  • match: collect — one row per driver; the first matched build’s key fills the (single-valued) slot, while every matched build’s full payload still rides inside the array column.

Driver wins on a name collision: if both the driver and a build input declare the same key field, the output keeps the driver’s value.

propagate_ck is a required field — every combine must spell out which mode it uses.

A failing match’s build-side dead letter follows the driver’s group. When the combine body fails for a driver row, that driver row is the trigger of the driver’s correlation group. The matched build record’s dead letter is held with the same group as a collateral (_cxl_dlq_trigger: false, category combine_output_row), written right after its driver’s row and carrying that driver’s _cxl_dlq_trigger_id, and it is written or rolled back exactly when the driver’s group is. It never condemns the build record’s own correlation group: another driver that matched the same build record keeps its output unless its own group failed. The build record is written once per failure: when several drivers fail against one build record, in one group or in several, each failing driver’s row is followed by its own copy of the build row, carrying that driver’s _cxl_dlq_trigger_id. A driver that fails against several build rows is written once per failure, each copy followed by the build row of that failure.

Composition interaction

A composition’s body operates on records flowing in from the parent pipeline; correlation identity flows into the composition inputs and back out the named ports unchanged. Compositions cannot declare their own correlation key — a key is a property of a source, not of the composition body that consumes a source’s records.

Debugging

Correlation grouping is tracked on internal columns you never write in YAML or CXL, and they are hidden from writer output by default. To surface them for debugging, set include_correlation_keys: true on a Sink node:

- type: sink
  name: debug
  input: any_node
  config:
    type: csv
    path: "./debug.csv"
    include_correlation_keys: true

The output then contains extra columns named $ck.<field> (literal prefix in the CSV header) for each declared correlation-key field.

To investigate DLQ collaterals: every collateral entry’s category is correlated, and the trigger entry in the same group carries the actual failure category and message.

See also

Document Envelope Context ($doc.*)

Many enterprise file formats wrap their record body in an envelope: named sections that surround the records and carry document-level metadata — a batch header with a run date and batch id, a trailer with a record count and checksum, or arbitrary sibling sections. Clinker exposes these sections to CXL through the $doc.<section>.<field> namespace.

nodes:
  - type: source
    name: payments
    config:
      name: payments
      type: xml
      path: data/payments.xml
      options:
        record_path: payments/Payment
      envelope:
        sections:
          BatchInfo:
            extract: { xml_path: "/payments/BatchInfo" }
            fields:
              batch_id: string
              run_date: date
          Summary:
            extract: { xml_path: "/payments/Summary" }
            fields:
              record_count: int
              checksum: string
      schema:
        - { name: amount, type: int }

Interactive companion: the document context explainer shows which records read each section, how each file becomes its own document, and what dlq_granularity: document rejects.

Like the rest of the pipeline config, the envelope: block is strict: an unknown key at any level (a misspelled sections:, extract:, or fields:) is rejected at plan parse time with a diagnostic naming the bad key, rather than being silently ignored.

A downstream transform reads any declared section field on every body record:

nodes:
  - type: transform
    name: tag
    input: payments
    config:
      cxl: |
        emit batch = $doc.BatchInfo.batch_id
        emit expected_total = $doc.Summary.record_count
        emit amount = amount

Section names are yours

The engine reserves no section names. BatchInfo and Summary above are arbitrary identifiers chosen by the pipeline author — Head / Foot, preamble / trailer, batch_metadata / eob_summary are all equally valid. A section name is whatever string you put in the sections: map; CXL exposes it verbatim as $doc.<that_name>.<field>.

Extracted sections are available throughout the body stream

Every extracted section is available to every body record. JSON and XML can pre-scan declared sections anywhere in the document, so a header at the top and a trailer at the bottom are both visible from the first record to the last. SWIFT likewise scans its service blocks before emitting fields. Multi-record CSV and fixed-width record_type extraction captures the leading header region only; it does not extract arbitrary trailing sections.

For formats with trailing-section extraction, a trailer field is available during body processing, not just at end-of-file. A Transform can compare every row, including the first, against the trailer’s count:

- type: transform
  name: check_count
  input: payments
  config:
    cxl: |
      emit amount = amount
      emit declared_count = $doc.Summary.record_count

Note that an extracted trailer section you read via $doc.* is distinct from the structural counts an EDI reader validates internally (the X12 SE/GE/IEA, EDIFACT UNT/UNZ, HL7 BTS/FTS segment counts). Those trailer counts are checked by the reader against the body it streamed, and a mismatch is a structural-integrity failure — see Malformed envelopes for how dlq_granularity: document dead-letters a malformed file instead of aborting the run.

The pre-scan reads the envelope-bearing segments of the file before body streaming begins. The amount of retained data depends on the declared sections, their payloads, and the format’s path to a trailing section:

  • JSON streams the pre-scan. The reader walks the document once and deserializes only the subtrees the declared sections point at — every other key (including a multi-megabyte body array) is parsed-and-skipped without being stored. The retained sections live in a bounded document index capped by max_index_bytes (see below), so retained section memory scales with the declared sections, not the document size.
  • XML streams the pre-scan too. The reader event-walks the document once and flattens only the declared section subtrees — every other element, including a multi-megabyte body, is event-walked and dropped without being materialized. The retained sections live in the same bounded document index capped by max_index_bytes. The reader still holds the file’s raw bytes for the lifetime of the read (one disk read backs both the body parser and the pre-scan), but the pre-scan no longer materializes the undeclared section subtrees.

For multi-record CSV, retained section names, field names, values and document containers carry allocation ownership against the run’s finite memory budget. Their charge lasts through downstream document use and any surviving aliases. A repeated header does not create an unbounded history of section snapshots; resource refusal remains fatal even under strategy: continue. CSV accepts the same UTF-8 and Latin-1 policy for section cells as for body cells.

Only the envelope sections live in the document context. This separation does not mean the executor retains only one body record: downstream Source and Combine paths can still materialize whole inputs. See what the memory budget measures. Other formats’ existing document-index limits remain as described below.

Bounding envelope retention with max_index_bytes

For JSON and XML sources, the document index is capped so a pathologically large declared section fails loud rather than exhausting memory. The cap is charged incrementally as each section is parsed — byte by byte while the section’s subtree is built — so even a single oversized declared section aborts mid-parse, before its whole subtree materializes, naming the offending section and the cap. (Undeclared siblings, including a multi-megabyte body, are skipped without being parsed into the index at all.)

- type: source
  name: events
  config:
    name: events
    type: json
    path: "./data/events.json"
    options:
      record_path: data.rows
      max_index_bytes: 64MB   # cap on retained envelope sections

The same max_index_bytes option applies to an XML source’s options: block, capping its envelope pre-scan identically.

max_index_bytes accepts a decimal size string (64MB, 500KB) or a bare byte count. It is optional; when omitted the reader applies a documented finite default of 64MB. Only the declared sections a program actually reads are retained, so envelope metadata sits far below this ceiling in practice — the cap exists to convert an unbounded mistake into a clear error.

Extract rules per format

Each section declares how the reader locates its payload:

Formatextract: keyValue
XMLxml_pathSlash-path to the section element, e.g. /doc/Head
JSONjson_pointerRFC 6901 pointer — empty (whole document) or leading /, e.g. /Head
EDIFACTsegmentA service-segment tag — only UNB
X12segmentA service-segment tag — only ISA (GS/ST surface as nested levels)
HL7 v2segmentA header-segment tag — only FHS (BHS/MSH surface as nested levels)
SWIFT MTsegmentA service block: "1", "2", "3", or "5" (or its default label)
Multi-record CSV / fixed-widthrecord_typeA header record-type tag, e.g. H

xml_path and the source-level record_path option are both slash-paths over XML but root differently: xml_path tolerates a leading / (/doc/Head is its documented form), while record_path rejects one. They locate different things and are deliberately not aligned — see XML Format → record_path and xml_path root differently.

Declaring an xml_path section against a JSON source (or vice versa), a segment extract against XML/JSON, a record_type extract against any format other than multi-record CSV / fixed-width, or any envelope: block at all on a plain (single-schema) CSV or fixed-width source is a configuration error that fails fast rather than silently producing empty sections.

A json_pointer must be a valid RFC 6901 pointer: either empty ("", naming the whole document) or a /-introduced path such as /Head or /batch/summary. A slashless value like Head is a typo — it would decode to zero segments and silently match the root document — so it is rejected at validation rather than resolving to the wrong metadata.

A plain (single-schema) CSV or fixed-width source carries no envelope — there is no header/trailer structure to extract. Declaring envelope: sections on one is a configuration error (E356): the sections would never be populated and every $doc.<section>.<field> against the source would resolve to null, so the compiler rejects it rather than accepting an inert declaration. Envelope extraction on a flat file — the record_type extract — applies only to a multi-record source, one declaring a discriminator: + records: block; declare that schema if the file genuinely carries header/trailer records.

Network (REST) sources carry no $doc context

A rest source pulls its records page by page over paginated HTTP — it has no single buffered document with head and tail sections, so it carries no envelope context. Envelope sections are a file-document concept, so the compiler rejects them on a REST source rather than letting them silently resolve to null:

  • Declaring an envelope: block on a rest source is an error (E349) — the declaration would be inert.
  • Reading $doc.<section>.<field> from a node fed by a rest source is an error (E349) — the access can never resolve.

Pull the document-level metadata into record fields through the API’s own response shape (record_path, split_to_rows) instead, so the value travels as a normal field rather than as document-envelope context.

EDIFACT segment extract

An EDIFACT source exposes its interchange header UNB as an envelope section. The section’s field names are the positional element keys e01, e02, … :

envelope:
  sections:
    interchange:
      extract: { segment: "UNB" }
      fields:
        e05: string          # interchange control reference

Only the UNB header is extractable. Trailer segments (UNT, UNZ) that arrive after the body are not envelope sections — their control counts are validated by the reader instead. A mismatch between a trailer’s declared count and the body the reader streamed is a structural-integrity failure: by default it aborts the run, and under a source’s dlq_granularity: document opt-in it dead-letters the whole file to the DLQ (see Malformed envelopes). A segment extract naming any tag other than UNB is rejected at startup. See EDIFACT Format for the full reference.

Multi-record record_type extract

A multi-record CSV or fixed-width source — one that declares a discriminator: and a records: list — exposes a header record type as an envelope section through the record_type extract. The tag names which of the source’s declared record types carries the section’s payload; the matched header row’s named fields become the section’s fields.

schema:
  discriminator: { start: 0, width: 1 }
  records:
    - { id: header,  tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
    - { id: detail,  tag: D, columns: [ { name: amount,   type: int, start: 1, width: 9 } ] }
    - { id: trailer, tag: T, columns: [ { name: count,    type: int, start: 1, width: 9 } ] }
  structure:
    - { record: trailer, count: count }
envelope:
  sections:
    head:
      extract: { record_type: H }       # the H header row → $doc.head.*
      fields:
        batch_id: string

Only a header record type — one whose rows precede the body at the file head — is extractable as a $doc section; the reader captures the first such row in a bounded pre-scan and excludes it from the body stream. A trailer record type (one named by a structure: constraint) arrives after the body it closes, so it is not an envelope section — its declared count is validated against the streamed body count instead, the same structural-integrity check the EDI trailers use. See CSV and Fixed-Width for the full reference.

A JSON example:

- type: source
  name: payments
  config:
    name: payments
    type: json
    path: data/payments.json
    options:
      record_path: records
    envelope:
      sections:
        Head:
          extract: { json_pointer: "/Head" }
          fields:
            batch_id: string
        Foot:
          extract: { json_pointer: "/Foot" }
          fields:
            count: int
    schema:
      - { name: amount, type: int }

against:

{
  "Head": { "batch_id": "RUN-001" },
  "records": [ { "amount": 10 }, { "amount": 20 } ],
  "Foot": { "count": 2 }
}

Typed fields

Each section’s fields: map declares the field name and its type, drawn from the same small vocabulary as source schemas: string, int, float, bool, date, date_time. The extracted raw value is coerced to the declared type at pre-scan time; a value that cannot coerce (e.g. a non-numeric string declared int) fails the source with a diagnostic naming the section, field, and offending value.

A field that the document does not carry resolves to null — $doc.* follows the same missing-value convention as $source.* and $pipeline.*. A section that the document does not carry at all is simply absent from the context; any $doc.<missing_section>.<field> resolves to null.

Declared-path validation

Every $doc.<section>.<field> reference is cross-checked at compile time against the schema the feeding source’s reader will actually serve. A reference that can never resolve — almost always a typo — is rejected at compile time, pointing at the node that made it, rather than resolving silently to null. How the path is checked depends on the source:

Closed-schema sources (XML, JSON) — the envelope: block is the complete schema: the reader extracts exactly the sections and fields it declares. A reference naming a section the source does not declare, or a field the declared section does not declare, is rejected with error E341 ($doc.Summry.total against a declared Summary). Run clinker explain --code E341 for the full write-up.

Multi-record CSV / fixed-width sources — a source whose schema declares discriminator: + records: exposes a header record type as a $doc section through the record_type extract. The reader coerces the matched header record’s columns through the section’s declared fields: and serves exactly those fields, so the section is closed just like an XML/JSON one: a reference naming an undeclared section, or a field the section does not declare, is rejected with error E341. (A plain single-schema CSV / fixed-width source has no such structure — declaring an envelope: on one is rejected with error E356.)

Segment/positional sources (X12, EDIFACT, HL7) — the file-level header (ISA/UNB/FHS) is declared through envelope: and is closed, but the reader also synthesizes nested envelope levels the config never names — X12’s functional_group / transaction_set, HL7’s batch / transaction_set — keyed by positional eNN / fNN elements bounded by the source’s max_elements / max_fields. A $doc path is checked against that synthesized vocabulary plus any section/field you declared, so a legitimate wire-derived path ($doc.functional_group.e06) is accepted while a misspelled section ($doc.functonal_group.e06) or an out-of-range positional element ($doc.transaction_set.e99) is rejected with error E348. Run clinker explain --code E348 for the full write-up.

REST sources carry no document and reject $doc outright — see Network (REST) sources carry no $doc context above (E349).

Plain (single-schema) CSV / fixed-width and SWIFT MT sources are not statically checked. A plain flat file synthesizes no $doc sections, and SWIFT serves declared sections under user-chosen or default block labels — neither fits the closed or positional model cleanly. (A multi-record CSV / fixed-width source is checked — see above.)

Indexed $doc access

A section field that holds an array or a map can be indexed inline, the same way any record value is — see Nested paths for the full bracket-index reference. Integer indices select array elements; string keys select map entries; the two compose into a chain:

cxl: |
  emit first_line = $doc.Header.line_items[0]          # array element
  emit run_date   = $doc.Header.meta["run_date"]       # map entry
  emit first_sku  = $doc.Header.line_items[0]["sku"]   # array-of-maps chain

An out-of-range array index or a missing map key resolves to null — it never errors or panics, and a null mid-chain short-circuits the rest of the chain to null rather than failing. This is the same missing-value convention $doc.<missing_field> and $source.* follow.

Only literal paths are readable

Every index segment must be a literal — a constant integer or string written in the program text. The section and field are always literal identifiers (the grammar requires it), so a literal index is the last piece a reader needs to know, before reading any input, exactly which envelope paths a run will consume. The pre-scan extracts precisely those statically-resolvable paths and nothing else.

A computed index — one derived from runtime data, such as $doc.Header.line_items[row_index] — is not statically resolvable: the reader cannot pre-scan a row-dependent element. clinker rejects it at compile time with a diagnostic pointing at the offending index, rather than reading it at run time. Use a literal index, or pull the value into a record field upstream and index that instead.

This compiles together with the declared-path rule above: a $doc read of a section or field the source does not declare is a compile-time error (E341), and a $doc read with a computed index is a compile-time diagnostic — neither reaches run time as a silent null. Only a literal path over a declared section is pre-scanned and readable.

One document per file

Each source file is its own document with its own envelope context. When a source matches multiple files (via glob: / paths:), each file gets a fresh document context with its own section values. Records from different files never share a context — a record’s $doc.* always reflects the file that record came from.

Document boundaries flow through the pipeline so that document-scoped operators fire at exactly the right point. A document-scoped operator fires exactly once per document, even when that document arrives across several inputs that a Merge or Combine brings together.

Per-document aggregation

A grouped or global Aggregate reading a multi-document source produces one set of grouped rows per document, not a single aggregate spanning every document. When a document closes, the Aggregate finalizes and emits the groups belonging to that document, then drops their state before the next document accumulates — so a glob: source over twelve monthly files through a group_by Aggregate yields twelve independent monthly roll-ups, and only one document’s groups are ever live at once (the others have already been emitted and freed, or have not yet started).

This applies only when a document boundary actually reaches the Aggregate. A plain single-file source is one document, so it still emits one aggregate. A Merge that combines several distinct single-document sources flushes those sources independently downstream — one roll-up per source document, exactly as feeding each source to its own Aggregate would. This holds for every Merge mode.

A Combine (join) preserves document boundaries on every strategy, so a per-document Aggregate downstream of a join also rolls up per driver document.

Nested (multi-level) envelopes

Some formats wrap their records in several envelope levels, one inside another. EDI X12 is the canonical example and the first format that implements this: an interchange (ISA/IEA) contains one or more functional groups (GS/GE), each containing one or more transaction sets (ST/SE), each containing the records. A single file can carry multiple interchanges back to back. See X12 Format for the full reference.

HL7 v2 is the second multi-level format: an optional file (FHS/FTS) contains optional batches (BHS/BTS), each containing one or more messages (MSH..), each containing the segment records. The tiers map onto the same nested levels — the FHS file header is a declared segment: "FHS" section, while the BHS batch and the MSH message surface automatically as the reader-supplied sections batch and transaction_set. Every tier is optional, and a level’s section exists only when its header segment is present in the input — the reader synthesizes a batch section only where it reads a BHS, and the message level only where it reads an MSH.

A bare MSH-led stream — messages with no BHS batch and no FHS file wrapper — therefore opens only the message level. $doc.transaction_set.* resolves against each message, but $doc.batch.* resolves to null: no BHS was read, so no batch section was synthesized. Whether a file carries a BHS batch wrapper is a property of the input bytes, not the config, so a $doc.batch.* path is never a compile error — it follows the same missing-value convention an absent $doc section follows (resolve null, never error), and populates as soon as a BHS wrapper is present:

# Bare MSH stream (no BHS/FHS) — only the message level opens:
- msg_type: $doc.transaction_set.f08   # MSH-9 message type — populated
- batch_id: $doc.batch.f01             # null — no BHS, so no batch tier

# BHS-wrapped stream — the BHS opens a batch level, so both populate:
- msg_type: $doc.transaction_set.f08   # MSH-9 message type — populated
- batch_id: $doc.batch.f01             # BHS field — now populated

See HL7 v2 Format for the full reference.

A reader for such a format opens and closes each nested level as it crosses the corresponding envelope boundary mid-file. Each level contributes its own sections to $doc. There is no new $doc syntax for nesting — every level’s sections are read through the same two-level $doc.<section>.<field> lookup. A record inside the innermost level sees every enclosing level’s sections at once. For X12 the interchange header is a declared segment: "ISA" envelope section (you choose its name), while the GS group and ST set surface automatically as the reader-supplied sections functional_group and transaction_set, each keyed by positional eNN elements:

cxl: |
  emit interchange_control = $doc.interchange.e13        # ISA13, declared section
  emit functional_id       = $doc.functional_group.e01   # GS01 (reader-supplied)
  emit transaction_type    = $doc.transaction_set.e01     # ST01 (reader-supplied)
  emit claim_amount        = amount                       # body field

A record streamed inside the ST level resolves the ST section, the enclosing GS section, and the outermost ISA section, all at once: each inner level inherits every enclosing level’s sections as siblings in one flat namespace. If two levels declare a section with the same name, the innermost wins for records inside it — the same shadowing rule a nested scope follows in any language. Picking distinct per-level names (as above) keeps every level independently visible.

The reader-supplied default names are not the only option: an X12 source can name the GS and ST levels itself and give each a typed field schema, so a nested level is addressable under a chosen name with coerced fields exactly like the declared ISA section. The declaration lives on the source’s options (not the envelope: block, which is reserved for the pre-scannable file-level header), and each nested level is named independently:

type: x12
options:
  group_section:
    name: functional_group       # your choice — the engine reserves no name
    fields: { e06: int }         # GS06 group control number, typed
  set_section:
    name: transaction_set        # your choice
    fields: { e01: int }         # ST01 transaction-set id, typed

Omit a level’s declaration and it falls back to its reader-supplied default name keyed by untyped positional eNN strings. See X12 Format for the full reference.

Boundaries nest correctly through the pipeline: each level opens before the records inside it and closes after them, in strict innermost-first order. A level that arrives across several branches is still handled once where a Merge or Combine brings those branches together — exactly like a single-level document.

Header-only interchanges

A multi-level envelope file can legitimately carry an interchange whose body is empty — envelope structure (an interchange header, and possibly inner group headers) with zero records inside. Such an interchange still opens a document and emits its open/close boundaries, so downstream operators and trailer-count validation observe it just like any other document. The interchange’s $doc.* sections are extracted and the boundaries flow even though no body record ever streams from it.

The same holds for an empty inner envelope — an open/close pair with no records between — and for an inner envelope that opens or closes after the file’s last body record. Every envelope boundary a reader signals is applied, whether or not a record follows it, so the document frame stays balanced end to end.

Fixed-width document output

A fixed-width Sink can echo any extracted section as its header or footer. The names select document context; they do not declare new input extraction rules. For example, these Sink config keys select sections named manifest and totals:

reconstruct_envelope: true
options:
  envelope:
    header_from_doc: manifest
    footer_from_doc: totals

Both names are author-defined. With a multi-record flat-file source, both sections must have been captured in the leading header region, even though totals is rendered at the end of the output document. A structural trailer count check does not make that trailing input record an extracted section.

Header and footer fields concatenate in stored order without the body’s fixed-width padding. Strings remain verbatim, numbers and booleans use scalar spellings, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null adds no text. Arrays and maps are unsupported: a structured field anywhere in the section rejects that entire header or footer before delivery. A missing section emits nothing; a present empty section emits only the configured LF/CRLF separator, or no bytes with line_separator: none.

Document start, each record, and document end are separate prepared operations. Only successful delivery advances document state or the body count; opening the next document starts its own count and selects its own sections. Earlier successful operations remain delivered if a later footer fails. A computed footer record-count field is unsupported for fixed-width (E346). See fixed-width output and output preparation.

Native JSON and XML output boundaries

JSON/XML reconstructed envelopes prepare document start, each body record, document end and finalization as separate complete operations. A document’s body count advances only when a record is delivered. Ending one document and opening the next resets that count and restores the next document’s own sections. Section names remain the names declared in your pipeline.

Reconstructed JSON envelopes retain their document grammar: array mode wraps the envelope documents in an array; NDJSON mode separates complete envelope documents with LF. Unlike ordinary compact NDJSON records, reconstructed pretty: true documents can span lines. XML places each Document wrapper inside one configured root, retaining header/footer section wrappers and emitting no XML declaration.

Explicitly opened empty documents still carry framing and a zero body count. In the CLI, a native source with no body records does not open a writer, so its output is an empty file even with reconstruction enabled. Do not infer this behavior from header-only message interchanges, whose readers emit explicit boundaries even without a body.

Ordinary native output supports correlation, splitting and per-file fan-out. Reconstructed envelopes combined with splitting, per-file fan-out, correlation or document-grain DLQ are rejected before execution (E347). These combinations cannot establish the required single document boundary.

Library writers distinguish draining already delivered bytes from finalizing a document. A byte drain adds no closing syntax. Once delivery fails, neither a drain, finalization nor teardown resumes the failed operation. See output preparation for failure effects.

Error Handling & DLQ

Clinker provides structured error handling with a dead-letter queue (DLQ) for records that fail processing. The error_handling: block at the top level of the pipeline YAML controls the behavior.

Configuration

error_handling:
  strategy: continue
  dlq:
    path: "./output/errors.csv"
    include_reason: true
    include_source_row: true

Strategies

error_handling.strategy is pipeline-wide – it is set once at the top level, not per node. It controls what happens when a record fails:

StrategyBehaviorExit code
fail_fastDefault. Abort the run on the first record failure.Non-zero, by the class of the aborting error (3 for an evaluation failure, 4 for an I/O failure – see Exit Codes)
continueRoute the failing record to the DLQ and keep processing.2 if any record was dead-lettered, 0 otherwise

There are exactly two, because the engine makes exactly one decision at each record failure: propagate it and stop, or dead-letter it and carry on.

fail_fast

The safest strategy. Any record-level error (type coercion failure, validation error, missing required field) halts the pipeline immediately, with a non-zero exit and no DLQ file. Use this when data quality is critical and you prefer to fix issues before reprocessing.

Some failures abort the run under either strategy, because they are not record-scoped: an unwritable output path, a config or CXL compile error, and the DLQ-rate ceiling (dlq.max_rate, E315/E316) all end the run regardless of the strategy.

CSV, JSON and XML resource failures are also fatal under either strategy. Memory or disk admission refusal, allocation failure, descriptor exhaustion and temporary-storage failure are not bad-record errors, so continue cannot turn them into successful output. A typed resource diagnostic preserves the kind of failure rather than reporting every case as a memory shortage. A failed destination can already have accepted a prefix; see output preparation.

Malformed JSON/XML input encoding is a data failure, including when discovered during schema discovery or envelope pre-scan. Under fail_fast, the CLI returns exit 4 and machine code source.data.invalid. It does not report a compilation error merely because no record has reached the pipeline. A late error can leave an already delivered prefix; an envelope pre-scan may discover it before any body records. Failed runs do not publish their staged normal output files.

Explicit cancellation ends an interrupted run with exit 130; it does not add a Sink error. If a real I/O or resource failure occurs alongside a shutdown request, the real failure retains its classification. Record and byte counters describe established progress, not rows merely attempted or prepared.

An executor invariant failure also aborts under either strategy with exit code 1. In particular, if a planned materialized input is unavailable when its consumer runs, Clinker stops instead of treating that input as a legitimate zero-row result. The message names the consuming node and planned producer (including the producer port when applicable) and says the input was not treated as empty. A source or stage that really emits zero rows remains valid; it carries an explicit empty buffer and completes normally. Report any missing- input internal error as an engine defect rather than routing it to the DLQ.

continue

The production workhorse. Bad records are written to the DLQ file with diagnostic metadata, and the pipeline continues processing remaining records. After the run completes, inspect the DLQ to understand and correct failures.

A pipeline that completes with DLQ entries exits with code 2 – this signals “pipeline completed successfully but some records were rejected.” It is not a crash or internal error. A continue run that dead-letters nothing exits 0, exactly like a clean fail_fast run.

Migrating from best_effort. The removed best_effort spelling was a third name for the continue behavior: it wrote the same DLQ entries and produced the same exit code, because the runtime never distinguished the two. Replace it with strategy: continue. A pipeline still carrying best_effort is rejected at config-validation time with a message naming the replacement.

Declared source-type failures are deliberately not lossy: under continue, the complete original record is written to the configured DLQ and no null, raw, or partially converted replacement enters the pipeline. This strategy therefore requires an error_handling.dlq block when such a failure occurs. fail_fast stops on the first failure without emitting a replacement. This includes fields renamed by a source schema: rejection retains the original decoded record and its values, even when conversion failed after other fields had already been examined.

An evaluation error is never false

A condition that fails to evaluate, such as one that divides by zero, has not said whether it holds. The engine never reads that failure as “false”, on any node: the record the condition was about is dead-lettered, and no decision that depends on the condition being false is taken for it.

  • A Transform filter that fails dead-letters the record; it is neither kept nor filtered out.
  • A Route branch condition that fails dead-letters the record, which takes no branch and not the default. In exclusive mode only the conditions up to the first true one are evaluated, so a later condition cannot fail.
  • A Combine where: that fails for a candidate build row dead-letters that pair. The driver is not unmatched, so on_miss does not fire; under match: first a failure on the deciding candidate is the driver’s only result, and under match: collect the driver writes no row. See Combine.

A pipeline that wants a failing condition treated as false says so in CXL, for example by guarding the division or coalescing the result with ?? false.

DLQ configuration

The DLQ is always written as CSV, regardless of the pipeline’s input/output formats.

  dlq:
    path: "./output/errors.csv"
    include_reason: true
    include_source_row: true
FieldRequiredDefaultDescription
pathNo–The pipeline-wide DLQ file. It receives the dead letters of every Source without its own per_source path. A dead letter with neither this path nor a per_source path for its Source is counted in the run’s dead-letter totals, sets exit code 2 and counts toward max_rate, but is written nowhere. The dlq: block itself is required to continue past a declared source-type failure.
include_reasonNotrueInclude _cxl_dlq_error_category and _cxl_dlq_error_detail columns.
include_source_rowNotrueInclude the failing record’s columns after the _cxl_dlq_* columns. Which record columns each DLQ file carries is fixed when the pipeline compiles; see How the DLQ columns are chosen. With false, only the _cxl_dlq_* columns are written.
max_rateNononeStop the run (E315, exit code 3) once the dead-lettered rows reach this fraction of the source rows read so far, both counted across the whole run. Must be greater than 0.0 and at most 1.0 (E318). Without it, the run is never stopped for its dead-letter rate. See Bounding how much can dead-letter.
min_recordsNo100How many source rows must have been read before max_rate is checked, so the first failures of a run cannot trip it on a tiny denominator. Also the default for each per_source min_records.
per_sourceNo–Settings for individual Sources, keyed by Source node name: a separate DLQ file, and a rate ceiling of their own. See Per-source DLQ settings.

Per-source DLQ settings

per_source gives a Source its own DLQ file, its own rate ceiling, or both. Each key is the name of a Source node:

error_handling:
  strategy: continue
  dlq:
    path: ./output/errors.csv
    max_rate: 0.05
    per_source:
      vendor_feed:
        path: ./output/vendor_feed_errors.csv
        max_rate: 0.20
        min_records: 500
      orders:
        max_rate: 0.01
FieldDefaultDescription
path–A separate DLQ file for this Source’s dead letters. They are written only there and do not appear in the pipeline-wide file. Without it, the Source’s dead letters go to the pipeline-wide path.
max_ratenoneStop the run (E316, exit code 3) once this Source’s dead-lettered rows reach this fraction of the rows read from this Source so far. Must be greater than 0.0 and at most 1.0 (E318).
min_recordsthe pipeline-wide min_records, else 100How many rows must have been read from this Source before its max_rate is checked.

A Source’s own max_rate is checked first, so a breach names that Source. The pipeline-wide max_rate, when set, still applies to the run as a whole. In the example, vendor_feed may dead-letter up to 20% of its own rows, but the run still stops when all dead letters together reach 5% of all rows read.

A key that does not name a declared Source is rejected at compile time (E317). Two DLQ paths that name one file are rejected too (E318), including paths that differ only in case on a case-insensitive filesystem, or in being written relatively and absolutely. A DLQ path that names the same file as a Sink’s path is rejected with E322.

Which record columns each file carries follows from the Sources routed to it; see How the DLQ columns are chosen.

How DLQ output is written

Dead-letter rows are written while the run executes, not collected until it ends. Each row is formatted under its file’s header, which is fixed when the pipeline compiles (see How the DLQ columns are chosen), and written into a staged copy of that file in the run’s publication attempt: in quarantine next to the destination by default, or under local_spool_dir with mode = "local_then_publish" (see Output publication). Each open DLQ file writes through one fixed 64 KiB buffer, so the memory the DLQ files use does not grow with the number of failures.

  • A DLQ file is created when its first row arrives. A file no row reaches is not created, and no empty file is published.

  • DLQ files are published only if the run succeeds, by the same publication step as the pipeline’s other outputs. A failed or interrupted run publishes no DLQ file.

  • The failures that are counted but have no destination (see path above) are never formatted or written.

  • Three kinds of dead letter are held in memory until the stage that found them finishes, and written then:

    • a join_values collision at a Sink that writes on its own thread;
    • an Aggregate add_record failure found while the Aggregate reads its input on its own thread;
    • a Combine output-row failure found while the Combine streams its driver on its own thread, or inside a grace-hash, sort-merge or IEJoin join.

    A Sink that writes on its own thread stops the run with an internal error once 65,536 collisions are waiting this way.

  • Under a correlation key or dlq_granularity: document, records are held until their group or document is decided. That is those features’ own state, described in their sections, not DLQ output; the rows they dead-letter are then written like any other. Under dlq_granularity: document the engine also keeps, for each rejected document, a compressed record of which rows it has already written, so a row held by several Sinks is written once. That record is charged to the memory budget. While a Sink writes a rejected document’s rows, the record grows by:

    • up to 16 bytes per row when the Sink receives the rows in the order their Source read them;
    • up to 96 bytes per row when it receives them in any other order, for example after a Sort on a data column;
    • up to 672 bytes per row for a document whose rows span its Source’s 4,294,967,296th row, in any order. A Source counts its rows across every file it reads.

    The first row of each document a Sink starts costs up to 920 bytes. The growth is released when the Sink’s pass ends, or every 65,536 rows of a document. Between releases one document’s record therefore grows by at most about 1 MiB when its rows arrive in the order they were read, about 6 MiB in any other order, and about 42 MiB when its rows span the 4,294,967,296th row. For example, 65,535 rows that alternate between a document’s first 65,536 rows and its next 65,536 grow the record to about 5.8 MB, and the release at the next row brings it back to 880 bytes. Once released, the record costs at most 4 bytes per written row, plus 96 bytes for every 65,536 rows of the document, written or not, plus under 1 KiB. That is about 2.3 KiB for a million-row document whose rows are contiguous. When only one row in each 65,536 is written, it is 88 bytes per written row plus 672 bytes, within the cap of about 100 bytes per written row. The record cannot spill: if one more row would not fit once every held row (below) has moved to disk, the run fails with E310. A failed document’s failing records are formatted as DLQ rows when they fail and held until the document is rejected. They are held in memory, charged to the memory budget, and move to one file in the spill directory only when the budget needs the memory, counting toward storage.spill.disk_cap_bytes (E320). With memory to spare nothing is written to disk. If even one more held row would not fit with every held row on disk, the run fails with E310.

Disk bounds how much DLQ output a run can produce: the free space at the staging location, and the publication attempt’s byte ceiling (storage.publication.max_attempt_bytes, and no more than retained_byte_limit), which every staged file of the run counts toward, DLQ files included. An attempt larger than that ceiling is refused at publication and nothing is published.

When the staging location fills, the run fails with an I/O error (exit code 4) and publishes nothing. After the destination’s own error text, the message names the DLQ file, the number of rows dead-lettered so far, and the stage and category with the most rows, and suggests a breaker:

<destination error>: dead-letter output ./output/errors.csv could not be written after 1048576 dead-lettered rows (most from stage transform:validate_orders, category validation_failure)

help: stop the run before dead letters fill the destination, for example:

  error_handling:
    type_error_threshold: 0.05
    dlq:
      max_rate: 0.05

These breakers bound how many rows can dead-letter; disk at the staging destination bounds their volume.

Bounding how much can dead-letter

Disk bounds the volume of dead-letter output; the breakers bound how many rows can dead-letter in the first place. Set one wherever a wrong schema could make most rows fail, for example when a feed can change its columns or types without notice:

error_handling:
  strategy: continue
  type_error_threshold: 0.05
  dlq:
    path: ./output/errors.csv
    max_rate: 0.05
  • dlq.max_rate (E315) and dlq.per_source.<name>.max_rate (E316) stop the run when the fraction of dead-lettered rows crosses the ceiling, once min_records rows have been read. The numerator counts dead-letter rows, collateral rows included, whether or not they have a DLQ file to go to. A source row counts once for each failure it took part in, so a row that failed twice counts twice, with or without a correlation key. Under dlq_granularity: document, a row that failed twice at the Source, in a Transform or in a Route is written once and counts once (see Document-level DLQ).
  • type_error_threshold (E368) stops the run when the fraction of declared source-type failures crosses the threshold. It catches a schema mismatch at the Source, before the failing rows reach later stages.

A rate ceiling is checked each time a dead letter is counted, so the row that crosses it is itself counted and written before the run stops. A stopped run exits with code 3 and publishes nothing. No breaker is set by default.

DLQ columns

Every DLQ record includes these metadata columns:

ColumnDescription
_cxl_dlq_idUUID v7 (time-ordered unique identifier), unique to the row. It is taken together with _cxl_dlq_timestamp, so ids order the same way as timestamps.
_cxl_dlq_trigger_idThe _cxl_dlq_id of the trigger row whose failure produced this row. That trigger row is always written in the same run. A trigger row points to itself, so _cxl_dlq_trigger is true exactly when this value equals _cxl_dlq_id; every row one failure produced carries the same value. See Pairing the rows one failure produced.
_cxl_dlq_timestampRFC 3339 timestamp of when the failure was observed, not of when the row was written. A collateral row (correlated, document_rejected) carries the time its correlation group or document was condemned, except that a rejected document’s other failing records carry the time each one failed. A group_size_exceeded row carries the time its group went over max_group_buffer.
_cxl_dlq_source_fileInput filename carried by that failing record’s $source.file provenance (or <merged> when no source-file provenance exists)
_cxl_dlq_source_nameName of the Source the failing record came from (or <merged> when the record carries no Source identity)
_cxl_dlq_source_row1-based row number in the source file
_cxl_dlq_triggering_fieldThe field whose evaluation failed, when the failure names one; empty for collateral rejections
_cxl_dlq_triggering_valueThe value the failure reported, when it carries one (for example the text that failed to convert)
_cxl_dlq_stageName of the transform or aggregate node where the error occurred
_cxl_dlq_routeRoute branch name (if the error occurred after routing)
_cxl_dlq_triggertrue when the row’s own failure dead-lettered it; false when another row’s failure took it along (a correlated or document_rejected row, or a Combine build row)
_cxl_dlq_source_recordOne of the record columns rather than a metadata column: present in any file a Source rejection can reach under strategy: continue, and filled only for a record-grained E345 rejection. Contains the fixed-width line text or a JSON array of decoded CSV cells, preserving the physical row without assigning it a declared record shape.

Timestamps need not increase down a file. A failure can be held before its row is written, for example in a correlation group that commits later, or by a stage listed under How DLQ output is written, so its row can follow rows observed after it. Sort on _cxl_dlq_timestamp or _cxl_dlq_id to read failures in the order they were observed.

When include_reason: true is set, two additional columns appear:

ColumnDescription
_cxl_dlq_error_categoryMachine-readable error classification
_cxl_dlq_error_detailHuman-readable error description

Pairing the rows one failure produced

Two rules decide which rows a DLQ file holds:

  • One row per failure. Every failure writes its own trigger row, with its own category, detail, triggering field and value, stage and route. A row that fails twice, on two Route branches or against two Combine build rows, is written twice. A Combine failure also writes the build row that contributed to it, right after its trigger. Under dlq_granularity: document this rule has an exception: a rejected document’s later failing record is written as a document_rejected row paired with the document’s first failure, and a record that failed twice at the Source, in a Transform or in a Route is written once (#1316; see Document-level DLQ).
  • Each condemned row once. A correlation group or document that fails adds each of its other rows once, as a collateral of its first failure. Under dlq_granularity: document, a record that an Aggregate, a Combine or a Reshape failure dead-lettered is written again, as a document_rejected row, only when a Sink on another branch also received it and its document is rejected. If its document is not rejected, that Sink can publish it (see Not covered).

A correlation key never removes, merges or relabels a failure row: the same failures are written with and without a key, and the key only adds the rows a failing group condemns and decides when rows are written.

One failure can dead-letter several rows:

  • a Combine body that fails writes the driver row and its matched build row, each under its own Source’s _cxl_dlq_source_name and _cxl_dlq_source_row, whichever join strategy ran. The build row is written once per failure, so a build record that two failing drivers matched is written twice, once after each driver’s row, and a driver that fails against three build rows is written three times, each copy followed by its own build row;
  • a failing row in a correlation group takes the rest of its group with it as correlated rows;
  • a group larger than max_group_buffer writes a group_size_exceeded row and the rest of its group as correlated rows (the group’s own failures keep their own trigger rows, see below);
  • under dlq_granularity: document, a failing record rejects the rest of its document as document_rejected rows.

Every row carries _cxl_dlq_trigger_id, the _cxl_dlq_id of the trigger row whose failure produced it, and that trigger row is always in the run’s output. A trigger row points to itself, so _cxl_dlq_trigger is true exactly when _cxl_dlq_trigger_id equals _cxl_dlq_id. Group by _cxl_dlq_trigger_id to see everything one failure took with it. In this excerpt (other columns omitted), the first two rows are a Combine driver row and its build row, and the last two are a correlation trigger and one of its collaterals:

_cxl_dlq_id,_cxl_dlq_trigger_id,_cxl_dlq_source_name,_cxl_dlq_error_category,_cxl_dlq_trigger
01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,orders,combine_output_row,true
01928f3a-6c10-7b22-9f07-51e6a8b4c302,01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,rates,combine_output_row,false
01928f3a-6c14-7c03-b2d8-0a9e7f615203,01928f3a-6c14-7c03-b2d8-0a9e7f615203,employees,type_coercion_failure,true
01928f3a-6c19-7d40-8c11-6e2b90d3f404,01928f3a-6c14-7c03-b2d8-0a9e7f615203,employees,correlated,false

A trigger whose failure wrote nothing else carries its own id and shares it with no other row. Every row has its own _cxl_dlq_id, so the value only repeats across the rows of one failure.

When a correlation group holds several failing rows, each failing row is a trigger and keeps its own id as its trigger id. The group’s correlated rows carry the trigger id of the group’s first failing row, the one whose error their _cxl_dlq_error_detail quotes. A rejected document has one trigger, its first failing record; its other records, including any that failed after it, carry that trigger’s id.

A group larger than max_group_buffer can hold failing rows too. Each of them is still written as its own trigger, with its own category and id, and before the rest of the group. The group’s other rows follow under one group_size_exceeded trigger, the first of them, and the remaining ones are correlated rows carrying its id. Each failure is written once, as its own trigger; a row that failed on one Route branch and reached a Sink on another is written as its own failure, and never again as correlated or as the overflow trigger. A group whose rows all failed writes no group_size_exceeded row, because the overflow took nothing with it that had not already failed.

The rows of one failure can land in different DLQ files: with per_source paths, a Combine build row goes to its own Source’s file while its driver row goes to the driver’s.

How the DLQ columns are chosen

Each DLQ file’s header is fixed when the pipeline compiles, before any record is read. It does not depend on which records failed, or on which stages they failed in: every time a pipeline writes a given DLQ file, that file has the same columns in the same order.

A header starts with the _cxl_dlq_* metadata columns, always in this order: _cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file, _cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field, _cxl_dlq_triggering_value, then _cxl_dlq_error_category and _cxl_dlq_error_detail when include_reason is on, then _cxl_dlq_stage, _cxl_dlq_route and _cxl_dlq_trigger.

With include_source_row on, the record columns follow. They come from every record shape that can reach that file:

  • the declared columns of each Source, and _cxl_dlq_source_record, when Source rejections are dead-lettered (strategy: continue);
  • the shape of the records entering each Transform, Route, Reshape, Aggregate, Combine and Sink that can dead-letter a record under the pipeline’s strategy (stages inside a composition count at the composition’s place in the pipeline);
  • the output columns of an Aggregate without group_by under strategy: continue, which go to the pipeline-wide file.

A shape is added to the file of each Source whose records can carry it: that Source’s per_source.<name>.path when it has one, otherwise the pipeline-wide path. A shape that carries no Source identity, such as the output of a Combine using match: first or match: all, is added to the pipeline-wide file.

The shapes then combine as follows:

  • One shape keeps its natural column order.
  • Several shapes give their first-seen union in plan order: each column appears once, at the position where it was first seen. Plan order is the node order --explain prints, which need not match the order the nodes are written in the YAML.
  • A column a record does not carry is an empty cell in that record’s row.
  • Engine-only sidecar columns ($widened and the $source.* stamps) are never written. Correlation-key columns ($ck.*) are.

For example, this pipeline has two Sources with different columns, one pipeline-wide DLQ file, and a Transform on each Source that can fail:

error_handling:
  strategy: continue
  dlq:
    path: rejects.csv
nodes:
- type: source
  name: orders
  config:
    schema:
      - { name: order_id, type: int }
      - { name: amount, type: int }
    # ...
- type: source
  name: refunds
  config:
    schema:
      - { name: refund_id, type: int }
      - { name: order_id, type: int }
      - { name: amount, type: int }
    # ...
# one Transform on each Source, then a Sink on each Transform

In this plan refunds comes before orders, so the record columns of rejects.csv are:

refund_id,order_id,amount,_cxl_dlq_source_record

An orders row writes an empty refund_id cell. order_id and amount appear once, although both Sources declare them. _cxl_dlq_source_record is in the header because a Source rejection can reach the file under continue. It is empty for these rows.

To see every DLQ file’s columns before a run, use clinker run pipeline.yaml --explain: its === Dead-Letter Output === section lists each file, the Sources routed to it and its full header (see Explain Plans).

Source-file provenance (_cxl_dlq_source_file) is read from each record, so one file’s path is never reused for a later row.

Error categories

The _cxl_dlq_error_category column contains one of these values:

CategoryDescription
missing_required_fieldA required field is absent from the record
type_coercion_failureA value could not be converted to the expected type
required_field_conversion_failureA required field exists but its value cannot be converted
nan_in_output_fieldA computation produced NaN
aggregate_type_errorAn aggregate function received an incompatible type
validation_failureA declarative validation check failed
aggregate_finalizeAn aggregate function failed during finalization: an integer sum outside the 64-bit range, a decimal total or quotient outside the decimal range, a weighted_avg whose weights total zero or whose row product is out of range, or a group holding both decimals and floats. The reason names the Aggregate and the emit, and gives the fix. Under continue the failed group goes to the dead-letter output. An aggregate in an Envelope footer: stops the run under every strategy.
correlatedA non-failing record was DLQ’d as collateral because another record in its correlation group failed
group_size_exceededA correlation-key group exceeded the configured max_group_buffer limit
document_rejectedA non-failing record was DLQ’d as collateral because another record in its document failed under a source’s dlq_granularity: document policy
late_recordA record arrived at a time-windowed aggregate after its event-time window had already closed
expansion_limit_exceededPer-input fan-out exceeded its authored ceiling. Transform max_expansion rejects before body rows emit; Source max_output_rows_per_input emits exactly its ceiling, then DLQs the original input on the first attempted row above it. Neither is silent truncation.
combine_output_rowA Combine output-stage eval failed for one driver row (probe-key or on_miss: null_fields body) or for one matched pair (residual or matched body). A failing residual is neither a match nor a miss: on_miss never fires for its driver, under match: all the driver’s other matches are still evaluated and emitted, under match: first a failure on the deciding candidate is the driver’s only result and a failure after it is never written, and under match: collect the driver writes no row; the entry carries the contributing-build lineage and rewinds both the driver and matched build source’s rollback cursor. The driver row and the matched build row each report their own Source in _cxl_dlq_source_name and their own row in _cxl_dlq_source_row, whichever join strategy ran. Routed to the DLQ under continue across every Combine join mode; fail_fast propagates the eval error
structural_validationA structural source rule failed: an envelope trailer’s declared count did not match its streamed body, a multi-record body appeared after its closing trailer, or a record type discriminator was unknown. Under dlq_granularity: document, the root cause has trigger: true and every already-streamed record of that file is document_rejected collateral. Under record-grained continue, E345 instead emits only the unknown row with _cxl_dlq_source_record.

Advanced options

Type error threshold

Abort the pipeline if the fraction of declared source-type failures exceeds a threshold:

  type_error_threshold: 0.05    # Abort if >5% of records fail

The cumulative ratio is:

declared source-type failures / decoded source rows observed

The rejected row appears once in both numerator and denominator. The same typed-error event is used for strategy routing, DLQ accounting, and this circuit breaker; unrelated validation, structural-document, and collateral DLQ entries do not enter the numerator. Equality is allowed: a threshold of 0.05 stops only when the ratio is strictly greater than 5%. 0.0 stops on the first type failure, while 1.0 never trips. Values must be finite and in [0.0, 1.0].

Correlation key

Declare correlation_key on the contributing Source’s config: block, not on error_handling:. Group DLQ rejections by a key field. When any record in a correlation group fails, records from the failing source’s contribution to that group are routed to the DLQ:

# Inside a Source's config:
correlation_key: order_id

For compound keys:

# Inside a Source's config:
correlation_key: [order_id, customer_id]

This is useful for transactional data where partial processing of a group is worse than rejecting the entire group. For example, if one line item in an order fails validation, you may want to reject the entire order.

Under multi-source ingest, the collateral fan-out narrows to the failing source: a src_b trigger does NOT DLQ records from src_a that share the same correlation key. Single-source pipelines see bit-identical behavior to today’s pipeline-wide collateral DLQ. See Per-source rollback narrowing for the full semantic and the two documented exceptions (max_group_buffer overflow and Combine output failures).

When a Combine output row fails under a correlation key, the failing driver row is the trigger of the driver’s correlation group. The dead letter for the matched build record is held with that group as a collateral (_cxl_dlq_trigger: false, category combine_output_row), written right after its driver’s row with its driver’s _cxl_dlq_trigger_id, and written or rolled back exactly when that group is. It never condemns the build record’s own correlation group, so another driver that matched the same build record keeps its output unless its own group failed. Each failure gets its own copy of the build row, paired with its own driver row: two failing drivers of one group each get one, and a driver that fails against several build rows is written once per failure, each copy followed by the build row of that failure. The held failure keeps its triggering field and value, as it does without a key.

Because every failure is written, a correlation key does not lower dlq_count, records_dlq or the max_rate numerators: they equal the counts the same failures give without a key, plus the rows the failing groups condemn.

For the full lifecycle and per-operator semantics (route, merge, aggregate, combine), see Correlation Keys.

Max group buffer

Limit the number of records buffered per correlation group:

  max_group_buffer: 100000     # Default: 100,000

A group that goes over this limit is dead-lettered whole. Its failing rows are written as their own triggers; its other rows are written under one group_size_exceeded row, as correlated rows. The limit counts held entries, not distinct rows: see Correlation Keys.

Document-level DLQ

By default a record failure dead-letters only that record (dlq_granularity: record). A source can instead reject the entire document any record of which fails, by declaring the granularity per source:

nodes:
  - type: source
    name: claims
    config:
      name: claims
      type: x12
      glob: ./claims/*.edi
      schema: [{ name: seg_id, type: string }]
      dlq_granularity: document   # record (default) | document

Under dlq_granularity: document and the continue strategy, a document is rejected when one of its records fails at the Source (against its declared type, or a structural rule), in a Transform, or in a Route. A record that fails in an Aggregate, a Combine or a Reshape is dead-lettered as its own row, as under record granularity, and does not reject its document (#1232; see Not covered). When a document is rejected:

  • the failing record becomes the root-cause DLQ entry (_cxl_dlq_trigger = true, carrying its original error category);
  • every other record of the document that reaches a Sink becomes a collateral entry (_cxl_dlq_trigger = false, category document_rejected); a record dropped before any Sink is not written;
  • every other record of the document that fails at the Source, in a Transform or in a Route is also a document_rejected collateral, written right after the root cause in the order the records failed. These records are stamped (_cxl_dlq_id, _cxl_dlq_timestamp) when they fail; the records a Sink held are stamped when the Sink rejects the document. Every collateral carries the root cause’s id in _cxl_dlq_trigger_id;
  • no record read from that document is written by any Sink, however many Sinks read it. Rows derived from the document’s records can still reach a Sink; see Not covered.

Clean documents in the same run stream through untouched, and records from sibling sources still on the default record granularity keep per-record semantics — the policy is per source.

Sinks run last. Every Sink runs after every other node, so a document’s verdict is final before any Sink writes one of its records. --explain lists the Sinks last. Dead-letter rows that other nodes write directly come before the rows a Sink writes.

Several Sinks. Each source row of a rejected document appears in the DLQ once, however many Sinks reached it. A row that failed is written as it was when it failed, ahead of any Sink’s copy of it. Any other row is written as it was held by the first Sink, in run order, that held it; a row that reached only a later Sink (a Route sent it there, say) is written by that Sink. Rows are matched by source row, so of the records one emit each makes from a source row, one is written. Every collateral, whichever Sink writes it, names the document’s root-cause entry in _cxl_dlq_trigger_id.

There is one exception, described under Not covered: a record that an Aggregate, a Combine or a Reshape failure dead-lettered (for a Combine, the driver row and the build row it matched) is written for that failure. It is written again, as a document_rejected row, only when a Sink on another branch also received it and its document is rejected; if its document is not rejected, that Sink can publish it (#1232). The rows a Combine or an Aggregate writes for a rejected document, and the rows that pass through a Reshape, are not held back at all.

This is the document-shaped analogue of correlation keys: use it when partial processing of a document (an EDI interchange, a batch file with a header/trailer) is worse than rejecting the whole document. Unlike correlation keys, which group across files by a key value, document-level DLQ scopes rejection to a single document’s records.

Document grain. The document is the outermost level — the source file. For a flat format (CSV, JSON, plain XML) each input file is one document. For a nested-envelope format (an X12 ISA → GS → ST interchange, an EDIFACT UNB → UNG → UNH) the document is the whole interchange / file, not an inner functional group or transaction set: a failure anywhere in the interchange rejects the entire interchange, including the transaction sets that validated cleanly. Reject the inner-level grain instead by partitioning the input so each interchange is its own file is not currently offered — the grain is fixed at the file.

DLQ rate. Each dead-letter row — the trigger and each collateral — counts once toward the configured DLQ max_rate, matching the correlated-collateral precedent, however many Sinks held it; a record written twice under the exception in Several Sinks counts twice. A rejected 1000-record document contributes 1000 when every one of its records reaches a Sink. It does not affect type_error_threshold, whose numerator contains only declared source-type failures.

dlq_count never counts a row that ok_count also counts, because no Sink writes a record read from a rejected document, with these exceptions, each described under Not covered:

  • a record that an Aggregate, a Combine or a Reshape failure dead-lettered, which a Sink on another branch can still write when its document is not rejected (#1232);
  • a rejected document’s rows that a Combine or an Aggregate writes, or that pass through a Reshape, which reach a Sink while the document’s records are dead-lettered;
  • a record whose joined or aggregated row fails in a node after the Combine or Aggregate: that failure is written while the record itself can still reach a Sink on another branch.

Memory. The engine buffers each open document’s records until its boundary, then flushes the document clean to the sink or rejects it and drops the buffer. Peak memory scales with the concurrently-open documents, not the total input; a single very large document spills its buffer to disk under the run’s memory budget rather than holding everything in RAM. See Streaming vs blocking for the spill model.

Sink restriction. Document-level DLQ flushes each whole document to a single output writer, so it cannot be combined with a per-source-file Sink (a {source_file} / {source_path} path template over a multi-file source). The two are rejected together at compile time (E343); use a single output path, or set dlq_granularity: record if per-file output is the requirement.

Strategy requirement. dlq_granularity: document requires error_handling.strategy: continue. It is incompatible with the default fail_fast: document-level dead-lettering keeps the run going past a bad document, which contradicts fail-fast’s abort-on-first-error. The combination is rejected at compile time (E344) — set strategy: continue to dead-letter bad documents, or keep fail_fast with the default dlq_granularity: record.

Correlation restriction. Document-level and correlation-key rejection are alternative atomic-disposition models. A document is keyed by its source file; a correlation group can span files and is keyed by authored field values. The engine does not define precedence or a combined writer boundary for those two populations, so a pipeline containing both dlq_granularity: document and any correlation_key is rejected at compile time (E370). Remove every correlation_key to keep document rejection, or set dlq_granularity: record to keep correlation rejection.

Composition restriction. A composition body may not declare a Sink when any Source uses dlq_granularity: document (E378): a body Sink runs inside its composition, where it cannot be held back until every document’s verdict is final. Declare the Sink at pipeline level instead. Where moving it through a new composition output port works today, the diagnostic prints that move ready to paste. Otherwise it names clinker explain --code E378, which shows how to declare the Sink’s work at pipeline level.

Not covered. Document rejection does not yet reach these rows:

  • A row failure inside an Aggregate, a Combine or a Reshape is dead-lettered per record and does not reject its document (#1232). When a Sink on another branch also received that record, it is written again as a document_rejected row if another failure rejects its document, and that Sink can publish it if none does.
  • The rows a Combine or an Aggregate writes are not held back by their document’s verdict, so a rejected document’s joined or aggregated rows still reach a Sink. A failure on such a row, in any node after the Combine or Aggregate, dead-letters only that row and does not reject its document, so the document’s other records, and that row’s own source record on another branch, can still be published (#1317).
  • A Reshape’s output rows do not keep their input row’s document, so a rejected document’s rows that pass through a Reshape still reach a Sink.
  • In a pipeline where any Source declares dlq_granularity: document, a CSV join_values collision with on_conflict: error at a Sink fails the run rather than dead-lettering the record (#933).

A document is identified by its file path, so two Sources reading the same file share one verdict.

Spilling stages. Document identity survives memory pressure end to end. The per-document buffer identifies each document before buffering and spills under the memory budget, and a Sort between the source and the output preserves each record’s document context — including the source file the grain keys on — across its own spill round-trip. A document whose records pass through a spilling Sort is therefore still grouped and rejected as one document under memory pressure, exactly as it would be in memory. The rows a hash Aggregate or a grace-hash Combine writes are not held back by their document’s verdict, spilled or not; see Not covered.

Malformed envelopes (structural validation)

Envelope formats carry their own structural-integrity claims: an X12 interchange declares a segment count in each SE/GE/IEA trailer, EDIFACT in each UNT/UNZ, HL7 batch/file in each BTS/FTS, and a multi-record flat file’s trailer record declares a body count via its structure: constraint. When the declared count does not match the body the reader actually streamed, the file is structurally invalid. A multi-record flat file can also break a non-count structural rule — a line whose record-type discriminator matches no declared records: entry (E345), or a body record appearing after the trailer that closes the document; these are classified separately from a count mismatch but carry the same disposition.

Under dlq_granularity: document, such a structural failure dead-letters the whole source file to the DLQ rather than aborting the run:

  • the file’s records dead-letter as one structural_validation root-cause entry (_cxl_dlq_trigger = true) plus a document_rejected collateral for every other already-streamed record of the file;
  • no record of the malformed file reaches the success sink.
nodes:
  - type: source
    name: claims
    config:
      name: claims
      type: x12
      glob: ./claims/*.edi
      schema: [{ name: seg_id, type: string }]
      dlq_granularity: document   # reuse the document opt-in — no separate config

The opt-in is the same dlq_granularity: document knob that governs per-record document rejection above; there is no separate validation: block. A malformed envelope is simply one more reason a source under the document policy condemns a whole document. E345 is the one structural class that also has a record-grained recovery: under strategy: continue and the default dlq_granularity: record, only the unknown-tag row is dead-lettered and the reader continues at the next physical row. The DLQ row exposes the unknown tag as record_type and the unguessed decoded input as _cxl_dlq_source_record (fixed-width line text, or a JSON array of decoded CSV cells).

Honest timing — rejected at the sink boundary, not before the first record. The trailer that carries the count arrives at the end of the file, after every body record it counts has already streamed through the DAG. Clinker is a bounded-memory streaming engine — it does not buffer the whole file up front to pre-validate it (that would defeat the streaming model). So the count mismatch is detected mid-stream, the file is marked failed, and the document-level DLQ buffer rejects every already-streamed record of the file at its close. The user-visible outcome is the same — no record of a malformed envelope is ever written to the output — but the rejection lands at the sink boundary, not literally before the file’s first record streams.

Grain — the whole file. An SE-level mismatch (one transaction set inside a larger interchange) rejects the entire interchange / file, not just that one transaction set, because the document grain is the outermost source file (see Document grain above). Split the input so each interchange is its own file if you need finer rejection.

Multiple files keep flowing. When a glob / paths source matches several files and one is malformed, only that file dead-letters — ingestion continues to the remaining files, so the clean files after a bad one still reach the sink. (This is unlike the default record granularity, where a count mismatch aborts the whole run and no file’s records are written.) Dead-lettering one malformed file never silently drops the good files around it.

Record-grained E345 is narrow. Under the default dlq_granularity: record, strategy: continue can recover only from an unknown multi-record discriminator because the reader has consumed exactly one bounded physical row and can resume unambiguously. Trailer-count mismatches and a body record after a document-closing trailer still abort at record granularity: neither belongs to one independently recoverable row. Genuine corruption (a truncated stream, a bad delimiter, a control-number echo mismatch, a segment after an X12/EDIFACT/HL7 envelope trailer) always aborts, even under the document opt-in. Under fail_fast, E345 also aborts at the offending line.

Cryptographic integrity (checksums / signatures) is not yet validated. Envelope formats can also carry a SHA-256 body hash, a JWS-signed JSON payload, or an XML Signature. Clinker extracts these envelope sections but does not yet verify them. Tracked for a future release.

Exit codes

CodeMeaning
0Pipeline completed successfully, no errors
1Configuration error – the pipeline never started
2Pipeline completed, but DLQ entries were produced
3Data error halted the run: a fail_fast evaluation/accumulator failure, or the DLQ-rate ceiling
4I/O, format, or spill failure

Exit code 2 is not a failure – it means the pipeline ran to completion and handled errors according to the configured strategy. Check the DLQ file for details. See Exit Codes & Error Diagnosis for the full reference and the orchestrator retry policy.

Complete example

pipeline:
  name: order_processing
  memory: { limit: "512M" }

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./data/orders.csv"
      correlation_key: order_id
      schema:
        - { name: order_id, type: int }
        - { name: customer_id, type: int }
        - { name: amount, type: float }
        - { name: email, type: string }

  - type: transform
    name: validate_orders
    input: orders
    config:
      cxl: |
        emit order_id = order_id
        emit customer_id = customer_id
        emit amount = amount
        emit email = email
      validations:
        - field: email
          check: "not_empty"
          severity: error
          message: "Customer email is required"
        - check: "amount > 0"
          severity: error
          message: "Order amount must be positive"

  - type: sink
    name: valid_orders
    input: validate_orders
    config:
      name: valid_orders
      type: csv
      path: "./output/valid_orders.csv"

error_handling:
  strategy: continue
  dlq:
    path: "./output/rejected_orders.csv"
    include_reason: true
    include_source_row: true
  type_error_threshold: 0.10

Nodes

A Clinker pipeline is a single flat nodes: list. Every entry carries a type: discriminator that selects a node kind — the unified node taxonomy. There is no separate “join section” or “filter section”: records flow through one homogeneous graph of typed nodes, wired together by input: / inputs:. This part documents the ten runtime node kinds; an eleventh, Composition, is a call-site that inlines a reusable sub-pipeline and is covered under Pipelines.

The pages in this part are ordered the way data flows through a DAG — a record enters at a Source, fans through the record-level and combining nodes, and leaves at a Sink:

NodeRoleArityStreaming vs blocking
SourceReads records from a file or network cursor; the entry point.0 → 1Streaming
TransformRecord-level CXL projection, filter, and lookup.1 → 0..N recordsStreaming
RoutePredicate-based fan-out into named branches.1 → NStreaming
MergeStreamwise concatenation of inputs that share a schema.N → 1Streaming
CombineN-ary record combining with mixed predicates (equi + range + arbitrary CXL).N → 1Blocking (build side)
AggregateGrouped or windowed reduction.1 → 1Blocking (or streaming when sorted)
ReshapeGroup-aware mutation and synthesis of records.1 input → 1 output portBlocking
CullPer-correlation-group removal on a group-level predicate, with a removed_to side-output port.1 → 2Blocking
EnvelopeFrames a body stream into per-document documents; a composable framing stage.1 → 1Streaming
SinkWrites records to an output destination; the exit point.1 → 0Streaming

Streaming vs blocking

Stateless nodes (Transform, Route, Merge, the Combine probe side, Sink) evaluate records one at a time without accumulating per-record state. Blocking nodes (Aggregate, sort, the grace-hash Combine build side) accumulate state inside the RSS budget and spill to disk rather than OOM the process. The Streaming vs. Blocking Stages page in the Operations Guide is the full memory model.

Wiring and naming

Every node needs a unique name: (no dots — the dot is reserved for port syntax). Single-input nodes use input:; Merge and Combine use inputs:. Route branches are consumed downstream as route_name.port. The Pipeline YAML Structure page covers the full wiring grammar, optional fields (description:, _notes:), and strict parsing rules.

Source Nodes

Source nodes read data from files and are the entry points of every pipeline. They have no input: field – they produce records, they do not consume them.

Basic structure

- type: source
  name: customers
  config:
    name: customers
    type: csv
    path: "./data/customers.csv"
    schema:
      - { name: customer_id, type: int }
      - { name: name, type: string }
      - { name: email, type: string }
      - { name: status, type: string }
      - { name: amount, type: float }

Schema declaration

The schema: field is required on every source node. Runtime ingestion does not guess types from data: declare each column’s name and CXL type explicitly. This schema drives compile-time type checking across the entire pipeline.

Each entry is a { name, type } pair:

schema:
  - { name: employee_id, type: string }
  - { name: salary, type: int }
  - { name: hired_at, type: date_time }
  - { name: is_active, type: bool }
  - { name: notes, type: { nullable: string } }

Available types

TypeDescription
nullThe null value only
stringUTF-8 text
int64-bit signed integer
float64-bit IEEE 754 floating point
decimalExact fixed-point number, subject to the declared precision and scale
boolBoolean (true / false)
dateCalendar date
date_timeDate with time component
arrayOrdered sequence of values
mapString-keyed object
anyUnknown type – field used in type-agnostic contexts
{ nullable: T }Nullable wrapper around a concrete inner type, for example { nullable: int }

Check the schema through the planner

The schema that governs execution is the one accepted by clinker-plan while compiling the complete pipeline. Run this after changing a source schema:

clinker run pipeline.yaml --explain text

Exit code 0 means the planner parsed the canonical YAML, resolved schema references and overlays, bound the pipeline, type-checked CXL, and produced a compiled plan. It does not prove that input bytes satisfy the declarations or that the resulting output is correct; execute representative input against an isolated destination and compare the result before a production run.

Workspace tooling may also present advisory schema analysis. That analysis can help find likely authoring mistakes, but it cannot admit or reject a pipeline. Its analyzed, partial, skipped, and failed statuses describe only how much the advisory model inspected. See Validation and Admission for the status meanings and known limits.

Without a column format:, date_time accepts the default offset-free forms and RFC 3339 timestamps with Z or an explicit numeric offset, such as 2026-01-31T08:27:00Z and 2026-01-31T10:57:00+02:30. Zoned timestamps are normalized to UTC before entering Clinker’s timezone-free date_time representation. Parsing is exact: surrounding whitespace, malformed calendar values, and out-of-range timestamps are rejected rather than trimmed, guessed, overflowed, or rounded. A column-level strftime format: is exclusive: only that authored format is tried.

A source column’s declared type must be concrete: numeric — the inference-only int | float union CXL resolves during type unification — is not a valid source column type. Declaring one is rejected at compile with E158; declare int or float explicitly.

clinker guess pipeline.yaml provides an authoring-only preview for columns explicitly marked numeric. It proposes concrete int or float declarations; it does not infer arbitrary schemas or change the runtime admission rules. The default preview is bounded and read-only. --check exhausts the admitted manifest; --write can publish one guarded, unambiguous inline edit. Resolve all numeric declarations to concrete types before executing the pipeline. See Numeric schema authoring and the guess command.

long_unique — storage hint for high-cardinality text

A string column may carry an optional long_unique: true flag. It is an advisory, opt-in hint, not a type change: it tells Clinker the column’s values are long and effectively unique — never repeated across records — so the run uses less memory for that column. Typical candidates are UUIDs rendered as text, street addresses, and free-text comment or note fields.

schema:
  - { name: ticket_id,  type: string, long_unique: true }   # 36-char UUID
  - { name: notes,      type: string, long_unique: true }   # free text
  - { name: department, type: string }                      # low-cardinality, default

The flag lowers memory use only. A value’s content and its comparison, grouping, join, sort, and output behavior are all unchanged — a long_unique value behaves identically to the same text in any other column. Omitting the flag (the common case) leaves the default behavior untouched. Set it only when you know a column is genuinely high-cardinality free text; on a column whose values repeat, leave it off.

source_name — read a differently-named physical column

A column may carry an optional source_name naming the physical input column it reads from, when that differs from the exposed name. The reader matches input fields by physical name and re-labels the value under name, so downstream CXL and the output see name carrying the physical column’s data.

schema:
  # read the physical `cust_id` column, expose it downstream as `customer_id`
  - { name: customer_id, type: string, source_name: cust_id }

Omitting source_name (the common case) reads the input field whose key equals name, unchanged from before. A channel schema patch’s rename op sets this alias automatically (see Channels).

Transport vs format

A source declaration has two independent layers:

  • Transport (transport:) selects where the records come from. Two transports exist: file — read bytes from the filesystem, resolved through one of the file matchers (path / glob / regex / paths) — and rest — pull records from a paginated HTTP endpoint under a hard page/record cap (see Network Sources (REST)). transport: is optional and defaults to file, so a source that omits it reads from disk exactly as before. rest needs the rest capability, which the released binary has; a build compiled without it refuses a rest source at validation with E223 rather than running the pipeline short one input (see Optional capabilities).
  • Format (type:) selects how the bytes decode into records: csv, json, xml, fixed_width, edifact, x12, hl7, swift.
- type: source
  name: orders
  config:
    name: orders
    transport: file        # optional; this is the default
    type: csv              # the on-disk format
    path: "./data/orders.csv"
    schema:
      - { name: order_id, type: int }

A file transport requires exactly one file matcher (path, glob, regex, or paths). Declaring none fails validation with E211; declaring more than one fails with E210. Both are reported at config-load time, before any file is opened.

Choosing files

Interactive companion: the file discovery explainer runs these steps on a sample folder as you change the settings.

Before reading, a file Source builds its list of files in a fixed order:

  1. Matcher. Paths are relative to the pipeline file’s folder.
    • path: names one file and paths: a list. A file that does not exist stops the run with E216, whatever on_no_match says.
    • glob: matches names inside one folder: * does not cross a /. Use ** to include subfolders (./data/**/*.csv). A leading dot is not special, so *.csv also matches .draft.csv. An invalid pattern is E212.
    • regex: searches the pipeline’s folder and, by default, every subfolder, and matches anywhere in each file’s path. That path begins with the folder part of the pipeline path as you typed it (clinker run pipelines/orders.yaml gives pipelines/data/orders_2024-01.csv), so anchor the end of the path, for example data/orders_\d{4}-\d{2}\.csv$. An unanchored pattern can pick up earlier outputs too. An invalid pattern is E213.
  2. exclude: – a list of glob patterns; a file whose name or full path matches any of them is dropped.
  3. Anything that is not a regular file is dropped.
  4. min_size: / max_size: – decimal units: 1KB is 1000 bytes, also B, MB, GB.
  5. modified_after: / modified_before: – a duration back from the time the pipeline is loaded (30s, 15m, 2h, 3d) or an RFC 3339 timestamp (2024-03-01T00:00:00Z).
  6. files.sort_by: name (default; the full path), created or modified, with files.sort_order: asc (default) or desc. The files of a paths: list are sorted too: the order they are written in is not the reading order.
  7. files.take_first: or files.take_last: keeps that many files from the sorted list (setting both is E218). With the default ascending name sort, take_first: 5 keeps the five earliest names.
  8. files.on_no_match: – when nothing is left: error (default, E216), warn (log a warning and produce no rows), or skip (produce no rows quietly).

files.recursive: controls whether regex: searches subfolders (true by default). It has no effect on glob:, which searches subfolders only where the pattern has **.

- type: source
  name: orders
  config:
    name: orders
    type: csv
    glob: ./data/orders_*.csv
    exclude: ["*_partial.csv"]
    modified_after: 30d
    files:
      sort_by: name
      sort_order: desc
      take_first: 3        # the three latest names
      on_no_match: warn
    schema:
      - { name: order_id, type: string }

With glob:, regex: or paths:, each file is read as its own document: a declared sort_order is checked per file, and an Aggregate rolls up per file (see Sort order and One document per file).

Format types

The type: field inside config: selects the on-disk format. Each format has its own reference page covering its options and decoding model:

type:FormatReference
csvDelimited text (RFC 4180)CSV Format
jsonArray / NDJSON / wrapper objectJSON Format
xmlElement-path-selected record elementsXML Format
fixed_widthColumn-positioned legacy extractsFixed-Width Format
edifactUN/EDIFACT interchangesEDIFACT Format
x12ANSI ASC X12 interchangesX12 Format
hl7HL7 v2.x pipe-and-hat messagesHL7 v2 Format
swiftSWIFT MT (FIN) messagesSWIFT MT Format

The same schema: rules apply regardless of format: the reader maps each decoded record onto the declared schema, and undeclared input fields fall under the on_unmapped policy below.

Declared-type failures

Source types are enforced before a record reaches buffering, sorting, or any downstream node. A value that cannot satisfy its authored type rejects the whole row exactly once; Clinker never substitutes the raw string, a sentinel, or an error-derived null.

Empty input has three distinct outcomes:

  • an empty value declared as string remains the empty string;
  • an empty non-string value declared nullable(T) becomes null;
  • an empty non-string value declared as non-nullable is a type error.

Parsing does not trim whitespace, guess locale conventions, or recognize case-insensitive null sentinels. Integer overflow, invalid/out-of-range dates, decimal precision overflow, and decimal values that would need rounding to the declared scale are type errors. Accepted decimal and date values retain their declared precision.

The E126 diagnostic identifies source, file, one-based row and column, field, and declared type. Its value preview is sanitized to one line and limited to 256 rendered UTF-8 bytes; controls, bidi characters, diagnostic delimiters, backslashes, and invalid UTF-8 are explicit indivisible escape tokens. When truncated, the preview ends in one … without splitting a token or Unicode scalar and reports the original byte length. The complete original record/value is retained only in the configured DLQ. See Error Handling & DLQ for strategy and threshold behavior.

Already-decoded values obey the same declarations without hidden coercions: string admits only a string, null admits only null, and any admits every supported native value while preserving its represented value. Numeric conversions must be exact; non-finite floating-point values and decimal values that would require rounding are rejected. With multiple: true, the scalar declaration is applied to every array element. If any element fails, the whole source record is rejected and the complete original array remains available through the DLQ.

Error strategies and complete populations

fail_fast aborts on the first declared source failure. continue and best_effort require a DLQ and use the configured threshold over the complete population:

rejected records / attempted records

The run aborts only when that ratio is strictly greater than the threshold. Equality is accepted, so a threshold of 0.1 admits exactly one rejected record in a population of ten but not two. A zero threshold aborts on the first rejection; a threshold of 1.0 admits an all-rejected population for the continuing strategies.

For ordered file sources, Clinker establishes the complete attempted and rejected population before any accepted record, punctuation, downstream side effect, or output byte is released. That population is applied exactly once whether execution remains resident, spills, or fuses a downstream node. A threshold violation therefore cannot leave a committed prefix of output.

on_unmapped — undeclared input fields

The per-source on_unmapped policy decides what to do with input fields the source’s schema: block does not name. Three modes — auto_widen (default), drop, reject:

- type: source
  name: orders
  config:
    name: orders
    type: csv
    path: "./data/orders.csv"
    on_unmapped:
      mode: auto_widen     # default; other values: drop, reject
    schema:
      - { name: order_id, type: string }
      - { name: amount, type: float }

See Auto-Widen & Schema Drift for the full specification: how undeclared columns flow through each downstream node type, the include_unmapped Sink flag, E315 merge-policy mismatch, and fixed-width behavior.

Sort order

If each physical input file is pre-sorted, declare the record order so the planner can admit order-dependent strategies such as streaming aggregation:

- type: source
  name: sorted_transactions
  config:
    name: sorted_transactions
    type: csv
    path: "./data/transactions_sorted.csv"
    schema:
      - { name: account_id, type: string }
      - { name: txn_date, type: date }
      - { name: amount, type: float }
    sort_order:
      - { field: "account_id", order: asc }
      - { field: "txn_date", order: asc }

Clinker binds these fields to the declared source schema, compares the typed values by the same rule every sort uses (see How values are ordered), and verifies each physical file independently before any record from that file reaches an order-dependent consumer. The declaration never means that a multi-file source is globally sorted: the last key in one file is not compared with the first key in the next file.

on_unsorted controls the result of the first adjacent inversion:

ValueBehavior
warn (default)Stably repair the complete physical file with the shared bounded-memory sort, emit one W307 warning for that file, then release it. A file already in order emits no warning.
errorReject the physical file without releasing an unverified prefix. The diagnostic identifies the source, file, adjacent rows, and keys.
    sort_order:
      - { field: "account_id", order: asc, null_order: last }
      - { field: "txn_date", order: desc, null_order: first }
    on_unsorted: warn

Source ordering accepts null_order: first or last; drop is rejected when the pipeline is planned because verifying order must not discard source records. The error gives one fix: delete null_order: drop and add a Transform after the Source whose whole config is the line it prints, config: { cxl: "filter not <field>.is_null()" }. With the line deleted, the Source declares its null keys last; if a file’s null keys arrive first, write null_order: first instead of deleting the line. That filter needs a field CXL can name as it is: one identifier of ASCII letters, digits and _, not starting with a digit and not a CXL keyword. For any other key, such as order id, filter or a flattened Address.City, the error prints no CXL and prints a source_name line for the column instead, such as source_name: "order id". Set the column’s name to a new identifier, add that line, use the new name wherever the pipeline names the column, and plan again: the error then prints the filter on the new name. Equal authored keys retain arrival order within the selected execution path. Clinker does not add a source identity, physical filename, or canonical-row tie-breaker.

Verification stages the complete sortable file event sequence behind the run’s memory arbitrator. It uses the existing stable resident/spill sort and bounded-fan-in merge machinery when repair spills, while preserving row identity, source/file provenance, and document context. Only flat sources and sources with one sortable frame per physical file are admitted; a format whose nested or repeated framing cannot be reordered losslessly fails during planning with a correction to remove sort_order or normalize the input.

If a downstream consumer needs one global order across all files, declare sort_order on the terminal Sink. Use enough output fields to define a total business order when byte-identical output matters.

The shorthand form is also accepted – a bare string defaults to ascending:

    sort_order:
      - "account_id"
      - { field: "txn_date", order: desc }

Watermarks

An event-time watermark declares which column on the source carries each record’s event time — the wall-clock instant the event happened, distinct from when Clinker read the row. When set, Clinker takes the column on every record, subtracts the source’s delay, and uses the result to track event-time progress so downstream time windows know when to close. The delay-corrected value is also stamped on every record as $source.event_time, the column a downstream time-windowed aggregate uses to assign records to windows.

- type: source
  name: clicks
  config:
    name: clicks
    type: csv
    path: "./data/clicks.csv"
    options:
      has_header: true
    watermark:
      column: event_ts       # must be date_time or date
      delay: 5s              # bounded out-of-order tolerance
      idle_timeout: 30s      # flip partitions to idle if quiet
    schema:
      - { name: user_id, type: string }
      - { name: event_ts, type: date_time }
      - { name: amount, type: int }

Fields:

  • column (required) — the schema column whose value is each record’s event time. The column’s declared type must be date_time or date. A column: that names a field absent from schema: raises E154; a column: whose declared type is neither raises E155.

  • delay (optional duration, default unset) — bounded out-of-order tolerance. Each record’s event time is shifted earlier by delay before being folded into the watermark, so the source’s effective watermark trails its observed max event time by this amount. Mirrors Flink’s BoundedOutOfOrdernessWatermarks. Without delay, the watermark advances strictly to the observed max — a single late record routes to the DLQ.

  • idle_timeout (optional duration, default unset) — if a source stays quiet longer than this, it stops holding back downstream window-close progress, so windows keep closing when one source pauses. Unset means the source never goes idle.

Durations use the suffixes ms, s, m, h, d. ms is matched before the single-character s, so 500ms reads as 500 milliseconds, not 500 seconds with a stray m.

A pipeline whose aggregate declares time_window: must have a watermark.column on every upstream-reachable source. Without it, event-time progress can never advance and the window can never close — the planner rejects this with E156.

Multi-value fields

A field that holds more than one value is declared on the schema column, not on the pipeline:

schema:
  - { name: order_id, type: string }
  - { name: tags, type: string, multiple: true }

multiple: true says the column holds zero or more values of its declared type. Reading collects every occurrence of the field into one array — a single occurrence is still an array, so downstream code never has to branch on how many values happened to arrive. A field absent from a record has no column at all and resolves to null, exactly as any other absent column does. CXL sees the column as an array; the declared type: describes each element and drives coercion. The declaration describes the shape of the data, so it serves both directions: a writer that can encode repetition reads the same declaration.

The split_to_rows, split_values, and join_values blocks below accept a compact shorthand (a bare field name, or a mapping that omits defaults). To see the fully-materialized form the engine actually runs — every default spelled out — print the canonical config with clinker config --resolved. It rewrites only those shorthand blocks and leaves the rest of the file untouched.

Both ends of the declaration are checked at compile, so a shape the formats cannot carry fails before a run starts rather than mid-stream:

FormatAs a sourceAs an output
jsonnative — an arraynative — an array
xmlnative — repeated child elementsnot yet (issue 916)
csvdelimited cell via split_valuesdelimited cell via join_values
fixed_widthdelimited cell via split_valuesnot yet (issue 918)
edifact, x12, hl7, swiftno — repetition is positionalno — repetition is positional

A multiple: true column reaching an output that cannot encode it is E359; one on a source that cannot produce it is E361. E359 covers an output’s own schema: block too — the attribute is direction-neutral, but the remaining writers do not encode repetition yet, so declaring it on such a sink would be accepted and ignored. Run clinker explain --code E361 for the full remediation of either.

csv and fixed_width read a multiple: true column through a split_values entry. Neither wire format repeats a field, but a cell’s text may hold several values separated by a delimiter. Declare that delimiter with a split_values entry and the reader parses the cell into the array the column holds. A multiple: true column no entry covers is rejected by E361 — the reader would have no delimiter and deliver the raw cell; either add the entry or leave the column single-valued and split it in a transform (tags.split(";")). The entry is read only on a single-schema source: a multi-record source of either format runs a backend that does not consume it. On the output side, a CSV sink joins a multiple: field into one delimited cell with join_values (defaulting to ; / on_conflict: error); fixed_width output is still pending (#918).

A split_values entry also recovers a CSV cell a sink wrote under join_values on_conflict: escape or encode_json: add escape: "\\" to un-escape an escaped delimiter, or json: true to read the whole cell as an embedded JSON array.

The segment formats are a permanent no, not a pending one. Repetition there is a positional coordinate rather than a list: a repeated composite is written as two axes interleaved in one element (11:B:1^12:B:2), which a flat array cannot represent without losing the component axis. The faithful shape is one column per coordinate — for HL7, that is what options.split_fields produces, with a writer that reassembles the wire field byte-for-byte.

One record per value: split_to_rows

split_to_rows fans a record out to one record per occurrence of a repeated field. Each entry is either a bare field name or a full mapping, and the two forms mix freely in one list:

- type: source
  name: invoices
  config:
    name: invoices
    type: json
    path: "./data/invoices.json"
    schema:
      - { name: invoice_id, type: int }
      - { name: customer, type: string }
      - { name: line_item, type: string }
      - { name: line_amount, type: float }
      - { name: line_no, type: int }
    split_to_rows:
      - tags                      # shorthand: field name, all defaults
      - field: line_items         # full form
        keep_empty: true
        mode: extract
        position_column: line_no
    max_output_rows_per_input: 10000
KeyDefaultMeaning
field—The repeated field, as a flattened dotted name
keep_emptytrueWhether a record whose field is empty or absent survives
modeextractextract — the occurrence becomes the record; split — the record shape is kept
position_columnnoneColumn receiving each occurrence’s 1-based position

The field is named as it appears in the input document, not as the schema exposes it: a column declared source_name: is addressed by that source_name. The same rule applies to split_values below.

keep_empty defaults to true. A record whose field holds an empty array, or carries no such field at all, is emitted with that field unset rather than disappearing. Several widely used engines drop the record instead; a vanished row is the costliest failure mode there is, so dropping is opt-in here.

mode: extract (the default) makes the occurrence the record: its own fields are lifted out from under the field name and every field outside the group is merged onto each output. {"orders": [{"id": 1}]} yields a top-level id, and repeated <Item><name> children yield name. When lifting lands an occurrence’s field on a name an outside field already occupies, the occurrence wins — it is the record, so its own value is not shadowed by the parent it was merged with. A position_column wins over both: you named it, so a field of the same name inside or outside the occurrence gives way to the index.

mode: split preserves the record shape: the occurrence’s fields keep their dotted path (orders.id, Item.name) and each output carries exactly one occurrence.

Entries apply in declaration order, so two entries multiply. Declaring the same field twice is rejected at compile (E358), as is fanning out a field the schema also declares multiple: true — the attribute collects the occurrences into one array, the fan-out spends them one per record, and a field cannot be both. On an XML source, two entries may not name nested element groups either (Item and Item.part): that reader assigns each element to one occurrence group by document position, and a nested pair leaves the inner group’s membership ambiguous.

max_output_rows_per_input is valid only alongside a non-empty split_to_rows block and bounds that cumulative product for one original JSON object or XML record element. Omit it, or set it to 0, for no ceiling. With a positive value N, the reader emits exactly the first N records in the same declaration order shown above. If an N+1 row is attempted, the reader stops that input and reports an expansion_limit_exceeded source failure. Under error_handling.strategy: continue, the complete decoded original input representation is routed to the DLQ; the first N records remain valid output. This is not silent truncation: the run records one explicit rejection naming the field, the configured ceiling, and the exact first violating count (N+1). Under fail_fast, the run fails at that boundary.

The check is lazy: the reader never materializes or pre-counts the Cartesian product. Its cursor holds the original input, the declared occurrence lists, and one output record. This source setting is separate from a Transform’s max_expansion; when both surfaces fan out, each enforces its own per-input boundary.

A JSON source accepts a nested pair to produce a two-level expansion, but only when the outer entry declares mode: split. mode: extract lifts the occurrence’s own keys to the top level, which removes the dotted path the inner entry addresses — the inner entry would then match nothing and fan nothing out, so the pairing is rejected (E358):

    split_to_rows:
      - { field: orders, mode: split }
      - { field: orders.items, mode: split }

Several values in one cell: split_values

split_values parses a delimited cell into the several values a multiple: column holds. It takes the same bare-name-or-mapping shorthand:

    split_values:
      - tags                      # shorthand: default delimiter `;`
      - field: codes              # full form
        delimiter: "|"
    schema:
      - { name: tags,  type: string, multiple: true }
      - { name: codes, type: string, multiple: true }

The delimiter defaults to ;. A split_values field the schema does not declare multiple: true is rejected at compile (E358): splitting produces several values, and only a multi-value column can hold them. So is an entry naming a column’s exposed name when that column reads a differently-named input field — the split runs against the document’s own field names, so name the source_name.

The entry is read by the JSON and XML readers, and — on a single-schema source — by the CSV and fixed-width readers. On a multi-record CSV or fixed-width source, or on any segment format, declaring it is rejected (E358) rather than silently ignored: those readers are never handed it, so the cell would arrive unsplit with nothing to say so.

Migrating from array_paths

array_paths: was the earlier form of these declarations. It is no longer read, and a source still carrying it is rejected at compile (E360) rather than running with the fan-out silently dropped. An explode path becomes a split_to_rows: entry, a delimited cell becomes a split_values: entry, and a path kept as an array becomes multiple: true on the schema column.

mode: extract (the default) reproduces the old projection — the element’s own fields lifted to the top level of each output record. It does not reproduce the old cardinality: explode dropped a record whose array was empty, while keep_empty defaults to true here and keeps it with the element’s fields unset. Add keep_empty: false to the entry to migrate row-for-row.

Format notes

split_to_rows is honored by the JSON and XML readers, over a file path and over a rest response body alike. split_values is honored by those two and also by the single-schema CSV and fixed-width readers. Declaring a knob on a reader that is never handed it — split_to_rows on any delimited-cell or segment format, split_values on a multi-record or segment source — is rejected at compile (E358) rather than accepted and inert, and a multiple: true column no split_values entry can cover is rejected by E361.

Composition body files are gated by the same four checks as the pipeline that calls them.

JSON — the field names the key holding the array (line_items, or order.line_items for an array nested under an object). A field present but holding a single object or scalar rather than an array counts as one occurrence, and is projected exactly as a one-element array would be — many producers unwrap a lone element, and XML cannot express the difference at all, so the two readers agree on this. A field that is absent, holds an empty array, or is explicitly null has no occurrence and is governed by keep_empty: an explicit null is how many producers write “no value”, and it counts as none rather than as one. For the same reason a multiple: true column holding an explicit null stays null rather than becoming [null], so size() over it reads the same as it does for a field the document omits.

XML — the field is the repeated child element’s dotted path relative to the record element (Item, or Items.Item when nested). Repetition and absence are indistinguishable in XML, so a record with no occurrence of the element is governed by keep_empty exactly as an empty array is. See XML Format for the full rules.

Transform Nodes

Transform nodes apply CXL expressions to each record, producing new fields, filtering records, or both. They process one record at a time in streaming fashion with constant memory overhead.

Basic structure

- type: transform
  name: enrich
  input: customers
  config:
    cxl: |
      emit full_name = first_name.concat(" ", last_name)
      emit tier = if lifetime_value >= 10000 then "gold" else "standard"
      filter status == "active"

The cxl: field is required and contains a CXL program. The three core CXL statements for transforms are:

  • emit – adds a field to the record, or replaces the field of the same name. The record’s other input fields are carried through to downstream nodes without being emitted; a Sink with include_unmapped: false is the one place that narrows them (see Sink Nodes).
  • filter – drops records that do not match the boolean condition.
  • let – binds a local variable for use in subsequent expressions (not emitted).
    cxl: |
      let margin = revenue - cost
      emit product_id = product_id
      emit margin = margin
      emit margin_pct = if revenue > 0 then margin / revenue * 100 else 0
      filter margin > 0

Analytic window

The analytic_window field enables cross-source lookups by joining a secondary dataset into the transform. The secondary source is loaded into memory and indexed by the join key.

- type: transform
  name: enrich_orders
  input: orders
  config:
    analytic_window:
      source: products
      on: product_id
      group_by: [product_id]
    cxl: |
      emit order_id = order_id
      emit product_name = $window.first().product_name
      emit quantity = quantity
      emit line_total = quantity * price

The $window.* namespace provides access to the windowed data. Functions like $window.first().<field>, $window.last().<field>, and $window.count() operate over the matched group. first() and last() return a whole record, so name the field to read; without one they return null. See Window Functions.

Validations

Declarative validation checks can be attached to a transform. They run against each record and either route failures to the DLQ (severity error) or log a warning and continue (severity warn).

- type: transform
  name: validate_orders
  input: raw_orders
  config:
    cxl: |
      emit order_id = order_id
      emit amount = amount
      emit email = email
    validations:
      - field: email
        check: "not_empty"
        severity: error
        message: "Email is required"
      - check: "amount > 0"
        severity: warn
        message: "Non-positive amount"
      - field: order_id
        check: "not_empty"
        severity: error

Validation fields

FieldRequiredDescription
fieldNoRestrict the check to a single field
checkYesValidation name (e.g. "not_empty") or CXL boolean expression
severityNoerror (default) routes to DLQ; warn logs and continues
messageNoCustom error message for DLQ entries
nameNoValidation name for DLQ reporting. Auto-derived from field + check if omitted
argsNoAdditional arguments as key-value pairs

Expansion cap (max_expansion)

When a transform body contains an emit each statement, every input record can fan out into multiple output records. The max_expansion field caps how many output records a single input record may produce – a safety bound against unexpectedly large arrays.

- type: transform
  name: explode_items
  input: orders
  config:
    max_expansion: 5000      # default: 10000
    cxl: |
      emit each it in items {
        emit order_id = order_id
        emit sku = it["sku"]
        emit price = it["price"]
      }
FieldTypeDefaultDescription
max_expansionu6410000Maximum cumulative output records per input record.

If a single input record’s emit each block produces more than max_expansion output records, the originating record routes to the DLQ with category expansion_limit_exceeded instead of producing a truncated or unbounded result. No partial output is emitted for that record – the cap is enforced eagerly so the writer never sees records from a runaway expansion.

When to tune

  • Lower (e.g. 100, 1000) when input arrays are bounded by a known business rule and you want hostile or malformed input to surface as a DLQ entry rather than as a flood of downstream records.
  • Higher (e.g. 100000, 1000000) when legitimate input carries large arrays – for example, an order with a long line-item list or an event carrying a per-second pricing curve.

The DLQ category expansion_limit_exceeded is distinct from generic CXL evaluation failures, so DLQ-side filters and metrics can target expansion runaway specifically. See Error Handling & DLQ for the wider DLQ contract.

Batch size (batch_size)

A streaming-eligible transform hands its output downstream in bounded batches rather than accumulating the whole stage before the next stage runs. batch_size sets how many events (records plus document-boundary punctuations) a batch holds. A per-transform batch_size overrides the pipeline-level pipeline.batch_size for this one stage; omit it to inherit the pipeline value (or the built-in default of 2048).

- type: transform
  name: enrich
  input: orders
  config:
    batch_size: 512         # override pipeline.batch_size for this stage
    cxl: |
      emit order_id = order_id
      emit total = quantity * unit_price
FieldTypeDefaultDescription
batch_sizeusizeinherits pipeline.batch_size (else 2048)Events per streaming batch for this transform. Must be >= 1.

A batch_size of 0 is rejected at config load (a zero-event batch never flushes). Smaller batches lower the in-flight memory of a streaming stage at the cost of more per-batch bookkeeping; larger batches amortize the bookkeeping at the cost of a larger live working set. The default suits typical record widths — tune it only when a profiling run shows a streaming stage’s per-batch footprint matters. See Streaming vs. Blocking Stages for which stages stream and which fully materialize.

Log directives

Log directives declare bounded structured diagnostic events during transform execution:

- type: transform
  name: process
  input: validated
  config:
    cxl: |
      emit id = id
      emit result = compute(value)
    log:
      - name: transform.record_processed
        level: info
        when: per_record
        every: 1000
        message: "Processed record"
        fields: [id]
      - name: transform.record_failed
        level: warn
        when: on_error
        message: "Record failed processing"
      - name: transform.started
        level: debug
        when: before_transform
        message: "Starting transform"

Log directive fields

FieldRequiredDescription
nameYesStable event name: a bounded dotted identifier using ASCII letters, digits, or underscores
levelYestrace, debug, info, warn, or error
whenYesbefore_transform, after_transform, per_record, or on_error
messageYesStatic event message, at most 1024 UTF-8 bytes. Interpolation is rejected; request record values with fields instead
everyFor per_recordPositive record interval. It is required for every per_record event, including explicit every: 1, and rejected for other timings
fieldsNoUp to 256 unique record field names requested as structured attributes. Available only for per_record and on_error events
conditionNoCXL boolean expression; the event fires only for records where it is true. Available only for when: per_record, and at most 512 UTF-8 bytes

A transform may declare at most 32 events and request at most 256 fields in aggregate across them. Event names and field selectors use the same grammar as deployment field policy: dot-separated segments beginning with an ASCII letter or underscore, followed by ASCII letters, digits, or underscores.

fields is the only channel by which record data reaches an event — message is static text. A selector naming a field the incoming record does not carry contributes nothing, so a directive whose selectors all miss would publish an event with no attributes at all, which reads exactly like a run whose records were empty.

Clinker refuses that when the pipeline compiles (E374). The rejection names the selector, lists the columns the input row does carry, and — when your spelling is close to one of them — gives you the corrected fields: line to paste:

[E374] transform `enrich` log[0].fields requests `orderId`, which the input
record does not carry; the upstream row has `order_id`, `amount`, `region` —
write `fields: [order_id]`

Selectors bind against the transform’s input row, for per_record and on_error alike: dispatch fires before this transform’s own cxl: block, so a column the transform produces cannot be requested. Request the columns it reads instead.

The check decides what the declared schema can decide. A column that reaches the transform through an open composition port is not visible to it, and a selector naming one of those is still checked only at run time — counted in the run’s admission accounting under the missing-field total.

One event name, one set of fields

An event name is the identity a collector groups records by, and it carries no node identity of its own. Two transforms may emit the same event — that is how one thing that happens in several places is reported as one thing — but they have to declare the same fields, or the collector receives two record shapes under one name and nothing downstream can separate them.

Clinker checks this across the whole pipeline, composition bodies included, and refuses a disagreement with E375:

[E375] transform "shape" in composition "customer_enrichment": `log` event
"transform.customer_seen" is also declared by transform "seen" with different
fields; one event name carries one set of fields — give both declarations the
same `fields`, or give one of them its own event name

A composition used twice is not a conflict: both instances declare the same event with the same fields, and agreement is what the rule asks for.

Logging only the records you care about

every thins a per-record event by count. condition selects it by content — use it when the interesting records are rare and you want all of them rather than every thousandth record:

    log:
      - name: transform.large_order
        level: info
        when: per_record
        every: 1
        condition: "amount > 1000"
        message: "large order"
        fields: [order_id, amount]

The two compose: every is applied first, then condition, so every: 100 with a condition logs every hundredth record that also matches.

A condition is CXL, checked when the pipeline compiles. It must resolve to a boolean, and it is evaluated against the transform’s input record — the one that arrived, before this transform’s own cxl: block runs. A field the transform only produces is therefore not in scope; write the condition in terms of the fields the transform reads.

A condition decides only whether an event fires. It cannot add anything to one: the values that leave the process are still exactly the fields you requested, each still subject to deployment policy. Narrowing a condition can never widen what is exported.

Transform declarations name events, request fields, and may gate a per-record event on its own input; they do not choose a destination, credentials, routing, redaction, or sampling policy. Each requested event-field pair is denied unless deployment observability policy explicitly allows, hashes, or replaces it. Telemetry delivery is bounded and best effort and cannot change transform results or published output — including a condition that fails to evaluate, which drops its event rather than failing the run.

Every transform event also carries the fixed correlation fields execution_id, batch_id, and pipeline_name. Unlike requested record fields, these are not default-deny and are not gated by field_policy: they are engine-supplied identity that never derives from a source, and they are what makes an exported event joinable to the machine stream and to the lineage events. A deployment that allows an event without also writing three correlation rules still gets telemetry it can correlate.

Because they are exported verbatim, choose a --batch-id that is safe to send to your collector — an identifier, not a tenant name or anything else you would not want retained there. A field_policy rule naming one of these three fields does not redact it. Source paths, records, secrets, and raw error text are never implicit attributes.

The former log_rule directive key and pipeline-level log_rules block are rejected. Move event identity and safe field requests into the transform’s log: entries as shown above; keep routing, privacy, credentials, and sampling in deployment policy.

Complete example

- type: source
  name: employees
  config:
    name: employees
    type: csv
    path: "./data/employees.csv"
    schema:
      - { name: employee_id, type: string }
      - { name: first_name, type: string }
      - { name: last_name, type: string }
      - { name: department, type: string }
      - { name: salary, type: int }
      - { name: hire_date, type: date }

- type: transform
  name: enrich_employees
  description: "Compute display name and tenure"
  input: employees
  config:
    cxl: |
      emit employee_id = employee_id
      emit display_name = last_name.concat(", ", first_name)
      emit department = department.upper()
      emit salary = salary
      emit annual_bonus = if salary >= 80000 then salary * 0.15
        else salary * 0.10
    validations:
      - field: employee_id
        check: "not_empty"
        severity: error
        message: "Employee ID is required"
      - check: "salary > 0"
        severity: warn
        message: "Salary should be positive"
    log:
      - name: transform.employee_processed
        level: info
        when: per_record
        every: 5000
        message: "Processing employees"
        fields: [employee_id]

Route Nodes

Route nodes split a stream of records into named branches based on CXL boolean conditions. Each branch becomes an independent output port that downstream nodes can wire to using port syntax.

Interactive companion: the Route and Merge explainer shows, record by record, which conditions are checked and where each record goes, and how a Merge rejoins the branches.

Basic structure

- type: route
  name: split_by_value
  input: orders
  config:
    mode: exclusive
    conditions:
      high: "amount.to_int() > 1000"
      medium: "amount.to_int() > 100"
    default: low

This creates three output ports: split_by_value.high, split_by_value.medium, and split_by_value.low.

Conditions

The conditions: field is an ordered map of branch names to CXL boolean expressions. Each expression is evaluated against the incoming record.

    conditions:
      priority: "urgency == \"high\" and amount > 500"
      standard: "urgency == \"medium\""
      bulk: "quantity > 100"
    default: other

Condition keys become the port names used in downstream input: wiring.

Compile-time checking

Branch conditions are typechecked when the pipeline is compiled – the plan-building pass you trigger with clinker run pipeline.yaml --explain or a bare --dry-run (see Validation & Dry Run). Each condition is checked against the concrete column types of the route’s input, so a condition that references an unknown column or compares incompatible types fails at compile time rather than partway through a run.

Compile failures surface as a CXL diagnostic keyed to the failure class – E202 for a branch condition that does not parse, E203 for one that references an unknown column, and E200 for one that compares incompatible types – and name the offending branch, for example split_by_value (branch high), so in a multi-branch route the error points at the specific branch rather than just the route node.

Default branch

The default: field is required. Records for which no condition is true are routed to the default branch. A condition whose result is null counts as not true.

A condition that fails to evaluate is not “no match”. Under error_handling.strategy: continue the record is dead-lettered and takes no branch, not even one whose condition held, and not the default; under fail_fast the run stops. In exclusive mode the conditions after the first true one are never evaluated, so a condition further down cannot fail for that record. See An evaluation error is never false.

Routing modes

Exclusive (default)

In exclusive mode, conditions are evaluated in declaration order and the first matching condition wins. A record appears in exactly one branch. Order matters – put more specific conditions first.

    mode: exclusive
    conditions:
      vip: "lifetime_value > 100000"
      high: "lifetime_value > 10000"
      medium: "lifetime_value > 1000"
    default: standard

A customer with lifetime_value = 50000 is not over 100000, so vip is not true; high is the first true condition and wins. A customer with lifetime_value = 150000 is true for all three conditions, but goes to vip alone, because vip is checked first. Listing medium first would send both customers to medium.

Inclusive

In inclusive mode, all matching conditions route the record. A single record can appear in multiple branches simultaneously.

    mode: inclusive
    conditions:
      needs_review: "amount > 10000"
      flagged: "status == \"flagged\""
      international: "country != \"US\""
    default: standard

A flagged international order over 10000 would appear in needs_review, flagged, and international – three copies routed to three branches.

Downstream wiring

Downstream nodes reference route branches using port syntax: route_name.branch_name. The default branch is reached the same way, by its name. Several nodes can read the same branch; each receives every record on it.

A node whose own name equals a branch name can also reference the Route by its bare name and receives that branch: a Sink named high with input: classify reads classify.high.

- type: route
  name: classify
  input: transactions
  config:
    mode: exclusive
    conditions:
      high: "amount > 1000"
      medium: "amount > 100"
    default: low

- type: transform
  name: high_value_processing
  input: classify.high
  config:
    cxl: |
      emit txn_id = txn_id
      emit amount = amount
      emit review_flag = true

- type: transform
  name: standard_processing
  input: classify.medium
  config:
    cxl: |
      emit txn_id = txn_id
      emit amount = amount

- type: sink
  name: low_value_out
  input: classify.low
  config:
    name: low_value_out
    type: csv
    path: "./output/low_value.csv"

Constraints

  • Give every condition a distinct name; each name is a separate port.
  • Give default a name that no condition uses. This is not currently checked when the pipeline is planned: a default with the same name as a condition is folded into that condition’s branch, so its records cannot be told apart from the matches.
  • Declare at least one condition. A Route with none is not rejected today; every record takes the default branch.

Complete example

pipeline:
  name: order_routing

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./data/orders.csv"
      schema:
        - { name: order_id, type: int }
        - { name: region, type: string }
        - { name: amount, type: float }
        - { name: priority, type: string }

  - type: route
    name: by_region
    input: orders
    config:
      mode: exclusive
      conditions:
        domestic: "region == \"US\" or region == \"CA\""
        emea: "region == \"UK\" or region == \"DE\" or region == \"FR\""
        apac: "region == \"JP\" or region == \"AU\" or region == \"SG\""
      default: other

  - type: sink
    name: domestic_orders
    input: by_region.domestic
    config:
      name: domestic_orders
      type: csv
      path: "./output/domestic.csv"

  - type: sink
    name: emea_orders
    input: by_region.emea
    config:
      name: emea_orders
      type: csv
      path: "./output/emea.csv"

  - type: sink
    name: apac_orders
    input: by_region.apac
    config:
      name: apac_orders
      type: csv
      path: "./output/apac.csv"

  - type: sink
    name: other_orders
    input: by_region.other
    config:
      name: other_orders
      type: csv
      path: "./output/other_regions.csv"

Merge Nodes

Merge nodes concatenate multiple upstream branches into a single stream. They are the counterpart to route nodes – where a route splits one stream into many, a merge joins many streams back into one.

Merge is for streamwise concatenation of inputs that share a schema. For record-level joining across inputs that have different schemas, see Combine Nodes.

Interactive companion: the Route and Merge explainer shows how each mode mixes its inputs, and what an inclusive Route’s copies look like after a Merge.

Basic structure

- type: merge
  name: combined
  inputs:
    - east_data
    - west_data
  config: {}

Note the key differences from other node types:

  • Uses inputs: (plural), not input: (singular).
  • The config: block is empty – all wiring is on the node header.
  • Using input: (singular) on a merge node is a parse error.

Wiring

The inputs: field is a list of upstream node references. These can be bare node names or port references from route nodes:

- type: merge
  name: rejoin
  inputs:
    - process_high
    - process_medium
    - classify.low           # Port syntax for a route branch
  config: {}

Downstream nodes wire to the merge as a normal single-input reference:

- type: sink
  name: final_output
  input: rejoin
  config:
    name: final_output
    type: csv
    path: "./output/combined.csv"

Modes

Merge’s cross-input ordering discipline is selected by config.mode. Two modes exist; concat is the default.

concat (default)

Predecessor records drain in declaration order: inputs[0] flows to output first, then inputs[1], then inputs[2], and so on. Within a single predecessor, its arrival order is preserved. Output is reproducible run-to-run for the same predecessor paths.

- type: merge
  name: combined
  inputs: [east, west]
  config:
    mode: concat

interleave

Records flow to output as they become available from any predecessor. Each input’s arrival order is preserved; cross-input order follows wall-clock arrival and is not promised.

- type: merge
  name: combined
  inputs: [east, west]
  config:
    mode: interleave

Seeded interleave — interleave_seed:

Snapshot tests and benchmarks that need reproducible cross-input ordering can opt into a deterministic schedule:

- type: merge
  name: combined
  inputs: [east, west]
  config:
    mode: interleave
    interleave_seed: 42

With a seed, the cross-input order is reproducible from run to run regardless of upstream timing. To get there, the Merge reads all of its inputs into memory before emitting, so a seeded interleave buffers more than the other modes — use it for tests and benchmarks, not high-volume production merges.

Choosing a mode

ModeOrderWhen to use
concatInputs emitted in declaration order, each fully drained before the next.Downstream depends on a stable, declaration-ordered sequence (byte-identical output, contiguous time partitions).
interleave (unseeded)Records emitted as they arrive; per-input order preserved, cross-input order varies.Lowest latency and the consumer is order-insensitive (e.g. an aggregator grouping by key, or a writer that doesn’t assert on row order).
interleave (seeded)Reproducible cross-input order.Tests and benchmarks that assert on exact row sequence. Buffers all inputs in memory.

For high-volume merges, prefer concat or unseeded interleave — both stream their inputs and let a slow downstream consumer naturally throttle the upstream readers, so memory stays bounded. A seeded interleave does not, because it buffers everything first.

Record ordering

Records arrive in the order described by the mode in use — see Modes and Choosing a mode above. Merge does not promote matching per-input sort_order declarations into one global sort: concatenating two independently sorted files can still put a high key before a lower key at the input boundary.

For concat and seeded interleave, exact sequence is a supported oracle for the same input paths and configuration. For unseeded interleave, compare the decoded record multiset and aggregate values; snapshotting incidental cross-input arrival order would assert behavior Clinker does not promise. If you need one sorted sequence regardless of Merge mode, declare sort_order on the downstream Sink. That terminal sort uses only the authored keys and does not invent a hidden identity tie-breaker.

Use cases

Reuniting route branches

The most common pattern is routing records through different processing paths and then merging them back together:

- type: route
  name: classify
  input: orders
  config:
    mode: exclusive
    conditions:
      high: "amount > 1000"
    default: standard

- type: transform
  name: process_high
  input: classify.high
  config:
    cxl: |
      emit order_id = order_id
      emit amount = amount
      emit surcharge = amount * 0.02
      emit tier = "premium"

- type: transform
  name: process_standard
  input: classify.standard
  config:
    cxl: |
      emit order_id = order_id
      emit amount = amount
      emit surcharge = 0
      emit tier = "standard"

- type: merge
  name: all_orders
  inputs:
    - process_high
    - process_standard
  config: {}

- type: sink
  name: result
  input: all_orders
  config:
    name: result
    type: csv
    path: "./output/all_orders.csv"

A Merge passes on every record from every input and does not remove duplicates. Branches of an inclusive Route can carry the same record, and rejoining them puts that record in the output once per branch it took. Use an exclusive Route when each record should come out once.

Unioning multiple sources

Merge nodes can combine records from multiple source files that share the same schema:

- type: source
  name: jan_sales
  config:
    name: jan_sales
    type: csv
    path: "./data/sales_jan.csv"
    schema:
      - { name: sale_id, type: int }
      - { name: amount, type: float }
      - { name: region, type: string }

- type: source
  name: feb_sales
  config:
    name: feb_sales
    type: csv
    path: "./data/sales_feb.csv"
    schema:
      - { name: sale_id, type: int }
      - { name: amount, type: float }
      - { name: region, type: string }

- type: merge
  name: all_sales
  inputs:
    - jan_sales
    - feb_sales
  config: {}

- type: aggregate
  name: totals
  input: all_sales
  config:
    group_by: [region]
    cxl: |
      emit total = sum(amount)
      emit count = count(*)

Schema constraints across inputs

Merge concatenates streams positionally against the merge node’s output_schema (taken from the first input). Every input must therefore agree on column shape — same column names, same on_unmapped policy, same correlation_key set.

Disagreement on the $widened auto_widen sidecar (one source uses auto_widen, another uses drop / reject) fails compile with E315. See Auto-Widen & Schema Drift → E315 for the full diagnostic shape and remediation.

Combine Nodes

Combine nodes are the N-ary record-combining operator. Every input is declared up front and bound to a qualifier; the where: expression matches records across inputs using qualified field references (e.g. orders.product_id == products.product_id); the cxl: body shapes the output row.

Combine is distinct from merge: merge concatenates upstream branches that share a schema, while combine joins records across inputs that have different schemas.

Interactive companion: the Combine playground shows, driver row by driver row, which build rows where: matches and what match: and on_miss: do with them, including range joins.

Basic structure

- type: combine
  name: enrich
  input:
    orders: orders         # qualifier: upstream node name
    products: products
  config:
    where: "orders.product_id == products.product_id"
    match: first
    on_miss: null_fields
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.product_name
      emit amount = orders.amount
    propagate_ck: driver

Note the differences from other node types:

  • Uses input: as a map, binding qualifier names to upstream node references. Other nodes use input: as a single string or inputs: as a list of strings.
  • Every field reference inside where: and cxl: must be qualified (<qualifier>.<field>). Bare field names are a compile error.
  • Using inputs: (plural list) on a combine node is a parse error.

Wiring

Each entry in the input: map binds a qualifier to an upstream node:

  input:
    orders: orders                  # qualifier "orders" -> source node "orders"
    products: products
    high_priority: classify.high    # qualifier "high_priority" -> route port

Qualifiers are local names used inside where: and cxl:; they do not need to match the upstream node name. Upstream references can be bare node names or port references from a route node.

Iteration order in the input: map is preserved and used as the default driver-selection order (see Choosing the driving input below).

Configuration fields

FieldRequiredDefaultDescription
whereYes–CXL boolean expression matching records across inputs. Must contain at least one cross-input equality or range conjunct (a predicate with neither is rejected at plan time — see Predicate requirements).
matchNofirstMatch cardinality: first, all, or collect.
on_missNonull_fieldsDriver-record handling on zero predicate matches: null_fields, skip, or error.
cxlYes (except under match: collect)–Emit statements defining the output row. Empty under match: collect.
driveNofirst inputExplicit driver-input qualifier. Overrides the iteration-order default.
strategyNoautoExecution strategy hint: auto or grace_hash.
propagate_ckYes–Selects which correlation-key columns ride onto the output. driver keeps the driver’s CK only; all unions every input’s CK columns; { named: [<field>, ...] } carries an explicit subset. See Correlation-key propagation below.
max_output_rowsNounlimitedOpt-in cap on the number of rows this combine may emit. When set, the run fails loud (diagnostic E325) the moment the output would exceed the cap — it never truncates to a partial result. See Output-size cap.

The where: predicate

The where: expression is a CXL boolean expression evaluated for every candidate record pair across inputs. It must contain at least one cross-input equality – an equality with field references from two different inputs:

  where: "orders.product_id == products.product_id"

Compound predicates combine multiple conjuncts with and. Each conjunct is classified by the planner:

  • Equi conjunct – a cross-input equality (a.x == b.y). Drives the hash lookup or sort-merge join.
  • Range conjunct – a cross-input ordered comparison (a.start <= b.ts and b.ts <= a.end). Handled by a range join (IEJoin), whether or not an equi conjunct also links the same two inputs.
  • Residual conjunct – any other CXL predicate (intra-input filter, function call, etc.). Applied as a post-filter after the equi/range match.
  where: |
    orders.product_id == products.product_id
    and orders.amount >= 100
    and products.region == "us-east"

Above: the equi conjunct drives the join; orders.amount >= 100 and products.region == "us-east" are applied as residuals.

Every combine predicate must carry at least one cross-input equality or range conjunct. A predicate with neither — a pure residual with no decomposable cross-input comparison — is rejected at plan time with diagnostic E313; there is no supported execution strategy for it. Pure-range predicates without an equi conjunct are fully supported via IEJoin.

Non-orderable range keys

A range conjunct compares values that must be orderable at runtime: integers, finite floats, exact decimals, dates, and datetimes. When a record’s range key evaluates to a non-orderable value — SQL NULL, a non-finite float (NaN/infinity), or any other type — that record can never satisfy the range comparison, so it is routed out of the range match rather than joined:

Decimal and mixed-numeric range keys. Exact fixed-point decimal range keys are fully supported on every join strategy — a monetary band join (amount >= tier.floor and amount < tier.ceiling) matches correctly. A range conjunct that mixes an integer and a float operand across the two inputs is also supported: the integer is compared as a float, exactly as the >=/< operators compare it elsewhere.

Decimal range keys are placed on a shared fixed-point grid with up to 18 fractional digits and an integer magnitude up to roughly 1.7 × 10²⁰ — well beyond any realistic monetary value. A decimal range value outside that grid (more than 18 fractional digits, which would truncate, or a magnitude that would overflow the grid) is not silently dropped: the run stops with diagnostic E326 naming the combine, so a wrong or empty result is never emitted. Rescale or narrow the compared values if you hit it.

Datetime range keys. date and datetime range keys compare at their native resolution — a datetime to the nanosecond, across the full representable calendar (no microsecond rounding; instants before 1677 or after 2262 stay exact). Two timestamps that differ only below the microsecond therefore match, sort, and group as distinct instants on every join strategy, so a sub-microsecond as-of or band lookup neither drops a boundary match nor merges two near-simultaneous events.

abs/min/max/clamp and rejected range keys (E327). abs, min, max, and clamp return the numeric supertype (int | float). When both operands of the range conjunct recover the same concrete type — abs(int) >= abs(int), or a min/max/clamp whose result can only be one type (all-int or all-float, recovered through nested calls and arithmetic such as abs(a.x + 1) or abs(min(a.i, b.i))) — the axis is exactly as safe as a plain matching-typed key and the join runs normally. Otherwise the conjunct cannot be reduced to one exact numeric axis and is rejected at plan time with E327, rather than routed to the join where it could silently drop rows or mismatch. E327 fires when such an operand stays genuinely ambiguous — a mixed pair like abs(int) >= abs(float), or a per-row-ambiguous result like min(int, float) that may return either the integer or the float operand — or on a non-orderable pairing such as a string comparison or a decimal compared against a float/numeric. Compare a matching-typed range key instead. (A supported int/float/decimal/date/datetime range key, including a mixed int/float field pair, never triggers E327.)

  • A driver record with a non-orderable range key is treated as a zero-match driver and handled by on_miss (null_fields / skip / error).
  • A build record with a non-orderable range key is dropped (it can match nothing).

This is a runtime routing decision on the data, distinct from the plan-time E313 rejection above (which is about the predicate shape, not the values).

Match modes

match: first

Emit one output row per driver record, using the first matching build-side record. Standard 1:1 enrichment. Default.

“First” means the earliest in the build input’s arrival order: of the build records that match a driver, the one the build input delivered first. The same order sets the row order of match: all within a driver and the element order of a match: collect array. It holds identically for every join strategy the planner may pick and whatever shape the where: predicate has. One exception remains. When the build input exceeds the memory limit, the Combine sets part of it aside on disk and splits that part into smaller pieces until each fits. If a piece still does not fit after splitting, the Combine processes it in chunks and currently decides first, collect and on_miss once per chunk rather than once per driver, so a driver can receive more than one first row or more than one collect row. A piece fails to fit in three cases: the build rows sharing one key value do not fit on their own; the build input has already been split into the maximum of 4,096 pieces; or splitting a piece again leaves most of its rows together, as happens when many rows share a key value or their key values happen to land in the same piece. So the exception can reach a driver even when no single key value is over the limit, for example with a large build input under a tight memory limit. This is a known defect.

  • To pick a different record, order the build input upstream. For example, to enrich each driver with the latest price, deliver the build input sorted descending on its date, such as with a descending sort_order on the build Source. A Combine has no ordering option of its own.
  • A correlation key orders its Source. Declaring correlation_key on the build Source makes the planner sort that Source’s rows by the key before they reach the Combine, unless its declared sort_order already begins with the key. That order is what “first” follows. See Correlation keys.
  • An unseeded interleave Merge has no fixed order. When the build input is such a Merge, its cross-input order follows arrival at run time, so “first” follows that unfixed order too. Use concat, or a seeded interleave, for a build input whose order must be the same on every run.
  config:
    where: "orders.product_id == products.product_id"
    match: first
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.product_name

The where: predicate selects the match; the cxl: body is a post-match projection that runs once on the chosen build record. Selection and projection are separate steps: if the body filters the row out (a filter that fails, or a body that emits nothing), that one output row is dropped. The combine does not fall back to a later matching build, and the driver is not treated as unmatched — it matched the predicate, the body just produced no row. on_miss (below) never fires for such a driver; it fires only when the predicate matched nothing at all. This holds identically for every join strategy the planner may pick.

Behavior change. This is a change in observable output for existing pipelines whose where: predicate carries a range, equi+range, or single-inequality comparison — the shapes the planner runs as a sort-merge join or an IEJoin (both the pure-range block-band path and the equi+range hash-partitioned path). On any of those three strategies, a driver that matched the predicate but whose body skipped every candidate was previously routed to on_miss — firing null_fields, skip, or error. It now silently produces no row, matching the pure-equality strategies (in-memory hash and grace hash), which already behaved this way and are unchanged. A pipeline that relied on the old routing (for example, on_miss: error tripping on a body-skipped driver) no longer sees it.

When where: fails to evaluate. A where: that raises an error for a candidate, such as a division by zero, has not said whether that candidate matches. It is neither a match nor a non-match. Under match: first the earliest candidate that is not a non-match decides. If it matched, the driver is enriched with it. If its where: failed, the driver is dead-lettered with that build row and writes no output row, even when a later candidate would match: taking the later one would publish a row that depends on an evaluation that failed. A candidate after the deciding one is not part of the driver’s result, so its failure is never written. This holds on every join strategy.

match: all

Emit one output row for every matching build-side record. 1:N fan-out – if a driver record matches three build records, three rows are emitted, in the build input’s arrival order (see match: first).

  config:
    where: "employees.department == benefits.department"
    match: all
    cxl: |
      emit employee_id = employees.employee_id
      emit benefit = benefits.benefit_name

Each output row depends only on its own pair, so a candidate whose where: fails to evaluate is dead-lettered with its build row while the driver’s other matches are still emitted.

match: collect

Gather every matching build-side record into a single Array-typed field on the output row. The driver record appears once; the build matches are aggregated into an array, in the build input’s arrival order (see match: first). The cxl: body must be empty under collect – the combine node synthesizes the output as { driver fields..., <build_qualifier>: Array }.

  config:
    where: "orders.product_id == products.product_id"
    match: collect
    cxl: ""

A per-group entry limit of 10,000 prevents unbounded growth.

A collect row states the complete set of a driver’s matches. If any candidate’s where: fails to evaluate, that set is unknown, so the driver writes no row, neither a partial array nor an empty one; each failing candidate is dead-lettered with its build row. Candidates past the 10,000-entry limit are still checked, so every failure among them is dead-lettered too.

Use collect when you need the set of matches as a single structured value; use all when you need a flat row per match.

Unmatched records (on_miss)

on_miss controls what happens to driver records with zero predicate matches — drivers for which no build-side record satisfied where:, because every candidate evaluated it to false or null, or because there was no candidate at all. Two kinds of driver are not misses and never reach on_miss, under any match mode and on every join strategy:

  • A driver that matched the predicate but whose cxl: body skipped or failed the row (see match: first). It simply produces no output row for that match, and a body failure is dead-lettered.
  • A driver any of whose candidates failed to evaluate where:. Each failure is dead-lettered, and whatever on_miss says, it does not fire: on_miss: error does not stop the run and on_miss: null_fields writes no null-filled row, because the driver was never shown to have no match.
ValueSemantics
null_fields (default)Build-side fields resolve to null. Driver record is still emitted. Equivalent to left-join.
skipDriver record is dropped. Equivalent to inner-join.
errorPipeline fails on the first unmatched driver record.
  config:
    where: "orders.product_id == products.product_id"
    on_miss: skip

on_miss: error is useful for strict referential integrity where any miss should halt processing. on_miss: skip is the inner-join shape. on_miss: null_fields is the left-join shape and the default.

Composite keys

Chain multiple cross-input equalities with and:

  config:
    where: |
      sales.department == targets.department
      and sales.region == targets.region
    cxl: |
      emit department = sales.department
      emit region = sales.region
      emit actual = sales.amount
      emit goal = targets.goal

All conjuncts must hold for a record pair to match.

Multi-input combine (three or more)

Combine accepts any number of inputs. Each pair of inputs that should be related needs an explicit cross-input equality:

- type: combine
  name: fully_enriched
  input:
    orders: orders
    products: products
    categories: categories
  config:
    where: |
      orders.product_id == products.product_id
      and products.category_id == categories.category_id
    match: first
    on_miss: null_fields
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.product_name
      emit category_name = categories.name
      emit amount = orders.amount
    propagate_ck: driver

The planner builds a join tree by walking equalities pairwise: starting from the driver, it adds one input at a time, always one linked by an equality to the inputs already joined. It is designed to prefer the input with the fewest rows at each step, but those row counts are not available yet when it runs, so in practice it takes the linked inputs in the order they are declared.

Choosing the driving input

The driver is the input whose records flow through one at a time during execution; the other inputs are materialized as build-side hash tables (or IEJoin index structures). By default the first input in the input: map is the driver.

Use drive: to override:

  config:
    where: "orders.product_id == products.product_id"
    drive: products
    cxl: |
      emit product_id = products.product_id
      emit product_name = products.product_name
      emit sample_order_id = orders.order_id

With drive: products, the pipeline emits one row per product enriched with a matching order, instead of one row per order enriched with its product. Pick the driver based on which side you want to iterate over (typically the larger stream, or the one whose ordering you want to preserve).

Strategy hint

ValueBehavior
auto (default)Planner picks a strategy from the predicate shape. Equalities only: an in-memory hash join, or grace hash when the build side’s estimated size is close to the memory limit. Equalities plus ranges: a range join (IEJoin) that also groups by the equal values. Ranges only: a sort-merge join when there is a single comparison between two int, float, date or datetime fields of the same type (not decimal) and both inputs already arrive sorted ascending on those fields, each from a single sorted file or stream; otherwise a range join (IEJoin). Both give the same result.
grace_hashForce grace hash join (disk-spilling partitioned hash). Applies only to pure-equi predicates; ignored on predicates with range conjuncts.

The choice is made when the pipeline is planned, from an estimate of the inputs’ size. Under auto, an equal-ids join runs in memory unless that estimate says the build side is too large. If the estimate is low or missing (for example, a glob: source whose files are not known in advance) and the build side turns out larger than the memory budget, the run stops with E310 MemoryBudgetExceeded; it does not switch to disk partway through (#1337). Set strategy: grace_hash when the build side may be larger than the memory budget: it partitions the build side to disk from the start.

Correlation-key propagation

Combine declares which correlation-key columns its output rows carry via the required propagate_ck field.

- type: combine
  name: enriched
  input:
    orders: orders
    products: products
  config:
    where: "orders.product_id == products.product_id"
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.name
    propagate_ck: driver        # driver-only (today's behavior)
    propagate_ck: all           # union of every input's $ck.* columns
    propagate_ck:
      named: [order_id]         # explicit subset (intersected with upstream)
  • driver – output carries only the driver input’s correlation-key columns. Build-side records contribute body fields, but their group identity is consumed by the match; when a match fails, the build record’s dead letter follows the failing driver’s correlation group.
  • all – output carries every input’s correlation-key columns. Use when the build side carries keys that downstream operators need to read.
  • named: [<field>, ...] – an explicit subset. Use to project a multi-field key down to a single field after a join.

Driver wins on a name collision: if both the driver and a build input declare the same key field, the output keeps the driver’s value. See the Correlation-key combine interaction reference for how each match mode fills the propagated key (especially match: collect).

propagate_ck is required on every combine; pipelines without an explicit value fail to compile. Existing pipelines migrate by adding propagate_ck: driver, which is bit-for-bit equivalent to today’s behavior.

Output-size cap (max_output_rows)

max_output_rows is an opt-in ceiling on how many rows a combine may emit. It defaults to unlimited; set it to guard against a permissive or mis-specified predicate that would explode a small pair of inputs into a huge result (for example, a range join over an unexpectedly hot key producing a near cross product):

  config:
    where: "orders.ts >= prices.effective_from"
    cxl: |
      emit order_id = orders.order_id
      emit price = prices.amount
    propagate_ck: driver
    match: all
    max_output_rows: 1000000

Semantics:

  • Fail-loud, never truncate. The moment the combine would emit more than the cap, the run stops with diagnostic E325 naming the node and the cap. It does not produce a capped or partial result — a silently truncated join would corrupt downstream data.
  • Independent of the memory budget. This is a result-size guard, not memory pressure. A runaway join can be perfectly bounded in memory (its output spills to disk) yet still produce far more rows than intended; max_output_rows caps the row count regardless of bytes.
  • Covers the whole output, on every strategy. The cap counts every emitted output row across all match modes and any on_miss rows, and is enforced identically whichever join strategy the planner picks (hash build-probe, grace-hash, sort-merge, or the IEJoin block-band).
  • collect counts driver rows. Under match: collect a combine emits one output row per driver row (each carrying an array of up to 10 000 collected matches), so max_output_rows bounds the driver-row count, not the number of collected array elements.
  • Dead-lettered rows are not counted. A matched pair whose cxl: body or residual raises a recoverable eval failure is routed to the DLQ, not the output, so it does not count toward the cap. Only that pair is dead-lettered: under match: all the driver’s other matches are still evaluated and emitted, whichever join strategy runs. A failing residual is neither a match nor a miss: under match: first a failure on the deciding candidate is the driver’s only result, and under match: collect it leaves the driver with no row (see match: first and match: collect). The driver row and the matched build row are dead-lettered, each with its own Source and row number; each failure writes its own pair, so a driver that fails against several build rows is dead-lettered once per failure, and a build row that several failing drivers matched is dead-lettered once for each of them.
  • N-ary combines cap the final output. A combine whose where: spans three or more inputs is decomposed into a chain of binary steps; max_output_rows guards the final combined output, not the intermediate chain steps.

If the large result is expected, raise the cap (or omit the field). If it is not, tighten the where: predicate. Run clinker explain --code E325 for the full remediation guide.

Memory considerations

Each non-driving (build-side) input is held in memory while the join runs, so plan for roughly 1.5–2× its file size — a 50 MB lookup table needs about 75–100 MB. An equal-ids join spills its build side to disk only when it runs as a grace hash join, which the planner chooses from its size estimate or which strategy: grace_hash forces. An in-memory join whose build side outgrows the budget stops the run with E310 instead of spilling (#1337). Set the budget with pipeline.memory.limit; see Memory Tuning.

Range and equi+range predicates (the IEJoin block-band strategy) are bounded on both input axes and the output: each side is drained into disk-backed, key-sorted blocks, and the emitted rows accumulate in a spillable sort buffer rather than a resident vector. So a range join whose inputs — or whose result — exceed the memory budget spills automatically and completes, rather than failing. Even a single hot key whose block-pair is a near cross product streams through a bounded nested loop instead of materializing the whole candidate set. Use max_output_rows if you want such a runaway result to stop rather than spill.

Document boundaries

A Combine passes document boundaries through to its output, so a per-document Aggregate after a join still rolls up per document. A driver source that carries several documents (a glob: over monthly files, say) produces one roll-up per driver document after the join, not one fold spanning all of them. A document carried on both the driver and the build side opens and closes exactly once downstream. See Document Context & Envelopes for the per-document aggregation model.

Complete example

pipeline:
  name: order_enrichment

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./data/orders.csv"
      schema:
        - { name: order_id, type: string }
        - { name: product_id, type: string }
        - { name: amount, type: float }

  - type: source
    name: products
    config:
      name: products
      type: csv
      path: "./data/products.csv"
      schema:
        - { name: product_id, type: string }
        - { name: product_name, type: string }
        - { name: category, type: string }

  - type: combine
    name: enrich
    input:
      orders: orders
      products: products
    config:
      where: "orders.product_id == products.product_id"
      match: first
      on_miss: null_fields
      cxl: |
        emit order_id = orders.order_id
        emit product_id = orders.product_id
        emit product_name = products.product_name
        emit category = products.category
        emit amount = orders.amount
      propagate_ck: driver

  - type: sink
    name: result
    input: enrich
    config:
      name: result
      type: csv
      path: "./output/enriched_orders.csv"

See also

  • Multi-Input Combine – recipe-style walkthrough with input data and expected output.
  • Merge Nodes – streamwise concatenation; the right operator when inputs share a schema and no per-record matching is needed.
  • Memory Tuning – memory budget, spill thresholds, and strategy overrides.

Aggregate Nodes

Aggregate nodes group records by one or more fields and compute summary values using CXL aggregate functions. They consume all input records in a group before emitting a single summary record per group.

Basic structure

- type: aggregate
  name: dept_totals
  input: employees
  config:
    group_by: [department]
    cxl: |
      emit total_salary = sum(salary)
      emit headcount = count(*)
      emit avg_salary = avg(salary)

Group-by fields pass through automatically – you do not need to emit them. In this example, the output records contain department, total_salary, headcount, and avg_salary.

Group-by fields

The group_by: field is a list of field names from the input schema. Records sharing the same values for all group-by fields are placed in the same group.

    group_by: [region, department]
    cxl: |
      emit total_salary = sum(salary)
      emit max_salary = max(salary)

This produces one output record per unique (region, department) combination.

Values are the same group key when they are equal under the rule sorting uses (see How values are ordered):

  • Numbers group by exact value, whatever their type: 1, 1.0 and the decimal 1 are one group, and distinct integers are never merged, however large.
  • -0.0 and 0.0 are one group.
  • Every NaN, of either sign, is one group, separate from the group of null values.
  • A group reports the value of its first-arriving row. An integer column is written as integers; a column holding both 5 and 5.0 writes 5 if the integer arrived first and 5.0 if the float did. The value is the same whether or not the Aggregate spilled to disk.

The same rule groups Cull and Reshape partition_by, a window’s group_by, correlation keys, distinct and output splitting.

Global aggregation

An empty group_by list treats the entire input as a single group, producing exactly one output record:

- type: aggregate
  name: grand_totals
  input: orders
  config:
    group_by: []
    cxl: |
      emit grand_total = sum(amount)
      emit record_count = count(*)
      emit avg_order = avg(amount)

Aggregate functions

The following aggregate functions are available in CXL:

FunctionDescription
sum(field)Sum of all values in the group
count(*)Number of records in the group
avg(field)Arithmetic mean
min(field)Minimum value
max(field)Maximum value
collect(field)Collect all values into an array
weighted_avg(value, weight)Weighted average

Strategy hint

The strategy: field controls how aggregation is executed:

- type: aggregate
  name: totals
  input: sorted_data
  config:
    group_by: [account_id]
    strategy: streaming
    cxl: |
      emit total = sum(amount)
StrategyBehavior
autoDefault. The optimizer chooses based on whether the input is provably sorted for the group-by keys.
hashForce hash aggregation. Works on any input ordering. Holds all groups in memory (with disk spill if memory budget is exceeded).
streamingRequire streaming aggregation. Processes one group at a time with O(1) memory per group. Compile-time error if the input is not provably sorted for the group-by keys.

When to use streaming

The optimizer chooses streaming aggregation automatically when the Aggregate’s input is still sorted and the first fields of that order are exactly the group_by fields, in any order: take as many fields from the front of the order as there are group_by fields, and they must be the same set. Order [department, day] streams group_by: [department] and group_by: [day, department], but not group_by: [day]. An Aggregate with no group_by always streams. Use strategy: streaming as an explicit assertion – it turns a silent fallback to hash aggregation into a compile error (CXL0419), which is useful for catching sort-order regressions.

The order has to survive every stage between the Source and the Aggregate. A Merge, a Combine, a Reshape or Cull, distinct, and a Transform that writes one of the sort fields all drop it. Writing a field back unchanged counts: emit department = department drops an order on department. A Transform carries every input field through, so there is no need to emit sort fields. The sort order explainer lets you build a chain and see where the order holds.

When to use hash

Hash aggregation works on unsorted input and is the safe default. It uses more memory but handles any data ordering. Memory-aware disk spill kicks in when RSS approaches the pipeline’s memory.limit.

Correlation-key interaction

When a pipeline’s sources declare correlation_key: fields, an aggregate behaves one of two ways depending on its group_by:

  • group_by covers the correlation key — if any record in a group fails, the whole group (including the aggregate output row) is sent to the DLQ.
  • group_by omits a correlation-key field — only the failing records are dropped and the affected groups are recomputed, so the surviving rows still produce a correct aggregate.

You do not configure this; the engine picks the behavior from your group_by. The one restriction is that the second case cannot be combined with strategy: streaming. See Correlation Keys for the full rules.

Time-windowed aggregates

When time_window: is set on the aggregate body, records are grouped not just by group_by but also by event-time window. Each record is placed into one or more windows by its event time (the $source.event_time value derived from the source’s watermark), and each window emits one row per group once it closes.

Every upstream-reachable source must declare a watermark: so the engine knows when a window is complete. Without one, no window ever closes and the planner rejects the pipeline with E156.

The engine emits user-declared columns only — window bounds do not appear in the output unless you compute and emit them yourself. The emit order is ascending window_start (deterministic), so output rows naturally group by window.

Tumbling windows

Non-overlapping fixed-size buckets. Each record lands in exactly one window [floor(t / size) * size, floor(t / size) * size + size).

time_window:
  tumbling: { size: 1h }

Input (tumbling_demo.csv):

user_id,event_ts,kind
u1,2026-05-14T10:05:00,click
u2,2026-05-14T10:30:00,click
u1,2026-05-14T10:42:00,click
u1,2026-05-14T11:03:00,click
u2,2026-05-14T11:15:00,click
u2,2026-05-14T11:50:00,click

Output with tumbling: { size: 1h }, group_by: [user_id], emit n = count(*):

user_id,n
u1,2
u2,1
u1,1
u2,2

Reading top-to-bottom: the first two rows are the [10:00, 11:00) bucket (u1’s 10:05 and 10:42, then u2’s 10:30); the next two are the [11:00, 12:00) bucket (u1’s 11:03, then u2’s 11:15 and 11:50). Each input record contributes to exactly one window.

Hopping windows

Overlapping fixed-size buckets advanced by slide. Each record lands in ceil(size / slide) windows: slide < size produces overlap, slide == size degenerates to tumbling, slide > size produces gaps where some records fall in zero windows.

time_window:
  hopping: { size: 1h, slide: 30m }

Input (hopping_demo.csv):

user_id,event_ts,amount
u1,2026-05-14T10:05:00,10
u1,2026-05-14T10:42:00,20
u1,2026-05-14T11:10:00,15

Output with group_by: [user_id], emit total = sum(amount), emit n = count(*):

user_id,total,n
u1,10,1
u1,30,2
u1,35,2
u1,15,1

Three input records, four output rows — each record fans into two overlapping size: 1h, slide: 30m windows:

  • [09:30, 10:30) — just 10:05 → total=10, n=1
  • [10:00, 11:00) — 10:05 + 10:42 → total=30, n=2
  • [10:30, 11:30) — 10:42 + 11:10 → total=35, n=2
  • [11:00, 12:00) — just 11:10 → total=15, n=1

Session windows

Per-key gap-bounded sessions. A new record extends its key’s current session if its event time is within gap of the session’s last event time; otherwise it starts a new session. The boundary is data-driven, not clock-aligned.

time_window:
  session: { gap: 10m }

Input (session_demo.csv):

user_id,event_ts,action
u1,2026-05-14T10:00:00,login
u1,2026-05-14T10:07:00,click
u1,2026-05-14T10:13:00,click
u1,2026-05-14T10:50:00,login
u1,2026-05-14T10:55:00,click

Output with group_by: [user_id], emit n = count(*):

user_id,n
u1,3
u1,2

u1’s first three rows form one session (10:00 → 10:07 → 10:13, consecutive gaps ≤ 10m). The 37-minute idle stretch exceeds gap, so 10:50 starts a fresh session that runs through 10:55. Two sessions, two output rows.

Allowed lateness

allowed_lateness extends how long a window stays open past its end before it emits, giving late-arriving records a grace period to still be counted. It is distinct from the source-side watermark.delay. Records that arrive after a window’s end + allowed_lateness route to the DLQ as LateRecord with stage label time_window:<aggregate-name>. See DLQ category: LateRecord for the DLQ row layout.

- type: aggregate
  name: hourly
  input: clicks
  config:
    group_by: [user_id]
    time_window:
      tumbling: { size: 1h }
    allowed_lateness: 30s
    cxl: |
      emit n = count(*)

Default (unset) means no grace beyond the watermark — windows close the instant min_across_sources crosses window_end. Set allowed_lateness when the source’s watermark.delay alone is too small to absorb the observed out-of-order tail.

Worked example: multi-source session window

This pipeline merges two independent login feeds and groups per-user events into gap-bounded sessions. When several sources feed one time-windowed aggregate, a window cannot close until every source has advanced past the window’s end — the slowest source paces the others, so no window emits before all of its records have arrived.

pipeline:
  name: multi_source_session

nodes:
  - type: source
    name: src_web
    description: Web login events.
    config:
      name: src_web
      type: csv
      path: ./data/session_logins.csv
      options:
        has_header: true
      watermark:
        column: event_ts
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: source, type: string }

  - type: source
    name: src_mobile
    description: Mobile login events.
    config:
      name: src_mobile
      type: csv
      path: ./data/session_mobile.csv
      options:
        has_header: true
      watermark:
        column: event_ts
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: source, type: string }

  - type: merge
    name: all_logins
    inputs: [src_web, src_mobile]

  - type: aggregate
    name: user_sessions
    input: all_logins
    config:
      group_by: [user_id]
      time_window:
        session: { gap: 5m }
      allowed_lateness: 30s
      cxl: |
        emit user_id = user_id
        emit logins = count(*)

  - type: sink
    name: results
    input: user_sessions
    config:
      name: results
      type: csv
      path: ./output/multi_source_session.csv

Both sources declare their own watermark.column independently, and each source’s records carry an event time regardless of which feed delivered them — so the aggregate does not care which column name a given source used. A session cannot emit until both src_web and src_mobile have advanced past the session’s end plus its allowed_lateness. Drop the watermark: block on either source and the planner rejects the pipeline with E156.

Run it from the repo:

cargo run -p clinker -- run examples/pipelines/multi_source_session.yaml

Complete example

- type: source
  name: transactions
  config:
    name: transactions
    type: csv
    path: "./data/transactions.csv"
    schema:
      - { name: account_id, type: string }
      - { name: txn_date, type: date }
      - { name: amount, type: float }
      - { name: category, type: string }
    sort_order:
      - { field: "account_id", order: asc }

- type: aggregate
  name: account_summary
  input: transactions
  config:
    group_by: [account_id]
    strategy: streaming
    cxl: |
      emit total_amount = sum(amount)
      emit txn_count = count(*)
      emit avg_amount = avg(amount)
      emit max_amount = max(amount)
      emit categories = collect(category)

- type: sink
  name: summary_output
  input: account_summary
  config:
    name: summary_output
    type: json
    path: "./output/account_summary.json"

The output is JSON because collect(category) emits an array. JSON serializes it as a native array; XML can instead write it as repeated child elements, and CSV can join it into a delimited cell. For a scalar-only output such as fixed-width, coerce it first with a downstream Transform (for example emit categories = categories.join(";")).

Reshape Nodes

Reshape nodes observe a whole correlation group and, per group, mutate the rows whose state caused a rule to fire while synthesizing new rows derived from those trigger rows. They are the node for “look at everything an entity did, then fix one record and insert the record that should have been there” — work no other node can do:

  • Aggregate reduces a group to one summary row.
  • Transform emits 0 or 1 record per input record.
  • Combine joins records across sources.

None of those produce new records derived from a group’s observed state while preserving the originals. Reshape does.

Reshape is a blocking grouping operator: it buffers every record of a group before any output row leaves, because a rule cannot decide what to synthesize until it has seen the whole group. It has a single output.

Basic structure

- type: reshape
  name: backfill_plans
  input: plans
  config:
    partition_by: [employee_id]
    order_by:
      - { field: plan_start, order: asc }
    rules:
      - name: fix_long_plan_years
        when: "plan_start - plan_end > 365"
        mutate:
          set:
            plan_end: "plan_start"
        synthesize:
          copy_from: trigger
          overrides:
            status: "'synthesized'"

For each employee_id group, every row where plan_start - plan_end > 365 (the trigger) has its plan_end rewritten and a new row synthesized from it with status overridden to synthesized.

partition_by

A list of field names. Records sharing the same values for all partition_by fields form one group, and every rule observes and acts within a single group. This is the correlation key the operator reasons over.

Whole-input grouping (partition_by: [])

An empty list is the degenerate case of that key: every record shares it, so the entire input forms one group and the rules apply across the whole dataset rather than per entity.

- type: reshape
  name: relabel
  input: rows
  config:
    partition_by: []          # one group: the whole input
    order_by:
      - { field: amount, order: asc }
    rules:
      - name: flag_large
        when: "amount > 100"
        mutate:
          set:
            label: "'large'"

order_by then sorts the whole input as a unit, and the no-cascade contract applies across every record at once. Two consequences follow from there being only one group:

  • The whole input must fit memory.limit at finalize, since a single group has to be resident when its rules fire (see Limits). Whole-input grouping is for datasets that fit the budget, not for bulk row-at-a-time work — a per-record mutation with no cross-record dependency belongs in a Transform, which streams.
  • A mutation conflict rolls back the whole group, so one conflict rolls back the entire run’s records rather than one entity’s.

Values Reshape cannot key

Reshape groups a record under a single null group whenever it cannot build a key from the partition value. Today that covers all of:

  • the column being absent from the record
  • an explicit null
  • an empty string ("")
  • an array- or map-valued cell (a multi-value column)

A NaN float, of either sign, is one group of its own, not part of the null group. Numbers group by exact value, as in Aggregate: 1, 1.0 and the decimal 1 are one group, and distinct integers never merge, however large.

So account="" and account=null land in the same Reshape group. Note that Cull does not fold empty strings — there, account="" and account=null are two groups. The difference is unintentional and tracked in #1022; until it is resolved, do not assume one node’s grouping matches the other’s for blank values.

If a blank-heavy column is your partition key, expect one large null group. Normalize blanks upstream (a Transform that maps "" to a real sentinel) when you want them grouped separately.

order_by

Optional. A list of sort fields applied within each group before rules run, so order-dependent synthesis is deterministic. Arrival order breaks ties.

Each entry is either a field name, which sorts ascending, or a map { field, order, null_order }:

KeyValuesDefault
fielda column of the inputrequired
orderasc or descasc
null_orderfirst or lastlast
order_by:
  - plan_start                                  # same as { field: plan_start }
  - { field: plan_end, order: desc, null_order: first }

Each group is ordered this way before its rules run, exactly as a Sink sort_order orders rows (see How values are ordered): a null goes where null_order puts it, last by default for asc and desc alike; numbers compare by their exact value across integers, floats and decimals; every NaN is one value after every number, and so comes first under desc; -0.0 equals 0.0; and rows the order calls equal keep their arrival order.

null_order: drop is rejected when the pipeline is planned. order_by only arranges the rows of a group; it never removes any. To leave out rows whose field is null, filter them in a Transform before the Reshape:

- type: transform
  name: started_only
  input: plans
  config:
    cxl: |
      filter not plan_start.is_null()

Rules

Each entry in rules: is a declarative rule with a name, a trigger predicate, and optional mutation and synthesis actions. Rules are evaluated in declaration order — but every rule observes the same original group snapshot (see No cascade below).

when — the trigger predicate

when is a CXL boolean expression evaluated against each row in the group. A row for which when is true is a trigger row for that rule: its mutate rewrites it, and its synthesize derives new rows from it. CXL boolean operators are and / or / not (Clinker’s expression language is not SQL).

mutate — in-place trigger-row mutation

mutate:
  set:
    plan_end: "plan_start"
    note: "concat(note, ' (corrected)')"

Each set: entry is field: <CXL expression>. The expression evaluates against the original trigger row and overwrites that field’s value on the row.

Two restrictions are enforced at compile time:

  • A set: target must already exist in the upstream schema. Reshape mutates existing columns; it does not add new ones. Emit the column from an upstream Transform first if you need it.
  • A set: may not write a partition_by field — group identity must survive Reshape.

synthesize — deriving new rows

synthesize:
  copy_from: trigger
  overrides:
    plan_date: "'2024-01-01'"
    status: "'synthesized'"

For each trigger row, synthesize emits one new row:

  • copy_from: trigger — the new row starts as a copy of the trigger row’s values, then overrides are applied on top.
  • copy_from: none — the new row’s columns start null; overrides must supply every column (enforced at compile time, so a synthesized row is never silently empty).

Each overrides: entry is field: <CXL expression>, evaluated against the trigger row.

Whichever copy_from a rule uses, a synthesized row keeps its trigger row’s $source.* identity — $source.name, $source.file, and $source.event_time. Downstream, $source.name on a synthesized row names the Source its trigger came from, and if a later stage dead-letters the row, the entry is attributed to that Source: it counts toward that Source’s per_source rate limit and is written to that Source’s per_source dead-letter file when one is configured. See Error Handling & DLQ.

No cascade

Every rule observes the original group state. A row mutated by rule A is not re-observed by rule B, and rule B’s when predicate sees the row’s original values, not rule A’s edits. This is a deliberate guarantee:

  • Determinism — cascade would make rule order silently change output.
  • Single observation — the group is observed once.

To sequence dependent transformations, chain two Reshape nodes in the DAG so the second observes the first’s output.

Mutation conflicts

If two rules write the same field on the same row, that is a mutation conflict. Some conflicts are caught at compile time when the rules’ selectors statically overlap; content-dependent collisions that cannot be proven at compile time are caught at runtime.

A runtime conflict routes a dead-letter-queue entry under the mutation_conflict category, and the whole correlation group rolls back — none of that group’s mutated or synthesized rows reach the output. The DLQ entry’s stage label is reshape:<node>:<rule_a>+<rule_b>, naming the colliding rule pair. See Error Handling & DLQ.

Known issue: today only the row where the conflict happened gets a DLQ entry. The group’s other rows are dropped without one, so they appear in neither the output nor the DLQ (#1349).

Audit stamps

Reshape stamps three engine-written columns on its output records so the provenance of a synthesized or mutated row is queryable downstream:

ColumnMeaning
$meta.synthetictrue on a synthesized row, false on an original (including a mutated trigger row)
$meta.synthesized_bythe <node>:<rule> that synthesized the row (empty on originals)
$meta.mutated_bycomma-separated <node>:<rule> labels of every rule that mutated the row (empty if none)

Like the $ck.* correlation columns, these $meta.* columns stay out of the default writer output — they are available for downstream CXL and audit, not silently dumped into your output files.

Memory model

Reshape is a blocking, grouped operator: it groups every input record by partition_by before any rule fires, because each rule must observe its whole correlation group (the no-cascade contract forbids folding a group incrementally). It therefore cannot stream — the full group set materializes before the first output row leaves.

That per-group buffer is governed by the same central memory arbitrator every other blocking operator polls (see Memory & Spill). As records are grouped, Reshape tracks the live in-memory footprint and, whenever the run crosses the soft spill threshold (80% of memory.limit by default), it spills buffered groups to disk:

  • What spills: the raw input records, never the post-processed output rows. On reload at finalize, mutation and synthesis re-run in memory exactly as they would have without spilling, so the output is identical whether a group stayed resident or round-tripped through disk — including its within-group row order, which is restored to arrival order after a reload. (Two caveats apply; see Limits below.) Spilling input records (rather than output rows) is also what keeps a copy_from: none synthesized row — built against the wider output schema — from being reconstructed against the wrong schema; input records all share one uniform schema.
  • Spill priority: 15, between grace-hash Combine (10) and external sort (20). A grouped record buffer costs more to evict than grace partitions (reload pays the re-synthesis CPU) but less than an external-sort merge. Reshape cannot back-pressure — once its predecessor has drained there is no upstream channel to pause — so under memory pressure it always spills its own buffer in-thread rather than pausing a producer.
  • Largest-first, stop at the threshold: when the budget trips, Reshape evicts the largest resident groups first and stops as soon as the resident footprint drops back under the soft threshold — it does not drain every group. The cross-group resident peak is bounded near the soft limit.
  • Skew (one giant group): a single correlation group whose resident tail alone exceeds the budget is spilled incrementally — sliced by the upper bits of each record’s admission sequence so successive spill waves evict fractions of the one group — while smaller groups stay resident. The ingest-time resident peak therefore stays bounded even under one giant skewed group.

Reshape’s buffer is byte-budget-only: it does not honor the error_handling.max_group_buffer record-count cap. That cap bounds the correlation-commit group buffer used by the retraction machinery, a different buffer; Reshape’s footprint is bounded by memory.limit and the spill path, not by a per-group record count.

The on-disk spill volume Reshape produces is surfaced per stage in clinker run --explain (the Estimated spill volume and Spill compression sections) and, after a run, in the actual per-stage spill totals.

Limits

Two current limitations qualify the “identical whether spilled or resident” guarantee above:

  • A single correlation group must fit the memory budget at finalize. The no-cascade contract requires the whole group to be resident when its rules fire, so even though cross-group and ingest-time peaks spill to disk, the finalize reload of one group needs that group to fit. Skew slicing bounds the ingest peak, but a single correlation group larger than memory.limit has no in-budget representation. Rather than risk an out-of-memory crash, the run fails loud with E310 MemoryBudgetExceeded, naming the Reshape node, the offending partition_by group, and its footprint against the budget:

    E310 backfill: arena exceeded budget (128000/8192) [one Reshape correlation
    group [employee_id="employee-00000"] does not fit; the reported use is that
    group's reload footprint alone. ...]
    

    Raising memory.limit is the only fix that leaves your output unchanged. Raise it clear of the reported figure — finalize also holds the run’s remaining groups, so that figure is a floor, not a target.

    The other two levers both change what you get, and are worth knowing only so you can weigh them deliberately:

    • Dropping columns this node does not read, in an upstream Transform, shrinks each buffered record. But Reshape’s output row is the upstream columns plus its $meta.* audit columns — it never drops a column itself — so anything you strip upstream is also missing from the written output.
    • Narrowing partition_by shrinks the group too, but that key defines the group the rules evaluate against, so a narrower key changes which records each rule sees and therefore changes your results. Treat the grouping key as a modelling decision, never as a memory knob.

    (A future two-pass finalize could lift this limit.)

    Run clinker explain --code E310 for remediation keyed to whichever memory surface overran.

  • Reshape rules cannot reference $doc document context. Because the spill round-trip does not yet preserve document envelope context, a $doc.* reference in a rule’s when, mutate.set, or synthesize expression would resolve to the real envelope for a resident group but to null for a spilled one — output that depends on the memory budget. A pipeline whose Reshape rules reference $doc is rejected at compile time. Move the $doc lookup into an upstream Transform that copies the value into a record column, then reference that column in the Reshape rule.

Cull Nodes

Cull nodes observe a whole correlation group and remove the entire group when a group-level predicate holds, routing the removed records to a side-output port instead of discarding them. They are the node for “look at everything an entity did, then set the whole entity aside for review or reprocessing” — work that operates on an aggregate property of a group, not on individual rows:

  • Route fans out per record on a row predicate.
  • Reshape mutates rows and synthesizes new ones within a group.
  • Aggregate reduces a group to one summary row.

None of those remove a whole correlation group based on an aggregate property of the group and emit the removed rows on a second stream. Cull does.

Cull is a blocking grouping operator: it buffers every record of a group before any output leaves, because a group-level predicate (e.g. “this group has more than 100 rows”) cannot be decided until the whole group is seen. It has two output ports: the main port (kept groups) and the removed_to side-output port (removed groups).

Basic structure

- type: cull
  name: flag_large_histories
  input: backfill
  config:
    partition_by: [employee_id]
    removed_to: review
    rules:
      - name: too_many_plans
        drop_group_when: "count(*) > 3"

For each employee_id group, if the group holds more than three rows the whole group is routed to the review side output; every other group flows to the main output.

partition_by

A list of field names. Records sharing the same values for all partition_by fields form one group, and the removal predicate observes one whole group at a time. This is the correlation key the operator reasons over. partition_by must cover every visible correlation-key field so group identity is preserved on both output ports.

Whole-input grouping (partition_by: [])

An empty list is the degenerate case of that key: every record shares it, so the entire input forms one group and drop_group_when decides the whole dataset at once — every record is kept, or every record is routed to removed_to.

- type: cull
  name: drop_bad
  input: events
  config:
    partition_by: []          # one group: the whole input
    removed_to: removed
    rules:
      - name: drop_any_error
        drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"

That pipeline routes all records to removed if any single record has status: error. Keyed by [account] the same rule would remove only the offending account’s records — so reach for whole-input grouping when the decision genuinely concerns the batch (an all-or-nothing gate on a delivery), not when you meant a per-entity rule.

The whole input must then fit memory.limit at finalize, since a single group has to be resident when its predicate runs (see the limit below). The sibling group-count bound does not apply: one group is as few as the decision state can be.

Values Cull cannot key

Cull groups a record under a single null group in two cases: the column is absent from the record, or its value is an explicit null.

Everything else is either its own group or a hard error:

  • An empty string ("") is its own group, distinct from account=null.
  • Numbers group by exact value, as in Aggregate: 1, 1.0 and the decimal 1 are one group, and distinct integers never merge, however large.
  • A NaN float, of either sign, is one group of its own, distinct from the null group.
  • An array- or map-valued cell aborts the run rather than grouping — a partition key must be a single scalar value. This abort currently presents as an internal error, but it is a data condition, not an engine defect: fix the offending column rather than treating the message as an engine invariant failure.

Reshape treats empty strings and multi-value cells differently: it folds both into its null group instead. The blank-versus-null divergence is tracked in #1022. The multi-value behavior is separate; until both rules are deliberately aligned or retained, do not assume one node’s grouping matches the other’s.

order_by

Optional. A list of sort fields that orders the rows of each group as they are written. It does not change which groups are removed: the removal rule is evaluated over the group in arrival order (#1264). Arrival order breaks ties.

Each entry is either a field name, which sorts ascending, or a map { field, order, null_order }:

KeyValuesDefault
fielda column of the inputrequired
orderasc or descasc
null_orderfirst or lastlast
order_by:
  - txn_date                                    # same as { field: txn_date }
  - { field: amount, order: desc, null_order: first }

A group’s rows are ordered exactly as a Sink sort_order orders rows (see How values are ordered): a null goes where null_order puts it, last by default for asc and desc alike; numbers compare by their exact value across integers, floats and decimals; every NaN is one value after every number, and so comes first under desc; -0.0 equals 0.0; and rows the order calls equal keep their arrival order.

null_order: drop is rejected when the pipeline is planned. order_by only arranges the rows of a group; it never removes any, and every record of a group is kept or removed together. To leave out rows whose field is null, filter them in a Transform before the Cull:

- type: transform
  name: dated_only
  input: transactions
  config:
    cxl: |
      filter not txn_date.is_null()

Rules

Each entry in rules: is a declarative removal rule with a name and a drop_group_when predicate. A group is removed when any rule’s predicate holds (the rules are OR-combined).

drop_group_when — the group-level removal predicate

drop_group_when is a CXL boolean expression evaluated in aggregate context over the whole group (group-by = partition_by). Because it is an aggregate expression, it uses CXL’s aggregate functions:

AggregateMeaning
count(*)number of rows in the group
sum(<expr>)sum of an expression over the group
min(<expr>) / max(<expr>)minimum / maximum over the group
avg(<expr>)mean over the group
rules:
  - name: too_many_plans
    drop_group_when: "count(*) > 3"
  - name: high_total
    drop_group_when: "sum(amount) > 10000"

CXL’s bare aggregate vocabulary is sum / count / min / max / avg / collect / weighted_avg — there is no bare any() aggregate. To express “remove the group if any row matches a condition”, sum an indicator and compare to zero:

rules:
  - name: drop_error_groups
    # Remove any account group containing at least one `error` row.
    drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"

Ordered comparisons (>, <, >=, <=) work over every comparable aggregate type, not just numbers — the predicate uses the same comparison rules as a Transform. Numbers, strings, and dates all order:

rules:
  - name: late_alphabet
    drop_group_when: "max(name) > 'M'"          # string ordering
  - name: recent_hire
    drop_group_when: "max(hired) >= #2020-01-01#" # date ordering (`#YYYY-MM-DD#` literal)

A group whose aggregate operand is null (for example max(...) over an all-null column) compares as false — a null operand never removes the group and never errors.

Comments in a predicate

A drop_group_when predicate may carry a # line-comment to explain the rule inline. Each rule’s predicate is parsed on its own, so a trailing comment applies only to that rule:

rules:
  - name: drop_error_groups
    drop_group_when: "sum(if status == 'error' then 1 else 0) > 0  # any error row removes the group"
  - name: high_total
    drop_group_when: "sum(amount) > 10000                          # large accounts"

The comment is source text only — it never changes the compiled decision. (A #YYYY-MM-DD# date literal is unaffected: it lexes as a date, not a comment.)

Output ports: main and removed_to

Cull has two producer-side output ports, the same mechanism a Route uses for its branches — not the dead-letter queue. Removed records are valid rows the operator deliberately partitions onto a second stream, not errors.

Downstream nodes draw from the two ports by reference:

  • The main output (kept groups) is referenced by the Cull node’s bare name: input: flag_large_histories.
  • The side output (removed groups) is referenced as <cull>.<removed_to>: input: flag_large_histories.review.
  - type: sink
    name: kept
    input: flag_large_histories            # main port — kept groups
    config: { name: kept, type: csv, path: kept.csv }

  - type: sink
    name: review
    input: flag_large_histories.review     # side-output port — removed groups
    config: { name: review, type: csv, path: review.csv }

removed_to must be a non-empty name distinct from the Cull node’s own name (enforced at compile time, so the two ports are always distinguishable).

A single downstream node may draw from both ports — for example a Merge recombining the kept and removed streams (inputs: [flag_large_histories, flag_large_histories.review]) — and receives the union of both ports’ records.

removed_to is not the DLQ

The removed_to port carries the unchanged upstream schema — Cull does not widen, and both ports emit exactly the input columns. Removed records are not DlqEntrys and never appear in the dead-letter queue or its counters; they flow down a normal data edge to whatever node draws the removed_to port (an audit sink, a reprocessing branch, another transform). Use the DLQ for errors; use a Cull side output for valid records you want to handle separately. See Error Handling & DLQ for the error path.

Memory model

Cull is a blocking, grouped operator: it groups every input record by partition_by before any record leaves, because the group-level drop_group_when predicate is an aggregate property of the whole group and cannot be folded into a per-record keep/remove decision. It therefore cannot stream — the full group set materializes before the first output row leaves.

That per-group buffer is governed by the same central memory arbitrator every other blocking operator polls (see Memory & Spill). As records are grouped, Cull tracks the live in-memory footprint and, whenever the run crosses the soft spill threshold (80% of memory.limit by default), it spills buffered groups to disk:

  • What spills: the raw input records. On reload at finalize, each group is re-split onto its output port exactly as it would have been without spilling, so the output is identical whether a group stayed resident or round-tripped through disk — including within-group row order, which is restored to arrival order after a reload. (The per-group removal decision is computed from an in-memory aggregate over the same records; that aggregate state is O(distinct groups) and is never spilled — only the raw records spill. It cannot spill because Cull has no upstream channel to pause, so instead it is bounded by a hard check: if that state plus the run’s other live charged memory would exceed memory.limit, the run fails loud with a memory-budget error rather than growing it unbounded. See the limit below.)
  • Spill priority: 15, between grace-hash Combine (10) and external sort (20), matching Reshape. Cull cannot back-pressure — once its predecessor has drained there is no upstream channel to pause — so under memory pressure it always spills its own buffer in-thread rather than pausing a producer.
  • Largest-first, stop at the threshold: when the budget trips, Cull evicts the largest resident groups first and stops as soon as the resident footprint drops back under the soft threshold — it does not drain every group.
  • Skew (one giant group): a single correlation group whose resident tail alone exceeds the budget is spilled incrementally — sliced by the upper bits of each record’s admission sequence — while smaller groups stay resident, so the ingest-time resident peak stays bounded even under one giant skewed group.

The on-disk spill volume Cull produces is surfaced per stage in clinker run --explain (the Estimated spill volume and Spill compression sections) and, after a run, in the actual per-stage spill totals.

Limit: a single group must fit the finalize budget

Cull evaluates its group-level predicate against the whole group at once, so even though cross-group and ingest-time peaks spill to disk, the finalize reload of one group needs that group to fit the memory budget. Skew slicing bounds the ingest peak, but a single correlation group larger than memory.limit has no in-budget representation. Rather than risk an out-of-memory crash, the run fails loud with E310 MemoryBudgetExceeded, naming the Cull node, the offending partition_by group, and its footprint against the budget:

E310 drop_big: arena exceeded budget (512000/8192) [one Cull correlation
group [account="BIG"] does not fit; the reported use is that group's reload
footprint alone. ...]

Raising memory.limit is the only fix that leaves your output unchanged. Raise it clear of the reported figure — finalize also holds the run’s remaining groups and the per-group decision map, so that figure is a floor, not a target.

The other two levers both change what you get, and are worth knowing only so you can weigh them deliberately:

  • Dropping columns this node does not read, in an upstream Transform, shrinks each buffered record. But Cull filters rows, never columns — both its ports carry the unchanged upstream schema — so anything you strip upstream is also missing from the main and removed_to outputs.
  • Narrowing partition_by shrinks the group too, but that key defines the group drop_group_when evaluates over. With a rule like count(*) > 100, splitting one account across a finer key drops each resulting group below the threshold, so an account that should have been removed is emitted on the main port instead — the run “works” and quietly returns a different result set. Treat the grouping key as a modelling decision, never as a memory knob.

Run clinker explain --code E310 for remediation keyed to whichever memory surface overran, or see the memory guide.

There is a second, symmetric bound on the number of groups. The per-group removal decision is held in an in-memory aggregate that is O(distinct groups) and — unlike the raw records — cannot spill. If a partition key is so high-cardinality that the decision state plus the run’s other live charged memory would exceed memory.limit (many small groups rather than one giant group), the run likewise fails loud with E310, rather than growing that state unbounded. Here, coarsening partition_by is a legitimate fix only if the coarser key is the grouping you actually meant — the same caveat as above applies. Otherwise, raise memory.limit.

Several distinct memory surfaces can raise E310 against a Cull node, so read the [...] detail to see which one overran. The ones documented here are Cull correlation group [...] (the giant-group case above), Cull drop-decision aggregate state (the group-count bound), Cull cross-region tee admission (a downstream stage in a different deferred region forces this node’s output to be parked in memory), and node-buffer materialization overlap (a consumer must collect one of Cull’s sequential port scans into a resident vector). Cull’s main and removed_to handoff buffers themselves are spill-eligible, including when a port fans out. Other surfaces the shared runtime charges may name this node too — the detail string is the authority, not this list.

Envelope Nodes

Envelope nodes frame a body stream into per-document documents. An Envelope is a discrete, composable stage you can place after any operator — a Transform, a Merge, a Combine, or an Aggregate — to declare “from here on, treat the records as belonging to framed documents.” It mirrors the message/EDI/XML envelope-wrapper pattern (the Enterprise Integration Patterns Envelope Wrapper, XProc’s p:wrap-sequence): the body is the payload, and the envelope is the document boundary around it.

This page documents the preserve and concat framing strategies and the orthogonal header: / footer: synthesis that layers on top of either.

Interactive companion: How many documents? shows the documents a Sink writes under preserve and concat, with and without a synthesized footer, how an Aggregate after the Envelope rolls up, and which shapes E347 and E355 reject.

Basic structure

- type: envelope
  name: framed
  body: merged
  config:
    strategy: preserve

The node reads its body: input and emits the same records, framed per document. A downstream Sink with reconstruct_envelope: true then writes one framed document per body grain.

Inputs: body, the wired header, and the not-yet-wired trailer

InputRequiredStatus
bodyyesThe records to frame into documents.
headernoA 1-row-per-grain header stream. A wired value replaces each body document’s header with the matching header record, attached by document grain.
trailernoA stream whose records append to each framed document. Accepted in config but not yet wired — a wired value is rejected at plan validation this release.

When you omit header:, an Envelope frames each body record using the body’s own ambient envelope — the document context every record already carries from its source.

Wiring a header: port replaces the document’s header

A wired header: input is a second stream carrying one header record per document, each on the same document grain as the body it frames. The node attaches a header to a body strictly by grain, so the header record reaching it must carry the body’s grain — and the node then replaces that document’s ambient header with the wired header record (the framing grain is preserved; only the header changes). This is transform-in-place header replacement: rewrite a header’s values upstream — for example, override a batch id or stamp a run date — and frame the body with the rewritten header instead of the source’s original.

nodes:
  # `rewrite_header` rewrites the source header's values while keeping each
  # record's grain, so the rewritten header still grounds to its body document.
  - type: envelope
    name: framed
    body: payments
    header: rewrite_header
    config: { strategy: preserve }

The header record must carry a body document’s grain. A grain-preserving Transform of the source’s promoted header keeps it; a replacement from a different source establishes it via a business-key join against the body. A header record whose grain matches no in-flight body document (or carries the synthetic, ungrounded grain a Transform stamps onto a record it builds from scratch) cannot be placed, so the run fails with E351 (run clinker explain --code E351 for the full write-up):

envelope "framed": a wired header record carries document grain <grain>, which
matches no in-flight body grain (or is a synthetic / ungrounded grain). The node
attaches a header to a body strictly by grain, so it cannot place a header that
grounds to no body document.

Exactly one header record may carry each body document’s grain. When the wired header stream carries two or more records on the same grain, the node has no rule to fold a second header onto an already-framed document, so it refuses to silently keep one and drop the rest. The run fails with E352 (run clinker explain --code E352 for the full write-up):

envelope "framed": the wired header input carries two or more records for
document grain <grain> — exactly one header record per document grain is
required.

Deduplicate the header stream to one record per grain upstream — an aggregate or distinct on the grain’s business key, or a Transform that emits a single rewritten header per source document.

A wired trailer: input is still rejected this release with a clear “not yet supported” message:

envelope node "framed": explicit `trailer` input wiring is not yet supported —
omit it to frame with the body's own envelope

strategy: preserve

preserve emits one framed document per body grain. It is a transparent framing stage: body records pass through with their document context and grain unchanged, and the document-boundary signals are forwarded verbatim. Inserting a preserve Envelope between a body stage and a Sink is byte-identical to today’s per-document framing — its value is being the explicit, composable stage that later strategies extend, not a change in output.

preserve is the default, so config: { strategy: preserve } and an empty config: {} are equivalent.

Framing is keyed on the document grain, never the source file

The grain is the level at which one logical document is reconstructed — and it is not always one-per-file:

  • A nested X12 interchange frames once per interchange. The GS functional-group and ST transaction-set levels inherit the interchange grain, so an ISA … IEA interchange is one framed document regardless of how many groups or transaction sets it nests.
  • An HL7 multi-message file frames once per message. Each MSH message opens its own grain, so a single file containing several messages produces several framed documents.

Because framing keys on the grain rather than the source file, splitting or combining files never silently changes the document count.

strategy: concat

concat does the opposite of preserve: it collapses a multi-document body into one framed document. Every body record is re-stamped onto a single consolidated document context, so the body opens and closes exactly once regardless of how many documents fed in. This is the strategy to use when several source documents — say two files joined by a Merge — should write as one consolidated document with a single header and footer.

nodes:
  - type: merge
    name: both
    inputs: [file_a, file_b]
  - type: envelope
    name: framed
    body: both
    config: { strategy: concat }
  - type: sink
    name: out
    input: framed
    config:
      name: out
      type: csv
      path: out.csv
      reconstruct_envelope: true

Re-stamping changes only the framing (the grain) and the ambient $doc.* view a record sees — it does not disturb per-record fields. In particular $source.file is a real column stamped when each record is read, so it still reports the record’s own originating file after a concat. Concat is lossless on per-record provenance; it changes only which document the record is framed inside.

The consolidated header, and the two-headers conflict

One consolidated document can carry only one envelope header. concat derives it from the headers of the documents that contribute body records, taking one header per document:

  • Every header agrees (or there is only one) → the consolidated document carries that common header.
  • No document carries a header → the consolidated document is headerless.
  • A headed document and a headerless document → the single header wins; the headerless document coexists with it (no conflict).

Only documents that contribute body records take part: a document that carries a header but no body records frames nothing once consolidated, so it never enters the comparison. Header identity is structural — two documents share a header when they declare the same sections, in the same order, with the same field values. Engine-added fields whose names start with $ (such as $raw, the raw segment text) are ignored, so two files whose headers differ only in raw content fold to one header, and the consolidated document keeps the first document’s full header, $raw included. A difference in a field you declare (for example an extracted control number) makes the headers distinct.

When the body carries two or more distinct non-empty headers, concat refuses to silently keep one and drop the rest. The run fails with E350 (run clinker explain --code E350 for the full write-up):

envelope "framed": concat collapses the body into one framed document, but the
body carried 2 distinct non-empty envelope headers — one document can frame only
one header, so concat will not silently drop the rest. Make the headers identical
upstream, or add a header-folding strategy that declares which header the
consolidated document keeps.

To resolve a conflict, either keep the documents separate with preserve, or make the headers identical upstream (project them to the same sections and values). header: synthesis (below) does not resolve it: concat compares the input headers and raises E350 before any synthesis runs (#1385).

header: and footer: are orthogonal to the framing strategy. The strategy decides how many output documents there are (preserve = one per body grain, concat = one consolidated); synthesis decides what header and footer each of those documents carries. The node computes a fresh header (declarative scalar expressions) and footer (streaming aggregates over the framed body) per output document, stamps them as named sections into the document’s envelope, and the same header_from_doc / footer_from_doc writer path renders them.

Both maps are keyed section -> field -> CXL expression. The inner field map preserves declaration order, which is the rendered cell order. A downstream Sink renders a section through header_from_doc / footer_from_doc, and those may name only a section that a feeding Source declares (E346), so in practice a synthesized section reuses a declared section name.

A synthesized section replaces the whole same-named section on the document; it does not add fields to it, so the section’s original fields (for example a source’s interchange.tag) are gone from that document. Other sections ride through untouched. Give header: and footer: different section names: if both name the same section, the header is applied last and replaces the footer, whose fields are lost without an error (#1388).

- type: envelope
  name: framed
  body: merged
  config:
    strategy: concat            # or preserve — synthesis works the same on either
    header:                     # section -> field -> scalar CXL
      group:                    # a different section from the footer's
        sender: $vars.sender_id
        created: $pipeline.run_date
    footer:                     # section -> field -> aggregate CXL
      interchange:
        record_count: count()
        total: sum(amount)

Header fields: evaluated once at document open

A header: field is a scalar expression evaluated once per output document, before the body streams. It may read only inputs known at document open — pipeline configuration ($vars), per-document provenance ($source), pipeline-scope state ($pipeline), and the ambient envelope ($doc.*). It may not read a body column, because there is no “current body record” when the header is emitted; a body-column reference is rejected at compile time with E353 (run clinker explain --code E353). Put body-derived values in a footer aggregate instead.

A footer: field folds the document’s body records into a footer value at the document’s close. The fold is an O(1) accumulator per open document, so it supports exactly the streaming distributive/algebraic aggregates over a bare field or *:

AggregateExample
countcount(), count(*)
sumsum(amount)
avgavg(amount)
minmin(amount)
maxmax(amount)

Anything else is rejected at compile time. A holistic or unbounded aggregate (collect, any, weighted_avg) or a composed/multi-argument aggregate (sum(amount * 1.1), weighted_avg(value, weight)) raises E354 (run clinker explain --code E354); a non-aggregate function (median, mode) fails earlier as an unknown function. Project a composed value in an upstream Transform, or compute a holistic value in an upstream Aggregate, then aggregate the bare column here.

Because the strategy sets the grain cardinality, the same footer: { interchange: { record_count: count() } } produces a different result on each strategy over the same body:

  • Under strategy: preserve, each body document frames its own output document, so each footer’s record_count is that document’s body count.
  • Under strategy: concat, the whole body collapses to one document, so the single footer’s record_count is the merged body count.

Placement

An Envelope is a normal single-input, single-output node — put it anywhere a record stream flows:

nodes:
  # … sources, a Combine that joins two streams into `merged` …
  - type: envelope
    name: framed
    body: merged
    config: { strategy: concat }   # preserve here is rejected with E347
  - type: sink
    name: out
    input: framed
    config:
      name: out
      type: csv
      path: out.csv
      reconstruct_envelope: true

After a Combine, Aggregate or Composition, a Sink with reconstruct_envelope: true needs a concat Envelope in between: those nodes’ rows carry no document of their own, and concat frames them as one. A preserve Envelope there is rejected with E347. For an Aggregate that is right, since its rows would pass through unframed; for a Combine the check is stricter than it needs to be, because the joined rows do keep their driver row’s document (#1384).

After a Combine this works: the joined rows keep their driver row’s document, and concat writes them as one framed document. After an Aggregate it plans without error but does not frame yet: Aggregate rows carry no source file, so the consolidated document has none either, and the Sink writes the rows with no header or footer (#603). Reshape output loses its document the same way (#1317).

Memory model

Both strategies re-park the body into the node’s own buffer slot, which the engine’s memory arbitrator governs and spills to disk under pressure — so neither strategy is bounded by total input size held in RAM.

preserve is a transparent framing pass-through: it forwards records and their document boundaries unchanged. concat additionally re-stamps each record onto the one consolidated document context and replaces the per-document boundaries with a single open/close pair; the header consolidation it does first groups the body records by document — one document’s worth of body records shares one grain and one header — so it does one pass over the records to collect the distinct headers, comparing only one envelope per document (the work is bounded by the number of documents, not the number of body rows).

header: / footer: synthesis adds, on top of the materialized body, one O(1) accumulator per footer field per open document — every allowed footer aggregate (count / sum / avg / min / max) holds a fixed-size state regardless of how many body rows it folds. So a document’s footer state is independent of its body-row count, and synthesis stays within the node’s bounded-memory model.

Sink Nodes

Sink nodes write processed records to files. They are the terminal nodes of a pipeline – every pipeline path must end at a sink (or records are silently dropped).

Terminal-node migration: type: output is retired and rejected with E376. Replace only the terminal discriminator with the paste-ready correction:

- type: sink

Composition and node output ports, produced artifacts, files and paths, serialization formats, stdout, command or machine output, writer results, and OpenLineage output datasets keep the word “output.” See Production Contracts for the compatibility boundary.

Use the same name for the Sink node and its config.name. At the current implementation boundary, mismatched names can produce a writer-mode error or empty published files instead of a planning diagnostic. The examples use matching names; planning success alone does not establish correct output for a mismatched pair.

Basic structure

- type: sink
  name: result
  input: transform_node
  config:
    name: result
    type: csv
    path: "./output/result.csv"

The type: field selects the output format: csv, json, xml, fixed_width, edifact, x12, hl7, or swift. The edifact, x12, and swift writers reconstruct one interchange/message envelope around emitted records; the hl7 writer re-emits HL7 v2 segments and optionally wraps them in batch/file envelopes. See EDIFACT Format, X12 Format, HL7 v2 Format, and SWIFT MT Format.

Structured single-writer outputs (edifact, x12, hl7, and swift) accept one concrete document grain per output file. A multi-file source or multi-input merge feeding one of these outputs is rejected instead of being silently written as one merged envelope. To write multiple structured documents, consolidate them deliberately with an Envelope node first or route each document to a separate output path.

Local and network-share destinations

An output path may be on a local filesystem or a mounted NFS/SMB share. Clinker detects the filesystem behind the actual destination and applies its contained-create and same-filesystem promotion rules there; users do not label their production paths with a CI profile. The committed filesystem matrix qualifies Clinker’s semantics against specific loopback NFSv4.1 and SMB3.1.1 mounts, but it cannot certify every vendor appliance, mount option, outage mode, or corporate network. Qualify representative production mounts before depending on atomic promotion during an outage or failover.

Clinker creates Unix output files with owner-only mode 0600. This prevents a new file from accidentally inheriting broad access in a shared drop zone. If a different service account or group must consume the result, arrange that access explicitly with the destination’s ACL/ownership policy; Clinker does not currently expose an output-mode setting.

For performance, keep spill files and optional staged input copies on a local disk when one is available. Blocking operators can create substantial random I/O, and performing that work directly on a network share adds latency and network traffic. The final output is still written as a hidden file on the destination filesystem and promoted there, so the completed file never relies on a cross-filesystem rename from local storage.

This commit lifecycle applies to single-file, per-source-file fan-out, and split: outputs. Clinker does not open or truncate an existing final while a replacement is running. Before publication, Clinker synchronizes and validates the complete output set, then promotes each hidden file directly to its final name. An overwrite is one atomic replacement rename: Clinker never moves the previous final out of the way first. It also never claims that a multi-file set can be rolled back after some replacements are already visible.

Publication is not one atomic filesystem operation for the whole set. Each individual rename is atomic, but a reader may briefly observe a mixture while the finite commit walks several destinations. If a promotion or directory sync fails, Clinker stops, exits 4, and reports three exact groups: finals that are visible and synchronized, finals that are visible but whose parent sync failed, and unpublished hidden partials. Already-visible finals stay visible; remaining old finals stay untouched. A process or machine crash in that window can leave the same mixed set plus .partial or .reservation siblings. Reconcile those named paths before retrying or consuming the set. Clinker does not create or use .backup files for output publication.

Every collision policy uses a hidden sibling reservation, including overwrite. This ensures only one live publisher may mutate a final destination. if_exists: error uses a no-replace promotion; if_exists: overwrite and clinker run --force replace only at successful promotion; if_exists: unique_suffix reserves candidate names until one wins. Reservations never expose zero-byte final placeholders. A reservation holds an operating-system lock and records its owner process. A later run reclaims it only after a short creation grace period and only when the lock is acquirable, proving that no live publisher owns it. If reservation cleanup fails after successful publication, Clinker exits 4 and names both the visible final and stale reservation as cleanup debt.

When unique_suffix can find no name at all because the destination itself refuses every candidate — the directory is not writable by this run — the diagnostic names the path you wrote, not the numbered candidate the search happened to stop on, and says that the destination rather than the name is what refused. Fix the directory’s permissions, or point path: somewhere this run may write.

Rendered fan-out paths are validated as new output paths. Directory traversal, an absolute result produced from a relative template, symbolic-link/reparse ancestors, and cross-filesystem promotion fail before a final is touched. Create the intended destination directories ahead of the run; Clinker does not follow rendered paths while creating missing fan-out parents.

{source_file} and {source_path} create one output route per discovered source file. Two source files that render to the same destination are rejected before any output is staged, with both source paths in the diagnostic. Escape a token as {{source_file}} or {{source_path}} when the braces are intended as literal filename text. Runtime source names and paths are inserted as opaque text: braces inside an actual filename are never interpreted as another token. When fan-out is combined with split:, every source has its own segment sequence: each starts at sequence 1 and rolls over independently.

With write_meta: true, Clinker writes a .meta.json sidecar for every actual committed final. A split output therefore gets one sidecar per segment, and a fan-out output gets one per rendered destination; no sidecar is written for the unrendered base template. Main outputs and sidecars share the same publication ledger, so a path collision between any two of them fails before publication and names both producers. Counters that are not known at sidecar-preparation time are omitted from the JSON rather than written as misleading zeroes.

When two paths are the same destination

Two Sink nodes — or a Sink node and a DLQ path — that resolve to one file are rejected at plan time with E322, before any record is read. Deciding that means deciding when two differently-spelled paths name one file, and that depends on the volume you are writing to, not on the text. Clinker measures the volume rather than guessing from its type, by creating and removing a probe file in the destination directory.

Two paths are the same destination when they differ only in:

  • . components, or relative versus absolute spelling. ./out/errors.csv and /data/out/errors.csv from /data are one file everywhere.
  • A symlinked parent directory. The existing part of each path is resolved, so a link and its target are one place.
  • Letter case — only on a volume that ignores case. The default macOS (APFS) and Windows (NTFS) volumes do; ext4, xfs, and btrfs do not. Where it applies, it covers the whole of Unicode, not just ASCII: Ärger.csv and ärger.csv are one file, as are Σ.csv and σ.csv. straße.csv and strasse.csv are always two files — no filesystem treats them as one.
  • Unicode normal form — only on a volume that ignores it. APFS and HFS+ do, in both their case-sensitive and case-insensitive variants; ext4 and NTFS do not. Where it applies, café.csv written with a precomposed é and the same name written as e plus a combining accent are one file. This is independent of the case rule: a case-sensitive APFS volume ignores normal form while still telling Café.csv and café.csv apart.

Because the last two depend on the volume, the same pipeline can be accepted on Linux and rejected on macOS. That is not an inconsistency — the two disks really do behave differently, and the rejection is the one that prevented a file from being written twice.

Two limits are worth knowing:

  • Clinker may report a collision on a volume that would in fact have kept the two files apart — for instance on an older Windows volume whose case table predates a character you used. The run stops with both paths named, and renaming either one clears it.
  • The reverse is possible on Windows, which folds some letters according to Turkish and Azeri rules that no locale-independent table reproduces. İ.csv and i.csv may be one file on such a volume while Clinker still sees two. If you write output paths that differ only in dotted or dotless I, give them distinct names.

Direct broadcast to several outputs

Several Sink nodes may name the same input. This is a broadcast: every Sink receives every upstream record, regardless of node declaration order. The run report counts one write per sink, so five input records feeding a CSV and a JSON Sink produce records_written: 10.

Use a Route node when outputs should receive different subsets. Writing a field such as _route does not select a destination; it is an ordinary output column unless a Route condition explicitly reads it.

Field control

Sink nodes can either pass every upstream field through to the writer or restrict output to the fields the upstream transform explicitly emitted. Several options control which fields appear and how they are named.

Unmapped input field passthrough

    include_unmapped: false    # Default: true

When true (the default), every field on an input record that the upstream transform did not explicitly emit still passes through to the output unchanged. This includes fields the source’s on_unmapped: auto_widen policy absorbed into the per-record $widened sidecar map – their contents expand back to top-level columns at the sink.

When false, only fields named by an emit statement in the upstream transform appear in the output. The $widened sidecar slot is stripped and undeclared input fields are dropped.

When true, how a carried-along column reaches the writer depends on the output format. Self-describing formats (JSON / NDJSON / XML) write each record’s own keys. A CSV output widens its header to the union of every record’s columns when it can materialize the batch, and otherwise — on a bounded-memory streaming path (a Merge, a fused Transform, a single-branch Route, a streaming-strategy Aggregate, or the probe side of a hash-build-probe Combine feeding the output), or an envelope-reconstructing path — fails loudly with a SchemaDrift error rather than dropping a column it cannot fit under its already-committed header. A fixed-width output has no room for an undeclared column and likewise raises SchemaDrift. See Auto-Widen & Schema Drift → Schema drift across records.

Migration notice

The default flipped from false to true in a recent release (see issue #90). Pipelines that relied on the previous behavior – where output records contained only the fields explicitly emitted upstream – must now set include_unmapped: false explicitly to restore that shape.

The flag composes independently with include_correlation_keys: true – see below. See Auto-Widen & Schema Drift -> Output controls for the full specification and cross-format flow examples.

Worked example

Suppose the upstream source emits records with order_id, customer_id, amount, and region, and a transform that emits only one derived field:

- type: transform
  name: classify
  input: orders
  config:
    cxl: |
      emit amount_bucket = if amount >= 1000 then "high" else "low"

With include_unmapped: true (the default), each output record carries order_id, customer_id, amount, region, and amount_bucket. With include_unmapped: false, each output record carries only amount_bucket. The transform’s CXL is unchanged in both cases – the Sink node decides the field set.

Include correlation-key shadow columns

    include_correlation_keys: true    # Default: false

When a source declares a correlation_key:, the engine tracks correlation-group identity on hidden columns that are stripped from output by default. Set include_correlation_keys: true to surface them in the writer output — typically for debugging correlation-group routing or auditing DLQ behavior. See Correlation Keys.

include_correlation_keys does not surface auto-widened columns – include_unmapped is the separate flag for that. The two are independent: each, both, or neither can be set.

Nested columns and writer capabilities

JSON writes map and array values recursively. XML maps ordinary keys to child elements, reserves unescaped @.../#text keys for attributes/text, and writes an array at a top-level column only when that output-facing column comes from a multiple: true declaration; arrays nested inside a map remain native XML structure. CSV likewise joins an array only for a compiled multiple: true column. An undeclared array at either writer is a routing error, not an implicit declaration. CSV, fixed-width, EDIFACT, X12, and HL7 reject a map that reaches a column slot. See JSON, XML, and Auto-Widen & Schema Drift.

Field mapping

mapping: declares the columns the file carries – which columns, under what names, in what order – without changing upstream CXL. It is a sequence, one item per output column:

    mapping:
      - order_id                  # carried through under its own name
      - sold_to: customer_id      # written as `sold_to`, read from `customer_id`
      - contact_email: customer_email
      - channel
      - sku

Two item shapes:

  • A bare column name emits that column unchanged. This is the common case, and it costs one line naming the column once.
  • A single-key pair renames. The output name is on the left, the source column on the right – the same side the bare form names. Reading an item left to right always tells you what appears in the file first.

The renames are the only items carrying a colon, so in a wide output they are found by scanning for structure rather than by comparing two names per line.

Order and selection

Declaration order is the output column order. Listed columns are written first, in the order the block declares them, whatever order they arrive in.

include_unmapped governs everything the block does not list. With include_unmapped: true (the default) unlisted columns are appended after the declared ones, in their existing relative order. With include_unmapped: false they are dropped, so the block becomes the complete statement of the output:

    include_unmapped: false
    mapping:
      - department
      - surname: last_name
      - first_name

Given upstream columns first_name, last_name, department, that writes exactly department,surname,first_name.

Every record carries every declared column. When a record does not supply an item’s source column, that column is still written, empty. The file’s shape follows the block, not the data — so a stream whose records differ in shape (a multi-record-type source, a column arriving through auto_widen, a composition body’s open row) still produces one stable column set in declaration order, rather than one that depends on which record happened to arrive first.

One upstream column may feed two output columns – - sku and - item_code: sku – because names must be unique on the output side, not the source side. Declaring the same output name twice is rejected (E364): a file cannot carry two columns under one header.

For the same reason, an output name that include_unmapped: true would also carry through is rejected. If upstream already has a sold_to column, writing - sold_to: customer_id under include_unmapped: true would put two sold_to columns in the file and readers would resolve the wrong one. Rename the mapped column, exclude the upstream one, or set include_unmapped: false.

Where the compiler cannot enumerate the upstream columns, the same collision reaches the run. The mapped value wins – the block is your explicit statement of what the file carries – and the displaced upstream column is named in a W366 warning at the end of the run. Applying one of the three fixes above silences it.

Diagnostics

A mapping: item naming a column that does not exist at that point in the pipeline is rejected at compile time (E365), with the available column list and a did you mean when the name is a near miss. Nothing is renamed silently.

The compiler cannot always see the column set. Inside a composition body the rows are open by construction, and under on_unmapped: auto_widen a column can reach the sink through the sidecar without being declared anywhere. There an item naming an unknown column compiles even when its name resembles a declared column: spelling similarity cannot prove that a dynamic field is absent. W365 reports it after the run if no written record supplied it.

What catches the rest is the end of the run: if no record supplied an item’s source column, that item wrote an empty column in every row, and the run reports it as W365, naming the column to correct. An item some records supply and others do not is a sparse column, not a mistake, and is not reported.

Both W365 and W366 are advisory. They print to standard error when the run finishes and do not change the exit code – the file is written and readable either way, and by the time a stream ends the run’s other outputs have already been flushed.

A column absent from the source’s schema: reaches the sink only through the auto_widen sidecar, which is expanded to top-level columns only under include_unmapped: true. A mapping: item may name such a column when that flag is set; under include_unmapped: false it cannot resolve and is rejected at compile time.

An empty block – mapping: {} or mapping: [] – is rejected (E364): it declares an output with no columns. To write every upstream column, remove the mapping: key rather than emptying it.

Writing the block as a YAML map instead of a sequence is rejected (E364); the message prints your own block already rewritten. Run clinker explain --code E364 for the migration, and read the direction note there before pasting: releases before this one documented output_name: source_field but executed the reverse, so the rewrite swaps each pair’s two sides to preserve what the pipeline was actually writing.

Excluding fields

Remove specific fields from output:

    exclude: [internal_id, _debug_flag, temp_calc]

exclude: matches incoming column names, and runs before mapping:. Two consequences:

  • The columns that survive keep their relative order. Upstream a, b, c, d with exclude: [b] writes a, c, d.
  • Naming a column that a mapping: item also produces is not a conflict – the exclusion removes the upstream column of that name and leaves the mapped one standing. That is the fix for the two-columns-under-one-header collision above: - sold_to: customer_id with exclude: [sold_to] writes one sold_to column, carrying customer_id’s value.

Excluding a column a mapping: item reads is a different matter, and is rejected (E364): the exclusion removes the column before the item can read it, so the item could never resolve.

Header control (CSV)

    include_header: true      # Default: true

Set to false to omit the CSV header row.

Null handling

    preserve_nulls: false     # Default: false

When false, null values are written as empty strings. When true, nulls are preserved in the output format’s native null representation (e.g., null in JSON).

Rounding decimals to a declared scale

A Sink node’s optional schema: may declare a column type: decimal with a scale. A decimal value landing in that column is rounded to the declared number of fractional places on write, using banker’s rounding — the same boundary contract a decimal source column applies on read.

    schema:
      - { name: dept,    type: string }
      - { name: total,   type: decimal, scale: 2 }
      - { name: average, type: decimal, scale: 2 }

Decimals compute at full precision inside the pipeline (division and avg keep every digit), so a declared output scale is how you pin a computed result to fixed places at the sink: avg(amount) over 1.00, 1.00, 2.00 writes 1.33 into a scale: 2 column, while sum(amount) — already at scale 2 — stays 4.00. This works for every format (CSV, JSON, fixed-width); an output column with no declared scale, or an output with no schema: block at all, keeps the full-precision value. Only decimal values in decimal-declared columns are affected — no other type is coerced. See Decimal — arithmetic rules for the full boundary-contract model.

The same rounding applies to a Sink node declared inside a composition body. When its schema: names an external .schema.yaml file, the path resolves relative to the composition file’s own directory (not the invoking pipeline’s).

Output format options

CSV

- type: sink
  name: csv_out
  input: processed
  config:
    name: csv_out
    type: csv
    path: "./output/result.csv"
    options:
      delimiter: "|"

delimiter is a single byte on the wire, so it must be exactly one ASCII character (for example ,, |, or \t). An empty, multi-character, or non-ASCII value is rejected at plan validation rather than silently truncated to its first byte.

JSON

- type: sink
  name: json_out
  input: processed
  config:
    name: json_out
    type: json
    path: "./output/result.json"
    options:
      format: ndjson           # array | ndjson
      pretty: true             # Pretty-print JSON
  • array (default) – writes a single JSON array containing all records.
  • ndjson – writes one JSON object per line.

JSON numbers cannot represent non-finite floats; a record carrying NaN or an infinity fails the write with a JSON error instead of silently becoming null. See JSON Format.

XML

- type: sink
  name: xml_out
  input: processed
  config:
    name: xml_out
    type: xml
    path: "./output/result.xml"
    options:
      root_element: "data"
      record_element: "row"
      attribute_prefix: "@"    # emit @-prefixed fields as XML attributes

Fields whose final path segment carries the attribute_prefix (default @, matching the XML source option) are emitted as XML attributes of their enclosing element, so attribute fields read from an XML source round-trip. See XML Format for details.

Fixed-width

- type: sink
  name: fw_out
  input: processed
  config:
    name: fw_out
    type: fixed_width
    path: "./output/result.dat"
    schema: "./schemas/output.schema.yaml"
    options:
      line_separator: crlf

Fixed-width output requires a format schema defining field positions and widths. Fields land at their declared byte ranges with gaps space-filled — see Fixed-Width Format for the layout semantics.

EDIFACT

- type: sink
  name: edi_out
  input: messages
  config:
    name: edi_out
    type: edifact
    path: "./out/result.edi"
    options:
      interchange: ["UNOA:1", "SENDER", "RECEIVER", "240101:1200", "REF1"]
      message_type: "ORDERS:D:96A:UN"
      write_una: false
      segment_newline: true

The EDIFACT writer reconstructs the interchange envelope around emitted records, recomputing the UNT/UNZ control counts and echoing the control references, and release-escapes any element data that carries a service character. The UNB header comes from interchange (literal elements) or interchange_from_doc (echoed from a $doc section). An interchange is a single envelope, so an edifact output cannot be combined with a split: block — the combination is rejected at config-validation time (E323). See EDIFACT Format for the full option reference, the record schema, and the round-trip semantics.

HL7 v2

- type: sink
  name: hl7_out
  input: messages
  config:
    name: hl7_out
    type: hl7
    path: "./out/result.hl7"
    options:
      file_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "FILE7"]
      batch_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "BATCH3"]
      segment_newline: true

The HL7 writer re-emits the MSH and body segments from the record stream, escaping any field data that carries a delimiter character (| → \F\, ^ → \S\, and so on). When a file_header (or file_header_from_doc) or batch_header is configured the writer wraps the messages in an FHS..FTS file or BHS..BTS batch and recomputes the closing BTS/FTS counts. A batch/file envelope is a single structure, so an hl7 output cannot be combined with a split: block — the combination is rejected at config-validation time (E339). See HL7 v2 Format for the full option reference, the record schema, the MSH off-by-one, and the round-trip semantics.

Sort order

Sort records before writing:

    sort_order:
      - { field: "name", order: asc }
      - { field: "amount", order: desc, null_order: last }
Sort optionValuesDefault
orderasc, descasc
null_orderfirst, last, droplast
  • first – nulls sort before all non-null values.
  • last – nulls sort after all non-null values.
  • drop – records with null sort keys are excluded from output.

drop is available only on a Sink sort_order, because only a Sink’s ordering decides which records are written. A Source sort_order, a Cull or Reshape order_by and a Transform analytic_window.sort_by only order records, so they accept first and last and reject drop when the pipeline is planned, pointing at a filter not <field>.is_null() Transform instead. When CXL cannot name the field as it is (a name with a space, a CXL keyword such as filter, or a flattened Address.City), the error prints no CXL and points at the Source schema’s source_name rename, which gives the column a name the filter can use.

drop removes records, so a run using it writes fewer records than it read and that is not a fault. A missing column counts as a null key: a record that never carried the sort field is dropped the same as one carrying an explicit null. With several dropping fields, a record is excluded if any of its keys is null, and counts once however many of them are.

The excluded records are counted, separately from records_dlq and from filter losses, so a short output can be attributed rather than guessed at. A run that dropped any reports the number on completion:

1234 record(s) excluded by null_order: drop

and the same number is written as records_null_dropped in the metrics spool when one is configured (see Metrics).

Under fan-out the count is per exclusion, not per source record: two Sinks that each declare a dropping sort_order each drop their own copy, so one source record excluded at both counts twice — the same multiplicity records_written carries. Subtracting this from records_total is therefore only sound on a pipeline with a single dropping Sink.

Nothing else records these records. Unlike a DLQ entry, a dropped record leaves no artifact to inspect afterwards – if you need to see which records were removed rather than only how many, route them out with a filter before the sort instead of declaring drop.

Shorthand: a bare string defaults to ascending with nulls last:

    sort_order:
      - "name"
      - { field: "amount", order: desc }

A Sink sort_order materializes all records that reach that terminal and re-establishes one order across them, including records from several physical files or Merge inputs. The guarantee is exactly the authored field sequence, direction, and null placement. drop is also part of the authored contract: records with a null in a sort key do not reach the writer.

For a split Sink, that order is global across the complete numbered split set, not restarted independently inside each file. Clinker sorts and applies null_order: drop before it rotates the writer. Each numbered file is therefore a contiguous slice of the one ordered sequence, and concatenating the files in sequence-number order recovers that sequence. Dropped rows do not count toward max_records, max_bytes, or the resulting number of split files.

The sort is stable. Equal authored keys retain their upstream arrival order within a given execution path, and the same path produces the same bytes in resident and forced-spill operation. Clinker does not add a source-row, filename, or canonical-record tie-breaker. If upstream strategies can produce different arrival orders, equal-key rows have no cross-strategy relative-order promise. Author enough fields for a total business order before using an exact byte comparison; otherwise validate the decoded record multiset and aggregate values instead.

How values are ordered

Every sort uses one rule for comparing two values: a Sink or Source sort_order, a Cull or Reshape order_by, a window’s sort_by, and the check that verifies a Source’s declared order. The rule does not depend on the memory limit, so a sort that spills to disk writes the same records in the same order as one that fits in memory.

  • Nulls are placed only by null_order: first, last or dropped. A missing column counts as a null.
  • Values of one type order naturally: numbers by value, strings by UTF-8 code point (no locale collation, so "Z" sorts before "a"), false before true, and dates and datetimes chronologically. A leap-second datetime sorts with the instant one second later that has the same fraction.
  • Integers, floats and decimals compare by their exact value, not through a rounded floating-point copy. The integer 1, the float 1.0 and the decimal 1.00 are equal. The integer 9007199254740993 sorts after the float 9007199254740992.0, although the two round to the same float. The decimal 0.1 sorts before the float 0.1, whose exact binary value is slightly larger.
  • Zero has one position: -0.0 and 0.0 are equal.
  • NaN is one value, whatever its sign. It sorts after every number, inf included, in ascending order, and so comes first in descending order.
  • Values of different types, which a column can hold when an expression’s branches produce different types, order by type: booleans, then numbers, then strings, then dates, then datetimes, then arrays, then maps.

Values the rule calls equal keep their arrival order, as described above, at every memory limit.

Physical writer boundaries

Planning derives the writer boundary from the finalized graph, not from how many Sink nodes appear in the YAML. The same ordering promise is therefore enforced at every physical byte-emission path:

  • ordinary single-file and split-file record output;
  • one output per physical source file;
  • reconstructed envelope output per document;
  • document DLQ output after the whole document is known to be clean;
  • deferred output per correlation group; and
  • incremental streaming output.

Complete-population modes apply the exact authored key at their population boundary using the same bounded-memory spill path. Incremental streaming cannot truthfully promise a terminal whole-population sort. If a finalized output mode is incompatible with an authored sort_order, planning rejects the pipeline instead of weakening the promise. The diagnostic names the Sink, mode, authored keys, and last reordering stage, and includes a corrected sort_order form that can be pasted into the source or upstream node.

File splitting

Split output into multiple files based on record count, byte size, or group boundaries:

- type: sink
  name: split_output
  input: processed
  config:
    name: split_output
    type: csv
    path: "./output/result.csv"
    split:
      max_records: 10000
      max_bytes: 10485760           # 10 MB
      group_key: "department"       # Never split mid-group
      naming: "{stem}_{seq:04}.{ext}"
      repeat_header: true           # Repeat CSV header in each file
      oversize_group: warn          # warn | error | allow

Split configuration fields

FieldRequiredDefaultDescription
max_recordsNo–Soft record count limit per file
max_bytesNo–Soft byte size limit per file
group_keyNo–Field name – never split within a group sharing this key value
namingNo"{stem}_{seq:04}.{ext}"File naming pattern. It must contain exactly one {seq:NN} token, where NN is a decimal width from 1 through 20. {stem} is the base name and {ext} is the file extension.
repeat_headerNotrueRepeat CSV header row in each split file
oversize_groupNowarnWhat to do when a single key group exceeds file limits

At least one of max_records or max_bytes should be specified for splitting to have any effect.

The naming grammar is strict: {stem}, {ext}, and the one required {seq:NN} token are the only placeholders. Unknown placeholders, a bare {seq}, non-numeric or out-of-range widths, and duplicate or missing sequence tokens are rejected during configuration validation. For example, {stem}_{seq:03}.{ext} renders sequence 7 as 007.

For formats whose output wraps the whole file in framing – a JSON array or an XML root element – each split file is a complete, independently valid document: the framing is closed at rotation and reopened for the next file.

When the Sink also declares sort_order, splitting happens after the complete Sink population has been ordered and null-key drops have been applied. Segment 1 receives the first surviving records, segment 2 the next records, and so on. The files are individually ordered and together form one ordered sequence when read by sequence number; split rotation never starts a new independent sort.

Oversize group policies

  • warn (default) – log a warning and allow the oversized file.
  • error – stop the pipeline.
  • allow – silently allow the oversized file.

When group_key is set, the split point is the first group boundary after the threshold is reached (greedy). Without group_key, files are split at the exact limit.

Streaming writes after an interleave Merge

When a single Sink sits directly after a Merge with mode: interleave whose inputs are all Sources, records are written to disk as they arrive rather than being buffered until the merge finishes. This keeps memory flat and lets a slow writer naturally pace the upstream readers.

- type: source
  name: src_a
  config: { type: csv, path: a.csv, schema: ... }
- type: source
  name: src_b
  config: { type: csv, path: b.csv, schema: ... }
- type: merge
  name: merged
  inputs: [src_a, src_b]
  config:
    mode: interleave        # required
- type: sink
  name: out
  input: merged
  config:
    name: out
    type: csv
    path: out.csv

This is automatic — there is no setting to enable it. It applies only to this exact shape: one interleave Merge of Sources feeding one non-splitting Sink, in a pipeline without correlation keys. Any other topology buffers as usual. Both paths preserve the same record multiset and writer semantics, but an unseeded interleave does not promise one exact cross-input row sequence. Add a Sink sort_order with a total business key when exact bytes are required.

Complete example

- type: sink
  name: department_reports
  input: enriched_employees
  config:
    name: department_reports
    type: csv
    path: "./output/employees.csv"
    # `include_unmapped: false` makes the mapping the whole output: these four
    # columns, in this order, and nothing else. Without it every unlisted
    # upstream column would still be appended after them, and an `exclude:`
    # would be needed to keep any of them out.
    include_unmapped: false
    mapping:
      - "Employee ID": employee_id
      - "Full Name": display_name
      - department
      - "Annual Salary": salary
    include_header: true
    sort_order:
      - { field: "department", order: asc }
      - { field: "display_name", order: asc }
    split:
      max_records: 5000
      group_key: "department"
      naming: "employees_{seq:03}.csv"
      repeat_header: true

CSV Format

CSV is the default file format. The reader decodes each CSV row (including quoted multiline cells) into a record whose fields are matched positionally (or by header name) against the source’s declared schema:; the writer reverses the process. CSV pairs with the file transport — see Source Nodes for the transport / format split and the schema rules every source shares.

- type: source
  name: orders
  config:
    name: orders
    type: csv
    path: "./data/orders.csv"
    schema:
      - { name: order_id, type: int }
      - { name: customer_id, type: int }
      - { name: amount, type: float }
      - { name: order_date, type: date }
    options:
      delimiter: ","         # default ","
      quote_char: "\""       # default "\""
      has_header: true        # default true
      encoding: "utf-8"      # default "utf-8"

Declared types

CSV has no native numeric or date types: decoding first produces text cells. Source ingestion then parses and validates each declared column against schema: before buffering, sorting, or downstream CXL evaluation. With type: int, the cell 42 reaches a Transform as an integer; with type: string, the cell 0042 remains text, including its leading zeroes. No explicit CXL cast is needed for a column already declared with its intended type. Casts such as .to_int() are useful when a pipeline deliberately keeps text in its schema and converts it later.

A value that fails its declared type rejects the row under the configured error policy; it never silently falls back to text. See Declared-type failures.

Options

All CSV options are optional. With no options: block, Clinker uses standard RFC 4180 defaults.

OptionDefaultDescription
delimiter,Field separator, exactly one ASCII byte. Set to \t for TSV, ; for semicolon-delimited exports.
quote_char"Quote character that escapes delimiters and newlines inside a field, exactly one ASCII byte.
has_headertrueWhen true, the first line names the columns and is consumed, not emitted. When false, fields bind to schema: positionally.
encodingutf-8Character set each field — including the header row — is decoded through. Supported values are utf-8 (the default) and iso-8859-1 (aliases latin-1, latin1). See Encoding.

delimiter and quote_char are each a single byte on the wire, so each must be exactly one ASCII character. An empty, multi-character, or non-ASCII value (for example "||" or "→") is rejected at plan validation — it is never silently truncated to its first byte.

Encoding

The reader decodes every field through the source’s declared encoding:

  • utf-8 (the default) is strict — a byte sequence that is not valid UTF-8 fails the run loudly rather than substituting replacement characters, so a mis-declared encoding is caught instead of silently corrupting data.
  • iso-8859-1 (Latin-1; also spelled latin-1 or latin1) maps each byte 0xNN to codepoint U+00NN, so high bytes such as 0xE9 (é) from legacy exports decode correctly.

The same options.encoding setting is available on a CSV Sink. It applies to headers, body fields, joined values, embedded JSON cells and reconstructed envelope rows. UTF-8 is the default on both sides. Names are case-insensitive; hyphens, underscores and spaces are ignored. UTF8, ISO8859_1, latin1, Latin-1 and l1 resolve to the corresponding canonical spelling.

Latin-1 is true ISO-8859-1: bytes 0x80..0x9F are the corresponding control codepoints, not punctuation from another code page. Characters above U+00FF (for example €) cannot be written in Latin-1 and cause an error; there is no replacement character or fallback. UTF-8 input strips its leading UTF-8 BOM; Latin-1 preserves an initial EF BB BF sequence as the three ordinary characters . CSV grammar is parsed from bytes before either header or body text is decoded.

An unsupported encoding is rejected during configuration admission with a precise error naming the value and a supported correction. Charset is part of the semantic plan identity; changing input or output charset changes that identity. Aliases and an omitted or explicit UTF-8 default resolve identically. Malformed UTF-8 in a header or body cell is an input-data error (source.data.invalid), including when the header is read to discover columns.

The closed encoding policy across formats is:

FormatEncoding policy
CSV, X12Authored options.encoding: UTF-8 or true ISO-8859-1.
JSON, XML, fixed-width, SWIFT MTUTF-8 only; no authored encoding override.
EDIFACTIn-band repertoire: UNOA/UNOB ASCII, UNOC ISO-8859-1, UNOY UTF-8; no authored override.
HL7In-band repertoire: blank/ASCII means ASCII, UNICODE UTF-8 means UTF-8; no authored override.

Single-schema and multi-record CSV sources support the same two encodings, including textual column headers, discriminator fields, body cells and declared envelope sections. Each file applies its own leading-BOM rule. A UTF-8 BOM is recognized only at the start of that file; the same bytes inside a field are data.

Header handling

With has_header: true, the header row’s names bind input columns to the schema: entries — column order in the file may differ from the schema. With has_header: false, binding is strictly positional, so the schema order must match the file’s column order.

Input columns the schema does not name are governed by the source’s on_unmapped policy, the same as every other format.

Multi-value cells (split_values)

A CSV cell holds one string, but that string may pack several values behind a delimiter (1,a;b;c). Declare the column multiple: true and add a split_values entry naming the field and its delimiter, and the reader parses the cell into an array:

- type: source
  name: orders
  config:
    name: orders
    type: csv
    path: ./orders.csv
    split_values:
      - { field: tags, delimiter: ";" }
    schema:
      - { name: order_id, type: string }
      - { name: tags, type: string, multiple: true }

tags reads as ["a", "b", "c"]. An empty cell is an empty array; a cell with no delimiter is a one-element array; each element is coerced to the column’s declared type:. A quoted cell is unquoted first, so a delimiter inside the quotes is not a boundary. A multiple: true column with no covering split_values entry is rejected at compile (E361). Multi-record CSV sources reject split_values and multiple: true columns with E358 and E361; these input options require a single-schema source. See split_values in the Source reference for the full grammar.

Writing CSV

On output, the writer emits one row per record with cells in the output schema’s column order — the same order as the header row — regardless of how an upstream node ordered the record’s fields. An output-schema column the record does not carry emits an empty cell, the same as an explicit null; with include_unmapped: false a record field the output schema does not name is not written. See Sink Nodes for header control, field mapping, and null handling.

- type: sink
  name: export
  input: orders
  config:
    name: export
    type: csv
    path: ./out/orders.csv
    options:
      encoding: iso-8859-1

Each output operation is prepared completely before any of its bytes reach the destination. The first body row and its automatic header are one operation. An unrepresentable cell therefore writes neither a partial row nor a stray header, and leaves previously accepted bytes and format state intact. Explicit document start and end operations are prepared separately. An I/O failure during delivery can leave a destination prefix and prevents further writes; successful preparation alone is not a delivered row. File publication is a separate contract described in Storage & Spill Location.

CSV quoting needs each complete cell. Its encoding workspace is admitted against the run’s resource budget before allocation and released after that cell; there is no authored field-size or record-size ceiling. A cell that cannot fit fails as a resource error rather than being truncated or dropped. Joined or embedded JSON cells account for the rendered text and encoded bytes while both are live. The complete operation also needs storage for its prepared bytes: it stays in memory unless an explicit spill location is available. Spill does not eliminate the minimum memory needed for a cell, policy, or schema mapping. See Memory Tuning for the accounting boundary and output preparation for storage and cleanup behavior.

Split output captures only the header actually delivered by a successful operation. With repeat_header: true, later files replay those same names; with repeat_header: false, only the first file emits the automatic header. include_header: false suppresses that header throughout. Reconstructed envelope rows do not become an automatic column header.

An empty stream emits no automatic header. A single empty or null cell is written as "" followed by a newline; adjacent empty cells are separated by the delimiter. Reading an ordinary empty cell yields an empty string, so CSV does not distinguish an empty string from a null unless the pipeline applies its own schema policy.

Writing multi-value cells (join_values)

A multiple: field is joined into one delimited cell on write — the write-side inverse of split_values. The default needs no configuration: values join with ;, and a value that itself contains the delimiter is a hard error rather than a cell that would split back wrongly. The planner carries the exact output-facing multiple: true column set through mapping and exclusion into the writer. An array reaching any other CSV column is rejected as a routing/type-contract error rather than joined implicitly.

- type: sink
  name: report
  input: orders
  config:
    name: report
    type: csv
    path: ./out/report.csv
    join_values:
      - tags                              # delimiter ";", on_conflict: error
      - { field: notes, delimiter: "|", on_conflict: escape, escape: "\\" }

A field with no join_values entry still joins, with the defaults. An entry overrides, per field:

  • delimiter — the separator written between values (default ;).
  • on_conflict — what to do when a value contains the delimiter:
    • error (default) — reject the record with the field and element position, preserving its original value for the DLQ, rather than emit a cell that splits back wrongly. This is what makes a defaulted delimiter safe. Under error_handling.strategy: continue, the offending record goes to the dead-letter queue (category multi_value_join_collision) and the run continues; under fail_fast it aborts. The exception is a pipeline where any Source declares dlq_granularity: document: there the collision fails the run (see Not covered, #933).
    • escape — prefix each delimiter (and each escape character) inside a value with escape (default \), so a matching split_values escape: recovers the original. Lossless. delimiter and escape must each be a single character.
    • encode_json — encode the whole field as an embedded JSON array, recovered by a matching split_values json: true. Preserves every value’s text exactly, including ones carrying the delimiter, quotes, or newlines — nothing is lost or mis-split. (A decimal/date/datetime element serializes as its JSON string form and reads back as a string, re-typed by the column’s declared type:, the same round trip every CSV cell takes.)

An empty field emits an empty cell — and, under the delimited policies (error, escape), a single empty-string value [""] emits an empty cell too, which reads back as zero values: the delimited encoding cannot tell an empty field from one empty value. Use encode_json when that distinction matters. A single non-empty value emits that value with no delimiter. The joined cell is quoted by the normal CSV rules when it contains the field delimiter, a quote, or a newline. Declaring join_values on a non-CSV output is rejected at compile (E362).

Round trip. on_conflict: escape and encode_json are recovered exactly by a matching source split_values entry:

# write side
join_values:
  - { field: tags, on_conflict: escape, escape: "\\" }
# read side (a later pipeline)
split_values:
  - { field: tags, escape: "\\" }

Header widening under auto-widen

When auto_widen is in effect and the Sink leaves include_unmapped at its default of true, different records can carry different carried-along columns. The header must still be shared by every row, so Clinker widens it to the union of every record’s columns in first-seen order: a column that first appears on a later record still gets its own header slot, and the earlier rows write an empty cell for it. This pre-scan runs on the buffered output path, where the record batch is materialized.

An output that streams under a bounded-memory budget cannot pre-scan the whole batch: a CSV output fused directly after a Merge/Transform, a single-branch Route, a streaming-strategy Aggregate, or the probe side of a hash-build-probe Combine, or one reconstructing an envelope (which suppresses the shared header and streams a headerless body), commits its columns to the first record. A later record carrying a column that first record lacked then fails the run with a SchemaDrift error naming the column, rather than silently dropping it. Declare the column in the source or output schema: so every record carries it, or route to a self-describing format (JSON / NDJSON / XML). A reconstruct_envelope CSV output therefore requires a stable body shape — every record must carry the same columns.

Multi-record files (header / trailer / body)

Some CSV exports interleave multiple record types in one file — a header row, many body rows, and a trailer row — each distinguished by a discriminator column. Declare these with a map-form schema: carrying a discriminator: and a records: list, instead of the single column-list schema:. Each record type names its tag (the discriminator value that identifies it) and its own columns:; the discriminator field must sit at the same column in every type (usually the first). The reader derives the runtime superset schema (a lead record_type column plus the union of every record type’s columns) automatically.

- type: source
  name: payments
  config:
    name: payments
    type: csv
    path: "./data/payments.csv"
    schema:                                     # one multi-record schema (map form)
      discriminator: { field: rec_type }        # the physical column carrying the type tag
      records:
        - { id: header,  tag: H, columns: [ { name: rec_type, type: string }, { name: batch_id, type: string } ] }
        - { id: detail,  tag: D, columns: [ { name: rec_type, type: string }, { name: id, type: int }, { name: amount, type: int } ] }
        - { id: trailer, tag: T, columns: [ { name: rec_type, type: string }, { name: count, type: int } ] }
      structure:
        - { record: trailer, count: count }     # validate T's count against the body count
    envelope:
      sections:
        head:
          extract: { record_type: H }          # the H record type surfaces as $doc.head.*
          fields:
            batch_id: string

The reader emits one record per CSV row, including quoted multiline cells, on a single superset schema whose lead record_type column carries the matched type’s id. A downstream Route discriminates on that column. Rows of different record types may carry different column counts (ragged rows) — the reader validates the column count per record type, not file-wide. A textual column-header row is skipped when has_header is true (the default), so a leading record_type,name,amount line is not mistaken for a record of an unknown type. Each declared field honors its own type / trim / pad, the same as a single-record CSV field.

  • Header rows declared as an envelope: section via the record_type extract surface as $doc.<section>.* and are excluded from the body stream (see Envelopes & Document Context).
  • Trailer rows named by a structure: constraint are validated as they stream — the declared count field is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body row after the trailer is rejected as content past the document close.
  • Blank lines (empty or whitespace-only, common after concatenation) are skipped rather than parsed.
  • An unknown discriminator value (a tag no records: entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run under fail_fast; under continue with the default record granularity it dead-letters only that physical row and continues with the next row; under dlq_granularity: document it condemns the whole file. A record-grained DLQ row carries a JSON array of the decoded CSV cells in _cxl_dlq_source_record, preserving empty cells without guessing which declared layout the unknown tag meant.

JSON Format

The JSON reader turns a JSON document into a record stream. It handles three physical shapes — a single array of objects, newline-delimited objects (NDJSON), or a wrapper object that nests the records under a path — and auto-detects the shape when you do not declare it. Each object is matched against the source’s declared schema:; see Source Nodes for the shared schema and transport rules.

- type: source
  name: events
  config:
    name: events
    type: json
    path: "./data/events.json"
    schema:
      - { name: event_id, type: string }
      - { name: timestamp, type: date_time }
      - { name: payload, type: string }
    options:
      format: object          # array | ndjson | object (auto-detect if omitted)
      record_path: "data"     # dot-separated keys to the records array
      max_index_bytes: 64MB   # cap on retained envelope sections (optional)

Text encoding

Input is strict UTF-8. One leading UTF-8 BOM is accepted and removed at each physical file open, including an envelope pre-scan. UTF-16 and UTF-32 BOMs and malformed UTF-8 are rejected; convert the file to UTF-8 before running it. There is no JSON encoding option or lossy fallback. A later invalid file does not erase records already delivered from earlier files. Validation follows the reader: it does not promise to discover malformed bytes beyond what it reads.

Output is UTF-8 without a BOM. Ordinary format: ndjson always writes one compact object followed by exactly one LF, including the last record; pretty: true does not expand ordinary NDJSON across lines. pretty still controls array output and reconstructed envelope documents. Envelope framing is documented separately.

Physical shapes

formatLayout
arrayThe file is a single JSON array of objects.
ndjsonOne JSON object per line (newline-delimited JSON).
objectA single top-level object; record_path locates the records array within it.

If format is omitted, Clinker auto-detects the shape from the file content. Declare it explicitly when the file is large enough that you want to skip detection, or when an object wrapper needs a record_path.

record_path

record_path is a dot-separated path of object keys, descended from the document root. data.rows selects the array at {"data": {"rows": [ … ]}}, and each of its elements becomes one record. This is the canonical statement of the grammar; other pages link here rather than restate it.

The rules, in full:

  • No $. root marker. It is not JSONPath. Write data.rows, not $.data.rows. Only the exact leading $. is rejected, so a key that merely starts with $ ($schema.rows) is still addressable.
  • No leading /. A leading slash is how a JSON Pointer is anchored; record_path is already anchored at the document root.
  • No empty segments — no doubled separator (data..rows) and no trailing one (data.).
  • Omitting record_path entirely lets the reader auto-detect the document shape. That is not the same as record_path: "", which is a path naming a key called “” and is rejected.

A value breaking any of these fails at compile time with E363, before any input is opened. The diagnostic names the corrected path where one can be derived.

record_path takes precedence over format:. When both are declared the reader navigates the path and streams the array it finds, whatever format: says — so pair record_path with format: object (or leave format: off). Declaring format: ndjson alongside a record_path does not read NDJSON.

Because a JSON key may contain any character, the two rejected prefixes give up a sliver of addressing: a top-level key literally named $ followed by a nested key, and a top-level key whose name starts with /, are not reachable through record_path.

Nested arrays

JSON records frequently embed arrays — line items on an invoice, tags on a product. Three source-level declarations decide what happens to them, all documented on the Source Nodes page:

  • split_to_rows fans the array out to one record per element. mode: extract (the default) hoists an object element’s keys onto the output record; mode: split keeps the record shape, flattening the element back under the field name (orders.id). An array of scalars keeps the value under the field’s own name under both modes.
  • A schema column declared multiple: true keeps the array as an array, and normalizes a lone scalar into a one-element array so the column’s shape never depends on what a particular document happened to carry.
  • split_values parses a delimited string cell into several values.

A record whose declared field holds an empty array, is explicitly null, or carries no such field at all, is preserved by default — keep_empty defaults to true, and setting it to false drops such a record. An explicit null is how many producers write “no value”, so it counts as no occurrence rather than one; for the same reason a multiple: true column holding an explicit null stays null rather than becoming [null].

A field that IS present but holds a single object or scalar rather than an array is one occurrence, projected exactly as a one-element array would be. Producers routinely unwrap a lone element, so a feed where some documents carry "line_items": [{…}, {…}] and others carry "line_items": {…} fans both out the same way and every output record ends up with the same columns. The XML reader, where a document cannot express the difference at all, already behaved this way.

Two declared fan-out fields apply in declaration order and multiply. A nested pair (orders then orders.items) produces the two-level expansion when the outer entry declares mode: split:

    split_to_rows:
      - { field: orders, mode: split }
      - { field: orders.items, mode: split }

Under mode: extract the outer entry lifts the occurrence’s keys to the top level, which removes the orders.items path the inner entry addresses — so that pairing is rejected at compile (E358) rather than silently fanning out only one level. A duplicated field is rejected too.

Set source-level max_output_rows_per_input: N to bound the cumulative product without materializing it. The reader emits the first N rows in stable order, then a first attempted row above the ceiling routes the original JSON object to the DLQ as expansion_limit_exceeded; 0 or omission is unlimited. See Source Nodes -> split_to_rows for the complete error and fail_fast behavior.

Flattened-name collisions

The reader dissolves nested objects into dotted keys, so {"a": {"b": 1}} becomes the field a.b. When two distinct keys flatten to the same name — for example a nested {"a": {"b": 1}} alongside a literal {"a.b": 2} in the same record — only one value could survive, and keeping one while dropping the other is silent data loss. The reader refuses the record instead, naming the colliding field. This mirrors the XML reader’s treatment of a repeated element: both formats now fail loud on an undeclared collision rather than one keeping the first value and the other the last. If the collision is intentional (both values belong together), declare the column multiple: true to collect them into an array in document order; otherwise rename one of the source keys so they no longer collide. As with XML, detection is per document at read time, so the run aborts under fail_fast and dead-letters the document under continue with dlq_granularity: document.

Detection covers two distinct source keys that flatten to the same dotted name — the nested {"a": {"b": 1}} plus literal {"a.b": 2} case above. It does not cover a key that is literally duplicated within one JSON object ({"tags": "x", "tags": "y"}): the JSON parser collapses such duplicates last-wins (keeping "y") before the record reaches collision detection, so that repeat is silently dropped rather than reported. A collision inside an array element that a split_to_rows: extract fan-out lifts to the top level (one element key clashing with a parent field or with another element key) is likewise not yet detected and still resolves last-wins — tracked by issue 920.

Bounding envelope retention: max_index_bytes

When a source declares an envelope: and a pipeline reads $doc.* paths from it, the JSON reader runs a streaming pre-scan that walks the document once and retains only the declared section subtrees — every other key, including a multi-megabyte body array, is parsed-and-skipped without being stored. The retained sections live in a bounded document index.

max_index_bytes caps that index. It is charged incrementally as each section is parsed, so even a single oversized declared section aborts mid-parse (naming the section and the cap) rather than risking an out-of-memory failure. It accepts a decimal size string (64MB, 500KB) or a bare byte count; optional, defaulting to 64MB. Only the declared sections a program actually reads are retained, so envelope metadata sits far below this ceiling in practice — the cap exists to convert an unbounded mistake into a clear error. See Document Envelope Context for the full model.

Non-finite floats

JSON numbers cannot represent NaN, +infinity, or -infinity. Writing a record (or an envelope section field) that holds a non-finite float to a JSON output fails with a bounded field diagnostic, rather than silently substituting null — a substituted null would be indistinguishable from a genuine source null on read-back. Filter such records or replace the value in a transform before the JSON output.

Writing JSON

A JSON output writes one object per record, in schema-column order, either as a single array (format: array, the default) or one object per line (format: ndjson).

- type: sink
  name: enriched
  input: processed
  config:
    name: enriched
    type: json
    path: "./output/enriched.json"
    preserve_nulls: false  # omit null columns; native map/array nulls remain values
    options:
      format: ndjson     # array | ndjson
      pretty: false      # indentation for arrays or reconstructed envelopes

Dotted column names become nested objects

A column name containing a . expands back into nesting, the same way the XML writer expands one into nested elements. Columns Address.City and Address.State write as one object:

{"Address":{"City":"Boston","State":"MA"},"name":"Ada"}

This is what makes a JSON-in / JSON-out pipeline reproduce its input shape: the reader flattened {"Address":{"City":…}} into the column Address.City, and the writer puts it back. It applies to every JSON output, with no option to turn it off — a flag would mean the same column name meant different things at different outputs.

Three points follow from the rule, all shared with the XML writer and specified in full on Field Paths:

  • Grouping. Columns sharing a prefix collect into one object, positioned where that prefix first appeared, even when the schema interleaves them.
  • Absent children. Under preserve_nulls: false a null column emits no key, and an object whose every descendant is absent emits no key at all rather than an empty {} — so it reads back as the absent column it stands for.
  • Values are untouched. A column holding a map or an array still serializes as that map or array. Expansion adds structure above the value, never inside it.

Native map and array values

CXL can construct maps, arrays, and array comprehensions directly; a JSON output writes those values recursively as native objects and arrays. Map key insertion order and array item order are preserved:

emit payload = {
  customer: customer_name,
  items: [{sku: item.sku, quantity: item.quantity} for item in line_items],
}

The output contains "payload" as an object with an "items" array. Native maps and arrays are values, so the default preserve_nulls: false does not remove null map entries or array items inside them; it controls null output columns and null leaves created by dotted-column expansion.

JSON and XML share one neutral-map key grammar. After CXL has decoded the string literal, an unescaped key is its ordinary logical spelling. Exactly one leading backslash marks a literal reserved-looking key only in these three forms: \@name, \#text, or \\name. The neutral decoder removes that one marker. In CXL source, where the string literal itself must escape the backslash, write "\\@name", "\\#text", or "\\\\name". Other leading-backslash forms are non-canonical and fail. JSON assigns no structural role to @name or #text, but it still uses this same decoder: "\\@literal" writes the JSON key "@literal".

Static and computed map keys follow the same rule. Two authored spellings that decode to the same logical key are a duplicate and fail rather than selecting a winner. Before writing any bytes for a record, the writer validates the entire neutral tree. Scalars have depth zero and each map or array adds one container; depth 64 is accepted and depth 65 is rejected. A failed nested value therefore cannot leave a partial JSON record in the output.

This recursive behavior is native to JSON/NDJSON and XML. Flat, positional, and message formats do not silently turn a map or array into JSON text. Use an explicit encoding the destination format declares—such as join_values for a multi-value flat field—or reshape the value before that output. Without one, the structured value is rejected before bytes for that record are written.

Keeping a literal . in a key

To emit a key that genuinely contains a ., escape the separator in the column name. The column a\.b writes the single key "a.b":

      schema:
        - { name: "a\\.b", type: string }   # emits {"a.b": …}
        - { name: "a.b",   type: string }   # emits {"a": {"b": …}}

A [ in a column name is currently literal but reserved; write \[ if you want it to stay literal indefinitely. See Field Paths for why.

Note that this is a write-side escape. A source key that literally contains a . still arrives from the reader as an unescaped column name (the Flattened-name collisions section above covers what the reader does), so {"a.b": 1} read and written back comes out as {"a": {"b": 1}}. Closing that is tracked by issue 920.

Column names that cannot both be written

Two columns can describe places that cannot both exist in one object — a column a holding a value alongside a column a.b that needs a to be an object. Rather than keep one and drop the other, the writer refuses the whole column set before emitting that record, identifying the offending column and the path rule to correct. Field Paths lists every clashing shape.

A column name carrying a malformed escape — a \ that is not part of \., \[, or \\, as in a column literally named C:\temp — is refused the same way, with escape guidance; write C:\\temp for a literal backslash.

Preparation and empty output

Each complete output operation is prepared within the run’s finite resources before delivery. An invalid value or resource refusal during preparation writes none of that operation. A destination failure during delivery can leave a prefix; the writer then stops and never retries or finalizes on teardown. Earlier delivered records remain delivered. See output preparation for spill, cancellation and the separate file-publication boundary.

A CLI source with no body records never opens its native writer and produces an empty file, including when envelope reconstruction is selected. This differs from explicitly finalizing a library array writer, which emits [] and an LF. An explicitly opened empty envelope document retains its declared framing and has a body count of zero.

XML Format

The XML reader selects record elements by a slash-separated path of element names and maps each one onto the source’s declared schema:. Child elements bind to fields by name; attributes bind under a configurable prefix. Namespaces are stripped by default so schema field names stay clean. See Source Nodes for the shared schema and transport rules.

- type: source
  name: catalog
  config:
    name: catalog
    type: xml
    path: "./data/catalog.xml"
    schema:
      - { name: product_id, type: int }
      - { name: name, type: string }
      - { name: price, type: float }
    options:
      record_path: "catalog/product"    # slash-separated element path
      attribute_prefix: "@"             # prefix for XML attribute fields
      namespace_handling: strip         # strip | qualify
      max_index_bytes: 64MB             # cap on retained envelope sections (optional)

Text encoding

XML input must contain valid UTF-8. One leading UTF-8 BOM is removed on every physical file open, including the envelope pre-scan. UTF-16/32 BOMs are rejected. An XML declaration may omit encoding or declare UTF-8; other encodings and conflicting declarations are rejected. Convert such input to UTF-8 rather than adding an encoding option. Names, attributes, text and CDATA are never decoded with replacement characters.

Validation applies to bytes the reader consumes. A pre-scan may find a late error before any body record is delivered; a streaming body can have already delivered earlier records. Each subsequent file establishes its own BOM and declaration policy. Metadata adjacent to a selected record does not become an extra row, and repeated matching containers preserve body order and empty rows.

Output is UTF-8 without a BOM or XML declaration. See native document boundaries for envelope and empty-output behavior.

Options

OptionDefaultDescription
record_path—Slash-separated path of element names selecting the elements that each become one record — see record_path. Omitted, every top-level element becomes one record.
attribute_prefix@Prefix that distinguishes an element’s attributes from its child elements when both map to schema fields.
namespace_handlingstripstrip removes namespace prefixes from element and attribute names; qualify preserves the namespace-qualified names.
max_index_bytes64MBCap on the bytes the envelope pre-scan retains while extracting declared $doc.* sections.

record_path

record_path is a slash-separated path of XML element names, matched level by level starting at the document element. catalog/product selects every <product> that is a child of the document element <catalog>. This is the canonical statement of the grammar; other pages link here rather than restate it.

The rules, in full:

  • The path is already anchored at the document element, so it carries no leading /. Write Orders/Order, not /Orders/Order.
  • No //. It is not XPath: there is no descendant-or-any-depth step. Name every enclosing element.
  • No empty segments — no doubled separator (Orders//Order) and no trailing one (Orders/).
  • No XPath predicates, axes, or wildcards (product[@id='7'], child::product, *). Select the elements by path and filter the records in a transform.
  • Every segment must be a legal XML element name. Under namespace_handling: qualify element names keep their prefix, so a qualified segment (ns:Order) is allowed and is what matches; under the default strip the prefix is gone and the segment is the local name.
  • Omitting record_path entirely makes every top-level element one record. That is not the same as record_path: "", which is a path naming an element called “” and is rejected.

A value breaking any of these fails at compile time with E363, before any input is opened. The diagnostic names the corrected path where one can be derived.

record_path and xml_path root differently

The envelope option extract: { xml_path: … } is also a slash-path over XML, but it tolerates a leading / — /doc/Head is its documented form. record_path rejects one.

The two are separate grammars addressing separate things: xml_path locates a single envelope section anywhere in the document, record_path locates the record elements the body streams. They are deliberately not aligned — writing record_path: "/catalog/product" is an error, and writing xml_path: "/doc/Head" is correct.

Truncated input

A truncated XML document — one whose input ends before an open element’s closing tag — is rejected with a format error rather than yielding the partial fields read so far. This holds for a record cut off mid-element, a skipped-over sibling subtree cut off before it closes, and an envelope section cut off during the pre-scan (which then attaches no $doc metadata). This matches the general contract that a truncated stream always aborts rather than silently dropping data.

Writing XML

The XML writer expands dotted field names to nested elements, by the same rule the JSON writer expands them into nested objects — grouping, ordering, absent-child pruning, and the \. escape for a literal dot are all specified once on Field Paths. What is specific to XML is layered on top of that decoding, not instead of it.

The attribute_prefix convention applies in reverse: a field whose final path segment carries the prefix is emitted as an XML attribute of its enclosing element instead of a child element. A top-level @id attaches to the record element’s start tag; a nested Address.@type attaches to the <Address> element. Records read from an XML source therefore round-trip — <Record id="7"><name>A</name></Record> reads and writes back unchanged, and the writer never emits an @-named element.

Each decoded segment must also be a well-formed XML Name, so a segment that begins with a digit or contains a space is rejected. A literal dot survives — . is a legal XML name character, so a column declared a\.b emits the single element <a.b> rather than nesting.

Two column names that cannot both be expanded — a column a holding a value alongside a column a.b needing a to be a container — are refused before any byte of the record is written, naming both columns. Earlier versions emitted two sibling <a> elements for that column set, which this reader then refused on the way back in.

- type: sink
  name: xml_out
  input: processed
  config:
    name: xml_out
    type: xml
    path: "./output/result.xml"
    preserve_nulls: false              # omit null elements; null attributes always omit
    options:
      root_element: "Root"              # default Root
      record_element: "Record"          # default Record
      attribute_prefix: "@"             # matches the source-side prefix
OptionDefaultDescription
root_elementRootName of the document root element wrapping all records.
record_elementRecordName of the element emitted per record.
attribute_prefix@Prefix marking a field as an attribute of its enclosing element. Set it to the same value as the source-side prefix when round-tripping; an empty string disables attribute classification (every field emits as an element).

Attribute handling details:

  • A null attribute field is dropped even under preserve_nulls: true — a null element round-trips as a self-closing tag, but an attribute has no form that reads back as null.
  • A field with children nested under an attribute-prefixed segment (e.g. @a.b) is rejected with a format error: an XML attribute is a leaf and cannot contain elements.
  • The attribute name (the segment after the prefix) must be a well-formed XML name — a letter, _, or : followed by letters, digits, _, -, ., or : (plus the XML 1.0 Unicode name ranges). A name with a space, =, quote, /, >, or a leading digit (e.g. @foo bar, @1st) is rejected with a format error rather than emitting a malformed start tag. Non-ASCII letters are accepted, so an attribute name read from a source document round-trips unchanged.
  • An element with only attribute fields and no children self-closes: Address.@type alone emits <Address type="home"/>.

Native map and array values

An element-valued CXL map is written recursively. Ordinary keys become child elements, an unescaped key beginning with attribute_prefix becomes an attribute on the current element, and the unescaped key #text becomes text in the current element. Map insertion order controls text/child order; attributes are collected onto the start tag. Arrays held under an ordinary key repeat that key as the element name.

emit payload = {
  "@kind": "event",
  "#text": "before",
  item: [
    {"@id": 1, "#text": "alpha"},
    {"@id": 2, "#text": "beta"},
  ],
  tail: "after",
}

writes:

<payload kind="event">before<item id="1">alpha</item><item id="2">beta</item><tail>after</tail></payload>

JSON and XML share one neutral-map key grammar. After CXL string decoding, ordinary keys are unescaped. Exactly one leading backslash marks a literal key only as \@name, \#text, or \\name, and the neutral decoder removes that one marker. Because the CXL string literal must encode the backslash too, the source spellings are "\\@name", "\\#text", and "\\\\name". Other leading-backslash forms are non-canonical and fail. An escape disables XML’s attribute or text classification; it does not make the decoded spelling a legal XML name. For example, a decoded @literal still cannot be an element name, while JSON can write it as an ordinary object key.

The rules are deliberately strict:

  • Attribute and #text values must be scalar or null; maps and arrays there are rejected.
  • A direct array inside another array is rejected because XML has no child name to repeat. Put the inner array under a map key to supply that name.
  • Every decoded ordinary key and attribute name must be a well-formed XML name.
  • Static and computed keys use identical decoding. Two spellings that decode to the same logical key—including attribute-looking or #text spellings—are a collision and reject rather than selecting a winner.
  • Scalars have depth zero and each map or array adds one container. Depth 64 is accepted; depth 65, malformed escapes, duplicate logical keys, and invalid names reject the whole record before its first byte is emitted. XML never silently falls back to JSON text.

With the default preserve_nulls: false, null child elements and null array items are omitted; with it enabled they emit as self-closing elements. Null attributes are always omitted.

The XML writer is deliberately two-pass per record. Its first borrowed pass validates the complete schema/value shape, XML names, and scalar roles before writing any bytes for that record. Its second pass encodes from the borrowed original record into a private prepared operation. Delivery begins only when the complete operation is ready. Authored strings remain borrowed; other scalars are formatted in a fixed 128-byte stack scratch buffer. The writer does not clone or materialize a second nested tree, and it retains no record values or rendered scalar capacity between calls. Its only memoized preparation heap state is the admitted schema-derived element plan, whose size is independent of record value widths; recursive calls are capped at 64 containers. This is separate from the XML reader’s optional envelope pre-scan described below.

Native recursion belongs to XML and JSON/NDJSON. Flat, positional, and message formats do not stringify maps or arrays implicitly. Declare an encoding that the destination supports—such as join_values for a multi-value flat field—or reshape the value before that output; otherwise the structured value is rejected before bytes for that record are written.

Writing multi-value fields (repeated elements)

A multiple: field is written as repeated child elements, one per value, in order — the XML counterpart to the CSV writer’s delimited join_values cell, and the write-side inverse of reading multiple: true. The default needs no configuration:

<Order><id>1</id><tags>a</tags><tags>b</tags></Order>

The planner carries the exact output-facing multiple: true column set through mapping and exclusion into the writer. A top-level array repeats only for one of those columns; an array reaching any other XML column is rejected as a routing/type-contract error rather than treated as an implicit declaration. Arrays nested inside a map remain part of XML’s native recursive structure. A field with one value emits exactly one element, byte-identical to a scalar field’s output; a field with an empty array emits nothing (no element, and no container even when one is configured); an empty-string value emits a self-closing item element (<tags/>).

A multiple: column that maps to an attribute field (a column whose name maps to an XML attribute, e.g. @tags, declared multiple: true) is rejected at compile with E359 — an XML attribute holds a single value and cannot repeat, and the writer emits repetition only as child elements. A runtime array reaching an attribute field is likewise rejected by the writer.

To rename the elements, add a join_values entry — the same block the CSV writer reads, sharing the field key. The XML writer reads two keys from it and ignores the CSV-only delimiter / on_conflict / escape:

- type: sink
  name: xml_out
  input: processed
  config:
    name: xml_out
    type: xml
    path: "./output/result.xml"
    join_values:
      - field: tags
        repeat_as: Tag      # per-item element name; defaults to the field name
        wrap_in: Tags       # optional container; omit for bare repeats
  • repeat_as — the element name emitted per item. Defaults to the field’s own element name.
  • wrap_in — a container element bracketing the repeated items. Omit it for bare repeats with no container.

A scalar value on a field that carries a join_values entry is treated as a one-element sequence: it receives the same repeat_as / wrap_in naming an array of length one would, so the emitted shape does not depend on whether a lone value arrived wrapped ([a]) or bare (a) — mirroring how the reader normalizes a lone scalar into a one-element array. A field with no entry emits the plain <field>value</field> element.

The two combine into the four arrangements, with no other key:

repeat_aswrap_inOutput for tags = [a, b]
——<tags>a</tags><tags>b</tags>
Tag—<Tag>a</Tag><Tag>b</Tag>
—Tags<Tags><tags>a</tags><tags>b</tags></Tags>
TagTags<Tags><Tag>a</Tag><Tag>b</Tag></Tags>

repeat_as and wrap_in must each be a well-formed XML name, validated the same way as the root_element / record_element names. Declaring join_values on an output format that is neither csv nor xml is rejected at compile (E362).

Round trip. A document read into a multiple: true column with the default naming writes back to the identical repeated elements — reading <Order><id>1</id><tags>a</tags><tags>b</tags></Order> into a tags column and writing it to an XML output with record_element: Order reproduces the input byte-for-byte.

Repeated elements

When a record element contains repeated child elements, two source-level declarations decide what happens to them, and both take the flattened dotted field name — see Source Nodes → Multi-value fields for the shared grammar. The XML-specific matching rules are below.

A declared field is the repeated element’s dotted path relative to the record element — the same form the flattened field names use. For a record element <Order> containing repeated <Item> children, the field is Item; for <Order><Items><Item>…, it is Items.Item.

One record per occurrence: split_to_rows

- type: source
  name: orders
  config:
    name: orders
    type: xml
    path: "./data/orders.xml"
    options:
      record_path: "Orders/Order"
    schema:
      - { name: id, type: int }
      - { name: "Item.name", type: string }
      - { name: "Item.qty", type: int }
    split_to_rows:
      - field: "Item"
        mode: split            # one output record per <Item> occurrence

Each output carries one occurrence’s fields plus every field outside the group, duplicated onto each record.

Under mode: split the occurrence’s fields keep their full dotted names (Item.name, Item.@sku), including the element’s attributes. Under the default mode: extract the declared field’s prefix is lifted off, so the same document yields name and qty; a repeated scalar element (<Tag>a</Tag>) has no remainder to lift and takes the declared field’s last segment, so Tags.Tag yields Tag under extract and stays Tags.Tag under split.

Lifting a prefix off can land an occurrence’s field on a name a field outside the group already occupies — <Order><name> alongside <Item><name>. The occurrence wins: under extract it is the record, so its own field is not shadowed by the parent it was merged with. Use mode: split when you need both values, which keeps them at name and Item.name.

A declared position_column wins over any field of that name, inside the occurrence or outside it. position_column: line_no against an <Item> that carries its own <line_no> child yields the occurrence’s index, not the document’s value — you named the column, so the index is what it holds.

An occurrence with no content (<Item></Item>) still emits a record, one carrying only the fields outside the group. A record with no occurrence of the element is governed by keep_empty: XML cannot distinguish an empty repetition from an absent element, and the default keep_empty: true passes the record through unchanged.

Entries apply in declaration order, so two declared fields multiply. Fields must name disjoint element groups — a duplicated field, or one extending another (Item and Item.part) — which is rejected at compile (E358), before the source opens. The disjointness rule is this reader’s: it assigns each element to one occurrence group by document position, which is sound exactly when the declared groups do not nest. A JSON source has no such constraint.

Set source-level max_output_rows_per_input: N to bound that cumulative product without constructing it in memory. The reader emits the first N rows in document/declaration order, then a first attempted row above the ceiling routes the original ordered XML field occurrences to the DLQ as expansion_limit_exceeded; 0 or omission is unlimited. See Source Nodes -> split_to_rows for the complete error and fail_fast behavior.

All occurrences in one field: multiple: true

Declaring a schema column multiple: true collects every occurrence of that flattened field into one array, in document order, instead of keeping only the first:

    schema:
      - { name: id, type: int }
      - { name: "Tag", type: string, multiple: true }

<Tag>a</Tag><Tag>b</Tag> yields ["a", "b"], and a single <Tag> still yields a one-element array. Declaring the flattened children of a repeated container (Item.name, Item.qty) collects each of them independently.

An empty occurrence — an empty-body <Tag></Tag> or a self-closing <Tag/> — is a real array element, collected in position as a null: <Tag>a</Tag><Tag></Tag><Tag>b</Tag> yields ["a", null, "b"] rather than squeezing the empty element out, so the array round-trips its per-item shape. (An empty text value reads as null, the same rule the reader applies elsewhere; the self-closing and empty-body forms behave identically.)

A field cannot be both collected and fanned out: naming a multiple: true column in split_to_rows is rejected at compile (E358).

A repeated element named by neither a split_to_rows entry nor a multiple: column is a loud error, not a silent drop. Keeping the first occurrence and discarding the rest would lose data without warning, so the reader refuses the record and names the offending field, pointing at the two ways to handle a repeat on purpose: declare the column multiple: true to collect every occurrence into an array, or add a split_to_rows entry to fan each occurrence out to its own record. Detection is per document at read time — a plan cannot know in advance that a particular document repeats a field. Under the default fail_fast strategy the run aborts with the diagnostic; under continue with dlq_granularity: document the offending document is routed to the dead-letter queue and the run continues.

Delimited text in one element: split_values

split_values parses <Tag>a;b;c</Tag> into ["a", "b", "c"]. The field must also be declared multiple: true.

Bounding envelope retention: max_index_bytes

When a source declares an envelope: and a pipeline reads $doc.* paths from it, the XML reader runs an event-driven streaming pre-scan that walks the document once and retains only the declared section subtrees — every other element, including a multi-megabyte body, is event-walked and dropped without being flattened into memory. The retained sections live in a bounded document index.

max_index_bytes caps that index. It is charged incrementally as each section is built, so even a single oversized declared section aborts mid-parse (naming the section and the cap) rather than risking an out-of-memory failure. It accepts a decimal size string (64MB, 500KB) or a bare byte count; optional, defaulting to 64MB. Only the declared sections a program actually reads are retained, so envelope metadata sits far below this ceiling in practice — the cap exists to convert an unbounded mistake into a clear error.

The reader holds no whole-document buffer: the body walks the document element-at-a-time, and the envelope pre-scan opens the source a second time to walk it independently — a file source is read twice, never buffered. Peak memory is the bounded section index plus a single live record, not the input size. See Document Envelope Context for the full model.

Preparation and delivery failures

Invalid XML names, illegal XML characters, unsupported nested shapes and resource refusal during preparation leave that operation’s destination bytes unchanged. Diagnostics identify a bounded offending field and the rule to correct. No replacement character or JSON-string fallback is written.

A destination may accept a prefix before failing. After that failure the writer refuses further work, including finalization, and dropping it never retries the prefix. Earlier delivered records remain delivered. The CLI publishes staged files only after successful execution; this is a separate boundary from writer delivery. See output preparation.

A CLI source with no body records produces an empty file because no writer is opened. Explicitly finalizing an unused library writer instead produces <Root></Root> with default names. An explicitly opened empty envelope retains its declared framing with a body count of zero.

Fixed-Width Format

Fixed-width files carry no delimiters — each field occupies a fixed column range on every line, the layout common to mainframe extracts and legacy COBOL exports. Because the byte layout is not self-describing, each column in a fixed-width source’s schema: carries its byte layout (start + width) alongside its CXL type — one unified declaration drives both the physical slice and compile-time type checking. See Source Nodes for the shared transport rules.

- type: source
  name: legacy_data
  config:
    name: legacy_data
    type: fixed_width
    path: "./data/mainframe.dat"
    schema:
      - { name: account_id,  type: string, start: 0,  width: 12 }
      - { name: balance,     type: float,  start: 12, width: 10 }
      - { name: status_code, type: string, start: 22, width: 2 }
    options:
      line_separator: crlf    # line-ending style

The column layout

Each column pins itself to a byte range with start (a 0-based offset) and width (a byte count); end (exclusive) may be given instead of width. Optional per-column formatting keys — justify, pad, trim, truncation — control padding and trimming on read and write. Because the same column declaration carries both the byte range and the CXL type, the physical layout and the types can never drift apart. A layout shared across pipelines can live in an external .schema.yaml file referenced by schema: layout.schema.yaml.

Writing fixed-width output

A fixed-width output node declares the same column layout in its schema:. The writer places every field at its declared byte range — start plus width (or end), resolved exactly as the reader slices — regardless of the order the columns are declared in, so a file written with a schema reads back under that same schema. Byte ranges the layout leaves undeclared (a gap between fields) are filled with spaces. A column that omits start continues at the previous column’s end, so a width-only schema lays its fields out sequentially. Two columns whose byte ranges overlap have no consistent layout; the writer rejects such a schema when the output opens, naming both columns and their ranges.

Widths are byte counts, matching how the reader slices. When a value is longer than its field, truncation cuts at a UTF-8 character boundary at or below the width, so a multi-byte character is never split: the emitted cell is always valid UTF-8 of exactly width bytes (it may hold fewer characters than the width when a trailing multi-byte character does not fit, with the freed bytes pad-filled). Because padding fills exact byte counts on write and is stripped one character at a time on read, pad must be a single-byte (ASCII) character — a multi-byte character (such as ·) or a multi-character string (such as "0 ") is rejected when the schema is resolved, on both the read and write sides. An absent or empty pad defaults to a space. Under truncation: error an over-long value is still a hard error before any slicing.

truncation: warn truncates and reports it; silent performs the same truncation and reports nothing. Numeric columns default to error, and other columns default to warn.

When a run finishes, each output that truncated under warn prints one W367 warning to standard error, naming every such column with the exact number of values cut, the longest original value in bytes, the column width, and the numbers of the first eight output records that were cut (records count from 1 across everything that output wrote, including every file of a split output; an output that writes one file per source file numbers its files one after another in file-path order, except that a file whose writer fails is numbered when it fails). The warning does not change the exit code. The report never copies a value, so it cannot leak data into logs, and its memory is fixed by the schema when the output opens: recording a truncation cannot fail, so warn never turns an over-long value into a rejected record. A record that is rejected or not delivered is not counted. Truncations are also counted by the clinker.sink.truncations metric. Run clinker explain --code W367 for the fixes.

A type: decimal output column with a scale rounds its values to that many fractional places on write (banker’s rounding), the same contract a decimal source column applies on read. This matters here: a computed decimal such as avg(amount) is a full-precision quotient that would overflow a narrow numeric field — a hard error under the default truncation: error for numeric fields — so declaring the field’s scale shrinks it to fixed places that fit. For example, avg over 1.00, 1.00, 2.00 written into { name: average, type: decimal, scale: 2, width: 6 } emits 1.33; without the scale the 28-digit quotient overflows the 6-byte field and fails.

Options

OptionDefaultDescription
line_separatorlfRecord separator: lf, crlf, or none for consecutive fixed-length records.

Under lf or crlf, the reader buffers each physical line only up to the declared record width plus a line-terminator allowance. A physical line wider than the declared width — trailing filler beyond the last declared field, or a schema that maps only a prefix of a wider fixed-length record — reads its declared-width portion; the remaining bytes are discarded up to the next line terminator and the reader continues with the following record. Because the buffered portion is capped, a malformed file (a corrupt or missing newline) cannot grow a single record until end of input: memory stays bounded regardless of how long the physical line runs. A final line with no trailing newline reads normally as long as its declared fields fit within the width.

Strict selected-cell input

Field offsets remain physical byte offsets. Each selected cell must be valid UTF-8 within its own byte range; a boundary that cuts through a multi-byte character fails rather than shifting the layout or inserting a replacement character. Undeclared gaps and discarded trailing bytes are not decoded, so invalid bytes in those ignored ranges do not invalidate a selected cell. A leading UTF-8 BOM is removed before applying the layout. There is no charset conversion for fixed-width input or output.

A typed numeric cell that cannot be parsed terminates the read, including under strategy: continue; that policy does not make numeric parse failures recoverable. This differs from the multi-record unknown-discriminator handling described below.

Multi-value cells (split_values)

A fixed-width field holds one value, but its text may pack several behind a delimiter within the field’s byte range. Declare the column multiple: true and add a split_values entry, and the reader splits the (padding-stripped) field text and coerces each part to the column’s declared type:

    split_values:
      - { field: tags, delimiter: ";" }
    schema:
      - { name: order_id, type: string, start: 0, width: 4 }
      - { name: tags, type: string, start: 4, width: 20, multiple: true }

Because the fixed-width reader is the sole coercion pass, each element is typed here — a multiple: true int field over 1;2;3 reads as [1, 2, 3]. A blank field is an empty array; a field with no delimiter is a one-element array. A multiple: true column with no covering split_values entry is rejected at compile (E361). The entry is read only on a single-schema source, not the multi-record reader below.

Repeating groups

A positional repeating group occupies a bounded sequence of fixed-width occurrence records. Declare the logical column as type: map with multiple: true, put the per-occurrence byte layout in fields, and give the group a finite positive occurs.max:

    schema:
      - { name: account_id, type: string, start: 0, width: 8 }
      - name: transactions
        type: map
        multiple: true
        start: 8
        fields:
          - { name: kind, type: string, start: 0, width: 1 }
          - { name: code, type: string, start: 1, width: 2 }
        occurs:
          min: 0
          max: 3
          fill: pad
          on_overflow: error
        count_field:
          name: transaction_count
          width: 1

Child start offsets are relative to one occurrence. As with top-level columns, a child may omit start to continue after the previous child. The count field, when present, is a leading physical cell inside the group’s byte range; the occurrence payload follows it. It controls how many occurrence maps the reader returns, but it is not a logical record column and cannot be named from CXL. In the example, 2A01B02 means two three-byte occurrences followed by one unused padded slot.

occurs.min defaults to zero and cannot exceed max. Both values count logical occurrences. max has no default and cannot be inferred: it is what bounds the resolved record width, the reader’s one-line buffer, and the writer’s one-record buffer. Missing, zero, overflowing, overlapping, hybrid, or recursively repeated layouts fail during normal pipeline compilation, before an input or destination is opened.

fill chooses the physical treatment of unused occurrences:

  • pad (the default) always reserves max occurrence slots and fills unused child cells with their declared padding. Without a count field, the reader infers only a trailing run of completely padded slots. A populated slot after an empty one is invalid, and the writer rejects an authored occurrence that itself renders entirely as padding because it could not be read back unambiguously. Add count_field when an all-empty occurrence is meaningful.
  • shift omits unused slots. A count field makes the following byte position explicit. Without one, a shifted group must be the last physical field; the reader otherwise cannot distinguish group bytes from the next field.

Overflow is an error by default. The diagnostic names the group, its declared maximum, and the supplied count without printing record values. To choose lossy output deliberately, set on_overflow: truncate and also select the retained end with keep: first or keep: last; keep is invalid with the default error policy. The writer validates and encodes the complete bounded record before the first destination write, so an invalid later occurrence or overflow cannot leave a partial record behind.

Repeating groups and delimiter-packed scalar cells are separate encodings. A bare multiple: true fixed-width column is not a positional group, and a split_values entry cannot stand in for fields plus occurs. A fixed-width sink that receives an array of records must declare the same named positional group in its output schema.

Scalar document headers and footers

With reconstruct_envelope: true, options.envelope.header_from_doc and options.envelope.footer_from_doc select arbitrary document sections. Their values are concatenated in section field order, without field padding, delimiters, or the body’s byte layout. Strings are verbatim, booleans use true/false, numbers use their natural scalar spelling, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null emits no text. Arrays and maps have no scalar envelope representation and fail before any bytes from that header or footer reach the destination.

An absent section emits nothing. A present section with no fields still emits its separator: LF, CRLF, or no bytes under line_separator: none. Header, body record, and footer are separate complete operations, so a bad footer cannot undo an earlier successful body. See fixed-width document output for section selection and input-extraction limits.

Library output and failures

Direct callers construct FixedWidthEncoder with their columns, FixedWidthWriterConfig, and finite WriterResources, then wrap it in PreparedWriter. MemoryOnlyResources::new requires an explicit nonzero budget. The resource-free writer constructor is unavailable. Read the truncation account through FormatWriter::truncation_summary() (or writer.encoder().truncation_summary()): None when nothing was truncated, otherwise per-column counts, longest lengths, and the first delivered record numbers.

Preparation validates the complete operation before destination writes. Preparation failure leaves committed state unchanged and permits a corrected retry. Once delivery starts, a destination can accept a prefix before failing; the writer then refuses continuation, including flush, and drop does not retry. flush_bytes() only drains the destination; flush() finalizes once and drains. See output preparation for the separate storage and publication guarantees.

Schema drift

Fixed-width is inert with respect to auto-widen: because every byte is accounted for by the format schema, there are no “unmapped” trailing columns to absorb. The on_unmapped policy has no effect on a fixed-width source.

Multi-record files (header / trailer / body)

Mainframe and banking extracts often interleave multiple record types in one file — a header line, many body lines, and a trailer line — each identified by a discriminator at a fixed byte position (commonly the first character). Declare these with a map-form schema: carrying a discriminator: byte range and a records: list, instead of a flat column list. Each record type names its tag (the discriminator value) and its own byte-positioned columns:; the reader synthesizes the lead record_type column automatically.

- type: source
  name: payments
  config:
    name: payments
    type: fixed_width
    path: "./data/payments.dat"
    schema:                                    # one multi-record schema (map form)
      discriminator: { start: 0, width: 1 }    # the type tag occupies byte 0
      records:
        - { id: header,  tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
        - { id: detail,  tag: D, columns: [ { name: id, type: int, start: 1, width: 5 }, { name: amount, type: int, start: 6, width: 4 } ] }
        - { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 5 } ] }
      structure:
        - { record: trailer, count: count }     # validate T's count against the body count
    envelope:
      sections:
        head:
          extract: { record_type: H }          # the H line surfaces as $doc.head.*
          fields:
            batch_id: string

The reader streams one record per line on a single superset schema whose lead record_type column carries the matched type’s id. A downstream Route discriminates on that column; the file is never buffered.

  • Header lines declared as an envelope: section via the record_type extract surface as $doc.<section>.* and are excluded from the body stream (see Envelopes & Document Context).
  • Trailer lines named by a structure: constraint are validated as they stream — the declared count field is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body line after the trailer is rejected as content past the document close.
  • Blank lines (empty or whitespace-only, common after concatenation) are skipped rather than rejected; a line whose declared field range is cut off mid-value is a truncation error, not a silently-partial read. Field parsing — type coercion, padding strip, justification — is shared with the single-record fixed-width reader, so a declared type parses identically on both paths.
  • An unknown discriminator value (a tag no records: entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run under fail_fast; under continue with the default record granularity it dead-letters only that physical line and continues with the next line; under dlq_granularity: document it condemns the whole file. A record-grained DLQ row carries the line text in _cxl_dlq_source_record without guessing which declared field layout the unknown tag meant.

EDIFACT Format

Clinker reads and writes UN/EDIFACT interchanges alongside CSV, JSON, XML, and fixed-width. An interchange is a finite file: it opens with an optional UNA service-string advice and a mandatory UNB header, wraps one or more UNH..UNT messages, and closes with a UNZ trailer. The reader streams one segment at a time and the writer reconstructs the envelope around emitted records. The reader decodes release-escape sequences into clean data values and the writer re-escapes them on output, so a reader → writer → reader round-trip preserves the data values and the envelope control references.

Delimiters and the UNA service string

Each segment is terminated by the segment terminator; within a segment, data elements split on the element separator and components on the component separator. A release character escapes a delimiter that occurs as literal data.

When the file begins with a 9-byte UNA prefix, its six service characters override the defaults in this fixed order: component, element, decimal, release, repetition, terminator. When UNA is absent, the syntax Level-A defaults apply:

RoleLevel-A default
Component separator:
Element separator+
Decimal notation.
Release / escape?
Repetitionspace (inactive)
Segment terminator'

UNA is optional — a parser that requires it would fail on the common no-UNA interchange, so Clinker assumes Level-A when it is absent.

Release character

The release character (default ?) marks the following byte as literal data rather than a delimiter: ?+ is a literal + inside an element, ?' is a literal apostrophe (not a terminator), and ?? is a literal ?. The reader decodes these sequences into clean data values, so a downstream CSV/JSON sink, a CXL string comparison, or a $doc field sees O'BRIEN, never the wire form O?'BRIEN. The writer re-escapes on output: any element value that carries the element separator, the segment terminator, or the release character is release-escaped automatically, so a value computed by a Transform or sourced from CSV — never EDIFACT-escaped to begin with — does not corrupt the interchange. A reader → writer → reader round-trip therefore preserves the data values exactly.

The component separator inside an element (e.g. the : in the composite UNOA:1) is kept as part of the element’s text and is not escaped — the positional element model works above component resolution, so a composite element round-trips unchanged. A literal colon in free-text data is the one ambiguity this introduces: because components are not split into separate fields, a : in a value re-reads as a component boundary. Repeating elements ride inside one element string intact and are likewise never truncated to their first repetition.

Newlines between segments

Some producers insert CR/LF after each segment terminator for readability. Those bytes are insignificant and are stripped between segments; CR/LF that appears inside an element is preserved.

Record shape

Each non-service segment becomes one record under a fixed positional schema:

ColumnMeaning
seg_idThe segment tag (BGM, NAD, …)
msg_refThe enclosing message reference (the UNH element 1)
msg_typeThe message type (the UNH element 2, full composite)
e01, e02, …The segment’s positional data elements (release sequences decoded)

Service segments (UNB, UNZ, UNH, UNT) are consumed by the reader to drive envelope state and validation — they are never emitted as body records. The UNH segment that opens a message is emitted as a body record (its seg_id is UNH), carrying the message reference in msg_ref and the message-type composite in msg_type, with its full positional element list also stamped onto e01, e02, … — so any UNH element past the message type (a common access reference, a message subset identification, and so on) is available as e03 onward and is reconstructed on write.

The number of eNN columns is controlled by the source max_elements option (default 32). A segment carrying more data elements than that is rejected with guidance rather than silently truncated. Absent trailing elements read as null.

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: edifact
      glob: ./inbox/*.edi
      options:
        max_elements: 48      # widen the positional schema for exotic segments
      schema:
        - { name: seg_id, type: string }
        - { name: msg_ref, type: string }
        - { name: e01, type: string }

Envelope sections over UNB

The interchange header UNB is extractable as a document envelope section, exposing its positional elements to CXL as $doc.<section>.<field>. Use the segment extract rule with the section field names matching the positional keys e01, e02, …:

envelope:
  sections:
    interchange:
      extract: { segment: "UNB" }
      fields:
        e05: string          # interchange control reference (UNB element 5)

A Transform can then read $doc.interchange.e05 on every body record.

Only the UNB header is extractable as an envelope section. Trailer segments (UNT, UNZ) arrive after the body and cannot become $doc fields without buffering the whole interchange — their control counts are instead validated inline by the reader (see below). A segment extract naming any tag other than UNB, or an xml_path / json_pointer extract against an EDIFACT source, is rejected at startup.

Control-count validation

The reader validates the structural integrity claims carried in the trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):

  • UNT segment count — must equal the actual number of segments in the message, counting the UNH and UNT themselves.
  • UNT message reference — must echo the opening UNH reference.
  • UNZ message count — must equal the actual number of UNH messages in the interchange.
  • UNZ control reference — must echo the UNB control reference.

Clinker locates the UNB control reference correctly even when the header carries an empty optional element or transmits the date/time of preparation as two separate parts rather than one combined element, so such interchanges still validate and round-trip with the trailer echoing the correct reference.

A missing UNZ at end of input is a truncation error; content after the UNZ trailer is rejected — including a lone stray release character with no following segment, which is treated as unterminated (truncated) content rather than silently dropped.

Routing a count mismatch to the DLQ

By default a UNT/UNZ count mismatch aborts the run. A source declaring dlq_granularity: document instead dead-letters the whole interchange / file to the DLQ — the file’s records become a structural_validation trigger plus document_rejected collaterals, and no record of the malformed file reaches the sink. The count is only known at the trailer, after the body has streamed, so the rejection lands at the sink boundary (no record is written out), not literally before the first record. The grain is the whole file. The control-reference echo mismatches and every other corruption (truncation, post-trailer content) always abort, even under the opt-in. See Malformed envelopes.

Writing EDIFACT

An EDIFACT Sink node reconstructs the envelope around emitted records. Records map by the same positional columns (seg_id, msg_ref, msg_type, eNN); trailing null/empty elements are trimmed so no fabricated delimiters appear, and a column the writer does not recognize is an error (project the record to the EDIFACT columns first). Engine-internal $-namespaced columns are excluded automatically.

nodes:
  - type: sink
    name: out
    input: messages
    config:
      name: out
      type: edifact
      path: ./out/result.edi
      options:
        interchange: ["UNOA:1", "SENDER", "RECEIVER", "240101:1200", "REF1"]
        message_type: "ORDERS:D:96A:UN"
        write_una: false
        segment_newline: true

Output options:

OptionMeaning
interchangeLiteral UNB data elements (release-escaped as needed on write).
interchange_from_docName of a $doc section to echo the UNB elements from (round-trip).
message_typeFallback UNH message type when a record carries no msg_type value.
write_unaEmit a leading UNA segment (default false).
segment_newlineWrite a newline after each segment terminator (default true).

Consecutive records are grouped into UNH..UNT messages on msg_ref transitions. The writer recomputes the UNT segment count and UNZ message count, and echoes the message and interchange control references, so the output passes its own count validation on re-read.

interchange_from_doc echoes the header from a record’s document context. That context is populated by a source’s UNB envelope section (declare a segment: "UNB" envelope section on the source) and travels with every body record through the pipeline — including to a sink that sits directly downstream of the source with no intervening Transform. The reader stashes the complete, ordered UNB element list (empty middle elements included), so the reconstructed header is faithful even when a middle element is empty and the user declares only the fields they care about. Supply interchange literal elements instead when the records have no source UNB section to echo.

Character set

EDIFACT names its body character repertoire in-band: the UNB header’s syntax identifier (data element S001, component 1 — the UNOA in UNOA:1) declares the repertoire for the whole interchange. There is therefore no encoding option on an EDIFACT source or sink; the reader discovers the repertoire from the UNB and the writer re-derives it from the UNB it emits, so a read → write round-trip is byte-faithful without any configuration.

UNB syntax levelRepertoire
UNOA, UNOBASCII; a byte >= 0x80 is an error
UNOCISO-8859-1 (Latin-1), one byte per character
UNOYUTF-8; invalid byte sequences are an error

Decoding stays streaming and per-segment — the interchange is never buffered whole. The syntax identifier is itself ASCII, so it is read straight from the raw UNB bytes before any text is decoded; the UNB and every body segment are then decoded through the negotiated repertoire. The UNB’s own sender and recipient identification elements may legitimately carry non-ASCII text under UNOC or UNOY, and they decode (and re-encode on output) under that repertoire too — so a UNOC interchange whose header or body carries Latin-1 high bytes (for example accented characters in a party name) parses without error, surfaces the correct text in $doc.UNB.*, and round-trips byte-for-byte.

The repertoire is enforced loudly. A UNOA/UNOB interchange whose body carries a high byte fails (“outside the ASCII repertoire”) rather than silently reinterpreting it; a UNOY interchange with invalid UTF-8 fails (“not valid UTF-8”); and a UNB declaring an unsupported syntax level (UNOD..UNOX) fails at startup with a precise error naming the level, never falling back to a guessed encoding or substituting replacement characters. On output the writer encodes element text through the same repertoire the UNB declares; a character the repertoire cannot represent (a non-ASCII character under UNOA/UNOB, or a codepoint above U+00FF under UNOC) is rejected rather than emitted truncated.

Limitations

  • Functional groups. A single UNB..UNZ interchange is supported; UNG/UNE functional-group segments are rejected with a precise error.
  • Output splitting. An interchange is a single UNB..UNZ envelope and cannot be divided across files. An edifact output combined with a split: block is rejected at config-validation time (diagnostic E323) rather than emitting a structurally corrupt interchange.
  • Rare degenerate headers. A few unusual header shapes — those that combine an empty date/time slot with a date-only date/time where the control reference normally sits — may not round-trip byte-for-byte on re-emit. Conformant headers and ordinary variations are unaffected.

X12 Format

Clinker reads and writes ANSI ASC X12 interchanges alongside CSV, JSON, XML, fixed-width, and EDIFACT. An X12 interchange is a finite file with a three-tier envelope: an ISA..IEA interchange wraps one or more GS..GE functional groups, and each functional group wraps one or more ST..SE transaction sets. The reader streams one segment at a time and the writer reconstructs the three envelope tiers around emitted records.

The three tiers surface as nested document-context levels: the ISA interchange becomes the file-level $doc document, and each GS group and ST set opens a nested level whose $doc sections layer over the enclosing tiers. A body record therefore sees every enclosing tier’s fields through one $doc.<section>.<field> lookup.

Delimiters and the ISA header

Unlike EDIFACT’s optional UNA service-string advice, X12 declares its delimiters in a fixed-length 106-byte ISA header. Three delimiter bytes live at structural positions within it:

RoleSource in the ISA
Element (data) separatorThe byte immediately after the ISA tag
Sub-element (component) sep.ISA16, the last single-byte ISA element
Segment terminatorThe byte immediately after ISA16

The reader reads these three bytes from the header rather than assuming a fixed delimiter set, so an interchange that uses */:/~, |/^/newline, or any other producer-chosen delimiters parses correctly. The ISA13 interchange control number is located as the 13th element of the header split on the discovered element separator — structurally, not by an absolute byte offset — so producer padding quirks do not misalign it.

On output, an X12 sink that echoes the header via interchange_from_doc also adopts the source’s discovered delimiter set, so a reconstructed interchange keeps the exact element separator, sub-element separator, and segment terminator bytes it arrived with (see Writing X12). The literal interchange option keeps the writer’s */:/~ defaults.

No escape character

X12 has no release/escape character (EDIFACT’s ? has no X12 equivalent). A data value that contains a delimiter byte is therefore unrepresentable. On output the writer rejects any element value carrying the element separator or the segment terminator with a precise error rather than silently corrupting the interchange; re-encode the value or choose delimiters the data does not contain.

The sub-element (component) separator inside an element (e.g. the : in a composite A:B:C) is kept as part of the element’s text and is not split — the positional element model works above component resolution, so a composite element round-trips unchanged.

Newlines between segments

Some producers insert CR/LF after each segment terminator for readability. Those bytes are insignificant and are stripped between segments; CR/LF that appears inside an element is preserved.

Record shape

Each non-service segment becomes one record under a fixed positional schema:

ColumnMeaning
seg_idThe segment tag (BEG, PO1, …)
group_refThe enclosing functional-group control number (GS06)
set_refThe enclosing transaction set control number (ST02)
set_typeThe transaction set identifier code (ST01, e.g. 850)
e01, e02, …The segment’s positional data elements

The reader stamps both envelope control numbers on every body record: group_ref from GS06 and set_ref from ST02. The same GS06 value also surfaces through $doc.functional_group.e06 for expressions that need the full functional-group envelope (see Envelope sections over the three tiers).

Service segments (ISA, IEA, GS, GE, SE) are consumed by the reader to drive the envelope and validation — they are never emitted as body records. The ST segment that opens a transaction set is emitted as a body record (its seg_id is ST), carrying the set reference and type.

The number of eNN columns is controlled by the source max_elements option (default 32). A segment carrying more data elements than that is rejected with guidance rather than silently truncated. Absent trailing elements read as null.

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: x12
      glob: ./inbox/*.x12
      options:
        max_elements: 48      # widen the positional schema for exotic segments
        encoding: iso-8859-1  # decode body element text as Latin-1
      schema:
        - { name: seg_id, type: string }
        - { name: group_ref, type: string }
        - { name: set_ref, type: string }
        - { name: e01, type: string }

Character set

X12 carries no in-band element that names the body character repertoire (unlike EDIFACT’s UNB syntax identifier). Element text is therefore decoded through the charset the source declares in its encoding option, defaulting to UTF-8. The ISA header’s control fields are ASCII, so the declared charset affects only the body element text.

encoding valueRepertoire
utf-8 (default)UTF-8; invalid bytes are an error
iso-8859-1 (aliases latin-1, latin1)ISO-8859-1 (Latin-1), one byte per char

Decoding stays streaming and per-element — the interchange is never buffered whole. An interchange whose body carries Latin-1 high bytes (for example accented characters in free-text name or address fields) parses without error once the source declares encoding: iso-8859-1, and the decoded element text matches the source bytes under Latin-1.

A source that omits the option and meets non-UTF-8 bytes fails explicitly (“segment is not valid UTF-8”) rather than corrupting the data silently; a source that declares an unsupported encoding fails at startup with a precise error naming the value. On output, set the same encoding on the X12 sink so the round-trip is byte-faithful; a character the chosen charset cannot represent (for example a non-Latin-1 codepoint under iso-8859-1) is rejected rather than emitted truncated.

Envelope sections over the three tiers

The interchange header ISA is extractable as a file-level document envelope section, exposing its positional elements to CXL as $doc.<section>.<field>. Use the segment extract rule with the section field names matching the positional keys e01, e02, …:

envelope:
  sections:
    interchange:
      extract: { segment: "ISA" }
      fields:
        e13: string          # interchange control number (ISA13)

The GS functional group and the ST transaction set surface automatically as the nested $doc sections functional_group and transaction_set, each keyed by positional eNN elements — no envelope declaration is needed for them. A Transform on any body record can read all three tiers at once:

emit isa13 = $doc.interchange.e13       # interchange control number
emit gs06  = $doc.functional_group.e06  # group control number (GS06)
emit st02  = $doc.transaction_set.e02   # set control number (ST02)

Naming and typing the nested levels

The ISA header is declared through envelope: because a bounded pre-scan resolves it before any body streams. The GS group and ST set exist only mid-file, so they cannot be declared the same way — instead, name them and give them a typed field schema under the X12 source’s options. The reader applies the declaration each time it crosses a group or set boundary, so $doc.<your-name>.<field> exposes the level’s elements under the name you chose, coerced to the types you declared — the same way a declared ISA field is typed and coerced:

type: x12
options:
  group_section:
    name: functional_group   # your choice — the engine reserves no name
    fields:
      e01: string            # GS01 functional identifier code
      e06: int               # GS06 group control number
  set_section:
    name: transaction_set    # your choice
    fields:
      e01: int               # ST01 transaction-set identifier code
      e02: string            # ST02 set control number
emit functional_id = $doc.functional_group.e01   # typed string
emit group_control = $doc.functional_group.e06    # typed int
emit txn_type      = $doc.transaction_set.e01     # typed int

The two levels are declared independently — name one, both, or neither. A declared field schema is the contract: only the elements it lists surface in the typed section, and an element the wire carries but the schema omits is absent from $doc. An element that cannot coerce to its declared type (declaring the alphabetic GS01 code as an int, say) fails the run with a precise error. Omit a level’s declaration and it keeps its default name (functional_group / transaction_set) keyed by untyped positional eNN strings, unchanged from before this option existed.

Only the ISA header is extractable through the envelope: block. Trailer segments (SE, GE, IEA) arrive after the body they close and cannot become $doc fields without buffering the whole interchange — their control counts are instead validated inline by the reader (see below). A segment extract naming any tag other than ISA, or an xml_path / json_pointer extract against an X12 source, is rejected at startup.

Control-count validation

The reader validates the structural integrity claims carried in the trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):

  • SE segment count (SE01) — must equal the number of segments in the transaction set, counting the ST and SE themselves.
  • SE set control number (SE02) — must echo the opening ST02.
  • GE transaction-set count (GE01) — must equal the number of ST sets in the functional group.
  • GE group control number (GE02) — must echo the GS06.
  • IEA functional-group count (IEA01) — must equal the number of GS groups in the interchange.
  • IEA control number (IEA02) — must echo the ISA13.

A missing IEA at end of input is a truncation error; content after the IEA trailer is rejected.

Routing a count mismatch to the DLQ

By default a SE/GE/IEA count mismatch aborts the run. A source declaring dlq_granularity: document instead dead-letters the whole interchange / file to the DLQ — the file’s records become a structural_validation trigger plus document_rejected collaterals, and no record of the malformed file reaches the sink. The count is only known at the trailer, after the body has streamed, so the rejection lands at the sink boundary (no record is written out), not literally before the first record. The grain is the whole file: an SE-level mismatch rejects the entire interchange, not just that transaction set. The control-number echo mismatches (SE02/GE02/IEA02) and every other corruption (truncation, post-trailer content) always abort, even under the opt-in. See Malformed envelopes.

Writing X12

An X12 Sink node reconstructs the three-tier envelope around emitted records. Records map by the same positional columns (seg_id, group_ref, set_ref, set_type, and eNN); trailing null/empty elements are trimmed so no fabricated delimiters appear, and a column the writer does not recognize is an error (project the record to the X12 columns first). Engine-internal $-namespaced columns are excluded automatically.

nodes:
  - type: sink
    name: out
    input: messages
    config:
      name: out
      type: x12
      path: ./out/result.x12
      options:
        interchange:
          ["00", "          ", "00", "          ", "ZZ", "SENDER         ",
           "ZZ", "RECEIVER       ", "240101", "1200", "U", "00401",
           "000000001", "0", "P", ":"]
        group_header: ["PO", "SENDER", "RECEIVER", "20240101", "1200", "1", "X", "004010"]
        set_type: "850"
        segment_newline: true

Output options:

OptionMeaning
interchangeLiteral ISA data elements (the 16 fixed-width ISA fields).
interchange_from_docName of a $doc section to echo the ISA elements from (round-trip).
group_headerLiteral GS01..GS08 elements (GS06 control number recomputed per group).
set_typeFallback ST01 set type when a record carries no set_type value.
segment_newlineWrite a newline after each segment terminator (default true).
encodingCharacter set element text is encoded through (default utf-8).

Consecutive records are grouped into ST..SE transaction sets on set_ref transitions and into GS..GE functional groups on group_ref transitions — the group discriminator is the outer-tier analog of set_ref. The writer recomputes the SE segment count, the per-group GE transaction-set count, and the IEA functional-group count, and echoes the set, group, and interchange control numbers, so the output passes its own count validation on re-read.

Multiple functional groups

A real interchange can carry several functional groups inside one ISA..IEA. The reader-stamped group_ref makes a direct X12-to-X12 pipeline preserve those boundaries: the writer opens a fresh GS..GE group every time the value changes, echoes it as the group control number (GS06/GE02), and recomputes that group’s GE01 transaction-set count. The configured group_header continues to supply the other GS fields for each reconstructed group. Records from a source without group_ref still collapse into one functional group.

nodes:
  - type: sink
    name: out
    input: orders
    config:
      name: out
      type: x12
      path: ./out/result.x12
      options:
        interchange_from_doc: interchange
        group_header: ["PO", "SENDER", "RECEIVER", "20240101", "1200", "1", "X", "004010"]

Records that share a group_ref value must arrive consecutively, exactly as records sharing a set_ref must: the writer streams and closes a group the moment the discriminator changes, so an interleaved stream would reopen a group it already closed. Sort upstream by group_ref (then set_ref) when the record order does not already guarantee it.

interchange_from_doc echoes the header from a record’s document context. That context is populated by a source’s ISA envelope section (declare a segment: "ISA" envelope section on the source) and travels with every body record through the pipeline — including to a sink that sits directly downstream of the source with no intervening Transform. The reader stashes the complete, ordered ISA element list together with the delimiter set it discovered from the header, so the reconstructed header is faithful and the whole output interchange — header, envelopes, and body — is emitted with the original element separator, sub-element separator, and segment terminator rather than the writer’s */:/~ defaults. Supply interchange literal elements instead when the records have no source ISA section to echo; that path keeps the default delimiters.

Limitations

  • Charset. Element text is decoded through the source’s encoding option (UTF-8 by default, or ISO-8859-1). An unsupported or unconfigured-but-non-UTF-8 repertoire is rejected explicitly rather than silently corrupted (see Character set above).
  • No escape character. X12 has no release mechanism, so a data value that contains a delimiter byte is rejected on output rather than corrupting the interchange.
  • Consecutive grouping. Both GS..GE functional groups (group_ref) and ST..SE transaction sets (set_ref) close the moment their discriminator changes, so records sharing a value must arrive together; sort upstream when the source order does not already guarantee it.
  • Output splitting. An interchange is a single ISA..IEA envelope and cannot be divided across files. An x12 output combined with a split: block is rejected at config-validation time (diagnostic E338) rather than emitting a structurally corrupt interchange.

HL7 v2 Format

Clinker reads and writes HL7 v2.x pipe-and-hat messages alongside CSV, JSON, XML, fixed-width, EDIFACT, and X12. An HL7 v2 file is a finite stream of carriage-return-terminated segments. The smallest unit is one message, which always begins with an MSH (message header) segment; messages may optionally be wrapped in a BHS..BTS batch and an FHS..FTS file envelope. The reader streams one segment at a time and the writer re-emits the segments, optionally reconstructing the batch/file envelopes.

The optional envelope tiers surface as nested document-context levels: an FHS file header becomes the file-level $doc document, and each BHS batch and MSH message opens a nested level whose $doc sections layer over the enclosing tiers. A body record therefore sees every enclosing tier’s fields through one $doc.<section>.<field> lookup. All tiers are optional — a bare stream of MSH messages with no batch or file wrapping is a valid HL7 v2 file.

Delimiters and the MSH header

HL7 declares its delimiters in the MSH header rather than assuming a fixed set. The byte immediately after the MSH tag is the field separator (MSH-1), and the four bytes that follow are the encoding characters (MSH-2), in this fixed order:

RoleSource in MSH-2Conventional byte
Component separatorfirst encoding character^
Repetition separatorsecond encoding char~
Escape characterthird encoding char\
Sub-component sep.fourth encoding char&

The reader reads these bytes from the header, so a message that uses the conventional |^~\& or any other producer-chosen delimiters parses correctly. The segment terminator is always a carriage return (0x0D) — unlike the field and encoding delimiters, it is never producer-chosen.

When a file opens with an FHS or BHS header instead of MSH, the delimiters are read from that header; HL7 requires the file’s FHS/BHS encoding characters to match its messages’ MSH.

The discovered delimiter set travels with each message through the pipeline, and an HL7 Output re-emits every message with the set its header declared — a custom-delimiter file round-trips byte-faithfully, never silently rewritten to the conventional |^~\& (see Writing HL7).

The MSH off-by-one

MSH-1 is the field separator, so it is implicit — it never appears as a data field. When the MSH segment is split on the field separator, the encoding-characters field (MSH-2) is the first data field. As a result, for any MSH field number N ≥ 2, the positional column is f<N-1>: MSH-2 is f01, the sending application MSH-3 is f02, the message type MSH-9 is f08, and the message control id MSH-10 is f09. The same off-by-one applies to FHS and BHS.

Escape sequences

Field data escapes a literal delimiter character with an escape sequence \X\, where X names the delimiter: \F\ field separator, \S\ component separator, \T\ sub-component separator, \R\ repetition separator, \E\ the escape character itself. The reader decodes these into their literal data byte, so downstream consumers — CSV/JSON output, CXL string predicates, $doc fields — see clean data, never the wire escapes. The writer re-escapes any delimiter byte in field data on output, so the reader → writer → reader round-trip is byte-faithful.

An application escape the positional reader does not decode (e.g. the formatting escape \.br\) is kept verbatim rather than dropped, so no data is lost.

The component, repetition, and sub-component separators inside a field (e.g. the ^ in a composite PATID^^^HOSP^MR) are kept as part of the field’s text and are not split — the positional field model works above component resolution, so a composite field round-trips unchanged.

Newlines between segments

Some producers (and CRLF-normalizing transports) add a line feed after each carriage-return terminator. Those bytes are insignificant and are stripped between segments. A producer that omits the trailing carriage return on the final segment is accepted — that shape is common in practice.

Record shape

Each segment becomes one record under a fixed positional schema:

ColumnMeaning
seg_idThe segment tag (MSH, PID, OBX, …)
set_refThe enclosing message’s control id (MSH-10)
set_typeThe enclosing message’s type (MSH-9, e.g. ADT^A01)
f01, f02, …The segment’s positional data fields

The MSH header segment is emitted as a body record (its seg_id is MSH), carrying the message’s fields positionally. Batch/file envelope segments (FHS, FTS, BHS, BTS) are consumed by the reader to drive the document levels and validate counts — they are never emitted as body records.

The number of fNN columns is controlled by the source max_fields option (default 64). A segment carrying more data fields than that is rejected with guidance rather than silently truncated. Absent trailing fields read as null.

nodes:
  - type: source
    name: messages
    config:
      name: messages
      type: hl7
      glob: ./inbox/*.hl7
      options:
        max_fields: 128       # widen the positional schema for large OBX segments
      schema:
        - { name: seg_id, type: string }
        - { name: set_ref, type: string }
        - { name: f01, type: string }

Component splitting (optional)

By default a composite field rides inside one fNN column with its component (^), repetition (~), and sub-component (&) separators intact — the positional model deliberately works above component resolution. When you want component-level access (the message code MSH-9.1 vs the trigger event MSH-9.2) without writing CXL string-splitting downstream, opt one or more fields into splitting with split_fields. The reader explodes the named field into structured columns, and an HL7 Output re-assembles the exact wire field from them, so an HL7→HL7 round-trip stays byte-identical.

options:
  split_fields:
    - { field: f08, components: 2 }                       # MSH-9 → message code + trigger
    - { field: f03, components: 5 }                       # PID-3 (CX) → its components
    - { field: f04, components: 2, subcomponents: 3 }     # also expose sub-components
    - { field: f13, components: 1, repetitions: 4 }       # repeating field → per-repetition

Each split fixes the column width on three structural axes: components (required, the ^ axis), subcomponents (default 1, the & axis), and repetitions (default 1, the ~ axis). The schema stays static — it never varies with per-record data. A field whose data carries more structure on any axis than the declaration reserves is rejected with guidance, the same posture as a max_fields overflow; raise the axis count or leave the field unsplit.

The exploded columns name the path from the field down to a leaf with the axis letters r, c, s, all 1-based, eliding the default index (1) on the repetition and sub-component axes so the common component-only case stays clean:

DeclarationColumns for f08
components: 2f08_c1, f08_c2
components: 1, subcomponents: 2f08_c1_s1, f08_c1_s2
components: 1, repetitions: 2f08_r1_c1, f08_r2_c1

The verbatim fNN column is replaced by the structured columns. Declare the exploded column names in the source schema: block (or rely on on_unmapped) the same way you would any other column:

options:
  split_fields:
    - { field: f08, components: 2 }
schema:
  - { name: seg_id, type: string }
  - { name: f08_c1, type: string }   # MSH-9.1 message code
  - { name: f08_c2, type: string }   # MSH-9.2 trigger event
emit code    = f08_c1
emit trigger = f08_c2

Splitting respects the escape rules: an escaped separator (e.g. \S\, a literal ^ in data) is not treated as a component boundary — the split runs on the raw bytes before the escape decodes, so the literal stays inside one component. On output the writer re-joins the leaves on the separators verbatim (never escaping ^/~/&) and still escapes any field-separator, escape, or carriage-return byte inside a leaf, so the round-trip is byte-faithful.

Envelope sections over the tiers

The file header FHS is extractable as a file-level document envelope section, exposing its positional fields to CXL as $doc.<section>.<field>. Use the segment extract rule with the field names matching the positional keys f01, f02, … :

envelope:
  sections:
    file:
      extract: { segment: "FHS" }
      fields:
        f07: string          # file name / id (FHS-8 under the off-by-one)

The BHS batch and the MSH message surface automatically as the nested $doc sections batch and transaction_set, each keyed by positional fNN fields — no envelope declaration is needed for them. A Transform on any body record can read all available tiers at once:

emit file_id  = $doc.file.f07            # FHS file id (declared section)
emit batch_id = $doc.batch.f07           # BHS batch id (auto section)
emit mtype    = $doc.transaction_set.f08 # MSH-9 message type (auto section)
emit ctrl     = $doc.transaction_set.f09 # MSH-10 control id (auto section)

Only the FHS header is extractable as a declared envelope section, and only when the file actually opens with one; a bare MSH-led file has no file-level envelope, so declaring an FHS section against it is rejected at startup. Trailer segments (BTS, FTS) arrive after the body they close and cannot become $doc fields without buffering the whole file — their counts are instead validated inline by the reader (see below). A segment extract naming any tag other than FHS, or an xml_path / json_pointer extract against an HL7 source, is rejected at startup.

Control-count validation

The reader validates the structural integrity claims carried in the batch/file trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):

  • BTS batch message count (BTS-1) — must equal the number of MSH messages in the batch. An empty BTS-1 disables the check (the count is optional in practice).
  • FTS file batch count (FTS-1) — must equal the number of BHS batches in the file. An empty FTS-1 disables the check.

Content after the FTS file trailer is rejected. A bare MSH-led file needs no trailers and validates with no count checks.

Routing a count mismatch to the DLQ

By default a BTS/FTS count mismatch aborts the run. A source declaring dlq_granularity: document instead dead-letters the whole file to the DLQ — the file’s records become a structural_validation trigger plus document_rejected collaterals, and no record of the malformed file reaches the sink. The count is only known at the trailer, after the body has streamed, so the rejection lands at the sink boundary (no record is written out), not literally before the first record. The grain is the whole file. Other corruption (truncation, post-trailer content) always aborts, even under the opt-in. An empty BTS-1/FTS-1 still disables the check entirely. See Malformed envelopes.

Writing HL7

An HL7 Sink node re-emits the MSH and body segments from the record stream, escaping any field data that carries a delimiter byte. Records map by the same positional columns (seg_id, fNN); trailing null/empty fields are trimmed so no fabricated delimiters appear, and a column the writer does not recognize is an error (project the record to the HL7 columns first). Split-leaf columns (f08_c1, f03_r2_c1_s3) are recognized too — the writer groups them by field and re-assembles the wire value from the column names alone, so no output option is needed to round-trip a split source. The reader-stamped set_ref/set_type echoes and engine-internal $-namespaced columns are excluded automatically.

The writer re-emits each message with the delimiter set its source header declared: the reader stamps the discovered field separator and encoding characters into the message’s document context, and the writer adopts that set at every MSH — field joining, escaping, and split re-assembly all use it, so a custom-delimiter message round-trips byte-faithfully. A record that did not come from an HL7 source (or whose document context was dropped upstream) is written with the conventional |^~\& set.

nodes:
  - type: sink
    name: out
    input: messages
    config:
      name: out
      type: hl7
      path: ./out/result.hl7
      options:
        file_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "FILE7"]
        batch_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "BATCH3"]
        segment_newline: true

Output options:

OptionMeaning
file_headerLiteral FHS fields; opens an FHS..FTS file envelope.
file_header_from_docName of a $doc section to echo the FHS fields from (round-trip).
batch_headerLiteral BHS fields; wraps the messages in a BHS..BTS batch.
segment_newlineWrite a newline after each segment terminator (default true).

When a file or batch header is configured, the writer recomputes the BTS batch message count and the FTS file batch count and emits the trailers at end of stream, so the output passes its own count validation on re-read. The MSH-2 encoding-characters field of each header is written verbatim so the delimiter declaration round-trips. With no header options set, the writer emits a bare stream of messages with no batch/file envelope.

file_header_from_doc echoes the FHS header from a record’s document context. That context is populated by a source’s FHS envelope section (declare a segment: "FHS" envelope section on the source) and travels with every body record through the pipeline. The reader stashes the complete, ordered FHS field list, so the reconstructed header is faithful. Supply file_header literal fields instead when the records have no source FHS section to echo.

Limitations

  • Charset. Field text is decoded as UTF-8. Non-UTF-8 messages are rejected explicitly rather than silently corrupted.
  • Positional fields by default; opt-in component columns. Components, repetitions, and sub-components ride inside one positional fNN field verbatim unless that field is named in split_fields (see Component splitting above), in which case the reader explodes it into structured columns and the writer re-assembles them. The reader always decodes the \X\ delimiter escapes regardless.
  • HL7 v3 and FHIR are out of scope. HL7 v3 is XML (use the xml format) and FHIR is JSON/REST (use the json format); this format handles HL7 v2.x pipe-and-hat encoding only.
  • Output splitting. A batch/file envelope is a single FHS..FTS structure and cannot be divided across files. An hl7 output combined with a split: block is rejected at config-validation time (diagnostic E339) rather than emitting a structurally corrupt file.

SWIFT MT Format

Clinker reads and writes SWIFT MT (FIN) messages alongside CSV, JSON, XML, fixed-width, EDIFACT, X12, and HL7 v2. A SWIFT MT message is a finite file built from brace-balanced blocks. The reader scans the message into retained fields, then emits one body field at a time and surfaces service blocks as document-envelope sections. The writer preserves field-value bytes while reconstructing the message with CRLF structural separators; it does not preserve the original file’s inter-block whitespace or field separators.

Block structure

A SWIFT MT message is a sequence of top-level blocks, each {n:...} where n is a numeric block id:

BlockRoleContents
1Basic headerA fixed header string (application id, BIC, …)
2Application headerInput/output direction, message type, recipient
3User headerOptional; may carry nested {tag:value} sub-blocks
4Message textThe message body — a run of :tag:value fields
5TrailerOptional; may carry nested {tag:value} sub-blocks

Unlike the flat delimiter-structured EDI formats (HL7, X12, EDIFACT), SWIFT framing is brace-balanced rather than terminator-delimited. Because of that, the nested sub-blocks of blocks 3 and 5 (for example {3:{108:MSGREF}}) are kept intact inside their parent block rather than mistaken for top-level blocks.

The -} text-block trailer

Block 4 is special. Its body is opaque line-structured free text — a field value (a :77E: envelope, a :79: narrative, an :86: information line) legitimately contains {, }, and even -} as data. So braces inside block 4 are treated as data, not framing: the block closes only on a line-anchored -} trailer — at the start of the block-4 body or after either LF or CR. An interior {, }, or a -} in the middle of a value is data, not a frame boundary. Both the framing braces and the closing -} trailer are stripped from the stored values, so a record carries clean tag/value data.

Whitespace between blocks

Producers insert CR/LF (and the \r\n that separates block-4 fields) for readability. Inter-block whitespace is insignificant and is skipped; the line breaks inside block 4 delimit the :tag:value fields.

Record shape

Each :tag:value line of block 4 becomes one record under a fixed positional schema — the same one-line-one-record model the X12 and HL7 readers use:

ColumnMeaning
blockThe block id the line came from (always 4 for body fields)
tagThe SWIFT field tag without its surrounding colons (20, 32A, 61)
valueThe field value, with continuation lines folded in verbatim

A multi-line field (a :50K: ordering-customer block, a :77E: / :86: narrative) keeps its continuation lines: any line of block 4 that does not begin a new :tag: is folded into the current field’s value with its line break preserved — including a blank line inside the value, so a narrative with an internal blank line round-trips faithfully. LF and CRLF are retained exactly inside values, including mixed separators, leading/trailing spaces, and trailing blank continuation lines. Only the separator before the next field or trailer is structural: one CRLF, LF, or (at the trailer) lone CR is removed. A data CR immediately followed by the structural LF is necessarily read as one CRLF separator; these bytes cannot express a separate data CR. A repeated tag (the :61: / :86: statement lines of an MT940, for instance) streams as one record per occurrence, in order.

The service blocks (1, 2, 3, 5) are consumed by the reader to serve envelope sections and drive the message-level document context — they are never emitted as body records.

nodes:
  - type: source
    name: payments
    config:
      name: payments
      type: swift
      glob: ./inbox/*.swift
      options:
        max_fields: 20000   # raise the block-4 field ceiling for large messages
      schema:
        - { name: block, type: string }
        - { name: tag, type: string }
        - { name: value, type: string }

The max_fields option caps the number of block-4 field lines a single message may carry (default 10000). A message exceeding it is rejected with guidance. This is a field-count guard, not a byte budget for retained reader data; it does not establish constant-memory input parsing.

Envelope sections over the service blocks

A SWIFT MT message is a single envelope: the four service blocks surface as file-level $doc sections, exposing each block’s text to CXL as $doc.<section>.body. Every body record can read the enclosing message’s headers through a $doc.<section>.body lookup.

Declare the sections on the source with the segment extract rule naming the block id. The whole block body surfaces under the field name body, because a SWIFT service block carries free-form text (a header string, nested {sub:tag} blocks) rather than positional elements:

envelope:
  sections:
    basic:
      extract: { segment: "1" }   # block 1, the basic header
    app:
      extract: { segment: "2" }   # block 2, the application header
    user:
      extract: { segment: "3" }   # block 3, the user header (nested sub-blocks kept verbatim)
    trailer:
      extract: { segment: "5" }   # block 5, the trailer

A Transform on any body record can read every enclosing block at once:

emit tag       = tag
emit value     = value
emit basic_hdr = $doc.basic.body     # block 1 header string
emit app_hdr   = $doc.app.body       # block 2 header string
emit user_hdr  = $doc.user.body      # block 3 body, nested {108:...} kept verbatim

The section names are entirely your choice — the engine reserves none. A segment extract may name a block either by its numeric id ("1", "3") or by the stable default label ("basic_header", "app_header", "user_header", "trailer"); both resolve the same block.

Block 4 is the message-text body streamed as records, not an envelope section — a segment: "4" extract is rejected at startup. An xml_path or json_pointer extract against a SWIFT source is likewise rejected, because those rules belong to the tree formats.

Malformed-message handling

A structurally broken message fails the run with a precise SWIFT error rather than producing garbled records:

  • Unbalanced brace — a block that never closes (or whose brace depth never returns to zero) is a truncation error naming the offending block.
  • Missing -} trailer — a block 4 that runs to end of input without its -} trailer is a truncation error.
  • Missing or non-numeric block id — a block without a numeric id after the { (or with no : separating the id from the body) is rejected.
  • Malformed :tag:value line — a block-4 line with no second colon closing the tag, or an empty tag, is rejected.
  • Repeated service block — a second {1:...} (or any repeated service block) in one message is rejected.
  • Invalid UTF-8 — block ids and bodies are validated before parsing; invalid bytes in headers, body text, or trailers are never replaced with substitute characters. A leading UTF-8 BOM is rejected: after optional ASCII whitespace, the next byte must be an opening {.

Initialization succeeds only after the complete message has parsed. On failure, partial fields and service blocks are discarded; later reads are terminal and cannot reveal partial records or sections. Malformed messages terminate execution under both fail_fast and continue.

A header-only message (no block 4, or an empty block 4) is valid: it produces no body records and drains cleanly.

Writing SWIFT MT

A SWIFT Sink node re-emits each block-4 record as a :tag:value line and re-frames the single message envelope around them: the service blocks 1/2/3 first, then block 4 ({4: … -}), then the optional block-5 trailer. Block-4 free text is opaque, so values are written verbatim with no escaping — an interior {, }, a mid-line -}, a folded continuation break, and an interior blank line all reproduce as data. The writer opens block 4 with {4:\r\n, appends \r\n after each :tag:value, and closes with -} before any block-5 trailer. Existing LF and CRLF inside values remain unchanged. Re-reading returns the same field values even when the original input used LF structural separators.

Records map by the tag and value columns. The block column is the constant 4 discriminator (an empty block is treated as block 4, so a Transform that projects only tag/value writes fine); a record carrying a block other than 4 is rejected, because service blocks are never emitted as records — they ride the document context.

Strings are written verbatim; other scalar values use their natural display spelling, and null renders empty text. Arrays and maps are rejected. A tag must be nonempty and contain no colon, CR, or LF.

nodes:
  - type: sink
    name: out
    input: messages
    config:
      name: out
      type: swift
      path: ./out/message.swift
      options:
        basic_header_from_doc: basic
        app_header_from_doc: app
        user_header_from_doc: user
        trailer_from_doc: trailer

Each service block is written from a literal body or echoed from a user-declared $doc section:

OptionMeaning
basic_headerLiteral block-1 body, written verbatim as {1:<body>}.
basic_header_from_docName of a $doc section to echo the block-1 body from.
app_headerLiteral block-2 body.
app_header_from_docName of a $doc section to echo the block-2 body from.
user_headerLiteral block-3 body (nested {sub:tag} content kept verbatim).
user_header_from_docName of a $doc section to echo the block-3 body from.
trailerLiteral block-5 body, written after block 4 closes.
trailer_from_docName of a $doc section to echo the block-5 body from.

The *_from_doc options name the section the user declared on the source — the engine reserves no section name. A literal *_header wins over its *_from_doc companion when both are set (the same rule applies to trailer); a service block with neither is omitted. The _from_doc echo reads the block body verbatim from the section’s body field — the same single-field shape the reader writes — so a SWIFT source’s service blocks (declared as segment envelope sections) round-trip unchanged when their section names are passed back to the writer here.

The first successfully delivered record supplies the document service bodies; its trailer is retained until finalization, even if later records carry a different context. A selected section must contain body. Missing sections or fields, structured values, and unbalanced braces in service bodies fail before delivering that operation, including the first message header.

The document context that carries the $doc sections rides on each body record, so the *_from_doc echoes require at least one block-4 record to read from. Explicitly finalizing a library writer with zero records emits {4:\r\n-}, wrapped only by literal-configured service blocks; document echoes are skipped because no record supplies context. The CLI opens its writer lazily: a zero-record run publishes an empty file, even when literal service options are configured.

Block-4 free text has no escape mechanism, so values are written verbatim. Almost any value round-trips faithfully, but two shapes are unrepresentable when the value is built from arbitrary records (CSV/JSON → Transform → SWIFT): a value whose continuation line — a line after a folded line break — begins with the block terminator -} after LF or CR (which would re-read as an early block close), or with a : tag marker after LF (which would re-read as a spurious field). A bare CR before : is literal data. The writer rejects such a value with a clear error rather than emitting silently-corrupt output. Values read from a SWIFT source can never take these shapes, so a read → write → read round-trip is always safe.

A SWIFT MT message is a single indivisible envelope, so a swift output cannot be combined with a split: block, including max_records splitting — the pairing is rejected at config-validation time (diagnostic E342).

Library output and failures

Direct callers use SwiftEncoder::new with a schema, SwiftWriterConfig, and finite WriterResources, then PreparedWriter::new. A standalone MemoryOnlyResources provider requires an explicit nonzero budget. The resource-free writer constructor is unavailable.

The first record’s headers and body are prepared together; finalization prepares the closing block and trailer together. Preparation failure leaves the destination and committed state unchanged and permits a corrected retry. Delivery failure can leave an accepted prefix and poisons continuation. flush_bytes() drains without closing the message; flush() finalizes once and drains, with no duplicate trailer on repeated calls. Drop does not finalize or retry writes. See output preparation for resource and publication boundaries.

Limitations

  • UTF-8 only. SWIFT MT messages are decoded as UTF-8; a non-UTF-8 block id or body is rejected explicitly rather than corrupted silently.
  • One message per input. Adjacent complete messages are rejected. Reader materialization keeps its existing allocation behavior; the finite writer budget does not account for the reader’s retained fields or service blocks.
  • Field-content parsing. The reader exposes each :tag:value line as a tag/value pair verbatim. Parsing a field’s internal structure (the sub-fields of a :32A: value-date/currency/amount, say) is a CXL concern downstream of the source, not a reader responsibility.

Network Sources (REST)

A Source reads from the filesystem by default. To pull records from a network endpoint instead, declare a transport: block on the Source. The transport selects where records come from; it sits above the on-disk type: (the format), which for a REST source still selects how the response bodies decode.

A network transport is a finite-pull source: it runs on its own thread, drives a synchronous client to cursor exhaustion, then exits. There is no daemon, no event loop, and no async runtime — the same single- process, run-to-drain model as a file pipeline. Finiteness is a hard property of the reader: a REST source enforces explicit page and record limits, so an unbounded endpoint cannot keep it running forever. If the server offers another continuation after max_pages, the source fails closed instead of reporting a truncated pull as successful completion.

A network source still requires a schema: block. That authored schema is the row-to-record target: the reader maps each decoded object onto it, coercing values leniently. A per-row value that cannot coerce is left unchanged at the reader and routed to the dead-letter queue at the Transform stage — identical to file-source semantics. A network source declares no file matcher (path / glob / regex / paths); declaring one is a configuration error (E219).

Because a network source has no file path, its $source.file provenance column and the {source_file} output template both resolve to a stable synthetic identifier, <source:NAME>, where NAME is the Source node’s name.

REST sources

A rest source issues paginated HTTP GETs against a base URL, decoding each response body through the declared json or xml format. (Other formats are rejected with E220 — a REST body is a multi-record document, not a flat CSV/fixed-width stream.)

nodes:
  - type: source
    name: orders_api
    config:
      name: orders_api
      type: json
      options:
        format: array        # each page body is a JSON array of objects
      transport:
        kind: rest
        url: https://api.example.com/v1/orders
        max_pages: 50         # HARD page cap — required
        pagination:
          strategy: link_header
        auth:
          scheme: bearer
          token: "${ORDERS_TOKEN}"
      schema:
        - { name: order_id, type: int }
        - { name: total,    type: float }
        - { name: placed_at, type: date_time }

Pagination strategies

The pagination.strategy selects how the reader advances pages and detects the last one. max_records bounds emitted records. max_pages bounds requests and requires the server to reach an actual terminal page; an offered continuation beyond that bound is an error.

  • none (default) — a single GET; the body is the whole result.

  • offset — ?offset=N&limit=L, advancing the offset by the page size each request. The last page is the one that returns fewer rows than limit. A page containing exactly limit rows is not proof of end-of-input, so Clinker must issue one more bounded request to observe a short or empty terminal page. Size max_pages to leave room for that probe; reaching the cap first fails with page_limit_reached instead of accepting a possibly truncated result.

    pagination:
      strategy: offset
      limit: 200
      offset_param: offset     # optional, defaults shown
      limit_param: limit
    
  • cursor_token — the reader reads a continuation token from a JSON pointer in each response and sends it back on the next request. Paging stops when the token field is absent or null.

    pagination:
      strategy: cursor_token
      cursor_param: page_token
      next_token_pointer: /meta/next_page   # RFC 6901 JSON pointer
    
  • link_header — the reader follows the URL in the response’s RFC 8288 Link: <…>; rel="next" header until no such link is present. Registered relation tokens are case-insensitive, so next, Next, and NEXT have the same meaning.

    pagination:
      strategy: link_header
    

Continuation and redirect safety

Every server-directed continuation or redirect is resolved against the effective response URL and normalized before another request is built. Only the original normalized origin is allowed. Cross-origin targets, HTTPS-to-HTTP downgrades, malformed or conflicting rel="next" metadata, redirect or continuation cycles, and traversal beyond the configured bounds fail before a foreign or repeated request is sent.

Normalization resolves . and .. exactly as RFC 3986 does, and changes nothing else about the path. An empty segment is a segment: /v1//items/../p names /v1//p, not /v1/p, because a doubled slash is a different resource on any server that does not collapse it. A path ending in .. resolves to the directory above, trailing slash included. Two targets that differ only in an empty segment therefore stay two pages, and a pull that visits both is not a continuation cycle.

Link continuation metadata is parsed and authorized only when pagination.strategy is link_header. Other strategies ignore it because their continuation authority comes from the configured offset, cursor, or single-request contract.

A Link header is read as bytes, so a parameter this reader never consults — a title or a type carrying an accented character, an emoji, or anything else outside ASCII — does not affect the pull. Only the target inside <…> is decoded, because only the target has to become a URL: a target that is not valid UTF-8 is reported as malformed metadata. When a header carries several comma-separated links and one of them cannot be parsed, the rest are still read, so a reply naming two different next pages is reported as the conflict it is rather than as unreadable metadata.

Authentication

auth.scheme selects the credential sent on every request:

  • none (default) — no auth header.

  • bearer — sends Authorization: Bearer <token>.

  • header — sends an arbitrary static header, e.g. an API key.

    auth:
      scheme: header
      name: X-API-Key
      value: "${API_KEY}"
    

Corporate proxies

REST sources use the process proxy environment: ALL_PROXY, HTTPS_PROXY, or HTTP_PROXY (including their lowercase forms), with NO_PROXY/no_proxy for bypass rules. This lets the standalone clinker run CLI reach a third-party vendor through a corporate forward proxy without a central orchestrator or a pipeline-specific proxy key. Keep proxy credentials out of pipeline YAML and avoid printing credential-bearing proxy URLs.

The current Rust TLS configuration trusts the bundled public Web PKI roots. A proxy that tunnels HTTPS works when the vendor certificate remains visible and chains to those roots. A TLS-inspecting proxy that substitutes a certificate from a private corporate CA is not currently supported by a Clinker trust-store setting, even if that CA is installed in the operating-system store. In that case Clinker fails closed with a TLS/proxy classification; do not disable certificate verification as a workaround.

Reliability and finiteness knobs

KeyDefaultMeaning
max_pages—Required. Hard ceiling on pages fetched, regardless of the server.
max_recordsnoneOptional hard ceiling on records emitted.
retries3Bounded retries on a transient failure (5xx, connect/timeout error, or a transient body-delivery timeout/reset). A 4xx is fatal — retrying cannot help.
timeout_secs30Per-request timeout. Bounds in-flight time so an interrupt lands within the shutdown window.

Request diagnostics report a failure class, attempt number, page number, HTTP status when available, and the query-free target path. Authorization headers, request bodies, response bodies, and URL query values are never included. A proxy or vendor error therefore remains actionable without copying credentials or signed query parameters into logs.

A retryable body-delivery failure discards the partial body and retries the whole page within the same bounded retry budget. What counts as retryable is one rule for the whole request, whichever phase observed the failure: a dropped, reset, or timed-out exchange is retried, and a failure that would arrive identically on every attempt is not. Body-size violations, TLS failures, an unroutable URL, a host that does not resolve, and local material this process cannot read are therefore fatal at once rather than retried — reported against attempt 1, and without spending timeout_secs once per remaining attempt on a condition that cannot change. retries is a ceiling on attempts worth making, not a number of attempts every failure receives.

A partial-page decode failure routes that page’s offending rows to the DLQ per-row, exactly like a file source; it does not abort the pull.

Shutdown

On SIGINT/SIGTERM the reader polls its cancellation handle at each page boundary and stops cleanly with a normal end-of-input — the same graceful drain a file source performs. The timeout_secs per-request bound caps how long a single in-flight request can delay that stop.

A request already in flight when the signal arrives is reported as a cancellation too, not as a failing endpoint — including a page whose body was being read when the connection dropped, which the reader would otherwise have retried. What the signal took away was that retry, so the run’s outcome is the cancellation. An endpoint that would have failed identically however many attempts remained is still reported as that endpoint failure, because no retry was lost: a supervisor re-queuing the batch would only repeat it.

Auto-Widen & Schema Drift

When an input file carries columns the source’s declared schema: block does not name, Clinker decides what to do with them via the per-source on_unmapped policy. The default is auto_widen, which preserves the extra columns end-to-end so schema drift never silently breaks a pipeline. This page covers the three modes, how undeclared columns flow downstream, the output controls, and the related diagnostics.

The three modes

- type: source
  name: orders
  config:
    name: orders
    type: csv
    path: "./data/orders.csv"
    on_unmapped:
      mode: auto_widen     # default; other values: drop, reject
    schema:
      - { name: order_id, type: string }
      - { name: amount, type: float }
  • auto_widen (default) — undeclared input fields are carried along with each record and re-expanded to top-level columns at the output (when the Sink node’s include_unmapped is left at its default of true). Nothing is silently lost, and you don’t have to declare every column up front. How a widened column reaches the output depends on the output format — self-describing formats (JSON / NDJSON / XML) carry it per record, tabular CSV widens its header to the union of every record’s columns, and fixed-width (a positional layout with no room for undeclared columns) fails loudly rather than dropping it. See Output controls below.
  • drop — undeclared input fields are silently stripped at read time. The source carries only its declared schema:.
  • reject — any record carrying a field not in the declared schema fails the source with a diagnostic naming the offending field. The strict choice when unexpected columns should be treated as errors.

CXL expressions can only read fields you declared in schema: — carried-along undeclared fields are not visible to CXL, only to the output. To use an undeclared field in an expression, add it to the source schema:.

How undeclared columns flow downstream

Carried-along columns follow these rules through each node type:

Node typeBehavior
TransformPassed through unchanged (transforms are row-preserving).
AggregateDropped — per-row extra columns have no meaning on a grouped row. To keep one, add it to group_by or emit it explicitly.
CombineThe driver’s carried columns ride through; build-side ones are dropped. To keep a build-side field, emit it explicitly in the combine body via <build_qualifier>.<field>.
Route / MergePassed through. Merge requires every input to share the same on_unmapped policy — mixing fails with E315 (see below).
CompositionThe body inherits the parent’s carried columns and whatever the body’s last node carries flows back out.
OutputExpanded to top-level columns when include_unmapped: true (the default); stripped when false. How a widened column reaches a CSV / XML / fixed-width writer depends on the format — see Schema drift across records.

Output controls

- type: sink
  name: out
  input: src
  config:
    name: out
    type: json
    path: out.json
    include_unmapped: true    # default: true

When true (the default), undeclared fields the source carried along are expanded back to top-level columns at the sink — useful for pass-through pipelines where every original column should reach the output. Set include_unmapped: false to write only the columns explicitly emitted upstream.

include_unmapped is independent of include_correlation_keys: each can be set on its own, and include_correlation_keys never surfaces auto-widened columns.

Cross-format flow

Expansion happens before the writer runs, so a CSV source with auto_widen feeding a JSON output with include_unmapped: true produces JSON objects whose keys include both the declared columns and the absorbed ones:

input.csv:    id,extra,city
              1,foo,Paris

output.json:  {"id": "1", "extra": "foo", "city": "Paris"}

Schema drift across records (tabular formats)

Different records can carry different auto-widened columns — for example a Merge of two sources where one carries region and the other category, so region appears only on the first source’s rows and category only on the second’s. Each output format handles that heterogeneity differently:

  • JSON / NDJSON / XML are self-describing: each record writes its own keys/elements, so a column present on only some records is simply absent from the others. Nothing is lost.
  • CSV needs one header shared by every row. When the output can be materialized (the common buffered path), Clinker pre-scans the batch and writes a header that is the union of every record’s columns, in first-seen order; rows that lack a later-appearing column write an empty cell for it. Nothing is lost.
  • On a bounded-memory CSV path — a streaming output fused directly after a Merge/Transform, a single-branch Route, a streaming-strategy Aggregate, or the probe side of a hash-build-probe Combine, or an envelope-reconstructing output — the writer commits its header to the first record before it has seen the rest, so a union is impossible. A later record carrying a column the header lacks then fails the run loudly with a SchemaDrift error naming the format and column, rather than silently writing a narrower row. Declare the column in the source (or output) schema: so every record carries it, or route to a self-describing format.
  • Fixed-width is positional — every column occupies a declared byte range, and there is no room for an undeclared one — so any carried-along column reaching a fixed-width output is a SchemaDrift error. Fixed-width sources never auto-widen (see below), so this only arises when a fixed-width output sits downstream of a source that does.

Writer errors on unexpanded columns

CSV and fixed-width writers can only write flat scalar columns. If an unexpanded map reaches one of those writers, the write fails with an UnserializableMapValue error naming the format and column. JSON and XML can write a user-visible nested value natively, although the engine-stamped $widened sidecar is still expanded or stripped by the Output projection and is never an author-facing XML structure.

The fix is to either leave include_unmapped at its default of true, so the columns are expanded to top-level before writing, or to convert the value to a scalar in CXL before emitting it. The error message lists both routes.

E315 — Merge inputs must agree on policy

Merge concatenates its inputs positionally, so every input must agree on column shape — same column names, same on_unmapped policy, same correlation_key set. If two upstream sources disagree on whether they carry auto-widened columns (one uses auto_widen, another uses drop / reject), compilation fails:

E315: merge "merged": input schemas disagree on the `$widened` auto_widen sidecar column.

The fix is to set every merge upstream source to the same on_unmapped policy.

Fixed-width sources

Fixed-width sources are positional — the reader only sees the byte ranges your schema defines, so there are never any “extra” columns to absorb. auto_widen has no effect on a fixed-width source; use on_unmapped: drop (or reject) to make that explicit and silence the informational log the engine emits otherwise.

Clinker Expression Language Overview

Clinker Expression Language (CXL) is a per-record ETL expression language. Every program operates on one record at a time, producing output fields, filtering records, or computing derived values.

Programs are sequences of emit, let, filter, and distinct statements that execute from top to bottom against the current record.

Core expression forms

IntentCXL form
Produce or rename a fieldemit alias = col
Keep only matching recordsfilter condition
Combine conditionsand / or / not (keywords)
Choose the first non-null valuea ?? b
Choose between valuesif ... then ... else ... or match { }

Boolean operators are keywords

CXL uses English keywords for boolean logic, not symbols:

$ cxl eval -e 'emit result = true and false' --field dummy=1
{
  "result": false
}

The operators &&, ||, and ! are syntax errors in CXL. Always use and, or, and not.

System namespaces use $ prefix

CXL provides built-in namespaces for accessing pipeline state, metadata, and window functions. All system namespaces are prefixed with $:

  • $pipeline.* – pipeline execution context (name, counters, provenance) and pipeline-scope declared state
  • $source.* – per-source context and source-scope declared state
  • $record.* – per-record scoped state (travels with the record, never an output column)
  • $window.* – window function calls
  • $vars.* – static, channel-overridable configuration
  • $config.* – a composition’s config parameters, read inside its body (constant-folded per instantiation)
$ cxl eval -e 'emit name = $pipeline.name'
{
  "name": "cxl-eval"
}

Compile-time type checking

CXL is statically type-checked, so type errors are caught before any data is processed. Run cxl check to validate a transform before a run. Errors come with source locations and fix suggestions.

$ cxl check transform.cxl
ok: transform.cxl is valid

If there are type errors, the checker reports them with spans:

error[typecheck]: cannot apply '+' to String and Int (at transform.cxl:12)
  help: convert one operand — use .to_int() or .to_string()

A minimal CXL program

emit greeting = "hello"
emit doubled = amount * 2
filter amount > 0

This program:

  1. Emits a constant string field greeting
  2. Emits doubled as twice the input amount
  3. Filters out records where amount is not positive

Try it:

$ cxl eval -e 'emit greeting = "hello"' -e 'emit doubled = amount * 2' \
    --field amount=5
{
  "greeting": "hello",
  "doubled": 10
}

Statement order matters

CXL statements execute sequentially. Later statements can reference fields produced by earlier emit or let statements:

$ cxl eval -e 'let tax_rate = 0.21' -e 'emit tax = price * tax_rate' \
    --field price=100
{
  "tax": 21.0
}

A filter statement short-circuits execution – if the condition is false, remaining statements do not run and the record is excluded from output.

Types & Literals

CXL has 10 value types. Every field value, literal, and expression result is one of these types.

Value types

TypeRust backingDescription
NullValue::NullMissing or absent value
Boolbooltrue or false
Integeri6464-bit signed integer
Floatf6464-bit double-precision float
DecimalDecimalExact base-10 fixed-point number for money/financials
StringFieldStrUTF-8 text
DateNaiveDateCalendar date without timezone
DateTimeNaiveDateTimeDate and time without timezone
ArrayOwnedValuesOrdered collection of values
MapOwnedMapKey-value pairs

Literal syntax

Integers

Standard decimal notation. Negative values use the unary minus operator.

$ cxl eval -e 'emit a = 42' -e 'emit b = -5' -e 'emit c = 0'
{
  "a": 42,
  "b": -5,
  "c": 0
}

Floats

Decimal notation with a dot. Must have digits on both sides of the decimal point.

$ cxl eval -e 'emit a = 3.14' -e 'emit b = -0.5'
{
  "a": 3.14,
  "b": -0.5
}

Strings

Double-quoted or single-quoted. Supports escape sequences: \\, \", \', \n, \t, \r.

$ cxl eval -e 'emit greeting = "hello world"'
{
  "greeting": "hello world"
}

Booleans

The keywords true and false.

$ cxl eval -e 'emit flag = true' -e 'emit neg = not flag'
{
  "flag": true,
  "neg": false
}

Dates

Hash-delimited ISO 8601 format: #YYYY-MM-DD#.

$ cxl eval -e 'emit d = #2024-01-15#'
{
  "d": "2024-01-15"
}

Null

The keyword null.

$ cxl eval -e 'emit nothing = null'
{
  "nothing": null
}

Arrays and comprehensions

Array literals accept full expressions and preserve their written order:

emit values = [order_id, amount * 2, null]

Use one for clause and an optional trailing if to construct an array from another array:

emit positive_doubles = [item * 2 for item in values if item > 0]

The source must be an array; null and scalar sources are errors. The binding is local to the item expression and predicate, cannot destructure, and cannot shadow an input field or surrounding let binding.

Maps

Map literals preserve key insertion order. Bare identifiers and quoted strings are static keys; brackets hold a computed expression whose result must be a non-null string:

emit payload = {
  customer: customer_name,
  items: [{sku: item.sku, quantity: item.quantity} for item in line_items],
  [dynamic_key]: dynamic_value,
}

Duplicate keys are errors, including two differently escaped spellings that decode to the same logical key. Nested maps and arrays are limited to 64 container levels, and CXL construction is limited to 10 MiB per input record; both limits fail the record instead of growing without bound.

Nested keys use one canonical escape grammar. After CXL string decoding, one leading backslash makes a reserved-looking key literal: \@name, \#text, or \\name. In CXL source each backslash in a quoted string is itself escaped, so write "\\@name", "\\#text", or "\\\\name". Other leading-backslash forms are rejected. Output formats decide how the neutral nested value is encoded. JSON removes the structural escape when writing the key; XML assigns roles to the unescaped forms as described in Writing XML.

The runnable examples/pipelines/nested_values.yaml pipeline sends one constructed value to both JSON and XML so the two native encodings can be compared directly.

Schema types

When declaring column types in YAML pipeline schemas, use these type names:

Schema typeCXL typeDescription
stringStringText values
intInteger64-bit integers
floatFloat64-bit floats
decimalDecimalExact base-10 fixed-point (money) — see below
boolBoolBoolean values
dateDateCalendar dates
date_timeDateTimeDate and time
arrayArrayOrdered collections
numericInt or FloatUnion type – accepts either
anyAnyUnknown type – no type constraints
nullable(T)Nullable(T)Wrapper – value may be null

Example YAML schema declaration:

schema:
  employee_id: int
  name: string
  salary: nullable(float)
  start_date: date

Type promotion

CXL automatically promotes types in mixed expressions:

Int + Float promotes to Float:

$ cxl eval -e 'emit result = 2 + 3.5'
{
  "result": 5.5
}

Null + T produces Nullable(T): Any operation involving null produces a nullable result.

$ cxl eval -e 'emit result = null + 5'
{
  "result": null
}

Nullable(A) + B unifies to Nullable(unified): When a nullable value meets a non-nullable value, the result type wraps the unified inner type in Nullable.

The decimal type

float is an IEEE-754 binary float: it cannot represent most base-10 fractions exactly, so 0.1 + 0.2 is 0.30000000000000004, not 0.3. That rounding is unacceptable for money. The decimal type is an exact base-10 fixed-point number — 0.10 + 0.20 is exactly 0.30 — and is the correct type for monetary amounts, prices, tax, and any figure that must round like decimal arithmetic on paper.

Declare a decimal column with type: decimal and a scale (the number of fractional digits). precision (total significant digits) is optional validation metadata:

schema:
  - { name: amount, type: decimal, scale: 2 }
  - { name: tax_rate, type: decimal, scale: 4 }

A decimal column parses its raw text into an exact value and rounds off any excess precision to the column scale (round-half-to-even, the unbiased “banker’s rounding” used in accounting), so a scale: 2 column stores 2.567 as 2.57. This is one edge of a boundary contract: a declared scale pins a value to that many places at the boundary it is declared on — a source column’s scale on read, an output column’s scale on write (see Aggregating decimals) — while decimals keep full precision inside the pipeline.

Arithmetic rules

  • decimal ⊗ decimal → decimal — exact.
  • decimal ⊗ int → decimal — the integer widens exactly, so amount + 1 and price * quantity stay exact decimals.
  • decimal ⊗ float is a type error. Mixing an exact decimal with a binary float would silently lose precision, so CXL rejects it and asks for an explicit cast. Choose the trade-off deliberately:
    • amount.to_float() * rate — opt into binary float precision.
    • rate.to_decimal() * amount — bring the float into exact decimal math (the float→decimal step is the one acknowledged lossy conversion).
  • Division and avg compute at full precision — the exact quotient, not a binary-float approximation. Inside the pipeline a computed decimal keeps every digit; it is pinned to a fixed number of places only at a boundary that declares a scale. Declaring the output column type: decimal with a scale rounds the value to that many places on write (banker’s rounding), exactly as a decimal source column rounds on read — so avg(amount) emitted into a scale: 2 output column writes 1.33, not the full quotient. With no declared output scale the full precision is preserved; use an explicit round in CXL when you need fixed places mid-pipeline.

Comparisons follow the same rule: decimal < int is fine, decimal < float requires a cast.

The branches of a conditional follow it too. An if, a match or a ?? whose branches are a decimal and a float does not compile, because its result would be a decimal on some rows and a float on others:

cannot mix decimal and float without an explicit cast: the branches of this `if` are a decimal (`amount`) and a float (`price`); declare `price` a decimal in its Source schema, `type: decimal` in place of `type: float`, so the branches have one numeric type

The message gives one fix. When the float is a Source column, declare it type: decimal in its Source schema: the reader parses the column’s text exactly, so if flag then amount else price is a decimal holding the values the file holds. (A JSON number read into a decimal column is still parsed through a float first; see #1299.) When the float is computed rather than read from a Source column, convert the decimal side instead: if flag then amount.to_float() else price * 2.0 is a float. Converting a float with .to_decimal() does not make it exact: the decimal keeps the float’s binary digits.

Casting

x.to_decimal() converts an int, string, or float into a decimal (try_decimal is the lenient form that yields null on failure). d.to_int(), d.to_float(), and d.to_string() convert a decimal back out.

Worked example — an exact invoice total

$ cxl eval -e 'emit total = ("19.99".to_decimal() * 3) + "4.80".to_decimal()'
{
  "total": "64.77"
}

19.99 * 3 = 59.97, + 4.80 = 64.77 — exact, with no binary-float drift. (In a pipeline, declare the source columns type: decimal instead of casting; JSON output renders a decimal as a scale-preserving string.)

Aggregating decimals

sum, avg, min, max, count, and distinct all work over a decimal column and stay exact — no binary float ever touches a running total:

  • sum(amount) returns a decimal: the exact total of the group’s values at the largest scale among them, rounded once (half to even) only when it does not fit a decimal at that scale. The sum of 1.00, -1.00 and 2 is 2.00 in any order, and a total of amounts with two decimal places is exact to the cent.
  • avg(amount) is sum(amount) / count(amount), a decimal at full division precision, and weighted_avg(v, w) is sum(v * w) / sum(w): the same digits and scale as those expressions give.
  • Only a group with no non-null value gives null. A decimal total outside the decimal range, a weighted_avg whose weights total zero or whose row product is out of range, and a group holding both decimals and floats each fail the group with an aggregate_finalize error that names the fix (see Aggregate functions).
  • min / max return the exact extremum, and count returns an integer.
  • Group-by and distinct keys are scale-normalized: two decimals that are numerically equal group together regardless of scale, so 2.50 and 2.5 fall in one group. This holds even when the aggregation spills to disk.

To pin an aggregate result to fixed places on the way out, declare the output column type: decimal with a scale: the value is rounded to that scale on write (banker’s rounding), so avg(amount) over 1.00, 1.00, 2.00 writes 1.33 into a scale: 2 output column while sum(amount) stays 4.00. An output column with no declared scale keeps the full-precision quotient. This applies to every format — CSV, JSON, and fixed-width — because the rounding happens as the record is projected onto the output. For fixed-width output it is often required: a full-precision quotient overflows a narrow numeric field, which is a hard error, whereas the rounded value fits.

weighted_avg also stays exact over decimals: a decimal value or weight (or both) gives sum(value * weight) / sum(weight) over exact totals, at full division precision. A zero total weight is an error, as x / 0 is. A decimal in one position mixed with a binary float in the other is a type error, matching the decimal ⊗ float arithmetic rule. Declare a float Source column type: decimal so the value and weight share one numeric domain, or, when the float is computed, convert the decimal argument with .to_float().

Type unification rules

When two types meet in an expression, CXL coerces them automatically:

  • Numbers combine: mixing an integer and a float gives a float (2 + 3.5 is 5.5).
  • A decimal and a float never combine, in an operator or in the branches of an if, match or ??: declare a float Source column type: decimal, or convert the decimal side with .to_float() when the float is computed.
  • Arithmetic and ordering comparisons with null give null. == and != never do (null == null is true), and and/or give a definite answer when the other side settles it. See Null Handling.
  • Mismatched types are an error: String + Int fails. Convert first with .to_int() or .to_string() so both sides are the same type.

Operators & Expressions

CXL provides arithmetic, comparison, boolean, null coalescing, and string operators. Boolean logic uses keywords (and, or, not), not symbols.

Arithmetic operators

OperatorDescriptionExample
+Addition (or string concatenation)2 + 3
-Subtraction10 - 4
*Multiplication3 * 5
/Division10 / 3
%Modulo (remainder)10 % 3
$ cxl eval -e 'emit result = 2 + 3 * 4'
{
  "result": 14
}

Multiplication binds tighter than addition, so 2 + 3 * 4 is 2 + (3 * 4) = 14, not (2 + 3) * 4 = 20.

$ cxl eval -e 'emit result = 10 % 3'
{
  "result": 1
}

Comparison operators

OperatorDescriptionExample
==Equalx == 0
!=Not equalx != 0
>Greater thanx > 10
<Less thanx < 10
>=Greater than or equalx >= 10
<=Less than or equalx <= 10
$ cxl eval -e 'emit result = 5 > 3' --field dummy=1
{
  "result": true
}

Boolean operators

CXL uses keywords for boolean logic. The symbols &&, ||, and ! are not valid CXL syntax.

OperatorDescriptionExample
andLogical ANDa and b
orLogical ORa or b
notLogical NOT (unary)not a
$ cxl eval -e 'emit result = true and not false'
{
  "result": true
}
$ cxl eval -e 'emit result = 5 > 3 or 10 < 2'
{
  "result": true
}

Null coalesce operator

The ?? operator returns its left operand if non-null, otherwise its right operand.

$ cxl eval -e 'emit result = null ?? "default"'
{
  "result": "default"
}
$ cxl eval -e 'emit result = "present" ?? "default"'
{
  "result": "present"
}

Like the branches of an if, the two sides of ?? must not be a decimal and a float: amount ?? price does not compile. As for if, the fix is to declare price with type: decimal in its Source schema, or, when the float side is computed rather than a Source column, to convert the decimal side with .to_float().

String concatenation

Use .concat() to join strings in compiled pipelines. Numeric + is not a substitute for explicit text concatenation at the planner’s type boundary.

$ cxl eval -e 'emit result = "hello".concat(" ", "world")'
{
  "result": "hello world"
}

Unary operators

OperatorDescriptionExample
-Numeric negation-x
notBoolean negationnot done
$ cxl eval -e 'emit result = -42'
{
  "result": -42
}

Method calls

Methods are called on a receiver using dot notation:

$ cxl eval -e 'emit result = "hello".upper()'
{
  "result": "HELLO"
}

Methods can be chained:

$ cxl eval -e 'emit result = "  hello  ".trim().upper()'
{
  "result": "HELLO"
}

Field references

Bare identifiers reference fields from the input record:

$ cxl eval -e 'emit result = price * qty' \
    --field price=10 \
    --field qty=3
{
  "result": 30
}

Qualified field references use dot notation for multi-source pipelines: source.field.

Operator precedence

From highest (binds tightest) to lowest:

PrecedenceOperatorsAssociativity
1 (highest). (method calls, field access)Left
2- (unary), notPrefix
3* / %Left
4+ -Left
5== != > < >= <=Left
6andLeft
7orLeft
8 (lowest)??Right

Use parentheses to override precedence:

$ cxl eval -e 'emit result = (2 + 3) * 4'
{
  "result": 20
}

Comments

Line comments start with # (when not followed by a digit – digit-prefixed # starts a date literal):

# This is a comment
emit total = price * qty  # inline comment
emit deadline = #2024-12-31#  # this is a date literal, not a comment

Statements

CXL programs are sequences of statements that execute top-to-bottom against each input record. Statement order matters – later statements can reference values produced by earlier ones.

emit

The emit statement produces an output field. Each emit becomes a column in the output record.

emit name = expression
$ cxl eval -e 'emit greeting = "hello"' -e 'emit doubled = 21 * 2'
{
  "greeting": "hello",
  "doubled": 42
}

Multiple emit statements build up the output record field by field:

$ cxl eval -e 'emit first = "Alice"' -e 'emit last = "Smith"' \
    -e 'emit full = first.concat(" ", last)'
{
  "first": "Alice",
  "last": "Smith",
  "full": "Alice Smith"
}

let

The let statement creates a local variable binding. The variable is available to subsequent statements but is NOT included in the output record.

let name = expression
$ cxl eval -e 'let tax_rate = 0.21' -e 'emit tax = 100 * tax_rate'
{
  "tax": 21.0
}

Note that tax_rate does not appear in the output – only emit statements produce output fields.

filter

The filter statement keeps a record only when its condition is exactly true; a condition that is false or null excludes it (see Null in conditions). When a filter excludes a record, remaining statements do not execute (short-circuit).

filter condition
$ cxl eval -e 'filter amount > 0' -e 'emit result = amount * 2' \
    --field amount=5
{
  "result": 10
}

When the filter condition is false or null, the entire record is dropped and no output is produced.

Filters can appear anywhere in the statement sequence. Place them early to skip unnecessary computation:

filter status == "active"
let discount = if tier == "gold" then 0.2 else 0.1
emit final_price = price * (1 - discount)

distinct

The distinct statement deduplicates records. The bare form deduplicates on all emitted fields. The by form deduplicates on a specific field.

distinct
distinct by field_name

In a pipeline, distinct tracks values seen so far and drops records that have already been emitted with the same key.

emit to a scoped namespace

An emit whose target is a $pipeline.*, $source.*, or $record.* name writes a producer-declared scoped variable instead of an output column. The variable must be listed in the writing Transform’s config.declares: block. $record.* is the per-record store that travels with the record but never serializes as an output column.

emit $record.quality_flag = if amount < 0 then "suspect" else "ok"

Read it downstream via the same namespace:

filter $record.quality_flag == "ok"

See Scoped Variables for the declaration model and the three scopes’ lifetimes.

trace

The trace statement emits debug logging. It has no effect on the output record. Trace messages are only visible when tracing is enabled at the appropriate level.

trace "processing record"
trace warn "unusual value detected"
trace info if amount > 10000 then "high value transaction"

Trace levels: trace (default), debug, info, warn, error. An optional guard condition (via if) limits when the trace fires.

Statement ordering

Statements execute sequentially. A statement can reference any field or variable defined by a preceding emit or let:

$ cxl eval -e 'let base = 100' -e 'let rate = 0.15' \
    -e 'emit subtotal = base * rate' \
    -e 'emit total = base + subtotal'
{
  "subtotal": 15.0,
  "total": 115.0
}

Referencing a name before it is defined is a resolve-time error:

emit total = base + tax    # error: 'base' is not defined yet
let base = 100
let tax = base * 0.21

use

The use statement imports a CXL module for reuse. See Modules & use for details.

use shared.dates as d
emit fy = d::fiscal_year(invoice_date)

Conditionals

CXL provides two conditional expression forms: if/then/else and match. Both are expressions – they return values and can be used anywhere an expression is expected.

If / then / else

The basic conditional expression:

if condition then value else alternative
$ cxl eval -e 'emit label = if amount > 100 then "high" else "low"' \
    --field amount=250
{
  "label": "high"
}

The else branch is optional. When omitted, records where the condition is false or null produce null. A null condition takes the else branch when there is one:

$ cxl eval -e 'emit bonus = if score > 90 then score * 0.1' \
    --field score=80
{
  "bonus": null
}

Branches of one numeric type

The two branches of an if must not be a decimal and a float, because the result would be a decimal on some rows and a float on others. Such an if does not compile:

cannot mix decimal and float without an explicit cast: the branches of this `if` are a decimal (`amount`) and a float (`price`); declare `price` a decimal in its Source schema, `type: decimal` in place of `type: float`, so the branches have one numeric type

The message gives one fix. When the float branch is a Source column, it is the column’s schema type: declare price with type: decimal in place of type: float, and the reader parses its text exactly, so both branches are decimals holding the values the file holds. (A JSON number read into a decimal column is still parsed through a float first; see #1299.) When the float branch is computed rather than read from a Source column, the fix converts the decimal branch with .to_float() instead, accepting binary float precision. An integer branch is fine beside either: it widens exactly into the decimal or float.

Chained conditionals

Chain multiple conditions with else if:

$ cxl eval -e 'emit tier = if amount > 1000 then "platinum"
    else if amount > 500 then "gold"
    else if amount > 100 then "silver"
    else "bronze"' \
    --field amount=750
{
  "tier": "gold"
}

Nested usage

Since if/then/else is an expression, it can be used inside other expressions:

$ cxl eval -e 'emit price = base * (if member then 0.8 else 1.0)' \
    --field base=100 \
    --field member=true
{
  "price": 80.0
}

Match

The match expression provides pattern matching. It comes in two forms: value matching (with a subject) and condition matching (without a subject).

Value form (with subject)

Match a subject expression against literal patterns:

match subject {
  pattern1 => result1,
  pattern2 => result2,
  _ => default
}
$ cxl eval -e 'emit label = match status {
    "A" => "Active",
    "I" => "Inactive",
    "P" => "Pending",
    _ => "Unknown"
  }' \
    --field status=A
{
  "label": "Active"
}

The wildcard _ is the catch-all arm. It matches any value not covered by preceding arms.

Condition form (without subject)

When no subject is provided, each arm’s pattern is evaluated as a boolean condition, and the first arm whose condition is exactly true wins. An arm whose condition is null is skipped:

match {
  condition1 => result1,
  condition2 => result2,
  _ => default
}
$ cxl eval -e 'emit tier = match {
    amount > 1000 => "high",
    amount > 100 => "medium",
    _ => "low"
  }' \
    --field amount=500
{
  "tier": "medium"
}

Practical examples

Tiered pricing:

emit discount = match {
  qty >= 1000 => 0.25,
  qty >= 100  => 0.15,
  qty >= 10   => 0.05,
  _           => 0.0
}

Status code mapping:

emit status_text = match http_code {
  200 => "OK",
  201 => "Created",
  400 => "Bad Request",
  404 => "Not Found",
  500 => "Internal Server Error",
  _   => "HTTP ".concat(http_code.to_string())
}

Region classification:

emit region = match country {
  "US" => "North America",
  "CA" => "North America",
  "MX" => "North America",
  "GB" => "Europe",
  "DE" => "Europe",
  "FR" => "Europe",
  _    => "Other"
}

Arms of one numeric type

As with if, the arms of a match must not include both a decimal and a float. The error names the first decimal arm and the first float arm, counted from 1, or the field when an arm is a bare field, and gives the same one fix: the Source schema type when the float arm is a float column, otherwise .to_float() on the decimal arm.

Match arms are evaluated in order

The first matching arm wins. Place more specific conditions before general ones:

# Correct: specific before general
emit category = match {
  amount > 10000 => "enterprise",
  amount > 1000  => "business",
  _              => "personal"
}

# Wrong: first arm always matches
emit category = match {
  amount > 0     => "personal",    # catches everything positive
  amount > 1000  => "business",    # never reached
  amount > 10000 => "enterprise",  # never reached
  _              => "unknown"
}

Built-in Methods

CXL provides built-in scalar methods organized into categories. Methods are called on a receiver value using dot notation: receiver.method(args).

Null propagation

Most methods return null when the receiver is null. This means null values flow through method chains without causing errors. The exceptions are documented in Introspection & Debug.

Method categories

String Methods (24 methods)

Text manipulation: case conversion, trimming, padding, searching, splitting, regex matching.

MethodDescription
upper, lowerCase conversion
trim, trim_start, trim_endWhitespace removal
starts_with, ends_with, containsSubstring testing
replaceFind and replace
substring, left, rightExtraction
pad_left, pad_rightPadding
repeat, reverseRepetition and reversal
lengthCharacter count
split, joinSplitting and joining
matches, find, captureRegex operations
format, concatFormatting and concatenation

Numeric Methods (8 methods)

Rounding, clamping, and comparison for integers and floats.

MethodDescription
absAbsolute value
ceil, floorCeiling and floor
round, round_toRounding to decimal places
clampConstrain to range
min, maxPairwise minimum/maximum

Date & Time Methods (13 methods)

Date component extraction, arithmetic, and formatting.

MethodDescription
year, month, dayDate component extraction
hour, minute, secondTime component extraction (DateTime only)
add_days, add_months, add_yearsDate arithmetic
diff_days, diff_months, diff_yearsDate difference
format_dateCustom date formatting

Conversion Methods (11 methods)

Type conversion in strict (error on failure) and lenient (null on failure) variants.

MethodDescription
to_int, to_float, to_string, to_boolStrict conversion
to_date, to_datetimeStrict date parsing
try_int, try_float, try_boolLenient conversion
try_date, try_datetimeLenient date parsing

Introspection & Debug (5 methods)

Type inspection, null checking, and debugging. These are the only methods that accept null receivers without propagating null.

MethodDescription
type_ofReturns the type name as a string
is_nullTests for null
is_emptyTests for empty string, empty array, or null
catchNull fallback (equivalent to ??)
debugPassthrough with tracing side effect

Path Methods (5 methods)

File path component extraction.

MethodDescription
file_nameFull filename with extension
file_stemFilename without extension
extensionFile extension
parentParent directory path
parent_nameParent directory name

Array Methods

Traversal and transformation over nested arrays. Closure-bearing methods take an arrow-syntax closure and evaluate it per element.

MethodDescription
filter, map, find, any, flat_mapClosure-bearing traversal
removeDrop the element at a given index
length, joinCross-listed on arrays (also defined on strings)

Map Methods

Builders and accessors for Value::Map payloads. All map methods return new maps – they never mutate the receiver.

MethodDescription
keys, valuesList map keys / values as arrays
mergeUnion of two maps (right wins on conflict)
setInsert / replace an entry, by single key or by a nested a.b[0].c path
remove_fieldDrop a single entry by top-level key
unsetDelete an entry by single key or by a nested a.b[0].c path (array index removes-and-shifts; missing path is a no-op)

String Methods

CXL provides 24 built-in methods for string manipulation. All string methods return null when the receiver is null (null propagation).

Case conversion

upper()

Converts all characters to uppercase.

$ cxl eval -e 'emit result = "hello world".upper()'
{
  "result": "HELLO WORLD"
}

lower()

Converts all characters to lowercase.

$ cxl eval -e 'emit result = "Hello World".lower()'
{
  "result": "hello world"
}

Whitespace trimming

trim()

Removes leading and trailing whitespace.

$ cxl eval -e 'emit result = "  hello  ".trim()'
{
  "result": "hello"
}

trim_start()

Removes leading whitespace only.

$ cxl eval -e 'emit result = "  hello  ".trim_start()'
{
  "result": "hello  "
}

trim_end()

Removes trailing whitespace only.

$ cxl eval -e 'emit result = "  hello  ".trim_end()'
{
  "result": "  hello"
}

Substring testing

starts_with(prefix: String) -> Bool

Tests whether the string starts with the given prefix.

$ cxl eval -e 'emit result = "hello world".starts_with("hello")'
{
  "result": true
}

ends_with(suffix: String) -> Bool

Tests whether the string ends with the given suffix.

$ cxl eval -e 'emit result = "report.csv".ends_with(".csv")'
{
  "result": true
}

contains(substring: String) -> Bool

Tests whether the string contains the given substring.

$ cxl eval -e 'emit result = "hello world".contains("lo wo")'
{
  "result": true
}

Find and replace

replace(find: String, replacement: String) -> String

Replaces all occurrences of find with replacement.

$ cxl eval -e 'emit result = "foo-bar-baz".replace("-", "_")'
{
  "result": "foo_bar_baz"
}

Extraction

substring(start: Int [, length: Int]) -> String

Extracts a substring starting at start (0-based character index). If length is provided, takes at most that many characters. If omitted, takes all remaining characters.

$ cxl eval -e 'emit result = "hello world".substring(6)'
{
  "result": "world"
}
$ cxl eval -e 'emit result = "hello world".substring(0, 5)'
{
  "result": "hello"
}

left(n: Int) -> String

Returns the first n characters.

$ cxl eval -e 'emit result = "hello world".left(5)'
{
  "result": "hello"
}

right(n: Int) -> String

Returns the last n characters.

$ cxl eval -e 'emit result = "hello world".right(5)'
{
  "result": "world"
}

Padding

pad_left(width: Int [, char: String]) -> String

Left-pads the string to the given width. Default pad character is a space.

$ cxl eval -e 'emit result = "42".pad_left(5, "0")'
{
  "result": "00042"
}
$ cxl eval -e 'emit result = "hi".pad_left(6)'
{
  "result": "    hi"
}

pad_right(width: Int [, char: String]) -> String

Right-pads the string to the given width. Default pad character is a space.

$ cxl eval -e 'emit result = "hi".pad_right(6, ".")'
{
  "result": "hi...."
}

Repetition and reversal

repeat(n: Int) -> String

Repeats the string n times.

$ cxl eval -e 'emit result = "ab".repeat(3)'
{
  "result": "ababab"
}

reverse() -> String

Reverses the characters in the string.

$ cxl eval -e 'emit result = "hello".reverse()'
{
  "result": "olleh"
}

Length

length() -> Int

Returns the number of characters in the string. Also works on arrays, returning the number of elements.

$ cxl eval -e 'emit result = "hello".length()'
{
  "result": 5
}

Splitting and joining

split(delimiter: String) -> Array

Splits the string by the delimiter, returning an array of strings.

$ cxl eval -e 'emit result = "a,b,c".split(",")'
{
  "result": ["a", "b", "c"]
}

join(delimiter: String) -> String

Joins an array of values into a string with the given delimiter. The receiver must be an array.

$ cxl eval -e 'emit result = "a,b,c".split(",").join(" - ")'
{
  "result": "a - b - c"
}

Regex operations

matches(pattern: String) -> Bool

Tests whether the string fully matches the given regex pattern.

$ cxl eval -e 'emit result = "abc123".matches("^[a-z]+[0-9]+$")'
{
  "result": true
}

find(pattern: String) -> Bool

Tests whether the string contains a substring matching the given regex pattern (partial match).

$ cxl eval -e 'emit result = "hello world 42".find("[0-9]+")'
{
  "result": true
}

capture(pattern: String [, group: Int]) -> String

Extracts a capture group from the first regex match. Default group is 0 (the full match).

$ cxl eval -e 'emit result = "order-12345".capture("order-([0-9]+)", 1)'
{
  "result": "12345"
}

Formatting and concatenation

format(fmt: String) -> String

Formats the receiver value as a string.

$ cxl eval -e 'emit result = 42.format("")'
{
  "result": "42"
}

concat(args: String…) -> String

Concatenates the receiver with one or more string arguments. Null arguments are treated as empty strings.

$ cxl eval -e 'emit result = "hello".concat(" ", "world")'
{
  "result": "hello world"
}

This is variadic – it accepts any number of string arguments:

$ cxl eval -e 'emit result = "a".concat("b", "c", "d")'
{
  "result": "abcd"
}

Numeric Methods

CXL provides 8 built-in methods for numeric operations. These methods work on both Integer and Float values (the Numeric receiver type). All return null when the receiver is null.

abs, ceil, floor, round, and round_to also accept an exact decimal receiver and return an exact decimal result — d.round(2) rounds d to two fractional digits using banker’s rounding, staying in the decimal domain rather than converting to a float. The type checker infers decimal for these calls as well, so the result composes with other decimals without a cast: amount.round_to(2) + fee typechecks when both columns are decimal.

abs() -> Numeric

Returns the absolute value. Preserves the original type (Int stays Int, Float stays Float).

$ cxl eval -e 'emit result = (-42).abs()'
{
  "result": 42
}
$ cxl eval -e 'emit result = (-3.14).abs()'
{
  "result": 3.14
}

ceil() -> Int

Rounds up to the nearest integer. Returns the value unchanged for integers.

$ cxl eval -e 'emit result = 3.2.ceil()'
{
  "result": 4
}
$ cxl eval -e 'emit result = (-3.2).ceil()'
{
  "result": -3
}

floor() -> Int

Rounds down to the nearest integer. Returns the value unchanged for integers.

$ cxl eval -e 'emit result = 3.8.floor()'
{
  "result": 3
}
$ cxl eval -e 'emit result = (-3.2).floor()'
{
  "result": -4
}

round([decimals: Int]) -> Float

Rounds to the specified number of decimal places. Default is 0 decimal places. On a decimal receiver the result is a decimal (exact banker’s rounding), not a float.

$ cxl eval -e 'emit result = 3.456.round()'
{
  "result": 3.0
}
$ cxl eval -e 'emit result = 3.456.round(2)'
{
  "result": 3.46
}

round_to(decimals: Int) -> Float

Rounds to the specified number of decimal places. Unlike round(), the decimals argument is required. On a decimal receiver the result is a decimal (exact banker’s rounding), not a float.

$ cxl eval -e 'emit result = 3.14159.round_to(3)'
{
  "result": 3.142
}

Use round_to when you want to be explicit about precision in financial or scientific calculations:

$ cxl eval -e 'emit price = 19.995.round_to(2)'
{
  "price": 20.0
}

clamp(min: Numeric, max: Numeric) -> Numeric

Constrains the value to the given range. Returns min if the value is below it, max if above, or the value itself if within range.

$ cxl eval -e 'emit result = 150.clamp(0, 100)'
{
  "result": 100
}
$ cxl eval -e 'emit result = (-5).clamp(0, 100)'
{
  "result": 0
}
$ cxl eval -e 'emit result = 50.clamp(0, 100)'
{
  "result": 50
}

min(other: Numeric) -> Numeric

Returns the smaller of the receiver and the argument.

$ cxl eval -e 'emit result = 10.min(20)'
{
  "result": 10
}
$ cxl eval -e 'emit result = 10.min(5)'
{
  "result": 5
}

max(other: Numeric) -> Numeric

Returns the larger of the receiver and the argument.

$ cxl eval -e 'emit result = 10.max(20)'
{
  "result": 20
}
$ cxl eval -e 'emit result = 10.max(5)'
{
  "result": 10
}

Practical examples

Clamp a percentage:

emit pct = (completed / total * 100).clamp(0, 100).round_to(1)

Absolute difference:

emit diff = (actual - expected).abs()

Floor division for batch numbering:

emit batch = (row_number / 1000).floor()

Date & Time Methods

CXL provides 13 built-in methods for date and time manipulation. These methods work on Date and DateTime values. All return null when the receiver is null.

Component extraction

year() -> Int

Returns the year component.

$ cxl eval -e 'emit result = #2024-03-15#.year()'
{
  "result": 2024
}

month() -> Int

Returns the month component (1-12).

$ cxl eval -e 'emit result = #2024-03-15#.month()'
{
  "result": 3
}

day() -> Int

Returns the day-of-month component (1-31).

$ cxl eval -e 'emit result = #2024-03-15#.day()'
{
  "result": 15
}

hour() -> Int

Returns the hour component (0-23). DateTime only – returns null for Date values.

$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime().hour()'
{
  "result": 14
}

minute() -> Int

Returns the minute component (0-59). DateTime only – returns null for Date values.

$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime().minute()'
{
  "result": 30
}

second() -> Int

Returns the second component (0-59). DateTime only – returns null for Date values.

$ cxl eval -e 'emit result = "2024-03-15T14:30:45".to_datetime().second()'
{
  "result": 45
}

Date arithmetic

add_days(n: Int) -> Date

Adds n days to the date. Use negative values to subtract. Works on both Date and DateTime.

$ cxl eval -e 'emit result = #2024-01-15#.add_days(10)'
{
  "result": "2024-01-25"
}
$ cxl eval -e 'emit result = #2024-01-15#.add_days(-5)'
{
  "result": "2024-01-10"
}

add_months(n: Int) -> Date

Adds n months to the date. Day is clamped to the last day of the target month if necessary.

$ cxl eval -e 'emit result = #2024-01-31#.add_months(1)'
{
  "result": "2024-02-29"
}
$ cxl eval -e 'emit result = #2024-03-15#.add_months(-2)'
{
  "result": "2024-01-15"
}

add_years(n: Int) -> Date

Adds n years to the date. Leap day (Feb 29) is clamped to Feb 28 in non-leap years.

$ cxl eval -e 'emit result = #2024-02-29#.add_years(1)'
{
  "result": "2025-02-28"
}

Date difference

diff_days(other: Date) -> Int

Returns the number of days between the receiver and the argument (receiver - other). Positive when the receiver is later.

$ cxl eval -e 'emit result = #2024-03-15#.diff_days(#2024-03-01#)'
{
  "result": 14
}
$ cxl eval -e 'emit result = #2024-01-01#.diff_days(#2024-03-15#)'
{
  "result": -74
}

diff_months(other: Date) -> Int

Returns the difference in months between two dates.

Note: This method currently returns null (unimplemented). Use diff_days and divide by 30 as an approximation.

diff_years(other: Date) -> Int

Returns the difference in years between two dates.

Note: This method currently returns null (unimplemented). Use diff_days and divide by 365 as an approximation.

Formatting

format_date(format: String) -> String

Formats the date/datetime using a chrono format string. See chrono format syntax.

Common format specifiers:

SpecifierDescriptionExample
%Y4-digit year2024
%m2-digit month03
%d2-digit day15
%HHour (24h)14
%MMinute30
%SSecond00
%BFull month nameMarch
%bAbbreviated monthMar
%AFull weekdayFriday
$ cxl eval -e 'emit result = #2024-03-15#.format_date("%B %d, %Y")'
{
  "result": "March 15, 2024"
}
$ cxl eval -e 'emit result = #2024-03-15#.format_date("%Y/%m/%d")'
{
  "result": "2024/03/15"
}

Practical examples

Fiscal year calculation (April start):

let d = invoice_date
emit fiscal_year = if d.month() < 4 then d.year() - 1 else d.year()

Age in days:

emit days_since = now.diff_days(created_date)

Quarter:

emit quarter = match {
  invoice_date.month() <= 3  => "Q1",
  invoice_date.month() <= 6  => "Q2",
  invoice_date.month() <= 9  => "Q3",
  _                          => "Q4"
}

ISO week format:

emit formatted = order_date.format_date("%Y-W%V")

Conversion Methods

CXL provides two families of conversion methods: strict (7 methods) and lenient (6 methods). Strict conversions raise an error on failure, halting pipeline execution. Lenient conversions return null on failure, allowing graceful handling of dirty data.

All conversion methods accept any receiver type (Any).

Strict conversions

Use strict conversions for required fields where invalid data should halt processing.

to_int() -> Int

Converts the receiver to an integer. Errors on failure.

  • Float: truncates toward zero
  • String: parses as integer
  • Bool: true becomes 1, false becomes 0
$ cxl eval -e 'emit result = "42".to_int()'
{
  "result": 42
}
$ cxl eval -e 'emit result = 3.9.to_int()'
{
  "result": 3
}

to_float() -> Float

Converts the receiver to a float. Errors on failure.

  • Integer: promotes to float
  • String: parses as float
$ cxl eval -e 'emit result = "3.14".to_float()'
{
  "result": 3.14
}
$ cxl eval -e 'emit result = 42.to_float()'
{
  "result": 42.0
}

to_decimal() -> Decimal

Converts the receiver to an exact decimal. Errors on failure.

  • Integer: converts exactly
  • String: parses base-10 exactly ("19.99" → 19.99, never via a binary float)
  • Float: converts via the binary value — this is the one lossy direction, made explicit precisely because decimal * float is otherwise a type error

Use to_decimal() to bring a value into exact decimal arithmetic. JSON output renders a decimal as a scale-preserving string.

$ cxl eval -e 'emit result = "0.10".to_decimal() + "0.20".to_decimal()'
{
  "result": "0.30"
}

to_string() -> String

Converts any value to its string representation. Never fails.

$ cxl eval -e 'emit result = 42.to_string()'
{
  "result": "42"
}
$ cxl eval -e 'emit result = true.to_string()'
{
  "result": "true"
}

to_bool() -> Bool

Converts the receiver to a boolean. Errors on failure.

  • String: "true", "1", "yes" become true; "false", "0", "no" become false (case-insensitive)
  • Integer: 0 is false, everything else is true
$ cxl eval -e 'emit result = "yes".to_bool()'
{
  "result": true
}
$ cxl eval -e 'emit result = 0.to_bool()'
{
  "result": false
}

to_date([format: String]) -> Date

Parses a string to a Date. Without a format argument, expects ISO 8601 (YYYY-MM-DD). With a format, uses chrono strftime syntax.

$ cxl eval -e 'emit result = "2024-03-15".to_date()'
{
  "result": "2024-03-15"
}
$ cxl eval -e 'emit result = "15/03/2024".to_date("%d/%m/%Y")'
{
  "result": "2024-03-15"
}

to_datetime([format: String]) -> DateTime

Parses a string to a DateTime. Without a format argument, expects ISO 8601 (YYYY-MM-DDTHH:MM:SS). With a format, uses chrono strftime syntax.

$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime()'
{
  "result": "2024-03-15T14:30:00"
}

Lenient conversions

Use lenient conversions for optional or dirty data fields. They return null instead of raising errors, making them safe to combine with ?? for fallback values.

try_int() -> Int

Attempts to convert to integer. Returns null on failure.

$ cxl eval -e 'emit a = "42".try_int()' -e 'emit b = "abc".try_int()'
{
  "a": 42,
  "b": null
}

try_float() -> Float

Attempts to convert to float. Returns null on failure.

$ cxl eval -e 'emit a = "3.14".try_float()' -e 'emit b = "N/A".try_float()'
{
  "a": 3.14,
  "b": null
}

try_decimal() -> Decimal

Attempts to convert to an exact decimal. Returns null on failure.

$ cxl eval -e 'emit a = "19.99".try_decimal()' -e 'emit b = "N/A".try_decimal()'
{
  "a": "19.99",
  "b": null
}

try_bool() -> Bool

Attempts to convert to boolean. Returns null on failure.

$ cxl eval -e 'emit a = "yes".try_bool()' -e 'emit b = "maybe".try_bool()'
{
  "a": true,
  "b": null
}

try_date([format: String]) -> Date

Attempts to parse a string as a Date. Returns null on failure.

$ cxl eval -e 'emit a = "2024-03-15".try_date()' \
    -e 'emit b = "not a date".try_date()'
{
  "a": "2024-03-15",
  "b": null
}

try_datetime([format: String]) -> DateTime

Attempts to parse a string as a DateTime. Returns null on failure.

$ cxl eval -e 'emit a = "2024-03-15T14:30:00".try_datetime()' \
    -e 'emit b = "invalid".try_datetime()'
{
  "a": "2024-03-15T14:30:00",
  "b": null
}

When to use each

Strict conversions (to_*) for:

  • Required fields that must be valid
  • Schema-enforced data where bad input should halt the pipeline
  • Fields already validated upstream

Lenient conversions (try_*) for:

  • Optional fields that may be missing or malformed
  • Dirty data with mixed formats
  • Fields where a fallback value is acceptable

Practical patterns

Safe numeric parsing with fallback:

emit amount = raw_amount.try_float() ?? 0.0

Parse dates from multiple formats:

emit parsed = raw_date.try_date("%Y-%m-%d")
    ?? raw_date.try_date("%m/%d/%Y")
    ?? raw_date.try_date("%d-%b-%Y")

Strict conversion for required fields:

emit employee_id = raw_id.to_int()    # halts on bad data -- correct behavior
emit salary = raw_salary.to_float()   # must be numeric

Lenient conversion for optional fields:

emit bonus = raw_bonus.try_float()    # null if missing or non-numeric
emit total = salary + (bonus ?? 0.0)  # safe arithmetic

Introspection & Debug

CXL provides 4 introspection methods and 1 debug method. The four introspection methods are the only methods that accept null receivers without propagating null – they are designed specifically for inspecting and handling null values. debug follows ordinary null propagation.

type_of() -> String

Returns the type name of the receiver as a string. Works on any value, including null.

Type name strings: "string", "int", "float", "decimal", "bool", "date", "datetime", "null", "array", "map".

$ cxl eval -e 'emit a = 42.type_of()' -e 'emit b = "hello".type_of()' \
    -e 'emit c = null.type_of()'
{
  "a": "int",
  "b": "string",
  "c": "null"
}

Useful for branching on dynamic types:

emit formatted = match value.type_of() {
  "int"   => value.to_string().concat(" (integer)"),
  "float" => value.round_to(2).to_string().concat(" (decimal)"),
  _       => value.to_string()
}

is_null() -> Bool

Returns true if the receiver is null, false otherwise. This is the primary way to test for null values – it is NOT subject to null propagation.

$ cxl eval -e 'emit a = null.is_null()' -e 'emit b = 42.is_null()'
{
  "a": true,
  "b": false
}

Use in filter statements:

filter not field.is_null()

is_empty() -> Bool

Returns true for empty strings, empty arrays, or null values. Returns false for all other values.

$ cxl eval -e 'emit a = "".is_empty()' -e 'emit b = "hello".is_empty()' \
    -e 'emit c = null.is_empty()'
{
  "a": true,
  "b": false,
  "c": true
}

Useful for filtering out blank or missing records:

filter not name.is_empty()

catch(fallback: Any) -> Any

Returns the receiver if it is non-null, otherwise returns the fallback value. This is the method equivalent of the ?? operator.

$ cxl eval -e 'emit a = null.catch("default")' \
    -e 'emit b = "present".catch("default")'
{
  "a": "default",
  "b": "present"
}

catch and ?? are interchangeable:

# These two are equivalent:
emit name = raw_name.catch("Unknown")
emit name = raw_name ?? "Unknown"

debug(label: String) -> Any

Passes the receiver through unchanged while emitting a trace log with the given label. Zero overhead when tracing is disabled. The return value is always the receiver, making it safe to insert into any expression chain. A null receiver returns null without logging.

$ cxl eval -e 'emit result = 42.debug("check value")'
{
  "result": 42
}

Insert debug anywhere in a method chain for inspection without affecting the output:

emit total = price.debug("price")
    * qty.debug("qty")

When tracing is enabled, this produces log lines like:

TRACE source_row=1 source_file=input.csv: price: Integer(100)
TRACE source_row=1 source_file=input.csv: qty: Integer(5)

Null-safe summary

MethodNull receiver behavior
type_of()Returns "null"
is_null()Returns true
is_empty()Returns true
catch(x)Returns x
debug(l)Returns null, logs nothing (propagation)
All other methodsReturn null (propagation)

Path Methods

CXL provides 5 built-in methods for extracting components from file path strings. All path methods take a string receiver and return a string. They return null when the receiver is null or when the requested component does not exist.

file_name() -> String

Returns the full filename (with extension) from the path.

$ cxl eval -e 'emit result = "/data/reports/sales.csv".file_name()'
{
  "result": "sales.csv"
}

file_stem() -> String

Returns the filename without the extension.

$ cxl eval -e 'emit result = "/data/reports/sales.csv".file_stem()'
{
  "result": "sales"
}

extension() -> String

Returns the file extension (without the leading dot).

$ cxl eval -e 'emit result = "/data/reports/sales.csv".extension()'
{
  "result": "csv"
}

Returns null when no extension is present:

$ cxl eval -e 'emit result = "/data/reports/README".extension()'
{
  "result": null
}

parent() -> String

Returns the parent directory path.

$ cxl eval -e 'emit result = "/data/reports/sales.csv".parent()'
{
  "result": "/data/reports"
}

parent_name() -> String

Returns just the name of the parent directory (not the full path).

$ cxl eval -e 'emit result = "/data/reports/sales.csv".parent_name()'
{
  "result": "reports"
}

Practical examples

Organize output by source directory:

emit source_dir = $pipeline.source_file.parent_name()
emit source_type = $pipeline.source_file.extension()

Extract file identifiers:

emit file_id = $pipeline.source_file.file_stem()
emit is_csv = $pipeline.source_file.extension() == "csv"

Route by file type:

let ext = input_path.extension()
emit format = match ext {
  "csv"  => "delimited",
  "json" => "structured",
  "xml"  => "markup",
  _      => "unknown"
}

Array Methods

CXL provides closure-bearing and non-closure array builtins for traversing and transforming nested arrays carried on a single record. The closure-bearing methods take an arrow-syntax closure and evaluate it once per element.

Null propagation

Every array method returns null when the receiver is null. The closure body is not invoked on a null receiver.

Closure-bearing methods

filter(it => Bool) -> Array

Returns a new array containing the elements for which the closure body evaluates to true.

- type: transform
  name: filter_items
  input: orders
  config:
    cxl: |
      emit kept = items.filter(it => it["price"] > 5)

For an input record where items is [{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}], kept is [{"sku":"a","price":10},{"sku":"b","price":20}].

map(it => T) -> Array

Returns a new array whose elements are the closure body’s value for each input element. The element type need not match the input element type.

    cxl: |
      emit skus = items.map(it => it["sku"])
      emit doubled_prices = items.map(it => it["price"] * 2)

skus is ["a", "b", "c"]; doubled_prices is [20, 40, 10].

find(it => Bool) -> Element | Null

Returns the first element for which the closure body evaluates to true. Returns null if no element matches.

    cxl: |
      emit first_premium = items.find(it => it["price"] > 15)

first_premium is {"sku":"b","price":20} for the running example.

any(it => Bool) -> Bool

Returns true if the closure body evaluates to true for at least one element. Returns false if no element matches (including on an empty array).

    cxl: |
      emit has_cheap = items.any(it => it["price"] < 10)

has_cheap is true.

flat_map(it => Array) -> Array

Like map, but the closure body returns an array per input element; the results are concatenated into a single flat array. A null body result contributes no elements; a non-array body result contributes a single element.

    cxl: |
      emit all_tags = items.flat_map(it => it["tags"])

For input items carrying tags arrays (e.g. [{"sku":"a","tags":["new"]},{"sku":"b","tags":["sale","new"]}]), all_tags is ["new","sale","new"].

Non-closure methods

remove(index: Int) -> Array

Returns a new array with the element at the given 0-based index removed. The original array is unchanged.

    cxl: |
      emit shifted = items.remove(1)

shifted is [{"sku":"a","price":10},{"sku":"c","price":5}] – index 0 is preserved, index 2 shifts down to index 1.

If the index is negative or out of range, remove returns the receiver array unchanged.

length() -> Int

Returns the number of elements in the array. length is also defined on strings (see String Methods).

    cxl: |
      emit item_count = items.length()

item_count is 3.

join(separator: String) -> String

Joins an array of values into a single string with the given separator between elements. Defined as a string method (see String Methods) but accepts array receivers.

    cxl: |
      emit sku_list = items.map(it => it["sku"]).join(", ")

sku_list is "a, b, c".

Bracket indexing vs .remove

Bracket indexing (items[0]) reads an element by position and returns null when out of range. .remove(idx) returns a new array with the element dropped; out-of-range indices leave the array unchanged. See Nested Paths for the index-access surface.

See also

  • Closures – the it => body form used by closure-bearing array methods.
  • Map Methods – builtins that operate on the map elements typically iterated by these array methods.
  • Nested Paths – bracket-index and dotted-path access through nested arrays and maps.
  • Emit Each – fan one input record into many output records, one per array element.

Map Methods

CXL provides six built-in methods for working with map values (key-value pairs). Maps arise naturally from JSON object inputs, from the set builder below, and from upstream emits that produce nested structures.

All map methods return new values – they never mutate the receiver. This is copy-on-write semantics: chaining .set then .remove_field produces a fresh map at each step, leaving the upstream binding untouched.

Null propagation

Every map method returns null when the receiver is null or is not a Value::Map.

Method reference

keys() -> Array

Returns the map’s keys as an array of strings, preserving insertion order.

- type: transform
  name: list_keys
  input: rows
  config:
    cxl: |
      emit field_names = profile.keys()

For an input record where profile is {"name":"Alice","tier":"gold","since":"2021-04"}, field_names is ["name","tier","since"].

values() -> Array

Returns the map’s values as an array, preserving insertion order. Value types are heterogeneous – the array carries each value as-is.

    cxl: |
      emit field_values = profile.values()

field_values is ["Alice","gold","2021-04"].

merge(other: Map) -> Map

Returns a new map containing every key from the receiver and from other. On conflicting keys, other’s value wins.

    cxl: |
      emit enriched = profile.merge(overrides)

For profile = {"name":"Alice","tier":"gold"} and overrides = {"tier":"platinum","since":"2021-04"}, enriched is {"name":"Alice","tier":"platinum","since":"2021-04"}.

set(key: String, value: Any) -> Map

Returns a new map with key set to value. If the key was already present, its value is replaced; insertion order is preserved.

    cxl: |
      emit stamped = profile.set("region", "us-east")

stamped is {"name":"Alice","tier":"gold","since":"2021-04","region":"us-east"}.

Nested paths

key may be a dotted/indexed path that descends into nested maps and arrays, so a single set writes into a deep document. Dots separate map keys; a [n] suffix indexes an array.

    cxl: |
      emit moved = profile.set("address.city", "NYC")
      emit relabel = order.set("items[0].sku", "A-100")
  • Auto-create. Missing intermediate map segments are created as empty maps, so a path can build structure that does not yet exist. {}.set("a.b.c", 7) returns {"a":{"b":{"c":7}}}. This is what lets set assemble a nested document from scratch (matching jq setpath and Bloblang assignment).
  • Type conflict -> null. If an intermediate segment already exists but is the wrong kind for the next step – descending into a key whose value is a scalar, indexing a map with [n], or naming a field on an array – the whole operation returns null. Nothing is partially written.
  • Array index past the end -> null. Indexing past the last element returns null for the whole operation; arrays are never silently grown. The path can only overwrite an array slot that already exists.
  • A bare key is a single key, not a path. "region" writes the top-level region. Only . and [n] introduce nesting; a key with neither behaves exactly as before.

For profile = {"name":"Alice","address":{"city":"LA"}}, profile.set("address.city", "NYC") is {"name":"Alice","address":{"city":"NYC"}} – the sibling name and any other address keys are preserved.

Known limitation. Because . and [ are path syntax, set cannot yet target a key whose name literally contains a . or [ (for example a JSON field literally named "a.b"). Column names solve this with a backslash escape — a\.b is one segment named a.b, per Field Paths — and the decided direction is for set / unset keys to read that same grammar, since both are flat strings addressing a path. Until they do: to write such a key, build it with merge and a map literal; to remove it, use remove_field, which matches the exact key string.

remove_field(key: String) -> Map

Returns a new map without key. If the key was absent, the receiver is returned unchanged.

    cxl: |
      emit slim = profile.remove_field("since")

slim is {"name":"Alice","tier":"gold"}.

unset(key: String) -> Map

Returns a new map with the entry addressed by key removed. unset is the deletion counterpart to set and reuses the same dotted/indexed path grammar: dots separate map keys, a [n] suffix indexes an array. A bare key (no . or [n]) drops a top-level entry, exactly like remove_field.

    cxl: |
      emit pruned = profile.unset("address.city")
      emit dropped = order.unset("items[0]")
  • Array element removes and shifts. unset("items[0]") deletes element 0 and shifts the remaining elements down, so the array shrinks by one (matching jq del). This is deliberately distinct from set("items[0]", null), which leaves a null hole — unset means delete.
  • Missing or conflicting path is a no-op. A path that does not resolve — a missing intermediate, a missing final key, an array index past the end, or a type conflict (a field segment against an array, an index segment against a map) — returns the receiver unchanged. This mirrors remove_field on an absent key, and is the opposite of set, which returns null on a conflicting path.
  • Copy-on-write. Like every map method, unset never mutates the receiver; the upstream binding is untouched.

For profile = {"name":"Alice","address":{"city":"LA","zip":"90001"}}, profile.unset("address.city") is {"name":"Alice","address":{"zip":"90001"}} – the sibling zip and the top-level name are preserved.

Worked example: chained set + remove_field

Map methods compose naturally because each returns a new map.

- type: transform
  name: rewrite_profile
  input: rows
  config:
    cxl: |
      emit profile =
        profile.set("region", "us-east").remove_field("internal_id")

For profile = {"name":"Alice","internal_id":"ix-77","tier":"gold"}, the emitted profile is {"name":"Alice","tier":"gold","region":"us-east"}. The internal_id slot is removed and the region slot is appended; both happen on a fresh map so the upstream record’s profile is unaffected for any other downstream branch.

Parentheses are required

All map methods are method calls and must be written with parentheses, even the zero-argument ones:

profile.keys()         -- ok
profile.keys           -- parses as a field lookup, not a method call

profile.keys parses as a dotted path – a lookup for a field literally named keys inside profile. That path almost certainly returns null. Always include the parentheses when invoking a map method.

Using map methods inside array closures

Map methods compose with closure-bearing array builtins when the array elements are themselves maps.

    cxl: |
      emit enriched_items = items.map(it => it.set("region", "us-east"))
      emit item_keys = items.map(it => it.keys())

Each it is a map; the closure body invokes a map method on it. enriched_items is an array where every element gained a region field. item_keys is an array of key-name arrays, one per element.

See also

  • Closures – arrow-syntax closures often invoke map methods on their it binding.
  • Array Methods – closure-bearing array methods commonly carry maps as their elements.
  • Nested Paths – bracket-index access (profile["name"]) reads a single key without producing a new map.

Window Functions

Window functions allow CXL expressions to access aggregated values across a set of records within an analytic window. Unlike aggregate functions (which collapse groups into single rows), window functions attach computed values to each individual record.

Window functions are accessed via the $window.* namespace and require an analytic_window: configuration on the transform node.

Interactive companion: the window functions explainer shows, for any row, which rows of its partition each function reads.

Configuring an analytic window

Window functions are only available in transform nodes that declare an analytic_window: section in YAML:

nodes:
  - name: ranked_sales
    type: transform
    input: raw_sales
    config:
      analytic_window:
        group_by: [region]
        sort_by:
          - field: amount
            order: desc
      cxl: |
        emit region = region
        emit amount = amount
        emit region_total = $window.sum(amount)
        emit running_total = $window.cumulative_sum(amount)
        emit rank_position = $window.row_number()

Window configuration fields

FieldDescription
group_byList of fields to partition the window by (the SQL PARTITION BY axis).
sort_byList of { field, order, null_order } ordering specifications: order is asc (default) or desc, and null_order is first or last (default last), placing null keys before or after every value. Values compare by the rule every sort uses; see How values are ordered. null_order: drop is rejected; see Nulls in sort_by.
sourceOptional explicit source-name reference for cross-source windows.
onOptional cross-source partition-lookup field.

Which rows a function reads

There is no frame option. A window function reads one of three things:

  • The whole partition. sum, avg, min, max, count, first_value, last_value, first(), last(), any, every, exists, not_exists, collect and distinct read every row of the record’s partition, whatever the record’s position. Every record in a partition gets the same $window.sum(amount).
  • The partition up to the current record. cumulative_sum is the running total, from the partition’s first record (in sort_by order) through the current one.
  • A position. row_number, rank and dense_rank give the current record’s place in sort_by order; lag(n) and lead(n) read the record n places before or after it.

Nulls in sort_by

sort_by only orders the rows of a partition; every row of the partition is still seen by the window functions and written by the Transform. A row whose key is null is placed by null_order: before every value with first, after every value with last, in either direction.

null_order: drop is rejected when the pipeline is planned:

transform "running": `null_order: drop` is not allowed on `analytic_window.sort_by` for field "amount": `sort_by` only orders the rows of a window partition, placing nulls `first` or `last`, and cannot remove a row. To remove the rows whose "amount" is null, delete `null_order: drop` and add a Transform before this node with `config: { cxl: "filter not amount.is_null()" }`.

The one fix is the filter the error prints: delete null_order: drop and add a Transform before the windowed one whose whole config is the printed line. That leaves rows with a null key out of every partition, and also removes them from the windowed Transform’s output:

- type: transform
  name: with_amount
  input: orders
  config: { cxl: "filter not amount.is_null()" }

For a field CXL cannot write as a bare name, the error prints a source_name: line instead; see source_name.

Aggregate window functions

These compute values over the record’s whole partition, except cumulative_sum, which stops at the current record.

$window.sum(field)

Sum of the field values across the whole partition. Null and non-numeric values are skipped; a sum of integers returns a Float.

emit running_total = $window.sum(amount)

$window.cumulative_sum(field)

Running total of the field values from the partition’s first record (in sort_by order) through the current record. Like $window.sum, a sum of integers returns a Float.

emit running_total = $window.cumulative_sum(amount)

$window.avg(field)

Average of the field values across the whole partition. Returns Float.

emit moving_avg = $window.avg(amount)

$window.min(field)

Minimum value in the partition.

emit window_min = $window.min(amount)

$window.max(field)

Maximum value in the partition.

emit window_max = $window.max(amount)

$window.count()

Number of records in the partition, the same for every record in it. Takes no arguments. For a record’s position, use $window.row_number().

emit window_size = $window.count()

$window.first_value(field)

Returns the value of field at the first record of the partition (ordered by sort_by). Equivalent to SQL FIRST_VALUE(field).

emit opening_amount = $window.first_value(amount)

$window.last_value(field)

Returns the value of field at the last record of the partition (ordered by sort_by), the same for every record in it.

emit closing_amount = $window.last_value(amount)

Ranking window functions

Zero-argument integer functions that return the current row’s rank within its partition.

$window.row_number()

1-indexed position of the current record within its partition.

emit row_idx = $window.row_number()

$window.rank()

SQL RANK(): rows that share the same sort_by tuple receive the same rank, and the next distinct row jumps by the size of the tie group.

emit sales_rank = $window.rank()

$window.dense_rank()

SQL DENSE_RANK(): ties share a rank with no gaps between distinct ranks.

emit sales_dense_rank = $window.dense_rank()

Positional window functions

These return a whole record by position within the partition. Name the field to read after the call, as in $window.lag(1).amount. Without a field name the call returns null on every record, and no error is raised.

$window.first().field

The first record of the partition, in sort_by order.

emit first_amount = $window.first().amount

$window.last().field

The last record of the partition, in sort_by order.

emit last_amount = $window.last().amount

$window.lag(n).field

The record n places before the current record. Returns null if there is no record at that offset.

emit prev_amount = $window.lag(1).amount
emit two_back = $window.lag(2).amount

$window.lead(n).field

The record n places after the current record. Returns null if there is no record at that offset.

emit next_amount = $window.lead(1).amount

Iterable window functions

These evaluate predicates or collect values across the window.

$window.any(predicate)

Returns true if the predicate is true for any record in the window.

emit has_high = $window.any(amount > 1000)

$window.every(predicate)

Returns true if the predicate is true for every record in the window.

emit all_positive = $window.every(amount > 0)

$window.exists(predicate)

Returns true if the predicate is true for at least one record in the window — a SQL-fluency alias of $window.any.

emit any_high = $window.exists(amount > 1000)

$window.not_exists(predicate)

Returns true if no record in the window satisfies the predicate. Equivalent to not $window.exists(predicate) and to $window.every(not predicate).

emit none_negative = $window.not_exists(amount < 0)

$window.collect(field)

Collects all values of the field in the window into an array.

emit all_amounts = $window.collect(amount)

$window.distinct(field)

Collects distinct values of the field in the window into an array.

emit unique_regions = $window.distinct(region)

$window.collect and $window.distinct emit arrays. JSON writes them as native arrays, XML as repeated child elements, and CSV as a delimited cell. Before a scalar-only sink such as fixed-width, coerce the value in a downstream Transform (for example emit regions = unique_regions.join(";")).

Complete example

nodes:
  - name: sales_analysis
    type: transform
    input: daily_sales
    config:
      analytic_window:
        group_by: [store_id]
        sort_by:
          - field: sale_date
            order: asc
      cxl: |
        emit store_id = store_id
        emit sale_date = sale_date
        emit daily_revenue = revenue
        emit store_avg = $window.avg(revenue)
        emit store_total = $window.sum(revenue)
        emit revenue_to_date = $window.cumulative_sum(revenue)
        emit prev_day_revenue = $window.lag(1).revenue
        emit day_over_day = revenue - ($window.lag(1).revenue ?? revenue)

For each sale this adds the store’s average and total over all of its days, its revenue to date, and the change from the previous day.

Correlation-key error handling

Window functions work correctly when a pipeline uses correlation keys for group-atomic error handling: if records are retracted from an upstream group, the window recomputes the affected partitions so its output stays consistent. There is nothing to configure.

Aggregate Functions

Aggregate functions operate across grouped record sets in aggregate nodes, collapsing multiple input records into summary rows. They are distinct from window functions, which attach computed values to each individual record.

Aggregate functions

CXL provides 7 aggregate functions. These are called as free-standing function calls (not method calls) within the CXL block of an aggregate node.

FunctionSignatureReturnsDescription
sum(expr)NumericInt, Float or Decimal (the input’s type)Sum of values
count(*)–IntCount of records in the group
avg(expr)NumericFloat, or Decimal for a decimal inputArithmetic mean
min(expr)AnyAnyMinimum value
max(expr)AnyAnyMaximum value
collect(expr)AnyArrayAll values collected into an array
weighted_avg(value, weight)Numeric, NumericFloat / DecimalWeighted arithmetic mean

YAML aggregate node

Aggregate functions are used inside the cxl: block of a node with type: aggregate. The node must declare group_by: fields.

nodes:
  - name: dept_summary
    type: aggregate
    input: employees
    config:
      group_by: [department]
      cxl: |
        emit total_salary = sum(salary)
        emit headcount = count(*)
        emit avg_salary = avg(salary)
        emit max_salary = max(salary)
        emit min_salary = min(salary)

Group-by fields pass through automatically

Fields listed in group_by: are automatically included in the output. You do NOT need to emit them – they are carried through as group keys.

In the example above, department is automatically present in every output record without an explicit emit department = department statement.

Function details

sum(expr) -> Int, Float or Decimal

Computes the sum of the expression across all records in the group. Null values are skipped.

cxl: |
  emit total_revenue = sum(price * quantity)

The result has the type of the values summed: integers give an integer, floats a float, decimals a decimal. Integers summed with floats give a float, and integers summed with decimals a decimal. An integer sum outside the 64-bit integer range is an error.

A float sum is the exact total of the group’s values, rounded once to the nearest float. It does not depend on the order rows arrive in or on memory.limit: the same group gives the same bytes whether the Aggregate holds every group in memory or spills and merges partial sums. A group that holds 1e16, 1.0 and -1e16 sums to 1, in any order, where adding the floats left to right gives 0 or 1 depending on the order. Integers mixed with floats are added exactly too, so an integer larger than 2^53 is not rounded before it is added. A NaN in the group makes the sum NaN, and +inf with -inf makes it NaN. Results can differ in the last bit from a version that rounded after each addition.

A decimal sum is the exact total of the group’s values, rounded once (half to even) only when it does not fit a decimal at its scale. Its scale is the largest scale among the group’s values, zeros and integers included, so the sum of 1.00, -1.00 and 2 is 2.00 whatever order the rows arrive in. It is an error only when the whole group’s exact total is outside the decimal range, ±79,228,162,514,264,337,593,543,950,335; a group whose running total passes outside the range and comes back is fine. The error’s fix aggregates the column as floats, sum(amount.to_float()) with your column in place of amount, when a binary float’s range and precision will do.

A group whose values are all null gives null. Null is never a substitute for a failure: a group that fails is an aggregate_finalize error (see Error categories), which under strategy: continue goes to the dead-letter output.

Decimal and float in one group

A decimal is never added to a float without an explicit conversion, in an aggregate as in amount + price. A sum, avg or weighted_avg whose values in one group include both a decimal and a float fails that group with:

decimal and float in one group: a decimal is never added to a float without an explicit conversion; declare the column that holds the floats `type: decimal` in its Source schema, so every value in the group is a decimal

When the typechecker can see the mix, for example sum(if flag then amount else price), the pipeline does not compile (E200; see Conditionals). The run-time error covers what it cannot see: a value whose type is only known at run time, such as an untyped column or a numeric result like amount.clamp(0, 100). Declaring the float column type: decimal in its Source schema keeps the total exact: the reader parses the column’s text as a decimal, so every value in the group is a decimal. (A JSON number read into a decimal column is still parsed through a float first; see #1299.) When the floats are computed upstream rather than read from a Source column, there is no column to retype: convert the decimal values with .to_float() instead, accepting binary float precision.

count(*) -> Int

Counts the number of records in the group. The argument is the wildcard *.

cxl: |
  emit num_orders = count(*)

avg(expr) -> Float or Decimal

Computes the arithmetic mean. Null values are skipped.

cxl: |
  emit avg_order_value = avg(order_total)

avg(x) is sum(x) / count(x): the group’s exact sum, rounded once as sum rounds it, divided by the number of non-null values, so it is exact over floats and decimals alike and does not depend on row order or memory.limit. Over decimals the result is a decimal, the quotient at full precision, so avg(amount) and sum(amount) / count(amount) give the same digits and the same scale. Over floats, and over integers mixed with floats, the result is a float. Over integers alone it is a float: the exact integer total, converted once to a float, divided by the count.

A decimal total outside the decimal range is an error, as for sum, and so is a group mixing decimals and floats. A group whose values are all null gives null.

min(expr) -> Any

Returns the minimum value in the group. Works on numeric, string, and date types. Null values are skipped; a group whose values are all null gives null.

cxl: |
  emit earliest_order = min(order_date)
  emit lowest_price = min(unit_price)

Values compare by the rule sorting uses (see How values are ordered):

  • Integers, floats and decimals compare by their exact value, so a column that holds both integers and floats (for example one built by an if whose branches return an integer and a float) is compared value by value.
  • NaN is the largest value, above inf.
  • Strings, dates and datetimes order as they do in a Sink sort.

The result does not depend on the order rows arrive in, or on the memory limit.

When several values in the group are equal under that rule, min returns the same one every time: an integer before a decimal before a float, the decimal with fewer fractional digits, and the float with a negative sign before one with a positive sign. So min of 1 and 1.0 is 1, min of the decimals 1.0 and 1.00 is 1.0, and min of -0.0 and 0.0 is -0.0.

max(expr) -> Any

Returns the maximum value in the group. Works on numeric, string, and date types. Null values are skipped; a group whose values are all null gives null.

cxl: |
  emit latest_order = max(order_date)
  emit highest_price = max(unit_price)

Values compare as they do for min: numbers by their exact value across integer, float and decimal, and NaN is the largest value, so a group that holds a NaN has NaN as its maximum. The result does not depend on the order rows arrive in, or on the memory limit.

When several values in the group are equal, max picks in the reverse order to min: a float before a decimal before an integer, the decimal with more fractional digits, and the float with a positive sign. So max of 1 and 1.0 is 1.0, max of the decimals 1.0 and 1.00 is 1.00, and max of -0.0 and 0.0 is 0.0.

collect(expr) -> Array

Collects all values of the expression into an array. Useful for building lists of values per group.

cxl: |
  emit all_order_ids = collect(order_id)

Because collect emits an array, JSON writes it as a native array, XML as repeated child elements, and CSV as a delimited cell. Coerce it to a scalar for a format such as fixed-width (for example emit ids = all_order_ids.join(";")).

weighted_avg(value, weight) -> Float or Decimal

Computes a weighted average: sum(value * weight) / sum(weight). Takes two arguments.

cxl: |
  emit weighted_price = weighted_avg(unit_price, quantity)

weighted_avg(v, w) is sum(v * w) / sum(w), with each row’s v * w computed as it is in any expression and both sums exact, rounded once as sum rounds them. When either the value or the weight is a decimal, the result is a decimal at full division precision, and it has the same digits and scale as sum(v * w) / sum(w). Otherwise it is a float. Over floats each row’s v * w is a float product, and the products and the weights are summed exactly, so the average does not depend on row order or memory.limit. Over integers alone the two exact totals are each converted once to a float and divided.

These groups fail with an aggregate_finalize error rather than writing a value:

  • Zero total weight. The group’s weights add up to exactly zero, so the average divides by zero, as x / 0 does in any expression. Rows whose weight is zero add nothing to the average, so the error’s fix drops them with a Transform before the Aggregate, config: { cxl: "filter qty != 0" } with your weight column in place of qty. A group whose non-zero weights cancel, such as a sale and its return, still totals zero after that filter and still fails.
  • A row’s product out of range. A row’s decimal value * weight is outside the decimal range. Retracting that row clears the error. The error’s fix computes the average in floats, weighted_avg(price.to_float(), qty.to_float()) with your columns in place of price and qty.
  • A total or the quotient out of range. A decimal total is outside the decimal range, or the quotient is because the weights nearly cancel.
  • Decimal and float in one group, in one row or across rows (see above).

Mixing a decimal with a binary float across the two arguments is a type error when the typechecker can see it. Declare the float column type: decimal in its Source schema so both arguments are decimals, or, when the float is computed rather than read from a Source column, convert the decimal argument with .to_float(). A group with no row whose value and weight are both non-null gives null.

Aggregates vs. windows

FeatureAggregate nodeWindow function
Record outputOne row per groupOne row per input record
Syntaxsum(field) (free-standing)$window.sum(field) (namespace)
Configurationtype: aggregate + group_by:type: transform + analytic_window:
Use caseSummarize groupsEnrich records with group context

An Aggregate’s sum, avg and weighted_avg are exact (see sum). A window function’s $window.sum and $window.avg are not: they still add in the order the rows of the partition are held, so a window sum over floats can differ in its last bits from the Aggregate’s sum of the same values.

Combining aggregates with expressions

Aggregate function calls can be mixed with regular CXL expressions in emit statements:

nodes:
  - name: category_stats
    type: aggregate
    input: products
    config:
      group_by: [category]
      cxl: |
        emit total_revenue = sum(price * quantity)
        emit avg_price = avg(price)
        emit margin_pct = (sum(revenue) - sum(cost)) / sum(revenue) * 100
        emit product_count = count(*)
        emit has_premium = max(price) > 100

Restrictions

  • let bindings in aggregate transforms are restricted to row-pure expressions (no aggregate function calls in let).
  • filter in aggregate transforms runs pre-aggregation – it filters input records before grouping.
  • distinct is not permitted inside aggregate transforms. Place a separate distinct transform upstream.

Complete example

pipeline:
  name: sales_summary
  nodes:
    - name: raw_sales
      type: source
      format: csv
      path: sales.csv

    - name: monthly_summary
      type: aggregate
      input: raw_sales
      group_by: [region, month]
      cxl: |
        emit total_sales = sum(amount)
        emit order_count = count(*)
        emit avg_order = avg(amount)
        emit top_sale = max(amount)
        emit all_reps = collect(sales_rep)

    - name: output
      type: sink
      input: monthly_summary
      format: json
      path: summary.json

This pipeline outputs JSON because all_reps = collect(sales_rep) emits an array, which the tabular writers (CSV/XML/fixed-width) reject; drop the collect binding or coerce it with a downstream Transform to keep a CSV sink.

Closures

CXL supports arrow-syntax closures as arguments to closure-bearing array builtins like filter, map, find, any, and flat_map. They give CXL a way to express element-by-element predicates and projections over nested arrays carried inside a single record – without writing a separate transform node per element.

Syntax

it => expression

A closure has one parameter, named it, and a single expression body. The arrow => separates them.

- type: transform
  name: filter_items
  input: orders
  config:
    cxl: |
      emit kept = items.filter(it => it["price"] > 5)

The body is an expression, not a block of statements. Use if/then/else or match if you need branching inside a closure.

    cxl: |
      emit price_buckets = items.map(it =>
        if it["price"] >= 100 then "premium"
        else if it["price"] >= 10 then "standard"
        else "value")

Parameter name

The parameter is always it. Other identifiers are not accepted as the closure binding:

items.filter(item => item["price"] > 5)   -- parse error
items.filter(it => it["price"] > 5)       -- ok

it is recognized in expression position only inside a closure body. Outside of one, it has no special meaning.

Lexical capture

Inside the closure body, the outer record’s fields and let bindings remain visible. For each iteration the closure parameter it is bound to the current element, the body evaluates, then it is removed before the next iteration.

    cxl: |
      let threshold = 10
      emit kept = items.filter(it => it["price"] > threshold)

Here the closure body reads both it (the current array element) and threshold (an outer let binding). The record’s fields are also reachable by name – a closure over items can still read customer_id, region, or any other field on the same record.

Where closures appear

Closures are valid only as method-call arguments to closure-bearing builtins. They cannot be assigned to variables, stored in fields, or passed to non-closure builtins:

let f = it => it * 2          -- rejected at resolve time
emit doubler = it => it * 2   -- rejected at resolve time

If you need to share a closure across multiple call sites, repeat the literal closure expression. CXL has no first-class function values.

Null propagation

Closure-bearing builtins applied to a null receiver return null without evaluating the body. The body is also never called on records where the array is null:

    cxl: |
      emit kept = items.filter(it => it["price"] > 5)
      -- when `items` is null, `kept` is null; the body never runs

This matches the null-propagation policy on every other builtin – see Null Handling for the wider rules.

Worked example: filter and map over a nested array

Suppose each input record carries an items array of objects, each with sku and price:

{"order_id":"O-1","items":[{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}]}

A transform that drops cheap items and projects the remaining SKUs:

- type: transform
  name: filter_items
  input: orders
  config:
    cxl: |
      emit order_id = order_id
      emit kept = items.filter(it => it["price"] > 5)
      emit kept_skus = items.filter(it => it["price"] > 5).map(it => it["sku"])

For the input above, the transform produces:

{
  "order_id": "O-1",
  "kept": [{"sku": "a", "price": 10}, {"sku": "b", "price": 20}],
  "kept_skus": ["a", "b"]
}

Bracket-index access (it["price"]) reaches into each map element. See Nested Paths for the full traversal surface.

See also

  • Array Methods – the closure-bearing builtins (filter, map, find, any, flat_map).
  • Map Methods – callable on map elements inside a closure body.
  • Nested Paths – bracket-index and dotted-path navigation through nested arrays and maps.
  • Emit Each – statement that fans one input record into many output records, using a binding similar to the closure parameter.

Nested Paths

CXL records can carry nested arrays and maps as field values (for example, a JSON input where each record has an items array of objects). Reaching into that structure uses two complementary forms: dotted paths and bracket indices.

Dotted paths

A dotted identifier path reads a static field name from a map.

doc.metadata.tenant

Each segment must be a valid identifier. Dotted paths are resolved at compile time – the typechecker walks the structure declared in the source schema and reports a missing-field error if any segment doesn’t exist.

- type: transform
  name: project_tenant
  input: events
  config:
    cxl: |
      emit tenant = doc.metadata.tenant
      emit user_id = doc.user.id

Use dotted paths for structures whose shape is fixed and known at authoring time.

Bracket indices

A bracket index reads a runtime-computed key. The receiver may be an array (integer index) or a map (string index).

items[0]
profile["name"]
items.map(it => it["sku"])

Bracket indices are dynamic – the index expression evaluates per record. The typechecker treats the result as Any and does not assert that the key is present.

Integer index on an array

- type: transform
  name: first_item
  input: orders
  config:
    cxl: |
      emit head = items[0]
      emit second = items[1]

For items = [{"sku":"a"},{"sku":"b"},{"sku":"c"}], head is {"sku":"a"} and second is {"sku":"b"}.

Out-of-range indices return null. Negative indices also return null (CXL does not support negative indexing).

String index on a map

    cxl: |
      emit name = profile["name"]
      emit tier = profile["tier"]

Missing keys return null – the lookup never raises an error. This is the same null-propagation policy closure builtins use on their receivers.

Mixing forms

The two forms compose in either order:

    cxl: |
      emit first_sku = items[0]["sku"]
      emit profile_email = users.profile["email"]

items[0]["sku"] is two bracket indices chained – an integer index against the array, then a string index against the resulting map. users.profile["email"] walks a dotted path to reach profile (a map field on users), then bracket-indexes into it for a runtime key.

Null propagation

Every nested-access form propagates null end-to-end. If the receiver is null, the result is null without evaluating the index expression:

    cxl: |
      emit sku = items[0]["sku"]
      -- when `items` is null, `sku` is null
      -- when `items[0]` is null, `sku` is also null

This matches the null behavior on dotted paths and on method-call receivers. Records with missing intermediate structure produce nulls in their derived fields rather than aborting the transform.

Method calls on indexed values

A bracket-indexed expression is a regular value, so it composes with any method or further index:

    cxl: |
      emit head_sku_upper = items[0]["sku"].upper()
      emit cheap_skus = items.filter(it => it["price"] < 10).map(it => it["sku"])

The first chain reads a string out of nested structure and uppercases it. The second filters an array of maps by a numeric field and projects the SKU strings out.

See also

  • Field Paths – the other surface for the same idea: how a flat column-name string spells a path, with a backslash escape instead of brackets.
  • Closures – closures over arrays of maps typically use bracket-index on the it binding.
  • Array Methods – traversal builtins that consume nested arrays.
  • Map Methods – builders and accessors for map values.
  • Null Handling – the wider null-propagation rules.

Field Paths

A column name is not opaque. An unescaped . inside it separates path segments, so the column Address.City addresses the path Address → City. Readers produce such names when they flatten nested input, writers expand them back into nested output, and the rule for reading them is the same everywhere in Clinker — one grammar, not one per format.

This page is the reference for that grammar. The places it applies:

  • Column names declared in a source or output schema:.
  • Column names a self-describing reader infers (the JSON reader flattens {"a":{"b":1}} to the column a.b; the XML reader flattens <a><b>1</b></a> the same way).
  • Column names a writer expands back into nesting — see Writing JSON and Writing XML.

It does not cover CXL expressions. Reaching into a value — a map or array held in one column — uses the expression forms on Nested Paths (profile.city, profile["a.b"], items[0]). The two are different surfaces over the same idea: a path is an ordered list of segments either way, but a column name spells it with . and a backslash escape, while an expression spells it with dots and brackets.

The rule

Reading a name left to right:

InputMeaning
.Ends the current segment, starts the next.
\.A literal . inside the current segment.
\\A literal \.
\[A literal [.
\ before anything elseAn error. \], \t, and a name ending in \ are all rejected.
Any other characterA literal character of the current segment — including a bare [, and ], @, $, /, and whitespace.

So Address.City is two segments, and a\.b is one segment named a.b.

Three consequences worth stating outright:

  • An unrecognized escape is an error, never a literal. A \ must be followed by ., [, or \ — nothing else, and not the end of the name. C:\temp is rejected, because silently treating \t as the two characters \ and t would make the encoding ambiguous, and silently dropping the \ would rename the column without saying so. Write C:\\temp. The same applies to \], which is covered below.
  • Empty segments are real. a..b is three segments — a, the empty name, and b. An empty key is a value a document can genuinely carry, so the grammar does not reject it.
  • A name may nest at most 64 levels deep. This matches the depth at which the JSON reader stops flattening, so any name that reader produces is a name a writer can expand. The XML reader flattens with no depth bound, so an extraordinarily deep XML document can produce a name past the cap; writing it back out fails with a clear error rather than recursing without limit.

Writing a literal bracket

[ is currently a literal character, so a column named a[0] works. It is nonetheless reserved: bracket indexing may later be given meaning inside a flat name, matching the [n] form CXL expressions already use. Writing \[ means “a literal [” today and will keep meaning exactly that. A bare [ is not guaranteed to.

So the column a[0] future-proofs as a\[0] — escaping only the opening bracket:

SpellingResult
a\[0]One segment named a[0]. Correct, and stable across the reserved-[ change.
a\[0\]Rejected — \] is not an escape.
a[0]One segment named a[0] today, but a bare [ is the form that is not guaranteed to keep that meaning.

Only the opening bracket is ever escaped. ] is never escaped and never needs to be: it would carry meaning only as the close of an unescaped [, so a literal ] standing on its own is already unambiguous. Escaping it would add noise without removing any ambiguity.

Writing \] is an error rather than a silently-accepted no-op, for the same reason C:\temp is: an escape that quietly meant nothing would let two different names decode to the same path, and one of them would be dropped. The error names the offending escape and the rule.

If a column name of yours contains [, escaping it now costs nothing and makes it future-proof.

Expansion on write

A writer that can express nesting rebuilds it from the decoded paths.

Grouping. Columns sharing a prefix collect into one container, positioned where that prefix first appeared — even when the schema interleaves them. Columns Address.City, name, Address.State write as:

{"Address":{"City":"Boston","State":"MA"},"name":"Ada"}

Absent children. A column omitted for a record (a null under preserve_nulls: false) contributes nothing, and a container whose every descendant is omitted emits no key at all rather than an empty object. That keeps the round trip honest: reading {"a":{}} produces no column, so writing no column must produce no "a".

Values are untouched. Expansion adds structure above a column’s value; a Value::Map or array held in that column still serializes as itself. A column Items.Item holding [1,2] writes as {"Items":{"Item":[1,2]}}.

No carve-outs. The rule reads the column-name string and nothing else, so it applies identically to engine-stamped columns. Under include_correlation_keys: true the column $ck.customer_id writes as {"$ck":{"customer_id":…}}.

When two names clash

Two columns can describe places that cannot both exist. The writer refuses the whole column set before writing a single byte, naming both columns — it never silently keeps one.

ColumnsWhy
a and a.ba holds a value and is also the container a.b sits inside.
a.b and a.b.cThe same clash, one level down.
a[b and a\[bTwo spellings of the identical path.

a.b and a\.b do not clash: they address a → b and the single segment a.b. That is what the escape is for.

When a clash comes from a . you meant literally, the diagnostic offers the escaped spelling:

XML writer cannot expand this output's column names into nested output: field
names `a.b` and `a.b.c` cannot both be written: `a.b` holds a value and is also
the container `a.b.c` nests inside. Rename one of them, or — if the `.` in `a.b`
is part of the name rather than a nesting separator — declare it as `a\.b`.

Where the grammar does not reach yet

Two surfaces read column names without this grammar. Both are tracked, and both are stated here so the current behavior is a known position rather than a surprise:

  • Flat writers emit the raw name. CSV and fixed-width have no nesting to expand into, so they write the column name verbatim — a column declared a\.b appears as the literal CSV header a\.b, backslash included. Whether flat writers should emit the decoded name instead is a separate decision.
  • Readers join without escaping. The JSON and XML readers join flattened path segments with a plain ., so a source key that literally contains a . ({"a.b": 1}) arrives as the column a.b — indistinguishable from a nested {"a":{"b":1}}, and it writes back nested. Closing that means escaping each key as the reader joins it, tracked by issue 920.

set and unset in CXL also use a path grammar of their own that does not yet support escaping — see the known limitation on Map Methods.

See also

  • Nested Paths — the expression-side forms for reaching into a value.
  • Writing JSON — expansion on the JSON write side.
  • Writing XML — expansion on the XML write side, plus the attribute convention layered over it.

Emit Each

The emit each statement fans one input record into multiple output records – one per element of an array on the input. The body emits the fields each output record carries. A trailing outer modifier preserves the trigger row when the array is empty or null.

Syntax

emit each <binding> in <source> {
  <statements>
}
  • <binding> is the identifier the body uses to refer to the current array element. The conventional name is it (same as the closure parameter), but any identifier is accepted.
  • <source> is any expression producing an array. Typically a field reference on the input record.
  • The body is a block of let and emit statements that produce one output record per iteration.

Worked example

Suppose each input record carries an items array of objects, each with sku and price:

{"order_id":"O-1","items":[{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}]}

A transform that fans each input into one record per item:

- type: transform
  name: explode
  input: orders
  config:
    cxl: |
      emit each it in items {
        emit order_id = order_id
        emit sku = it["sku"]
        emit price = it["price"]
      }

For the input above, the transform produces three output records:

{"order_id":"O-1","sku":"a","price":10}
{"order_id":"O-1","sku":"b","price":20}
{"order_id":"O-1","sku":"c","price":5}

The body reads both it (the current element) and order_id (an outer record field). Outer-record fields remain visible inside the body for every iteration.

Cardinality

If the source array has N elements, emit each produces exactly N output records. Empty array sources produce zero records. A null source also produces zero records – no DLQ entry, no error – mirroring the explode-on-null convention used elsewhere in CXL.

Known issue: a top-level plain emit each over an empty array or null currently passes the input row through as one record instead of producing zero records (#1350). This comes from reading the engine’s code; it has not been confirmed with a pipeline run. Until it is fixed, put filter not items.is_empty() (with your own field for items) before the block if you rely on the row being dropped: is_empty() is true for both an empty array and null.

When fan-out nests, the cardinalities multiply: an outer array of M elements whose inner arrays have N elements each produces up to M×N records. The cumulative max_expansion cap bounds that product.

A non-array, non-null source raises a runtime type-mismatch error and routes the originating record to the DLQ.

Preserving the trigger row: outer

A trailing outer modifier switches emit each to its outer-join variant. The grammar is identical except for the keyword after the source:

emit each <binding> in <source> outer {
  <statements>
}

The only behavioral difference is what happens when the source is null or an empty array. Plain emit each drops the trigger row entirely (zero output records). The outer variant instead emits the trigger row once, with <binding> bound to null:

Sourceemit each ...emit each ... outer
3-element3 records3 records (identical)
empty array0 records (see the known issue above)1 record, binding = null
null0 records (see the known issue above)1 record, binding = null

This is the shape SQL engines spell LATERAL VIEW OUTER EXPLODE (Spark, Hive) or an outer UNNEST (DuckDB): “for each tag on this article emit a tagged row, but keep articles that have no tags.”

Using the worked example above with an order that carries no items:

{"order_id":"O-2","items":[]}
- type: transform
  name: explode_outer
  input: orders
  config:
    cxl: |
      emit each it in items outer {
        emit order_id = order_id
        emit sku = it["sku"]
        emit price = it["price"]
      }

produces a single record that keeps order_id while the per-item fields read through the null binding:

{"order_id":"O-2","sku":null,"price":null}

Outer-record fields (like order_id) and any emit statements preceding the block still apply to the preserved trigger row, so an outer row is never bare.

The source type rule is slightly wider than plain emit each: a statically-null source is accepted (it is the case the variant exists to handle), alongside arrays and Any. Everything else in this page — the cumulative max_expansion cap, the nesting rules, the body-statement restrictions — applies unchanged to the outer variant. The two variants compose freely: an outer block may nest inside a plain emit each block and vice versa.

Output schema

The body’s emit statements define the output record’s field set, the same way emit does in a regular transform body. Fields the body does not emit fall under the Sink node’s include_unmapped policy (see Sink Nodes).

Fields written by the body shadow same-named fields on the originating input record.

Nested fan-out: fan-out within fan-out

An emit each body may itself contain emit each blocks — fan-out within fan-out for one trigger row. This is the canonical “for each article, for each section, for each tag, emit a row” shape:

emit each section in article["sections"] {
  emit each tag in section["tags"] {
    emit article_id = article_id
    emit section = section["name"]
    emit tag = tag
  }
}

For one input article, this produces one output record per (section, tag) pair. The inner binding (tag) reads the current inner element; the outer binding (section) and any outer-record field (article_id) stay visible inside the inner body. A field name reused as both an outer and inner binding shadows lexically — the inner binding wins inside the inner body, and the outer value is restored when the inner block finishes.

Emits are positional: an emit placed in the outer body before a nested block applies to every leaf record that block produces, but an emit placed after a nested block does not retroactively reach the records that block already emitted. Put the fields shared across leaves above the nested block.

Plain and outer blocks compose in any order. An inner plain emit each over an empty or null array contributes no records for that branch, while an inner emit each ... outer preserves one trigger row (inner binding bound to null) — exactly the per-level semantics from the single-level table, applied at each level.

Nesting is bounded to 32 levels so that adversarially deep input cannot exhaust the parser stack; legitimate document fan-out is only a few levels deep. Beyond that bound, parsing fails with a “nesting too deep” diagnostic.

The flat-array workaround (precompute a flattened array with .flat_map and use a single emit each) is still available and may be clearer for a simple two-level cartesian product, but is no longer required.

Body-statement restrictions

Within the body, let, emit, trace, and nested emit each / emit each ... outer are accepted. filter and distinct are rejected at evaluation time – a body filter would split work between branches the engine can’t represent. Move filter/distinct logic into a downstream transform, or pre-filter the source array with .filter before the emit each block.

Safety cap: max_expansion

By default, emit each can fan one input record into at most 10,000 output records. This limit is counted cumulatively across all nesting levels, so nested fan-out cannot multiply past it. A record that exceeds the limit routes to the DLQ with category expansion_limit_exceeded instead of producing an unbounded result. Override the limit with the max_expansion field in the transform config.

See Transform Nodes -> Expansion Cap for the YAML field and tuning guidance.

See also

System Variables

CXL provides several system variable namespaces prefixed with $. These give CXL expressions access to pipeline execution context, user-defined variables, per-record metadata, and the current time.

$pipeline.* – Pipeline context

Pipeline variables are accessed via $pipeline.member_name. Some are frozen at pipeline start; others update per record.

Stable (frozen at pipeline start)

VariableTypeDescription
$pipeline.nameStringPipeline name from YAML config
$pipeline.execution_idStringUUID v7, unique per pipeline run
$pipeline.batch_idStringFrom --batch-id CLI flag, or auto-generated UUID v7
$pipeline.start_timeDateTimeFrozen at pipeline start, deterministic within a run
$ cxl eval -e 'emit name = $pipeline.name' \
    -e 'emit exec = $pipeline.execution_id'
{
  "name": "cxl-eval",
  "exec": "00000000-0000-0000-0000-000000000000"
}

Counters

These members exist, but nothing updates them while a pipeline runs, so each one reads 0 in every expression (#1346). Do not use them to count progress: a condition such as $pipeline.total_count % 10000 == 0 is true for every record. The run’s real counts are reported when it ends, on the summary line and in the metrics file; see Where did my rows go? and Metrics & Monitoring.

VariableTypeNames the count of
$pipeline.total_countIntRecords read
$pipeline.ok_countIntRecords that reached an output
$pipeline.dlq_countIntRecords sent to the dead-letter queue
$pipeline.filtered_countIntRecords excluded by filter statements
$pipeline.distinct_countIntRecords excluded by distinct statements

$source.* – Per-record source lineage

$source.* exposes engine-stamped columns that travel with every record from its origin Source node downstream through merges, combines, and transforms. They identify where the record came from and when in event-time it happened. All three columns are filtered out of default Output projections — reference them explicitly with emit if you need them in your output schema.

VariableTypeDescription
$source.fileStringPath of the input file the current record was read from.
$source.nameStringName of the Source node that produced the current record. Survives through merge / combine so downstream nodes can branch on origin.
$source.event_timeDateTimeEngine-stamped event time, delay-corrected by the source’s watermark.delay. Null when the source has no watermark: block, or when the per-record value did not parse.
filter $source.name == "src_web"
emit origin = $source.name
emit ingest_file = $source.file
emit ts = $source.event_time

$source.event_time is the column a time-windowed aggregate reads to assign records to windows. It is only populated for records from a source that declares watermark: — otherwise it holds Null.

$vars.* – User-defined variables

User-defined variables are declared in the YAML pipeline config under pipeline.vars: and accessed via $vars.name in CXL expressions.

YAML declaration

pipeline:
  name: invoice_processing
  vars:
    high_value_threshold: 10000
    tax_rate: 0.21
    output_currency: "USD"
    fiscal_year_start_month: 4

CXL usage

filter amount > $vars.high_value_threshold
emit tax = amount * $vars.tax_rate
emit currency = $vars.output_currency

Variables provide a clean way to externalize configuration from CXL logic. Combined with channels, different variable sets can parameterize the same pipeline for different environments or clients.

$config.* – Composition config parameters

$config.<param> reads a composition’s declared config parameter from inside that composition’s body. It is only available in a composition body — a top-level pipeline declares no config schema, so $config.* there is a compile error.

Each parameter is declared in the composition’s _compose.config_schema: block, then read from the body’s CXL:

# in fraud_check.comp.yaml
_compose:
  name: fraud_check
  config_schema:
    threshold: { type: float, default: 0.8 }
nodes:
  - type: transform
    name: flag
    input: inp
    config:
      cxl: |
        emit order_id = order_id
        emit flagged = score >= $config.threshold

Unlike $vars.* (which flows to the executor as a runtime value), $config.<param> is constant-folded at compile time: each reference is replaced by the value resolved for that instantiation, so two call sites of the same composition with different config: compile to different bodies. The resolution precedence, highest first, is a channel/group config: clobber, then the call site’s config:, then the signature default.

Because the value is resolved per instantiation, overriding a config knob via a channel or group config: value clobber changes what the composition body computes — the override is applied to execution, and the winning layer is still recorded in the provenance side-table for channels resolve / explain --field.

$record.* – Per-record scoped state

$record.* is a per-record key-value store that travels with the record through the pipeline but never serializes as an output column. It is the mechanism for tagging records with quality flags, routing hints, or audit information that should not appear in the final output unless explicitly re-emitted as a regular column.

Each $record variable is declared in the writing Transform’s config.declares: block (scope: record) and written from that Transform’s CXL:

Writing record state

- type: transform
  name: classify
  input: orders
  config:
    declares:
      - { name: quality, scope: record, type: string }
    cxl: |
      emit order_id = order_id
      emit $record.quality = if amount < 0 then "suspect" else "ok"

Reading record state

Any downstream node reads it via $record.<key>:

filter $record.quality == "ok"
emit audit_quality = $record.quality

See Scoped Variables for the full declaration model and the pipeline / source / record lifetimes.

now – Current time

The now keyword returns the current wall-clock time as a DateTime value. It is evaluated fresh per record, so each record gets the actual time of its processing.

$ cxl eval -e 'emit timestamp = now'
{
  "timestamp": "2026-04-11T15:30:00"
}

now is useful for timestamping records:

emit processed_at = now
emit days_old = now.diff_days(created_date)

Note: now is a keyword, not a function call. Write now, not now().

Complete example

pipeline:
  name: order_enrichment
  vars:
    discount_threshold: 500
    tax_rate: 0.08

  nodes:
    - name: orders
      type: source
      format: csv
      path: orders.csv

    - name: enrich
      type: transform
      input: orders
      cxl: |
        emit order_id = order_id
        emit amount = amount
        emit discount = if amount > $vars.discount_threshold then 0.1 else 0.0
        emit tax = amount * $vars.tax_rate
        emit total = amount * (1 - discount) + tax
        emit processed_at = now
        emit source_file = $source.file
        emit pipeline_run = $pipeline.execution_id

    - name: output
      type: sink
      input: enrich
      format: csv
      path: enriched_orders.csv

Null Handling

Null values in CXL represent missing or absent data. CXL uses null propagation – most operations on null produce null – with specific tools for detecting and handling nulls.

Interactive companion: the null explainer shows, step by step, how an expression is worked out when a field is null, and whether a filter keeps the record.

Null propagation

When a method receives a null receiver, it returns null without executing. This is called null propagation and applies to all methods except the introspection methods.

$ cxl eval -e 'emit result = null.upper()'
{
  "result": null
}

Propagation flows through method chains:

$ cxl eval -e 'emit result = null.trim().upper().length()'
{
  "result": null
}

Null propagation exceptions

Four methods are exempt from null propagation and actively handle null receivers:

MethodNull behavior
is_null()Returns true
type_of()Returns "null"
is_empty()Returns true
catch(x)Returns x

debug(label) is not one of them: on a null receiver it returns null without logging, like any other method.

$ cxl eval -e 'emit a = null.is_null()
emit b = null.type_of()
emit c = null.catch("fallback")'
{
  "a": true,
  "b": "null",
  "c": "fallback"
}

Null coalesce operator (??)

The ?? operator returns its left operand if non-null, otherwise its right operand. It is the primary tool for providing default values.

$ cxl eval -e 'emit a = null ?? "default"
emit b = "present" ?? "default"'
{
  "a": "default",
  "b": "present"
}

Chain multiple ?? operators for fallback chains:

$ cxl eval -e 'emit result = null ?? null ?? "last resort"'
{
  "result": "last resort"
}

Three-valued logic

Boolean operations with null follow three-valued logic: read null as “unknown”, and a result is decided only when the known side settles it.

and

LeftRightResult
truenullnull
falsenullfalse
nulltruenull
nullfalsefalse
nullnullnull

The key insight: false and null is false because the result is false regardless of the unknown value.

or

LeftRightResult
truenulltrue
falsenullnull
nulltruetrue
nullfalsenull
nullnullnull

The key insight: true or null is true because the result is true regardless of the unknown value.

not

OperandResult
truefalse
falsetrue
nullnull

Arithmetic with null

Any arithmetic operation involving null produces null:

$ cxl eval -e 'emit result = 5 + null'
{
  "result": null
}

Comparison with null

== and != never produce null. Two nulls are equal, and a null is not equal to any other value:

$ cxl eval -e 'emit a = null == null
emit b = null != 100'
{
  "a": true,
  "b": true
}

Every other comparison (<, >, <=, >=) involving null produces null, as arithmetic does:

$ cxl eval -e 'emit result = null > 100'
{
  "result": null
}

Null in conditions

A condition that comes out null is treated as “not true”:

  • filter keeps a record only when its condition is exactly true. A null condition drops the record, just as false does.
  • if takes the else branch when its condition is null, or produces null when there is no else.
  • A match without a subject skips an arm whose condition is null. A match with a subject compares with ==, so a null => … arm matches a null subject.

Because not null is also null, a record whose amount is null fails both filter amount > 100 and filter not (amount > 100). And because != never produces null, filter amount != 100 keeps it. Test for null explicitly with is_null(), or give a default first with ??, when it matters which way an empty value goes.

To test for null, use is_null():

$ cxl eval -e 'emit result = null.is_null()'
{
  "result": true
}

Practical patterns

Fallback values with ??

emit name = raw_name ?? "Unknown"
emit amount = raw_amount ?? 0
emit active = is_active ?? false

Safe conversion with try_* and ??

emit price = raw_price.try_float() ?? 0.0
emit qty = raw_qty.try_int() ?? 1

Explicit null testing

filter not amount.is_null()
emit has_email = not email.is_null()

Catch method (equivalent to ??)

emit name = raw_name.catch("Unknown")

Conditional null handling

emit status = if amount.is_null() then "missing"
    else if amount < 0 then "invalid"
    else "ok"

Filter blank or null

# Filter out records where name is null or empty string
filter not name.is_empty()

Null-safe chaining

When working with fields that may be null, place the null check early or use ??:

# Safe: coalesce first, then transform
emit normalized = (raw_name ?? "").trim().upper()

# Safe: test before use
emit name = if raw_name.is_null() then "N/A" else raw_name.trim()

Modules and use

CXL modules organize reusable constants and pure, single-expression functions. Module files use the .cxl extension and are admitted while Clinker plans the pipeline. Execution uses the admitted declarations stored in the compiled plan; it does not read module files again.

Module files

# rules/shared/finance.cxl
let tax_rate = 0.21
let default_currency = "USD"

fn tax(amount) = amount * tax_rate
fn normalize_currency(value) = value.trim().upper()

A module may contain:

  • let constants whose expressions depend only on other constants and pure CXL operations;
  • fn declarations with named parameters and one expression body; and
  • use declarations for other modules.

Functions cannot contain statements such as emit, filter, or distinct. Recursive function calls and cyclic module imports are rejected during planning.

Where use is recognized

Planning resolves module imports from every field that carries executable CXL, including:

  • a Transform’s primary expression, validation checks, and per-record log conditions;
  • an Aggregate’s expression;
  • every Route condition;
  • a Combine predicate and body;
  • Envelope header and footer expressions;
  • Reshape rule conditions, mutations, and synthesized overrides;
  • Cull group-drop conditions; and
  • the same fields inside reachable composition bodies.

Ordinary strings do not participate in module resolution. Node names, validation and log messages, output paths, and other descriptive text cannot introduce an import merely by containing text that resembles use.

Importing and using a module

Module identities and member access both use dot notation:

use shared.finance as finance

emit tax = finance.tax(amount)
emit currency = finance.default_currency

The alias is optional. Without as, the last identity segment is the alias:

use shared.finance
emit tax = finance.tax(amount)

There is no :: member syntax and no wildcard import. A missing member, calling a constant, or reading a function without parentheses is a planning error with the offending module and member named in the diagnostic.

Direct imports and private dependencies

Pipeline CXL can access only modules it imports directly. A module may import another module by its absolute logical identity:

# rules/app/invoice.cxl
use shared.finance as finance

let standard_rate = finance.tax_rate
fn invoice_tax(amount) = finance.tax(amount)
# pipeline transform
use app.invoice as invoice
emit tax = invoice.invoice_tax(amount)

shared.finance is included in the admitted transitive closure, but it is private to app.invoice. The pipeline must add its own use shared.finance if it needs to address that module directly. Dependencies are never re-exported.

Rules-root selection

Clinker selects exactly one rules root for non-catalog module identities. The precedence is:

  1. explicit clinker run --rules-path <DIR>;
  2. pipeline.rules_path in the pipeline YAML;
  3. [catalog].rules_root in clinker.toml; then
  4. the workspace-relative rules/ default.

There is no search path and no first-match shadowing. Every relative candidate is anchored to the selected workspace, not the process working directory or the pipeline file’s directory. See the CLI reference and typed workspace catalog.

An explicit [catalog.rules] entry maps a logical rule identity to a particular workspace-contained file and takes priority over the derived <rules-root>/<identity segments>.cxl path for that identity.

Planning bounds and diagnostics

Planning loads only the direct imports and their transitive dependencies. Each canonical module is parsed once. The default closure limits are:

LimitDefault
One module file1 MiB
Unique modules64
Import depth32
Total closure source16 MiB

Planning fails before execution for a missing or unreadable module, invalid UTF-8 or CXL, duplicate declarations or aliases, an import/function cycle, or a closure that exceeds a bound. Cycle diagnostics show the complete discovered chain so the import edge to remove is visible.

After loading the complete reachable closure, planning validates both declaration graphs:

  • constant dependencies must be acyclic; and
  • function calls must be acyclic, including direct, mutual, and cross-module recursion.

Cycle diagnostics report the complete chain with the relevant call or declaration locations. Imported calls are also checked at the authored call site. The diagnostic names the logical module and member when the member is not a function, the argument count is wrong, or the expanded function body is ill-typed. For example, if shared.numbers.add takes two arguments, the corrected call is:

use shared.numbers as numbers
emit total = numbers.add(left, right)

Source-file lifetime

Module files are an input to planning, not a runtime dependency. Once planning succeeds, the compiled plan owns the immutable parsed declarations for every admitted direct and transitive module. The same plan can execute repeatedly if those source files are renamed, changed, or removed after planning. Changes take effect only after compiling a new plan.

Removing or changing a required file before planning still fails admission. This boundary prevents a checked plan from silently executing different module code and keeps execution independent of filesystem path authority.

Complete example

# rules/etl/clean.cxl
let max_amount = 999999.99

fn normalize_name(name) = name.trim().upper()
fn safe_amount(raw) = raw.try_float() ?? 0.0
fn flag_suspicious(amount, threshold) =
  if amount > threshold then "review" else "ok"
# pipeline CXL
use etl.clean as clean

emit customer = clean.normalize_name(raw_customer)
emit amount = clean.safe_amount(raw_amount)
filter amount <= clean.max_amount
emit review_flag = clean.flag_suspicious(amount, 10000)

The cxl CLI Tool

The cxl command-line tool validates, evaluates, and formats CXL source files. It is the standalone companion to the Clinker pipeline engine, useful for testing expressions, validating transforms, and debugging CXL logic.

Commands

cxl check

Parse, resolve, and type-check a .cxl file. Reports errors with source locations and fix suggestions.

$ cxl check transform.cxl
ok: transform.cxl is valid

On errors:

error[parse]: expected expression, found '}' (at transform.cxl:12)
  help: check for missing operand or extra closing brace
error[resolve]: unknown field 'amoutn' (at transform.cxl:5)
  help: did you mean 'amount'?
error[typecheck]: cannot apply '+' to String and Int (at transform.cxl:8)
  help: convert one operand — use .to_int() or .to_string()

cxl eval

Evaluate CXL expressions against provided data and print the result as JSON.

Inline expression:

$ cxl eval -e 'emit result = 1 + 2'
{
  "result": 3
}

From a file with field values:

$ cxl eval transform.cxl \
    --field Price=10.5 \
    --field Qty=3

From a file with JSON input:

$ cxl eval transform.cxl --record '{"price": 10.5, "qty": 3}'

Multiple inline statements:

$ cxl eval -e 'let tax = 0.21
emit net = price * (1 - tax)' --field price=100
{
  "net": 79.0
}

cxl fmt

Parse and pretty-print a .cxl file in canonical format with normalized whitespace and consistent styling.

$ cxl fmt transform.cxl

Output is printed to stdout. Redirect to overwrite:

$ cxl fmt transform.cxl > transform.cxl.tmp && mv transform.cxl.tmp transform.cxl

Input data

–field name=value

Provide individual field values as key-value pairs. Values are automatically type-inferred:

InputInferred typeExample
Integer patternInt--field count=42
Decimal patternFloat--field price=10.5
true / falseBool--field active=true
nullNull--field value=null
Anything elseString--field name=Alice
$ cxl eval -e 'emit t = amount.type_of()' --field amount=42
{
  "t": "int"
}
$ cxl eval -e 'emit t = name.type_of()' --field name=Alice
{
  "t": "string"
}

–record JSON

Provide a full JSON object as input. Mutually exclusive with --field.

$ cxl eval -e 'emit total = price * qty' \
    --record '{"price": 10.5, "qty": 3}'
{
  "total": 31.5
}

JSON types map directly:

JSON typeCXL type
nullNull
true / falseBool
integer numberInt
decimal numberFloat
"string"String
[array]Array
{object}Map

Output format

Output is always JSON. Each emit statement produces a key-value pair:

$ cxl eval -e 'emit a = 1
emit b = "two"
emit c = true'
{
  "a": 1,
  "b": "two",
  "c": true
}

Date and DateTime values are serialized as ISO 8601 strings:

$ cxl eval -e 'emit d = #2024-03-15#'
{
  "d": "2024-03-15"
}

Exit codes

CodeMeaning
0Success (or warnings only)
1Parse, resolve, type-check, or evaluation errors
2I/O error (file not found, invalid JSON, etc.)

Pipeline context in eval mode

When running cxl eval, a minimal pipeline context is provided:

VariableValue
$pipeline.name"cxl-eval"
$pipeline.execution_idZeroed UUID
$pipeline.batch_idZeroed UUID
$pipeline.start_timeCurrent wall-clock time
$pipeline.source_fileFilename or "<inline>"
$pipeline.source_row1
nowCurrent wall-clock time (live)

Practical usage

Quick expression testing:

$ cxl eval -e 'emit result = "hello world".upper().split(" ").length()'
{
  "result": 2
}

Validate a transform file:

$ cxl check transforms/enrich_orders.cxl && echo "Valid"

Test conditional logic:

$ cxl eval -e 'emit tier = match {
    amount > 1000 => "high",
    amount > 100 => "med",
    _ => "low"
  }' \
    --field amount=500
{
  "tier": "med"
}

Test date operations:

$ cxl eval -e 'emit year = d.year()
emit month = d.month()
emit next_week = d.add_days(7)' \
    --record '{"d": "2024-03-15"}'

Test null handling:

$ cxl eval -e 'emit safe = raw.try_int() ?? 0' --field raw=abc
{
  "safe": 0
}

CLI Reference

Clinker ships two command-line tools: clinker (the pipeline runner) and cxl (the expression checker/evaluator/formatter, covered in the CXL CLI chapter). This page is the complete reference for clinker.

clinker run

Execute a pipeline.

clinker run [OPTIONS] <CONFIG>

Positional arguments

ArgumentDescription
<CONFIG>Path to the pipeline YAML configuration file (required)

Options

FlagDefaultDescription
--memory-limit <SIZE>YAML memory.limit, else 512MMemory budget for the execution. Uses the same grammar as the YAML memory.limit: a byte count with an optional binary (1024-based) K/M/G suffix (K = 1024 bytes, M = 1024², G = 1024³), where a bare integer is bytes. Other forms — a decimal GB, an explicit GiB, or a fractional value such as 1.5G — are rejected. When the limit is approached, aggregation operators spill to disk rather than crashing. When passed, this value overrides any memory.limit set in the pipeline YAML; when omitted, the YAML value applies (or the 512M default when the YAML is also silent). An empty or whitespace-only value — as an ops wrapper produces when it forwards an unset variable, e.g. --memory-limit "$CLINKER_MEM" with CLINKER_MEM unset — is treated the same as omitting the flag. A non-empty malformed value (for example the decimal 4GB rather than the binary 4G) is rejected at the CLI boundary with an error naming --memory-limit and echoing the value, so a typo fails loudly instead of silently falling back to the default and shrinking a larger YAML budget. Because the flag simply populates pipeline.memory.limit, a startup budget error (E312) for a value you passed via --memory-limit refers to that same limit.
--threads <N>YAML pipeline.concurrency.threads, else number of CPUsPositive capacity applied independently to the Rayon CPU-kernel pool and to concurrent Source schema/read work across top-level and composition-body Sources. It is not a total operating-system thread limit: Source workers and the Rayon pool remain distinct. The selected value is recorded in execution metrics. Zero is rejected before the config is opened.
--batch-id <ID>UUID v7Logical-batch correlation available as pipeline.batch_id, in {batch_id} output-path templates, machine events, and opt-in output provenance sidecars. Supplying it does not override the fresh UUIDv7 execution ID and does not provide deduplication, resume, or exactly-once behavior. It is not currently a field in the metrics-spool payload.
--machine ndjson-v1–Opt into the clinker.run schema-1 lifecycle on stdout. Requires a non-empty --batch-id; conflicts with plan/dry-run output and with --lineage - or --lineage-events -. File-based lineage remains compatible: a plan-only --lineage <FILE> export shares this stream’s identity and closes it with an explicit empty publication inventory, since it runs no attempt. Every line is one compact JSON object; human diagnostics move to stderr. Consumers must concurrently drain both pipes, reject unsupported schema majors, accept only additive schema-1 fields, and reconcile exactly one supported terminal with the actual process status and current-attempt artifact evidence. EOF, malformed output, forced termination, or a missing/duplicate terminal is incomplete, never success. See Running Clinker Directly or Under a Supervisor.
--explain [FORMAT]textPrint the execution plan and exit without processing data. Accepted formats: text, json, dot. With json or dot, standard output carries only the document and human diagnostics move to stderr, so a consumer can redirect stdout straight into a parser; with text they stay together on stdout. See Explain Plans.
--lineage <PATH>–Preflight the workspace lineage identity policy, build column lineage, and write it as OpenLineage NDJSON, then exit without processing data. Give a file path, or - for stdout. The export is the whole invocation, so one that cannot be delivered exits non-zero rather than reporting success: a destination the exporter cannot write exits 4, and an event the [observability.lineage] byte caps reject exits 1. Each diagnostic names the destination, states which of the two failed, and prints the configuration change where one applies. A failed export leaves no partial file behind, so a following upload step cannot pick up a stale one. Both this flag and --lineage-events need the lineage capability, which the released binary has; a build compiled without it refuses the flag at validation rather than exiting zero having emitted nothing (see Optional capabilities). See Column Lineage.
--lineage-events <PATH>–Preflight the workspace lineage identity policy, run the pipeline, and emit live OpenLineage run events (a START at run begin, then a terminal COMPLETE / FAIL / ABORT with real timing and row counts) as NDJSON to a file path, or - for stdout. Cannot be combined with --lineage, --explain, --dry-run, or -n. With -, normal run output can interleave with the event stream; use a file for clean NDJSON. See Live run events.
--dry-run–With no -n, performs complete config, overlay, CXL, schema, DAG, resource, and publication-configuration validation, prints resolved outputs, and exits without opening or reading a Source and without opening or publishing a Sink.
-n, --dry-run-n <N>–Bounded preview. Requires --dry-run and a positive N. Clinker checks the limit before every read and reads at most N records from each declared Source, including Sources inside composition bodies. Records drain in stable plan order to the explicit preview stream; configured Sink paths are never opened or published. Preview emits no live run telemetry or lineage lifecycle.
--dry-run-output <FILE>stdoutDestination for bounded-preview bytes. Requires --dry-run-n; without it, the option is rejected before config access. All preview Sinks write through this one explicit destination using their configured formats.
--rules-path <DIR>selected workspace’s rules/Select the CXL module rules root for this run. Precedence is explicit CLI value, then pipeline.rules_path, then [catalog].rules_root, then the workspace-relative rules/ default. A relative value is anchored to the workspace selected by --base-dir or workspace discovery, not the process working directory. One root is selected; Clinker does not search multiple roots. See Modules and use and the typed workspace catalog.
--base-dir <DIR>–Base directory for resolving relative paths in the YAML config. Defaults to the directory containing the config file.
--allow-absolute-paths–Permit absolute file paths in the pipeline YAML. By default, absolute paths are rejected to encourage portable configs.
--env <NAME>–Sets CLINKER_ENV in the current process before the pipeline loads. The current run path does not otherwise consume that value for channel selection; select a channel explicitly with --channel.
--quiet–Suppresses the “applied overlay” summary. Other stdout, tracing, warnings, and errors are not uniformly silenced.
--force–Overrides an output’s if_exists: error policy and permits overwrite. Outputs using the default if_exists: overwrite already overwrite without this flag; unique_suffix keeps its own collision behavior.
--log-level <LEVEL>infoClosed logging level: error, warn, info, debug, or trace. Any other spelling is rejected.
--metrics-spool-dir <DIR>–Directory for per-execution metrics files. See Metrics & Monitoring.
--channel <ID>–Apply a logical id from [catalog.channels]. The selected file must also have a [catalog.pipelines] id listed in the channel manifest. Matching groups are target-bounded before labels narrow them.
--group <NAME>–Force-include a group overlay by name (repeatable). The selected pipeline or one of its admitted compositions must appear in the group’s explicit targets: set. Use clinker channels resolve to preview the effective plan.
--no-auto-groups–Suppress selector-derived group membership; only groups named with --group apply.

--error-threshold is retired and rejected. Configure the typed pipeline policy instead; the CLI diagnostic prints this paste-ready replacement:

error_handling:
  type_error_threshold: 0.05

Credential profile foundation

The current binary does not yet accept a credential-profile option or a credential-profile configuration table. Referenced credentials therefore do not activate a source, destination, or observability exporter through this surface. Do not pass an environment, channel, or group as a substitute: those selectors never choose credentials, and there is no default or sentinel profile.

The run-local foundation that later preflight wiring will call is already bounded. Its default ceilings are 64 named profiles, 256 provider registrations across those profiles, 1 MiB of decoded profile/provider definition state, and 256 simultaneously retained handles. Admission checks all definition counts and bytes before a profile can resolve a requirement. Each live lease reports its exact retained bytes to the run memory arbitrator before provider allocation. The registry has no inbound producer and is not a backpressure target. An arbitrator spill callback queues a request and reports zero bytes freed synchronously; the next registry-owned checkpoint revokes and releases the partial set in reverse acquisition order and unregisters the registry. Cap, memory, and provider failures follow the same fail-closed cleanup path; an explicit run coordinator may pause acquisition until resume.

These are foundation limits, not newly available command-line behavior. A later complete preflight surface must add the one explicit profile selector, credential-required omission checks, and consumer activation together before the option can appear in the options table above.

Examples

# Basic execution
clinker run pipeline.yaml

# Production run with memory budget and forced overwrite
clinker run pipeline.yaml --memory-limit 512M --force --log-level warn

# Validate without processing
clinker run pipeline.yaml --dry-run

# Preview at most 25 records from each declared Source without publishing Sinks
clinker run pipeline.yaml --dry-run -n 25 --dry-run-output preview.csv

# Compile and explain without reading data
clinker run pipeline.yaml --explain text

# Show execution plan as Graphviz
clinker run pipeline.yaml --explain dot | dot -Tpng -o plan.png

# Run with a batch ID available to templates and provenance sidecars
clinker run pipeline.yaml --batch-id "daily-2026-04-11"

# Emit the bounded schema-1 lifecycle for a supervisor
clinker run pipeline.yaml --machine ndjson-v1 --batch-id "daily-2026-04-11"

The standalone command remains the common case and requires no supervisor. Machine mode adds a child-process control stream; it does not add scheduling, retry, heartbeat, or process-tree management to Clinker. A supervising parent must heartbeat independently of advisory progress and start a fresh process with a new execution ID for every retry.

Typed resource failures carry registry-owned messages and policy_required retry advice. A supervisor must decide whether a fresh attempt is safe; temporary-storage or delivery failure can follow bytes already accepted by a destination and does not establish rollback.

Resource conditionFailure codeCategory
Memory budget refusedruntime.resource.memory_budget_exceededinfrastructure
Allocation or allocation layout unsatisfiedruntime.resource.allocation_failedinfrastructure
Spill quota refusedruntime.resource.spill_cap_exceededinfrastructure
Descriptor quota refusedruntime.resource.descriptor_cap_exceededinfrastructure
Temporary storage or readback failedruntime.resource.storage_failedinfrastructure
Continuation refused after delivery failedruntime.resource.delivery_poisonedinfrastructure
Resource finalized or ownership authority violatedruntime.invariant.unknowninternal_invariant

Explicit resource cancellation is an aborted run, reported as cancelled in machine mode with exit 130 and no failure classification. It does not depend on telemetry delivery or a pending shutdown signal. A genuine resource or data failure retains its classification even when a shutdown signal is pending; malformed source data remains source.data.invalid with do_not_retry advice.


clinker guess

Preview, exhaustively check, or safely write concrete int or float replacements for inference-only numeric columns.

clinker guess [OPTIONS] <CONFIG>

With no selector, guess reads the base pipeline. --channel <ID> selects one cataloged channel plus its target-admitted derived groups; --group <NAME> selects one explicit, target-admitted group without a channel. The two selectors conflict. A missing or ambiguous selector is an error rather than a fallback to the base pipeline, so the report always describes exactly one effective configuration.

FlagDescription
<CONFIG>Pipeline YAML containing the source-schema numeric leaves to inspect.
--channel <ID>Select one cataloged channel and its derived, target-admitted groups.
--group <NAME>Select one explicit, target-admitted group without a channel.
--field <NODE.COLUMN>Narrow the preview to one numeric source field. Repeatable; repeated selectors are deduplicated in request order. In a multi-record schema, one selector covers every same-named literal numeric owner across records: while leaving same-named concrete declarations unchanged. Unknown, malformed, or entirely concrete fields are rejected.
--checkExhaust the frozen, capped manifest and exit 3 if any selected owner remains unresolved.
--writeExhaust evidence and edit exactly one inline, literal, single-owner numeric leaf after guarded compare-and-swap revalidation. Mutually exclusive with --check.
--base-dir <DIR>Workspace root holding clinker.toml and the channel/group roots. Defaults to the pipeline file’s directory.

The command constructs readers through the same CSV, JSON, and XML option and schema-coercion path as runtime ingest. It emits one deterministic JSON document containing the selected configuration, bounded coverage, parser-owned numeric evidence, unresolved reasons, proposed types, and an exact semantic YAML patch. Preview and check never edit. Write reports written only after guarded publication completes; otherwise it reports not_written and leaves the patch for manual application. It never edits an overlay.

Each field report carries an owners array. Every owner has its own evidence, votes, proposed type, unresolved reasons, and exact address. A single-record field has one canonical /v1/schema/sources/.../columns/... owner. Multi-record owners use /v1/schema/sources/.../records/.../columns/..., in authored record order, so same-named leaves never collapse into a fake single-record address or share a proposal derived from another record type.

Discovery freezes the complete deterministic manifest up to 4,096 files and reports its fixed-size path/order/size identity without listing every path. Preview admits at most four files and 8 MiB of discovered file sizes globally, round-robin across sources while preserving each source’s stable prefix. Every selected reader enforces its discovered length and fails if the file is opened at another size, truncates, or grows; it never hands a format reader bytes past the admitted boundary. Preview then reads at most 1,024 records globally. Coverage retains at most four file details per source and reports sampled, truncated, uncovered, and unreported aggregate counts for the rest. Multi-pass formats can report physical bytes_read above the admitted input size because each pass is counted. --check reads every file and record in that same frozen manifest instead of applying the preview budgets.

Candidate storage is capped at 100,000 source-schema leaves, matching the canonical YAML parser’s 100,000-node limit; each YAML document is also capped at 32 MiB. Each parser-owned NumericObservation retains at most 128 numeric lexeme bytes, and each exact schema owner retains at most eight evidence items. These limits appear in the JSON report. Candidate, observation, evidence, coverage, and write-snapshot collections are fixed-bounded independently of input size. Write retains one fixed-size digest per file in the 4,096-file manifest and hashes through a fixed 1 MiB buffer; it does not register a runtime memory consumer.

Write is eligible only for one resolved owner in the base pipeline whose exact authored bytes and direct YAML provenance identify a literal inline numeric leaf. It preserves all sibling bytes and proves by canonical reparse that only that type leaf changed in the resolved staged configuration. Overlays, external/generated or synthetic ownership, aliases, interpolation, symlinks, non-local inputs, multiple owners, unresolved evidence, and input/config drift are patch-only outcomes.

The CLI holds an advisory fs4 lock on the stable sibling <CONFIG>.clinker-guess.lock file through sibling-temp flush/fsync, final byte/semantic/input revalidation, atomic rename, and directory fsync. The lock file remains beside the config so every cooperating replacement locks the same inode. It must remain a regular non-symlink file and is owner-only on Unix. This is a fail-closed compare-and-swap for cooperating writers using the same lock, not a kernel-enforced conditional rename against a writer that ignores advisory locking; such a writer can race after the final comparison.

Exit 0 means a complete preview was written, including an unresolved preview, an exhaustive check resolved every owner, or one safe write was published. Selection, configuration, and field errors exit 1; unresolved checks and patch-only/no-edit writes exit 3; source discovery, input I/O, reader, publication, signal-handler, and stdout failures exit 4; interruption exits 130 before any report is emitted or before publication begins.

# Preview every inference-only numeric leaf in the base pipeline
clinker guess pipeline.yaml

# Preview one field under a cataloged channel
clinker guess pipeline.yaml --channel acme --field orders.amount

# Preview one explicit group without a channel
clinker guess pipeline.yaml --group enterprise

# Exhaust the frozen manifest and make unresolved owners fail the command
clinker guess pipeline.yaml --check

# Exhaust evidence and edit one safe directly-authored owner
clinker guess pipeline.yaml --field orders.amount --write

clinker explain

Inspect one compiled field’s provenance or discover registry-owned diagnostic descriptors.

clinker explain <CONFIG> --field <PATH> [OPTIONS]
clinker explain --list [--status <STATUS>] [--category <CATEGORY>]
clinker explain --code <CODE>

Exactly one of --field, --list, or --code is required. A pipeline path is required only for --field and is rejected for the two static discovery modes.

FlagDescription
--field <PATH>Explain one exact or unambiguous shorthand field address in the compiled pipeline.
--listPrint every registered descriptor in stable code order.
--status <STATUS>With --list, require active or retired-reserved.
--category <CATEGORY>With --list, require one of configuration, composition, source-and-expression, execution-and-format, terminal-authoring, security, or advisory.
--code <CODE>Print one registered descriptor and its optional longer detail page.
--channel <ID>With --field, apply the selected channel before compiling provenance.
--group <NAME>With --field, force-include a group overlay (repeatable).
--no-auto-groupsWith --field, suppress selector-derived groups.
--base-dir <DIR>Workspace root used by the field-provenance compile path; defaults to ..

All seven descriptor fields—code, severity, status, category, retryability, meaning, and correction—come from the same leaf registry in both discovery views. Closed enum values in the descriptor use lowercase kebab-case; filter spellings come from the same enum tables used to parse those filters. A detail page can add examples but cannot define whether a code exists. Unknown or empty filters, no-match combinations, unknown codes, and conflicting modes exit nonzero. See Explain Plans for examples and for the separate clinker run --explain plan display.


clinker metrics collect

Sweep per-execution metrics files from a spool directory into a single NDJSON archive.

clinker metrics collect [OPTIONS]

Options

FlagDescription
--spool-dir <DIR>Spool directory to sweep (required).
--output-file <FILE>NDJSON archive destination (required). If the file exists, new entries are appended.
--delete-after-collectRemove spool files after they have been successfully written to the archive.
--dry-runPreview which files would be collected without writing anything.

Examples

# Collect and archive, then clean up spool
clinker metrics collect \
  --spool-dir /var/spool/clinker/ \
  --output-file /var/log/clinker/metrics.ndjson \
  --delete-after-collect

# Preview what would be collected
clinker metrics collect \
  --spool-dir ./metrics/ \
  --output-file ./archive.ndjson \
  --dry-run

clinker channels

Inspect and validate the channel/group multi-tenant overlay system.

clinker channels resolve <TARGET> [OPTIONS]
clinker channels lint [OPTIONS]
clinker channels group members <GROUP> [OPTIONS]
clinker channels label set <KEY>=<VALUE> <CHANNEL_ID>... [OPTIONS]

clinker channels resolve

Renders the effective post-overlay plan for one target — the DAG plus per-value provenance (which layer supplied each value, and which group injected which node). This answers “what does tenant X actually run?”.

FlagDefaultDescription
<TARGET>–Path to the base pipeline (or composition) YAML to resolve (required).
--channel <ID>–Logical channel id from [catalog.channels]. Matching groups are derived only after explicit target admission.
--group <NAME>–Force-include a group overlay by name (repeatable), subject to the same target set as automatic selection.
--no-auto-groups–Suppress selector-derived group membership.
--base-dir <DIR>.Workspace root holding clinker.toml and the channel/group roots.

Exits non-zero when the overlay raises an error (e.g. a config key matching no parameter), so resolve doubles as a targeted check for one tenant.

clinker channels lint

Compiles every target declared by every [catalog.channels] entry and reports failures — the CI safety net for base-change blast radius. It uses the same logical identity and target-scope checks as run and explain; a basename or current working directory never supplies target identity.

FlagDefaultDescription
--base-dir <DIR>.Workspace root to lint.

Exits non-zero if any combination fails to compile or apply. Dangling splice anchors (an op referencing a missing node) and config keys matching no parameter are reported per combination.

clinker channels group members

Lists the channels whose labels currently satisfy a group’s selector — “who is in this group right now?”. Because membership is derived from labels, this evaluates the group’s match: selector against each channel’s manifest labels through the same derivation the overlay resolver uses.

FlagDefaultDescription
<GROUP>–Group name (the group.name of a *.group.yaml).
--base-dir <DIR>.Workspace root holding clinker.toml and the channel/group roots.

A group with no match: selector is explicit-only and reports no derived members. A channel whose labels make the selector ill-typed or reference an undeclared label is reported as a selector error (never a silent non-match), and the command exits non-zero when any such error occurs.

clinker channels label set

Stamps (or overwrites) one label across the named channels by editing each channel’s channel.cfg.yaml manifest in place. Idempotent: re-running with the same value writes nothing. Only the manifest’s labels: block is rewritten; other keys and comments are preserved. The channel manifest must already exist with a non-empty channel.targets list; label set will not create a targetless manifest.

FlagDefaultDescription
<KEY>=<VALUE>–Label assignment. KEY must be an identifier (letters, digits, _) so a selector can reference it. VALUE is typed by YAML scalar inference (true/false → bool, integers → int, decimals → float, otherwise string).
<CHANNEL_ID>...–One or more channel ids (tenant folder names) to stamp.
--base-dir <DIR>.Workspace root holding clinker.toml and the channel root.

Because group membership is attribute-derived, label set is the maintenance operation for group membership: set a label once and every group whose selector matches gains the channel — no membership list to hand-edit.

Examples

# What does tenant `globex` actually run for this pipeline?
clinker channels resolve pipeline/order_fulfillment.yaml --channel globex

# Preview a group overlay standalone (no channel)
clinker channels resolve pipeline/order_fulfillment.yaml --group enterprise

# Compile every channel/group overlay in the workspace and report failures
clinker channels lint

# Which channels are currently in the `enterprise` group?
clinker channels group members enterprise

# Onboard two tenants into the enterprise tier in one shot
clinker channels label set tier=enterprise globex acme-corp

clinker refactor

Structural refactors that span a base pipeline and every channel/group overlay that references it.

clinker refactor rename-node <TARGET> <OLD> <NEW> [OPTIONS]

clinker refactor rename-node

Renames a base node and propagates the rename to every overlay reference. The overlay op model addresses base nodes by name, so renaming a node otherwise breaks every overlay that referenced it. This command rewrites, in one operation:

  • the base node’s name and every consumer’s input: / inputs: / body:/header:/trailer: reference;
  • a Combine’s named-input map (qualifier key and/or upstream value) and — when the Combine draws from the renamed node under a same-named qualifier — its where: / cxl: bodies, rewritten via the CXL parser so only true source qualifiers are touched (a method receiver like region.contains(...) is left alone);
  • across every target-admitted group / channel-manifest / per-target overlay file: op target, after, before, injected alias, explicit input, rewire keys and values, an inline node, a set config.cxl value’s CXL, and top-level config dotted-path prefixes (old.param → new.param).
FlagDefaultDescription
<TARGET>–Path to the base pipeline (or composition) YAML that declares the node.
<OLD>–Current node name (must exist in the target).
<NEW>–New node name — identifier only (letters, digits, _); must not already exist in the target.
--dry-run–Print the diff of every file that would change without writing anything.
--base-dir <DIR>.Workspace root holding clinker.toml and the channel/group roots.

Ambiguity is guarded: renaming to a name that already exists in the target, or renaming a node that does not exist, is a hard error. A Combine where:/cxl: body that must be rewritten but does not parse aborts the whole operation before anything is written. After a real (non-dry-run) run the command re-runs channels lint so an incomplete rename fails loudly.

Scope is catalog- and target-bounded. Per-target files are matched by their logical channel.target; channel manifests are admitted only when channel.targets contains the selected pipeline; and groups are admitted only when group.targets names that pipeline or a composition in its resolved closure. Filenames and basenames never establish identity, and a selector match cannot widen the refactor beyond the group’s declared target set.

Files are rewritten by re-serializing their YAML: key order is preserved, but comments and incidental scalar styling are normalized. Use --dry-run to review the exact on-disk diff first.

Examples

# Preview a rename across the base pipeline and every overlay that references it
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --dry-run

# Apply it, then re-lint the workspace
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases

clinker config

Inspect a pipeline configuration file.

clinker config --resolved <CONFIG>

clinker config –resolved

Prints the config with the multi-value shorthand expanded to canonical form. The bare-field forms of split_to_rows:, split_values:, and join_values: are rewritten to full mappings with every default spelled out — so you can see exactly what the engine runs:

  • a bare - line_items under split_to_rows: becomes - { field: line_items, keep_empty: true, mode: extract };
  • a bare - tags under split_values: becomes - { field: tags, delimiter: ";" };
  • a bare - tags under join_values: becomes - { field: tags, delimiter: ";", on_conflict: error, escape: "\\" }.

The rewrite is surgical: only those shorthand blocks change. Comments, key order, indentation, and every other surface are preserved byte-for-byte, so the output parses to a plan semantically identical to the input, and running config --resolved on the result is a no-op. Schema columns are already canonical (multiple: true is always written explicitly), so the schema block is left untouched.

This is config canonicalization for the pipeline file itself. It is distinct from clinker channels resolve, which renders the effective post-overlay plan for a specific tenant.

A few surfaces are deliberately left as written rather than expanded, since regenerating them would lose information: a shorthand block that carries an interior comment or blank line between its items is passed through unchanged (so the comment is never dropped), and a value written as a YAML alias (*anchor) is left in place — the anchor it points to is expanded at its definition, so the alias still resolves to the expanded value. The output uses the input file’s line endings (LF or CRLF).

FlagDefaultDescription
<CONFIG>–Path to the pipeline YAML config file (required). The file is validated before it is rewritten, so a malformed config fails with a config error rather than emitting a half-expanded document.
--resolved–Print the fully-expanded canonical form to stdout. Currently the only mode; required.

Examples

# Show the fully-expanded canonical form
clinker config --resolved pipeline.yaml

# Materialize the shorthand into a new file
clinker config --resolved pipeline.yaml > pipeline.canonical.yaml

See Source Nodes → Multi-value fields for the shorthand these forms expand from.


clinker attempts

Inspect and clean up retained publication attempts owned by a pipeline:

clinker attempts list <PIPELINE> [--path-execution-id <ID>] [--continuation <TOKEN>] [--show-paths] [--format text|json]
clinker attempts inspect <PIPELINE> --execution-id <ID> [--path-execution-id <ID>] [--show-paths] [--format text|json]
clinker attempts purge <PIPELINE> (--execution-id <ID> | --expired) [--path-execution-id <ID>] [--execute] [--continuation <TOKEN>] [--show-paths] [--format text|json]

<PIPELINE> is required and must be a traversal-free, workspace-relative .yaml or .yml path. Every invocation reloads and compiles that pipeline, then derives its finite destination-parent roots from the compiled config. There is no option for supplying a storage root, deletion path, or safety override.

When the original run used path- or overlay-affecting options, repeat them on the attempt command: --base-dir, --allow-absolute-paths, --rules-path, --channel, repeatable --group, and --no-auto-groups. Output templates that use run identity also require the matching --path-execution-id, --batch-id, or --timestamp. The path identity is deliberately distinct from inspect and purge’s --execution-id selector, so purge --expired can reconstruct an execution-scoped destination without changing its selector. Attempt operations replay file-source discovery for {source_file} and {source_path} fan-out and anchor a pipeline without --base-dir at the pipeline’s own directory, matching run. These values recompile typed ValidatedPath roots and never grant authority to a caller-supplied deletion path.

list and inspect never mutate retained state. purge is also non-mutating by default: it reports the attempts that the current retention policy admits. Only --execute performs bounded cleanup. Live locks, invalid ownership, unsupported filesystem entries, ambiguous clocks, and unreadable manifests remain keep decisions even with --execute.

Command or flagBehavior
listLists retained attempts across all existing roots owned by the freshly compiled pipeline.
inspect --execution-id <ID>Reports one canonical execution ID across those roots.
purge --execution-id <ID>Previews one logical execution; add --execute to remove only positively owned, eligible files.
purge --expiredPreviews all policy-expired attempts admitted by the bounded page; add --execute to clean them.
--continuation <TOKEN>Resumes the exact plan-, root-, and selector-bound page emitted by a partial result. JSON resume_argv is authoritative; the text command applies platform quoting to the raw opaque token.
--show-pathsAdds sanitized workspace-relative attempt paths. Machine-local prefixes and sensitive-looking components remain redacted.
--format jsonEmits one compact JSON object with stable field order and logical identifiers. The default is deterministic human-readable text.

Default output is path-free. It contains the logical root ID, execution ID, lifecycle state, eligibility, artifact IDs, cleanup debt, and exact bounds. The compact JSON form carries the same fields plus shell-independent recovery_argv and resume_argv arrays. Neither form includes record values, credentials, secrets, or raw debug data.

Safety refusals and incomplete cleanup exit with status 4 and use the stable E371 or E372 data. The report includes its logical failure code, registry-owned retry advice, and a pasteable workspace-relative recovery command, for example:

diagnostic: E371
failure: attempt.retention.manifest_invalid
retry: policy_required
recover: clinker attempts inspect pipelines/orders.yaml --execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159

Examples:

# Path-free, non-mutating inventory
clinker attempts list pipelines/orders.yaml

# Inspect one retained execution as compact JSON
clinker attempts inspect pipelines/orders.yaml \
  --execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159 \
  --format json

# Preview expired cleanup, then perform the same bounded selection
clinker attempts purge pipelines/orders.yaml --expired
clinker attempts purge pipelines/orders.yaml --expired --execute

See Storage & Spill Location for retention, bounds, destination qualification, and cleanup ordering.


Environment Variables

VariableDescription
CLINKER_ENVActive environment name. Equivalent to --env. Used by when: conditions in channel overrides to select environment-specific configuration.
CLINKER_METRICS_SPOOL_DIRDefault metrics spool directory. Overridden by --metrics-spool-dir.

Precedence (highest to lowest): CLI flag, environment variable, YAML config value.

Validation and Admission

clinker-plan is the authority that admits a pipeline to execution. A pipeline is executable only after the planner has parsed canonical YAML, bound schemas and compositions, type-checked Clinker Expression Language (CXL), and produced a CompiledPlan.

Canonical planner validation

clinker run pipeline.yaml --explain text

This compiles the complete pipeline through clinker-plan, the sole authority that can admit it for execution. The command checks:

  • YAML structure and required fields
  • CXL syntax and compile-time type checking
  • Schema compatibility between connected nodes
  • DAG wiring (no cycles, dangling inputs, or missing nodes)
  • Plan-time source and output configuration gates

No runtime readers are opened and no output files are created. Planning may inspect available file metadata or evaluate matchers for cost estimates. The command exits with code 0 only after the planner produces a CompiledPlan, and with code 1 for a configuration, schema, or plan diagnostic. Admission does not prove that later input decoding or I/O will succeed, and the rendered plan does not prove output correctness.

Composition resource descriptors and bindings are part of this admission. The planner checks the bounded [catalog.resources] table, declared _compose.resources_schema slots, call-site and overlay logical identities, kind/capability compatibility, required slots, fixed locks, and recursive composition bodies without resolving credentials or opening handles.

An ordinary composition call containing alias: or outputs: fails during strict YAML parsing with E377 at the authored location. Replace alias: with the composition node’s name:. Declare ports under _compose.outputs and refer to them downstream as <composition-node-name>.<port>. The separate add.alias field remains valid only within an overlay add operation.

Bare --dry-run performs the same planner compilation without rendering the plan:

clinker run pipeline.yaml --dry-run

Prefer --explain text when reviewing a schema change because the resulting plan is visible evidence of what the planner admitted. See Explain Plans for text, JSON, and DOT plan output.

Guessing numeric types and repeated source fields

numeric is an authoring-only placeholder. Runtime planning still rejects it with E158; use clinker guess to collect the real readers’ parser evidence and produce an exact patch, then review and apply that patch before compilation.

clinker guess pipeline.yaml
clinker guess pipeline.yaml --field orders.amount --field orders.tax
clinker guess pipeline.yaml --channel production --check
clinker guess pipeline.yaml --field orders.amount --write

With no selector, the base pipeline is inspected. Exactly one --channel ID or --group NAME selects an effective configuration; the two options conflict. Repeatable --field node.column selectors narrow literal numeric leaves or select one concrete column from a single-record CSV, JSON, or XML source for multiplicity review. Numeric selectors retain their existing meaning and can represent more than one authored multi-record leaf; the report gives every exact owner address separately. Concrete multiplicity candidates are limited to directly authored columns that do not already declare multiple: true.

The default preview is deterministic, bounded, and read-only. It freezes the configured stable file order (name ascending by default) and reports the fixed-size identity of its normalized paths, order, and sizes, then allocates four file opens, 1,024 records, and 8 MiB of admitted file sizes globally in round-robin source/file order. Each selected reader is pinned to its discovered length: a pre-open mismatch, truncation, or growth fails instead of expanding the preview, and no format reader receives bytes beyond that admitted length. Multi-pass formats may report more physical bytes_read because each bounded pass rereads the same admitted input. The manifest itself is capped at 4,096 files; narrow a matcher or use files.take_first / files.take_last if the selected set is larger. The YAML/configuration cap limits candidates to 100,000 source-schema leaves. Per owner, at most eight representative observations are retained, each with at most 128 bytes of numeric lexeme evidence. Coverage retains at most four file details per source and reports aggregate sampled, truncated, uncovered, and unreported counts for the rest. Multiplicity inference retains counters and a fixed interpretation set, never field values or a raw sample corpus. More than 16 distinct CSV delimiter candidates is review-only. These fixed bounds are also printed in the JSON report when they apply.

The manifest identity covers normalized path, configured order, and discovered size; it is not a content hash or a compare-and-swap proof. Preview and check perform no edit. Write additionally streams an exact BLAKE3 snapshot of every file in the capped manifest before evidence collection and compares it again after collection and immediately before publication.

--check uses the same frozen, capped manifest but reads every selected file and record. It is exhaustive over that manifest rather than subject to the preview’s open/record/byte sampling budgets. --write is equally exhaustive and edits only when exactly one resolved owner is directly authored in the base pipeline. Numeric evidence may replace one literal numeric leaf. Multiplicity evidence may set one column’s multiple: true and, for CSV, add one complete split_values entry using the proven delimiter and activated escape. Both are one owner mutation. An overlay, external/generated owner, alias, interpolation, existing conflicting split declaration, no-op already-multiple column, symlink, non-local input, multiple owners, unresolved evidence, or changed snapshot leaves the pipeline untouched and reports the patch with exit 3. The edit is reparsed and compared with a typed expected configuration, so comments, ordering, spans, and every unrelated scalar remain unchanged.

Publication holds an advisory fs4 lock on the stable sibling <CONFIG>.clinker-guess.lock file through a sibling-temp flush/fsync, final exact byte/semantic/input revalidation, atomic replacement, and parent directory fsync. The lock file remains beside the config so cooperating replacements keep one lock inode across renames; it must remain regular, non-symlinked, and owner-only on Unix. This catches any change visible at the final comparison, but is not a kernel-enforced content-conditional rename: a writer that ignores the advisory lock can still race after the final check. Use one cooperating configuration writer per file.

Numeric votes come only from the parser-owned observations used by the shared runtime reader construction. Exact integers vote int; finite, representation- safe values vote float. Mixed integer/float evidence resolves to float only when every integer is exactly representable there. Numeric defaults vote through the schema parser. Accepted missing/null/empty states abstain but remain reported, forbidden absence is a conflict, and all-no-value evidence remains unresolved. No confidence threshold or statistical guess is used.

Repeated-value evidence

Multiplicity is proved per logical record; counts from separate records are never added together. The production reader runs against a temporary schema clone so it can retain ordered repeated values for observation without changing the effective pipeline:

  • XML becomes conclusive when one record contains two or more sibling elements at the selected path. A sibling in each of two records is still unconfirmed.
  • JSON becomes conclusive when one record contains an array longer than one. Null, empty, and one-element arrays remain unconfirmed.
  • CSV becomes conclusive only when exactly one delimiter/activated-escape interpretation parses and re-encodes every non-null cell to the original bytes in the source’s declared character set, and at least one cell produces multiple values. Two surviving interpretations are review-only.

For example, each selected tags field below uses the same existing schema surface:

schema:
  - name: tags
    type: string

Conclusive XML has two siblings in one row:

<root><row><tags>a</tags><tags>b</tags></row></root>

Conclusive JSON has an array longer than one:

[{"tags": []}, {"tags": ["a"]}, {"tags": ["a", "b"]}]

Conclusive CSV has one lossless interpretation:

tags
a|b
plain

The corresponding safe CSV edit reuses the normal multi-value syntax:

split_values:
  - field: tags
    delimiter: "|"
schema:
  - name: tags
    type: string
    multiple: true

These inputs remain review-only or unconfirmed and cannot write:

[{"tags": []}, {"tags": ["a"]}, {"tags": ["b"]}]
tags
a|b;c
d|e;f

Run the exhaustive gate before requesting a write:

clinker guess pipeline.yaml --field values.tags --check
clinker guess pipeline.yaml --field values.tags --write
ExitMeaning
0Preview completed, including a preview with unresolved owners; exhaustive check resolved every owner; or write published its one safe edit.
1Configuration or selection error.
3Exhaustive check is unresolved, or write emitted a patch but did not safely edit.
4Source discovery, reader, I/O, signal-handler, or report-output failure.
130Interrupted before a complete report could be emitted.

Inspect outcome, every owner-level unresolved_reasons entry, coverage, and the emitted patch. A preview exit of 0 is not proof that every selected field resolved; use --check when the exit status must enforce that condition.

Bounded execution preview

clinker run pipeline.yaml --dry-run -n N executes a bounded sample after planning. N must be positive; the runtime checks the limit before each read and reads at most N records from each declared Source, including Sources inside composition bodies. This is a source-record limit, not an output-row limit: filtering, aggregation, joins, and fan-out can change the output count.

Configured Sink paths are never opened or published. Instead, all preview Sinks serialize in stable plan order to stdout, or to the one explicitly selected --dry-run-output PATH. That explicit destination is written; choose it deliberately. Preview emits no live run telemetry or lineage lifecycle.

Without -n, --dry-run performs planning only and reads no input records. --dry-run-output requires -n, and -n requires --dry-run. A successful sample validates only the sampled data; it does not prove that the rest of the input will pass. See CLI options.

Known limitation: At revision 3b343a4e, some bounded previews fail during cleanup with completed node-buffer scope retained ..., even when the same pipeline completes in an ordinary run. Treat that nonzero exit as a failed preview. Planning-only validation remains available; use an ordinary run with disposable output paths to verify actual results.

Advisory workspace schema analysis

clinker-schema is a separate advisory authoring library. It reuses the canonical typed YAML parser for pipeline structure and schema references, but it is not called by clinker run. Its warnings and coverage status cannot admit or reject execution, and an advisory result never overrides the planner.

StatusExact meaning
analyzedEvery applicable facet represented by the advisory model was inspected. This is not planner acceptance.
partialSome applicable content was inspected, but an unsupported shape or bounded limit left a gap. Read the attached reasons, then run the canonical planner check.
skippedThe artifact had no applicable advisory content, such as no external schema reference or no fields to inspect. This is not success.
failedThe artifact could not be inspected safely or structurally, such as a read, parse, or reference-resolution failure. Diagnose the reason and still use the planner as the execution authority.

Known advisory limits are explicit rather than silently treated as valid:

  • The advisory schema model covers linked external .schema.yaml metadata. Inline column lists, generated schemas, multi-record schemas, and planner-owned external schema shapes remain planner concerns; an applicable unsupported external shape reports partial.
  • Transform field scanning is a conservative heuristic. It does not reproduce CXL parsing, name resolution, schema flow, composition binding, or type checking; when linked schema content is otherwise analyzable, a transform therefore makes that advisory coverage partial.
  • Array element types and object shapes without declared child fields are not modeled completely. Format matching is also unavailable for fixed-width and SWIFT sources, and the advisory Parquet token is not a supported pipeline source format.
  • Analysis retains at most 10,000 field descriptors to depth 64, 4,096 schema references, and 1,024 reasons. Reaching a bound reports partial rather than dropping the gap.

After reading an advisory report, make the authoring decision against the canonical result:

clinker run pipeline.yaml --explain text
  1. Run clinker run pipeline.yaml --explain text for planner admission and inspect the compiled DAG.
  2. Use bare --dry-run only when a quiet canonical compile check is preferable.
  3. Run representative data against an isolated destination and inspect it.
  4. Run the full job only after checking the representative result and destination policy.

Explain Plans

The --explain flag prints the execution plan – the DAG of nodes, their connections, and the parallelism strategy the optimizer has chosen – without reading any data.

Text format

clinker run pipeline.yaml --explain
# or explicitly:
clinker run pipeline.yaml --explain text

The text format shows a human-readable summary of the execution plan:

Execution Plan: customer_etl
============================

Node 0: customers (Source, parallel: file-chunked)
  -> transform_1

Node 1: transform_1 (Transform, parallel: record)
  -> route_1

Node 2: route_1 (Route, parallel: record)
  -> [high] output_high
  -> [default] output_standard

Node 3: output_high (Sink, parallel: serial)

Node 4: output_standard (Sink, parallel: serial)

Key information shown:

  • Node index and name – the topological position in the DAG. Under dlq_granularity: document every Sink is listed after every other node, which is the order the run dispatches them in (see Document-level DLQ).
  • Node type – Source, Transform, Aggregate, Route, Merge, Sink, Composition
  • Parallelism strategy – how the optimizer plans to execute the node
  • Connections – downstream nodes, with port labels for route branches
  • Buffer class (Physical Properties section) – buffer: streaming for a node that hands its output straight to a single downstream consumer, or buffer: materialized for one that holds a whole stage’s output in an inter-stage buffer. See Streaming vs. Blocking Stages for the distinction.

The buffer class is a pre-runtime signal for memory pressure: a materialized node holds its rows against pipeline.memory.limit and may spill to disk once the budget is tight, while a streaming node holds only a small in-flight slice. Use the annotation alongside --memory-limit / pipeline.memory.limit to predict which stages will dominate memory before running the pipeline.

When a Sink declares sort_order, text output also includes a terminal writer decision:

=== Sink Writer Ordering ===

sink.export:
  terminal_order: customer_id asc, created_at desc
  disposition: deferred_sort
  boundary_mode: records_only
  partition_scope: global_split_sequence

disposition: proven_terminal_sort means the final upstream Sort already establishes the exact authored order; its name appears as proven_by. A deferred_sort is enforced over the complete writer population and can use the bounded spill path. boundary_mode names the physical write path. partition_scope says how wide the promise is: for example, global_split_sequence is one order across all numbered split files, while per_source_file is an independent order for each fan-out destination. The section is absent when no terminal order was authored.

Dead-letter output

When the pipeline has an error_handling.dlq block, the text output ends with a === Dead-Letter Output === section. It lists every DLQ file the run can write, with the header the compiled plan fixed for it, so you can check the columns before any data is read:

=== Dead-Letter Output ===

  rejects.csv
    sources: (pipeline-wide fallback)
    columns: _cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file, _cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field, _cxl_dlq_triggering_value, _cxl_dlq_error_category, _cxl_dlq_error_detail, _cxl_dlq_stage, _cxl_dlq_route, _cxl_dlq_trigger, order_id, order_total, _cxl_dlq_source_record
  refunds_rejects.csv
    sources: refunds
    columns: _cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file, _cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field, _cxl_dlq_triggering_value, _cxl_dlq_error_category, _cxl_dlq_error_detail, _cxl_dlq_stage, _cxl_dlq_route, _cxl_dlq_trigger, refund_id, refund_amount, reason, _cxl_dlq_source_record

Each entry gives:

  • The path that names the file, as written in the YAML.
  • sources: the Sources whose per_source.<name>.path routes to this file. The pipeline-wide path is labelled (pipeline-wide fallback): it takes the rows of every Source without a per_source path, and rows that carry no Source identity.
  • columns: the file’s complete header, in order. A row that lacks one of these columns writes an empty cell.

The files are listed with the pipeline-wide file first, then each per_source file in Source-name order. The section is absent when the pipeline has no dlq block. How the columns are chosen is described in Error Handling.

JSON format

clinker run pipeline.yaml --explain json

Standard output carries only the JSON document: plan warnings and other human diagnostics are written to stderr in this format and in dot, so redirecting stdout into a parser is safe. Read stderr as well if you want to see them – under --explain text they remain on stdout alongside the plan.

Produces a machine-readable JSON object for programmatic consumption. Useful for:

  • CI pipelines that need to assert plan properties
  • Custom dashboards that visualize execution plans
  • Diffing plans between config versions

An authored Sink order adds a writer_boundaries array. Each entry carries the Sink name, structured terminal_order fields and directions, the terminal_order_label, disposition, optional proven_by, boundary_mode, and partition_scope. The key is omitted when no Sink declares an order.

A pipeline with an error_handling.dlq block adds a dead_letter object with the same information as the text dead-letter section. Its buckets array has one entry per DLQ file, in the same order:

"dead_letter": {
  "buckets": [
    {
      "path": "rejects.csv",
      "sources": [],
      "fallback": true,
      "header": ["_cxl_dlq_id", "_cxl_dlq_trigger_id", "_cxl_dlq_timestamp", "...", "order_id", "order_total", "_cxl_dlq_source_record"]
    },
    {
      "path": "refunds_rejects.csv",
      "sources": ["refunds"],
      "fallback": false,
      "header": ["_cxl_dlq_id", "_cxl_dlq_trigger_id", "_cxl_dlq_timestamp", "...", "refund_id", "refund_amount", "reason", "_cxl_dlq_source_record"]
    }
  ]
}

sources lists the per_source names routed to the file. fallback is true for the pipeline-wide file. header is the complete header, including the _cxl_dlq_* columns (shortened to "..." above). The key is omitted when the pipeline has no dlq block.

# Compare plans before and after a config change
clinker run old.yaml --explain json > plan_old.json
clinker run new.yaml --explain json > plan_new.json
diff plan_old.json plan_new.json

Graphviz DOT format

clinker run pipeline.yaml --explain dot

Produces a Graphviz DOT graph. Pipe it to dot to render an image:

# PNG
clinker run pipeline.yaml --explain dot | dot -Tpng -o pipeline.png

# SVG (scalable, good for documentation)
clinker run pipeline.yaml --explain dot | dot -Tsvg -o pipeline.svg

# PDF
clinker run pipeline.yaml --explain dot | dot -Tpdf -o pipeline.pdf

This requires the graphviz package to be installed on the system.

The resulting diagram shows:

  • Nodes as labeled boxes with type and parallelism annotations
  • Edges as arrows with port labels where applicable
  • Branch/merge fan-out and fan-in structure
  • Terminal order, writer disposition, boundary mode, and partition scope on a Sink that declares sort_order

When to use explain

  • During development – verify the DAG shape matches your mental model before writing test data.
  • After adding route or merge nodes – confirm branch wiring is correct.
  • When tuning parallelism – check which strategy the optimizer selected for each node.
  • In code review – generate a DOT diagram and include it in the PR for visual confirmation.

Explain parses the YAML and builds the plan without opening runtime readers or processing records. Planning may inspect source metadata or matchers for cost estimates, but it does not create pipeline outputs.

clinker run pipeline.yaml --explain       # parse, compile, print the plan
clinker run pipeline.yaml --dry-run       # parse and compile without printing the plan

Both commands perform the same compile-time checks: schema binding, CXL type checking, DAG wiring, and plan-time source and output gates. --explain also renders the compiled plan; bare --dry-run is the quieter validation form. Neither command opens runtime readers, processes records, or creates pipeline outputs.

Retraction section

If at least one Aggregate has a group_by that omits a correlation-key field, the output includes a === Retraction === block. It lists which aggregates and windows use group-atomic retraction (see Correlation Keys) and a rough per-row memory estimate for each, so you can gauge the memory cost before a production run. The block is absent on pipelines that don’t use this mode.

Exact group sizes are unknown until the pipeline runs, so treat the estimates as a planning aid and confirm the live shape with clinker metrics collect after the first run.

Statistics

When the plan carries column statistics, the output ends with a === Statistics === section. Each figure is tagged with where it came from:

  • Row counts — an estimate per source. A [file metadata] figure is estimated from the input file’s size before any record is read; a [exec sketch] figure is an exact count measured during an actual run. These row counts are what the optimizer uses to pick a Combine’s join strategy.
  • Column sketches — distinct-value counts and frequent-value hints that a Combine gathers over its join keys while records flow, used to speed up matching.

A statistic that was never gathered renders as null rather than a fabricated zero — for example, a multi-file glob source or a network source whose size cannot be read adds no Statistics section at all.

Field provenance

clinker explain <pipeline> --field <path> traces where a single resolved value comes from across every configuration layer, printing the winning layer plus each shadowed layer and its source span. The path arity selects what is traced:

  • <node>.<param> (two parts) — a composition config parameter, resolved across composition defaults and channel/group overlays.
  • <source>.<column>.<attribute> (three parts) — a source-schema attribute (type, scale, precision, format, width, required, …), resolved across the schema-provenance layers Base < Pipeline < Group < Channel. Base is the source’s own declared schema:; the higher layers are the patch_schema overlay ops each channel/group applies.
# Where does the `scale` on the orders source's `amount` column come from?
clinker explain pipeline.yaml --field orders.amount.scale

# Resolve the same attribute with a channel overlay applied first.
clinker explain pipeline.yaml --field orders.amount.scale --channel acme_prod
Field: orders.amount.scale

  Resolved value: 2

  Provenance chain (outermost to innermost):
  [WON] Channel               →  2  (line 12)
        Pipeline              →  0  (shadowed)  (line 5)
        Base                  →  0  (shadowed)

The [WON] marker names the layer whose value survives; shadowed layers show what they proposed. An unknown source, column, or attribute is rejected with a hint listing the valid names at that level.

Reading a plan-time failure

A pipeline that fails a plan-time check never reads any input. The failure is printed before the run starts, and it carries four things:

E363

  × source "src": `record_path` "$.rows" starts with the JSONPath root marker
  │ `$.`, which is not part of the grammar; `record_path` is a dot-separated
  │ path of object keys, descended from the document root (for example
  │ `data.rows`). Write "rows" instead
   ╭─[pipeline.yaml:4:1]
 3 │ nodes:
 4 │   - type: source
   · ────────┬───────
   ·         ╰── declared here
 5 │     name: src
   ╰────
  help: `record_path` on a `json` source is a dot-separated path of object
        keys descended from the document root: no `$.` JSONPath root marker,
        no leading `/`, and no empty segments. It takes precedence over
        `format:`, so pair it with `format: object` or leave `format:` off.
        Omit `record_path` entirely and the reader auto-detects the document
        shape. Run `clinker explain --code E363` for the full grammar.
  • The code (E363) heads the report. Where a page exists for it, hand it to clinker explain --code for the worked example.
  • The message names the offending input and the rule it broke.
  • The source line is quoted from your YAML, with the offending node underlined.
  • The help: paragraph names the fix. When the gate does not already say so, a See: clinker explain --code <CODE> line is appended.

Warnings are reported the same way but marked ⚠ rather than ×, so an advisory is distinguishable from the diagnostic that stopped the run.

The same report is printed under --explain, which compiles the plan before printing it.

Two notes on where the snippet comes from:

  • A pipeline that pulls in a composition body is reported without the quoted source line. A plan-time diagnostic carries a line number but not which file it belongs to, so rather than risk underlining an unrelated line, the report gives the code, message and help alone.
  • A channel/group overlay suppresses the snippet only when it rewrites the compiled config through structural ops, source patches, or composition config: values. A selection that contributes only runtime vars leaves the pipeline document unchanged, so its snippet remains safe and is retained.
  • Bare --dry-run compiles the plan and prints the same report without reading source data.

Looking up diagnostic codes

clinker explain --list enumerates every registered diagnostic in stable code order. Each entry includes its code, severity, status, category, retryability, meaning, and correction. Closed enum values in this descriptor use lowercase kebab-case. Narrow the list with exact filters:

clinker explain --list
clinker explain --list --status retired-reserved
clinker explain --list --category source-and-expression

The status vocabulary is closed:

StatusMeaning
activeThe code describes a condition in the current authoring surface.
retired-reservedThe old condition is no longer accepted, but its identifier remains permanently reserved and still explains the paste-ready correction.

Categories are configuration, composition, source-and-expression, execution-and-format, terminal-authoring, security, and advisory. Unknown or empty filter values fail; a valid combination that matches no code also fails instead of printing an ambiguous empty result.

clinker explain --code <CODE> prints the same registry-owned descriptor as the list view, followed by a longer detail page when one exists:

clinker explain --code E15Y   # retraction-mode aggregate incompatible with strategy: streaming
clinker explain --code E376   # retired type: output spelling; use type: sink

Not every registered code has a longer page. A registered code without one is still valid and prints its complete descriptor plus Detail page: none; only a code absent from the registry is unknown. The See: clinker explain --code <CODE> line is appended to a diagnostic only when the longer page exists.

List and code discovery are static authoring metadata. They do not compile a pipeline, inspect records, or render runtime values or secrets. Use only the clinker explain spelling: there is no separate diagnostic command.

Column Lineage

The --lineage flag builds the pipeline’s column-level lineage – which source columns each output column is derived from, and which source columns influence the output as a whole – and writes it as OpenLineage events. Like --explain, it compiles the plan and exits without reading any data, so the lineage is derived statically from the pipeline definition.

# Write to a file
clinker run pipeline.yaml --lineage lineage.ndjson

# Write to stdout (pipe into other tooling)
clinker run pipeline.yaml --lineage -

There are two emission modes:

  • --lineage – a static, plan-derived export. It compiles the plan and exits without reading data, so it runs instantly and describes the pipeline’s lineage rather than a specific execution.
  • --lineage-events – live run-lifecycle emission. It runs the pipeline and emits a START when the run begins and a terminal COMPLETE / FAIL / ABORT when it ends, carrying real timing and row counts. See Live run events below.

Both modes share the same column-lineage facet and the same on-the-wire OpenLineage shape; the live mode wraps it in real run-lifecycle events. In external identity mode, complete events can also cross the independently bounded delivery worker described below. That worker owns only the selected file or stdout sink; it does not share the OTLP Collector worker or its memory arena.

Dataset identity preflight

Both flags require an explicit [observability.lineage] identity policy in the workspace clinker.toml. The default identity_mode = "external" requires one exact binding for every emitted Source and Sink node. A binding uses either a canonical datasource or a complete catalog namespace/name pair:

[observability.lineage]
identity_mode = "external"

[[observability.lineage.dataset]]
node = "source_customers"
canonical_datasource = "s3://warehouse/customers"

[[observability.lineage.dataset]]
node = "output_customers"
catalog_namespace = "analytics"
catalog_name = "customers_clean"

A source or output declared inside a composition body needs its own binding, keyed by the call site it belongs to:

[[observability.lineage.dataset]]
node = "enrich_orders.reference_prices"
canonical_datasource = "s3://warehouse/prices"

Body node names live in their own scope and may legally repeat a top-level name, so the key is <composition node>.<body source> rather than the bare name — two call sites of one body can be pointed at different files, and each gets its own identity.

. joins a call site to a body node, and a key never has to disambiguate that join from a node’s own name: a . in a node name is refused at plan time with E010, for every node kind and inside composition bodies too. node = "enrich.ref" therefore always addresses the source ref inside composition node enrich.

A \ that belongs to a node’s own name is written \\ in the key, so the key format stays unambiguous on its own rather than by relying on the naming rule. The same escape covers . — node = "enrich\\.ref" (in TOML, \\ is a literal backslash; the literal string 'enrich\.ref' says the same thing) would address a node whose own name is enrich.ref — but no pipeline the planner accepts can produce that key. Node names without \ — nearly all of them — are unaffected.

A node whose key cannot be written as a binding — over 128 bytes once the call site is joined to it — is refused by name, naming the limit. The correction is to rename the pipeline nodes the key is built from; there is no binding that can carry an over-long key.

# is reserved in an authored dataset name (catalog_name, or the name half of a canonical datasource) because it separates a multi-record source’s record types from their base dataset. Without the restriction a name like payments#detail would collide with record type detail of a source bound to payments, and the two would merge in the catalogue — attributing one dataset’s columns to the other. Namespaces are unaffected.

Clinker validates all required bindings before opening the lineage sink or, for --lineage-events, discovering sources and creating output attempts. Missing, duplicate, partial, ambiguous, or invalid bindings fail as observability.configuration.invalid; rejected values and physical paths are not copied into the diagnostic. The complete observability policy, including the required OTLP and authentication tables, is documented under lineage identity.

The destination file is emptied once the exporter has started, before any event is written. Point each run at a path you are willing to overwrite, and copy a record you want to keep before re-running against it. This applies to both --lineage and --lineage-events.

What that means for a run that produces no events:

  • Refused before the exporter starts — an invalid pipeline, a rejected configuration, a lineage binding that does not resolve — the file is untouched and still holds the previous run’s events.
  • Refused after the exporter starts, or fails before its first event, the file is empty. An empty file means this run wrote nothing; it never means the previous run’s result still stands.
  • A plan-only --lineage export that wrote nothing removes the destination, so no zero-byte artifact is left for a later step to publish. This applies only to a regular file: a destination that always reports zero length, such as /dev/null or a FIFO, is a successful export and is never removed. It is also skipped when the export ran out of flush time, because the exporter may still be writing.

If a consumer must distinguish “this run produced no lineage” from “an older run’s lineage is still here”, give each run its own destination path rather than relying on the state of a shared one.

Path-derived dataset names remain available only through the exact local compatibility spelling below. This mode is visibly labeled on stderr and is for local diagnostics, not external delivery:

[observability.lineage]
identity_mode = "local_diagnostic_paths"

Output format

The output is NDJSON (one JSON object per line) conforming to the OpenLineage 2-0-2 core spec. A run is described by a START event followed by a COMPLETE event that share one runId:

{"eventType":"START","run":{"runId":"019f030d-0b3e-7ee1-86ec-1bb5b4a2776b","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"namespace":"clinker","name":"audit_join","facets":{"clinker_pipeline":{"sourceHash":"7fd096a9..."},"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, ...}
{"eventType":"COMPLETE","run":{"runId":"019f030d-0b3e-7ee1-86ec-1bb5b4a2776b","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"namespace":"clinker","name":"audit_join","facets":{"clinker_pipeline":{"sourceHash":"7fd096a9..."},"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "inputs":[...], "outputs":[{"namespace":"analytics","name":"audit_report","facets":{"columnLineage":{ ... }}}]}
  • runId is a UUID v7 minted for this export and shared by both events. Because --lineage is a static, plan-derived export, the START/COMPLETE pair describes the pipeline’s lineage, not an executed data run — no rows are processed and the two events share one timestamp. A separate clinker run mints its own runId. (For real timing and row counts tied to an actual execution, use --lineage-events.)
  • Correlation is copied from one immutable CLI lifecycle snapshot: runId is the generated execution ID, the clinker-defined clinker_batch run facet carries the caller/generated batchId, and the clinker_semanticPlan job facet carries the effective fingerprint algorithm, schema version, and digest. Static and live events do not independently mint or parse these identities.
  • job.namespace is clinker; job.name is the pipeline name. The pipeline’s content hash rides in the clinker_pipeline job facet (sourceHash), not the job name – so the name stays stable across edits while runs of the same definition remain correlatable.
  • inputs are the source datasets; outputs are the sink datasets. External mode uses the exact configured canonical or catalog identities, so relocating a pipeline does not change its lineage graph. Explicit local_diagnostic_paths compatibility mode instead uses the file namespace with resolved paths (and falls back to the clinker namespace plus the node name for a network source).
  • The dataset namespace/name identifies the stable collection. A concrete logical partition or location would be emitted as the standard role-specific input/output subset facet — which rides under inputFacets on an input and outputFacets on an output, the positions its schema names, not under the dataset-level facets — and an explicitly authorized alias as the standard symlinks facet, which is a plain dataset facet and does ride under facets; neither is ever inferred from worker paths, attempt paths, hashes, or process context. No pipeline emits either facet today – the workspace config exposes no subset or symlink fields, so nothing can authorize one.
  • The columnLineage facet is attached to each output dataset on the COMPLETE event.

Reading the columnLineage facet

The facet has two parts, mirroring the OpenLineage ColumnLineageDatasetFacet:

"columnLineage": {
  "fields": {
    "amount": { "inputFields": [
      { "namespace":"file", "name":".../audit_orders.csv", "field":"amount",
        "transformations":[{"type":"DIRECT","subtype":"IDENTITY"}] }
    ]}
  },
  "dataset": [
    { "namespace":"file", "name":".../audit_orders.csv", "field":"order_id",
      "transformations":[{"type":"INDIRECT","subtype":"JOIN"}] }
  ]
}
  • fields – DIRECT (value-derivation) lineage, keyed per output column: the source columns each output column’s value is computed from. A rename (emit full = name), a multi-hop chain, or a path through a composition body (including nested compositions) collapses to the originating source column. A column whose value derives from an envelope read ($doc.<section>.<field>, bare / indexed / inside a larger expression) gets a DIRECT input field on the originating source dataset whose field is the rendered $doc.… path – so envelope-derived columns trace back to the document section they came from.
  • dataset – INDIRECT (influence) lineage for the dataset as a whole: source columns that shaped which rows exist, via filtering, joining, grouping, or sorting – collected once rather than duplicated across every column.

Each transformation carries a type (DIRECT / INDIRECT) and a subtype (IDENTITY, TRANSFORMATION, AGGREGATION, JOIN, GROUP_BY, FILTER, SORT, CONDITIONAL).

Multi-record sources

A multi-record flat file carries several record shapes in one physical file, discriminated by a lead record_type column. Record types differ in their columns, not in which rows they select, so each is treated as its own logical dataset rather than as a subset of one flat superset dataset:

  • Each record type is a dataset named <dataset>#<id> – the source’s bound dataset identity with the record type’s id as a # fragment. Under identity_mode = "external" that is the configured canonical or catalog identity (namespace s3://payments-lake, name raw/payments#detail), so no filesystem path enters the name; under local_diagnostic_paths it is the resolved file path (.../payments.txt#detail). Its columns are exactly that record type’s declared columns, so an output column that derives from a detail-record field traces to …#detail, and one from a header field traces to …#header.
  • A column declared by several record types (unified into one superset column) lists each owning #<id> dataset as an input field, so a derived output column traces to every record type it could have come from.
  • The engine-stamped record_type discriminator lead column belongs to the container rather than to any one record type, so it stays on the base dataset (no fragment) – a Route that branches on record_type still references {<base>, record_type}.
  • The run’s inputs list the base dataset followed by each #<id> record-type dataset, in record-type declaration order. Declaring them is load-bearing, not cosmetic: a lineage consumer resolves a columnLineage input field only against datasets the run declared as inputs, so a record-type dataset left out of inputs would have its column edges silently dropped on ingest.

A record type’s parent / join_key – the intra-file hierarchy linking a child record type to its parent – is not emitted as a lineage edge, since no plan operation performs that join.

Live run events

--lineage-events <PATH> runs the pipeline and emits OpenLineage run events tied to that actual execution, as NDJSON to a file path (or - for stdout):

clinker run pipeline.yaml --lineage-events events.ndjson

Unlike --lineage (which exits before reading data), this processes data, so it cannot be combined with --lineage, --explain, --dry-run, or -n.

Prefer a file path for a clean stream. With - (stdout), the run’s own stdout output — for example the per-stage spill-volume summary — interleaves with the event lines, so stdout is not pure NDJSON. Writing to a file keeps the events unmixed.

A run emits a START when it begins, then exactly one terminal event when it ends:

  • START – offered to the lineage path before the run body executes. In local diagnostic mode it is written synchronously; in external mode it is admitted non-blockingly to the bounded lineage queue and can be dropped under the configured policy. It carries the input and output datasets by identity plus the shared batch and semantic-plan correlation facets; no completed dataset facets exist yet.
  • COMPLETE – the run finished. It carries the input datasets and the output datasets with their columnLineage facets, exactly like the static export.
  • FAIL – the run errored. It carries the standard OpenLineage errorMessage run facet and the clinker-defined clinker_failure facet. Both are derived from the same bounded, sanitized classification used by machine supervision; the latter adds stable code, category, and retryAdvice fields.
  • ABORT – the run was interrupted (e.g. a SIGINT/SIGTERM shutdown) and drained what it could before unwinding.
{"eventType":"START","eventTime":"2026-07-03T17:00:00Z","run":{"runId":"019f...","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"facets":{"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "inputs":[...], "outputs":[...]}
{"eventType":"COMPLETE","eventTime":"2026-07-03T17:00:04Z","run":{"runId":"019f...","facets":{"clinker_batch":{"batchId":"batch-42"},"clinker_runStats":{"recordsRead":1000,"recordsWritten":970,"recordsDlq":30,"durationMs":4210}}},"job":{"facets":{"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "outputs":[{"...":"...","facets":{"columnLineage":{ ... }}}]}

Key differences from the static export:

  • runId is the run’s execution_id (a UUID v7) — the same identity used across clinker’s provenance sidecars and metrics spool, so an orchestrator can correlate the lineage events with the run’s other artifacts.
  • Both events carry the same clinker_batch run facet and clinker_semanticPlan job facet. Together with runId, their batch ID, execution ID, and semantic fingerprint tuple come from the run’s single CLI-owned lifecycle source and exactly match the optional machine stream when both are enabled.
  • The START and terminal events carry distinct eventTimes (run begin and run end), not one shared timestamp.
  • The terminal event carries a clinker_runStats run facet — a clinker-defined facet with recordsRead, recordsWritten, recordsDlq, and durationMs. Counts are pipeline-wide run totals, not per-output.
  • On FAIL, the run also carries the standard errorMessage run facet (ErrorMessageRunFacet 1-0-0) plus clinker_failure, with the shared sanitized message and stable failure code/category/retry advice.

Every started run that reaches a handled executor or publication boundary records one terminal snapshot; output-commit errors close as FAIL from that same source. Terminal emission is best-effort after the run’s authoritative publication decision. A lineage admission drop or sink write, flush, or deadline failure is reported on standard error and does not fail a run whose outputs already landed. A process crash can still leave only a delivered START event.

External delivery boundary

External identity mode is also the boundary for the independently bounded lineage worker. Each complete event is serialized under lineage.max_event_bytes and offered without blocking to a queue capped by lineage.queue_bytes; a full queue or oversized event drops the newest event. One synchronous worker owns the selected file or stdout sink, and shutdown waits no longer than lineage.flush_timeout_ms. Its typed outcome distinguishes normal shutdown, write failure (including permission errors), flush failure, and deadline expiry, with accepted, dropped, and full counters reported separately.

Within that deadline the worker stops taking new events off the queue halfway through, keeping the rest of the budget to finish the event it is already writing. A destination too slow to keep up therefore receives a file that is short — missing its last events — rather than one that ends inside a half-written record, which an NDJSON reader cannot parse. Nothing is added to the deadline: the whole flush still ends within lineage.flush_timeout_ms. A destination that stops accepting bytes altogether cannot be waited on, so in that one case the file may end mid-record, and the delivery outcome reports that separately from its counters. It has no access to the telemetry arena or Collector worker.

What the run prints

A delivery that lost events, ended on anything other than a normal shutdown, or left its destination inside a record prints one line on standard error:

clinker: lineage delivery outcome: status=deadline-exceeded error_kind=none accepted=2 dropped=0 full=0 records_complete=false

records_complete is the completeness of the file, where the counters beside it are the completeness of the export. It is true on every normal shutdown and on the slow-destination path above — a short file is still valid NDJSON — and false only when the worker was abandoned inside a write. That is the one state a consumer cannot determine for itself: a truncated NDJSON file simply ends, with nothing in it to say more was coming, and the counters cannot tell you either, because a run that gave up on a slow destination reports the same accepted total whether or not the last record made it out whole.

Being a condition rather than a count, it breaks the clean-run silence on its own: a run that dropped no events still prints this line if it left the destination unreadable. A normal shutdown that dropped nothing and ended on a record boundary prints nothing at all.

The plan-only --lineage export, whose whole invocation is the export, says the same thing in prose when it misses its flush deadline, and its correction follows from it. A short export ends “on a record boundary, short by the events that never got out” and is simply re-run. An export that “ends inside a record and is not readable as NDJSON” must be discarded rather than published as this run’s lineage, and only then re-run.

A sink write or flush failure leaves the same two files, and says so the same way. Where the deadline path ran out of time on an export that was otherwise going fine, here the destination itself refused, so two separate facts are reported: what the destination was left holding, and where the retry should point. An export that “ends on a record boundary and is readable as NDJSON” is reported without a disposal instruction; one that “ends inside a record and is not readable as NDJSON” must be discarded first, and the correction says so ahead of the retry advice. The retry advice itself is unchanged — a permanent refusal (permission denied, read-only filesystem, a directory) asks for a different destination, anything else asks for a re-run — because re-running against a destination that has just refused a write may refuse it again.

Nothing is said about a destination that is not there: a failure that wrote no bytes at all leaves an empty file, which this path removes so a publish step cannot upload it as the run’s lineage, and the diagnostic then describes no file.

This external worker is not a second identity mode and does not make local paths suitable as catalog identity. The explicit local_diagnostic_paths mode remains a synchronous compatibility path for local file or console inspection and cannot enter the external delivery worker.

Lineage uses logical dataset bindings, not working-directory or attempt paths. It copies the batch ID, execution ID, semantic fingerprint, and terminal facts from the same immutable lifecycle snapshot used by machine supervision and OTLP, while keeping its delivery result independent. A lineage fault cannot change final or DLQ bytes, exit status, machine terminal payload, publication inventory, visible final set, or retained failed-attempt evidence. Event values also remain outside this identity-only payload; telemetry field policy is enforced before Collector queue admission. If [observability] is absent, neither the lineage worker nor the Collector worker exists.

When to use

  • Impact analysis – before changing a source schema, see which outputs and columns depend on it.
  • Auditing & governance – feed the OpenLineage events into a catalog (e.g. Marquez) to track data provenance.
  • Review – attach the lineage of a new pipeline to a PR to confirm the intended derivations.

Because --lineage reads no data, it runs instantly and works on a pipeline whose inputs do not yet exist.

Limitations

Lineage is derived from the compiled plan, so a few constructs are approximated:

  • A column-grain $doc read is traced as DIRECT lineage (see fields above) in a transform projection, a combine body, a composition body, and an aggregate emit, attributed only to a source whose envelope declares the section. A $doc read in an influence predicate – a route condition, a cull drop_group_when, or a combine where – is surfaced as INDIRECT influence (FILTER for route and cull, JOIN for combine). Two $doc cases remain uncovered: a whole-section envelope echo (an output header/footer regenerated from a source document section, with no output column or expression); and any $doc reference in a Reshape rule, which the compiler rejects outright (Reshape re-runs its rules after a per-group spill that drops envelope context), so there is no Reshape envelope lineage to produce.
  • A match: collect combine declared without a projection body produces coarse column lineage: each collected column derives (as TRANSFORMATION) from every build-side column, because there is no body expression to pin the exact source column.
  • INDIRECT influence covers route/cull predicates, join keys, aggregate grouping, and correlation sort over record columns (and $doc envelope terms in route/cull/combine predicates, as above). An aggregate’s pre-aggregation row filter, a transform-inline filter, and Reshape order_by / partition_by are not (yet) attributed as influence.
  • Constant and count(*) columns (which have no source input) are omitted from fields; engine-stamped columns ($ck.*, $meta.*, $source.*) are skipped, mirroring the default writer.

Memory Tuning

Clinker is designed to be a good neighbor on shared servers. Rather than consuming all available memory, it works within a configurable budget and reaches for back-pressure or disk spill before it runs out.

What the budget measures

Memory attribution and process RSS measure different things. Clinker samples RSS to detect pressure, while individual operators report the data they retain. Allocator overhead, thread stacks, native I/O workspace and startup allocations also contribute to RSS. Setting a budget therefore does not promise that the process’s resident size will always equal its accounted data size.

CSV input decoding and output preparation use the run’s finite resource budget. Growing decoded cells, owned record values and retained document metadata carry their accounting with them until the last owner releases the allocation. Output policy, schema mappings, captured headers and prepared operation storage are also admitted before allocation. Replacing a buffer accounts for both old and new storage while they overlap; sharing a value does not release its charge.

An operation that cannot obtain its working memory fails as a resource error. CSV does not silently truncate a large cell, impose a separate authored cell-size limit, or treat resource refusal as a record eligible for the DLQ. An explicitly configured spill location can hold prepared output bytes, but cell rendering and metadata still require memory. See output preparation.

This is not a whole-process allocation or constant-memory guarantee. The CSV parser’s raw buffers and the intermediate JSON tree used to parse JSON-encoded cells remain outside this admission boundary. Unchanged format readers and later legacy record copies also have separate or incomplete accounting. Source and Combine paths can still retain whole inputs; input-sized residency remains tracked in #1183. A successful small fixture does not establish that a larger input fits. The existing YAML and CLI tuning controls below remain the controls for this budget.

The memory: block

All pipeline-level memory tuning lives under a single optional block:

pipeline:
  name: my_pipeline
  memory:
    limit: "1G"          # optional — defaults to 512M
    backpressure: pause  # optional — defaults to pause

The entire block is optional. A pipeline with no opinions about memory writes nothing:

pipeline:
  name: my_pipeline

…and gets the runtime defaults (512 MB hard limit, backpressure: pause).

Individual fields are also optional. Setting just one is fine:

pipeline:
  name: my_pipeline
  memory:
    limit: "2G"

Setting the memory limit

CLI flag (highest priority):

clinker run pipeline.yaml --memory-limit 512M

YAML config:

pipeline:
  memory:
    limit: "512M"

When --memory-limit is passed it overrides pipeline.memory.limit for that run; omit the flag and the YAML value applies unchanged, falling back to the 512 MB default only when neither is set. An empty or whitespace-only flag value — as an ops wrapper produces when it forwards an unset variable, so --memory-limit "$CLINKER_MEM" expands to --memory-limit "" — is treated exactly like omitting the flag: the YAML value (or default) applies, rather than the run aborting. Suffixes are binary (1024-based): K = 1024 bytes, M = 1024², G = 1024³; a bare integer is bytes. (This differs from the decimal KB/MB/GB used by min_size/max_size, which are 1000-based.)

Default: 512 MB.

Invalid values: the two entry points treat a malformed limit differently. An empty or unparseable memory.limit in the YAML (for example a stray non-numeric value) falls back to the 512 MB default. A non-empty malformed --memory-limit flag, by contrast, is rejected up front with a config error that names --memory-limit and echoes the value you passed — so a typo such as the decimal 4GB (the binary suffix is 4G) fails loudly instead of silently collapsing to the default and shrinking a larger budget set in your YAML. (An empty or whitespace-only flag value is not malformed: it is treated as if the flag were omitted, as noted above.) Either way, a value whose size is well-formed but too large to represent — its scaled byte count exceeds the maximum a 64-bit counter can hold — is rejected rather than wrapping to a small budget (the YAML overflow error names memory.limit; the flag overflow error names --memory-limit). Pick a limit that fits your host’s real memory.

A well-formed but undersized value is a different case: it is not a malformed flag, so it clears the boundary check, and whether it aborts the run depends on the backpressure policy. Under a producer-pausing policy (pause, the default, or both) a value below the process’s baseline resident memory is rejected at startup by the budget gate as E312. Under the non-pausing spill policy that startup gate does not fire: the run proceeds and relies on spilling to stay within the budget rather than aborting. Because --memory-limit simply populates pipeline.memory.limit, that E312 — which names the limit and echoes the offending byte value — refers to the same limit you passed via the flag.

Choosing a backpressure policy

When memory use approaches the limit (the soft threshold is 80 % of limit), something has to give up memory. The backpressure knob chooses what:

ValueBehavior
pause (default)Where possible, pause an upstream reader so it stops producing until pressure eases; when a paused reader is about to be needed, first spill downstream state and then proceed, so a pause never stalls the run.
spillNever pause a producer — always free memory by spilling a stage to disk.
bothPause where possible, otherwise spill whichever stage is holding the most memory.

pause is the right default for most pipelines: pausing a fast Source feeding a slow downstream stage is cheaper than writing its buffered records to disk. Reach for spill or both only when you have a specific reason to prefer a different posture — for example, both when one large stage dominates the budget and you want it spilled first.

How pause and resume work

Under pause (and both), a producer paused because memory crossed the soft threshold is resumed automatically once memory recedes — it is never left parked. Pause and resume use two watermarks to avoid flapping (a hysteresis band):

  • Pause when live memory rises above the soft threshold, 0.80 × limit.
  • Resume when live memory falls back below the lower resume watermark, resume_threshold × limit (default 0.70 × limit).

Between the two watermarks nothing changes, so a normal batch-to-batch swing in memory cannot make a producer flap between paused and resumed on every poll.

A paused reader also never blocks the run. When the engine reaches a stage that needs a paused reader’s records, it first sheds reclaimable downstream state to disk and then resumes the reader and proceeds — so pause throttles producers under pressure but degrades to spill-and-continue at the point it would otherwise wait, rather than stalling.

resume_threshold

resume_threshold tunes the low watermark of that band, as a fraction of the hard limit:

pipeline:
  memory:
    limit: "1G"
    resume_threshold: 0.65   # optional — defaults to 0.70
  • Default: 0.70 (omit the field to take it).
  • Valid range: strictly greater than 0 and strictly less than the 0.80 soft threshold, so the resume point always sits below the pause point. A value outside (0, 0.80) — including 0.0, a negative, or anything ≥ 0.80 — is rejected at plan time with E324. A misspelled key is rejected as an unknown field.
  • Lower widens the band: a paused producer stays paused longer (smoother, but slower to re-open).
  • Higher narrows the band: faster to re-open, but more prone to flapping.

Only the pausing policies (pause, both) use this watermark; under spill no producer is ever paused, so it has no effect.

Streaming batch size (batch_size)

pipeline.batch_size sets how many events (records plus document-boundary punctuations) a streaming-eligible stage hands off to its downstream consumer at a time over a back-pressured channel. For a fused stage (Source → Transform → Sink, Merge.interleave of Sources) it bounds the in-flight working set to one batch rather than the whole stage, because the stage pulls records off a live upstream channel without ever building a full result. The other streaming stages build their full result first and stream it in batches; there the knob sizes only the inter-stage slice, not the producer’s footprint. The knob is optional; omit it to use the built-in default of 2048 events. See Streaming vs. Blocking Stages for the distinction.

pipeline:
  name: orders_rollup
  batch_size: 1024          # optional; default 2048

A per-transform override is available on a Transform’s config.batch_size (see Transform Nodes); it takes precedence over the pipeline value for that one stage. A batch_size of 0 is rejected at config load. The knob affects only the memory profile of streaming stages, never their output — blocking stages (sort, hash Aggregate, Combine build side) ignore it and continue to fully materialize. See Streaming vs. Blocking Stages for the full model.

Behavior under memory pressure

You don’t manage memory by hand — the engine does it within the budget you set. What this means in practice:

  • Spillable stages always complete if disk space is available, regardless of input size. When a blocking stage (sort, hash Aggregate, a grace hash or range Combine) outgrows the budget, it spills to disk instead of failing.

    • An in-memory equal-ids Combine does not spill. The planner picks the in-memory hash join or the disk-spilling grace hash join from its size estimate before the run starts. If an in-memory join’s build side outgrows the budget, the run stops with E310 MemoryBudgetExceeded (#1337). Set strategy: grace_hash on a Combine whose build side may not fit; see Combine Nodes.
    • Range and equality+range Combine also spill. A Combine whose where: joins the two inputs on an inequality (<, <=, >, >=) — a pure band join such as orders.amount >= bands.lo and orders.amount < bands.hi, or an equality-plus-range join such as orders.region == bands.region and orders.amount >= bands.lo — runs the block-band strategy, which is spill-bounded on both axes: it external-sorts each input side to disk and accumulates its matched output in a spillable sort as well, so it completes on inputs and result sizes larger than the budget. When an equality key is present, records are grouped by that key (via its hash) before the range walk and only same-key pairs are joined; a single very common key is spread across many disk-backed blocks and joined a bounded pair at a time, so even a heavily skewed key stays within the budget. The join still aborts with E310 MemoryBudgetExceeded only as a last resort — when a single indivisible unit of work (one pair of input blocks plus the join’s scratch arrays) cannot fit the hard limit even on its own. If you hit that, give the pipeline more headroom or narrow the predicate.
  • Performance degrades gracefully. Under pressure you’ll see slower execution and possibly disk I/O — not a crash.

  • The limit is a soft ceiling, not a hard wall. Momentary spikes may briefly exceed it before the engine reacts. Only if memory blows past the limit outright does the run abort with E310 MemoryBudgetExceeded, which names the stage that overran.

  • Shared handoff buffers spill and re-scan sequentially. A stage that fans out to several consumers, or feeds a composition input port, records the exact reader count for each output port and may spill that shared slot like any other materialized buffer. Readers run one at a time over immutable backing; spilled and mixed buffers open one file at a time for each scan, and the final reader takes the authoritative slot. Clinker never pre-forks one copy per destination. A consumer that must collect the scan into a full resident vector reserves that materialization before allocating it. If the projected overlap exceeds the hard limit, it aborts with E310 MemoryBudgetExceeded, names that consumer, and reports NodeBuffer as the budget category. More spill space can keep the shared backing off-heap, but it cannot eliminate an individual operator’s required resident working set.

  • One oversized correlation group trips E310 in Reshape and Cull. Both apply their rules to a whole group at once, so a group must fit the budget at finalize even though cross-group and ingest-time peaks spill. A group larger than memory.limit has no in-budget representation, so the run aborts with E310 MemoryBudgetExceeded naming the node and the offending partition_by group. Raising memory.limit clear of the reported figure is the only fix that leaves your output unchanged (finalize also holds the run’s other groups, so that figure is a floor, not a target). Dropping unread columns in an upstream Transform also shrinks the group, but both nodes write every input column through, so those columns leave the output too. Do not narrow partition_by to clear it — that key defines the group the rules evaluate over, so a narrower key makes the run succeed by changing your results. Both nodes also hold per-group bookkeeping that is O(distinct groups) and cannot spill — the spill path evicts buffered records, never group entries — but they handle an extreme partition cardinality differently, and the difference is the diagnostic you get:

    • Cull fails loud. Its in-memory drop-decision state is combined with the run’s other live charged memory and checked against the budget on every admission, so a cardinality that would breach memory.limit aborts with E310 MemoryBudgetExceeded naming the decision state instead of a group.
    • Reshape has no such gate. Its group map and group-order list grow with the distinct-group count unchecked, so a high-cardinality Reshape can be OOM-killed rather than reporting E310; the resulting status is platform-dependent. Keep the group count well within the budget on a Reshape; there is no engine check to catch it for you. This gap is tracked in #1027.
  • clinker explain --code E310 covers every surface. The rendered diagnostic names which one overran in its [...] detail; the explain page keys its remediation to that.

  • A Reshape partition value it cannot key is folded into the null group, silently. Where Cull aborts on an array or map partition_by value, Reshape groups that record under null — alongside every record whose partition column is missing, explicitly null, or an empty string — and the run exits 0. A NaN is not among them: in both nodes every NaN is one group of its own. Nothing in the output marks the merge, so if any of those shapes can occur in your partition column, normalize or filter it in an upstream Transform. See Values Reshape cannot key.

  • These are limits, not bugs. Every case above is an ordinary operational condition with a documented code, so it reports through E310 rather than as an internal error. Four conditions do still present as an internal error today despite being about your data or your host, not an engine defect:

    • an array or map value in a Cull partition_by column (Reshape folds these into the null group instead, per the bullet above);
    • a runtime data error inside a Cull drop_group_when expression;
    • a runtime data error inside a Reshape rule’s when, mutate.set, or synthesize expression — the direct analogue of the Cull case above. Like it, this aborts the run rather than routing the record to the dead-letter queue, even under strategy: continue;
    • a spill read/write failure in a Reshape or Cull group buffer, including the spill volume filling up; #1021 tracks preserving its spill diagnostic and exit classification.

    Any other internal error is a Clinker defect worth reporting.

Some stages stream (they hold only a small in-flight slice of records) and some materialize (they hold a whole stage’s worth before emitting). clinker run --explain annotates each node with buffer: streaming or buffer: materialized so you can see which stages will dominate the budget before you run. See Streaming vs. Blocking Stages for which is which.

Dead-letter output

Dead-lettered rows are not collected in memory until the run ends. Each row is formatted under its DLQ file’s header, which is fixed when the pipeline compiles, and written straight into a staged copy of that file. Each open DLQ file costs one fixed 64 KiB write buffer, and the number of DLQ files is fixed by the pipeline. Apart from the held failures described below, a run whose every row fails therefore uses about as much memory for dead letters as a run where none do. That is why the DLQ writers are not charged to the memory budget: they do not grow with input.

What does grow with failures is the DLQ files themselves, and they are bounded by disk: free space at the staging location, and the publication attempt’s byte ceiling (storage.publication.max_attempt_bytes). To stop a run before it produces a large DLQ, set a breaker: dlq.max_rate and dlq.per_source.<name>.max_rate (E315/E316), or type_error_threshold (E368). See How DLQ output is written, Bounding how much can dead-letter and How the DLQ columns are chosen.

Some failures are held in memory, uncharged, until the stage that found them finishes, and are written then: join_values collisions at a Sink that writes on its own thread, Aggregate add_record failures found on the Aggregate’s input thread, and Combine output-row failures found on the Combine’s driver thread or inside a grace-hash, sort-merge or IEJoin join. Records a correlation key holds until their group is decided are that feature’s own state, not DLQ output.

Under dlq_granularity: document three things are held, all charged to the memory budget:

  • Each Sink holds every open document’s records in a buffer that spills to disk when the budget needs the memory, until the document’s verdict is final.
  • A failed document’s failing records are held as their dead-letter rows until the document is rejected. They move to one file in the spill directory when the budget needs the memory, and count toward storage.spill.disk_cap_bytes (E320). If one more held row would not fit even with every held row on disk, the run fails with E310.
  • For each rejected document, a compressed record of the rows already written, so a row several Sinks held is written once. It never spills; if it would pass the limit once every held row is on disk, the run fails with E310.

See Document-level DLQ and, for what each of these costs, How DLQ output is written.

Sizing guidelines

WorkloadRecommended limitNotes
Small files (<10 MB)128MMinimal memory pressure
Medium files (10–50 MB)256MCovers most ETL jobs
Large files or complex aggregations512M (default) – 1GMultiple group-by keys, large cardinality
Multiple large group-by keys1G+High-cardinality distinct values

Target workload: Clinker is optimized for 1–5 input files of up to 100 MB each, processing 10K–2M records per run.

Aggregation strategy interaction

Memory consumption depends heavily on the aggregation strategy the optimizer selects:

  • Hash aggregation accumulates state in a hash map. Memory usage is proportional to the number of distinct group-by values. With high-cardinality keys, this can consume significant memory before spill triggers.

  • Streaming aggregation processes groups in order and emits results as each group completes. Memory usage is minimal (proportional to a single group’s state) but requires the input to be sorted by the group-by keys.

  • strategy: auto (the default) lets the optimizer choose based on the declared sort order of the input. If the data arrives sorted by the group-by keys, streaming aggregation is selected automatically.

To influence strategy selection:

  - type: aggregate
    name: rollup
    input: sorted_data
    config:
      group_by: [department]
      strategy: streaming    # force streaming (input MUST be sorted)
      cxl: |
        emit total = sum(amount)

Only force streaming when you are certain the input is sorted by the group-by keys. If the data is not sorted, results will be incorrect. Use auto when in doubt.

Oversized single rows

An aggregate that keeps min, max, or another value-buffering binding holds each contributing row’s raw values until the group finalizes. If a single input row’s buffered footprint is larger than the entire memory.limit, no amount of spilling can hold it — spill would only re-read the same oversized row. The engine surfaces this per-row overflow rather than absorbing it:

  • With error_handling.strategy: fail_fast (the default), the run aborts with E310 MemoryBudgetExceeded, naming the aggregate stage and reporting the offending row’s byte footprint against the budget.
  • With strategy: continue, the offending record is routed to the dead-letter queue under the aggregate_finalize category, and the run proceeds.

This is almost always a sign the budget is set far too low for the record shape — raise memory.limit so a typical row fits comfortably.

Compositions

A composition (a reusable sub-pipeline included via use:) does not get its own memory budget — its operators share the parent pipeline’s budget and spill to the same temporary directory. A spilled composition input is scanned sequentially, and its materialization stays charged continuously as ownership moves into the body; it is not briefly dropped from the accounting or charged twice. During body-input schema conversion, the old input and new records are both charged for the short interval when both allocations exist. If that materialization would exceed the hard limit, E310 names the composition call-site directly. If a later budget overrun happens inside the composition, the error names that same call-site (e.g. enrich_call) so you can locate it, prefixing the message with in composition "enrich_call": ... when the overrun is internal to the body.

Monitoring memory usage

Use the metrics system to track peak_rss_bytes across runs:

clinker run pipeline.yaml --metrics-spool-dir ./metrics/

The metrics file includes peak_rss_bytes, which shows the maximum resident memory during execution. If this consistently approaches your memory limit, consider increasing the budget or restructuring the pipeline to reduce intermediate state.

Shared server considerations

On servers running JVM applications, memory is often at a premium. Recommendations:

  • Set --memory-limit or memory.limit explicitly rather than relying on the default. Know your budget.
  • Use --threads to limit CPU contention alongside memory limits.
  • Monitor peak_rss_bytes in production metrics to right-size the limit over time.
  • Schedule large pipelines during off-peak hours when JVM heap pressure is lower.

Storage & Spill Location

Blocking operators — Aggregate, sort, and grace-hash Combine — accumulate state in memory up to the configured budget, then spill to disk when a soft or hard memory threshold trips, rather than running the process out of memory. By default those spill files land in the operating system’s temporary directory. The [storage] block in clinker.toml lets you redirect them.

Output preparation

CSV, JSON, XML, fixed-width and SWIFT output in the CLI and executor prepares each complete output operation before delivering its bytes. For CSV, the first body row and its automatic header share one operation; explicit document start and end are separate operations. The same finite-resource preparation API is available to library integrations. Fixed-width prepares document headers, body records and footers separately. SWIFT prepares its service headers together with the first body record, then prepares the closing block and trailer at finalization.

Prepared output bytes stay in memory unless storage.spill.dir supplies an explicit spill location. This differs from the operator spill default described below: output preparation does not silently use the operating system’s temporary directory. Configured spill uses the run’s disk budget and a finite descriptor allowance. It does not remove the memory required for a rendered CSV cell, retained header, format configuration and schema plans, fixed-width warning history, or a SWIFT document trailer. Fixed-width retains every committed warning; resource refusal fails the next operation rather than discarding history. See Memory Tuning.

Failure before delivery writes none of that operation’s bytes. Once delivery starts, ordinary I/O can accept a prefix before failing. The writer then refuses further operations, including flush, so the original failure is not hidden by later calls. Prepared bytes and accepted destination bytes are different counts; a partly delivered operation does not become a committed record.

Temporary files remain charged until their removal is confirmed, including when cleanup must be retried. Dropping a handle or requesting cancellation does not turn failed cleanup into free disk capacity. Resource telemetry can be dropped when its fixed arena is full; admission, output and cleanup do not depend on those signals being retained.

These rules add no storage setting and make no atomic-publication promise for an arbitrary destination. Output publication governs the separate file-publication boundary. EDIFACT, X12 and HL7 retain their existing writer paths. These five codecs do not establish all-format migration or admission of existing reader/parser allocations.

The [storage] block

Storage settings are a property of the workspace, not of an individual pipeline, so they live in clinker.toml at the workspace root rather than in the per-pipeline YAML:

[storage.spill]
dir = "/var/clinker/spill"   # optional; operator default = OS temp dir
disk_cap_bytes = "10GB"      # optional; default = unlimited
compress = "auto"            # optional; auto | off | on   (default = auto)

[storage.staging]
enabled  = false             # opt-in; default off
dir      = "/var/clinker/staging"   # required when enabled
patterns = ["/mnt/nfs/data/**"]     # which sources to stage

[storage.publication]
mode = "direct"                    # direct | local_then_publish
destination_profile = "local"     # local | nfs_v4_1 | smb_3_1_1
failed_retention_seconds = 86400   # 24 hours; zero is allowed

The whole block is optional. With no clinker.toml, or a clinker.toml that omits [storage], blocking operators spill to the OS temp directory. Prepared output preparation stays in memory unless storage.spill.dir is set.

Table names are checked

Clinker reads five top-level tables — [catalog], [storage], [observability], [channel], and [group] — and passes over any other top-level table, so a clinker.toml may carry tables meant for other tooling.

A table name that is a misspelling of one of those five is refused instead, naming the table you wrote and the one it was mistaken for. Absence is how a workspace says “off” — omitting [observability] disables telemetry — so a name that misses by a letter would otherwise turn a policy off silently:

clinker.toml table [observabilty] is not a table clinker reads, and is a
misspelling of [observability]; nothing under it would have been applied —
write `[observability]`, or rename the table so it is not mistakable for one

Keys inside a recognized table are strict already: an unknown key there is refused outright.

Output publication and retained attempts

[storage.publication] is the only author-facing block for output publication, destination qualification, retained attempts, and bounded cleanup. It is optional; the defaults use destination-local quarantine (mode = "direct"), the local filesystem profile, and 24-hour retention for failed attempts. NFS and SMB profiles are accepted only when the operating system identifies the matching filesystem family; an unidentifiable remote drive is rejected rather than guessed.

[storage.publication]
mode = "direct"
destination_profile = "local"
failed_retention_seconds = 86400
creation_grace_seconds = 300
max_attempt_bytes = "4GB"
retained_byte_limit = "8GB"
retained_attempt_limit = 8
min_free_bytes = "2GB"
sweep_entry_limit = 1000
sweep_byte_limit = "8GB"
sweep_time_limit_ms = 2000
SettingDefaultHard limitMeaning
failed_retention_seconds86,400 (24 hours)604,800 (7 days)How long incomplete, abandoned, or otherwise failed terminal attempts remain. Zero is valid and makes them immediately eligible after the live lock is released.
creation_grace_seconds3003,600Grace period before a non-terminal attempt can be considered abandoned.
max_attempt_bytes4 GB16 GBMaximum admitted estimate for one publication attempt.
retained_byte_limit8 GB64 GBAggregate retained-attempt byte ceiling used by policy.
retained_attempt_limit8128Aggregate retained-attempt count ceiling.
min_free_bytes2 GB64 GBAdditional free-space headroom required by admission.
sweep_entry_limit1,00010,000Maximum directory entries considered by one cleanup page.
sweep_byte_limit8 GB64 GBMaximum regular-file bytes considered by one cleanup page. Must be at least max_attempt_bytes + 4,194,304B so one maximum attempt and its bounded manifest can always make progress.
sweep_time_limit_ms2,00030,000Maximum monotonic elapsed time for one cleanup page.

The capacity observation is advisory. It is a one-time comparison of the attempt estimate plus min_free_bytes against observed free space; it reserves no blocks or quota. A later write or synchronization can still fail with ENOSPC or EDQUOT, and Clinker retains exact attempt state rather than claiming publication succeeded.

Configuration resolution rejects a sweep budget that cannot inspect one maximum-sized attempt plus the bounded 4 MiB manifest. The diagnostic reports the exact minimum and a paste-ready sweep_byte_limit setting; this prevents a valid admitted attempt from becoming permanently too large for cleanup.

Aggregate admission counts physical manifest-owned staging and quarantine files across every destination and local-spool root. Temporary local and destination copies both count while both exist; an uninspectable size fails admission instead of being treated as zero. Count and byte inventory, expiry cleanup, the limit check, and attempt-root creation are serialized across the same root set, so concurrent runs cannot both consume the final retained slot. Before releasing that serialization boundary, Clinker records the admitted estimate in every owned attempt root. A later process charges at least one copy of that reservation for the execution until exact artifact sizes replace it. Admission locks live only inside the internal .clinker-attempts namespace, so an authored output leaf cannot replace the mutex. Lowering retained_attempt_limit does not hide attempts that were admitted by an earlier configuration. Listing and expired purging continue to page over the bounded physical namespace and report policy debt until the retained count is back within the new limit.

Publication modes and destination profiles

direct writes each artifact into restrictive quarantine on its destination filesystem, synchronizes it, and only then promotes it to the final leaf. local_then_publish requires local_spool_dir; Clinker writes and verifies the local copy, copies and verifies it into destination-local quarantine, and promotes only that destination copy. There is no copy fallback that writes directly to a visible final, and no automatic mode fallback. A destination profile mismatch fails closed.

The supported destination profile is a qualification claim, not a spelling that makes an arbitrary mount safe. NFSv4.1 and SMB3.1.1 support requires an actual mounted destination qualified with the selected profile, including real late-failure behavior. Hosted qualification also runs the ordinary publication admission API from independent processes against each mounted profile, in opposite multi-root order, and requires exactly one process to enter the final retained-count and retained-byte slot without deadlock. This production lock proof is separate from the filesystem’s byte-range/OFD lock probe. Release evidence for low space must observe a real ENOSPC. Injected EDQUOT is useful seam coverage but is non-qualifying unless a real quota is provisioned and the filesystem reports EDQUOT during the mounted test.

Publication truth is per artifact. Earlier artifacts can be synchronized and visible while a later promotion fails, so an output set is never described as atomically published across artifacts. Quarantine ownership is scoped to one execution ID; attempts never share writable staging ownership.

Retention and metadata-last cleanup

Every retained attempt has a restrictive owner manifest and live lock. Cleanup accepts only roots derived from a freshly compiled workspace-relative pipeline, canonical execution IDs or the typed expired selector, and bound opaque continuations. It never accepts a raw deletion path.

Cleanup keeps data on any ambiguity. It opens entries through retained, no-follow handles, verifies the supported manifest and allowed children, acquires the live lock, checks wall-clock eligibility, and stays within the configured entry, byte, and monotonic-time limits. Owned artifact bytes are removed first. Clinker then revalidates the directory, removes the owner manifest and liveness metadata last, and removes the now-empty attempt root. A crash or refusal before that sequence finishes remains explicit cleanup debt for a bounded retry.

Use the non-mutating operator surface to inspect or preview retained state:

clinker attempts list pipelines/orders.yaml
clinker attempts inspect pipelines/orders.yaml \
  --execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159
clinker attempts purge pipelines/orders.yaml --expired
clinker attempts purge pipelines/orders.yaml --expired --execute

Repeat any path or overlay identity used by the run (--base-dir, absolute-path permission, rules root, channel/groups, --path-execution-id, batch identity, or timestamp) so the command recompiles the same typed owned roots. File-backed fan-out roots normally replay source discovery. A retained failure also keeps a bounded, plan-bound receipt containing the logical source identities used by {source_file} and {source_path} plus path-free identifiers for every owned output or spool root. If a source is later removed or its directory is renamed, attempt commands re-render those authored templates and require the resulting validated roots to match the receipt exactly before listing, inspection, or purge. Successful publication removes the receipt with the rest of the attempt. Attempt commands do not accept a raw cleanup path. Continuation tokens are opaque raw values; JSON output provides authoritative structured recovery/resume argument arrays for shell-independent automation.

Output is path-free by default and reports logical root, execution, and artifact IDs. --show-paths adds only sanitized workspace-relative paths; machine-local prefixes, sensitive-looking components, credentials, secrets, record values, and raw staging detail remain redacted. Successful operations exit 0. Invalid selectors, pipelines, and continuations exit 1. Bounded partial work or any cleanup debt exits 4 with E371/E372 retry and workspace-relative recovery guidance. See the attempt command reference and exit-code contract.

storage.spill.dir — where spill files go

When dir is set, the per-run spill directory (clinker-spill-<random>/) is created under that path, and every blocking operator writes its spill files there. When dir is omitted, the per-run directory is created under the OS temp directory (std::env::temp_dir, typically $TMPDIR or /tmp).

The directory is validated once at startup, before any input is read. If the path does not exist, is a file, or is not writable, the run fails immediately with a diagnostic naming the setting:

storage.spill.dir /var/clinker/spill does not exist; create it or point at an existing volume

Validating up front — rather than at the first spill — means a misconfigured spill volume fails fast, while the run is cheap to abandon, instead of after minutes of work. (This is the trap DuckDB fell into when its temp-directory setting was honored only lazily, duckdb/duckdb#9401.)

Why redirect spill off /tmp

On many Linux hosts — especially systemd-managed ones — /tmp is mounted as tmpfs, which is backed by RAM (and swap), not disk. Spilling there does not actually free physical memory: the spill bytes stay resident, defeating the whole point of the memory budget. If df -T /tmp reports a tmpfs filesystem, point storage.spill.dir at a path on a real block device so spilling moves pressure off RAM and onto disk.

Prefer local spill for network-share pipelines

When sources and outputs live on NFS or SMB, point storage.spill.dir at a real local disk if possible. Spill workloads include repeated reads, writes, merges, and synchronization; running them on the share usually multiplies latency and network I/O without improving the durability of the final output. Likewise, optional source staging can copy matched share inputs to local disk before execution. Output commit remains destination-local: Clinker writes and flushes a hidden file on the output share, then promotes it on that same filesystem so the final rename does not cross devices.

Inspecting the resolved spill root

clinker run --explain prints the resolved spill root and where it came from, so you can confirm the setting took effect before committing to a run:

Spill root: /var/clinker/spill [storage.spill.dir]

…or, with no configuration:

Spill root: /tmp [OS temp dir (default)]

The same --explain output reports the resolved disk cap on the next line:

Spill disk cap: 10737418240 bytes [storage.spill.disk_cap_bytes]

…or, with no cap configured:

Spill disk cap: unlimited (default)

Finally, --explain reports the resolved compression decision per spill-writing operator, so you can see which spills will be LZ4-framed (lz4) and which will be written raw (off) before the run starts. Under auto the choice varies by operator width:

Spill compression: Auto [storage.spill.compress]
  Aggregate 'totals' → lz4
  Sort 'by_amount' → off

Only operators that actually write spill files appear here: the external sort, the hash Aggregate, the grace-hash / sort-merge Combine, and the pure-range (block-band) IEJoin Combine, which external-sorts each side and writes its min/max-tagged blocks to disk and spills its matched-output sort runs the same way. The remaining in-memory join strategies — the inline hash build/probe and the equi+range IEJoin (hash-partitioned range join) — run their kernel entirely in RAM and never open a spill file, so spill compression does not apply to them and they are omitted from this list, even though they carry a spill priority for memory arbitration.

storage.spill.disk_cap_bytes — cap concurrent spill

By default a run will spill as much as it needs, limited only by the physical space on the spill volume. disk_cap_bytes sets a budget on the spill the run holds at once: the on-disk size of the spill files live at any moment. When that footprint would cross the cap, the run aborts with a dedicated diagnostic instead of continuing to fill the volume. Because the cap tracks what is concurrently on disk, an operator that deletes intermediate spill files as it consumes them (such as the merge that folds a heavily fragmented external sort back together) does not count those transient files twice — only the disk a run actually occupies at once is charged against the cap.

[storage.spill]
dir = "/mnt/fast-ssd/clinker-spill"
disk_cap_bytes = "50GB"

The value accepts the same human-readable byte-size grammar as the source size filters — a bare integer is bytes, and KB/MB/GB suffixes use decimal units (1GB = 1,000,000,000 bytes), matching du, df, and the AWS CLI. Omitting the key leaves spill unlimited, exactly as before.

The cap is a policy ceiling, deliberately independent of both the memory budget and the physical volume size. A run can sit well inside its memory.limit and still exhaust local disk through an unbounded stream of spill files; the cap lets an operator bound that on a shared volume. It is the guard DataFusion shipped without (apache/datafusion#15358) until production runs filled volumes.

storage.spill.compress — LZ4 compression policy

Spill files are postcard-encoded record streams. By default each stream is wrapped in an LZ4 frame, which shrinks large spilled runs. But LZ4 carries a per-frame fixed cost — clearing the compressor’s internal state on every frame reset — and on small spills that cost can outweigh the byte savings. The LZ4 v1.8.2 release notes call this out directly, and Pentaho Kettle ships explicit guidance to turn spill compression off for small rows.

compress controls the policy:

[storage.spill]
compress = "auto"   # auto | off | on   (default = auto)
ModeBehavior
auto (default)Compress only when a spilled batch is projected large enough to amortize LZ4’s per-frame cost — both ≥ 4 KiB and ≥ 1024 rows. Below either threshold the batch is written raw. The projection comes from the operator’s schema width and the run’s batch_size, so the decision is made per blocking operator.
offNever compress. Postcard records are written straight to disk with no LZ4 frame. Cheapest for small spills; largest on-disk size.
onAlways compress with an LZ4 frame. The pre-knob behavior, best for spills of large, compressible rows.

Each spill file records its compression choice in a one-byte header tag, so the read path always dispatches to the right decoder regardless of the mode the file was written with — changing the knob between runs never breaks re-reading an earlier run’s files.

The 4 KiB / 1024-row thresholds mark the empirical crossover: below them the LZ4 frame’s fixed cost dominates the small amount of compressible payload, and writing raw is faster end-to-end (the spill_compression benchmark sweeps batch sizes from 256 B to 64 KiB and confirms auto tracks the faster of on / off across the range). Most pipelines should leave compress at auto; set on when spilling wide, highly compressible rows to a space-constrained volume, and off when spills are dominated by many small batches.

Observability — what the planner will do before you run

clinker run --explain is plan-only (it reads no input and spills nothing), so it is the safe place to see what a run would do to the spill volume and to the staging dir before committing to it. On top of the resolved spill root, disk cap, and compression decision documented above, --explain surfaces three storage-observability sections, and a real clinker run reports the matching actuals at end-of-run so you can calibrate the estimate.

A note on byte units. Three different unit conventions appear across the storage surface, and it helps to know which is which before comparing figures:

  • Config values you write (disk_cap_bytes = "10GB") use decimal units — 1GB = 1,000,000,000 bytes — matching du, df, and the AWS CLI (see the disk-cap grammar).
  • The === Estimated Spill Volume === section humanizes with binary suffixes — K/M/G = KiB/MiB/GiB — so it lines up with the predicted_peak figure on each stage’s Physical Properties line, which uses the same humanizer.
  • The cap-headroom line and the post-run actuals print raw bytes with no suffix, so the cap-minus-estimate subtraction and the estimate-vs-actual comparison are exact rather than rounded.

When you calibrate the estimate against the post-run actual, convert the binary estimate suffix to bytes first (1K = 1024 bytes, 1M = 1,048,576 bytes) so you are comparing the same unit the actuals report.

Estimated spill volume per stage

The === Estimated Spill Volume === section lists one line per spill-writing stage (hash Aggregate, external sort, grace-hash / sort-merge Combine, and the pure-range block-band IEJoin Combine) with its plan-time spill-volume estimate, followed by a total. The remaining in-memory join strategies (inline hash build/probe, equi+range IEJoin) never write spill files, so they do not appear here and do not inflate the total:

=== Estimated Spill Volume ===

Estimated spill volume (per blocking stage):
  [aggregation:hash] dept_totals → 1K
  [sort] by_amount → 4K
  Total: 5K

Each figure is the operator’s coarse predicted peak live state — the same predicted_peak the Physical Properties arbitration line shows — and bytes render in binary units (K/M/G = KiB/MiB/GiB). Summing rather than maxing is the conservative choice for a preflight: two blocking operators can be live and spilled at the same time, so their footprints add.

A streaming-only pipeline (no blocking operator) has nothing that spills, so the section is omitted entirely.

Unknown stages. The estimate is seeded from input file sizes resolved at plan time. A stage whose volume cannot be known before the run renders unknown instead of a misleading 0B, and the total notes that unknown stages are excluded:

  [aggregation:hash] dept_totals → unknown
  Total (known stages): 0B (excludes stages whose volume is unknown at plan time
  — a network source, a missing or unreadable input, or a glob/regex matcher
  whose discovery fails)

The seed is known for every file-backed matcher whose files can be sized at plan time: a single-file path: source, an explicit paths: list, and a glob: or regex: matcher. A glob/regex seed runs the same discovery resolver the run uses — applying its exclude, min_size/max_size, modified_after/before, take, and sort filters — and sums the matched files’ sizes, so the estimate names exactly the bytes the run will read with no second implementation to drift. A glob/regex that matches nothing seeds zero (rendered as unknown, since there is no spill volume to preview). The seed is genuinely unknown for a network source, for a missing or unreadable input file, and for a glob/regex matcher whose discovery itself fails (an invalid pattern, or no match under on_no_match: error) — the run surfaces the same error at startup. Check the post-run actuals below to calibrate any estimate.

Staging plan per source

When storage.staging is enabled, the === Staging Plan === section reports, for each source (and each discovered file under a multi-file matcher): whether it would be staged, the resolved content-addressed staged path, and — under on_existing = reuse — the reuse-if-fresh cache decision (hit if a committed prior copy still matches the live source, miss if it would be re-staged):

=== Staging Plan ===

Source 'orders':
  /data/in/orders-2024.csv → staged: yes, path: /mnt/local/staging/3f2a…b1.staged, reuse: hit
  /data/in/orders-2025.csv → staged: yes, path: /mnt/local/staging/9c4e…07.staged, reuse: miss

The reuse prediction runs the exact freshness check (mtime + size against the committed manifest) the real run makes, read-only — --explain copies nothing. A source that matches no staging pattern reports staged: no (no pattern match, reads in place); a network source reports not stagable (network source reads in place). When staging is disabled the section states that every source reads in place.

Cap headroom

When a spill cap is configured, --explain reports the headroom (cap minus estimate) with the same per-invocation disclaimer the startup cap-headroom preflight carries, and the same 80% warning:

Cap headroom: 5000000000 bytes free (5000000000 estimated of 10000000000 cap, 50%)
  [per invocation — does NOT account for sibling invocations sharing the spill
  volume under partition-and-run]

Machine-readable form — --explain json

clinker run --explain json emits the whole plan as JSON for tooling (the canvas, dashboards, CI gates). The same storage observability the text form prints lives under a structured storage_summary object, so a consumer reads per-stage spill estimates and the cap / staging summary without re-parsing prose:

{
  "schema_version": "1",
  "nodes": [ ... ],
  "node_properties": { ... },
  "storage_summary": {
    "spill_root": { "path": "/mnt/fast-ssd/clinker-spill", "source": "storage.spill.dir" },
    "spill_disk_cap_bytes": 1000000000,
    "estimated_spill": {
      "per_stage": [
        { "node_name": "dept_totals", "display_name": "[aggregation:hash] dept_totals", "estimate_bytes": 1024 },
        { "node_name": "by_amount", "display_name": "[sort] by_amount", "estimate_bytes": 4096 }
      ],
      "total_known_bytes": 5120,
      "any_unknown": false
    },
    "spill_compression": {
      "mode": "auto",
      "per_operator": [
        { "node_name": "dept_totals", "display_name": "[aggregation:hash] dept_totals", "compression": "lz4" },
        { "node_name": "by_amount", "display_name": "[sort] by_amount", "compression": "off" }
      ]
    },
    "cap_headroom": {
      "headroom_bytes": 999994880,
      "estimated_bytes": 5120,
      "cap_bytes": 1000000000,
      "pct_of_cap": 0.000512,
      "over_threshold": false
    },
    "staging": { "enabled": false, "sources": [] }
  }
}

The fields mirror the text sections one-for-one: estimated_spill is the === Estimated Spill Volume === section (a stage whose volume is unknown at plan time carries estimate_bytes: null and sets any_unknown: true), spill_compression is the Spill compression: projection, cap_headroom is the cap-headroom line (omitted when no cap is configured or the estimate is zero), and staging is the === Staging Plan === section. The JSON and DOT formats emit only their machine payload — the human-readable === Resolved Outputs === preamble the text form prints is suppressed so the output parses cleanly.

Post-run actuals — calibrating the estimate

A real clinker run that spills prints a per-stage actual spill-volume section at end-of-run, so you can compare it against the --explain estimate for the same stage — the calibration loop that turns a coarse pre-run estimate into a trustworthy one over repeated runs:

=== Spill Volume (actual, per stage) ===
  dept_totals → 1048576 bytes
  by_amount → 4194304 bytes
  Total: 5242880 bytes (compare against the --explain estimate)

The per-stage breakdown sums to the pipeline-wide cumulative spill total. A run that stayed within memory spilled nothing and prints no section. A large estimate-vs-actual delta is the single highest-leverage signal when a pipeline starts spilling unexpectedly (the failure mode behind Polars’ documented 13.5× spill amplification, where an optimizer interaction turned 30 GB of input into 400 GB of spill with no per-stage visibility).

Note on the --explain compression projection. The per-operator spill-compression decision shown under Spill compression: is projected from the same column count the operator’s runtime spill writer sees, so the projected auto verdict matches the file the run actually writes. A hash Aggregate and a grace-hash / sort-merge Combine project against their output schema (engine-stamped identity columns included), exactly the width their dispatch arms resolve compression against; an enforcer sort projects against the width of the records flowing into it — its upstream’s emitted schema — which is the width its sort buffer reads at runtime. The read path also dispatches on each spill file’s own one-byte header tag, so re-reading is robust regardless.

Distinguishing the runtime storage-abort conditions

A run that fails while spilling or staging emits one of several distinct diagnostics so a single glance at the error tells you exactly what to fix — instead of every disk and memory problem rendering as one ambiguous “out of memory” message (the trap DuckDB hit in duckdb/duckdb#14142, where a temp-dir cap was reported as “Out of Memory Error … 187.3 GiB/187.3 GiB used” and users inspected df only to find free space). The aborts split along two axes: the spill side (in-memory operator state landing on disk) and the staging side (matched source files copied to local disk before they are read).

Spill aborts

ConditionCodeWhat happenedWhat to do
Out of memoryE310An operator’s in-RAM state crossed the hard memory.limit (a true RSS overrun).Raise memory.limit, reduce input, or let the operator spill.
Spill cap exceededE320Cumulative spill bytes crossed storage.spill.disk_cap_bytes. The volume may still have free space — you hit the configured budget.Raise disk_cap_bytes, point storage.spill.dir at a larger volume, or reduce the spill footprint.
Spill volume fullE321The OS reported the spill volume out of space (ENOSPC). The physical disk filled.Free space on the volume, or move storage.spill.dir to a larger mount.
Spill directory unavailable(Spill)The spill directory went bad mid-run — unmounted, remounted read-only, deleted by a cleaner, or permissions revoked.Remount/restore the volume; stop the over-eager cleaner.

The key separations:

  • E310 vs E320 — an OOM is an in-RAM overrun; a cap-exceeded is a disk-budget stop. A run can hit E320 while comfortably inside its memory envelope, so conflating the two would point you at the wrong knob.
  • E320 vs E321 — E320 is the budget you set; E321 is the disk itself running dry. If you removed disk_cap_bytes, an over-large run would no longer trip E320 and would instead spill until the volume filled (E321).

(A future per-operator memory-reservation surface will add a fifth, reservation-exhausted condition; it is not part of the engine yet.)

Staging-copy aborts

When storage.staging is enabled, copying a matched source to local disk can fail in three distinct ways. Like the spill split, each has its own code so a content-corruption problem never renders as a budget problem and vice versa. Staging runs before any record flows, so these surface as startup-style validation failures.

ConditionCodeWhat happenedWhat to do
Staged copy corruptE335The local copy’s BLAKE3 digest did not match the source — the transport (e.g. a soft-mount NFS share) delivered different bytes than the source holds.Re-run over a healthy transport, harden the mount, or stage from a stable snapshot. Do not set verify = "none" to silence it — that hides corruption, not fixes it.
Staging cap exceededE336The cumulative bytes staged this run would cross storage.staging.disk_cap_bytes. The volume may still have free space — you hit the configured budget, not a full disk.Raise disk_cap_bytes, point storage.staging.dir at a larger volume, narrow storage.staging.patterns, or remove the cap.
Staged copy already existsE337A staged copy of this source already exists and on_existing = error refuses to touch it.Remove the existing copy, or switch on_existing to overwrite (re-stage) or reuse (reuse a fresh copy).

The same cap-vs-full-disk separation applies here as on the spill side: E336 is the budget you set (mirroring E320), so it must not render as an out-of-space message — a physically full staging volume instead surfaces as a staging I/O error (mirroring E321). E335 is distinct from a generic staging I/O error: an I/O error means the OS reported a fault, whereas E335 means the copy completed cleanly yet still does not match the source.

Startup storage validation

Before a run spawns its first source-ingest thread — after the plan compiles but before any input is read or any byte is spilled or staged — Clinker runs a single comprehensive validation pass over the resolved [storage] configuration. It rejects configurations that are physically wrong for the job, each with a stable diagnostic code, the offending clinker.toml field, and a clinker explain --code <CODE> pointer. Validating up front fails a misconfigured volume while the run is still cheap to abandon, rather than after minutes of work when the first spill or staged copy hits the bad volume.

CodeRejected configurationWhy
E330storage.spill.dir on an in-memory filesystem (Linux tmpfs / ramfs, Windows RAM disk).Spilling there keeps the bytes in RAM, so it frees no physical memory and defeats the memory budget.
E331storage.spill.dir on a network filesystem (NFS / SMB / CIFS / FUSE).A spill target on a soft-mounted share risks silent truncation and mmap data loss — the failure modes spill exists to avoid.
E332storage.staging.dir on a network filesystem.Staging copies inputs off a flaky share; a staging dir that is itself on a share reintroduces the fragility staging exists to escape.
E333storage.staging.dir on the same physical device as a matched (staged) source.The copy moves no I/O off the source volume, so it buys nothing while still spending time and space. Applies only to matched sources.
E334storage.spill.dir equal to storage.staging.dir.Spill files and staged source copies are sized and cleaned up differently; sharing one directory makes accounting and cleanup ambiguous.

The filesystem-class checks (E330–E332) read the volume type through one cross-platform detection layer, so they behave identically on Linux, macOS, and Windows: Linux matches the statfs f_type magic, macOS matches the f_fstypename string, and Windows maps GetDriveTypeW. (macOS has no native tmpfs, so E330 only ever fires on Linux and Windows.) The same-device check (E333) compares the device id on Linux/macOS and the volume serial number on Windows — the very same probe the staging same-volume rule uses, so there is one consistent notion of “same device” across the whole run.

Free-space preflight

Separately from the runtime disk cap (E320) and the full-volume surface (E321), the startup pass runs a free-space preflight: it queries the bytes available on the spill volume and compares them to the run’s estimated spill footprint (the sum of every blocking operator’s predicted peak state, the same estimate --explain surfaces). When the spill volume looks too small, the run prints a warning and continues:

W330: spill volume /var/clinker/spill has 2000000000 bytes free but the run is
estimated to spill up to 8000000000 bytes; the run may abort with a full-volume
error (E321) at the final spill — point storage.spill.dir at a larger volume or
reduce the spill footprint (raise memory.limit, partition the input)

This is advisory, not fatal: the estimate is a coarse upper bound (it ignores spill compression and the streaming drain), so the run may well finish within the available space. The warning exists so a long pipeline that would die at its final spill surfaces that risk before it runs for an hour, rather than after. The free-space query uses a cross-platform probe (statvfs on Unix, GetDiskFreeSpaceExW on Windows) that returns a 64-bit byte count, so the historical 32-bit f_bavail truncation never affects the comparison.

Cap-headroom preflight

When storage.spill.disk_cap_bytes is configured, the same startup pass also runs a cap-headroom preflight: it compares the run’s estimated spill volume to the configured cap and warns when the estimate reaches 80% of the cap. Unlike the free-space preflight (which probes the physical volume), this checks the run against the policy ceiling you set, so it fires even on a volume with plenty of free space:

W331: this run is estimated to spill up to 9000000000 bytes, which is 90% of the
configured spill cap storage.spill.disk_cap_bytes (10000000000 bytes); the run
may abort with a spill-cap error (E320) before it finishes — raise disk_cap_bytes
or reduce the spill footprint (raise memory.limit, partition the input). This
headroom is per invocation: if you partition the input and run several clinker
invocations against the same spill volume and cap, they share the cap, so the
real headroom is smaller than this figure

Like W330, this is advisory, not fatal — the estimate is a coarse upper bound, so a run that compresses well or never trips its memory budget may finish comfortably under the cap. It fires on a normal clinker run (before ingestion, at startup), not only under --explain, so an operator sees the signal on the real run even when they did not explicitly inspect the plan first.

Per-invocation accounting. The cap and the headroom figure are scoped to a single clinker invocation. Under the partition-and-run model — where you split a large input by file or key and launch several clinker processes that share one spill volume and one disk_cap_bytes — the physical spill volume is shared by every sibling, so the real headroom is smaller than any one invocation’s figure. The warning text states this explicitly rather than silently presenting a per-invocation number as a whole-volume guarantee. Clinker is single-process by design (one invocation = one OS process), so the engine cannot see its siblings; the disclaimer is the honest stance.

Mid-run spill failures

The startup check guarantees the spill directory is writable when the run begins, but it can still go bad mid-run — an NFS share remounts read-only, a volume unmounts, an over-eager temp-file cleaner deletes the directory, or permissions are revoked. When a spill write fails because the directory has vanished or become read-only, the run aborts cleanly with a distinct diagnostic rather than a generic I/O error or a panic:

spill directory /var/clinker/spill became unavailable mid-run: No such file or directory
(the directory may have been unmounted, remounted read-only, deleted by an
external cleaner, or had its permissions revoked)

This surfaces the directory-level cause directly, so the fix (remount the volume, stop the cleaner, restore permissions) is obvious from the message.

Crash purge of orphaned spill directories

A run’s spill directory (clinker-spill-<random>/) is normally removed when the run ends — a clean exit, a run that aborts with a fatal error, or even a panic all delete it. But a SIGKILL, the Linux OOM-killer, or a power loss kills the process before that cleanup runs, leaking the directory and every spill file inside it. Over many crashed runs that fills the spill volume.

To prevent that, a run cleans up orphaned spill directories at startup — but only when a spill directory is explicitly configured (storage.spill.dir), before it creates its own. It removes only directories left by dead runs and never touches one a concurrent run is still using.

When storage.spill.dir is not set, the spill root defaults to the OS temp directory (std::env::temp_dir, typically $TMPDIR or /tmp), and no startup purge runs there. In the default case a run still cleans up its own spill directory on every exit short of a hard kill; a directory leaked into the OS temp directory by a hard kill is the operating system’s temp-reaper’s responsibility, not Clinker’s. The purge is confined to a configured spill root because Clinker owns that volume but does not own the shared OS temp directory.

storage.staging — opt-in source staging

Reading source files directly from a network share (NFS, SMB) couples every run to the share’s availability and quirks: a soft-mount can silently truncate a read, and latency multiplies across many small files. Source staging copies matched source files to a local volume before the pipeline reads them, so the run works from stable local copies. It is off by default and activated per workspace by pattern match — pipelines that don’t opt in behave exactly as before.

[storage.staging]
enabled        = true
dir            = "/var/clinker/staging"   # required when enabled
patterns       = [
    "/mnt/nfs/data/**",
    "//fileserver/share/**",
]
disk_cap_bytes = "50GB"   # optional; cap on bytes copied per run (default unlimited)
verify         = "blake3" # optional; blake3 | none   (default blake3)
on_existing    = "overwrite" # optional; overwrite | reuse | error (default overwrite)
cleanup        = "on_success" # optional; on_success | always | never (default on_success)
KeyDefaultMeaning
enabledfalseMaster switch. When false, patterns is ignored and every source reads in place.
dir—Local directory the copies are written under. Required when enabled.
patterns[]Glob patterns selecting which source paths to stage. A source is staged only when enabled and its path matches at least one pattern. Empty ⇒ nothing is staged.
disk_cap_bytesunlimitedCumulative cap on bytes copied per run. Same byte-size grammar as the spill cap ("50GB", bare integers are bytes).
verifyblake3Post-copy integrity check. blake3 hashes source and copy and requires a match — the only check that catches a soft-mount’s silent truncation. none skips the check.
on_existingoverwriteWhat to do when a staged copy of this source already exists from a prior run: overwrite re-copies unconditionally; reuse reuses the existing copy only when it is still fresh (the source’s modification time and size match what was recorded when it was staged), otherwise re-copies; error fails the run rather than touch the existing copy. See The staging cache below.
cleanupon_successWhen staged copies are deleted relative to the run’s outcome: on_success removes them after a clean exit but keeps them after a failure so the operator can inspect the exact inputs the failed run saw; always removes them regardless; never keeps them as a persistent reuse cache for a later reuse run. See Cleanup.

Pattern matching

patterns uses the same glob grammar as a source’s exclude: list. Each pattern is tested against both the full path and the basename, so /mnt/nfs/** matches a deep path by its full path while *.csv matches any CSV by basename. ** crosses directory boundaries; * does not.

Startup validation

When enabled, staging is validated once at startup, before any input is opened, so a misconfiguration fails the run immediately rather than at the first copy. The run is refused when:

  • dir is unset.
  • dir does not exist, is a file, or is not writable (probed with a real create-and-delete, so a read-only mount or restrictive ACL is caught).
  • a patterns entry is not a valid glob.
  • dir sits on the same volume as a matched source. Staging within one volume copies bytes without moving I/O off the slow share — a well-documented anti-pattern — so it is refused up front rather than left to surface as a confusingly slow pipeline. The check compares the source’s and the staging dir’s storage volume (the device id on Linux/macOS, the volume mount root on Windows); point dir at a local disk on a different volume.

The same-volume rule applies only to matched sources: a source the patterns don’t select reads in place, so its volume is irrelevant.

How a file is staged

Staging copies the matched source to your local staging directory once, then verifies the copy against the source (with verify = blake3, the default, a content mismatch fails the run with E335). From then on the pipeline reads from the local copy. The same source always resolves to the same staged file, so a later run can find and reuse a prior copy.

The staging cache (on_existing)

Because staged copies live at stable paths, a copy from a prior run is still on disk when the next run starts (unless cleanup removed it). on_existing decides what happens when that prior copy is found:

ModeBehavior
overwrite (default)Always re-stage. The prior copy is removed and the source is copied fresh. The safe default: a copy from a crashed run must not be trusted.
reuseReuse the prior copy only when it is still fresh — the source’s current modification time and size both match what was recorded when it was staged. A fresh match skips the copy entirely (no bytes read off the share, nothing charged against the disk cap). A changed mtime or size means the source was rewritten, so the copy is stale and is re-staged.
errorFail the run with a clear diagnostic if a staged copy already exists, rather than overwrite or reuse it. For workflows that want an explicit “the cache is already populated” stop.

reuse is the mode that turns staging into a cache: re-running the same pipeline over an unchanged network share copies nothing on the second run. The freshness check is mtime + size, not a re-hash, so it is cheap.

Staging is safe to run from several clinker invocations at once over a shared staging volume: a source is copied exactly once no matter how many runs race for it, a run always reads a complete copy, and no run fails because a sibling was reading, cleaning up, or re-staging the same source.

Cleanup (on_success | always | never)

cleanup decides when a run’s staged copies are removed, keyed on the run’s outcome:

ModeBehavior
on_success (default)Remove the copies after a clean exit; keep them after a failure (or an interrupted / DLQ-producing run) so the operator can inspect the exact inputs the run saw and re-run without re-fetching.
alwaysRemove the copies when the run ends, success or failure.
neverKeep the copies indefinitely as a persistent reuse cache. Combine with on_existing = reuse to make repeated runs over a stable source copy-free. The operator reclaims the staging dir manually (or lets the next run’s crash purge eventually reap stale entries).

Each staged file’s manifest is removed alongside it, so cleanup never leaves a manifest pointing at a staged file that is gone.

Crash purge of orphaned artifacts

A SIGKILL, the Linux OOM-killer, or a power loss can kill a run before its cleanup runs, leaving half-finished staging artifacts behind. To stop those from accumulating, every run cleans up leftover artifacts from dead runs at startup, before it stages anything. A complete staged copy is the reuse cache and is always kept; only incomplete leftovers are reclaimed.

File permissions

Staged copies hold verbatim source records — potentially PII, credentials, or financial data — so on Unix they are created with owner-only permissions. On Windows staged files inherit the staging directory’s permissions, so restrict the directory if the volume is shared with other users.

Crash durability and the parent-directory fsync

Staged copies survive a crash: a later run finds a complete file or nothing at all, never a half-written one.

Streaming vs. Blocking Stages

Every node in a pipeline is one of two kinds at runtime, and the difference is what keeps Clinker’s memory bounded:

  • Streaming stages pass records through without holding the whole input. Their memory footprint stays small no matter how large the input is.
  • Blocking stages must see their entire input before they can produce any output, so they accumulate state. They stay within the memory budget and spill to disk when it gets tight, rather than holding everything in RAM.

Peak memory includes all concurrently live operator state, source queues, writer buffers, and retained intermediate records. The shared budget and spill policies govern that combined working set; the largest blocking stage alone is not a peak-memory bound.

Interactive companion: the streaming vs. blocking explainer classifies every stage of a few pipeline shapes as you change their settings, and shows the --explain lines.

Which stages stream

A stage streams when two things hold: it is one of the stages listed below, and its output goes to exactly one consumer that can take a stream, which is a Sink, the input of an Aggregate, or the driver side of a hash Combine. Any other stage, a stage that feeds two consumers, and a stage that roots an analytic window keep their output in a buffer instead. For example, in Source → Transform → Transform → Sink only the first Transform streams.

Two shapes stream and hold only one batch at a time, however large the input:

  • Source → Transform → Sink chains, where the Transform has no window and its Source feeds only it. Records flow straight from the reader through the transform to the writer.
  • Merge in interleave mode without an interleave_seed, whose inputs are all Sources, each feeding only the Merge.

These hand their output straight to their one consumer, but still build their own result first:

  • Route with only one branch wired to a downstream stage. The default: branch counts: a Route with one condition and a wired default has two consumers, and gives each its own buffer.
  • Merge in concat mode, in seeded interleave mode, or in interleave mode fed by other stages.
  • Aggregate with strategy: streaming — when the input is pre-sorted on the group key, each group is emitted as soon as the key advances. (See Aggregate Nodes.)
  • A hash Combine’s output, and its driver side, which streams in against the already-built lookup table.
  • A range Combine’s output, once it has sorted both sides.
  • Sink — a Sink writes each record to its writer as it arrives. A Sink with sort_order, split or a per-source-file path takes no stream, so the stage before it keeps a buffer.

Document boundaries (the signals behind $doc.*) flow inline with records through streaming stages, so a document’s close always trails its last record.

Which stages block

A stage blocks when its result depends on records it has not seen yet:

  • sort — the full input must be present before the first sorted record is known.
  • Hash Aggregate — a group’s final value depends on every member, so the group table retains aggregate state for every live group. (A streaming-strategy Aggregate over pre-sorted input is the exception above.)
  • A Combine’s build side — the lookup table is built in full before any driver record is matched. The probe side streams; the build side materializes.
  • Time-windowed and correlation-key Aggregates — these hold their group state for windowing or for the correlation commit, so they materialize.

A blocking stage keeps its accumulated state inside pipeline.memory.limit and spills to disk when the budget gets tight.

Seeing the classification

clinker run <pipeline>.yaml --explain annotates every node with its class in the Physical Properties section:

sink.report:
  buffer: streaming

aggregation.dept_totals:
  buffer: materialized

buffer: streaming marks a stage that holds only a small in-flight slice; buffer: materialized marks one that holds a whole stage’s output and may spill it. The annotation follows the rules the executor applies at runtime, with two known gaps:

  • A Sink with reconstruct_envelope: true turns streaming output off at runtime, but --explain still reports buffer: streaming for that Sink and for a Transform feeding it.
  • Under a correlation key, --explain reports each Sink as buffer: streaming, though the correlation commit writes its rows.

See Explain Plans and Memory Tuning.

Under dlq_granularity: document, as under a correlation key, streaming handoffs are off for the whole pipeline: no Transform, Merge, Route, Aggregate or Combine hands its output to a streaming consumer, so none of them reports buffer: streaming. Each Sink reports buffer: materialized, because it holds every open document’s records until the document’s verdict is final. A Source read by a single Transform still hands its records straight to that Transform, and keeps buffer: streaming.

Tuning the batch size

The number of records a streaming stage hands downstream at a time is set by pipeline.batch_size (default 2048), with an optional per-transform override. Smaller batches lower in-flight memory at the cost of more per-batch overhead; larger batches do the reverse. The batch size changes only the memory profile of streaming handoffs — never their output, and never the behavior of blocking stages.

Optimizing Pipelines

Clinker keeps memory bounded and spills to disk automatically, so most pipelines run fine with no tuning at all. When you do need a pipeline to run faster or in less memory, a handful of authoring choices do nearly all the work. This page is the practical checklist; the engine mechanics behind each tip live in the separate Engine Internals book.

Let stages stream instead of buffer

The cheapest pipeline is one where records flow straight through without being held in memory. A Source → Transform → Sink chain streams end to end — no intermediate stage is materialized. You get this automatically; the things that break it are fan-out (a Route with several branches, an output that forks) and blocking operators (sort, hash aggregation, the build side of a Combine).

Practical implication: keep the hot path simple. A filter-and-reshape job that’s just Source → Transform → Sink already runs at minimal memory. See Streaming vs. Blocking Stages for which operators stream and which block.

Make aggregation stream with sort_order

A hash Aggregate holds one entry per distinct group key in memory — fine for low-cardinality keys, expensive for high-cardinality ones. If your input is already sorted on the group-by keys, declare it:

- type: source
  name: txns
  config:
    type: csv
    path: ./data/transactions_sorted.csv
    sort_order:
      - { field: account_id, order: asc }
    schema:
      - { name: account_id, type: string }
      - { name: amount, type: float }

- type: aggregate
  name: per_account
  input: txns
  config:
    group_by: [account_id]
    cxl: |
      emit total = sum(amount)

With a matching sort_order, the optimizer switches the aggregate to streaming — it emits each group as the key advances and holds only one group at a time, regardless of cardinality. To make the requirement explicit (and turn a silent fallback to hash aggregation into a compile error), set strategy: streaming. See Aggregate Nodes → Strategy hint.

sort_order is checked for each file as it is read. With on_unsorted: warn (the default) a file that is out of order is sorted before it is released and a W307 warning names it; with on_unsorted: error the file is rejected. A wrong declaration costs time rather than correctness, but declare it only when the data really is sorted. See Source Nodes → Sort order.

The order only reaches the Aggregate if every stage in between keeps it: a Merge, a Combine, distinct, or a Transform that writes one of the sort fields (even emit account_id = account_id) drops it. The sort order explainer shows which stages keep it.

Choose the Combine driver side deliberately

A Combine holds each non-driving (build-side) input in memory as a lookup table, then streams the driver against it. So:

  • Put the smaller relation on the build side and drive with the larger stream — you iterate the big input once and keep only the small one resident. Plan for roughly 1.5–2× the build file’s size in memory.
  • The driver also sets output order and which side’s correlation identity propagates, so pick it for those reasons too. See Combine Nodes and Correlation Keys → Combine interaction.

A large build side isn’t a failure — the join spills to disk automatically — but spilling is slower than staying in memory, so sizing the driver right is the main lever.

Size the memory budget

The default budget is 512 MB. Raise it when a pipeline does high-cardinality aggregation or large joins and you have the RAM; lower it to be a good neighbor on a shared box. The budget is a target, not a hard wall — stages spill rather than fail when they exceed it.

pipeline:
  name: my_pipeline
  memory:
    limit: "1G"

Full sizing guidance and the backpressure knob are in Memory Tuning.

Reduce intermediate state

Less data in flight means less to buffer and spill:

  • Filter early. Drop records you don’t need in the first Transform, before they reach a blocking stage.
  • Project narrowly. Emit only the fields downstream stages actually use; carrying wide records through a sort or aggregate costs memory per row.
  • Aggregate before joining when you can — feeding a small rolled-up relation into a Combine is cheaper than joining raw rows and aggregating after.

Confirm with --explain and metrics

Before running, clinker run pipeline.yaml --explain annotates each node with buffer: streaming or buffer: materialized, so you can see which stages will dominate memory. After a run, the metrics spool reports peak_rss_bytes — if it consistently approaches your limit, raise the budget or cut intermediate state. See Explain Plans.

Metrics & Monitoring

Clinker writes per-execution metrics as JSON files to a spool directory. These files can be collected into an NDJSON archive for ingestion into monitoring systems.

Interactive companion: Where did my rows go? follows every row of ten small pipelines into the end-of-run counters, and shows which rows are in no counter at all.

Enabling metrics

There are three ways to enable metrics collection, listed from highest to lowest priority:

CLI flag:

clinker run pipeline.yaml --metrics-spool-dir ./metrics/

Environment variable:

export CLINKER_METRICS_SPOOL_DIR=./metrics/
clinker run pipeline.yaml

YAML config:

pipeline:
  metrics:
    spool_dir: "./metrics/"

When metrics are enabled, each execution writes one JSON file to the spool directory, named <execution_id>.json.

Metrics schema

Each metrics file follows schema version 3. The collector rejects spool files written under an older schema version, so upgrading clinker across a schema bump means draining the spool first.

{
  "execution_id": "01912345-6789-7abc-def0-123456789abc",
  "schema_version": 3,
  "pipeline_name": "customer_etl",
  "config_path": "/opt/clinker/pipelines/daily_etl.yaml",
  "hostname": "prod-etl-01",
  "started_at": "2026-04-11T10:00:00Z",
  "finished_at": "2026-04-11T10:00:05Z",
  "duration_ms": 5000,
  "exit_code": 0,
  "records_total": 50000,
  "records_ok": 49950,
  "records_written": 49950,
  "records_dlq": 50,
  "records_null_dropped": 0,
  "execution_mode": "Streaming",
  "peak_rss_bytes": 134217728,
  "thread_count": 4,
  "input_files": ["./data/customers.csv"],
  "output_files": ["./output/enriched.csv"],
  "dlq_path": "./output/errors.csv",
  "error": null,
  "retraction": {
    "groups_recomputed": 0,
    "partitions_dispatched": 0,
    "iterations": 0,
    "degrade_fallback_count": 0,
    "synthetic_ck_columns_emitted_total": 0,
    "synthetic_ck_fanout_lookups_total": 0,
    "synthetic_ck_fanout_rows_expanded_total": 0
  },
  "per_source_record_counts": { "customers": 50000 },
  "per_source_dlq_counts": { "customers": 50 }
}

Field reference

FieldTypeDescription
execution_idstringUUID v7 or custom --batch-id value
schema_versionintegerSchema version of this payload; currently 3
pipeline_namestringThe name from the pipeline YAML
config_pathstringAbsolute path to the config file
hostnamestringMachine hostname
started_atstringISO 8601 UTC timestamp
finished_atstringISO 8601 UTC timestamp
duration_msintegerWall-clock duration in milliseconds
exit_codeintegerProcess exit code (see Exit Codes)
records_totalintegerRecords read from every Source, including records rejected while being read (for example a value that does not fit its declared type) and the lookup inputs of a Combine. per_source_record_counts splits it by Source
records_okintegerDistinct source records that reached at least one output. Under inclusive Route fan-out one input matching N branches counts once. An Aggregate output row counts as one record of its group (the first one read), so the other records of the group are not counted; a Combine output row counts as its driver record, so lookup records are not counted; a whole-input Aggregate (group_by: []) over no records still writes one row and counts 1
records_writtenintegerTotal writes across all sinks. Equals records_ok for single-output exclusive pipelines; exceeds it under inclusive Route fan-out or multiple Sinks
records_dlqintegerRows written to the dead-letter queue, collateral rows included. A source row counts once for each failure it took part in, with or without a correlation key
records_null_droppedintegerRecords excluded by a Sink null_order: drop sort field (see Sort order). Counts exclusions, not distinct source records: like records_written, one source record dropped at two Sinks counts twice. Absent from spool files written before this counter existed, where it reads as 0
execution_modestringDAG-derived execution summary: Streaming (no full-stage materialization required) or TwoPass (a blocking stage forces an accumulation pass)
peak_rss_bytesinteger/nullPeak resident set size in bytes, sampled across chunk boundaries on Linux, macOS, and Windows. null on platforms where RSS sampling is unavailable
thread_countintegerThread pool size used
input_filesarrayPaths to all source files
output_filesarrayPaths to all output files written
dlq_pathstring/nullPath to the DLQ file, or null if none
errorstring/nullError message on exit 1/3/4, or null on success (exit 0) and partial success (exit 2)
retractionobjectCorrelation-key retraction counters (see below). All-zero on strict pipelines, which never enter the relaxed loop
per_source_record_countsobjectIngest record count per Source node, keyed by node name. A source that read zero records is present with a count of 0
per_source_dlq_countsobjectDLQ entry count per Source node; sources with zero DLQ entries are absent. The values sum to at most records_dlq — see the note below

The sum of per_source_dlq_counts values is at most records_dlq, and can be less: a failure in a Combine emit or a post-aggregate row is not traceable to a single declared source, so it is counted in records_dlq but not in this per-source breakdown. For pipelines whose dead-letters all originate at a declared source, the two match exactly.

The retraction object carries the relaxed correlation-key retraction orchestrator’s counters: groups_recomputed, partitions_dispatched, iterations, degrade_fallback_count, synthetic_ck_columns_emitted_total, synthetic_ck_fanout_lookups_total, and synthetic_ck_fanout_rows_expanded_total. Every field is 0 on strict pipelines and on relaxed pipelines that never trigger a retraction. See Correlation Keys for the underlying mechanism.

Collecting metrics

The spool directory accumulates one file per execution. Use clinker metrics collect to sweep them into an NDJSON archive:

clinker metrics collect \
  --spool-dir ./metrics/ \
  --output-file ./metrics/archive.ndjson \
  --delete-after-collect

This appends all spool files to the archive (one JSON object per line) and removes the originals. The NDJSON format is compatible with most log aggregation and monitoring tools.

Preview without writing:

clinker metrics collect \
  --spool-dir ./metrics/ \
  --output-file ./metrics/archive.ndjson \
  --dry-run

Integration with monitoring systems

Grafana / Prometheus

Parse the NDJSON archive with a log shipper (Promtail, Filebeat, Vector) and create dashboards tracking:

  • duration_ms – execution time trends
  • records_dlq – data quality over time
  • peak_rss_bytes – memory utilization

Datadog

Ship NDJSON to Datadog Logs, then create metrics from log attributes:

# Example: tail the archive and ship to Datadog
tail -f ./metrics/archive.ndjson | datadog-agent log-stream

ELK Stack

Filebeat can ingest NDJSON directly:

# filebeat.yml
filebeat.inputs:
  - type: log
    paths:
      - /var/log/clinker/metrics.ndjson
    json.keys_under_root: true

Simple alerting with jq

For environments without a full monitoring stack, use jq to query the archive directly:

# Find all runs with DLQ entries in the last 24 hours
jq 'select(.records_dlq > 0)' metrics/archive.ndjson

# Find runs that exceeded 400MB RSS
jq 'select(.peak_rss_bytes > 419430400)' metrics/archive.ndjson

# Average duration by pipeline
jq -s 'group_by(.pipeline_name) | map({
  pipeline: .[0].pipeline_name,
  avg_ms: (map(.duration_ms) | add / length)
})' metrics/archive.ndjson

Workspace OTLP and lineage policy

Deployment observability is optional and disabled when clinker.toml has no [observability] table. It is workspace policy, not pipeline YAML, and does not participate in the compiled plan’s semantic fingerprint. A present table is one complete policy: callers may supply a complete resolved replacement only when the workspace table is absent; individual fields are never merged.

The workspace loader validates this policy without opening a source, output, attempt directory, worker, credential provider, or network connection. It keeps the Collector endpoint as length-bounded raw text exactly as authored. The one shape it requires of that text is the shape it requires of every other authored string in the table: non-empty, within its byte cap, with no surrounding whitespace and no embedded control character. Padding and a carriage return are not part of an endpoint under any parse, and refusing them here names observability.otlp.endpoint and hands you a pasteable correction, where the later network boundary can only report that some endpoint was unusable.

The network admission boundary parses the text itself later, before any delivery effect; scheme, authority, credentials, paths, query strings, fragments, normalization, and the fixed OTLP signal routes are deliberately not decided by the workspace parser. Collector reachability is not a configuration admission check.

A complete fixed-capacity example is:

[observability]
arena_bytes = "4MB"
ordinary_lane_bytes = "3MB"
high_severity_lane_bytes = "1MB"
max_batch_bytes = "256KB"
max_attributes_per_event = 32
max_attribute_bytes = "4KB"
drop_policy = "drop_newest"
sample_every = 1
rate_limit_per_second = 1000
rate_limit_burst = 1000
flush_timeout_ms = 15000

[observability.otlp]
endpoint = "https://collector.example.com"
connect_timeout_ms = 1000
request_timeout_ms = 5000
retry_max_attempts = 3
retry_total_timeout_ms = 10000
max_response_bytes = "64KB"

[observability.otlp.auth]
mode = "none"

[observability.lineage]
queue_bytes = "1MB"
max_event_bytes = "64KB"
drop_policy = "drop_newest"
flush_timeout_ms = 5000
identity_mode = "external"

[[observability.lineage.dataset]]
node = "source_customers"
canonical_datasource = "s3://warehouse/customers"

[[observability.lineage.dataset]]
node = "output_customers"
catalog_namespace = "analytics"
catalog_name = "customers_clean"

[[observability.field_policy]]
event = "run.completed"
field = "records_written"
action = "allow"

[[observability.field_policy]]
event = "transform.customer_seen"
field = "customer_id"
action = "hash"

[[observability.field_policy]]
event = "transform.customer_seen"
field = "email"
action = "replace"
replacement = "[redacted]"

Byte-size strings use decimal units (1KB = 1,000 bytes and 1MB = 1,000,000 bytes). The fixed defaults and hard ceilings are:

KeyDefaultHard ceiling or relationship
arena_bytes"4MB""64MB"; equals the exact sum of both lane caps
ordinary_lane_bytesthree quarters of the arena"64MB" and disjoint from the high-severity lane
high_severity_lane_bytesone quarter of the arena"64MB" and disjoint from the ordinary lane
max_batch_bytes"256KB""1MB" and no larger than either lane
max_attributes_per_event32256
max_attribute_bytes"4KB""64KB"
sample_every11,000,000
rate_limit_per_second1,0001,000,000
rate_limit_burst1,0001,000,000
flush_timeout_ms15,00060,000
otlp.connect_timeout_ms1,00060,000 and no greater than request timeout
otlp.request_timeout_ms5,00060,000 and no greater than retry total
otlp.retry_max_attempts310
otlp.retry_total_timeout_ms10,00060,000 and no greater than flush timeout
otlp.max_response_bytes"64KB""1MB"
lineage.queue_bytes"1MB""64MB", reserved independently of the telemetry arena
lineage.max_event_bytes"64KB""1MB" and no larger than its lineage queue
lineage.flush_timeout_ms5,00060,000

Every byte default above is the quantity its own spelling parses to, so writing a default out in full changes nothing.

Sizing the arena

The two lanes partition the arena exactly: no telemetry byte is charged twice, and none of the arena is unreachable. You may write any of the three and leave the rest to be worked out from what you wrote:

  • arena_bytes alone — the lanes split it three-to-one, whatever its size. arena_bytes = "8MB" gives a 6 MB ordinary lane and a 2 MB high-severity one.
  • One lane alone — the arena stays at its default and the other lane takes the remainder.
  • Both lanes — the arena is their sum.
  • All three — the equality is checked rather than adjusted, and a disagreement is refused before the run starts.

The arena is the budget in every case: a lane is never allowed to grow it. A lane that does not fit inside the arena is refused, naming both keys.

An exported attribute longer than max_attribute_bytes is cut to fit and marked with a trailing …, and the marker is charged against the same cap, so a marked value is never longer than an unmarked one would have been. The mark matters because a bare prefix is not obviously a prefix: an amount of 123456789 cut to four bytes reads as 1234, and a timestamp cut short is still a well-formed timestamp. A dashboard reading 1… fails visibly; one reading 1234 charts a wrong number. Note that the mark is a signal, not a guarantee — a free-text value that genuinely ends in … is indistinguishable from a truncated one.

Node names are exported verbatim under the same treatment. A name is never dropped for the characters it contains, so a Transform named with a space or a non-ASCII character still produces a span.

Both delivery paths admit with drop_policy = "drop_newest"; there is no blocking, unbounded, or disk-spool spelling. The telemetry arena contains two disjoint lanes: trace, debug, and info signals occupy the ordinary lane, while warn and error occupy the high-severity lane. The lineage queue is a separate reservation and cannot be expressed as an alias of either telemetry lane or the arena.

sample_every = N keeps one in every N signals, counted within each lane separately. The lanes are disjoint precisely so ordinary volume cannot crowd out problems, and sampling honours that: with sample_every = 10, a Transform emitting nine per-record info events for every error still keeps one in ten of its errors, and that fraction does not change when the info volume does. A run that raises its per-record logging therefore does not quietly thin out its error reporting.

What the machine terminal reports about export

Under --machine ndjson-v1 the terminal event carries an observability object summarising what the exporter did. It holds one counter group per signal — logs, metrics, and traces, each with accepted, rejected, attempts, and failures — plus flush_complete.

Read flush_complete before reading the counters. When it is true the counts are the run’s final accounting. When it is false the exporter did not get to the end of the flush: either it ran past flush_timeout_ms, or it could not take the signal arena from the pipeline before giving up on it. Either way the counts are what had been recorded at that point, deliveries may still have been in flight, signals may remain that were never sent, and a low accepted means “we stopped counting” rather than “the collector refused them”. Those call for different responses, and only one of them is a collector problem.

Any 2xx answer counts as delivered. A collector’s own success status is 200, and its body is where a partial success declares the records it refused — those are what rejected counts. A gateway in front of a collector may answer 202 Accepted or 204 No Content instead, and there is no such body to read: the whole chunk counts under accepted, because the answer says it was taken. A rejection is a 4xx or a 5xx.

A delivery cut short after its request was fully sent is not sent again, even with attempts left in the budget. The collector may already hold that batch — a reply that is only slow cannot be told apart from one that was lost — and repeating it would ingest the same log records twice and count the same monotonic sums twice, which is wrong rather than merely unconfirmed. Such a batch is counted under failures, having spent fewer attempts than retry_max_attempts allowed: its delivery is unconfirmed, not known to have failed. A collector that habitually answers slowly needs a larger otlp.request_timeout_ms, not more attempts.

A retryable status (429, 502, 503, 504) waits before the next attempt. The wait starts at a fixed 100 ms and doubles before each retry after the first, so a collector recovering from load is not handed the whole attempt budget inside one of its recovery windows.

A Retry-After on that answer overrides the backoff whenever it asks for longer: the collector said how long “later” is, and the export waits at least that. Only the delay-in-seconds spelling is read. If the field asks for longer than otlp.retry_total_timeout_ms leaves — or arrives in the HTTP-date spelling, which this exporter does not parse — the delivery stops there rather than re-sending sooner than the collector asked or sleeping out the whole deadline. It is counted under failures with the throttle as its cause and retry_with_backoff as its advice; retrying belongs to whatever supervises the run. otlp.retry_total_timeout_ms remains the ceiling on all of it, so no requested wait can extend an export.

Delivery outcomes never change execution, publication, or the process exit status; the summary is an observation about the export, not about the run.

What the machine terminal reports about admission

The per-signal groups above are export-side, and an exporter can only count what reached it. A run that discarded most of its signals at the arena would otherwise report accepted = N, rejected = 0, flush_complete = true — a clean, complete-looking export of a silently truncated dataset. The admission object beside them says what the arena took and what it refused, before any export:

"admission": {
  "counts_complete": true,
  "accepted": 9,
  "dropped": {
    "sampled": 8, "rate_limited": 0, "queue_full": 0, "contended": 0,
    "oversize": 0, "invalid_identity": 0, "undecodable": 0
  },
  "lanes": {
    "ordinary":      { "sampled": 4, "queue_full": 0, "retained_bytes": 0, "capacity_bytes": 32000 },
    "high_severity": { "sampled": 4, "queue_full": 0, "retained_bytes": 0, "capacity_bytes": 32000 }
  },
  "fields": { "denied": 0, "truncated": 0, "limit_dropped": 0, "missing": 0 },
  "arena_recoveries": 0,
  "retained_bytes": 0, "peak_retained_bytes": 2477, "capacity_bytes": 64000
}

counts_complete says whether the rest of the object is a final accounting. These counters are read from the producer’s arena, and the arena keeps changing while the exporter drains it — undecodable is credited at drain. A flush that ran to completion joined the exporter first, so nothing was left to credit and counts_complete is true. A flush that expired on flush_timeout_ms detached an exporter that is still draining, so the read landed mid-drain and counts_complete is false: every number below it is whatever had been reached, and a low one means “we could not finish counting” rather than “nothing was lost”. The flush stays bounded either way — a finishing run does not wait on an unresponsive collector — so this flag, not a longer wait, is what keeps a truncated view from looking complete.

accepted counts the logs and spans the arena took. Metric points are coalesced into fixed counters rather than admitted as signals, so none of them appear here.

Each key under dropped is one reason a signal never became exportable. They are named as OpenTelemetry error.type values — a full arena is queue_full, not full — so these map onto SDK self-observability metrics without a rename. undecodable is the one member of the set that is also counted in accepted: those signals were admitted, then could not be read back at drain.

fields is a different kind of number and must not be added into a loss total. Those counts describe what became of the fields of records that were accepted — values denied, values truncated, attributes dropped at the per-event cap, and values a directive requested that the record did not carry. They reduce what a record says; they never discard one.

The first three are policy doing what you configured it to do. missing is not: a transform’s log directive asked for a column and the record did not have one. Most such requests are refused when the pipeline compiles (E374). What reaches this counter is the case the planner cannot decide — a selector naming a column that arrives through an open composition port. It is credited where the signal is built, so under a sampling policy it sees one miss per sampled event rather than one per record: read it as “this is happening”, not as how often.

arena_recoveries is neither a drop nor a quantity of anything lost. It counts the times the arena resumed from a poisoned lock — telemetry panicked while holding its own guard, and the arena carried on rather than taking the run down with it. A non-zero value says every counter beside it was produced by a subsystem that faulted mid-run. Treat it as a defect report against Clinker, and read the run’s other telemetry numbers with that in mind; the pipeline’s own results are unaffected, because telemetry never changes them.

Reconciling admission against export

Every signal counted under dropped is one the collector never saw. A delivery accounts for every item in its chunk — a chunk travels whole, so accepted + rejected equals the items handed to the exporter whatever the outcome was. That gives an exact identity for a run where flush_complete is true:

admission.accepted
  = (logs.accepted   + logs.rejected)
  + (traces.accepted + traces.rejected)
  + admission.dropped.undecodable
  - 1

The - 1 is the run-lifecycle span, which the exporter synthesizes at the final flush rather than drawing from the arena. metrics has no term because metric points are not admitted signals. When flush_complete is false the export counters are a partial accounting by definition and the identity does not apply; admission.counts_complete is false on that same run, because the arena side was read while the detached exporter was still draining it.

Any shortfall against that is arena loss, and dropped says which kind.

Why the lanes are split

sampled and queue_full are the two refusals that can cost an error while the ordinary lane is the thing under pressure, so both are attributed per lane. That is what makes the sampling guarantee above checkable from a run’s own accounting: with sample_every = 10, lanes.high_severity.sampled holds at one in ten of the high-severity signals produced however far lanes.ordinary.sampled climbs beside it. A single total cannot show that, and an author reading one would have no way to tell what share of their errors survived.

rate_limited and contended have no per-lane spelling. The rate limiter and the arena lock are properties of the shared arena rather than of a lane.

retained_bytes is what the lane still held when the run finished, and capacity_bytes is its reservation. Note that max_batch_bytes is the per-slot bound, so a lane holds lane_bytes / max_batch_bytes signals between drains — that ratio, not the byte size alone, is what queue_full is measured against.

Telemetry loss without --machine

A run without --machine ndjson-v1 discards the terminal object entirely, so when anything was dropped Clinker writes one line to standard error:

clinker: telemetry admission outcome: accepted=9 dropped=8 sampled=8 rate_limited=0 queue_full=0 contended=0 oversize=0 invalid_identity=0 undecodable=0 ordinary_sampled=4 ordinary_queue_full=0 high_sampled=4 high_queue_full=0 missing_fields=0 arena_recoveries=0 counts_complete=true

It mirrors the lineage delivery line, including its suppression rule: a run that dropped nothing prints nothing. A line that appeared on every run reading all zeroes is noise an operator learns to skip, and the one run that did lose signals would be skipped with it.

missing_fields and arena_recoveries break that silence on their own. Both sit outside dropped, and neither is anything an operator asked for: an attribute the collector never received, and telemetry having panicked under its own guard. The denied and truncated field counters stay silent by contrast, because they are policy doing exactly what it was configured to do — and they are not on this line at all.

The suppression is on the counters being final and clean, not on their reading zero. A run whose flush expired prints the line whatever the numbers say, with counts_complete=false: all-zero counts taken mid-drain are not evidence of a clean run, and staying silent on them would report a run that may well have lost signals as one that certainly did not.

Like the lineage line, this is an observation. Telemetry loss does not change execution, publication, the machine terminal result, or the exit status.

Runtime ownership and failure isolation

The deployment path keeps capability ownership narrow:

BoundaryOwned capability
Workspace plan/configSecret-free raw endpoint text plus numeric, capacity, retry, and deadline bounds.
NetworkThe sole endpoint admission, a private admitted-endpoint proof, fixed OTLP signal routes, and transport.
ExecutorThe real log, metric, and trace producers plus the fixed-memory telemetry arena.
LineageCanonical/catalog dataset identity, authorized subset and symlink facts, and independently bounded event delivery.
CLIPre-effect composition, one immutable lifecycle-fact source, worker lifecycle, and separate typed delivery outcomes.

The workspace policy loader owns only the secret-free raw endpoint string. For an enabled run, the CLI’s first capability transition calls the network crate’s sole endpoint-admission API. That API accepts one HTTPS origin and derives exactly three routes: /v1/logs, /v1/metrics, and /v1/traces. Relative, malformed, HTTP, credential-bearing, path-bearing, query-bearing, fragment-bearing, and already signal-specific endpoint text is rejected — and so is text that is not exactly the origin it names: surrounding whitespace, or any embedded control character such as a carriage return, is refused rather than trimmed or ignored, so no endpoint value can smuggle a header into a request. Each is rejected as observability.otlp.endpoint with a pasteable HTTPS-origin correction before source discovery, output attempts, arena reservation, worker construction, or network effects. Rejected text is not echoed.

After admission, the CLI combines the admitted origin with the configured request, retry, response, arena, and flush bounds in one immutable run-local bundle. auth.mode = "none" is the supported production capability today and sends no credential headers. auth.mode = "reference" remains a logical, secret-free policy name, but the run fails before exporter effects until the AUTH-01 credential applicator supplies that capability; the applicator will not be allowed to change the admitted origin or fixed routes.

Logs, metrics, and traces share one finite telemetry arena and exporter worker, but retain distinct typed per-signal delivery outcomes and fixed aggregate counters. Those producers are the executor’s actual lifecycle, runtime, and terminal producers; the transport does not invent equivalent events. The OpenLineage worker has its own queue, byte cap, sink, deadline, counters, and typed outcome. The arena allocation and both worker spawns complete before source discovery, staging, publication-attempt creation, sink writes, or a lineage START; inability to create either worker fails admission without those effects as observability.delivery.failed, with exit 4 and retry_with_backoff. Invalid endpoint, authentication, identity, and bounds policy remains observability.configuration.invalid with do_not_retry. Both paths copy the same batch ID, execution ID, semantic-plan algorithm/version/digest, and terminal facts from one immutable lifecycle snapshot; neither path reconstructs or owns those facts. Collector partial acceptance, rejection, transport failure, shutdown, or flush expiry, and lineage drop, sink failure, or deadline expiry remain optional observations: they do not change final or DLQ bytes, process status, the machine terminal result, publication inventory, visible finals, or retained failed-attempt evidence. The machine terminal exposes aggregate per-signal counters only; it does not flatten or replace either typed delivery outcome.

Field policy is applied before telemetry enters the arena. Denied values never reach Collector request bodies, OpenLineage events, counters, diagnostics, or machine records. With no [observability] table, Clinker performs no endpoint admission, arena reservation, worker creation, or exporter I/O.

External --lineage and --lineage-events exports serialize each complete OpenLineage event within lineage.max_event_bytes before attempting immediate admission to the byte-bounded lineage queue. A full queue drops the newest event instead of delaying the finite job. One lineage-only synchronous worker owns the destination and receives no output, DLQ, publication, or machine-mode authority. At completion Clinker waits no longer than lineage.flush_timeout_ms; dropped events, write or flush failure, and a deadline-exceeded worker are reported separately on standard error and do not change the authoritative ETL/publication result. The explicit local_diagnostic_paths compatibility mode remains a local synchronous file or console export and cannot use this external delivery path.

Exported signal shapes

Clinker speaks OTLP/JSON over the three fixed routes. What it puts on the wire:

Traces. One run is one trace. The lifecycle span clinker.run is the trace root and every Transform and Sink span is a child of it, so a collector can reconstruct the whole run from any single span. Each span carries a trace id, a span id, and both startTimeUnixNano and endTimeUnixNano.

A Transform is one span, emitted when the Transform finishes and covering the interval it ran for. It is not a start record followed by an end record. That shape was not exportable: a span requires both timestamps, so a start-only record is not a valid span, and two independently admitted halves can be sampled or dropped separately, leaving a collector holding one half of a pair it cannot use. If you want to observe that a Transform has begun while it is still running, that signal is the clinker.transform.started metric below, which is recorded before the work runs and exported on the normal metric cadence — the span’s startTimeUnixNano then tells you exactly when it began.

Metrics. The Transform counters — clinker.transform.started, clinker.transform.completed, clinker.transform.records, and clinker.transform.errors — are exported as monotonic sums with delta aggregation temporality. Each exported point is the count accumulated since the previous export, not a running total, and carries the startTimeUnixNano and timeUnixNano bounding that interval. Sum the deltas to get the run total; reading any single point as an absolute value will understate the run, often by a large factor on a long one.

Each real Sink writer work unit likewise emits one clinker.sink.started and exactly one of clinker.sink.completed, clinker.sink.failed, or clinker.sink.interrupted. clinker.sink.records, clinker.sink.errors, and clinker.sink.bytes report the rows handled, errors observed, and serialized bytes accepted by that writer boundary. clinker.sink.truncations counts the values a fixed-width writer cut to fit a truncation: warn column — the same count the end-of-run W367 warning reports. Synchronous, fused streaming, and correlation-deferred writers all use the same counter names and one closed clinker.sink span. A failed flush can therefore report bytes accepted before the destination rejected the flush; the failed terminal counter remains the authoritative outcome.

Each dead-letter (DLQ) file a run writes is a work unit of its own. It emits one clinker.dead_letter.started when the file receives its first row, and exactly one of clinker.dead_letter.completed (the file was written in full and handed to publication), clinker.dead_letter.failed (a write or flush failed, or the run failed), or clinker.dead_letter.interrupted (the run was interrupted). clinker.dead_letter.records and clinker.dead_letter.bytes report the rows written to the file and the bytes its writer accepted, header included. Each unit closes one clinker.dead_letter span whose clinker.logical_node attribute is dead_letter[<n>], where <n> is the file’s position, counting from 0, in the === Dead-Letter Output === section of clinker run --explain; the path never appears. A DLQ file that receives no rows emits nothing, and neither do dead letters that have no destination: count those with records_dlq. Preview runs write no DLQ file and emit no dead-letter signals.

clinker.correlation.group_overflows counts the correlation-key groups that went over error_handling.max_group_buffer, one per group when the run commits it. It counts a group whose rows all failed on their own too, which writes no group_size_exceeded row, so it can exceed the number of group_size_exceeded rows in the DLQ. See Correlation Keys.

An instrument that recorded nothing in an interval carries no points for it. That is an ordinary interval, not a malformed export: the batch is delivered and the points its other instruments did record arrive intact.

Transforms that a correlated commit re-runs. A pipeline with a relaxed correlation-key aggregate converges at commit time: the engine re-runs the transforms downstream of that aggregate, retracting the rows that turned out to fail, until the result stops changing. Only the converged result is published, and the exported signals describe it the same way — one span covering the whole convergence, one clinker.transform.started, one clinker.transform.completed, and record and error counts taken from the converged pass rather than added up across the discarded ones. An every: cadence continues across the passes rather than restarting on each. So a transform inside a convergence reports the rows the run actually carried, and its counters stay summable alongside every other transform’s.

A convergence that does not finish — a failure in one of the re-run transforms, or an interrupting signal — still reports every transform it had passed over. Each gets one span with an ERROR status covering the interval it ran for, one clinker.transform.started, and the record and error counts the interrupted pass had reached. It gets no clinker.transform.completed: nothing completed, and that counter is what tells you whether everything did.

Logs. Each emission of an authored log: directive becomes one OTLP log record: severityText from the directive’s level, the directive’s message as the body, the event name as the clinker.event attribute, the three run-correlation attributes described under Authentication and privacy, and whichever requested record fields the field policy allowed, hashed, or replaced.

Request size. Delivery is bounded per request, not per record. A drained batch that would exceed one request’s byte budget is split across as many requests as it needs rather than discarded. max_batch_bytes bounds one stored record inside the arena; it is not the request size, and the two are not the same number.

Authentication and privacy

Authentication is always explicit. Credential-free delivery uses exactly:

[observability.otlp.auth]
mode = "none"

Referenced delivery retains one provider-neutral logical name:

[observability.otlp.auth]
mode = "reference"
reference = "telemetry/production"

The reference is not a credential. A later run-local authentication provider must resolve it before effects. Inline headers, bearer/basic values, environment-variable names, and mixed auth variants are rejected; omission does not mean anonymous delivery. Diagnostics name the authored field and show a safe corrected table without echoing an endpoint, credential value, record value, or physical path.

Event fields are denied by default. Each [[observability.field_policy]] entry selects exactly one dotted event/field pair and one allow, hash, or replace action. replacement is required only for replace, and duplicate rules for the same pair are invalid. A replacement is written verbatim into the exported record, so it is held to the same shape as the other authored strings here: non-empty, bounded, with no surrounding whitespace and no embedded control character.

Field policy governs record fields — values a Transform selected out of the data being processed. It does not govern run correlation. Every exported log record carries clinker.execution_id, clinker.batch_id, and clinker.pipeline_name unconditionally, with no [[observability.field_policy]] entry required for them.

The reason is worth stating plainly, because the opposite behaviour was a defect rather than a policy: those three values are identifiers Clinker generates for the run itself, not data read from a source. A record-field privacy policy has nothing to decide about them. Subjecting them to it meant a workspace that declared an event without also writing three correlation rules exported log records with no correlation at all — telemetry that could not be joined to the machine stream’s execution_id, to the lineage events, or to another run of the same pipeline — and it counted all three as privacy denials on every event, inflating the arena’s denied-field accounting by three per event. Correlation is now carried outside the policy entirely, so neither happens.

A [[observability.field_policy]] rule whose field happens to be named execution_id, batch_id, or pipeline_name still parses. It governs a record column of that name, if a directive requests one; it has no effect on the correlation attributes above.

Lineage identity

identity_mode = "external" is the default. Every externally emitted source or Sink node needs exactly one binding: either canonical_datasource, or the complete catalog_namespace/catalog_name pair. Missing, duplicate, partial, and mixed bindings fail validation; Clinker does not synthesize an external identity from a working directory, worker path, temporary root, URL, attempt identifier, or path hash.

The stable collection identity remains the dataset namespace/name. When the runtime has explicitly authorized a concrete logical partition or location, lineage represents it with the standard input/output subset facet rather than changing that collection name. Explicitly authorized aliases use the standard symlinks facet. The current resolved workspace config has no author-facing subset or symlink fields, so Clinker does not infer either fact from local paths, attempt directories, hashes, or process context.

The only path-derived compatibility mode is the exact, explicit value below. It accepts no external dataset bindings and is for labeled local diagnostics, not external delivery:

[observability.lineage]
identity_mode = "local_diagnostic_paths"

Operational recommendations

  • Always enable metrics in production. The overhead is negligible (one small JSON write at the end of each run).
  • Run metrics collect --delete-after-collect on a schedule (e.g., hourly) to prevent spool directory growth.
  • Use --batch-id with meaningful identifiers to correlate metrics across retries and environments.
  • Alert on records_dlq > 0 to catch data quality regressions early.
  • Track peak_rss_bytes trends to anticipate when memory limits need adjustment.

Exit Codes & Error Diagnosis

Clinker uses structured exit codes to communicate the outcome of a pipeline run. These codes are designed for integration with schedulers, cron, CI systems, and monitoring tools.

Exit code reference

CodeMeaningDescription
0SuccessPipeline completed successfully, or an attempt operation completed without cleanup debt. A purge preview that safely selects nothing is also successful.
1Configuration or argument errorInvalid YAML, CXL syntax error, type mismatch, DAG wiring problem, invalid attempt selector, invalid continuation, a --lineage export rejected by the configured observability caps, a --lineage destination that will refuse every identical retry — one the process may not write, that does not exist, that is read-only, or that is not a file. Fix the pipeline configuration or command arguments.
2Partial successPipeline ran to completion, but some records were routed to the dead-letter queue. Check the DLQ file.
3Evaluation errorCXL runtime error during record processing (e.g., division by zero, type coercion failure).
4Infrastructure or retained cleanup debtFile/format failure, disk full, a --lineage export that failed in a way a retry may resolve — a reader that went away, a write that timed out, a volume that was out of space, or a flush that exceeded its deadline — an environment that refused the SIGINT/SIGTERM handler the run requires, in which case the run stops before reading or writing anything and the environment is what must change, or an attempt operation stopped with bounded, ambiguous, live, or otherwise retryable cleanup debt. This status never means completed-with-DLQ.
130CancelledGraceful SIGINT or SIGTERM cancellation won before publication, or a required --machine lifecycle record could not be written before publication. Final paths for the current attempt remain unchanged.

Known issue: a command-line usage error (an unknown flag or a missing argument) currently exits 2, the partial-success code, instead of 1, except under clinker attempts (#1372). A scheduler that treats 2 as “completed with dead letters” cannot tell the two apart yet; check stderr for a usage message.

For an ordinary standalone run, these statuses are the complete process result. A --machine ndjson-v1 consumer must additionally require exactly one supported terminal event whose result and embedded exit, where present, match the actual child status and current-attempt artifact evidence. EOF, malformed or unsupported output, a duplicate or missing terminal, forced termination, or any mismatch is an incomplete attempt even if an older final already exists. See Running Clinker Directly or Under a Supervisor.

Understanding exit code 2

Exit code 2 is not a crash. It means:

  • The pipeline started and ran to completion.
  • All viable records were processed and written to output files.
  • Some records could not be processed and were diverted to the dead-letter queue.

Your scheduler should distinguish completed-with-DLQ from an aborted run. Whether downstream work may continue is a data-quality policy decision. The DLQ contains the problematic records and their rejection diagnostics.

To bound the tolerated type-error ratio, configure the YAML policy:

error_handling:
  strategy: continue
  type_error_threshold: 0.05
  dlq:
    path: ./output/errors.csv
    format: csv

The threshold is a ratio from 0 to 1, not a record count. Exceeding the configured type-error ratio produces exit 3. It does not count every possible DLQ reason. The retired --error-threshold flag is rejected; see error handling for the policy’s population.

Diagnosing failures

Exit code 1: Configuration error

The error message includes a span-annotated diagnostic pointing to the exact location of the problem:

Error: CXL type error in node 'transform_1'
  --> pipeline.yaml:25:15
   |
25 |   emit total = amount + name
   |                ^^^^^^^^^^^^^ cannot add Int and String

Action: Fix the YAML or CXL expression indicated in the diagnostic, then re-run with --dry-run to confirm the fix.

Exit code 2: Partial success (DLQ entries)

Check the DLQ file for details:

# The DLQ path is shown in the run output and in metrics
cat output/errors.csv

Common causes:

  • Null values in fields that a CXL expression does not handle
  • Data that does not match the declared schema (e.g., non-numeric value in an integer column)
  • Coercion failures between types

Action: Review the DLQ records, fix the data or add null handling to CXL expressions, and re-run.

Exit code 3: Evaluation error

A CXL expression failed at runtime. The error message includes the failing expression and the record that triggered it:

Error: division by zero in node 'compute_ratio'
  expression: emit ratio = total / count
  record: {total: 500, count: 0}

Action: Add guard conditions to the CXL expression:

emit ratio = if count == 0 then 0 else total / count

Exit code 4: Infrastructure or retained cleanup debt

File system or format errors:

Error: file not found: ./data/customers.csv
  --> pipeline.yaml:8:12

Common causes:

  • Input file does not exist or path is wrong
  • Permission denied on input or output directories
  • Output file already exists (use --force to overwrite)
  • Disk full during output writing
  • Input file format does not match the declared type (e.g., invalid CSV)
  • A retained-attempt query reached its entry, byte, or monotonic-time bound
  • Cleanup kept an attempt because ownership, liveness, manifest, clock, or filesystem evidence was ambiguous

For clinker attempts, exit 4 includes path-free cleanup debt plus E371 (unsafe or invalid attempt refused) or E372 (cleanup incomplete or budget exhausted) when the debt belongs to a logical execution. Follow the emitted retry advice and pasteable clinker attempts inspect <workspace-relative.yaml> --execution-id <id> command. If a continuation is present, paste its exact resume command to continue the bounded page. Do not treat status 4 as authority to delete a directory manually.

Action: Fix file paths, permissions, or disk space for run failures. For attempt operations, inspect the named logical execution and resolve the stated debt before retrying or executing another bounded purge.

Exit code 130: Cancelled

Exit 130 means the attempt stopped before publication and the current attempt’s final paths are unchanged. Two things produce it:

  • A SIGINT or SIGTERM that won the cancellation gate before the first final rename.
  • Under --machine ndjson-v1, a required lifecycle record that could not be written. The run refuses to publish an outcome it cannot report, so a broken control pipe stops the attempt rather than promoting silently.

A discardable machine record — a periodic progress observation — is not in that set. Losing one is reported on stderr as machine progress channel failed and the run continues to its real outcome; it never converts a completed run into a cancellation or discards computed output. A supervisor should therefore read 130 as “nothing was published”, never as “an advisory record went missing”.

Cancellation is also recorded as a cancellation everywhere else it is reported: the OpenLineage terminal is ABORT and the --machine terminal is cancelled, regardless of which source observed the signal first. A source that notices cancellation while a request is in flight — a REST page read, for instance — produces the same terminals as a file source that drains and stops at a chunk boundary.

Action: Re-run the attempt from the beginning of the input with the same --batch-id. There is no resume or checkpoint state to recover; see Retry and identity boundaries. If the run was not cancelled by an operator, check stderr for a machine control-channel write failure and confirm the consumer is draining stdout.

Plan-time diagnostic codes

The process exit codes above tell a scheduler whether the run succeeded. The E### codes below appear inside the structured Error: messages a configuration error (exit code 1) prints, and identify the specific compile-time check that rejected the pipeline. The codes below cover the event-time watermark and time-windowed aggregate surface (issue #61); related code sets live in Pipeline Variables, Channels, and Correlation Keys.

CodeTriggerRemediation
E154A source declares watermark.column: <col> but <col> is not present in that source’s schema: block.Add the column to schema:, or remove the watermark: block.
E155A source declares watermark.column: <col> and the column exists, but its declared CXL type is not date_time or date.Change the column’s type: to date_time or date, or point watermark.column at a column that already has one of those types.
E156An aggregate declares time_window: but at least one upstream-reachable source does not declare watermark.column.Add watermark: { column: <event-time-column> } to each listed source, or remove time_window: from the aggregate. Without a watermark on every upstream source, min_across_sources never advances past None and the window can never close.
E157A source declares an external schema: file (schema: path.schema.yaml) that could not be read or parsed as a SourceSchema.Fix the file path or its contents. A schema file is a bare column list or a multi-record discriminator:/records: map — it may not itself point at another schema file.
E158A source column’s declared type is (or wraps) the inference-only numeric union.Declare a concrete int or float. numeric is int | float resolved during type unification and never carries into a compiled source schema.
E159A source pairs a generated schema with a non-EDI format.generated (engine-synthesized positional columns) is valid only for the EDI-family formats (edifact, x12, hl7, swift). Declare an explicit column list for any other format.

See Source Nodes → Watermarks and Aggregate Nodes → Time-windowed aggregates for the field semantics each code is enforcing.

DLQ category: LateRecord

When a time-windowed aggregate sees a record whose event time falls inside an already-closed window (window_end + allowed_lateness < min_across_sources), the engine routes the record to the DLQ instead of attempting to fold it into a finalized accumulator. Mirrors Flink’s sideOutputLateData and Spark Structured Streaming’s late-data drop.

The DLQ row carries:

  • _cxl_dlq_error_category = late_record
  • _cxl_dlq_stage = time_window:<aggregate-name>
  • _cxl_dlq_error_detail — the closed window’s [start, end) bounds as i64 nanoseconds since the Unix epoch

Tune watermark.delay (source-side, applies before any aggregate) or allowed_lateness (operator-side, applies per aggregate) to absorb expected out-of-order tails before they reach this path.

Scheduler integration

For running Clinker under a workflow orchestrator (Temporal, Airflow, Dagster) — mapping these exit codes onto a retry policy, plus the cancellation and output-atomicity guarantees — see Running Under a Workflow Orchestrator.

Cron script

#!/bin/bash
set -euo pipefail

PIPELINE=/opt/clinker/pipelines/daily_etl.yaml
METRICS_DIR=/var/spool/clinker/

# `set -e` would end the script as soon as clinker exits non-zero, before the
# `case` below runs. `|| EXIT=$?` captures the status without stopping.
EXIT=0
clinker run "$PIPELINE" \
  --memory-limit 512M \
  --log-level warn \
  --metrics-spool-dir "$METRICS_DIR" \
  --force || EXIT=$?

case $EXIT in
  0)
    echo "$(date): Success" >> /var/log/clinker/daily_etl.log
    ;;
  2)
    echo "$(date): Warning - DLQ entries produced" >> /var/log/clinker/daily_etl.log
    mail -s "Clinker ETL Warning: DLQ entries" ops@company.com < /dev/null
    ;;
  *)
    echo "$(date): FAILURE (exit code $EXIT)" >> /var/log/clinker/daily_etl.log
    mail -s "Clinker ETL FAILURE (exit $EXIT)" ops@company.com < /dev/null
    ;;
esac

exit $EXIT

CI pipeline (GitHub Actions)

- name: Run ETL pipeline
  run: clinker run pipeline.yaml --dry-run
  # Exit code 1 fails the build on config errors

- name: Smoke test with real data
  run: clinker run pipeline.yaml --dry-run -n 100
  # Catches runtime evaluation errors

Systemd

Systemd Type=oneshot services interpret non-zero exit codes as failures. To allow exit code 2 (partial success) without triggering service failure:

[Service]
Type=oneshot
SuccessExitStatus=2
ExecStart=/opt/clinker/bin/clinker run /opt/clinker/pipelines/daily_etl.yaml --force

Known issue: a command-line usage error, such as a mistyped flag or a missing argument, currently exits 2 instead of 1 (except under clinker attempts), so SuccessExitStatus=2 would also count a typo in ExecStart as a success (#1372). After editing the unit, run its ExecStart command once by hand and check that it does not print a usage error.

Production Deployment

Clinker is a single statically-linked binary with no runtime dependencies. Deployment is straightforward: copy the binary to the server.

Installation

# Copy the binary
scp target/release/clinker user@server:/opt/clinker/bin/

# Verify it runs
ssh user@server /opt/clinker/bin/clinker --version

No JVM, no Python, no container runtime required.

/opt/clinker/
  bin/
    clinker                    # The binary
  pipelines/
    daily_etl.yaml             # Pipeline configs
    weekly_report.yaml
  data/                        # Input data (or symlinks to data locations)
  output/                      # Output files
  rules/                       # CXL module files (for use statements)
  metrics/                     # Metrics spool directory

Create a dedicated user:

sudo useradd --system --home-dir /opt/clinker --shell /usr/sbin/nologin clinker
sudo chown -R clinker:clinker /opt/clinker

Systemd service

For scheduled one-shot execution:

[Unit]
Description=Clinker ETL - Daily Customer Processing
After=network.target

[Service]
Type=oneshot
ExecStart=/opt/clinker/bin/clinker run /opt/clinker/pipelines/daily_etl.yaml \
  --memory-limit 512M \
  --log-level warn \
  --metrics-spool-dir /var/spool/clinker/ \
  --force
WorkingDirectory=/opt/clinker
User=clinker
Group=clinker
SuccessExitStatus=2

# Resource limits
MemoryMax=1G
CPUQuota=200%

# Logging
StandardOutput=journal
StandardError=journal
SyslogIdentifier=clinker-daily

[Install]
WantedBy=multi-user.target

Pair with a systemd timer for scheduling:

[Unit]
Description=Run Clinker daily ETL at 2 AM

[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true

[Install]
WantedBy=timers.target
sudo systemctl enable --now clinker-daily.timer

Note: SuccessExitStatus=2 tells systemd that exit code 2 (partial success with DLQ entries) is not a service failure. See Exit Codes for the full reference.

Known issue: a command-line usage error, such as a mistyped flag or a missing argument, currently exits 2 instead of 1 (except under clinker attempts), so SuccessExitStatus=2 would also count a typo in ExecStart as a success (#1372). After editing the unit, run its ExecStart command once by hand and check that it does not print a usage error.

Cron scheduling

# Run daily at 2 AM, log to syslog
0 2 * * * /opt/clinker/bin/clinker run \
  /opt/clinker/pipelines/daily_etl.yaml \
  --log-level warn --force \
  2>&1 | logger -t clinker

# Collect metrics hourly
0 * * * * /opt/clinker/bin/clinker metrics collect \
  --spool-dir /var/spool/clinker/ \
  --output-file /var/log/clinker/metrics.ndjson \
  --delete-after-collect

Environment-based configuration

Use the CLINKER_ENV variable or --env flag to activate environment-specific overrides:

# Production
CLINKER_ENV=production clinker run pipeline.yaml

# Staging
CLINKER_ENV=staging clinker run pipeline.yaml

Combined with channel overrides in the pipeline YAML, this allows a single pipeline definition to target different file paths, connection strings, or thresholds per environment.

Logging

Log levels for production

LevelUse case
warnRecommended for production cron jobs. Prints warnings and errors only.
infoDefault. Includes progress messages. Useful during initial deployment.
errorMinimal output. Only prints when something fails.
debugTroubleshooting. Generates significant output.
traceDevelopment only. Extremely verbose.

Directing logs

To syslog via logger:

clinker run pipeline.yaml --log-level warn 2>&1 | logger -t clinker

To a log file:

clinker run pipeline.yaml --log-level warn 2>> /var/log/clinker/etl.log

Systemd journal captures stdout and stderr automatically when running as a service.

DLQ monitoring

When a pipeline exits with code 2, records that could not be processed are written to the dead-letter queue file. Set up a daily check:

#!/bin/bash
# Check for DLQ files produced today
DLQ_DIR=/opt/clinker/output/
DLQ_FILES=$(find "$DLQ_DIR" -name "*_errors.csv" -mtime 0 -size +0c)

if [ -n "$DLQ_FILES" ]; then
    echo "DLQ entries found:" | mail -s "Clinker DLQ Alert" ops@company.com <<EOF
The following DLQ files were produced today:

$DLQ_FILES

Review the files and address data quality issues.
EOF
fi

Batch ID for tracing

Use --batch-id with a meaningful, consistent naming scheme:

# Date-based
clinker run pipeline.yaml --batch-id "daily-$(date +%Y-%m-%d)"

# Include environment
clinker run pipeline.yaml --batch-id "prod-daily-$(date +%Y-%m-%d)"

The batch ID appears in metrics output and log lines, making it easy to correlate a specific run across logs, metrics, and DLQ files. On retries, use a different batch ID (e.g., append -retry-1) to distinguish attempts.

Upgrades

To upgrade Clinker:

  1. Validate the new version against your pipelines:
    /opt/clinker/bin/clinker-new run pipeline.yaml --dry-run
    
  2. Replace the binary:
    cp clinker-new /opt/clinker/bin/clinker
    
  3. Verify:
    /opt/clinker/bin/clinker --version
    

There is no configuration migration. Pipeline YAML files are forward-compatible within the same major version.

Running Clinker Directly or Under a Supervisor

Clinker is a finite synchronous CLI. The ordinary command is the primary path and needs no worker, service, scheduler, or orchestration component:

clinker run pipeline.yaml

An external scheduler or workflow runner can opt into a versioned child-process protocol when it needs machine-readable lifecycle evidence:

clinker run pipeline.yaml --machine ndjson-v1 --batch-id logical-batch

Machine mode does not change who owns the run. Clinker still validates, executes, cancels, and publishes one finite attempt. The parent process owns scheduling, retry and backoff, deadlines, heartbeats, process lifetime, and direct-child reaping. Clinker embeds no orchestrator SDK, worker, daemon, or service runtime.

Stream ownership and compatibility

In machine mode, stdout contains only compact UTF-8 NDJSON and is flushed after each event. Human diagnostics and tracing use non-ANSI stderr. The parent must drain stdout and stderr concurrently; reading one pipe to completion before the other can deadlock when an OS pipe fills.

Every event carries these fields:

FieldMeaning
protocolAlways clinker.run.
schemaProtocol major version. ndjson-v1 emits 1.
eventLifecycle kind such as started, plan_resolved, progress, publication_artifacts, completed, failed, or cancelled.
seqZero-based sequence, increasing by exactly one within this process.
batch_idCaller-supplied logical-batch correlation retained across retries.
execution_idFresh non-overridable UUIDv7 generated for this process.
plan_identitypending, resolved with the semantic plan fingerprint, or unavailable after admission failure.

The resolved fingerprint covers the compiled topology, schemas, composition and CXL dependencies, winning channel/group config values, and all four runtime-variable scopes. A composition contributes the body each call site actually binds, so two channels that patch one shared body differently are two plans. It excludes deployment-only file locations and layer source formatting, so relocating equivalent inputs does not invalidate a pinned plan while a value that can change execution does. Where rejected records are written is a location and is excluded; whether they are written is not — a pipeline with no error_handling.dlq.path and no per-source override discards them, and reads as a different plan from one that keeps them.

The version field carries the fingerprint schema, currently 3. A digest is comparable only with another digest of the same version: when the schema changes, the same pipeline yields a different digest, and the version is how a consumer holding a pinned value tells that apart from a changed plan.

Progress events add a bounded logical phase, kind, elapsed time, counts, and truncation flags. They never contain records, secrets, source URLs, or physical paths. Periodic records follow Clinker’s own clock rather than internal engine activity, so a run inside one long operation keeps producing them; they stay advisory and bounded and never replace the parent’s own heartbeat. Progress records below specifies the counts and what a consumer may conclude from them. Failed terminals add a stable failure code, broad category, sanitized message, and retry_with_backoff, do_not_retry, or policy_required advice. Before a publication-aware terminal, bounded publication_artifacts records carry the path-free inventory in ordered chunks. Artifact entries contain only artifact_id, kind, and state. The terminal carries the attempt’s completeness, cleanup-debt count, total artifact count, and counts by artifact state. Every NDJSON record, including a maximum-cardinality inventory chunk and its terminal summary, is at most 16 KiB.

One invocation that reaches a terminal without running an attempt is supported: the plan-only --lineage <FILE> export, which preflights the identity policy, writes its document, and returns before any data is read. It shares this stream’s execution_id and batch_id, so the exported document is correlatable with the invocation that produced it, and it closes with completed / success / exit 0 carrying an explicit empty publication — zero artifacts, every state count zero, no cleanup debt. An absent inventory would be read on that row as publication complete for a run that published nothing; an empty one says the same thing the reconciliation table already has vocabulary for. Its document is written and flushed before that terminal is attempted, so a terminal the pipe refuses there is reconciled exactly as a published run’s is: exit 4 with infrastructure.delivery.unreportable_outcome and policy_required, never retry_with_backoff, because the export already exists on disk. Modes that write their own document to standard output — --explain, --dry-run, -n, --lineage -, --lineage-events - — are refused at admission with exit 1 before any record is written.

A required lifecycle record that cannot be delivered while nothing has yet been read, written, or staged ends the run at exit 130 with a cancelled terminal, and the attempt’s final paths are unchanged. That applies to started, to the planning transition, and to plan_resolved, on a run and on the plan-only --lineage <FILE> export alike: a supervisor that stops reading during plan compile learns the same thing about the same condition whichever it asked for.

After plan resolution, Clinker starts the required machine-progress worker before source discovery, staging, attempt creation, sink writes, or lifecycle START. If that worker cannot be created, the stream ends with exactly one infrastructure.runtime.transient failed terminal and exit 4; no run effect has started, and a supervisor may retry with backoff.

A consumer must reject an unsupported schema major. Within schema 1 it may ignore additive fields and unknown nonterminal event kinds, but those additions carry no completion or failure meaning. Missing required fields, malformed UTF-8 or JSON, non-monotonic sequence, identity changes, duplicate terminals, or EOF without a terminal make the attempt incomplete.

Records reach the stream whole or not at all, and a parent that reads slowly does not change what the stream says. A record the pipe refuses is not written, takes no seq with it, and is not what the next record is numbered after, so a lost advisory observation leaves the numbering dense rather than shifting every record after it. A record the pipe accepts only part of is completed by the next write of this stream — never restarted — so a momentarily full pipe cannot produce two copies of one record or two terminals. What was already delivered is likewise not repeated: a terminal retried after a refusal sends only the inventory chunks the pipe has not taken, so each chunk index appears exactly once ahead of the terminal that counts them.

Terminal, exit, and artifact reconciliation

Neither a terminal event nor a process status is sufficient alone. Accept a controlled outcome only when the supported terminal family, its embedded exit where present, the actual child status, and current-attempt artifact evidence agree:

Terminal evidenceRequired process statusArtifact interpretationAdapter result
completed, result success, exit 00Publication is complete; every reported artifact is individually complete.Success.
completed, result completed_with_dlq, exit 22Publication is complete and includes the reported complete DLQ artifact.Completed under the caller’s data-quality policy.
failed, embedded exit 1, 3, or 4The same exitUse the exact reported publication state. When publication is absent the state could not be reported; infer nothing about the visible set from its absence. A reported visible subset, if any, consists only of individually complete artifacts.Failure; apply the typed retry advice and caller policy.
cancelled130Graceful cancellation won before publication; final paths for this attempt remain unchanged.Cancellation.
No terminal, malformed stream, unsupported major, duplicate terminal, forced termination, or mismatched exitAnyDo not infer current-attempt success from a pre-existing final or a visible complete subset.Incomplete attempt.

Exit 4 is intentionally broad; the typed failed terminal distinguishes retry advice without requiring the parent to parse rendered diagnostics. EOF alone never proves success. A control-pipe failure can prevent terminal delivery, so even an otherwise plausible exit remains incomplete without matching terminal evidence.

That is also why a failed terminal delivery is never reported as a transient runtime fault. When a run has published and only the terminal saying so cannot be written, a retried terminal that does get through reports exit 4 with infrastructure.delivery.unreportable_outcome and policy_required — never retry_with_backoff — and carries the publication the refused terminal carried. The failure is on the reporting channel, not in execution: the finals are visible, the lineage and OTLP terminals for the same run recorded its completion, and re-running the batch would duplicate published data. A supervisor reconciles it as a failure whose artifact evidence is complete and whose repetition is a policy decision, not an automatic one. When neither terminal reaches the stream the attempt is incomplete by the table above, which is the same reconciliation it always was.

The terminal family follows the exit code, and the completed family covers only exit 0 and exit 2. Any other non-cancellation exit is written as a failed terminal carrying the run’s own failure classification, including after an earlier failed emission that could not be encoded and so left the single terminal slot free. A non-zero exit is never restated as result success; if no terminal can be encoded at all, the stream ends without one and the attempt is incomplete by the table above.

Publication is atomic per artifact, not for the artifact set. A failure or uncontrolled stop during multi-artifact publication can leave an exact subset of newly promoted, individually complete finals visible. Consumers that need a complete set must wait for reconciled success and verify the expected artifact inventory. Reassemble artifact chunks only when every chunk index from zero to chunk_count - 1 is present in sequence before the terminal and the assembled count matches the terminal’s artifact_count and state_counts. Never treat a partial terminal stream as set-wide success.

Language-neutral adapter loop

  1. Launch one child with a stable caller-owned batch_id, piped stdout and stderr, and an overall attempt deadline. Retain only bounded sanitized tails for diagnostics.
  2. Drain stdout and stderr concurrently. Parse stdout incrementally as UTF-8 NDJSON and validate the first record’s protocol major, identities, and sequence.
  3. Heartbeat the external scheduler on an independent cadence. Report only the latest sanitized identity, sequence, logical phase/counts, and snapshot age; do not wait for or translate a Clinker progress event into a heartbeat.
  4. Enforce a total Start-to-Close-style deadline for the whole process. A no-progress timeout is an additional explicit deployment policy, not a substitute for the overall deadline.
  5. On cancellation or deadline, deliver the platform’s real graceful signal to the direct child, keep both pipes draining, and start a separately bounded cancellation grace period. If grace expires, force termination exactly once. A forced stop is incomplete, not cancelled or successful.
  6. Always wait for and reap the direct child before joining both drain tasks. Accept an outcome only through the terminal, process-status, and artifact reconciliation table above.
  7. On retry, launch a completely fresh process from the beginning of the input with the same batch_id. Require a new execution_id; retain no Clinker progress event as checkpoint or resume state.

The heartbeat interval must be below the scheduler’s heartbeat timeout. The overall attempt deadline must cover ordinary execution and publication; the grace period is a separate bounded interval for cooperative cancellation before forced termination.

Progress records

A progress event carries a progress object and a truncation object:

{"event":"progress","seq":7,
 "progress":{"phase":"executing","kind":"periodic","elapsed_ms":1031,
             "records_read":200000,
             "bytes_read":27000835,"bytes_total":81777788,
             "files_done":0,"files_total":1},
 "truncation":{"detail":false,"events":false}}

kind is transition for a lifecycle edge (planning, executing, finalizing, publishing) and periodic for an advisory observation inside a phase.

The counts

records_read is the number of source records read so far, across every source. It never decreases within a run, and the last one a run emits is the number of records that run read. It has no companion total, and no total will be added. The terminal carries no record count of its own: a progress record is the only place this stream reports one, so there is nothing to reconcile it against and no disagreement to arbitrate. A source is read as a stream, so its record count is not established until its last record has been read: any “records remaining” Clinker could publish mid-run would be a guess presented as a measurement. A supervisor answers is this run moving by comparing records_read between two events, which is the question the record is here to answer. When will it finish is not a question this stream answers.

bytes_read and bytes_total are the denominator to prefer. bytes_total is the summed on-disk size of every input, read from file metadata before anything is opened, so it is measured rather than estimated. bytes_read advances within a file, not only at file boundaries, which is what makes it useful on the common single-large-file run where every other count sits still until the end.

bytes_total is never 0. A run with no bytes to read publishes null instead, because a zero total is not a denominator — dividing by it yields NaN, not 0%, and a run with nothing to read is not a run that is nought per cent through anything.

bytes_total is null in three cases. First, when any source cannot supply a size — a network source, or a path whose metadata will not read. Second, when any source is read more than once: json and xml sources re-open their input to pre-scan the envelope before streaming the body, so their bytes cross the counter twice and the count becomes IO performed rather than input consumed. Rather than publish a total the count will overrun, Clinker withdraws it. Third, as above, when the run’s inputs are empty. bytes_read is still emitted in every one of those cases and still rises monotonically — it remains a usable liveness signal, just not a fraction.

bytes_read counts file-backed input only. A source with no bytes on disk — a network source, for instance — contributes nothing to it, so a run reading only from such a source reports bytes_read: 0 for its whole life while records flow normally. A null bytes_total is the signal that this may be so: when bytes_total is null, judge liveness from records_read, not from bytes. Reading a still 0 byte count as a stalled run is the one mistake this field invites.

Like files_total, bytes_total is null on the earliest records of a run, before source discovery has established it. It is written once and never changes afterwards — including never reverting to null — so a consumer may cache it on first sight, and its arrival mid-stream is normal rather than an identity change.

files_done and files_total are a second, independent denominator, useful where the byte one is withdrawn. A source’s file set is enumerated at startup, so the count is known rather than estimated. files_total is null when any source reads from something other than an enumerated file set.

Neither denominator is a substitute for the other and they are never combined. A denominator covering only part of a run’s work is withdrawn rather than published, because nothing on the wire would say which part it covered.

files_total is also null on the earliest records of a run, before source discovery has completed. It is written once and never changes afterwards, so a consumer may cache it on first sight; it becoming non-null mid-stream is normal and is not an identity change.

Two things a consumer must not do with these counts:

  • Do not treat either ratio as a completion percentage. Both measure input consumed, not work finished. They reach 100% while sort merges, aggregate finalization, and output publication are still running, which is exactly the shape of a progress bar that sits at 100% and appears to hang.
  • Do not divide by a null total. Absence means unknown for this run, not zero and not an error. Clinker publishes no percentage of its own for this reason: a ratio against an absent denominator is the one value that turns a missing number into a wrong one.

The record cap and the cadence floor

Periodic records are bounded twice. At most one is emitted per second, and at most 128 per run. After the 128th, Clinker emits exactly one further record with truncation.events set to true, and then stops emitting periodic records for the rest of the run. Transitions are never capped, so finalizing and publishing still arrive, as does the terminal.

Because the two bounds compose, any run longer than roughly two minutes stops producing periodic progress well before it ends. This is deliberate: it is what keeps the stream bounded regardless of how long a run takes.

truncation.events is the in-band signal for it. A consumer that sees it should conclude that the periodic stream has ended and the run is continuing normally. It must not conclude that the run has stalled, and must not treat the silence that follows as evidence of a hung process. A supervisor that kills or retries a run on progress silence will kill healthy long runs, and for a pipeline that writes output, a retry means the work is done twice.

Liveness is the parent’s own responsibility, on the parent’s own clock — see step 3 of the adapter loop. The run’s real outcome is the terminal record, which always arrives.

Cancellation and process trees

Clinker installs handlers for SIGINT and SIGTERM. If graceful cancellation wins the atomic gate before publication, it exits 130 and leaves finals unchanged. If publication wins first, later signals do not relabel or erase already complete visible artifacts; the bounded promotion finishes with completed or failed truth.

A cancellation is reported as a cancellation on every surface that reports it, and which source noticed the signal does not change that. A file source drains to a chunk boundary and stops; a source that observes cancellation while a request is already in flight — a REST page read, for instance — unwinds from inside that read. Both produce exit 130, a cancelled machine terminal, and an OpenLineage ABORT. Cancellation is never reported as an engine failure class, so alerting keyed on the lineage or OTLP terminal does not page for an operator-initiated stop.

Exit 130 also covers one non-signal case: a required machine lifecycle record that could not be written before publication. Clinker refuses to publish an outcome it cannot report, so a broken control pipe stops the attempt with finals unchanged. Discardable records are excluded from that rule — a lost periodic progress observation is reported on stderr and the run continues to its real outcome, and never converts a completed run into a cancellation.

Signal-handler installation is admission-critical. If installation fails, Clinker exits with infrastructure status 4 before opening the machine protocol stream or touching sources, staging, attempts, outputs, or lineage.

The direct-child contract is exercised on Linux with a real SIGTERM rather than a closed control pipe or an in-process cancellation shortcut. The cooperative case verifies exit 130, a matching cancelled terminal, unchanged final paths, and child reaping while both bounded drains remain live. The uncooperative case verifies that the grace interval is independent of the overall attempt deadline, force happens once only after that interval, the direct child is reaped before drain joining, and the attempt remains incomplete. Process groups and descendant ownership are deliberately outside that direct-child proof.

The adapter owns the platform termination domain. On POSIX, an adapter that creates descendants may place them in a process group and signal that group. On Windows, such an adapter may use a Job Object, including kill-on-close policy. These are adapter responsibilities: Clinker does not ship a process-group manager or Job Object owner. The parent must still reap the direct Clinker child after graceful or forced termination.

Retry and identity boundaries

batch_id is correlation only. Each invocation generates a fresh execution_id and starts input from the beginning. Clinker has no cross-attempt checkpoint, resume cursor, deduplication state, distributed transaction, or exactly-once guarantee. Safe application retry depends on stable input and the destination’s chosen if_exists policy.

Cancellation and forced termination preserve failed-attempt evidence under the configured storage retention policy. An incomplete attempt keeps its staging directory and manifest without changing pre-existing finals; the manifest’s eligible_after timestamp, 24 hours by default, controls when ordinary cleanup may reclaim it. An immediate retry does not resume or mutate that directory: it starts a new attempt with a fresh execution_id and its own staging state.

Progress is advisory liveness evidence, not a durable checkpoint or external heartbeat. Machine events are control evidence, not a secret-bearing event bus or compliance log. Ordinary users remain free to run the standalone command without --machine or any supervisor.

See also

CSV-to-CSV Transform

This recipe reads employee data from a CSV file, computes salary tiers using CXL expressions, and writes the enriched result to a new CSV file.

Input data

employees.csv:

id,name,department,salary
1,Alice Chen,Engineering,95000
2,Bob Martinez,Marketing,62000
3,Carol Johnson,Engineering,88000
4,Dave Williams,Sales,71000
5,Eva Brown,Marketing,58000
6,Frank Lee,Engineering,102000

Pipeline

salary_tiers.yaml:

pipeline:
  name: salary_tiers

nodes:
  - type: source
    name: employees
    config:
      name: employees
      type: csv
      path: "./employees.csv"
      schema:
        - { name: id, type: int }
        - { name: name, type: string }
        - { name: department, type: string }
        - { name: salary, type: int }

  - type: transform
    name: classify
    input: employees
    config:
      cxl: |
        emit id = id
        emit name = name
        emit department = department
        emit salary = salary
        emit level = if salary >= 90000 then "senior" else "junior"
        emit salary_band = match {
          salary >= 100000 => "100k+",
          salary >= 90000 => "90-100k",
          salary >= 70000 => "70-90k",
          _ => "under 70k"
        }

  - type: sink
    name: report
    input: classify
    config:
      name: report
      type: csv
      path: "./output/salary_report.csv"

error_handling:
  strategy: fail_fast

Run it

# Validate first
clinker run salary_tiers.yaml --dry-run

# Preview output
clinker run salary_tiers.yaml --dry-run -n 3

# Full run
mkdir -p output
clinker run salary_tiers.yaml

At revision 3b343a4e, bounded preview of this example can fail with an internal node-buffer cleanup error; the ordinary run produces the output below. See the preview limitation.

Expected output

output/salary_report.csv:

id,name,department,salary,level,salary_band
1,Alice Chen,Engineering,95000,senior,90-100k
2,Bob Martinez,Marketing,62000,junior,under 70k
3,Carol Johnson,Engineering,88000,junior,70-90k
4,Dave Williams,Sales,71000,junior,70-90k
5,Eva Brown,Marketing,58000,junior,under 70k
6,Frank Lee,Engineering,102000,senior,100k+

Key points

Schema declaration. The source node declares the schema explicitly with typed columns. This enables compile-time type checking of CXL expressions – if you write salary + name, the type checker catches the error before any data is read.

Emit statements. A Transform adds or replaces the fields it emits. Other input columns can pass through. To select the public output columns deliberately, use the Sink’s mapping/exclude settings and include_unmapped: false; an emit list alone is not a data-leakage boundary.

Match expressions. The match block evaluates conditions top to bottom and returns the value of the first matching arm. The _ wildcard is the default case and must appear last.

Error handling. The fail_fast strategy aborts the pipeline on the first record error. For production pipelines processing dirty data, consider continue instead, which routes the failing record to the dead-letter queue and keeps going – see Error Handling & DLQ.

Variations

Filtering records

Add a filter statement to exclude records:

  - type: transform
    name: classify
    input: employees
    config:
      cxl: |
        filter salary >= 60000
        emit id = id
        emit name = name
        emit salary = salary

Records where salary < 60000 are dropped silently – they do not appear in the output or the DLQ.

Computed columns with type conversion

      cxl: |
        emit id = id
        emit name = name
        emit monthly_salary = (salary.to_float() / 12.0).round_to(2)
        emit salary_display = "$".concat(salary.to_string())

The .to_float() conversion is required because salary is declared as int and division by a float literal requires matching types.

Multi-Input Combine

This recipe enriches order records with product metadata from a separate catalog stream using a combine node. Combine is a first-class N-ary operator: every input is declared up front, and the where expression uses qualified field references (orders.product_id, products.product_id) to express the join.

Input data

orders.csv:

order_id,product_id,quantity,unit_price
ORD-001,PROD-A,5,29.99
ORD-002,PROD-B,2,149.99
ORD-003,PROD-A,1,29.99
ORD-004,PROD-C,10,9.99
ORD-005,PROD-B,3,149.99

products.csv:

product_id,product_name,category
PROD-A,Widget Pro,Hardware
PROD-B,DataSync License,Software
PROD-C,Cable Kit,Hardware

Pipeline

order_enrichment.yaml:

pipeline:
  name: order_enrichment

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./orders.csv"
      schema:
        - { name: order_id, type: string }
        - { name: product_id, type: string }
        - { name: quantity, type: int }
        - { name: unit_price, type: float }

  - type: source
    name: products
    config:
      name: products
      type: csv
      path: "./products.csv"
      schema:
        - { name: product_id, type: string }
        - { name: product_name, type: string }
        - { name: category, type: string }

  - type: combine
    name: enrich
    input:
      orders: orders
      products: products
    config:
      where: "orders.product_id == products.product_id"
      match: first
      on_miss: null_fields
      cxl: |
        emit order_id = orders.order_id
        emit product_id = orders.product_id
        emit product_name = products.product_name
        emit category = products.category
        emit quantity = orders.quantity
        emit unit_price = orders.unit_price
        emit line_total = orders.quantity.to_float() * orders.unit_price
      propagate_ck: driver

  - type: sink
    name: result
    input: enrich
    config:
      name: result
      type: csv
      path: "./output/enriched_orders.csv"

Run it

clinker run order_enrichment.yaml --dry-run
clinker run order_enrichment.yaml --dry-run -n 3
clinker run order_enrichment.yaml

Expected output

output/enriched_orders.csv:

order_id,product_id,product_name,category,quantity,unit_price,line_total
ORD-001,PROD-A,Widget Pro,Hardware,5,29.99,149.95
ORD-002,PROD-B,DataSync License,Software,2,149.99,299.98
ORD-003,PROD-A,Widget Pro,Hardware,1,29.99,29.99
ORD-004,PROD-C,Cable Kit,Hardware,10,9.99,99.90
ORD-005,PROD-B,DataSync License,Software,3,149.99,449.97

How combine works

A combine node declares every input in its input: map, binding each upstream stream to a qualifier used inside expressions:

- type: combine
  name: enrich
  input:
    orders: orders        # qualifier: upstream_node
    products: products
  config:
    where: "orders.product_id == products.product_id"
    propagate_ck: driver

The config: block carries four fields that shape behavior:

  • where – a CXL boolean expression. Every field reference must be qualified with its input name. The expression must contain at least one cross-input equality (e.g. orders.product_id == products.product_id); additional range or arbitrary conjuncts can be combined with and.
  • match – first (default), all, or collect. See below.
  • on_miss – null_fields (default), skip, or error. Applies only to records on the driving input that find no match.
  • cxl – emit statements that shape the output row. Under match: collect, this field must be empty; the combine node auto-derives the output schema.

Match modes

match: first

Emit one output row per driver record, using the first matching build-side record. This is the standard 1:1 enrichment. When no match exists, the behavior is governed by on_miss.

config:
  where: "orders.product_id == products.product_id"
  match: first

match: all

Emit one output row for every matching build-side record. This is 1:N fan-out – if a driver record matches three build records, three rows are emitted.

- type: combine
  name: expand_benefits
  input:
    employees: employees
    benefits: benefits
  config:
    where: "employees.department == benefits.department"
    match: all
    cxl: |
      emit employee_id = employees.employee_id
      emit benefit = benefits.benefit_name
    propagate_ck: driver

An employee in a department with three benefits produces three output records.

match: collect

Gather every matching build-side record into a single Array-typed field on the output row. The driver record appears once; the build matches are aggregated into a list. The cxl: body must be empty under match: collect – the combine node synthesizes the output as { driver fields..., <build_qualifier>: Array }.

- type: combine
  name: gather
  input:
    orders: orders
    products: products
  config:
    where: "orders.product_id == products.product_id"
    match: collect
    cxl: ""
    propagate_ck: driver

Use collect when you need the set of matches as a single structured value (e.g. every price history row for an order). Use all when you need one flat row per match.

Unmatched records (on_miss)

on_miss controls what happens to driver records with zero matches:

config:
  where: "orders.product_id == products.product_id"
  on_miss: null_fields   # default: emit with build fields set to null
config:
  where: "orders.product_id == products.product_id"
  on_miss: skip          # inner-join semantics: drop unmatched drivers
config:
  where: "orders.product_id == products.product_id"
  on_miss: error         # fail the pipeline on first unmatched driver

Use skip for inner-join semantics, null_fields for left-join semantics, and error for strict referential integrity where any miss should halt processing.

Composite keys

Chain multiple equalities with and to combine on more than one field. Each conjunct is a separate cross-input equality:

- type: combine
  name: match_by_region
  input:
    sales: sales
    targets: targets
  config:
    where: |
      sales.department == targets.department
      and sales.region == targets.region
    cxl: |
      emit department = sales.department
      emit region = sales.region
      emit actual = sales.amount
      emit goal = targets.goal
    propagate_ck: driver

Both equalities must hold for a record pair to match.

Equi plus residual filter

The where clause can mix equi predicates with additional filter conjuncts. Non-equality conjuncts are applied as a residual filter after the equi match:

- type: combine
  name: high_value_enrichment
  input:
    orders: orders
    products: products
  config:
    where: |
      orders.product_id == products.product_id
      and orders.amount >= 100
    match: first
    on_miss: skip
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.product_name
      emit amount = orders.amount
    propagate_ck: driver

The equi conjunct drives the hash lookup; the amount >= 100 conjunct is evaluated as a post-filter. At least one cross-input equality is required in every combine.

Multi-input combine (three or more)

Combine accepts any number of inputs. Each pair of inputs that should be related needs an explicit equality in the where clause:

- type: combine
  name: fully_enriched
  input:
    orders: orders
    products: products
    categories: categories
  config:
    where: |
      orders.product_id == products.product_id
      and products.category_id == categories.category_id
    match: first
    on_miss: null_fields
    cxl: |
      emit order_id = orders.order_id
      emit product_name = products.product_name
      emit category_name = categories.name
      emit amount = orders.amount
    propagate_ck: driver

Input order in the input: map is preserved, and downstream reasoning treats the first input as the default driving side unless a drive: hint overrides it.

Choosing the driving input

By default the planner picks a driving (probe) input and builds hash tables for the rest. Use drive: to force a specific input to be the driver – typically the larger stream, or the one whose ordering you want to preserve:

- type: combine
  name: product_driven
  input:
    orders: orders
    products: products
  config:
    where: "orders.product_id == products.product_id"
    match: first
    drive: products
    cxl: |
      emit product_id = products.product_id
      emit product_name = products.product_name
      emit sample_order_id = orders.order_id
    propagate_ck: driver

With drive: products, the pipeline emits one row per product enriched with a matching order, instead of one row per order enriched with its product.

Memory considerations

Build-side inputs are materialized in memory as hash tables keyed by the equi columns. For each non-driving input, plan for roughly 1.5-2x the raw CSV size in heap. A 50 MB product catalog typically uses 75-100 MB of hash-table memory. Tune with --memory-limit; see Memory Tuning for spill thresholds and strategy overrides.

Document boundaries

When the driver carries document boundaries – a glob: source where each file is its own document – the Combine forwards those boundaries to its output, so a per-document Aggregate after the join rolls up per driver document. See Combine – Document boundaries.

Routing to Multiple Outputs

This recipe splits a stream of order records into separate output files based on business rules. High-value orders go to one file, standard orders to another.

Input data

orders.csv:

order_id,customer,amount,region
ORD-001,Acme Corp,15000,US
ORD-002,Globex,450,EU
ORD-003,Initech,8500,US
ORD-004,Umbrella,22000,APAC
ORD-005,Stark Ind,950,US
ORD-006,Wayne Ent,3200,EU

Pipeline

order_routing.yaml:

pipeline:
  name: order_routing
  vars:
    high_value_threshold: { type: int, default: 5000 }

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: csv
      path: "./orders.csv"
      schema:
        - { name: order_id, type: string }
        - { name: customer, type: string }
        - { name: amount, type: float }
        - { name: region, type: string }

  - type: route
    name: split_by_value
    input: orders
    config:
      mode: exclusive
      conditions:
        high: "amount >= $vars.high_value_threshold"
      default: standard

  - type: sink
    name: high_value_output
    input: split_by_value.high
    config:
      name: high_value_output
      type: csv
      path: "./output/high_value.csv"

  - type: sink
    name: standard_output
    input: split_by_value.standard
    config:
      name: standard_output
      type: csv
      path: "./output/standard.csv"

Run it

clinker run order_routing.yaml --dry-run
mkdir -p output
clinker run order_routing.yaml

Expected output

output/high_value.csv:

order_id,customer,amount,region
ORD-001,Acme Corp,15000,US
ORD-003,Initech,8500,US
ORD-004,Umbrella,22000,APAC

output/standard.csv:

order_id,customer,amount,region
ORD-002,Globex,450,EU
ORD-005,Stark Ind,950,US
ORD-006,Wayne Ent,3200,EU

How routing works

Port syntax

Route nodes produce named output ports. Downstream nodes reference these ports using dot syntax: split_by_value.high and split_by_value.standard.

The port names come from two places:

  • Condition names in the conditions map (here, high)
  • The default field (here, standard)

Exclusive mode

With mode: exclusive, each record goes to exactly one branch. Conditions are evaluated top to bottom – the first matching condition wins, and the record is sent to that port. Records that match no condition go to the default port.

Pipeline variables

The threshold is defined in pipeline.vars and referenced in the CXL expression as $vars.high_value_threshold. This makes it easy to adjust the threshold without editing the route condition, and channel overrides can change it per environment.

Variations

Multiple branches

Route nodes can have any number of named branches:

  - type: route
    name: split_by_region
    input: orders
    config:
      mode: exclusive
      conditions:
        us: "region == \"US\""
        eu: "region == \"EU\""
        apac: "region == \"APAC\""
      default: other

  - type: sink
    name: us_output
    input: split_by_region.us
    config:
      name: us_output
      type: csv
      path: "./output/us_orders.csv"

  - type: sink
    name: eu_output
    input: split_by_region.eu
    config:
      name: eu_output
      type: csv
      path: "./output/eu_orders.csv"

  # ... additional outputs for apac, other

Transform before output

Insert a transform between the route and output to shape the data differently per branch:

  - type: transform
    name: enrich_high_value
    input: split_by_value.high
    config:
      cxl: |
        emit order_id = order_id
        emit customer = customer
        emit amount = amount
        emit priority = "URGENT"
        emit review_required = true

  - type: sink
    name: high_value_output
    input: enrich_high_value
    config:
      name: high_value_output
      type: csv
      path: "./output/high_value.csv"

Combining routing with aggregation

Route first, then aggregate each branch independently:

  - type: aggregate
    name: high_value_summary
    input: split_by_value.high
    config:
      group_by: [region]
      cxl: |
        emit total = sum(amount)
        emit count = count(*)

This produces a per-region summary of high-value orders only.

Aggregation & Rollups

This recipe demonstrates grouping records and computing summary statistics. The pipeline filters active sales records, then rolls them up by department.

Input data

sales.csv:

id,department,amount,status,rep
1,Engineering,5000,active,Alice
2,Marketing,3000,active,Bob
3,Engineering,7000,active,Carol
4,Sales,4000,inactive,Dave
5,Marketing,2000,active,Eva
6,Engineering,9500,active,Frank
7,Sales,6000,active,Grace
8,Marketing,1500,inactive,Hank

Pipeline

dept_rollup.yaml:

pipeline:
  name: dept_rollup

nodes:
  - type: source
    name: sales
    config:
      name: sales
      type: csv
      path: "./sales.csv"
      schema:
        - { name: id, type: int }
        - { name: department, type: string }
        - { name: amount, type: float }
        - { name: status, type: string }
        - { name: rep, type: string }

  - type: transform
    name: active_only
    input: sales
    config:
      cxl: |
        filter status == "active"

  - type: aggregate
    name: rollup
    input: active_only
    config:
      group_by: [department]
      cxl: |
        emit total = sum(amount)
        emit count = count(*)
        emit average = avg(amount)
        emit maximum = max(amount)
        emit minimum = min(amount)

  - type: sink
    name: report
    input: rollup
    config:
      name: report
      type: csv
      path: "./output/dept_totals.csv"
      sort_order: [{ field: department, order: asc }]

Run it

clinker run dept_rollup.yaml --dry-run
mkdir -p output
clinker run dept_rollup.yaml

Expected output

output/dept_totals.csv:

department,total,count,average,maximum,minimum
Engineering,21500,3,7166.666666666667,9500,5000
Marketing,5000,2,2500,3000,2000
Sales,6000,1,6000,6000,6000

One row per department, ordered by the Sink’s explicit sort_order. The average column is a binary float and is not implicitly rounded to two decimal places. The inactive records (Dave’s $4000, Hank’s $1500) are excluded by the filter.

How aggregation works

Group-by keys

The group_by field lists the columns that define each group. Records with the same values for all group-by columns are aggregated together. The group-by columns appear automatically in the output – you do not need to emit them.

Aggregate functions

Available aggregate functions in CXL:

FunctionDescription
sum(expr)Sum of values
count(*)Number of records
avg(expr)Arithmetic mean
min(expr)Minimum value
max(expr)Maximum value
first(expr)First value encountered
last(expr)Last value encountered

Per-document aggregation

Per-document aggregation works after any upstream that forwards document boundaries — a document-aware source (a glob: / paths: source that treats each file as its own document, or an enveloped format like XML or EDI), a Merge, or a Combine. The Aggregate produces one set of grouped rows per document rather than a single aggregate spanning every file. Each document’s groups finalize and emit at that document’s close boundary, so a glob over twelve monthly files yields twelve independent monthly roll-ups. A plain single-file source is one document and still emits a single aggregate. This holds whether the Aggregate’s upstream streams or materializes its output — both flush per document.

A Merge of distinct single-document sources forwards each source’s per-document close on every mode — concat, seeded interleave, and the fused unseeded all-Source interleave fast path — so each source flushes its own roll-up. And per-document aggregation also works after a Combine on any join strategy — for example a driver glob: source joined to a lookup table, then a group-by Aggregate, yields one roll-up per driver document:

nodes:
  - type: source            # driver: each monthly file is a document
    name: orders
    config: { name: orders, type: csv, glob: "./orders/*.csv", schema: [ ... ] }
  - type: source            # small lookup table
    name: products
    config: { name: products, type: csv, path: "./products.csv", schema: [ ... ] }
  - type: combine
    name: enrich
    input: { orders: orders, products: products }
    config: { where: "orders.product_id == products.product_id", match: first, on_miss: skip, propagate_ck: driver, cxl: "..." }
  - type: aggregate
    name: monthly_totals    # one roll-up per driver document (per month)
    input: enrich
    config: { group_by: [category], cxl: "..." }

See Envelopes & Document Context for the boundary rules across every Merge mode and Combine strategy.

Strategy selection

Clinker offers two aggregation strategies:

  • Hash aggregation (default): Builds an in-memory hash map keyed by the group-by columns. Works with any input order. Memory usage is proportional to the number of distinct groups.

  • Streaming aggregation: Processes records in order, emitting each group’s result as soon as the next group starts. Requires input sorted by the group-by keys. Uses minimal memory regardless of the number of groups.

The default strategy (auto) selects streaming when the optimizer can prove the input is sorted by the group-by keys, and hash otherwise. You can force a strategy:

    config:
      group_by: [department]
      strategy: streaming   # requires sorted input

See Memory Tuning for details on memory implications.

Variations

Multiple group-by keys

    config:
      group_by: [department, region]
      cxl: |
        emit total = sum(amount)
        emit count = count(*)

Produces one row per unique (department, region) combination.

Pre-aggregation transform

Compute derived fields before aggregating:

  - type: transform
    name: prepare
    input: sales
    config:
      cxl: |
        filter status == "active"
        emit department = department
        emit amount = amount
        emit is_large = amount >= 5000

  - type: aggregate
    name: rollup
    input: prepare
    config:
      group_by: [department]
      cxl: |
        emit total = sum(amount)
        emit large_count = sum(if is_large then 1 else 0)
        emit small_count = sum(if not is_large then 1 else 0)

Aggregation followed by routing

Aggregate first, then route the summary rows:

  - type: aggregate
    name: rollup
    input: active_only
    config:
      group_by: [department]
      cxl: |
        emit total = sum(amount)

  - type: route
    name: split_by_total
    input: rollup
    config:
      mode: exclusive
      conditions:
        large: "total >= 10000"
      default: small

This routes departments with over $10,000 in total sales to one output and the rest to another.

No group-by (grand total)

Omit group_by to aggregate all records into a single output row:

    config:
      cxl: |
        emit grand_total = sum(amount)
        emit record_count = count(*)
        emit average_amount = avg(amount)

Time-windowed rollups

To group records into event-time buckets, declare a watermark: on every source and a time_window: on the aggregate. Each window emits one rollup per group when it closes. Three patterns cover the common shapes; all three ship as runnable pipelines under examples/pipelines/.

Tumbling: hourly click counts

Non-overlapping one-hour buckets per user. Use when each record should contribute to exactly one reporting bucket.

examples/pipelines/tumbling_clicks.yaml:

pipeline:
  name: tumbling_clicks

nodes:
  - type: source
    name: clicks
    description: Per-user click stream with an event-time column.
    config:
      name: clicks
      type: csv
      path: ./data/tumbling_clicks.csv
      options:
        has_header: true
      watermark:
        column: event_ts
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: kind, type: string }

  - type: aggregate
    name: hourly_clicks
    description: Per-user click count, bucketed by event-time hour.
    input: clicks
    config:
      group_by: [user_id]
      time_window:
        tumbling: { size: 1h }
      cxl: |
        emit user_id = user_id
        emit n = count(*)

  - type: sink
    name: results
    input: hourly_clicks
    config:
      name: results
      type: csv
      path: ./output/tumbling_clicks.csv

error_handling:
  strategy: fail_fast

Run:

cargo run -p clinker -- run examples/pipelines/tumbling_clicks.yaml

Each hour-aligned bucket emits one row per user_id once its time window has passed. Records that arrive out-of-order land in the DLQ as late_record — add delay: on the source or allowed_lateness: on the aggregate if the input has a known out-of-order tail.

Hopping: 1-hour sums advanced every 5 minutes

Overlapping one-hour windows that move forward every 5 minutes. Use for moving averages and rolling sums where one record should contribute to multiple overlapping reports.

examples/pipelines/hopping_sliding_5m_1h.yaml:

pipeline:
  name: hopping_sliding_5m_1h

nodes:
  - type: source
    name: clicks
    config:
      name: clicks
      type: csv
      path: ./data/hopping_clicks.csv
      options:
        has_header: true
      watermark:
        column: event_ts
        delay: 5s
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: amount, type: int }

  - type: aggregate
    name: sliding_amount
    input: clicks
    config:
      group_by: [user_id]
      time_window:
        hopping:
          size: 1h
          slide: 5m
      allowed_lateness: 30s
      cxl: |
        emit user_id = user_id
        emit total = sum(amount)
        emit n = count(*)

  - type: sink
    name: results
    input: sliding_amount
    config:
      name: results
      type: csv
      path: ./output/hopping_sliding_5m_1h.csv

error_handling:
  strategy: fail_fast

Run:

cargo run -p clinker -- run examples/pipelines/hopping_sliding_5m_1h.yaml

Each record fans into ceil(size / slide) = 12 overlapping windows, so the output row count is roughly 12× the active-window record count. The source’s delay: 5s plus the aggregate’s allowed_lateness: 30s give the pipeline 35 seconds of total grace beyond strict event-time order before a record drops to the DLQ.

Session: per-user multi-source login sessions

Variable-duration windows bounded by inactivity, computed across two independent sources. Use for activity grouping where the window length is data-driven rather than clock-aligned.

examples/pipelines/multi_source_session.yaml:

pipeline:
  name: multi_source_session

nodes:
  - type: source
    name: src_web
    description: Web login events.
    config:
      name: src_web
      type: csv
      path: ./data/session_logins.csv
      options:
        has_header: true
      watermark:
        column: event_ts
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: source, type: string }

  - type: source
    name: src_mobile
    description: Mobile login events.
    config:
      name: src_mobile
      type: csv
      path: ./data/session_mobile.csv
      options:
        has_header: true
      watermark:
        column: event_ts
      schema:
        - { name: user_id, type: string }
        - { name: event_ts, type: date_time }
        - { name: source, type: string }

  - type: merge
    name: all_logins
    inputs: [src_web, src_mobile]

  - type: aggregate
    name: user_sessions
    input: all_logins
    config:
      group_by: [user_id]
      time_window:
        session: { gap: 5m }
      allowed_lateness: 30s
      cxl: |
        emit user_id = user_id
        emit logins = count(*)

  - type: sink
    name: results
    input: user_sessions
    config:
      name: results
      type: csv
      path: ./output/multi_source_session.csv

error_handling:
  strategy: fail_fast

Run:

cargo run -p clinker -- run examples/pipelines/multi_source_session.yaml

Each source declares its own watermark.column independently. A session can’t emit until both src_web and src_mobile have caught up past session_end + allowed_lateness, so the rollup waits for the slower source before closing. Drop the watermark: block on either source and the pipeline is rejected at plan time with E156.

When to pick each

KindBucket shapeTypical use
tumblingDisjoint, clock-aligned, fixed widthHourly metrics, daily rollups, billing periods.
hoppingOverlapping, clock-aligned, fixed widthMoving averages, sliding sums, anomaly detection where each record should affect multiple reports.
sessionVariable width, gap-bounded, per-keyUser sessions, telemetry burst grouping, activity envelopes where the window length is data-driven.

Slowly-Changing Dimensions (SCD Type 2)

This recipe splits over-long dimension records into closed historical rows plus a continuation row, the shape an SCD Type 2 backfill produces. It uses a Reshape node to both mutate the record that needs splitting and synthesize the continuation row, per correlation group.

It ships as a runnable pipeline at examples/pipelines/scd_type2.yaml.

The problem

Each subject (here, an employee) has a history of dimension records — benefit plans, addresses, price tiers — each with a validity window. A record whose window runs longer than it should (it was never split when the underlying fact changed) needs to be closed at the boundary, with a fresh continuation record carrying the rest of the window forward. That is a per-group transformation: the whole employee’s history is one correlation group, and the split must observe the original group, not a half-mutated one.

Reshape fits exactly: it groups by partition_by, and within each group a rule can both mutate the trigger record and synthesize a derived one — all against the original group snapshot (the no-cascade contract).

Input data

data/scd_plans.csv — each employee’s plan history, with start/end day-numbers:

employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,100,baseline
E002,200,150,baseline
E002,2000,300,baseline
E003,300,250,baseline
E004,400,380,baseline
E004,1500,400,baseline
E005,500,450,baseline

E001, E002, and E004 each have one row whose window exceeds the one-year boundary (plan_start - plan_end > 365); the others are already short enough.

Pipeline

scd_type2.yaml:

pipeline:
  name: scd_type2_backfill
  # A small budget forces the spill path even on this tiny fixture, so the
  # example also demonstrates bounded-memory Reshape. Raise or remove it for
  # production volumes.
  memory: { limit: "16K", backpressure: spill }

nodes:
  - type: source
    name: plans
    config:
      name: plans
      type: csv
      path: ./data/scd_plans.csv
      options:
        has_header: true
      schema:
        - { name: employee_id, type: string }
        - { name: plan_start, type: int }
        - { name: plan_end, type: int }
        - { name: status, type: string }

  - type: reshape
    name: backfill
    input: plans
    config:
      partition_by: [employee_id]
      order_by:
        - { field: plan_start, order: asc }
      rules:
        - name: split_long_plan
          when: "plan_start - plan_end > 365"
          mutate:
            set:
              plan_end: "plan_start"        # close the over-long window
          synthesize:
            copy_from: none
            overrides:
              employee_id: "employee_id"
              plan_start: "plan_start"
              plan_end: "plan_end"          # the rest of the window
              status: "'synthesized'"

  - type: sink
    name: out
    input: backfill
    config:
      name: out
      type: csv
      path: ./output/scd_type2.csv

error_handling:
  strategy: continue

Run it

cargo run -p clinker -- run examples/pipelines/scd_type2.yaml

Expected output

output/scd_type2.csv:

employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,1000,baseline
E001,1000,100,synthesized
E002,200,150,baseline
E002,2000,2000,baseline
E002,2000,300,synthesized
E003,300,250,baseline
E004,400,380,baseline
E004,1500,1500,baseline
E004,1500,400,synthesized
E005,500,450,baseline

Eight input rows produce eleven output rows: the three trigger rows have their plan_end closed at plan_start, and each emits one status=synthesized continuation row. The untriggered rows (E003, E005, the short rows of every employee) pass through unchanged.

How it works

One rule, two actions

The single rule’s when predicate selects the trigger rows. For each trigger row:

  • mutate.set rewrites plan_end to plan_start, closing the over-long window at its start boundary. The mutated row keeps its identity and is emitted in place.
  • synthesize derives a brand-new continuation row. copy_from: none starts from an all-null base and every column is supplied by an overrides expression, so the continuation row is fully constructed from the trigger’s values rather than copied. It is marked status=synthesized so downstream stages can tell originals from engine-derived rows.

Because Reshape applies every rule against the original group snapshot, the mutate and the synthesize both read the trigger row as it arrived — the mutation never feeds back into the synthesis.

Audit provenance

Reshape stamps $meta.synthetic, $meta.synthesized_by, and $meta.mutated_by on every output row (see Audit stamps). These stay out of the default CSV output but are available for downstream CXL — a follow-on Route or Transform can filter on $meta.synthetic to handle generated rows separately.

Bounded memory and spill

The example’s memory.limit: "16K" with backpressure: spill is deliberately tiny so the run exercises Reshape’s disk-spill path on a small fixture. Reshape buffers each employee’s group, and when the budget trips it spills the raw input records to disk and re-runs synthesis on reload — the output is identical whether a group stayed in memory or round-tripped through disk. The per-stage spill volume appears in clinker run --explain and in the post-run spill summary. For real workloads, drop the artificial limit (the default budget is 512 MB) and Reshape stays in memory until it genuinely needs to spill.

Two limits apply: a single correlation group must still fit the memory budget at finalize (the no-cascade contract reloads the whole group to apply its rules — a group larger than the budget fails loud rather than crashing), and Reshape rules cannot reference $doc document context while spill is in play (such a pipeline is rejected at compile time). Each employee’s group in this example is tiny, so neither limit is reached here. See Reshape’s memory model and Memory & Spill for the full picture.

Idempotence

Re-running this pipeline over its own output does not re-trigger: a closed row has plan_start - plan_end == 0, and a synthesized continuation row likewise sits inside the boundary, so the when predicate fires only on genuinely over-long windows. That makes the backfill safe to apply repeatedly.

Backfill, Then Cull for Review

This recipe chains two grouping operators: a Reshape node backfills each subject’s history, then a Cull node sets aside whole subjects that need manual review — routing them to a second output stream instead of dropping them.

It ships as a runnable pipeline at examples/pipelines/employee_plan_backfill.yaml.

The problem

You have run an SCD-style backfill over each employee’s benefit-plan history (closing over-long windows and synthesizing continuation rows). After backfilling, some employees end up with a large or otherwise unusual plan history that an analyst should eyeball before it lands in the clean dataset. You want two outputs:

  • a clean stream for the employees whose history looks fine, and
  • a review stream for the flagged employees — their records intact, not discarded, not errored.

That is a per-group decision based on an aggregate property of the whole group (“this employee has more than three plan rows”), and the flagged records belong on a second valid data stream, not in the dead-letter queue. Cull fits exactly: it groups by partition_by, evaluates a group-level drop_group_when predicate, and emits removed groups on a first-class removed_to side-output port.

Input data

data/employee_plans.csv — each employee’s plan history, with start/end day-numbers:

employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,100,baseline
E002,200,150,baseline
E003,50,40,baseline
E003,300,250,baseline
E003,600,550,baseline
E003,900,850,baseline

E001 has one over-long window (1000 - 100 > 365); E003 has four plan rows.

The pipeline

nodes:
  - type: source
    name: plans
    config:
      name: plans
      type: csv
      path: ./data/employee_plans.csv
      schema:
        - { name: employee_id, type: string }
        - { name: plan_start, type: int }
        - { name: plan_end, type: int }
        - { name: status, type: string }

  - type: reshape
    name: backfill
    input: plans
    config:
      partition_by: [employee_id]
      order_by:
        - { field: plan_start, order: asc }
      rules:
        - name: split_long_plan
          when: "plan_start - plan_end > 365"
          mutate:
            set:
              plan_end: "plan_start"
          synthesize:
            copy_from: trigger
            overrides:
              status: "'synthesized'"

  - type: cull
    name: flag_large_histories
    input: backfill
    config:
      partition_by: [employee_id]
      removed_to: review
      rules:
        - name: too_many_plans
          drop_group_when: "count(*) > 3"

  - type: sink
    name: out
    input: flag_large_histories         # main port — kept employees
    config: { name: out, type: csv, path: ./output/employee_plans_clean.csv }

  - type: sink
    name: review
    input: flag_large_histories.review  # side-output port — flagged employees
    config: { name: review, type: csv, path: ./output/employee_plans_review.csv }

How it works

  1. Reshape (backfill) groups by employee_id and, for the over-long window, closes it (plan_end = plan_start) and synthesizes a continuation row marked status=synthesized. E001 gains a synthesized row, ending up with three rows.
  2. Cull (flag_large_histories) groups the backfilled rows by employee_id again and evaluates count(*) > 3 over each whole group. E003 has four rows, so the whole E003 group is routed to the review side-output port; E001 (three rows) and E002 (one row) flow to the main output.
  3. Two outputs draw the two ports: out references the Cull node by name (the main port, kept employees); review references flag_large_histories.review (the side-output port, flagged employees).

The main output (employee_plans_clean.csv) carries E001 (with its synthesized continuation row) and E002; the review output (employee_plans_review.csv) carries all four of E003’s rows. Both streams carry the unchanged input schema — Cull does not widen, and the flagged records are valid rows on a normal data edge, never DLQ entries.

Expressing group-level conditions

drop_group_when is an aggregate predicate over the whole group. CXL’s bare aggregates are sum / count / min / max / avg / collect / weighted_avg — there is no bare any(). To flag a group when any row matches a condition, sum an indicator and compare to zero:

rules:
  # Flag the whole employee for review if any plan row is still flagged
  # `status == 'error'` after backfill.
  - name: any_error
    drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"

See the Cull node reference for the full predicate vocabulary, the producer-side port model, and the bounded-memory spill behavior.

File Splitting

This recipe demonstrates splitting large output files into smaller chunks, optionally keeping related records together.

Basic record-count splitting

Split output into files of at most 5,000 records each:

pipeline:
  name: monthly_report

nodes:
  - type: source
    name: transactions
    config:
      name: transactions
      type: csv
      path: "./data/transactions.csv"
      schema:
        - { name: id, type: int }
        - { name: date, type: string }
        - { name: department, type: string }
        - { name: amount, type: float }
        - { name: description, type: string }

  - type: sink
    name: split_output
    input: transactions
    config:
      name: split_output
      type: csv
      path: "./output/report.csv"
      split:
        max_records: 5000
        naming: "{stem}_{seq:04}.{ext}"
        repeat_header: true

Output files

output/report_0001.csv  (5000 records + header)
output/report_0002.csv  (5000 records + header)
output/report_0003.csv  (remaining records + header)

Naming pattern variables

VariableDescriptionExample
{stem}Base filename without extensionreport
{ext}File extensioncsv
{seq:04}Zero-padded sequence number (width 4)0001

The path field provides the template: ./output/report.csv means stem is report and ext is csv.

Header behavior

When repeat_header: true, each output file includes the CSV header row. This is the recommended setting – each file is self-contained and can be processed independently.

Grouped splitting

Keep all records with the same group key value in the same file:

      split:
        max_records: 5000
        group_key: "department"
        naming: "{stem}_{seq:04}.{ext}"
        repeat_header: true
        oversize_group: warn

With group_key: "department", the splitter ensures that all records for a given department land in the same output file. A new file starts only at a group boundary (when the department value changes), even if the current file has not reached max_records yet.

Oversize group policy

If a single group contains more records than max_records, the oversize_group setting controls behavior:

PolicyBehavior
warn (default)Log a warning and write all records for the group into one file, exceeding the limit
errorStop the pipeline with an error
allowSilently allow the oversized file

For example, if max_records is 5,000 but the Engineering department has 7,000 records, the warn policy produces a file with 7,000 records and logs a warning.

Byte-based splitting

Split by file size instead of record count:

      split:
        max_bytes: 10485760  # 10 MB per file
        naming: "{stem}_{seq:04}.{ext}"
        repeat_header: true

The splitter estimates the current file size and starts a new file when the limit is approached. The actual file size may slightly exceed the limit because the current record is always completed before splitting.

Byte-based splitting works with every output format. For formats that wrap the whole file in framing – a JSON array or an XML root element – each rotation closes the current file’s framing and reopens it in the next, so every chunk is a complete, independently valid document (its own [ ... ] array or <Root> ... </Root> tree), never a fragment.

Combined limits

Use both max_records and max_bytes together – whichever limit is reached first triggers a new file:

      split:
        max_records: 10000
        max_bytes: 5242880   # 5 MB
        naming: "{stem}_{seq:04}.{ext}"
        repeat_header: true

This is useful when record sizes vary widely. Short records might produce a tiny file at 10,000 records, while long records might hit the byte limit well before 10,000.

Full pipeline example

A complete pipeline that reads a large transaction file, filters it, and splits the output:

pipeline:
  name: split_transactions

nodes:
  - type: source
    name: transactions
    config:
      name: transactions
      type: csv
      path: "./data/all_transactions.csv"
      schema:
        - { name: id, type: int }
        - { name: date, type: string }
        - { name: department, type: string }
        - { name: category, type: string }
        - { name: amount, type: float }

  - type: transform
    name: current_year
    input: transactions
    config:
      cxl: |
        filter date.starts_with("2026")

  - type: sink
    name: chunked
    input: current_year
    config:
      name: chunked
      type: csv
      path: "./output/transactions_2026.csv"
      split:
        max_records: 5000
        group_key: "department"
        naming: "{stem}_{seq:04}.{ext}"
        repeat_header: true
        oversize_group: warn
clinker run split_transactions.yaml --force

Practical considerations

  • Downstream consumers. Splitting is useful when the receiving system has file size limits (e.g., an upload API that accepts files up to 10 MB) or when parallel processing of chunks is desired.

  • Record ordering. Records within each output file maintain their original order from the pipeline. Across files, the sequence number ({seq}) indicates the order.

  • Group key sorting. For group_key to work correctly, the input should ideally be sorted by the group key. If the input is not sorted, records for the same group may appear in multiple files. Pre-sort with a transform if needed, or accept the split-group behavior.

  • Overwrite behavior. Use --force when re-running a pipeline with splitting enabled. Without it, the pipeline aborts if any of the output chunk files already exist.

Intra-Record Closures

This recipe shows the complete intra-record fan-out shape: an NDJSON source where each record carries an array of line items, a transform that filters items by price and then fans each remaining item into its own output record, and a flat NDJSON sink ready for downstream billing.

The pieces involved:

  • Arrow-syntax closures for predicates and projections.
  • Array methods (filter, map) for in-place transformation.
  • Bracket-index access (it["sku"]) for reading fields off each map element.
  • emit each for fan-out.
  • The Sink node’s include_unmapped flag for controlling which fields reach the sink.

Input data

orders.ndjson – one JSON object per line, each carrying a nested items array:

{"order_id":"O-1","customer":"alice@example.com","items":[{"sku":"a","price":10,"qty":2},{"sku":"b","price":20,"qty":1},{"sku":"c","price":3,"qty":5}]}
{"order_id":"O-2","customer":"bob@example.com","items":[{"sku":"a","price":10,"qty":1},{"sku":"d","price":50,"qty":1}]}

Each record has two order-level fields (order_id, customer) and an items array whose elements are maps with sku, price, and qty.

Goal

For each order:

  1. Drop items priced under $5 (a sub-threshold cutoff).
  2. Fan the surviving items into one output record each, carrying the order-level identifiers plus the per-item fields.
  3. Compute the per-line revenue (unit_price * qty) for each output record.

Pipeline

billing_lines.yaml:

pipeline:
  name: billing_lines

nodes:
  - type: source
    name: orders
    config:
      name: orders
      type: json
      options:
        format: ndjson
      path: "./orders.ndjson"
      schema:
        - { name: order_id, type: string }
        - { name: customer, type: string }
        - { name: items, type: any }

  - type: transform
    name: filter_lines
    input: orders
    config:
      cxl: |
        emit order_id = order_id
        emit customer = customer
        emit item_count = items.length()
        emit kept = items.filter(it => it["price"] >= 5)

  - type: transform
    name: explode
    input: filter_lines
    config:
      max_expansion: 10000
      cxl: |
        emit each it in kept {
          emit order_id = order_id
          emit customer = customer
          emit sku = it["sku"]
          emit unit_price = it["price"]
          emit qty = it["qty"]
          emit line_total = it["price"] * it["qty"]
        }

  - type: sink
    name: lines_out
    input: explode
    config:
      name: lines_out
      type: json
      path: "./output/billing_lines.ndjson"
      options:
        format: ndjson
      include_unmapped: false
      exclude: [items, kept]

error_handling:
  strategy: continue

Run it

# Validate first
clinker run billing_lines.yaml --dry-run

# Preview the first few output records
clinker run billing_lines.yaml --dry-run -n 3

# Full run
clinker run billing_lines.yaml

Expected output

output/billing_lines.ndjson:

{"order_id":"O-1","customer":"alice@example.com","sku":"a","unit_price":10,"qty":2,"line_total":20}
{"order_id":"O-1","customer":"alice@example.com","sku":"b","unit_price":20,"qty":1,"line_total":20}
{"order_id":"O-2","customer":"bob@example.com","sku":"a","unit_price":10,"qty":1,"line_total":10}
{"order_id":"O-2","customer":"bob@example.com","sku":"d","unit_price":50,"qty":1,"line_total":50}

Order O-1’s three input items collapse to two output records (the sku=c line was filtered out because its price was below $5). Order O-2’s two items both survive the filter and produce two output records.

How it works

Filter stage. The filter_lines transform reads each order, runs items.filter(it => it["price"] >= 5) to drop sub-threshold items, and stashes the survivors in a kept field. The closure body uses bracket indexing (it["price"]) because each it is a map; bracket indexing returns null for missing keys without aborting. The same record also carries an item_count projection so downstream nodes could route or audit on the original (pre-filter) item count.

Explode stage. The explode transform contains one emit each block over kept. For each surviving item, the body emits a flat record with the order-level identifiers (order_id, customer) repeated, plus the per-item fields lifted out of it. The body has no filter or nested emit each – those are forbidden inside the block; pre-filter upstream as we did, or post-filter in a downstream transform.

include_unmapped: false. The default Sink policy is to pass every unmapped input field through. Here we set it to false so the order-level items array (carried through from the source), the item_count projection, and the intermediate kept array (used only as the fan-out source) do not leak into the per-line output. The exclude: [items, kept] list provides a belt-and-suspenders defense against future renaming.

max_expansion: 10000. Caps how many output records a single input order may produce. The default is 10000; we set it explicitly here so the value is visible in the YAML. Orders with arrays larger than the cap route to the DLQ with category expansion_limit_exceeded (see Transform Nodes -> Expansion Cap).

Variations

Pass through every input field

Remove include_unmapped: false (or set it to true) and the original order-level fields plus the intermediate kept array will appear on every output record. Useful when downstream consumers expect a complete record context, or when you need to audit what was filtered.

Emit a single record per order with the kept-items array

Drop the explode transform and route filter_lines directly to the Sink. Each output record stays at order grain, with kept carrying the post-filter array. This is the same pipeline minus the fan-out step.

Reach for .flat_map instead of two transforms

When the per-element transformation is simple enough to fit in a single closure body, flat_map collapses the filter + project + explode pattern into one expression. It produces a flat array, which downstream nodes still see as a single field on the input record; the explicit emit each is what produces multiple output records.

Rewrite a nested field in place with .set

When you want to keep the record at order grain but mutate a value buried inside it, the set map method takes a dotted/indexed path and rewrites a single leaf, leaving every sibling untouched:

    cxl: |
      emit order = order.set("items[0].sku", "A-100").set("ship.region", "us-east")

The first set overwrites the SKU of the first item; the second writes ship.region, auto-creating the ship map if the order had no ship field yet. Because set is copy-on-write, this builds a fresh order document without disturbing the upstream binding. A path that conflicts with the existing shape (descending into a scalar, or an array index past the end) yields null for that set rather than partially writing – guard with catch if a path may not match every record.

See also