Clinker
Clinker is a pure-Rust, bounded-memory batch DAG executor for CSV, JSON, XML, and fixed-width data. It reads finite inputs, drives them through a directed acyclic graph of transformation nodes one record at a time, and exits when the inputs are drained. It ships as a single static binary with no interpreter, no runtime, and no install dependencies.
Pipelines are declared in YAML. Data transformation logic is written in CXL, a custom expression language purpose-built for ETL. Together they replace legacy tools like Informatica, SSIS, Talend, and NiFi with something deterministic, lightweight, and easy to reason about.
What Clinker is, plainly
A finite batch executor with per-record streaming evaluation, not a
long-running stream processor. A pipeline run is a job: Sources read until
EOF, the DAG drains, the process exits. Within a run, stateless operators
(Transform, Route, most Combine probe-side work, Sink) evaluate records one
at a time without accumulating per-record state. Every stage is charged
against the configured RSS budget. Fused Source → Transform → Sink paths
run streaming with no per-stage materialization; non-fused boundaries
(Route fan-out, Merge fan-in, Composition bodies, diamond DAGs) materialize
records into per-stage buffers that charge against the same envelope. The
engine spills buffers to disk at 80% of the limit and fails fast with
E310 MemoryBudgetExceeded at the hard limit, naming the offending
producer. Blocking operators (Aggregate, sort, grace-hash Combine)
accumulate state inside that same budget and spill to disk when soft
and hard memory thresholds trip, rather than OOM-killing the process.
If you have used Flink, Kafka Streams, or Beam in unbounded mode: Clinker is not that. There are no watermarks against wall-clock time, no infinite-source semantics, no exactly-once delivery across restarts. The closest prior art is Pentaho Kettle / Apache Hop, Embulk, Singer, Benthos in batch mode, and Vector running file-to-file – finite ETL jobs with per-record evaluation and a hard memory ceiling.
Three pillars of what Clinker is:
- Finite inputs. Files (CSV, JSON, XML, fixed-width, EDIFACT, X12,
HL7 v2, SWIFT MT) are the canonical shape. Finite-cursor network
sources (paginated REST APIs with hard page/record caps) fit the same
model – they exhaust their cursor and EOF.
Unbounded sources (Kafka topics, Kinesis streams, Server-Sent Events,
webhooks,
tail -f-style file followers) are out of scope and will remain so. - Finite jobs. A pipeline run begins when you invoke
clinker run, drains the DAG, and exits with a status code. No long-running daemon, no service surface, no infinite event loop. - Single process. One
clinkerbinary invocation is one operating- system process. Parallelism happens inside the process via threads (std::thread, Rayon). Clinker does not spawn worker processes, does not coordinate a cluster, and does not shuffle data between machines. Scale by giving the host more cores, more RAM, and more disk – the DuckDB / Polars / Kettle model. If a single host genuinely can’t fit the work, partition the input by file or by key and run multipleclinkerinvocations from a shell script; that’s a five-line bash script, not an architectural addition.
Why Clinker?
Single binary, zero dependencies. Download it, run it. No JVM, no Python, no package manager. Runs on Linux, macOS, and Windows out of the box — CI builds and tests on all three, and the spill, staging, and RSS-sampling layers have platform-specific paths so behavior is consistent across them.
Good neighbor on busy servers. Clinker enforces a strict memory ceiling (default 512 MB) so it can run alongside JVM applications, databases, and other services without competing for RAM. Aggregation spills to disk when memory pressure rises.
Reproducible output. Given the same input and pipeline, Clinker produces byte-identical output across runs. No nondeterminism from thread scheduling, hash randomization, or floating-point reordering.
Operability-first design. Per-stage metrics, dead-letter queues for error records, explain plans for understanding execution, and structured exit codes for scripting. Built for production from day one.
Two binaries:
| Binary | Purpose |
|---|---|
clinker | Run pipelines against real data |
cxl | Check, evaluate, and format CXL expressions interactively |
A taste of Clinker
Here is a complete pipeline that reads a customer CSV, filters to active customers, classifies them into tiers, and writes the result:
pipeline:
name: customer_etl
nodes:
- type: source
name: customers
config:
name: customers
type: csv
path: "./data/customers.csv"
schema:
- { name: customer_id, type: int }
- { name: first_name, type: string }
- { name: last_name, type: string }
- { name: status, type: string }
- { name: lifetime_value, type: float }
- type: transform
name: enrich
input: customers
config:
cxl: |
filter status == "active"
emit customer_id = customer_id
emit full_name = first_name.concat(" ", last_name)
emit tier = if lifetime_value >= 10000 then "gold" else "standard"
- type: sink
name: result
input: enrich
config:
name: result
type: csv
path: "./output/enriched_customers.csv"
Run it:
clinker run customer_etl.yaml
That is the entire workflow. No project scaffolding, no configuration files, no compile step. One YAML file, one command.
Next steps
- Installation – download the binary and verify it works
- Your First Pipeline – build and run a pipeline step by step
- Key Concepts – understand the mental model behind Clinker pipelines
Non-Goals
This page lists what Clinker is deliberately not. These are architectural commitments — design surfaces Clinker will not grow into, not just features that haven’t been built yet.
If you arrived here because you were considering Clinker for one of the scenarios below, the answer is “a different tool is the right fit.” Each non-goal is paired with the kind of tool that is the right fit.
Not an unbounded stream processor
Clinker reads sources that have an end. A pipeline run is a finite job: Sources read until EOF, the DAG drains, the process exits.
Out of scope:
- Kafka topics, Kinesis streams, Pub/Sub subscriptions (long-running consumers without a natural end).
- Server-Sent Events, WebSocket subscriptions, webhooks-as-input.
tail -f-style file followers.- Watermarking against wall-clock time.
- Exactly-once delivery across process restarts.
- Stateful infinite-stream windowing (tumbling / sliding / session windows over event time without a finite boundary).
Right fit instead: Apache Flink, Kafka Streams, Apache Beam in unbounded mode, Vector with streaming sources, Benthos with streaming inputs, Apache NiFi.
Not a multi-process or distributed engine
One clinker run invocation is one operating-system process. Clinker
does not spawn worker processes, does not coordinate a cluster, and does
not shuffle data between machines.
Out of scope:
- Worker-process pools on a single machine.
- Multi-machine sharded execution.
- Network shuffle between executors.
- Cluster managers (Kubernetes operators, YARN, Mesos integrations).
- Distributed memory accounting.
- Partial-failure recovery across worker boundaries.
Right fit instead: Apache Spark, Trino / Presto, Apache Flink in cluster mode, Apache Beam on Dataflow, Hadoop MapReduce.
Scaling Clinker: give the host more cores, more RAM, more disk — the
DuckDB / Polars / Kettle / Hop model. If a single host genuinely can’t
fit the work, partition the input by file or by key and run multiple
clinker invocations from a shell script. That’s a five-line script,
not an architectural addition.
Not a long-running service
Clinker is a CLI binary, not a server. There is no daemon mode, no HTTP control plane, no JDBC/ODBC listener, no UI server, no scheduled job runner inside Clinker itself.
Out of scope:
- HTTP API exposing pipeline execution.
- Built-in cron / scheduler / orchestrator.
- Persistent connection pool living across pipeline runs.
- A long-lived process accepting new pipeline submissions over a socket.
Right fit instead:
- For scheduling: cron, systemd timers, Airflow, Dagster, Prefect, Temporal.
- For HTTP-fronted ETL: any of the above orchestrators wrapping
clinker runinvocations. - For interactive queries against finite data: DuckDB, Polars, or any embedded query engine.
Orchestration by Temporal (and the others above) is supported via a
shell-out contract — the orchestrator runs clinker run as a child
process and reads its exit code, logs, and metrics. Clinker embeds no
Temporal client or worker; that coupling is a decided non-goal
(issue #622). See
Running Under a Workflow Orchestrator for
the exit-code, cancellation, and output-atomicity guarantees that contract
depends on.
Not an OLAP / SQL query engine
Clinker is a per-record expression engine with explicit nodes: in
a DAG. It does not parse SQL, does not optimize joins via cost-based
optimization across the whole pipeline, and does not present a relational
table model.
Out of scope:
- SQL parsing (the CXL language is the surface; no
SELECT ... FROMis accepted). - Cost-based join reordering across more than the local Combine node.
- Materialized views or query caching.
- Interactive query latencies under a second.
- ANSI-SQL semantics for NULL, type coercion, or aggregate behavior.
Right fit instead: DuckDB, ClickHouse, DataFusion, Trino, Postgres, or any RDBMS. If you want SQL-driven transformation over files, DuckDB is the closest single-binary alternative to Clinker for the cases where SQL is the right surface.
Not a connector marketplace
Clinker ships with a deliberately small set of source and sink types: CSV, JSON, XML, fixed-width, EDIFACT, X12, HL7 v2, and SWIFT MT files, plus a finite-cursor REST source. Writing to a network endpoint is not supported: a REST Output sink (issue #224) and finite-cursor SQL sources and sinks (#225, #226) are tracked but unbuilt. There is no plugin registry, no third-party connector store, no SaaS-API catalog.
Out of scope:
- Hundreds of pre-built SaaS integrations (Salesforce, HubSpot, Stripe, etc.).
- A central registry of community-maintained connectors.
- Schema discovery against arbitrary external APIs.
- Change-data-capture (CDC) sources.
Right fit instead: Airbyte, Fivetran, Stitch, Singer with its tap ecosystem, dlt (data load tool).
Not a streaming-CDC engine
Clinker treats each pipeline run as a fresh, finite pass over the input. It does not maintain a persistent log of source changes, does not replicate row-level changes from a database, and does not produce an append-only stream of inserts / updates / deletes.
Out of scope:
- Postgres logical replication subscriptions.
- MySQL binlog tailing.
- Debezium-style CDC stream production.
- Maintaining a target database in continuous sync with a source.
Right fit instead: Debezium, Maxwell, AWS DMS, Striim, Estuary Flow, or vendor-native CDC like Snowflake Streams.
What Clinker is
For the positive framing, see the Introduction and Key Concepts. The short version:
- A pure-Rust, single-binary, bounded-memory batch DAG executor for finite file and finite-cursor inputs.
- Per-record evaluation through a directed acyclic graph of Source, Transform, Aggregate, Route, Merge, Combine, Output, and Composition nodes.
- Pipelines declared in YAML, transformation logic written in CXL (a custom per-record expression language).
- One process, finite job, EOF-then-exit. Disk spill under memory pressure rather than OOM.
Installation
Clinker is a single static binary with no runtime dependencies. Download it,
put it on your PATH, and you are ready to go.
Binaries
Clinker ships two binaries:
clinker– the pipeline executor. This is the main tool you use to validate and run pipelines against data.cxl– the CXL expression checker, evaluator, and formatter. Use it during development to test expressions interactively, check types, and format CXL blocks.
Verify installation
After placing the binaries on your PATH, confirm they work:
clinker --version
clinker 0.1.0
cxl --version
cxl 0.1.0
Both commands should print a version string and exit. If you see
command not found, check that the directory containing the binaries is in
your PATH.
Building from source
Clinker requires Rust 1.91+ (edition 2024). If you have a Rust toolchain installed, build and install both binaries directly from the repository:
# Clone the repository
git clone https://github.com/rustpunk/clinker.git
cd clinker
# Install the pipeline executor
cargo install --path crates/clinker
# Install the CXL expression tool
cargo install --path crates/cxl-cli
This compiles release-optimized binaries and places them in ~/.cargo/bin/,
which is typically already on your PATH.
To verify the build:
cargo test --workspace
This runs the full test suite (approximately 1100 tests) and confirms everything is working correctly on your system.
Rust toolchain
The repository includes a rust-toolchain.toml that pins the exact Rust
version. If you use rustup, it will automatically download the correct
toolchain when you build.
| Requirement | Value |
|---|---|
| Rust edition | 2024 |
| Minimum version | 1.91 |
| C dependencies | None |
A working C compiler is not one of the requirements: nothing in the build
graph runs one. TLS for the rest source and the OTLP exporter goes through
rustls with the graviola provider, which ships as Rust and inline assembly
rather than the C and per-architecture assembly that other providers build,
and content hashing uses blake3’s pure-Rust SIMD paths. CI builds the
workspace, including its test and benchmark targets, with every C-compiler
environment variable pointed at a program that fails, so a dependency that
starts needing one is caught rather than noticed by whoever first builds
without a compiler installed. That job also builds a crate that deliberately
compiles C and requires it to fail, so the check is known to still work and not
merely to be passing.
Graviola supports x86_64 and aarch64. Those are the architectures Clinker
is built and tested on; another one needs a different rustls provider, and the
ones available today build C.
Optional capabilities
Three parts of Clinker are Cargo features of the clinker crate. All three are
on by default — a downloaded binary, or one built with plain
cargo install --path crates/clinker, has every one of them, and nothing in
this documentation assumes otherwise.
| Feature | What it adds |
|---|---|
rest | The rest source transport (transport: rest on a source node). |
otlp | OTLP/HTTP export of logs, metrics and spans to a collector. |
lineage | OpenLineage emission: --lineage and --lineage-events. |
They exist for deployments that want a smaller binary or a narrower dependency
graph — a build with no rest and no otlp links no HTTP client and no TLS
stack at all:
# File sources and lineage only: no network transport is compiled in.
cargo install --path crates/clinker --no-default-features --features lineage
A build without a capability still parses every construct the full one does.
What it will not do is run one silently: a pipeline that declares a rest
source, a clinker.toml that sets observability.otlp.endpoint, or a
--lineage flag is refused at validation — before any source is opened or any
output written — with a diagnostic that names what was asked for and says the
capability is not in this binary.
Turning otlp off also stops telemetry being recorded: the fixed arena is
reserved for an exporter to drain, so with no exporter there is nothing to
reserve it for, and the --machine ndjson-v1 terminal then carries no
observability field at all.
Verify a Release
Verify both the archive checksum and its build attestation before running a downloaded Clinker binary. A checksum detects changed bytes; the attestation also binds those bytes to this repository and its release workflow.
Download one platform archive
Replace vX.Y.Z and the archive name with the release you intend to install:
gh release download vX.Y.Z \
--repo rustpunk/clinker \
--pattern 'clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz*'
Each native archive has a sibling .sha256 file. The release also contains a
SHA256SUMS inventory covering all supported archives.
Check the SHA-256 digest
On Linux:
sha256sum --check clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz.sha256
On macOS:
shasum -a 256 -c clinker-vX.Y.Z-aarch64-apple-darwin.tar.gz.sha256
On Windows PowerShell:
$archive = "clinker-vX.Y.Z-x86_64-pc-windows-msvc.zip"
$expected = (Get-Content "$archive.sha256").Split()[0].ToLowerInvariant()
$actual = (Get-FileHash -Algorithm SHA256 $archive).Hash.ToLowerInvariant()
if ($actual -ne $expected) { throw "Clinker archive checksum mismatch" }
Stop if the checksum command fails. Do not extract or run the archive.
Verify build provenance
With a current GitHub CLI and network access:
gh attestation verify \
clinker-vX.Y.Z-x86_64-unknown-linux-gnu.tar.gz \
--repo rustpunk/clinker
The verification must identify rustpunk/clinker as the source repository and
the release archive itself as the attested subject. An attestation proves where
and how the bytes were built; it is not a claim that the program is free of
security defects.
If either checksum or provenance verification fails, keep the archive quarantined and report the release tag, archive name, and failing command.
Your First Pipeline
This walkthrough builds a pipeline from scratch, runs it, and explores the tools Clinker provides for validating and understanding pipelines before they touch real data.
1. Create sample data
Save the following as employees.csv:
id,name,department,salary
1,Alice Chen,Engineering,95000
2,Bob Martinez,Marketing,62000
3,Carol Johnson,Engineering,88000
4,Dave Williams,Sales,71000
2. Write the pipeline
Save the following as my_first_pipeline.yaml:
pipeline:
name: salary_report
nodes:
- type: source
name: employees
config:
name: employees
type: csv
path: "./employees.csv"
schema:
- { name: id, type: int }
- { name: name, type: string }
- { name: department, type: string }
- { name: salary, type: int }
- type: transform
name: classify
input: employees
config:
cxl: |
emit id = id
emit name = name
emit department = department
emit salary = salary
emit level = if salary >= 90000 then "senior" else "junior"
- type: sink
name: report
input: classify
config:
name: report
type: csv
path: "./salary_report.csv"
This pipeline has three nodes:
employees(source) – reads the CSV file and declares the schema.classify(transform) – passes all fields through and adds alevelfield based on salary.report(sink) – writes the result to a new CSV file.
The input: field on each consumer node wires the DAG together. Data flows
from employees through classify to report.
3. Validate before running
Before processing any data, check that the pipeline is well-formed:
clinker run my_first_pipeline.yaml --dry-run
Dry-run parses the YAML, resolves the DAG, and type-checks all CXL expressions
against the declared schemas. If there are errors – a typo in a field name, a
type mismatch, a missing input: reference – Clinker reports them with
source-location diagnostics and stops. No data is read.
4. Preview records
To see what the output will look like without writing files, preview a few records:
clinker run my_first_pipeline.yaml --dry-run -n 2
This reads at most 2 records from this source, runs them through the pipeline,
and writes the preview to stdout without opening salary_report.csv. For a
pipeline with several Sources, the limit applies separately to each Source;
output counts can differ after filters, joins, or aggregates. Use
--dry-run-output preview.csv to select an explicit preview destination.
At revision 3b343a4e, this example’s bounded preview can fail with an internal
node-buffer cleanup error. Its ordinary run produces the output below. See
the preview limitation.
5. Understand the execution plan
To see how Clinker will execute the pipeline:
clinker run my_first_pipeline.yaml --explain
The explain plan shows the DAG topology, the order nodes will execute, per-node parallelism strategy, and schema propagation through the pipeline. This is valuable for understanding complex pipelines with routes, merges, and aggregations.
6. Run it
clinker run my_first_pipeline.yaml
Clinker reads employees.csv, applies the transform, and writes
salary_report.csv. The output:
id,name,department,salary,level
1,Alice Chen,Engineering,95000,senior
2,Bob Martinez,Marketing,62000,junior
3,Carol Johnson,Engineering,88000,junior
4,Dave Williams,Sales,71000,junior
Alice’s salary of 95,000 meets the threshold, so she is classified as
senior. Everyone else is junior.
What just happened
The pipeline executed as a streaming process:
- The source node read
employees.csvone record at a time. - Each record flowed through the
classifytransform, which evaluated the CXL block to produce the output fields. - The Sink node wrote each transformed record to
salary_report.csv.
At no point was the entire dataset loaded into memory. This is how Clinker processes files of any size under its memory ceiling.
Next steps
- Key Concepts – understand the building blocks of Clinker pipelines
- Pipeline YAML Structure – full reference for pipeline configuration
- CXL Overview – learn the expression language in depth
Key Concepts
This page covers the mental model behind Clinker pipelines. If you have experience with other ETL tools, most of this will feel familiar – but pay attention to where Clinker diverges, especially around Clinker Expression Language (CXL), per-record evaluation, and the memory budget.
Batch jobs, not unbounded streams
A Clinker run is a finite batch job. Source nodes read their files until EOF, the DAG drains, and the process exits. There are no watermarks against wall-clock time, no infinite-source semantics, no exactly-once delivery across restarts. If you have used Flink, Kafka Streams, or Beam in unbounded mode: Clinker is not that.
The word “streaming” in Clinker’s documentation always refers to per-record
evaluation within a single batch run – records flow through the graph one
at a time rather than being materialized as a whole table – not to
long-running stream-processor semantics. Internal identifiers in the codebase
(function names like streaming_output_task, config fields like
strategy: streaming, error messages, log lines) use the word in the same
row-by-row sense; if you see it in a stack trace, it is not Flink leaking
through.
Finite inputs only
Clinker reads sources that have an end. Files are the canonical shape, and
finite-cursor network sources (paginated REST APIs with hard page/record
caps) fit the same model – they exhaust their cursor and EOF. Unbounded
sources (Kafka, Kinesis, Server-Sent Events, webhooks, tail -f-style file
followers) are explicitly out of scope and will remain so.
Single process, ever
One clinker run invocation is one OS process. Parallelism happens inside
that process via threads. Clinker does not spawn worker processes, does not
coordinate a cluster, and does not shuffle data between machines. Scale by
giving the host more cores, more RAM, more disk – the DuckDB / Polars /
Kettle model. If a single host genuinely can’t fit the work, partition the
input by file or by key and run multiple clinker invocations from a shell
script; that’s a five-line script, not an architectural addition to
Clinker.
For the full list of what Clinker deliberately does not do, see Non-Goals.
Pipelines are DAGs
A pipeline is a directed acyclic graph of nodes. Data flows from sources, through processing nodes, to outputs. There are no cycles – a node cannot consume its own output, directly or indirectly.
You define the graph with each node kind’s input fields. Most consumers use a
single input:, Merge uses an inputs: list, Combine uses a named input:
map, and Envelope names its body plus optional header and trailer
streams. Clinker resolves these references, validates that the graph is
acyclic, and determines execution order automatically.
From YAML to a run
Clinker first parses and validates the pipeline, binds schemas, compiles CXL,
and produces a typed CompiledPlan. Execution entry points require that plan,
so invalid pipelines are rejected before records are read. The current runtime
then recompiles the configuration embedded in the supplied plan and executes
the newly validated artifacts. It does not yet execute the stored DAG and
other stored artifacts directly.
The locked D-01 through D-11 contract changes that lifecycle: the supplied
CompiledPlan becomes authoritative, survives execution, and may be reused
sequentially in one process while only a defined runtime envelope may refresh.
Phase 5 owns that correction and any versioned persistent plan cache. Direct
stored-plan execution and persistent caching are not current capabilities. See
Stored-plan execution and cache identity
for the current status and owner of each contract.
The nodes: list
Every pipeline has a single flat list of nodes. Each node has a type:
discriminator that determines its behavior. The eleven node types are:
| Type | Purpose |
|---|---|
source | Read data from a file (CSV, JSON, XML, fixed-width) |
transform | Apply CXL logic to reshape, filter, or enrich records |
aggregate | Group records and compute summary values (sum, count, etc.) |
route | Split a stream into named ports based on conditions |
merge | Concatenate multiple streams that share a schema |
combine | Join records across N inputs with cross-input predicates |
reshape | Mutate or synthesize records within correlation groups |
cull | Remove whole correlation groups to a side-output port |
envelope | Frame body records with document headers and trailers |
output | Write data to a file |
composition | Embed a reusable sub-pipeline |
You can have as many nodes of each type as your pipeline requires. The only constraint is that the resulting graph must be a valid DAG.
CXL is per-record ETL
CXL is Clinker’s per-record ETL expression language. Each record flows through a CXL block independently. A program maps, filters, or enriches the current record in statement order; it is not a table query or join planner.
The core statements:
emit name = expr– produce a field in the output record, adding it or replacing the field of the same name. The record’s other fields are carried through unchanged, so there is no need to writeemit id = id; a Sink can narrow what it writes (see Sink Nodes).let name = expr– bind a local variable for use in later expressions. Local variables do not appear in the output.filter condition– discard the record if the condition is false. A filtered record produces no output and is not counted as an error.distinct/distinct by field– deduplicate records.distinctdeduplicates on all output fields;distinct by fielddeduplicates on a specific field.
CXL uses and, or, and not for boolean logic – not && or ||. String
concatenation uses +. Conditional expressions use
if ... then ... else ... syntax.
System namespaces use a $ prefix: $pipeline.*, $source.*, $record.*,
$window.*, and $vars.*. These provide access to pipeline and source
context, per-record scoped state, window function state, and static
configuration respectively.
Per-record evaluation and the memory budget
Within a run, stateless operators evaluate records one at a time: a CXL block sees exactly one record, with no table-level context and no per-record state carried across records. Clinker does not load an entire file into memory before processing it. This is what “streaming” means in Clinker – row-by-row evaluation inside a finite batch job, not Flink-style unbounded stream processing.
“One at a time” describes the evaluation model, not the transport. For
efficiency, records move between stages in bounded batches (default 2048
events, tunable per stage or pipeline-wide via
batch_size) rather
than as individual messages – but each record is still evaluated
independently against the CXL block. The batch is a handoff unit, not a
window: an operator never needs the whole batch to process any single
record. (Blocking operators – Aggregate, sort, grace-hash Combine – are
the exception that do accumulate across records; see below.)
Per-record evaluation keeps per-row memory usage bounded for the
stateless parts of the graph (Transform, Route, Merge, most Combine
probe-side work, Output). Every stage is charged against the configured
RSS budget. Fused Source → Transform → Sink paths run streaming, with
no per-stage materialization, so a 100 GB CSV passes through with the
same footprint as a 100 KB CSV. A stage that hands its output to a single
downstream Sink also avoids a charged inter-stage buffer –
single-branch Route, non-fused Merge, streaming Aggregate, and the Combine
probe-side stream their result straight to the writer (see
Streaming vs. Blocking Stages).
The remaining boundaries – multi-branch Route fan-out, output that forks
to several consumers, Composition bodies, diamond DAGs – materialize
records into per-stage buffers that charge against the same budget
envelope. Every materialized buffer can spill past the soft threshold,
including buffers shared by several readers and Route/Cull output-port
buffers. Readers run sequentially over the same immutable memory-or-spill
backing; each opens one cursor, and the final reader takes the authoritative
buffer regardless of declaration or dispatch order. A consumer that needs a
full resident vector reserves that materialization first. If the overlap would
exceed the hard limit, the engine fails before allocating with a structured
E310 MemoryBudgetExceeded diagnostic that names the consumer.
Use clinker run --explain to see which nodes will materialize
(buffer: materialized) versus which will stream (buffer: streaming)
before runtime – that label is the canonical “which stages charge the
budget” signal. See the --explain reference and
the memory-tuning page.
Stateful operators must accumulate. Aggregate, sort, and grace-hash Combine cannot emit until they have seen enough input – sums need every addend, a full sort needs the last row, a hash join needs the build side complete. These operators run inside a configured RSS budget (default 512 MB) and degrade gracefully under pressure rather than OOM:
- Aggregate uses hash aggregation by default and spills partitions to disk when soft/hard memory thresholds trip. When the input is already sorted by the group key, the planner picks streaming aggregation, which requires only constant memory.
- Sort spills runs to disk and merges them.
- Combine picks among in-memory hash join, grace hash join (spilled), and IEJoin / sort-merge depending on predicates and memory pressure. A pure-range Combine (band join with no equality key) runs the block-band IEJoin, which external-sorts each side and spills a matched-output sort to disk so both its input and its result stay inside the budget.
The memory ceiling is a first-class promise. Clinker is designed to share a server with JVM applications, databases, and other services without competing for RAM.
Input wiring
Consumer nodes reference their upstream via the input: field:
- type: transform
name: enrich
input: customers # reads from the node named "customers"
Route nodes produce named output ports. Downstream nodes reference a specific port using dot notation:
- type: route
name: split_by_region
input: customers
config:
routes:
us: region == "US"
eu: region == "EU"
default: other
- type: sink
name: us_output
input: split_by_region.us # reads from the "us" port
Merge nodes accept multiple inputs using inputs: (plural):
- type: merge
name: combined
inputs:
- us_transform
- eu_transform
Combine, Envelope, and Composition have their own multi-input shapes. See Pipeline YAML Structure for the complete wiring table.
Schema declaration
Source nodes require an explicit schema: that declares every column’s name
and type:
config:
schema:
- { name: customer_id, type: int }
- { name: email, type: string }
- { name: balance, type: float }
- { name: created_at, type: date }
Clinker uses these declarations to type-check CXL expressions at compile time, before any data is read. If a CXL block references a field that does not exist in the upstream schema, or applies an operation to an incompatible type, the error is caught during validation – not at row 5 million of a production run.
Supported types include int, float, string, bool, date, and
datetime.
Error handling
A pipeline picks one error handling strategy, in the top-level
error_handling: block:
| Strategy | Behavior |
|---|---|
fail_fast | Stop the pipeline on the first error (default) |
continue | Route error records to a dead-letter queue file and continue |
When using continue, Clinker writes rejected records to a DLQ file alongside
the output. Each DLQ entry includes the original record, the error category,
the error message, and the node that rejected it. This makes diagnosing
production issues straightforward: check the DLQ, fix the data or the
pipeline, and rerun. A run that dead-letters at least one record exits with
code 2 rather than 0, so a scheduler can tell a clean run from a partial one.
See Error Handling & DLQ for the DLQ columns, the error categories, and the per-source options.
Interactive Guides
Some parts of Clinker are easier to understand by trying them than by reading about them. Each guide below is a single page that runs in your browser, on a desktop or a phone. You change a setting or tap a record, and the page shows what the engine does with it. The guides don’t run a pipeline; each one reproduces the rules described on the reference page it links back to.
Where does a null go?
How CXL works out an expression when a field is empty: and, or and not
with null, why null == null is true, and why a filter drops a record whose
condition comes out null. Reference: Null Handling.
Correlation keys
How one failing line takes the rest of its order to the DLQ, and how an
Aggregate’s group_by decides between rejecting a whole group and recomputing
totals without the failed line. Reference: Correlation Keys.
Document context
Where $doc.* values come from, how each file becomes its own document with
its own Aggregate roll-up, and what dlq_granularity: document rejects.
Reference: Document Envelope Context.
How many documents?
What an Envelope node’s preserve and concat do to the documents a Sink
writes, how a synthesized footer differs between them, how an Aggregate after
the Envelope rolls up, and which pipeline shapes E347 and E355 reject.
Reference: Envelope Nodes.
Which files get read?
How a Source turns glob:, regex:, paths: and its filters into the list of
files it reads, and in what order. Reference:
Source Nodes → Choosing files.
Is my data still sorted?
Which stages keep a Source’s declared sort_order and which drop it, and when
an Aggregate can stream instead of holding every group. Reference:
Aggregate Nodes and Source Nodes.
Route and Merge
Where each record goes when a Route’s conditions are true, not true, or fail, in exclusive and inclusive mode, and how a Merge rejoins the branches. Reference: Route Nodes and Merge Nodes.
Combine playground
Which build rows where: matches for each driver row, and what match:,
on_miss: and drive: do with them, including range joins. Reference:
Combine Nodes.
Window functions
Which rows of a partition each $window.* function reads, and why
$window.sum is a partition total while $window.cumulative_sum is the
running total. Reference: Window Functions.
Where did my rows go?
What the end-of-run line N total, N ok, N written, N dlq counts, why the
numbers often don’t add up, and which rows are in no number at all: filtered
and duplicate rows, rows folded into Aggregate groups, Combine misses and Route
branches nothing reads. Reference: Metrics & Monitoring.
Streaming vs. blocking
Which stages stream and which hold their output in a buffer, for several
pipeline shapes, with the --explain lines each one produces. Reference:
Streaming vs. Blocking Stages.
For engine developers
The Clinker Engine Internals book has three more:
- a memory system explainer, linked from its Memory Arbitration & Scheduling chapter. It covers the memory budget, back-pressure, spilling to disk and the scheduler, with a simulator.
- a range-join explainer, linked from its Combine Internals chapter. It runs the block-band IEJoin on small inputs: sorting and slicing into blocks, pruning block pairs, the kernel step by step, and the nested-loop fallback.
- a retraction-loop explainer, linked from its Retraction Protocol chapter.
It replays the commit loop that corrects an Aggregate whose
group_byleaves out a correlation-key field, on three of the engine’s test pipelines.
Pipeline YAML Structure
A Clinker pipeline is a single YAML file with three top-level sections: pipeline (metadata), nodes (the processing graph), and optionally error_handling.
Top-level shape
pipeline:
name: my_pipeline # Required — pipeline identifier
memory: # Optional — see ops/memory.md
limit: "256M" # Optional (K/M/G suffixes), default 512M
backpressure: pause # Optional, default `pause`
vars: # Optional typed static configuration
threshold: { type: int, default: 500 }
label: { type: string, default: "Monthly Report" }
date_formats: ["%Y-%m-%d"] # Optional — custom date parsing formats
rules_path: "./rules/" # Optional — CXL module search path
concurrency: # Optional
threads: 4
chunk_size: 1000
metrics: # Optional
spool_dir: "./metrics/"
nodes: # Required — flat list of pipeline nodes
- type: source
name: raw_data
config:
name: raw_data
type: csv
path: "./data/input.csv"
schema:
- { name: id, type: int }
- { name: value, type: string }
- type: transform
name: clean
input: raw_data
config:
cxl: |
emit id = id
emit value = value.trim()
- type: sink
name: result
input: clean
config:
name: result
type: csv
path: "./output/result.csv"
error_handling: # Optional
strategy: fail_fast
Pipeline metadata
The pipeline: block carries global settings that apply to the entire run.
| Field | Required | Description |
|---|---|---|
name | Yes | Pipeline identifier. Used in logs and metrics. |
memory | No | Memory-arbitrator tuning. Nested fields: limit (RSS budget, K/M/G suffixes, default 512M) and backpressure (spill/pause/both, default pause). See Memory Tuning. |
vars | No | Typed static configuration accessible in CXL via $vars.*. Each key declares type and an optional default; see Scoped Variables. |
date_formats | No | List of strftime-style patterns for date parsing. |
rules_path | No | Directory for CXL use module resolution. |
concurrency | No | threads and chunk_size for parallel chunk processing. |
metrics | No | spool_dir for per-run JSON metric files. |
date_locale | No | Unsupported. Any explicit value is rejected with E119. Use explicit date_formats entries. |
log_rules | No | Unsupported. Any explicit value is rejected with E124; configure runtime logging outside pipeline YAML. |
include_provenance | No | Unsupported. Any explicit value is rejected with E125. Use write_meta: true on each intended Output. |
These three names are admitted only far enough to produce precise, spanned
diagnostics. Empty strings/maps and false are still explicit values and are
rejected before execution; omission is the only accepted form.
Reserved metadata contract
The current status and locked owner are explicit:
| Field | Current status | Locked target and owner |
|---|---|---|
date_locale | Rejected (E119) | Remove it and express supported parsing with date_formats:. |
log_rules | Rejected (E124) | Remove it; runtime telemetry policy is not authored in pipeline YAML. |
include_provenance | Rejected (E125) | Remove it and set write_meta: true on each Output that needs a provenance sidecar. |
For provenance sidecars that work today, set write_meta: true on an Output
node. D-24 keeps that spelling. See
Approved exceptions and rejected placeholders.
The nodes list
Every pipeline has a flat nodes: list. Each entry is a node with a type: discriminator that determines its kind:
| Type | Role |
|---|---|
source | Reads data from a file |
transform | Applies CXL expressions to each record |
aggregate | Groups and summarizes records |
route | Splits records into named branches by condition |
merge | Concatenates multiple upstream branches that share a schema |
combine | Joins records across N inputs with where: predicates |
reshape | Mutates or synthesizes records within correlation groups |
cull | Removes whole correlation groups to a side-output port |
envelope | Frames body records with optional document header and trailer streams |
output | Writes records to a file |
composition | Imports a reusable transform fragment |
Node naming
Every node must have a name: field. Names must be unique within the pipeline and must not contain dots – the dot addresses something other than a node: a route branch (split.high, see below) and a node inside a composition call site (enrich.ref). The rule covers every node kind, including sources, outputs, and composition call sites, and it covers nodes declared inside a .comp.yaml body. Names are used for wiring, logging, and diagnostics.
A dotted name is refused at plan time with E010, which names the node and the name to use instead:
node name "enrich.ref" is invalid: '.' is reserved for branch references and
composition call-site paths; rename the node to "enrich_ref" (use underscores
or hyphens) and update every reference to it
Wiring by node kind
Input fields live at the node’s top level, alongside name: and type:. Their
shape is specific to the node kind:
| Node kind | Input shape |
|---|---|
source | No input field |
transform, aggregate, route, reshape, cull, output | One upstream reference in input: |
merge | Ordered list of upstream references in inputs: |
combine | Qualifier-to-upstream map in singular input: |
envelope | Required body: plus optional header: and trailer: upstream references |
composition | Required primary input: plus an inputs: map binding every required composition port; the map is authoritative for DAG wiring |
Single upstream – used by ordinary one-input consumers:
- type: transform
name: clean
input: raw_data # References the source node named "raw_data"
config: ...
Port syntax – for consuming a specific branch from a route node, use node.port:
- type: sink
name: high_value_out
input: split.high # Consumes the "high" branch of route node "split"
config: ...
Multiple upstreams – merge nodes use inputs: (plural) instead of input::
- type: merge
name: combined
inputs:
- east_processed
- west_processed
config: {}
Qualified inputs – Combine uses singular input: with a map whose keys
become CXL qualifiers:
- type: combine
name: enriched
input:
orders: clean_orders
products: product_catalog
config:
where: "orders.product_id == products.product_id"
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_name = products.name
propagate_ck: driver
Envelope and Composition have additional port semantics. See Envelope Nodes and Compositions before wiring those node kinds.
Source nodes have no input field. They are entry points – adding an input: field to a source is a parse error.
Using the wrong field or value shape for a node kind is caught at parse time by strict deserialization.
Optional fields on all nodes
Every node type supports these optional fields:
description:– human-readable text for documentation. Ignored by the engine._notes:– arbitrary metadata (JSON object). Ignored by the engine and available to external tooling.
- type: transform
name: enrich
description: "Add customer tier based on lifetime value"
_notes:
color: "#4a9eff"
position: { x: 300, y: 200 }
input: customers
config:
cxl: |
emit tier = if lifetime_value >= 10000 then "gold" else "standard"
Strict parsing
All config structs use deny_unknown_fields. If you misspell a field name – for example, writing inputt: instead of input: or stratgy: instead of strategy: – the YAML parser rejects it immediately with a diagnostic pointing to the typo. This catches configuration errors before any data processing begins.
Environment variable: CLINKER_ENV
The CLINKER_ENV environment variable can be used for conditional logic outside of pipelines (e.g., selecting channel directories or controlling CLI behavior). It is not directly referenced within pipeline YAML but is available to the channel and workspace systems.
Scoped Variables
Clinker’s scoped-variable system lets a pipeline read and write
named values at three lifetimes: the pipeline run, the source, and
the record. Each variable is declared on the Transform that writes
it via that Transform’s declares: block (type, scope, optional
default), written by the same Transform’s CXL with
emit $<scope>.<name> = ..., and read inline from any downstream
node via the $pipeline.*, $source.*, and $record.* namespaces.
The three scopes
| Scope | Lifetime | Reset | Reader namespace |
|---|---|---|---|
pipeline | Entire pipeline run | Never (per run) | $pipeline.<key> |
source | One per source file (Arc<str>-keyed) | Per source-file | $source.<key> |
record | A single record as it flows through nodes | Per record | $record.<key> |
Record-scope variables are the per-record private store: every transform
along the row’s path can read them, but they never serialize as output
columns unless explicitly re-emitted as a regular column. They are written
with emit $record.<key> = ... from a transform that declares them.
Declaring variables
A scoped variable is declared on the Transform that writes it, in that
Transform’s config.declares: list. Each entry is named, scoped, typed,
and optionally given a default that satisfies reads firing before the
writer has run:
- type: transform
name: enrich
input: orders
config:
declares:
- { name: cutoff_date, scope: pipeline, type: date, default: "2024-01-01" }
- { name: ingest_label, scope: source, type: string, default: "prod" }
- { name: fuzzy_score, scope: record, type: float }
cxl: |
emit id = id
emit $pipeline.cutoff_date = "2024-01-01"
emit $source.ingest_label = $source.file.file_stem()
emit $record.fuzzy_score = fuzzy_match(name, $pipeline.canonical_name)
Allowed types: int, float, string, bool, date, date_time.
Each (scope, name) pair must be declared on exactly one Transform —
the same pair declared on two Transforms is rejected at config-validation
time, ahead of compilation. $pipeline, $source, and $record are flat
shared namespaces; declare each name once and read it from every consumer.
The pipeline’s top-level vars: block is a separate, flat registry
for static configuration read via $vars.<key> — it does not carry the
nested pipeline: / source: / record: scopes:
pipeline:
name: order_processing
vars:
fuzzy_threshold: { type: float, default: 0.85 } # read as $vars.fuzzy_threshold
Built-in members of each scope ($source.file, $source.name, $source.row,
$source.path, $source.count, $source.batch,
$source.ingestion_timestamp; $pipeline.start_time,
$pipeline.name, $pipeline.execution_id, $pipeline.batch_id,
$pipeline.total_count, $pipeline.ok_count, $pipeline.dlq_count,
$pipeline.filtered_count, $pipeline.distinct_count) are reserved —
declaring a user variable with one of those names is rejected at
parse time.
$source.count semantics
$source.count is the per-source record total for the Source that produced the
current record. The total isn’t known until the source finishes, so you can’t
use it during per-record evaluation: a read on a mid-stream record (in a
Transform, Route, Window, or Merge) resolves to Null, while reads after the
source has finished (such as a terminal aggregate emit) resolve to the per-source
total.
This means a streaming denominator like value / $source.count yields
Null on mid-stream records. If you need a running row counter, declare
a scope: source variable on a Transform and increment it from that
Transform’s CXL instead.
Reading variables
CXL access is identical for declared and built-in keys:
- type: transform
name: filter_recent
input: orders
config:
cxl: |
emit id = id
filter received_at > $pipeline.cutoff_date
emit batch = $source.batch_id
emit confidence = $record.fuzzy_score
Reads of undeclared keys are rejected with E203 (CXL name resolution failed) at compile time, with a “did you mean” suggestion that scans the declared registry.
Writing variables
A scoped variable is written by the Transform that declares it: list the
variable in the Transform’s declares: block and assign it from the same
Transform’s CXL with emit $<scope>.<name> = <expr>. The Transform still
processes records normally — declaring and writing a scoped var is
additive to its ordinary emit/filter logic.
- type: transform
name: capture_header
input: salesforce_in
config:
declares:
- { name: batch_id, scope: source, type: string }
- { name: ingestion_label, scope: source, type: string }
cxl: |
emit id = id
emit $source.batch_id = batch
emit $source.ingestion_label = $source.file.file_stem()
- type: transform
name: row_score
input: enrich
config:
declares:
- { name: fuzzy_score, scope: record, type: float }
cxl: |
emit id = id
emit $record.fuzzy_score = fuzzy_match(name, $pipeline.canonical_name)
An emit $<scope>.<name> write to a variable the Transform does not
declare is rejected at compile time. Requiring the declares: entry keeps
the dependency between writers and readers visible at plan time.
Init phase: pre-runtime population
Set phase: init on a Transform to pre-compute a $pipeline.* or
$source.* value from a config-file source before the main run starts:
- type: source
name: config_src
config:
name: config_src
type: csv
path: config.csv
schema:
- { name: cutoff, type: int }
- type: aggregate
name: max_agg
input: config_src
config:
group_by: []
cxl: |
emit cap = max(cutoff)
- type: transform
name: precompute_cutoff
input: max_agg
config:
phase: init
declares:
- { name: cutoff_date, scope: pipeline, type: int }
cxl: |
emit cap = cap
emit $pipeline.cutoff_date = cap
Init-phase nodes must be terminal — no runtime-phase node may consume from an init-phase Transform. (Init-phase nodes can chain through init-only descendants for compositions.) Use disjoint Sources for init vs runtime when you need both: a Source shared between an init and a runtime branch only feeds the init pass.
Compile-time validation
Scoped variables are checked before the run starts. Every reference and every writer is validated, and every flow from a writer to its readers is checked against the pipeline. Each code below tells you what to fix.
| Code | What it catches |
|---|---|
| E109 | Channel targets a composition but carries vars: overrides. |
| E116 | Channel var changes an existing type, or any default mismatches its declared type. |
| E117 | Channel var name shadows a reserved system field for that scope. |
| E118 | Channel vars.source.<src> references an unknown source-node name. |
| E164 | An init-phase Transform has a runtime descendant. |
| E171 | A reader is not a transitive DAG descendant of its writer. |
| E172 | Bare $source.<custom> read downstream of a Merge or Combine. |
| E173 | Composition body reads a parent scoped var without opting in. |
| E174 | Composition _compose.scoped_vars declares a different type than the parent. |
| E175 | An init-phase node reads a runtime-only writer’s variable. |
| E203 | A reference to an undeclared scoped variable (resolver-level failure). |
Cross-Transform duplicate declares: (the same (scope, name) declared
on two Transforms) is rejected before the run starts. $pipeline,
$source, and $record are flat shared namespaces; declare each name
once and reference it from every consumer.
Each diagnostic points at the exact place you read or wrote the variable, plus the conflicting writer or parent declaration, so the report lands where you can act on it rather than in some unrelated configuration block.
Post-merge access: qualified $source.<input>.<key>
After a Merge or Combine, the bare $source.<custom> form is
ambiguous: each record carries its own source’s value, but the
reader’s intent is usually to compare across inputs. E172 rejects
the unqualified form and the qualified form is the legal alternative:
- type: transform
name: read_after_merge
input: merged
config:
cxl: |
emit id = id
emit lt = $source.left_input.left_label
emit rt = $source.right_input.right_label
The <input_name> segment matches the named input on the Combine
(its IndexMap key) or the upstream node name on the Merge.
Composition opt-in
A composition body cannot see parent scoped variables by default —
the seal is enforced by E173. To pass values across the boundary,
the composition declares the schema of parent vars it consumes in
its _compose.scoped_vars block:
# read_pipeline_var.comp.yaml
_compose:
name: read_pipeline_var
inputs:
inp:
schema:
- { name: id, type: int }
outputs:
out: tap
scoped_vars:
pipeline:
cutoff:
type: int
nodes:
- type: transform
name: tap
input: inp
config:
cxl: |
emit id = id
emit cutoff_seen = $pipeline.cutoff
The parent must declare cutoff with the matching type; mismatches
raise E174.
What scoped variables are not
These are intentional non-features:
- No persistence across runs. State is in-memory only. A pipeline run starts with declaration defaults; the writes don’t survive the process.
- No undeclared writes. A Transform may only write a scoped
variable it lists in
declares:; anemit $pipeline.xto an undeclared name is a compile error. Requiring the declaration keeps every writer visible at plan time and the writer→reader dependency explicit in the DAG. - No dynamic var creation. The set of variables is closed at plan time, by design. This bounds memory and makes the validation matrix above tractable.
Channel overrides
A channel can both override a pipeline’s declaration defaults and add
new entries across all four registries ($vars.*, $pipeline.*,
$source.*, $record.*). Each registry has its own sub-block under
vars: on a .channel.yaml, and each entry uses the same
{ type, default } shape that pipeline-side declarations use:
# Pipeline declarations
pipeline:
name: orders
vars:
fuzzy_threshold: { type: float, default: 0.85 } # $vars.*
nodes:
- type: source
name: orders_src
config: { name: orders_src, type: csv, path: in.csv,
schema: [{ name: id, type: int }] }
- type: transform
name: enrich
input: orders_src
config:
declares:
- { name: cutoff_date, scope: pipeline, type: date, default: "2024-01-01" }
- { name: ingest_label, scope: source, type: string, default: "prod" }
- { name: tier, scope: record, type: string, default: "bronze" }
cxl: |
emit id = id
# channel/acme-prod/orders.channel.yaml
channel:
target: ../../pipelines/orders.yaml
vars:
static:
fuzzy_threshold: { type: float, default: 0.95 }
pipeline:
cutoff_date: { type: date, default: "2026-01-01" }
source:
orders_src:
ingest_label: { type: string, default: "acme-prod" }
record:
tier: { type: string, default: "platinum" }
The overlay lives in the tenant’s folder (channel/acme-prod/) and is applied
with --channel acme-prod; the channel.target field is authoritative.
Override semantics (entry name already declared) require the channel’s
type to match the declared type — mismatches produce E116. Add
semantics (entry name not yet declared) extend the registry with a new
declaration. In both cases, a default that does not match the entry’s type
also produces E116. $source overrides are keyed by source-node name; an
unknown source name produces E118. The reserved-name guard
(E117) blocks channels from shadowing system fields like
$pipeline.execution_id or $source.path. Channels that target a
.comp.yaml may not carry vars: (E109 if they do).
See Channels for the full overlay rules and the channel manifest reference.
Channels
Channels make one pipeline serve many tenants. A single base pipeline is authored once; each tenant (a channel) layers its own configuration, variable defaults, and structural changes on top — without copying or editing the base YAML. The system is built for scale: thousands of per-tenant channels against one pipeline, with strict validation and per-value provenance.
A channel is a tenant. A group is a reusable overlay shared by many
channels — selected automatically from a channel’s labels, or invoked by name.
Groups and target files can contribute value clobber (config: / vars: /
resources:)
and an ordered op list (overrides:). A channel-wide manifest is narrower:
it may contain only labels plus declared config and variables.
Workspace layout
Channels live in a channel-centric workspace. A clinker.toml at the workspace
root declares the layout roots; the rest is folders of YAML:
workspace/
clinker.toml # declares the [channel] and [group] roots
pipeline/ *.yaml # base pipelines (the pipeline-default layer)
composition/ *.comp.yaml # reusable sub-pipelines
schema/ *.schema.yaml # shared schemas
group/ *.group.yaml # group overlays: selector, priority, overrides
channel/<tenant>/ # one cataloged channel resource folder
channel.cfg.yaml # required manifest: identity, targets, labels, wide values
orders.yaml # filename is descriptive; channel.target is authoritative
The channel id is the stable logical key in [catalog.channels]. A
--channel tenant.globex invocation resolves that catalog entry directly and
then selects the target file by its declared logical pipeline id. Neither a
folder name, a file basename, nor the current working directory is an identity.
clinker.toml roots
[channel]
root = "channel" # per-channel folders live under <root>/<channel-id>/
shard = "none" # enumeration layout: none (default) | first-char | hash
[group]
root = "group" # *.group.yaml definitions live here
Both tables are optional; omitting them defaults [channel].root to channel,
[channel].shard to none, and [group].root to group. shard is an
enumeration-ergonomics choice for very large channel trees (it splits the folder
fan-out); a channel is always looked up by computed path regardless of shard
scheme, so shard never changes resolution semantics.
Typed workspace catalog
The same clinker.toml declares stable logical identities in separate catalog
namespaces for rules, schemas, compositions, pipelines, channels, and typed
composition resources.
[catalog]
rules_root = "rules"
[catalog.rules]
"shared.dates" = "rules/shared/dates.cxl"
[catalog.schemas]
"shared.dates" = "schema/shared/dates.schema.yaml"
[catalog.compositions]
"etl.normalize" = "composition/normalize.comp.yaml"
[catalog.pipelines]
"daily.orders" = "pipeline/orders.yaml"
[catalog.channels]
"tenant.globex" = "channel/globex"
[catalog.resources.shared_orders]
kind = "file"
path = "data/orders.csv"
access = "read"
Logical identities are kind-scoped. The rule and schema named shared.dates
above are distinct typed resources; asking for one kind never substitutes an
entry from another kind. A missing identity or a reference through the wrong
kind fails planning and names the catalog table that must contain it.
catalog.resources is additionally descriptor-typed. The current file kind
accepts only path and access (read, write, or read-write), is admitted
under fixed catalog entry and descriptor-byte caps, and contains no credential
fields. Resource bindings use the logical key (shared_orders above), never
the path.
Every catalog path and rules root is anchored to the selected workspace (from
--base-dir or workspace discovery). Parent traversal, an absolute path outside
the workspace, and a symlink whose canonical target escapes the workspace are
rejected. Duplicate identities within one kind are rejected, as are two catalog
identities—even across kinds—that alias the same canonical file. These checks
happen before compilation, so neither lexical aliases nor symlink aliases can
create a hidden second authority for one file.
For CXL modules, an explicit rule entry is used when present; otherwise the
logical identity maps beneath one selected rules root. Root precedence is
explicit CLI --rules-path, then pipeline.rules_path, then
[catalog].rules_root, then the workspace-relative rules/ default. Selection
chooses one root rather than searching several. Planning admits the bounded
direct/transitive module closure into the compiled plan, after which execution
does not reopen module source files. See Modules and use
and the --rules-path reference.
The layer model
Every value and every op is attributed to exactly one layer. Layers apply in a fixed semantic order — never lexical or file order:
pipeline-default < group(s) by priority < channel-wide < channel-per-target
- pipeline-default — the base pipeline’s own configuration.
- group(s) by priority — every group applied to the run, ordered by
priority(higher priority applies later and thus wins). - channel-wide — the channel manifest (
channel.cfg.yaml): overlays that apply to every pipeline this channel runs. - channel-per-target — the per-target overlay file
(
<target>.channel.yaml): the highest-precedence layer.
Clobber, never deep-merge
A higher layer’s value replaces the lower layer’s value wholesale. There is
no deep-merge and no list-append: overriding a list swaps the entire list. To
override individual elements, model them as a keyed map (which the config: and
overrides: surfaces already are), not a list — so each element is addressed and
replaced by key. Every resolved value maps 1:1 back to the single layer that
supplied it, and channels resolve / explain --field report that layer.
Structural ops (overrides:) apply in a total order — layer precedence first,
then declaration order within a layer. Collisions are errors, never silent
no-ops: adding a node whose name already exists, or targeting a missing or
already-removed node, fails with a diagnostic anchored to the offending op.
Overlays are resolved before executable compilation. Structural op streams
are concatenated in total order and folded over the base node list. Clinker then
compiles that target once for typed candidate validation; only when every
config: candidate passes name, type, ambiguity, and fixed-lock checks does the
winning config map enter executable compilation. Scoped vars: are likewise
validated before executor initialization. One invocation produces one validated
effective plan.
Value clobber: config, vars, and resources
The value-clobber surface carries scalar overrides. It appears identically on a group, a channel manifest, and a per-target overlay.
config: overrides composition config knobs, keyed by node.param dotted
paths (the composition node’s name, then the parameter name):
config:
scorer.threshold: { value: 0.95 } # override the `threshold` knob of `scorer`
The override changes executed behavior, not just the rendered provenance: the
composition body reads the knob as $config.<param>,
which the planner constant-folds to the resolved value for that instantiation at
compile time. The winning layer is still recorded in the provenance side-table,
so channels resolve / explain --field continue to report which layer supplied
the value.
A config: key that matches no parameter in the compiled plan is a hard error
(E113) — a misspelled or stale key aborts the run rather than
silently doing nothing.
Rebinding a composition resource
resources: changes which logical catalog resource supplies one declared
composition slot. The key is composition-node.slot; the leaf uses the same
{ value, fixed } shape as other clobbers:
# channel.cfg.yaml, a group file, or a per-target file
resources:
lookup.orders: { value: tenant_orders }
The base composition call must already declare orders under
_compose.resources_schema, and tenant_orders must exist under
[catalog.resources] with the required kind and capabilities. Group,
channel-wide, and per-target candidates use the ordinary precedence order and
retain every attempted layer plus the winner. fixed: true locks a lower
binding against higher layers just as it does for config values.
Only a scalar logical identity is accepted as value. Inline descriptors and
credential/profile/secret/token selectors are rejected at the strict YAML
leaf. An overlay cannot introduce a slot, address an internal nested slot, or
change ports, composition names, or config through this surface. Resource
rebinding changes the semantic plan fingerprint but does not resolve
credentials or open runtime handles.
Locking a value: fixed
fixed is metadata on the value it locks, never a sibling map. A config leaf
uses { value, fixed }; a variable leaf uses { type, default, fixed }.
fixed defaults to false. Unknown spellings and a misplaced top-level
fixed: block fail at the authored line with the corrected leaf form.
# channel.cfg.yaml — the channel-wide manifest
channel:
name: tenant.globex
targets: [daily.orders]
config:
scorer.threshold: { value: 0.9, fixed: true }
# order_fulfillment.channel.yaml — the per-target overlay (a higher layer)
channel: { target: daily.orders }
config:
scorer.threshold: { value: 0.95 } # rejected: channel-wide locked this key
The per-target candidate is invalid because the channel-wide value is fixed;
the diagnostic points to the per-target leaf and the run does not start. The
resolved provenance remains 0.9, and channels resolve marks that winning
layer (fixed). Invalid candidates are validated even when another layer would
win, so a typo or type mismatch cannot hide behind precedence.
vars: overrides or adds scoped-variable defaults, using the same four scopes a
pipeline’s own vars: block uses ($vars.* / $pipeline.* / $source.* /
$record.*). Each leaf is the same { type, default } shape a pipeline
declaration uses:
vars:
static: # $vars.*
currency: { type: string, default: "USD", fixed: true }
pipeline: # $pipeline.*
cutoff_date: { type: date, default: "2026-01-01" }
source: # $source.<src>.* — outer key is the source-node name
orders:
ingest_label: { type: string, default: "prod" }
record: # $record.*
tier: { type: string, default: "bronze" }
See Variables for the scoped-variable model these overlay.
Structural ops: overrides
The overrides: surface is an ordered list of discrete, name-addressed ops
applied to the base pipeline’s node list before compilation. Each op is a
mapping with an op: discriminant. Unknown keys, or keys that belong to a
different op kind, are rejected at parse time.
The op vocabulary is add / remove / replace / set / bypass /
patch_schema.
add — splice in a node
Insert a new node, either inline or as a composition reference. The splice
anchor is exactly one of after: / before: / an explicit input:.
overrides:
# Inline transform, spliced after `normalize` (its former consumers now read `stamp`):
- op: add
node:
type: transform
name: stamp
input: normalize
config:
cxl: "emit order_id = order_id"
after: normalize
# A composition, named by `alias`, with a config knob for the injected node:
- op: add
composition: ../composition/fraud_check.comp.yaml
alias: fraud_check
after: normalize
config:
threshold: 0.8
after: X reads from X and repoints X’s former consumers onto the new node;
before: X feeds X, taking over X’s former upstream. An inline node with no
splice anchor keeps its own declared input:. Adding a node whose name already
exists is an error.
remove — delete a node and rewire
Delete a node by name, repointing its named consumers through an explicit
rewire: map so no dangling reference is left behind:
overrides:
- op: remove
target: legacy_audit
rewire:
route_priority.input: product_lookup # <consumer>.input: <new upstream>
Each rewire: key is a <node>.input path; each value is the replacement
upstream. Any consumer still referencing the removed node afterward is an error,
as is removing a node that does not exist.
bypass — remove a linear node
Sugar for remove on a 1-in/1-out node: it auto-rewires the node’s sole
consumer onto its sole upstream.
overrides:
- op: bypass
target: legacy_audit
bypass only applies to a single-input, single-consumer node; a fan-in/fan-out
node must use the explicit remove op with a spelled-out rewire: map.
replace — swap a node’s definition
Replace a whole node by name, keeping its identity (and therefore every consumer
edge) intact. The replacement node’s own name: must equal target:.
overrides:
- op: replace
target: normalize
node:
type: transform
name: normalize
input: orders
config:
cxl: "emit order_id = upper(order_id)"
set — set one field within a node
Set a single field within a named node by path. The currently addressable path
is config.cxl — the primary CXL body of a transform / aggregate /
combine node — so replacing a stage’s logic wholesale is a set, not a
special case:
overrides:
- op: set
target: route_priority
field: config.cxl
value: >
emit _route = if priority_level == "urgent"
then "priority_report" else "fulfilled_orders"
Here _route is an ordinary audit field; it does not select an Output. Direct
Outputs sharing route_priority each receive every record. To partition rows
by destination, add a Route node with conditions that read
the field (or express the conditions directly on the Route).
Any other field path is a hard error, never a silent no-op.
patch_schema — shape a source’s columns
Add / rename / modify / remove columns on a source node’s declared schema, via a column-name-keyed map (the map key is the column name). Each column carries exactly one op:
overrides:
- op: patch_schema
target: orders
schema:
amount: { type: float, scale: 2 } # modify: set any subset of attrs
cust_id: { rename: customer_id } # rename (a physical->logical alias)
order_notes: remove # drop an existing column (bare scalar)
region: { add: { type: string } } # add a new column (map key = new name)
The modify leaf is a bare attribute map: it sets any subset of the column’s
attributes (type, scale, precision, format, width, …), leaf-replace,
keeping every attribute it does not name. A typo’d attribute is rejected rather
than silently appended. The same grammar applies identically at every override
layer (pipeline / group / channel).
The keyed-map shape (rather than a list) is deliberate: a column op is addressed
and leaf-replaced by name, with first-class rename / remove / add, exactly
matching the source-config schema patch grammar so the
two surfaces resolve columns and their diagnostics identically.
rename is a source-column alias, not a bare relabel: the reader still binds
the original physical column and re-labels its value under the new name, so
downstream CXL and the output see the new name carrying the original column’s
data. A missing column, an add that collides with an existing name, or a rename
onto an existing name are all errors (E231–E233).
To see which layer set a given attribute on a patched column, trace it with
clinker explain <pipeline> --field <source>.<column>.<attribute> (optionally
--channel <name>); the output names the winning Base < Pipeline < Group < Channel layer and each shadowed one. See
Field provenance.
Groups and selectors
A group (group/<name>.group.yaml) is a reusable overlay layer that sits
between the pipeline default and the channel layers. It carries the same two
surfaces every layer carries — config: / vars: value clobber and an
overrides: op list:
group:
name: enterprise
targets:
pipelines: [daily.orders]
compositions: [etl.normalize]
match: 'tier == "enterprise"' # optional selector; higher priority wins
priority: 20
config:
scorer.threshold: { value: 0.8 }
overrides:
- op: add
node:
type: transform
name: fraud_stamp
input: normalize
config:
cxl: "emit order_id = order_id"
after: normalize
A group plays two roles under one concept:
- Selector-derived — when
match:is present, the group is applied automatically to every channel whose labels satisfy the CXL boolean. Multiple matching groups are ordered bypriority(higher wins; the default priority is0). - Standalone / explicit — when
match:is absent, the group is never auto-selected; it applies only when invoked by name with--group. Groups are channel-agnostic — their overrides never read channel labels — so any group can run standalone against the base pipeline, with or without a channel.
Every group owns a non-empty explicit targets: set of catalog pipeline and/or
composition ids. A selector only narrows that set: a matching label can never
make the group global. Forced --group use is target-bounded by the same set.
Selectors are label-only CXL
match: is a bare CXL boolean expression evaluated in a
restricted label-only context: the only names in scope are the channel’s
labels. $record / $source / $pipeline / $vars / $doc, window and
aggregate calls, now, and wildcards are all rejected, so a selector is a pure,
deterministic predicate over labels.
match: 'region == "west" and tier == "enterprise"'
Labels are typed from their YAML/JSON scalar kind (string, bool, int, float), so
the typechecker rejects label/literal type mismatches. A selector that
references a label a channel does not declare is a hard error, never a silent
false — a typo surfaces as an unresolved-identifier error rather than
quietly excluding the channel.
The channel manifest
channel.cfg.yaml declares the channel identity, its non-empty pipeline target
set, identity labels, and optional channel-wide values:
channel:
name: tenant.globex
targets: [daily.orders]
labels: { region: west, tier: enterprise } # identity — drives group selectors
config:
scorer.threshold: { value: 0.9, fixed: true }
vars:
static:
currency: { type: string, default: "USD", fixed: false }
Labels are identity, never a pipeline override. The manifest and its target
set are required. Channel-wide overrides: and sources: are forbidden because
they would apply graph/source/schema changes without a single admitted target;
move those operations into the corresponding target file.
The per-target overlay
A target file overlays exactly one manifest-declared catalog pipeline and its
admitted composition closure. The channel.target: logical id is authoritative;
the filename has no identity semantics:
channel:
target: daily.orders
config:
scorer.threshold: { value: 0.95 }
overrides:
- op: patch_schema
target: orders
schema:
tax_exempt: { add: { type: bool } }
Complete admission and execution identity
Channel loading is fail closed. Clinker canonicalizes the workspace root and candidate path, rejects traversal and symlink escapes, opens each admitted file once, verifies its post-open identity, and reads it into one bounded byte buffer. UTF-8 validation, parsing, and content identity all use that exact buffer; the bytes cannot be swapped between validation and hashing.
Planning validates the whole admitted catalog before selecting a requested pipeline or target:
- every manifest target must have exactly one target overlay;
- every declared target is parsed and validated, including targets not selected for this run;
- the complete reachable pipeline and composition closure is validated for every target; and
- group discovery, channel discovery, or file I/O errors abort admission rather than silently skipping an entry.
The planned execution identity includes the selected pipeline bytes and every applied layer in precedence order: defaults, ordered groups, the channel-wide overlay, and the selected target overlay. Group priority, declaration sequence, and whether membership was derived or explicit are part of that identity. Changing applied bytes or their order changes the identity; changing an unapplied overlay does not.
CLI surface
Running with overlays
# Run as a tenant: resolves catalog identities and derives target-bounded groups.
clinker run pipeline/order_fulfillment.yaml --channel globex --base-dir .
# Force-include a group by name, with or without a channel.
clinker run pipeline/order_fulfillment.yaml --group enterprise --base-dir .
run resolves the overlay stack from the workspace (rooted at --base-dir,
default the current directory) and folds the resolved overrides into the plan
before execution. Overlay flags shared across run and explain:
| Flag | Meaning |
|---|---|
--group <NAME> | Force-include a group overlay by name (repeatable), provided its explicit target set admits the selected pipeline or composition closure. |
--no-auto-groups | Suppress selector-derived group membership; only explicit --group overlays apply. |
--channel <ID> | Apply a logical id from [catalog.channels]; the selected [catalog.pipelines] id must appear in its manifest targets. Derives only target-admitted matching groups. |
explain --field <node.param> --group <NAME> reports the same overlay stack for
provenance lookups, mirroring run.
Inspecting overlays
channels resolve renders the effective post-overlay DAG for one target under a
chosen channel and/or groups, with per-value provenance — which layer supplied
each value and which group injected which node:
# Resolve the effective plan for the globex channel (derives matching groups from its labels)
clinker channels resolve pipeline/order_fulfillment.yaml --channel globex --base-dir .
# Preview a group overlay standalone (no channel)
clinker channels resolve pipeline/order_fulfillment.yaml --group enterprise --base-dir .
Here --channel is a logical id from [catalog.channels]. The selected pipeline
must likewise appear in [catalog.pipelines]; resolve never guesses identity
from the filename. Matching groups are considered only after their explicit
target sets admit that pipeline or one of its composition dependencies.
channels lint compiles every cataloged channel target and reports every
failure through the same resolver used by run and explain:
clinker channels lint --base-dir .
Membership and labels
# List the channels a group's selector currently matches
clinker channels group members enterprise --base-dir .
# Stamp/overwrite a label across one or more channels (idempotent)
clinker channels label set tier=enterprise globex initech --base-dir .
channels label set takes a key=value assignment; the value is typed by YAML
scalar inference (true/false → bool, integers → int, decimals → float, else
string) so numeric and boolean labels compare correctly against selectors. The
channel manifest must already exist with its explicit channel.targets list;
the command never creates a targetless manifest.
Renaming a base node
refactor rename-node renames a base node and propagates the rename to every
overlay that references it (splice anchors, target:, rewire: keys) across the
workspace:
# Preview every file that would change
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --dry-run
# Apply it, then re-lint
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --base-dir .
clinker channels lint --base-dir .
The new name must be letters, digits, and _ only.
Source config patches
Independent of the overlay op engine, a channel file can patch a source
node’s parsed config directly through a sources: block, applied before
validation and compile so the run behaves exactly as if the source YAML had been
hand-edited. This is the same column-keyed schema grammar patch_schema reuses,
plus multi-value and per-format option patches.
sources:
transactions: # source-node name (unknown -> E230)
options:
record_path: batch_records # set a scalar per-format option (bad key -> E235)
split_to_rows: # keyed by field name
items: { mode: split, position_column: line_no } # add-or-modify
tags: { position_column: ~ } # clear one attribute
line_items: remove # drop an entry (unknown field -> E234)
max_output_rows_per_input: 10000 # replace the source ceiling; 0 disables it
split_values: # keyed by field name
codes: { delimiter: "|" } # add-or-modify an entry
tags: { delimiter: ~ } # reset to the default delimiter
notes: remove # drop an entry (unknown field -> E234)
schema: # keyed by column name
amount: { type: float, scale: 2 }
cust_id: { rename: customer_id }
order_notes: remove
region: { add: { type: string } }
All ops are keyed and leaf-replace — there is no deep-merge. On an existing
split_to_rows / split_values entry a partial map is a modify: an omitted key
keeps its current value, and a new entry takes the same defaults hand-written
config would. Because an omitted key means “keep current”, clearing an attribute
that is already set needs its own form — an explicit YAML null. On
position_column that removes the attribute; on delimiter, which always holds
some separator, it restores the ; default. options are
merged onto the source’s current options and re-validated through the format’s
option struct, so an unknown or mistyped key is rejected exactly as in
hand-written config. A schema rename is a source-column alias — the same
alias a base column can declare directly with source_name::
schema:
# read the physical `cust_id` column, expose it downstream as `customer_id`
- { name: customer_id, type: string, source_name: cust_id }
Format-structure patches (X12 / HL7 v2)
Beyond the format-agnostic ops above, a sources: patch can reshape the
format-layer structures an X12 or HL7 source declares in its options: block —
with keyed add/modify/remove grammar instead of blob-replacing the whole
options map:
sources:
interchange: # an X12 source
group_section: # the GS functional-group declaration
name: fg # rename the section (omit to keep)
fields:
e04: int # set/add a typed field
e05: remove # drop a declared field
set_section: remove # drop the whole ST declaration
messages: # an HL7 v2 source
split_fields: # keyed by positional field name
f08: { components: 3 } # add-or-modify a composite split
f03: remove # drop a declared split
group_section / set_section patch the X12 nested-envelope declarations (the
GS functional-group and ST transaction-set levels); split_fields patches
the HL7 composite-field splits, keyed by positional field name and resolved by
wire position (f8 and f08 address the same split). Each op applies only to
a source of the matching format (anything else is E238). The set form is a
partial modify on an existing declaration — an omitted name or axis width
keeps its current value — and creates the declaration when absent, in which
case name (X12) or components (HL7) is required (E240). Removing a
declaration, field, or split the source does not carry is E239. These ops
apply after the options merge, so they layer on top of an options value
that replaces the same declaration in one patch.
Multi-record patches (discriminator-driven flat files)
A multi-record flat file interleaves several record layouts in one file, each
identified by a discriminator tag. A sources: patch reshapes that layout with
records: (keyed by record-type id) and a discriminator: merge, so a tenant’s
record set can differ from the base without editing the pipeline:
sources:
ledger: # a multi-record source
discriminator: { start: 2 } # move the tag byte range (partial merge)
records:
detail: { tag: X } # retag; a nested `columns:` reshapes fields
trailer: remove # drop a record type
header: # add a record type (map key = its id)
add:
tag: H
columns:
- { name: hdr_id, type: string, start: 1, width: 8 }
A records entry follows the same keyed grammar as schema: a bare remove
drops the record type, { add: { tag, columns, ... } } declares a new one, and
a bare attribute map modifies an existing one. A modify sets any subset of the
record type’s tag / parent / join_key / description and carries a nested
columns: map that runs the column-op grammar (modify / rename / add /
remove) against that record type’s own fields. The discriminator: op merges
field by field onto the current discriminator — a named field overwrites, an
omitted one is kept — and the merged result must be a byte range (start +
optional width) XOR a field.
These ops apply only to a multi-record schema (E241). Modifying or removing an unknown record-type id is E242, adding an id that already exists is E243, a merged discriminator that is neither pure byte-range nor pure field is E244, and a discriminator tag shared by two record types after the patch is E245.
Sources inside a composition body
A plain sources: key names a top-level source node. To patch a source
declared inside a composition body, qualify the key with the composition
call-site node name: <composition-node>.<source>. The composition body is
expanded during compile, so the patch is applied to the body’s source when the
body is bound — before the body typechecks — exactly as a top-level patch shapes
a top-level source before it binds:
sources:
enrich.lookups: # source `lookups` inside composition node `enrich`
schema:
code: { rename: lookup_code }
Resolution is one level deep: the qualifier must name a composition node in the
pipeline (an unknown composition — or a nested a.b.c key naming a source inside
a nested composition body — is E230), and the source half must name a source
node declared in that composition’s body (an unknown one is E230, naming the body
file). A plain unqualified key still targets a top-level source, and a name that
matches no top-level source still fails with E230 — now hinting at the qualified
form when the pipeline has compositions.
Note: an authored body Source must link to one declared composition resource slot with
resource: <slot>; a directpath,glob,regex, orpathsmatcher is rejected. The source patch above changes schema/reader configuration but does not select the resource. Planning binds the slot and compiles a call-scoped Source instance. On a data run, a credential-freefilebinding opens through the CLI-prepared catalog factory and streams via the executor’s bounded Source path. Credential-bearing bindings remain unsupported and fail before opening because no profile-selection surface is available.
When a patch changes the effective source config, the run’s pipeline identity differs from the base and from other patched variants, so their outputs and lineage do not collide.
Diagnostics
| Code | Meaning |
|---|---|
| E103 | A config: candidate has the wrong value type or attempts to override a lower fixed value, or a resource binding names an unknown/undeclared/incompatible slot or catalog identity. Every candidate is checked at its own leaf, including one a later layer would shadow. |
| E107 | A channel/group variable candidate disagrees with the pipeline declaration or its default does not match the declared type. |
| E110 | A variable candidate shadows a reserved scoped-variable name. |
| E111 | A vars.source candidate names no source in the selected pipeline. |
| E113 | A config: / override key matches no composition parameter in the compiled plan. A misspelled or stale key aborts the run instead of silently doing nothing. |
| E114 | An overlay op failed to apply (missing splice anchor, duplicate node name, missing/removed target, invalid set field, invalid bypass node). The diagnostic is anchored to the offending op’s source span, not the base pipeline. |
| E118 | A shorthand node.param candidate is ambiguous in the selected composition closure; use the exact target-specific node path. |
| E230 | A source patch (sources.<src> or patch_schema) targets a source that does not exist: an unknown top-level source, an unknown composition for a qualified <composition>.<source> key, a <composition>.<source> naming no source in that composition’s body, or a nested (a.b.c) key. |
| E231 | A schema rename / modify / remove of a column that does not exist. |
| E232 | A schema add of a column name that already exists. |
| E233 | A schema rename whose target name collides with an existing column. |
| E234 | A split_to_rows / split_values remove of a field with no matching entry. |
| E235 | An options patch sets an unknown or mistyped option key for the source’s format. |
| E236 | A renamed/aliased column’s exposed name collides with a real input field, which would mislocate that field. Raised at read time. |
| E237 | A schema patch on a multi-record / generated / external-file schema — column ops apply only to a single-record column list. |
| E238 | A group_section / set_section patch on a non-X12 source, or a split_fields patch on a non-HL7 source. |
| E239 | A remove of a nested-section declaration, declared section field, or field split the source does not carry. |
| E240 | A malformed format-structure patch: creating a nested section without a name, adding a split without components, a split key that is not a positional fNN name, or a zero axis width. |
| E241 | A records / discriminator patch on a single-record / generated / external-file schema — these ops apply only to a multi-record schema. |
| E242 | A records modify / remove of a record-type id the source does not declare. |
| E243 | A records add of a record-type id that already exists. |
| E244 | A merged discriminator that is neither a pure byte range (start + optional width) nor a pure field. |
| E245 | Two record types share a discriminator tag after the patch, which would make the reader’s discriminator dispatch ambiguous. |
Compositions
Compositions are reusable pipeline fragments that can be imported into multiple pipelines. They encapsulate common transform patterns – date derivations, address normalization, currency conversion – into self-contained, testable units.
Using a composition
A composition node in your pipeline references an external .comp.yaml file:
- type: composition
name: risk
input: orders
use: "./compositions/risk_score.comp.yaml"
inputs:
inp: orders
config:
threshold: 0.5
The use: field points to the composition definition file. The inputs: map
binds each declared composition input port to an upstream node. The top-level
input: is also required by the current node shape, but inputs: is the
authoritative port wiring. The config: block passes parameters that customize
this invocation.
Resolving the use: path
A use: value names a .comp.yaml in the workspace. It is resolved
relative to the directory of the pipeline file being compiled, then
against the set of .comp.yaml files discovered under the workspace root,
finally falling back to a filename match. A use: that resolves to no
.comp.yaml — a typo, a wrong relative prefix, or a file that does not
exist — fails compilation with a spanned E103 diagnostic naming the
composition node. The whole run aborts loudly; it does not silently drop
the composition and write an empty output. The same holds for the other
composition-binding errors (E102–E108): an ill-bound call site fails
compile rather than producing a run that writes zero records. Run
clinker explain --code E103 for details.
Composition definition file
A .comp.yaml file declares its interface in _compose: and its executable
subgraph in nodes::
# compositions/risk_score.comp.yaml
_compose:
name: risk_score
inputs:
inp:
schema:
- { name: order_id, type: string }
- { name: amount, type: float }
outputs:
out: scored
config_schema:
threshold:
type: float
default: 0.5
range: [0.0, 1.0]
nodes:
- type: transform
name: scored
input: inp
config:
cxl: |
emit order_id = order_id
emit amount = amount
emit high_value = amount >= $config.threshold * 2000.0
Composition fields
| Field | Required | Description |
|---|---|---|
_compose.name | Yes | Composition identifier |
_compose.inputs | Yes | Named input ports and their minimum required schemas |
_compose.outputs | Yes | Output port aliases pointing to body nodes or route ports |
_compose.config_schema | No | Typed configuration parameters, defaults, and constraints |
_compose.scoped_vars | No | Explicit scoped-variable names the sealed body may read from its caller |
_compose.resources_schema | No | Typed resource slots; see Resource bindings |
nodes | Yes | Unified node list for the sealed composition body |
Reading config parameters in the body
A composition body reads its own config parameters as $config.<param>. The planner constant-folds each reference to the value resolved for that instantiation — the call site’s config: value, or a channel/group config: override, or the declared default — so the same composition used with different config: compiles to different bodies. Because the resolution happens per instantiation, a channel or group config: override changes what the body computes, not just the reported provenance.
Explaining config provenance
Every resolved composition parameter retains its base value and each attempted group, channel-wide, and per-target override. The winning layer, shadowed values, fixed locks, and source spans remain attached to the stable compiled node identity. Inspect a value with either a unique shorthand or its exact versioned address:
clinker explain pipeline.yaml --field 'risk.threshold'
clinker explain pipeline.yaml \
--field '/v1/config/nodes/risk/fields/threshold'
The output includes the canonical exact address and lists layers in a stable
order. Exact addresses include every enclosing composition call. For example,
two sibling calls may each contain a local node named shared:
/v1/config/calls/left/nodes/shared/fields/threshold
/v1/config/calls/right/nodes/shared/fields/threshold
In that case shared.threshold fails with E118 instead of selecting one by
insertion order. The diagnostic lists both exact addresses in deterministic
order; copy the intended --field correction. An unknown query fails with
E117 and lists only same-field candidates, never unrelated nodes or fields. An
empty query fails with E116.
Address segments use RFC 6901 escaping: ~ becomes ~0 and / becomes ~1.
Unicode remains unchanged. This makes rendering and parsing lossless, including
after provenance serialization and repeated inspection of a compiled plan.
Source-schema provenance uses the parallel exact form
/v1/schema/sources/<source>/columns/<column>/attributes/<attribute>; the
three-part source.column.attribute shorthand remains available.
Body validation
Nodes inside a composition body are validated with the same node-scoped
config checks as top-level pipeline nodes. A body node that would be
rejected at the top level — an envelope wiring the not-yet-supported
trailer: port, a transform declaring a reserved variable name or a
default that does not match its declared type, an invalid log
directive, or a batch_size: 0 — fails compilation with an E115
diagnostic naming the composition call site, the body file, and the
violation. Run clinker explain --code E115 for details.
A body source or sink that sets a CSV delimiter or quote_char
which is not exactly one ASCII byte is likewise rejected at compile time,
not first at run, with the same one-byte rule top-level nodes get.
A body source or sink whose schema: names an external
.schema.yaml file has that path resolved relative to the composition
file’s own directory (not the invoking pipeline’s), and the file’s columns
are inlined before the body binds. A body Sink therefore rounds
decimal columns to their declared scale
at the write boundary exactly as a top-level Sink does.
A body sink cannot be combined with a Source that declares
dlq_granularity: document: document-level dead-lettering needs every Sink at
pipeline level, and the pipeline fails compilation with E378. Run
clinker explain --code E378 for how to move the Sink out of the body.
Executable example corpus
The five fragments under examples/pipelines/compositions/ are executable
examples of the current authoring surface:
- Clean Names
- Fiscal Date Fields
- Order Classification
- Shipping Cost
- Validate Email
Each fragment uses _compose.inputs, _compose.outputs, config_schema, and
the unified nodes list. The two date-dependent examples require an explicit
as_of_date configuration value so their results do not depend on the day the
test runs.
The composition example test recursively inventories every .comp.yaml file
in that directory. Its case manifest must name exactly that discovered set:
empty or missing inventories, duplicate keys, missing or extra cases, and paths
that escape the corpus directory fail with distinct diagnostics before any
example runs. Each case is then loaded through the production composition
loader, placed in a generated pipeline, and executed by the real clinker
binary. The test checks the exit status, record counters, and output bytes.
clinker run --explain proves that a generated pipeline compiles. It does not
prove that the example produces the documented result; the executable corpus’s
byte comparison is the behavior check.
Advanced wiring
For compositions with multiple input ports, bind every declared port by name:
- type: composition
name: enrich_address
input: orders
use: "./compositions/order_product_enrich.comp.yaml"
inputs:
orders: orders
products: product_catalog
The primary input: must be present for the current YAML node shape. The
planner builds composition edges from the named inputs: map, so that map must
contain every required port declared by _compose.inputs.
Downstream nodes consume a composition’s declared output ports using the usual
node.port syntax. If there is only one output port, the bare composition node
name selects it.
Resource bindings
A composition declares the external capabilities it needs as typed slots. The
currently admitted kind is file; it requires read capability and a run-local
file opener:
_compose:
name: order_lookup
inputs:
input: { schema: [{ name: order_id, type: string }] }
outputs: { out: order_reference }
config_schema: {}
resources_schema:
orders:
kind: file
required: true
nodes:
- type: source
name: order_reference
config:
name: order_reference
type: csv
resource: orders
schema: [{ name: order_id, type: string }]
resource: is an explicit body-Source-to-slot link. The slot must be declared
by the enclosing _compose.resources_schema; the Source name and its format do
not select a resource implicitly. A resource-backed body Source must not also
declare path, glob, regex, or paths, because the bound catalog resource
is its only external target. Every authored body Source must declare
resource:. Composition input ports are separate synthetic roots: to consume
caller-provided rows, declare _compose.inputs.<port> and set a downstream
node’s input: <port> instead of authoring a Source node for that port.
Top-level Sources are unchanged: they continue to require exactly one direct
matcher. resource: on a top-level Source is rejected until a separate
top-level binding surface is designed.
The workspace supplies a concrete, secret-free descriptor under a logical
identity in clinker.toml:
[catalog.resources.shared_orders]
kind = "file"
path = "data/orders.csv"
access = "read"
The call site binds only the declared slot to that logical identity:
- type: composition
name: lookup
input: orders
use: ./compositions/order_lookup.comp.yaml
inputs: { input: orders }
resources: { orders: shared_orders }
Resource descriptors are strict. Unknown kinds or fields, inline objects, unknown catalog identities, undeclared slots, missing required slots, and kind/capability mismatches fail planning. A call site cannot contain a path, credential profile, secret, token, or opened handle. File descriptors must remain inside the workspace and the catalog is admitted under fixed entry and descriptor-byte limits.
Planning retains the winning logical identity and every attempted overlay
layer for each binding. That identity also participates in the semantic plan
fingerprint. For each authored body Source, planning compiles a distinct
call-site-scoped instance carrying only the slot, logical identity, resource
kind, required capabilities, opener family, run lifetime, provenance, and
stable logical dataset identity. It does not retain the catalog’s physical
path. During clinker run, the CLI resolves that credential-free file
requirement at the workspace edge and transfers an opaque single-use reader
factory to the executor. The executor acquires the complete compiled group,
opens all of its Sources before starting any of them, and streams their finite
records through the ordinary bounded Source path. An open/read failure or
interruption closes every opened session and releases the group. This surface
still does not select credentials: a group requiring credentials fails before
runtime effects because no credential-profile option exists yet.
Ordinary composition calls do not have outputs: or alias: fields. Either
key fails with E377 at its authored location. Use _compose.outputs for the
public output contract and the composition node’s name for its
caller-visible namespace. add.alias remains valid only inside an overlay
add operation, where it names the inserted node.
Call-site fields
| Field | Required | Description |
|---|---|---|
input | Yes | Primary upstream required by the current node shape |
use | Yes | Path to the .comp.yaml definition |
inputs | Yes for declared ports | Map of composition input ports to upstream node references |
config | No | Parameter overrides (key-value pairs) |
resources | No | Declared slot to logical [catalog.resources] identity; scalar values only |
outputs | Rejected | E377: declare ports under _compose.outputs and use node.port downstream |
alias | Rejected | E377: use this composition node’s name as the namespace |
Contract status
The bounded catalog, typed file slot/binding, overlay provenance, stable logical dataset identity, and E377 call-surface rejection implement the planning half of D-12 through D-16. Credential references, credential resolution, and runtime handle activation are not implemented; a resource binding is therefore a validated planning contract, not permission to perform I/O.
See the canonical composition-resource and call-site contract for status, evidence, compatibility impact, and the AUTH-01 boundary.
Complete example
pipeline:
name: order_pipeline
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
schema:
- { name: order_id, type: string }
- { name: amount, type: float }
- type: composition
name: risk
input: orders
use: "./compositions/risk_score.comp.yaml"
inputs:
inp: orders
config:
threshold: 0.5
- type: sink
name: result
input: risk
config:
name: result
type: csv
path: "./output/scored_orders.csv"
Correlation Keys
A correlation key declares a set of records from a single source as an atomic group: if any record in the group fails validation or processing, the whole group is sent to the DLQ. This is the right shape for transactional data where partial processing is worse than total rejection – the canonical example is an order with multiple line items where one bad line should reject the entire order.
This page describes how to declare a correlation key and how it behaves through each node that can fan out, fan in, group, or join records.
Interactive companion: the correlation keys explainer lets you make lines fail and see what reaches the output and the DLQ, with and without an Aggregate.
Declaration
Correlation keys are declared per source. Each source’s config: block carries an optional correlation_key: field naming the column (or list of columns) whose value identifies a record’s correlation group within that source.
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: ./data/orders.csv
correlation_key: order_id
schema:
- { name: order_id, type: string }
- { name: amount, type: int }
- type: source
name: customers
config:
name: customers
type: csv
path: ./data/customers.csv
correlation_key: [customer_id, region] # multi-column key
schema:
- { name: customer_id, type: string }
- { name: region, type: string }
- { name: name, type: string }
- type: source
name: sensor_readings
config:
name: sensor_readings
type: csv
# No correlation_key: record-level errors land in the DLQ as
# standalone entries with no group atomicity.
schema:
- { name: ts, type: date_time }
- { name: value, type: float }
A record’s correlation group is identified by the tuple of values for that source’s listed fields. Records sharing the same tuple within the same source belong to the same group. There is no pipeline-level correlation key — declare it on each contributing source.
The group identity is captured at ingest, so rewriting the key column in a later Transform does not change a record’s group — anonymizing or transforming order_id downstream still keeps the original grouping intact.
A source whose declared correlation_key: field names a column not present in its own schema: block is rejected at compile time with diagnostic E153. The fix is to add the field to the schema or remove it from correlation_key:.
Row order
Declaring correlation_key on a Source changes the order its rows reach the rest of the pipeline. When the Source’s declared sort_order already begins with the correlation-key fields, in either direction, that declared order is kept as it is. Otherwise the planner inserts a sort after the Source, ascending on the correlation key and then on the remaining declared sort_order fields, so that each group’s rows are adjacent. Rows with equal keys keep their order in the file.
Every consumer that depends on row order sees this order: Sink output order, an Aggregate’s first-arriving value, Cull and Reshape ties, and which build record a Combine’s match: first picks. Adding a correlation key to an existing pipeline can therefore change those results even when no record fails.
DLQ semantics
When a record fails inside a correlation group:
- The failing record produces a trigger DLQ entry. Its category reflects the actual failure (e.g.
type_error,validation_failed). - Every other record from the same source in that group produces a collateral DLQ entry, carrying the category
correlated. - Records belonging to other (clean) groups proceed normally.
A record with a null value for the correlation-key field is treated as its own group: it has no peers, so DLQ atomicity does not span multiple records.
The dlq_count counter sums triggers and collaterals. It counts one trigger per failure, so a row that fails twice in a group counts twice, as it does without a key.
Group buffering
The engine buffers records per correlation group until either the group completes or a failure triggers a flush. The max_group_buffer: field on the pipeline-level error_handling: block caps per-group buffering across every source’s groups:
error_handling:
max_group_buffer: 100000 # Default: 100,000
A group that goes over the cap is dead-lettered whole when the run commits it. It is not a hard error: the run continues.
- Each row of the group that failed on its own is written as its own trigger, with its own category and
_cxl_dlq_id. - The group’s other rows are written under one
group_size_exceededtrigger, the first of them; the rest arecorrelatedrows that carry its id as their_cxl_dlq_trigger_id. - A group whose rows all failed writes only those failures, with no
group_size_exceededrow.
Each failure is written once, with its Combine build row if it has one (see Combine interaction), each condemned row once, and every row counts toward dlq_count and the DLQ rate limits. The group_size_exceeded row’s _cxl_dlq_timestamp is when the group went over the cap, and its error detail states the cap and how many entries the group held.
The cap counts the entries a group holds, not its distinct rows: a row counts once for each Sink it reaches, and each failure counts once. A row that an inclusive Route sends to two Sinks counts twice, and so does a row that fails on one branch and reaches a Sink on another. A failing Combine match counts once, although it holds both the driver row and the matched build row.
Going over the cap does not stop a group from buffering. The group keeps buffering its rows until the run commits it, so today the cap decides how a large group is dead-lettered but does not bound the memory it uses.
Per-operator interactions
Route interaction (fan-out)
A correlation group can span multiple route branches. Group atomicity is preserved across branches: if any record in the group fails (in any branch’s transform, or in the route predicate itself), the entire group is rejected from every branch.
For an inclusive route where one record reaches both branches, a single failure DLQ’s that source row exactly once — not once per branch. A record that fails on both branches has two failures and is written twice, once as each branch’s trigger with its own stage, as it is without a key.
Merge interaction (fan-in)
Merge concatenates upstream branches that share a schema. Records keep their correlation identity through the merge, so rows from different sources that share the same key value become one correlation group downstream: a failure on any one of them DLQ’s the whole group across both sources.
Per-source rollback narrowing
When two sources contribute records to the same correlation group, a failure originating from one source does not collaterally DLQ records from the other source. The collateral fan-out is scoped to the failing source’s records only.
For example, with [src_a, src_b] → merge → transform → out where both declare correlation_key: id, an error that fires on a src_b row produces a trigger for that row while the src_a row sharing the same id is spared and reaches the output. Single-source pipelines behave exactly as a pipeline-wide collateral DLQ would, since every co-grouped record shares the one source.
Two cases stay group-wide rather than narrowing per source:
max_group_bufferoverflow DLQ’s every record in the overflowing group — no single source is to blame for the overflow.- Combine output failures DLQ the synthesized output row, which has no single-source attribution. The exception concerns the output row only: the matched build record’s dead letter follows the failing driver’s group, as described under Combine interaction, and never widens the narrowing to the build record’s source.
Aggregate interaction
When an aggregate’s group_by covers every correlation-key field, the aggregate stays on the strict path: each emitted row inherits the correlation identity of its inputs, and any DLQ trigger in the group rolls back every record in the group, including the aggregate output row.
- type: aggregate
name: order_totals
input: orders # correlation_key: order_id
config:
group_by: [order_id] # covers the key
cxl: |
emit total = sum(amount)
When an aggregate’s group_by omits a correlation-key field, the engine automatically retracts only the failing records and recomputes the affected groups, so the surviving contributions still produce a correct aggregate row. You do not configure this — the engine picks the path from the group_by content. (One restriction: this mode cannot be combined with strategy: streaming, which is rejected at compile time.)
- type: aggregate
name: dept_totals
input: orders # correlation_key: order_id
config:
group_by: [department] # omits the key — surviving rows recomputed
cxl: |
emit total = sum(amount)
Combine interaction
Every combine declares propagate_ck: to select which correlation-key fields its output rows carry:
propagate_ck: driver— output inherits only the driver input’s correlation identity. The common case; today’s strict-correlation pipelines stay on this setting.propagate_ck: all— output carries the union of correlation-key fields across every input. Use when the build side carries keys that downstream operators need to read.propagate_ck: { named: [<field>, ...] }— output carries exactly the named subset. Use to project a multi-field key down after a join.
- type: combine
name: enriched
input:
o: orders # driver (correlation_key: employee_id)
d: departments # build side
config:
where: "o.employee_id == d.employee_id"
match: first
on_miss: skip
cxl: |
emit employee_id = o.employee_id
emit amount = o.amount
emit dept = d.dept
propagate_ck: driver
How match mode fills the propagated key:
match: first— the single matched build’s key fills the slot.match: all— one output row per matched build, each carrying its own build’s key.match: collect— one row per driver; the first matched build’s key fills the (single-valued) slot, while every matched build’s full payload still rides inside the array column.
Driver wins on a name collision: if both the driver and a build input declare the same key field, the output keeps the driver’s value.
propagate_ck is a required field — every combine must spell out which mode it uses.
A failing match’s build-side dead letter follows the driver’s group. When the combine body fails for a driver row, that driver row is the trigger of the driver’s correlation group. The matched build record’s dead letter is held with the same group as a collateral (_cxl_dlq_trigger: false, category combine_output_row), written right after its driver’s row and carrying that driver’s _cxl_dlq_trigger_id, and it is written or rolled back exactly when the driver’s group is. It never condemns the build record’s own correlation group: another driver that matched the same build record keeps its output unless its own group failed. The build record is written once per failure: when several drivers fail against one build record, in one group or in several, each failing driver’s row is followed by its own copy of the build row, carrying that driver’s _cxl_dlq_trigger_id. A driver that fails against several build rows is written once per failure, each copy followed by the build row of that failure.
Composition interaction
A composition’s body operates on records flowing in from the parent pipeline; correlation identity flows into the composition inputs and back out the named ports unchanged. Compositions cannot declare their own correlation key — a key is a property of a source, not of the composition body that consumes a source’s records.
Debugging
Correlation grouping is tracked on internal columns you never write in YAML or CXL, and they are hidden from writer output by default. To surface them for debugging, set include_correlation_keys: true on a Sink node:
- type: sink
name: debug
input: any_node
config:
type: csv
path: "./debug.csv"
include_correlation_keys: true
The output then contains extra columns named $ck.<field> (literal prefix in the CSV header) for each declared correlation-key field.
To investigate DLQ collaterals: every collateral entry’s category is correlated, and the trigger entry in the same group carries the actual failure category and message.
See also
- Error Handling & DLQ – general DLQ configuration, fail-fast vs continue, type-error thresholds.
- Aggregate Nodes – group-by semantics and the strategy hint.
- Combine Nodes – driver selection and match modes.
- Sink Nodes –
include_correlation_keysand other field-control flags.
Document Envelope Context ($doc.*)
Many enterprise file formats wrap their record body in an envelope:
named sections that surround the records and carry document-level
metadata — a batch header with a run date and batch id, a trailer with
a record count and checksum, or arbitrary sibling sections. Clinker
exposes these sections to CXL through the $doc.<section>.<field>
namespace.
nodes:
- type: source
name: payments
config:
name: payments
type: xml
path: data/payments.xml
options:
record_path: payments/Payment
envelope:
sections:
BatchInfo:
extract: { xml_path: "/payments/BatchInfo" }
fields:
batch_id: string
run_date: date
Summary:
extract: { xml_path: "/payments/Summary" }
fields:
record_count: int
checksum: string
schema:
- { name: amount, type: int }
Interactive companion: the document context explainer shows which records read each section, how each file becomes its own document, and what dlq_granularity: document rejects.
Like the rest of the pipeline config, the envelope: block is strict:
an unknown key at any level (a misspelled sections:, extract:, or
fields:) is rejected at plan parse time with a diagnostic naming the
bad key, rather than being silently ignored.
A downstream transform reads any declared section field on every body record:
nodes:
- type: transform
name: tag
input: payments
config:
cxl: |
emit batch = $doc.BatchInfo.batch_id
emit expected_total = $doc.Summary.record_count
emit amount = amount
Section names are yours
The engine reserves no section names. BatchInfo and Summary
above are arbitrary identifiers chosen by the pipeline author — Head
/ Foot, preamble / trailer, batch_metadata / eob_summary are
all equally valid. A section name is whatever string you put in the
sections: map; CXL exposes it verbatim as $doc.<that_name>.<field>.
Extracted sections are available throughout the body stream
Every extracted section is available to every body record. JSON and XML
can pre-scan declared sections anywhere in the document, so a header at the
top and a trailer at the bottom are both visible from the first record to
the last. SWIFT likewise scans its service blocks before emitting fields.
Multi-record CSV and fixed-width record_type extraction captures the
leading header region only; it does not extract arbitrary trailing sections.
For formats with trailing-section extraction, a trailer field is available during body processing, not just at end-of-file. A Transform can compare every row, including the first, against the trailer’s count:
- type: transform
name: check_count
input: payments
config:
cxl: |
emit amount = amount
emit declared_count = $doc.Summary.record_count
Note that an extracted trailer section you read via $doc.* is distinct
from the structural counts an EDI reader validates internally (the X12
SE/GE/IEA, EDIFACT UNT/UNZ, HL7 BTS/FTS segment counts). Those
trailer counts are checked by the reader against the body it streamed, and a
mismatch is a structural-integrity failure — see Malformed
envelopes for
how dlq_granularity: document dead-letters a malformed file instead of
aborting the run.
The pre-scan reads the envelope-bearing segments of the file before body streaming begins. The amount of retained data depends on the declared sections, their payloads, and the format’s path to a trailing section:
- JSON streams the pre-scan. The reader walks the document once and
deserializes only the subtrees the declared sections point at —
every other key (including a multi-megabyte body array) is
parsed-and-skipped without being stored. The retained sections live in
a bounded document index capped by
max_index_bytes(see below), so retained section memory scales with the declared sections, not the document size. - XML streams the pre-scan too. The reader event-walks the document
once and flattens only the declared section subtrees — every other
element, including a multi-megabyte body, is event-walked and dropped
without being materialized. The retained sections live in the same
bounded document index capped by
max_index_bytes. The reader still holds the file’s raw bytes for the lifetime of the read (one disk read backs both the body parser and the pre-scan), but the pre-scan no longer materializes the undeclared section subtrees.
For multi-record CSV, retained section names, field names, values and document
containers carry allocation ownership against the run’s finite memory budget.
Their charge lasts through downstream document use and any surviving aliases.
A repeated header does not create an unbounded history of section snapshots;
resource refusal remains fatal even under strategy: continue. CSV accepts the
same UTF-8 and Latin-1 policy for section cells as for body cells.
Only the envelope sections live in the document context. This separation does not mean the executor retains only one body record: downstream Source and Combine paths can still materialize whole inputs. See what the memory budget measures. Other formats’ existing document-index limits remain as described below.
Bounding envelope retention with max_index_bytes
For JSON and XML sources, the document index is capped so a pathologically large declared section fails loud rather than exhausting memory. The cap is charged incrementally as each section is parsed — byte by byte while the section’s subtree is built — so even a single oversized declared section aborts mid-parse, before its whole subtree materializes, naming the offending section and the cap. (Undeclared siblings, including a multi-megabyte body, are skipped without being parsed into the index at all.)
- type: source
name: events
config:
name: events
type: json
path: "./data/events.json"
options:
record_path: data.rows
max_index_bytes: 64MB # cap on retained envelope sections
The same max_index_bytes option applies to an XML source’s options:
block, capping its envelope pre-scan identically.
max_index_bytes accepts a decimal size string (64MB, 500KB) or a
bare byte count. It is optional; when omitted the reader applies a
documented finite default of 64MB. Only the declared sections a
program actually reads are retained, so envelope metadata sits far below
this ceiling in practice — the cap exists to convert an unbounded
mistake into a clear error.
Extract rules per format
Each section declares how the reader locates its payload:
| Format | extract: key | Value |
|---|---|---|
| XML | xml_path | Slash-path to the section element, e.g. /doc/Head |
| JSON | json_pointer | RFC 6901 pointer — empty (whole document) or leading /, e.g. /Head |
| EDIFACT | segment | A service-segment tag — only UNB |
| X12 | segment | A service-segment tag — only ISA (GS/ST surface as nested levels) |
| HL7 v2 | segment | A header-segment tag — only FHS (BHS/MSH surface as nested levels) |
| SWIFT MT | segment | A service block: "1", "2", "3", or "5" (or its default label) |
| Multi-record CSV / fixed-width | record_type | A header record-type tag, e.g. H |
xml_path and the source-level record_path option are both slash-paths over
XML but root differently: xml_path tolerates a leading / (/doc/Head is its
documented form), while record_path rejects one. They locate different things
and are deliberately not aligned — see
XML Format → record_path and xml_path root differently.
Declaring an xml_path section against a JSON source (or vice versa),
a segment extract against XML/JSON, a record_type extract against
any format other than multi-record CSV / fixed-width, or any envelope:
block at all on a plain (single-schema) CSV or fixed-width source is a
configuration error that fails fast rather than silently producing empty
sections.
A json_pointer must be a valid RFC 6901 pointer: either empty ("",
naming the whole document) or a /-introduced path such as /Head or
/batch/summary. A slashless value like Head is a typo — it would
decode to zero segments and silently match the root document — so it is
rejected at validation rather than resolving to the wrong metadata.
A plain (single-schema) CSV or fixed-width source carries no envelope
— there is no header/trailer structure to extract. Declaring envelope:
sections on one is a configuration error (E356): the sections would
never be populated and every $doc.<section>.<field> against the source
would resolve to null, so the compiler rejects it rather than
accepting an inert declaration. Envelope extraction on a flat file — the
record_type extract — applies only to a
multi-record
source, one declaring a discriminator: + records: block; declare that
schema if the file genuinely carries header/trailer records.
Network (REST) sources carry no $doc context
A rest source pulls its records page by page over paginated HTTP — it
has no single buffered document with head and tail sections, so it
carries no envelope context. Envelope sections are a file-document
concept, so the compiler rejects them on a REST source rather than
letting them silently resolve to null:
- Declaring an
envelope:block on arestsource is an error (E349) — the declaration would be inert. - Reading
$doc.<section>.<field>from a node fed by arestsource is an error (E349) — the access can never resolve.
Pull the document-level metadata into record fields through the API’s own
response shape (record_path, split_to_rows) instead, so the value travels
as a normal field rather than as document-envelope context.
EDIFACT segment extract
An EDIFACT source exposes its interchange header UNB as an envelope
section. The section’s field names are the positional element keys
e01, e02, … :
envelope:
sections:
interchange:
extract: { segment: "UNB" }
fields:
e05: string # interchange control reference
Only the UNB header is extractable. Trailer segments (UNT, UNZ)
that arrive after the body are not envelope sections — their control
counts are validated by the reader instead. A mismatch between a trailer’s
declared count and the body the reader streamed is a structural-integrity
failure: by default it aborts the run, and under a source’s
dlq_granularity: document opt-in it dead-letters the whole file to the DLQ
(see Malformed envelopes).
A segment extract naming any tag other than UNB is rejected at startup. See
EDIFACT Format for the full reference.
Multi-record record_type extract
A multi-record
CSV or fixed-width source — one that declares a discriminator: and a
records: list — exposes a header record type as an envelope section
through the record_type extract. The tag names which of the source’s
declared record types carries the section’s payload; the matched header
row’s named fields become the section’s fields.
schema:
discriminator: { start: 0, width: 1 }
records:
- { id: header, tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
- { id: detail, tag: D, columns: [ { name: amount, type: int, start: 1, width: 9 } ] }
- { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 9 } ] }
structure:
- { record: trailer, count: count }
envelope:
sections:
head:
extract: { record_type: H } # the H header row → $doc.head.*
fields:
batch_id: string
Only a header record type — one whose rows precede the body at the
file head — is extractable as a $doc section; the reader captures the
first such row in a bounded pre-scan and excludes it from the body
stream. A trailer record type (one named by a structure:
constraint) arrives after the body it closes, so it is not an
envelope section — its declared count is validated against the streamed
body count instead, the same structural-integrity check the EDI trailers
use. See CSV
and Fixed-Width
for the full reference.
A JSON example:
- type: source
name: payments
config:
name: payments
type: json
path: data/payments.json
options:
record_path: records
envelope:
sections:
Head:
extract: { json_pointer: "/Head" }
fields:
batch_id: string
Foot:
extract: { json_pointer: "/Foot" }
fields:
count: int
schema:
- { name: amount, type: int }
against:
{
"Head": { "batch_id": "RUN-001" },
"records": [ { "amount": 10 }, { "amount": 20 } ],
"Foot": { "count": 2 }
}
Typed fields
Each section’s fields: map declares the field name and its type, drawn
from the same small vocabulary as source schemas: string, int,
float, bool, date, date_time. The extracted raw value is coerced
to the declared type at pre-scan time; a value that cannot coerce
(e.g. a non-numeric string declared int) fails the source with a
diagnostic naming the section, field, and offending value.
A field that the document does not carry resolves to null — $doc.*
follows the same missing-value convention as $source.* and
$pipeline.*. A section that the document does not carry at all is
simply absent from the context; any $doc.<missing_section>.<field>
resolves to null.
Declared-path validation
Every $doc.<section>.<field> reference is cross-checked at compile time
against the schema the feeding source’s reader will actually serve. A
reference that can never resolve — almost always a typo — is rejected at
compile time, pointing at the node that made it, rather than resolving
silently to null. How the path is checked depends on the source:
Closed-schema sources (XML, JSON) — the envelope: block is the
complete schema: the reader extracts exactly the sections and fields it
declares. A reference naming a section the source does not declare, or a
field the declared section does not declare, is rejected with error E341
($doc.Summry.total against a declared Summary). Run
clinker explain --code E341 for the full write-up.
Multi-record CSV / fixed-width sources — a source whose schema declares
discriminator: + records: exposes a header record type as a $doc
section through the record_type extract. The reader coerces the matched
header record’s columns through the section’s declared fields: and serves
exactly those fields, so the section is closed just like an XML/JSON one: a
reference naming an undeclared section, or a field the section does not
declare, is rejected with error E341. (A plain single-schema CSV /
fixed-width source has no such structure — declaring an envelope: on one
is rejected with error E356.)
Segment/positional sources (X12, EDIFACT, HL7) — the file-level header
(ISA/UNB/FHS) is declared through envelope: and is closed, but the
reader also synthesizes nested envelope levels the config never names —
X12’s functional_group / transaction_set, HL7’s batch /
transaction_set — keyed by positional eNN / fNN elements bounded by
the source’s max_elements / max_fields. A $doc path is checked
against that synthesized vocabulary plus any section/field you declared, so
a legitimate wire-derived path ($doc.functional_group.e06) is accepted
while a misspelled section ($doc.functonal_group.e06) or an out-of-range
positional element ($doc.transaction_set.e99) is rejected with error
E348. Run clinker explain --code E348 for the full write-up.
REST sources carry no document and reject $doc outright — see
Network (REST) sources carry no $doc context
above (E349).
Plain (single-schema) CSV / fixed-width and SWIFT MT sources are not
statically checked. A plain flat file synthesizes no $doc sections, and
SWIFT serves declared sections under user-chosen or default block labels —
neither fits the closed or positional model cleanly. (A multi-record CSV /
fixed-width source is checked — see above.)
Indexed $doc access
A section field that holds an array or a map can be indexed inline, the same way any record value is — see Nested paths for the full bracket-index reference. Integer indices select array elements; string keys select map entries; the two compose into a chain:
cxl: |
emit first_line = $doc.Header.line_items[0] # array element
emit run_date = $doc.Header.meta["run_date"] # map entry
emit first_sku = $doc.Header.line_items[0]["sku"] # array-of-maps chain
An out-of-range array index or a missing map key resolves to null — it
never errors or panics, and a null mid-chain short-circuits the rest of
the chain to null rather than failing. This is the same missing-value
convention $doc.<missing_field> and $source.* follow.
Only literal paths are readable
Every index segment must be a literal — a constant integer or string written in the program text. The section and field are always literal identifiers (the grammar requires it), so a literal index is the last piece a reader needs to know, before reading any input, exactly which envelope paths a run will consume. The pre-scan extracts precisely those statically-resolvable paths and nothing else.
A computed index — one derived from runtime data, such as
$doc.Header.line_items[row_index] — is not statically resolvable: the
reader cannot pre-scan a row-dependent element. clinker rejects it at
compile time with a diagnostic pointing at the offending index, rather
than reading it at run time. Use a literal index, or pull the value into
a record field upstream and index that instead.
This compiles together with the declared-path rule above: a $doc read
of a section or field the source does not declare is a compile-time
error (E341), and a $doc read with a computed index is a
compile-time diagnostic — neither reaches run time as a silent
null. Only a literal path over a declared section is pre-scanned and
readable.
One document per file
Each source file is its own document with its own envelope context.
When a source matches multiple files (via glob: / paths:), each file
gets a fresh document context with its own section values. Records from
different files never share a context — a record’s $doc.* always
reflects the file that record came from.
Document boundaries flow through the pipeline so that document-scoped operators fire at exactly the right point. A document-scoped operator fires exactly once per document, even when that document arrives across several inputs that a Merge or Combine brings together.
Per-document aggregation
A grouped or global Aggregate reading a multi-document source produces
one set of grouped rows per document, not a single aggregate spanning
every document. When a document closes, the Aggregate finalizes and emits
the groups belonging to that document, then drops their state before the
next document accumulates — so a glob: source over twelve monthly files
through a group_by Aggregate yields twelve independent monthly
roll-ups, and only one document’s groups are ever live at once (the
others have already been emitted and freed, or have not yet started).
This applies only when a document boundary actually reaches the
Aggregate. A plain single-file source is one document, so it still emits
one aggregate. A Merge that combines several distinct single-document
sources flushes those sources independently downstream — one roll-up
per source document, exactly as feeding each source to its own Aggregate
would. This holds for every Merge mode.
A Combine (join) preserves document boundaries on every strategy, so a
per-document Aggregate downstream of a join also rolls up per driver
document.
Nested (multi-level) envelopes
Some formats wrap their records in several envelope levels, one inside another. EDI X12 is the canonical example and the first format that implements this: an interchange (ISA/IEA) contains one or more functional groups (GS/GE), each containing one or more transaction sets (ST/SE), each containing the records. A single file can carry multiple interchanges back to back. See X12 Format for the full reference.
HL7 v2 is the second multi-level format: an optional file (FHS/FTS)
contains optional batches (BHS/BTS), each containing one or more
messages (MSH..), each containing the segment records. The tiers map
onto the same nested levels — the FHS file header is a declared
segment: "FHS" section, while the BHS batch and the MSH message
surface automatically as the reader-supplied sections batch and
transaction_set. Every tier is optional, and a level’s section exists
only when its header segment is present in the input — the reader
synthesizes a batch section only where it reads a BHS, and the message
level only where it reads an MSH.
A bare MSH-led stream — messages with no BHS batch and no FHS file
wrapper — therefore opens only the message level. $doc.transaction_set.*
resolves against each message, but $doc.batch.* resolves to null:
no BHS was read, so no batch section was synthesized. Whether a file
carries a BHS batch wrapper is a property of the input bytes, not the
config, so a $doc.batch.* path is never a compile error — it follows the
same missing-value convention an absent $doc section follows (resolve
null, never error), and populates as soon as a BHS wrapper is present:
# Bare MSH stream (no BHS/FHS) — only the message level opens:
- msg_type: $doc.transaction_set.f08 # MSH-9 message type — populated
- batch_id: $doc.batch.f01 # null — no BHS, so no batch tier
# BHS-wrapped stream — the BHS opens a batch level, so both populate:
- msg_type: $doc.transaction_set.f08 # MSH-9 message type — populated
- batch_id: $doc.batch.f01 # BHS field — now populated
See HL7 v2 Format for the full reference.
A reader for such a format opens and closes each nested level as it
crosses the corresponding envelope boundary mid-file. Each level
contributes its own sections to $doc. There is no new $doc syntax for
nesting — every level’s sections are read through the same two-level
$doc.<section>.<field> lookup. A record inside the innermost level sees
every enclosing level’s sections at once. For X12 the interchange header is
a declared segment: "ISA" envelope section (you choose its name), while
the GS group and ST set surface automatically as the reader-supplied
sections functional_group and transaction_set, each keyed by positional
eNN elements:
cxl: |
emit interchange_control = $doc.interchange.e13 # ISA13, declared section
emit functional_id = $doc.functional_group.e01 # GS01 (reader-supplied)
emit transaction_type = $doc.transaction_set.e01 # ST01 (reader-supplied)
emit claim_amount = amount # body field
A record streamed inside the ST level resolves the ST section, the enclosing GS section, and the outermost ISA section, all at once: each inner level inherits every enclosing level’s sections as siblings in one flat namespace. If two levels declare a section with the same name, the innermost wins for records inside it — the same shadowing rule a nested scope follows in any language. Picking distinct per-level names (as above) keeps every level independently visible.
The reader-supplied default names are not the only option: an X12 source
can name the GS and ST levels itself and give each a typed field
schema, so a nested level is addressable under a chosen name with coerced
fields exactly like the declared ISA section. The declaration lives on
the source’s options (not the envelope: block, which is reserved for
the pre-scannable file-level header), and each nested level is named
independently:
type: x12
options:
group_section:
name: functional_group # your choice — the engine reserves no name
fields: { e06: int } # GS06 group control number, typed
set_section:
name: transaction_set # your choice
fields: { e01: int } # ST01 transaction-set id, typed
Omit a level’s declaration and it falls back to its reader-supplied
default name keyed by untyped positional eNN strings. See
X12 Format for the full
reference.
Boundaries nest correctly through the pipeline: each level opens before the records inside it and closes after them, in strict innermost-first order. A level that arrives across several branches is still handled once where a Merge or Combine brings those branches together — exactly like a single-level document.
Header-only interchanges
A multi-level envelope file can legitimately carry an interchange whose
body is empty — envelope structure (an interchange header, and possibly
inner group headers) with zero records inside. Such an interchange still
opens a document and emits its open/close boundaries, so downstream
operators and trailer-count validation observe it just like any other
document. The interchange’s $doc.* sections are extracted and the
boundaries flow even though no body record ever streams from it.
The same holds for an empty inner envelope — an open/close pair with no records between — and for an inner envelope that opens or closes after the file’s last body record. Every envelope boundary a reader signals is applied, whether or not a record follows it, so the document frame stays balanced end to end.
Fixed-width document output
A fixed-width Sink can echo any extracted section as its header or footer.
The names select document context; they do not declare new input extraction
rules. For example, these Sink config keys select sections named manifest
and totals:
reconstruct_envelope: true
options:
envelope:
header_from_doc: manifest
footer_from_doc: totals
Both names are author-defined. With a multi-record flat-file source, both
sections must have been captured in the leading header region, even though
totals is rendered at the end of the output document. A structural trailer
count check does not make that trailing input record an extracted section.
Header and footer fields concatenate in stored order without the body’s
fixed-width padding. Strings remain verbatim, numbers and booleans use scalar
spellings, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null adds
no text. Arrays and maps are unsupported: a structured field anywhere in the
section rejects that entire header or footer before delivery. A missing
section emits nothing; a present empty section emits only the configured
LF/CRLF separator, or no bytes with line_separator: none.
Document start, each record, and document end are separate prepared operations. Only successful delivery advances document state or the body count; opening the next document starts its own count and selects its own sections. Earlier successful operations remain delivered if a later footer fails. A computed footer record-count field is unsupported for fixed-width (E346). See fixed-width output and output preparation.
Native JSON and XML output boundaries
JSON/XML reconstructed envelopes prepare document start, each body record, document end and finalization as separate complete operations. A document’s body count advances only when a record is delivered. Ending one document and opening the next resets that count and restores the next document’s own sections. Section names remain the names declared in your pipeline.
Reconstructed JSON envelopes retain their document grammar: array mode wraps
the envelope documents in an array; NDJSON mode separates complete envelope
documents with LF. Unlike ordinary compact NDJSON records, reconstructed
pretty: true documents can span lines. XML places each Document wrapper inside one configured root, retaining
header/footer section wrappers and emitting no XML declaration.
Explicitly opened empty documents still carry framing and a zero body count. In the CLI, a native source with no body records does not open a writer, so its output is an empty file even with reconstruction enabled. Do not infer this behavior from header-only message interchanges, whose readers emit explicit boundaries even without a body.
Ordinary native output supports correlation, splitting and per-file fan-out. Reconstructed envelopes combined with splitting, per-file fan-out, correlation or document-grain DLQ are rejected before execution (E347). These combinations cannot establish the required single document boundary.
Library writers distinguish draining already delivered bytes from finalizing a document. A byte drain adds no closing syntax. Once delivery fails, neither a drain, finalization nor teardown resumes the failed operation. See output preparation for failure effects.
Error Handling & DLQ
Clinker provides structured error handling with a dead-letter queue (DLQ) for records that fail processing. The error_handling: block at the top level of the pipeline YAML controls the behavior.
Configuration
error_handling:
strategy: continue
dlq:
path: "./output/errors.csv"
include_reason: true
include_source_row: true
Strategies
error_handling.strategy is pipeline-wide – it is set once at the top level, not per node. It controls what happens when a record fails:
| Strategy | Behavior | Exit code |
|---|---|---|
fail_fast | Default. Abort the run on the first record failure. | Non-zero, by the class of the aborting error (3 for an evaluation failure, 4 for an I/O failure – see Exit Codes) |
continue | Route the failing record to the DLQ and keep processing. | 2 if any record was dead-lettered, 0 otherwise |
There are exactly two, because the engine makes exactly one decision at each record failure: propagate it and stop, or dead-letter it and carry on.
fail_fast
The safest strategy. Any record-level error (type coercion failure, validation error, missing required field) halts the pipeline immediately, with a non-zero exit and no DLQ file. Use this when data quality is critical and you prefer to fix issues before reprocessing.
Some failures abort the run under either strategy, because they are not record-scoped: an unwritable output path, a config or CXL compile error, and the DLQ-rate ceiling (dlq.max_rate, E315/E316) all end the run regardless of the strategy.
CSV, JSON and XML resource failures are also fatal under either strategy. Memory or disk
admission refusal, allocation failure, descriptor exhaustion and temporary-storage
failure are not bad-record errors, so continue cannot turn them into successful
output. A typed resource diagnostic preserves the kind of failure rather than
reporting every case as a memory shortage. A failed destination can already have
accepted a prefix; see output preparation.
Malformed JSON/XML input encoding is a data failure, including when discovered
during schema discovery or envelope pre-scan. Under fail_fast, the CLI returns
exit 4 and machine code source.data.invalid. It does not report a compilation
error merely because no record has reached the pipeline. A late error can leave
an already delivered prefix; an envelope pre-scan may discover it before any
body records. Failed runs do not publish their staged normal output files.
Explicit cancellation ends an interrupted run with exit 130; it does not add a
Sink error. If a real I/O or resource failure occurs alongside a shutdown request,
the real failure retains its classification. Record and byte counters describe
established progress, not rows merely attempted or prepared.
An executor invariant failure also aborts under either strategy with exit code
1. In particular, if a planned materialized input is unavailable when its
consumer runs, Clinker stops instead of treating that input as a legitimate
zero-row result. The message names the consuming node and planned producer
(including the producer port when applicable) and says the input was not
treated as empty. A source or stage that really emits zero rows remains valid;
it carries an explicit empty buffer and completes normally. Report any missing-
input internal error as an engine defect rather than routing it to the DLQ.
continue
The production workhorse. Bad records are written to the DLQ file with diagnostic metadata, and the pipeline continues processing remaining records. After the run completes, inspect the DLQ to understand and correct failures.
A pipeline that completes with DLQ entries exits with code 2 – this signals “pipeline completed successfully but some records were rejected.” It is not a crash or internal error. A continue run that dead-letters nothing exits 0, exactly like a clean fail_fast run.
Migrating from
best_effort. The removedbest_effortspelling was a third name for thecontinuebehavior: it wrote the same DLQ entries and produced the same exit code, because the runtime never distinguished the two. Replace it withstrategy: continue. A pipeline still carryingbest_effortis rejected at config-validation time with a message naming the replacement.
Declared source-type failures are deliberately not lossy: under continue,
the complete original record is written to the
configured DLQ and no null, raw, or partially converted replacement enters the
pipeline. This strategy therefore requires an error_handling.dlq block when
such a failure occurs. fail_fast stops on the first failure without emitting a
replacement. This includes fields renamed by a source schema: rejection retains
the original decoded record and its values, even when conversion failed after
other fields had already been examined.
An evaluation error is never false
A condition that fails to evaluate, such as one that divides by zero, has not said whether it holds. The engine never reads that failure as “false”, on any node: the record the condition was about is dead-lettered, and no decision that depends on the condition being false is taken for it.
- A Transform
filterthat fails dead-letters the record; it is neither kept nor filtered out. - A Route branch condition that fails dead-letters the record, which takes no
branch and not the
default. Inexclusivemode only the conditions up to the first true one are evaluated, so a later condition cannot fail. - A Combine
where:that fails for a candidate build row dead-letters that pair. The driver is not unmatched, soon_missdoes not fire; undermatch: firsta failure on the deciding candidate is the driver’s only result, and undermatch: collectthe driver writes no row. See Combine.
A pipeline that wants a failing condition treated as false says so in CXL, for
example by guarding the division or coalescing the result with ?? false.
DLQ configuration
The DLQ is always written as CSV, regardless of the pipeline’s input/output formats.
dlq:
path: "./output/errors.csv"
include_reason: true
include_source_row: true
| Field | Required | Default | Description |
|---|---|---|---|
path | No | – | The pipeline-wide DLQ file. It receives the dead letters of every Source without its own per_source path. A dead letter with neither this path nor a per_source path for its Source is counted in the run’s dead-letter totals, sets exit code 2 and counts toward max_rate, but is written nowhere. The dlq: block itself is required to continue past a declared source-type failure. |
include_reason | No | true | Include _cxl_dlq_error_category and _cxl_dlq_error_detail columns. |
include_source_row | No | true | Include the failing record’s columns after the _cxl_dlq_* columns. Which record columns each DLQ file carries is fixed when the pipeline compiles; see How the DLQ columns are chosen. With false, only the _cxl_dlq_* columns are written. |
max_rate | No | none | Stop the run (E315, exit code 3) once the dead-lettered rows reach this fraction of the source rows read so far, both counted across the whole run. Must be greater than 0.0 and at most 1.0 (E318). Without it, the run is never stopped for its dead-letter rate. See Bounding how much can dead-letter. |
min_records | No | 100 | How many source rows must have been read before max_rate is checked, so the first failures of a run cannot trip it on a tiny denominator. Also the default for each per_source min_records. |
per_source | No | – | Settings for individual Sources, keyed by Source node name: a separate DLQ file, and a rate ceiling of their own. See Per-source DLQ settings. |
Per-source DLQ settings
per_source gives a Source its own DLQ file, its own rate ceiling, or both.
Each key is the name of a Source node:
error_handling:
strategy: continue
dlq:
path: ./output/errors.csv
max_rate: 0.05
per_source:
vendor_feed:
path: ./output/vendor_feed_errors.csv
max_rate: 0.20
min_records: 500
orders:
max_rate: 0.01
| Field | Default | Description |
|---|---|---|
path | – | A separate DLQ file for this Source’s dead letters. They are written only there and do not appear in the pipeline-wide file. Without it, the Source’s dead letters go to the pipeline-wide path. |
max_rate | none | Stop the run (E316, exit code 3) once this Source’s dead-lettered rows reach this fraction of the rows read from this Source so far. Must be greater than 0.0 and at most 1.0 (E318). |
min_records | the pipeline-wide min_records, else 100 | How many rows must have been read from this Source before its max_rate is checked. |
A Source’s own max_rate is checked first, so a breach names that Source.
The pipeline-wide max_rate, when set, still applies to the run as a whole.
In the example, vendor_feed may dead-letter up to 20% of its own rows, but
the run still stops when all dead letters together reach 5% of all rows read.
A key that does not name a declared Source is rejected at compile time (E317). Two DLQ paths that name one file are rejected too (E318), including paths that differ only in case on a case-insensitive filesystem, or in being written relatively and absolutely. A DLQ path that names the same file as a Sink’s path is rejected with E322.
Which record columns each file carries follows from the Sources routed to it; see How the DLQ columns are chosen.
How DLQ output is written
Dead-letter rows are written while the run executes, not collected until it
ends. Each row is formatted under its file’s header, which is fixed when the
pipeline compiles (see
How the DLQ columns are chosen), and written into
a staged copy of that file in the run’s publication attempt: in quarantine
next to the destination by default, or under local_spool_dir with
mode = "local_then_publish" (see
Output publication).
Each open DLQ file writes through one fixed 64 KiB buffer, so the memory the
DLQ files use does not grow with the number of failures.
-
A DLQ file is created when its first row arrives. A file no row reaches is not created, and no empty file is published.
-
DLQ files are published only if the run succeeds, by the same publication step as the pipeline’s other outputs. A failed or interrupted run publishes no DLQ file.
-
The failures that are counted but have no destination (see
pathabove) are never formatted or written. -
Three kinds of dead letter are held in memory until the stage that found them finishes, and written then:
- a
join_valuescollision at a Sink that writes on its own thread; - an Aggregate
add_recordfailure found while the Aggregate reads its input on its own thread; - a Combine output-row failure found while the Combine streams its driver on its own thread, or inside a grace-hash, sort-merge or IEJoin join.
A Sink that writes on its own thread stops the run with an internal error once 65,536 collisions are waiting this way.
- a
-
Under a correlation key or
dlq_granularity: document, records are held until their group or document is decided. That is those features’ own state, described in their sections, not DLQ output; the rows they dead-letter are then written like any other. Underdlq_granularity: documentthe engine also keeps, for each rejected document, a compressed record of which rows it has already written, so a row held by several Sinks is written once. That record is charged to the memory budget. While a Sink writes a rejected document’s rows, the record grows by:- up to 16 bytes per row when the Sink receives the rows in the order their Source read them;
- up to 96 bytes per row when it receives them in any other order, for example after a Sort on a data column;
- up to 672 bytes per row for a document whose rows span its Source’s 4,294,967,296th row, in any order. A Source counts its rows across every file it reads.
The first row of each document a Sink starts costs up to 920 bytes. The growth is released when the Sink’s pass ends, or every 65,536 rows of a document. Between releases one document’s record therefore grows by at most about 1 MiB when its rows arrive in the order they were read, about 6 MiB in any other order, and about 42 MiB when its rows span the 4,294,967,296th row. For example, 65,535 rows that alternate between a document’s first 65,536 rows and its next 65,536 grow the record to about 5.8 MB, and the release at the next row brings it back to 880 bytes. Once released, the record costs at most 4 bytes per written row, plus 96 bytes for every 65,536 rows of the document, written or not, plus under 1 KiB. That is about 2.3 KiB for a million-row document whose rows are contiguous. When only one row in each 65,536 is written, it is 88 bytes per written row plus 672 bytes, within the cap of about 100 bytes per written row. The record cannot spill: if one more row would not fit once every held row (below) has moved to disk, the run fails with E310. A failed document’s failing records are formatted as DLQ rows when they fail and held until the document is rejected. They are held in memory, charged to the memory budget, and move to one file in the spill directory only when the budget needs the memory, counting toward
storage.spill.disk_cap_bytes(E320). With memory to spare nothing is written to disk. If even one more held row would not fit with every held row on disk, the run fails with E310.
Disk bounds how much DLQ output a run can produce: the free space at the
staging location, and the publication attempt’s byte ceiling
(storage.publication.max_attempt_bytes, and no more than
retained_byte_limit), which every staged file of the run counts toward,
DLQ files included. An attempt larger than that ceiling is refused at
publication and nothing is published.
When the staging location fills, the run fails with an I/O error (exit code 4) and publishes nothing. After the destination’s own error text, the message names the DLQ file, the number of rows dead-lettered so far, and the stage and category with the most rows, and suggests a breaker:
<destination error>: dead-letter output ./output/errors.csv could not be written after 1048576 dead-lettered rows (most from stage transform:validate_orders, category validation_failure)
help: stop the run before dead letters fill the destination, for example:
error_handling:
type_error_threshold: 0.05
dlq:
max_rate: 0.05
These breakers bound how many rows can dead-letter; disk at the staging destination bounds their volume.
Bounding how much can dead-letter
Disk bounds the volume of dead-letter output; the breakers bound how many rows can dead-letter in the first place. Set one wherever a wrong schema could make most rows fail, for example when a feed can change its columns or types without notice:
error_handling:
strategy: continue
type_error_threshold: 0.05
dlq:
path: ./output/errors.csv
max_rate: 0.05
dlq.max_rate(E315) anddlq.per_source.<name>.max_rate(E316) stop the run when the fraction of dead-lettered rows crosses the ceiling, oncemin_recordsrows have been read. The numerator counts dead-letter rows, collateral rows included, whether or not they have a DLQ file to go to. A source row counts once for each failure it took part in, so a row that failed twice counts twice, with or without a correlation key. Underdlq_granularity: document, a row that failed twice at the Source, in a Transform or in a Route is written once and counts once (see Document-level DLQ).type_error_threshold(E368) stops the run when the fraction of declared source-type failures crosses the threshold. It catches a schema mismatch at the Source, before the failing rows reach later stages.
A rate ceiling is checked each time a dead letter is counted, so the row that crosses it is itself counted and written before the run stops. A stopped run exits with code 3 and publishes nothing. No breaker is set by default.
DLQ columns
Every DLQ record includes these metadata columns:
| Column | Description |
|---|---|
_cxl_dlq_id | UUID v7 (time-ordered unique identifier), unique to the row. It is taken together with _cxl_dlq_timestamp, so ids order the same way as timestamps. |
_cxl_dlq_trigger_id | The _cxl_dlq_id of the trigger row whose failure produced this row. That trigger row is always written in the same run. A trigger row points to itself, so _cxl_dlq_trigger is true exactly when this value equals _cxl_dlq_id; every row one failure produced carries the same value. See Pairing the rows one failure produced. |
_cxl_dlq_timestamp | RFC 3339 timestamp of when the failure was observed, not of when the row was written. A collateral row (correlated, document_rejected) carries the time its correlation group or document was condemned, except that a rejected document’s other failing records carry the time each one failed. A group_size_exceeded row carries the time its group went over max_group_buffer. |
_cxl_dlq_source_file | Input filename carried by that failing record’s $source.file provenance (or <merged> when no source-file provenance exists) |
_cxl_dlq_source_name | Name of the Source the failing record came from (or <merged> when the record carries no Source identity) |
_cxl_dlq_source_row | 1-based row number in the source file |
_cxl_dlq_triggering_field | The field whose evaluation failed, when the failure names one; empty for collateral rejections |
_cxl_dlq_triggering_value | The value the failure reported, when it carries one (for example the text that failed to convert) |
_cxl_dlq_stage | Name of the transform or aggregate node where the error occurred |
_cxl_dlq_route | Route branch name (if the error occurred after routing) |
_cxl_dlq_trigger | true when the row’s own failure dead-lettered it; false when another row’s failure took it along (a correlated or document_rejected row, or a Combine build row) |
_cxl_dlq_source_record | One of the record columns rather than a metadata column: present in any file a Source rejection can reach under strategy: continue, and filled only for a record-grained E345 rejection. Contains the fixed-width line text or a JSON array of decoded CSV cells, preserving the physical row without assigning it a declared record shape. |
Timestamps need not increase down a file. A failure can be held before its
row is written, for example in a correlation group that commits later, or by
a stage listed under How DLQ output is written,
so its row can follow rows observed after it. Sort on _cxl_dlq_timestamp or
_cxl_dlq_id to read failures in the order they were observed.
When include_reason: true is set, two additional columns appear:
| Column | Description |
|---|---|
_cxl_dlq_error_category | Machine-readable error classification |
_cxl_dlq_error_detail | Human-readable error description |
Pairing the rows one failure produced
Two rules decide which rows a DLQ file holds:
- One row per failure. Every failure writes its own trigger row, with its
own category, detail, triggering field and value, stage and route. A row
that fails twice, on two Route branches or against two Combine build rows,
is written twice. A Combine failure also writes the build row that
contributed to it, right after its trigger. Under
dlq_granularity: documentthis rule has an exception: a rejected document’s later failing record is written as adocument_rejectedrow paired with the document’s first failure, and a record that failed twice at the Source, in a Transform or in a Route is written once (#1316; see Document-level DLQ). - Each condemned row once. A correlation group or document that fails
adds each of its other rows once, as a collateral of its first failure.
Under
dlq_granularity: document, a record that an Aggregate, a Combine or a Reshape failure dead-lettered is written again, as adocument_rejectedrow, only when a Sink on another branch also received it and its document is rejected. If its document is not rejected, that Sink can publish it (see Not covered).
A correlation key never removes, merges or relabels a failure row: the same failures are written with and without a key, and the key only adds the rows a failing group condemns and decides when rows are written.
One failure can dead-letter several rows:
- a Combine body that fails writes the driver row and its matched build row,
each under its own Source’s
_cxl_dlq_source_nameand_cxl_dlq_source_row, whichever join strategy ran. The build row is written once per failure, so a build record that two failing drivers matched is written twice, once after each driver’s row, and a driver that fails against three build rows is written three times, each copy followed by its own build row; - a failing row in a correlation group takes the rest of its group with it as
correlatedrows; - a group larger than
max_group_bufferwrites agroup_size_exceededrow and the rest of its group ascorrelatedrows (the group’s own failures keep their own trigger rows, see below); - under
dlq_granularity: document, a failing record rejects the rest of its document asdocument_rejectedrows.
Every row carries _cxl_dlq_trigger_id, the _cxl_dlq_id of the trigger row
whose failure produced it, and that trigger row is always in the run’s
output. A trigger row points to itself, so _cxl_dlq_trigger
is true exactly when _cxl_dlq_trigger_id equals _cxl_dlq_id. Group by
_cxl_dlq_trigger_id to see everything one failure took with it.
In this excerpt (other columns omitted), the first two rows are a Combine
driver row and its build row, and the last two are a correlation trigger and
one of its collaterals:
_cxl_dlq_id,_cxl_dlq_trigger_id,_cxl_dlq_source_name,_cxl_dlq_error_category,_cxl_dlq_trigger
01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,orders,combine_output_row,true
01928f3a-6c10-7b22-9f07-51e6a8b4c302,01928f3a-6c10-7b21-8a4e-3f1c2d9e0a01,rates,combine_output_row,false
01928f3a-6c14-7c03-b2d8-0a9e7f615203,01928f3a-6c14-7c03-b2d8-0a9e7f615203,employees,type_coercion_failure,true
01928f3a-6c19-7d40-8c11-6e2b90d3f404,01928f3a-6c14-7c03-b2d8-0a9e7f615203,employees,correlated,false
A trigger whose failure wrote nothing else carries its own id and shares it
with no other row. Every row has its own _cxl_dlq_id, so the value only
repeats across the rows of one failure.
When a correlation group holds several failing rows, each failing row is a
trigger and keeps its own id as its trigger id. The group’s correlated rows
carry the trigger id of the group’s first failing row, the one whose error
their _cxl_dlq_error_detail quotes. A rejected document has one trigger, its
first failing record; its other records, including any that failed after it,
carry that trigger’s id.
A group larger than max_group_buffer can hold failing rows too. Each of
them is still written as its own trigger, with its own category and id, and
before the rest of the group. The group’s other rows follow under one
group_size_exceeded trigger, the first of them, and the remaining ones are
correlated rows carrying its id. Each failure is written once, as its own
trigger; a row that failed on one Route branch and reached a Sink on another
is written as its own failure, and never again as correlated or as the
overflow trigger. A group whose rows all failed writes no
group_size_exceeded row, because the overflow took nothing with it that
had not already failed.
The rows of one failure can land in different DLQ files: with
per_source paths, a Combine build row goes to its own Source’s file while
its driver row goes to the driver’s.
How the DLQ columns are chosen
Each DLQ file’s header is fixed when the pipeline compiles, before any record is read. It does not depend on which records failed, or on which stages they failed in: every time a pipeline writes a given DLQ file, that file has the same columns in the same order.
A header starts with the _cxl_dlq_* metadata columns, always in this order:
_cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file,
_cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field,
_cxl_dlq_triggering_value, then _cxl_dlq_error_category and
_cxl_dlq_error_detail when include_reason is on, then _cxl_dlq_stage,
_cxl_dlq_route and _cxl_dlq_trigger.
With include_source_row on, the record columns follow. They come from every
record shape that can reach that file:
- the declared columns of each Source, and
_cxl_dlq_source_record, when Source rejections are dead-lettered (strategy: continue); - the shape of the records entering each Transform, Route, Reshape, Aggregate, Combine and Sink that can dead-letter a record under the pipeline’s strategy (stages inside a composition count at the composition’s place in the pipeline);
- the output columns of an Aggregate without
group_byunderstrategy: continue, which go to the pipeline-wide file.
A shape is added to the file of each Source whose records can carry it: that
Source’s per_source.<name>.path when it has one, otherwise the pipeline-wide
path. A shape that carries no Source identity, such as the output of a
Combine using match: first or match: all, is added to the pipeline-wide
file.
The shapes then combine as follows:
- One shape keeps its natural column order.
- Several shapes give their first-seen union in plan order: each column
appears once, at the position where it was first seen. Plan order is the
node order
--explainprints, which need not match the order the nodes are written in the YAML. - A column a record does not carry is an empty cell in that record’s row.
- Engine-only sidecar columns (
$widenedand the$source.*stamps) are never written. Correlation-key columns ($ck.*) are.
For example, this pipeline has two Sources with different columns, one pipeline-wide DLQ file, and a Transform on each Source that can fail:
error_handling:
strategy: continue
dlq:
path: rejects.csv
nodes:
- type: source
name: orders
config:
schema:
- { name: order_id, type: int }
- { name: amount, type: int }
# ...
- type: source
name: refunds
config:
schema:
- { name: refund_id, type: int }
- { name: order_id, type: int }
- { name: amount, type: int }
# ...
# one Transform on each Source, then a Sink on each Transform
In this plan refunds comes before orders, so the record columns of
rejects.csv are:
refund_id,order_id,amount,_cxl_dlq_source_record
An orders row writes an empty refund_id cell. order_id and amount
appear once, although both Sources declare them. _cxl_dlq_source_record is
in the header because a Source rejection can reach the file under continue.
It is empty for these rows.
To see every DLQ file’s columns before a run, use
clinker run pipeline.yaml --explain: its === Dead-Letter Output ===
section lists each file, the Sources routed to it and its full header (see
Explain Plans).
Source-file provenance (_cxl_dlq_source_file) is read from each record, so
one file’s path is never reused for a later row.
Error categories
The _cxl_dlq_error_category column contains one of these values:
| Category | Description |
|---|---|
missing_required_field | A required field is absent from the record |
type_coercion_failure | A value could not be converted to the expected type |
required_field_conversion_failure | A required field exists but its value cannot be converted |
nan_in_output_field | A computation produced NaN |
aggregate_type_error | An aggregate function received an incompatible type |
validation_failure | A declarative validation check failed |
aggregate_finalize | An aggregate function failed during finalization: an integer sum outside the 64-bit range, a decimal total or quotient outside the decimal range, a weighted_avg whose weights total zero or whose row product is out of range, or a group holding both decimals and floats. The reason names the Aggregate and the emit, and gives the fix. Under continue the failed group goes to the dead-letter output. An aggregate in an Envelope footer: stops the run under every strategy. |
correlated | A non-failing record was DLQ’d as collateral because another record in its correlation group failed |
group_size_exceeded | A correlation-key group exceeded the configured max_group_buffer limit |
document_rejected | A non-failing record was DLQ’d as collateral because another record in its document failed under a source’s dlq_granularity: document policy |
late_record | A record arrived at a time-windowed aggregate after its event-time window had already closed |
expansion_limit_exceeded | Per-input fan-out exceeded its authored ceiling. Transform max_expansion rejects before body rows emit; Source max_output_rows_per_input emits exactly its ceiling, then DLQs the original input on the first attempted row above it. Neither is silent truncation. |
combine_output_row | A Combine output-stage eval failed for one driver row (probe-key or on_miss: null_fields body) or for one matched pair (residual or matched body). A failing residual is neither a match nor a miss: on_miss never fires for its driver, under match: all the driver’s other matches are still evaluated and emitted, under match: first a failure on the deciding candidate is the driver’s only result and a failure after it is never written, and under match: collect the driver writes no row; the entry carries the contributing-build lineage and rewinds both the driver and matched build source’s rollback cursor. The driver row and the matched build row each report their own Source in _cxl_dlq_source_name and their own row in _cxl_dlq_source_row, whichever join strategy ran. Routed to the DLQ under continue across every Combine join mode; fail_fast propagates the eval error |
structural_validation | A structural source rule failed: an envelope trailer’s declared count did not match its streamed body, a multi-record body appeared after its closing trailer, or a record type discriminator was unknown. Under dlq_granularity: document, the root cause has trigger: true and every already-streamed record of that file is document_rejected collateral. Under record-grained continue, E345 instead emits only the unknown row with _cxl_dlq_source_record. |
Advanced options
Type error threshold
Abort the pipeline if the fraction of declared source-type failures exceeds a threshold:
type_error_threshold: 0.05 # Abort if >5% of records fail
The cumulative ratio is:
declared source-type failures / decoded source rows observed
The rejected row appears once in both numerator and denominator. The same
typed-error event is used for strategy routing, DLQ accounting, and this
circuit breaker; unrelated validation, structural-document, and collateral
DLQ entries do not enter the numerator. Equality is allowed: a threshold of
0.05 stops only when the ratio is strictly greater than 5%. 0.0 stops on
the first type failure, while 1.0 never trips. Values must be finite and in
[0.0, 1.0].
Correlation key
Declare correlation_key on the contributing Source’s config: block, not on
error_handling:. Group DLQ rejections by a key field. When any record in a correlation group fails, records from the failing source’s contribution to that group are routed to the DLQ:
# Inside a Source's config:
correlation_key: order_id
For compound keys:
# Inside a Source's config:
correlation_key: [order_id, customer_id]
This is useful for transactional data where partial processing of a group is worse than rejecting the entire group. For example, if one line item in an order fails validation, you may want to reject the entire order.
Under multi-source ingest, the collateral fan-out narrows to the failing source: a src_b trigger does NOT DLQ records from src_a that share the same correlation key. Single-source pipelines see bit-identical behavior to today’s pipeline-wide collateral DLQ. See Per-source rollback narrowing for the full semantic and the two documented exceptions (max_group_buffer overflow and Combine output failures).
When a Combine output row fails under a correlation key, the failing driver
row is the trigger of the driver’s correlation group. The dead letter for the
matched build record is held with that group as a collateral
(_cxl_dlq_trigger: false, category combine_output_row), written right after
its driver’s row with its driver’s _cxl_dlq_trigger_id, and written or rolled
back exactly when that group is. It never condemns the build record’s own
correlation group, so another driver that matched the same build record keeps
its output unless its own group failed. Each failure gets its own copy of the
build row, paired with its own driver row: two failing drivers of one group
each get one, and a driver that fails against several build rows is written
once per failure, each copy followed by the build row of that failure. The
held failure keeps its triggering field and value, as it does without a key.
Because every failure is written, a correlation key does not lower
dlq_count, records_dlq or the max_rate numerators: they equal the
counts the same failures give without a key, plus the rows the failing
groups condemn.
For the full lifecycle and per-operator semantics (route, merge, aggregate, combine), see Correlation Keys.
Max group buffer
Limit the number of records buffered per correlation group:
max_group_buffer: 100000 # Default: 100,000
A group that goes over this limit is dead-lettered whole. Its failing rows
are written as their own triggers; its other rows are written under one
group_size_exceeded row, as correlated rows. The limit counts held
entries, not distinct rows: see
Correlation Keys.
Document-level DLQ
By default a record failure dead-letters only that record (dlq_granularity: record). A source can instead reject the entire document any record of which fails, by declaring the granularity per source:
nodes:
- type: source
name: claims
config:
name: claims
type: x12
glob: ./claims/*.edi
schema: [{ name: seg_id, type: string }]
dlq_granularity: document # record (default) | document
Under dlq_granularity: document and the continue strategy, a document is rejected when one of its records fails at the Source (against its declared type, or a structural rule), in a Transform, or in a Route. A record that fails in an Aggregate, a Combine or a Reshape is dead-lettered as its own row, as under record granularity, and does not reject its document (#1232; see Not covered). When a document is rejected:
- the failing record becomes the root-cause DLQ entry (
_cxl_dlq_trigger = true, carrying its original error category); - every other record of the document that reaches a Sink becomes a collateral entry (
_cxl_dlq_trigger = false, categorydocument_rejected); a record dropped before any Sink is not written; - every other record of the document that fails at the Source, in a Transform or in a Route is also a
document_rejectedcollateral, written right after the root cause in the order the records failed. These records are stamped (_cxl_dlq_id,_cxl_dlq_timestamp) when they fail; the records a Sink held are stamped when the Sink rejects the document. Every collateral carries the root cause’s id in_cxl_dlq_trigger_id; - no record read from that document is written by any Sink, however many Sinks read it. Rows derived from the document’s records can still reach a Sink; see Not covered.
Clean documents in the same run stream through untouched, and records from sibling sources still on the default record granularity keep per-record semantics — the policy is per source.
Sinks run last. Every Sink runs after every other node, so a document’s verdict is final before any Sink writes one of its records. --explain lists the Sinks last. Dead-letter rows that other nodes write directly come before the rows a Sink writes.
Several Sinks. Each source row of a rejected document appears in the DLQ once, however many Sinks reached it. A row that failed is written as it was when it failed, ahead of any Sink’s copy of it. Any other row is written as it was held by the first Sink, in run order, that held it; a row that reached only a later Sink (a Route sent it there, say) is written by that Sink. Rows are matched by source row, so of the records one emit each makes from a source row, one is written. Every collateral, whichever Sink writes it, names the document’s root-cause entry in _cxl_dlq_trigger_id.
There is one exception, described under Not covered: a record that an Aggregate, a Combine or a Reshape failure dead-lettered (for a Combine, the driver row and the build row it matched) is written for that failure. It is written again, as a document_rejected row, only when a Sink on another branch also received it and its document is rejected; if its document is not rejected, that Sink can publish it (#1232). The rows a Combine or an Aggregate writes for a rejected document, and the rows that pass through a Reshape, are not held back at all.
This is the document-shaped analogue of correlation keys: use it when partial processing of a document (an EDI interchange, a batch file with a header/trailer) is worse than rejecting the whole document. Unlike correlation keys, which group across files by a key value, document-level DLQ scopes rejection to a single document’s records.
Document grain. The document is the outermost level — the source file. For a flat format (CSV, JSON, plain XML) each input file is one document. For a nested-envelope format (an X12 ISA → GS → ST interchange, an EDIFACT UNB → UNG → UNH) the document is the whole interchange / file, not an inner functional group or transaction set: a failure anywhere in the interchange rejects the entire interchange, including the transaction sets that validated cleanly. Reject the inner-level grain instead by partitioning the input so each interchange is its own file is not currently offered — the grain is fixed at the file.
DLQ rate. Each dead-letter row — the trigger and each collateral — counts once toward the configured DLQ max_rate, matching the correlated-collateral precedent, however many Sinks held it; a record written twice under the exception in Several Sinks counts twice. A rejected 1000-record document contributes 1000 when every one of its records reaches a Sink. It does not affect type_error_threshold, whose numerator contains only declared source-type failures.
dlq_count never counts a row that ok_count also counts, because no Sink writes a record read from a rejected document, with these exceptions, each described under Not covered:
- a record that an Aggregate, a Combine or a Reshape failure dead-lettered, which a Sink on another branch can still write when its document is not rejected (#1232);
- a rejected document’s rows that a Combine or an Aggregate writes, or that pass through a Reshape, which reach a Sink while the document’s records are dead-lettered;
- a record whose joined or aggregated row fails in a node after the Combine or Aggregate: that failure is written while the record itself can still reach a Sink on another branch.
Memory. The engine buffers each open document’s records until its boundary, then flushes the document clean to the sink or rejects it and drops the buffer. Peak memory scales with the concurrently-open documents, not the total input; a single very large document spills its buffer to disk under the run’s memory budget rather than holding everything in RAM. See Streaming vs blocking for the spill model.
Sink restriction. Document-level DLQ flushes each whole document to a single output writer, so it cannot be combined with a per-source-file Sink (a {source_file} / {source_path} path template over a multi-file source). The two are rejected together at compile time (E343); use a single output path, or set dlq_granularity: record if per-file output is the requirement.
Strategy requirement. dlq_granularity: document requires error_handling.strategy: continue. It is incompatible with the default fail_fast: document-level dead-lettering keeps the run going past a bad document, which contradicts fail-fast’s abort-on-first-error. The combination is rejected at compile time (E344) — set strategy: continue to dead-letter bad documents, or keep fail_fast with the default dlq_granularity: record.
Correlation restriction. Document-level and correlation-key rejection are
alternative atomic-disposition models. A document is keyed by its source file;
a correlation group can span files and is keyed by authored field values. The
engine does not define precedence or a combined writer boundary for those two
populations, so a pipeline containing both dlq_granularity: document and any
correlation_key is rejected at compile time (E370). Remove every
correlation_key to keep document rejection, or set dlq_granularity: record
to keep correlation rejection.
Composition restriction. A composition body may not declare a Sink when any
Source uses dlq_granularity: document (E378): a body Sink runs inside its
composition, where it cannot be held back until every document’s verdict is
final. Declare the Sink at pipeline level instead. Where moving it through a
new composition output port works today, the diagnostic prints that move ready
to paste. Otherwise it names clinker explain --code E378, which shows how to
declare the Sink’s work at pipeline level.
Not covered. Document rejection does not yet reach these rows:
- A row failure inside an Aggregate, a Combine or a Reshape is dead-lettered
per record and does not reject its document
(#1232). When a Sink on
another branch also received that record, it is written again as a
document_rejectedrow if another failure rejects its document, and that Sink can publish it if none does. - The rows a Combine or an Aggregate writes are not held back by their document’s verdict, so a rejected document’s joined or aggregated rows still reach a Sink. A failure on such a row, in any node after the Combine or Aggregate, dead-letters only that row and does not reject its document, so the document’s other records, and that row’s own source record on another branch, can still be published (#1317).
- A Reshape’s output rows do not keep their input row’s document, so a rejected document’s rows that pass through a Reshape still reach a Sink.
- In a pipeline where any Source declares
dlq_granularity: document, a CSVjoin_valuescollision withon_conflict: errorat a Sink fails the run rather than dead-lettering the record (#933).
A document is identified by its file path, so two Sources reading the same file share one verdict.
Spilling stages. Document identity survives memory pressure end to end. The per-document buffer identifies each document before buffering and spills under the memory budget, and a Sort between the source and the output preserves each record’s document context — including the source file the grain keys on — across its own spill round-trip. A document whose records pass through a spilling Sort is therefore still grouped and rejected as one document under memory pressure, exactly as it would be in memory. The rows a hash Aggregate or a grace-hash Combine writes are not held back by their document’s verdict, spilled or not; see Not covered.
Malformed envelopes (structural validation)
Envelope formats carry their own structural-integrity claims: an X12 interchange declares a segment count in each SE/GE/IEA trailer, EDIFACT in each UNT/UNZ, HL7 batch/file in each BTS/FTS, and a multi-record flat file’s trailer record declares a body count via its structure: constraint. When the declared count does not match the body the reader actually streamed, the file is structurally invalid. A multi-record flat file can also break a non-count structural rule — a line whose record-type discriminator matches no declared records: entry (E345), or a body record appearing after the trailer that closes the document; these are classified separately from a count mismatch but carry the same disposition.
Under dlq_granularity: document, such a structural failure dead-letters the whole source file to the DLQ rather than aborting the run:
- the file’s records dead-letter as one
structural_validationroot-cause entry (_cxl_dlq_trigger = true) plus adocument_rejectedcollateral for every other already-streamed record of the file; - no record of the malformed file reaches the success sink.
nodes:
- type: source
name: claims
config:
name: claims
type: x12
glob: ./claims/*.edi
schema: [{ name: seg_id, type: string }]
dlq_granularity: document # reuse the document opt-in — no separate config
The opt-in is the same dlq_granularity: document knob that governs per-record document rejection above; there is no separate validation: block. A malformed envelope is simply one more reason a source under the document policy condemns a whole document. E345 is the one structural class that also has a record-grained recovery: under strategy: continue and the default dlq_granularity: record, only the unknown-tag row is dead-lettered and the reader continues at the next physical row. The DLQ row exposes the unknown tag as record_type and the unguessed decoded input as _cxl_dlq_source_record (fixed-width line text, or a JSON array of decoded CSV cells).
Honest timing — rejected at the sink boundary, not before the first record. The trailer that carries the count arrives at the end of the file, after every body record it counts has already streamed through the DAG. Clinker is a bounded-memory streaming engine — it does not buffer the whole file up front to pre-validate it (that would defeat the streaming model). So the count mismatch is detected mid-stream, the file is marked failed, and the document-level DLQ buffer rejects every already-streamed record of the file at its close. The user-visible outcome is the same — no record of a malformed envelope is ever written to the output — but the rejection lands at the sink boundary, not literally before the file’s first record streams.
Grain — the whole file. An SE-level mismatch (one transaction set inside a larger interchange) rejects the entire interchange / file, not just that one transaction set, because the document grain is the outermost source file (see Document grain above). Split the input so each interchange is its own file if you need finer rejection.
Multiple files keep flowing. When a glob / paths source matches several files and one is malformed, only that file dead-letters — ingestion continues to the remaining files, so the clean files after a bad one still reach the sink. (This is unlike the default record granularity, where a count mismatch aborts the whole run and no file’s records are written.) Dead-lettering one malformed file never silently drops the good files around it.
Record-grained E345 is narrow. Under the default dlq_granularity: record, strategy: continue can recover only from an unknown multi-record discriminator because the reader has consumed exactly one bounded physical row and can resume unambiguously. Trailer-count mismatches and a body record after a document-closing trailer still abort at record granularity: neither belongs to one independently recoverable row. Genuine corruption (a truncated stream, a bad delimiter, a control-number echo mismatch, a segment after an X12/EDIFACT/HL7 envelope trailer) always aborts, even under the document opt-in. Under fail_fast, E345 also aborts at the offending line.
Cryptographic integrity (checksums / signatures) is not yet validated. Envelope formats can also carry a SHA-256 body hash, a JWS-signed JSON payload, or an XML Signature. Clinker extracts these envelope sections but does not yet verify them. Tracked for a future release.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Pipeline completed successfully, no errors |
| 1 | Configuration error – the pipeline never started |
| 2 | Pipeline completed, but DLQ entries were produced |
| 3 | Data error halted the run: a fail_fast evaluation/accumulator failure, or the DLQ-rate ceiling |
| 4 | I/O, format, or spill failure |
Exit code 2 is not a failure – it means the pipeline ran to completion and handled errors according to the configured strategy. Check the DLQ file for details. See Exit Codes & Error Diagnosis for the full reference and the orchestrator retry policy.
Complete example
pipeline:
name: order_processing
memory: { limit: "512M" }
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
correlation_key: order_id
schema:
- { name: order_id, type: int }
- { name: customer_id, type: int }
- { name: amount, type: float }
- { name: email, type: string }
- type: transform
name: validate_orders
input: orders
config:
cxl: |
emit order_id = order_id
emit customer_id = customer_id
emit amount = amount
emit email = email
validations:
- field: email
check: "not_empty"
severity: error
message: "Customer email is required"
- check: "amount > 0"
severity: error
message: "Order amount must be positive"
- type: sink
name: valid_orders
input: validate_orders
config:
name: valid_orders
type: csv
path: "./output/valid_orders.csv"
error_handling:
strategy: continue
dlq:
path: "./output/rejected_orders.csv"
include_reason: true
include_source_row: true
type_error_threshold: 0.10
Nodes
A Clinker pipeline is a single flat nodes: list. Every entry carries a
type: discriminator that selects a node kind — the unified node
taxonomy. There is no separate “join section” or “filter section”:
records flow through one homogeneous graph of typed nodes, wired together
by input: / inputs:.
This part documents the ten runtime node kinds; an eleventh,
Composition, is a call-site that inlines a
reusable sub-pipeline and is covered under Pipelines.
The pages in this part are ordered the way data flows through a DAG — a record enters at a Source, fans through the record-level and combining nodes, and leaves at a Sink:
| Node | Role | Arity | Streaming vs blocking |
|---|---|---|---|
| Source | Reads records from a file or network cursor; the entry point. | 0 → 1 | Streaming |
| Transform | Record-level CXL projection, filter, and lookup. | 1 → 0..N records | Streaming |
| Route | Predicate-based fan-out into named branches. | 1 → N | Streaming |
| Merge | Streamwise concatenation of inputs that share a schema. | N → 1 | Streaming |
| Combine | N-ary record combining with mixed predicates (equi + range + arbitrary CXL). | N → 1 | Blocking (build side) |
| Aggregate | Grouped or windowed reduction. | 1 → 1 | Blocking (or streaming when sorted) |
| Reshape | Group-aware mutation and synthesis of records. | 1 input → 1 output port | Blocking |
| Cull | Per-correlation-group removal on a group-level predicate, with a removed_to side-output port. | 1 → 2 | Blocking |
| Envelope | Frames a body stream into per-document documents; a composable framing stage. | 1 → 1 | Streaming |
| Sink | Writes records to an output destination; the exit point. | 1 → 0 | Streaming |
Streaming vs blocking
Stateless nodes (Transform, Route, Merge, the Combine probe side, Sink) evaluate records one at a time without accumulating per-record state. Blocking nodes (Aggregate, sort, the grace-hash Combine build side) accumulate state inside the RSS budget and spill to disk rather than OOM the process. The Streaming vs. Blocking Stages page in the Operations Guide is the full memory model.
Wiring and naming
Every node needs a unique name: (no dots — the dot is reserved for port
syntax). Single-input nodes use input:; Merge and Combine use inputs:.
Route branches are consumed downstream as route_name.port. The
Pipeline YAML Structure page covers the full
wiring grammar, optional fields (description:, _notes:), and strict
parsing rules.
Source Nodes
Source nodes read data from files and are the entry points of every pipeline. They have no input: field – they produce records, they do not consume them.
Basic structure
- type: source
name: customers
config:
name: customers
type: csv
path: "./data/customers.csv"
schema:
- { name: customer_id, type: int }
- { name: name, type: string }
- { name: email, type: string }
- { name: status, type: string }
- { name: amount, type: float }
Schema declaration
The schema: field is required on every source node. Runtime ingestion does not guess types from data: declare each column’s name and CXL type explicitly. This schema drives compile-time type checking across the entire pipeline.
Each entry is a { name, type } pair:
schema:
- { name: employee_id, type: string }
- { name: salary, type: int }
- { name: hired_at, type: date_time }
- { name: is_active, type: bool }
- { name: notes, type: { nullable: string } }
Available types
| Type | Description |
|---|---|
null | The null value only |
string | UTF-8 text |
int | 64-bit signed integer |
float | 64-bit IEEE 754 floating point |
decimal | Exact fixed-point number, subject to the declared precision and scale |
bool | Boolean (true / false) |
date | Calendar date |
date_time | Date with time component |
array | Ordered sequence of values |
map | String-keyed object |
any | Unknown type – field used in type-agnostic contexts |
{ nullable: T } | Nullable wrapper around a concrete inner type, for example { nullable: int } |
Check the schema through the planner
The schema that governs execution is the one accepted by clinker-plan while
compiling the complete pipeline. Run this after changing a source schema:
clinker run pipeline.yaml --explain text
Exit code 0 means the planner parsed the canonical YAML, resolved schema
references and overlays, bound the pipeline, type-checked CXL, and produced a
compiled plan. It does not prove that input bytes satisfy the declarations or
that the resulting output is correct; execute representative input against an
isolated destination and compare the result before a production run.
Workspace tooling may also present advisory schema analysis. That analysis can
help find likely authoring mistakes, but it cannot admit or reject a pipeline.
Its analyzed, partial, skipped, and failed statuses describe only how
much the advisory model inspected. See Validation and
Admission for the
status meanings and known limits.
Without a column format:, date_time accepts the default offset-free forms
and RFC 3339 timestamps with Z or an explicit numeric offset, such as
2026-01-31T08:27:00Z and 2026-01-31T10:57:00+02:30. Zoned timestamps are
normalized to UTC before entering Clinker’s timezone-free date_time
representation. Parsing is exact: surrounding whitespace, malformed calendar
values, and out-of-range timestamps are rejected rather than trimmed, guessed,
overflowed, or rounded. A column-level strftime format: is exclusive: only
that authored format is tried.
A source column’s declared type must be concrete: numeric — the
inference-only int | float union CXL resolves during type unification — is
not a valid source column type. Declaring one is rejected at compile with
E158; declare int or
float explicitly.
clinker guess pipeline.yaml provides an authoring-only preview for columns
explicitly marked numeric. It proposes concrete int or float declarations;
it does not infer arbitrary schemas or change the runtime admission rules.
The default preview is bounded and read-only. --check exhausts the admitted
manifest; --write can publish one guarded, unambiguous inline edit. Resolve
all numeric declarations to concrete types before executing the pipeline.
See Numeric schema authoring and the
guess command.
long_unique — storage hint for high-cardinality text
A string column may carry an optional long_unique: true flag. It is an
advisory, opt-in hint, not a type change: it tells Clinker the column’s
values are long and effectively unique — never repeated across records — so
the run uses less memory for that column. Typical candidates are UUIDs
rendered as text, street addresses, and free-text comment or note fields.
schema:
- { name: ticket_id, type: string, long_unique: true } # 36-char UUID
- { name: notes, type: string, long_unique: true } # free text
- { name: department, type: string } # low-cardinality, default
The flag lowers memory use only. A value’s content and its
comparison, grouping, join, sort, and output behavior are all unchanged — a
long_unique value behaves identically to the same text in any other column.
Omitting the flag (the common case) leaves the default behavior untouched. Set
it only when you know a column is genuinely high-cardinality free text; on a
column whose values repeat, leave it off.
source_name — read a differently-named physical column
A column may carry an optional source_name naming the physical input
column it reads from, when that differs from the exposed name. The reader
matches input fields by physical name and re-labels the value under name, so
downstream CXL and the output see name carrying the physical column’s data.
schema:
# read the physical `cust_id` column, expose it downstream as `customer_id`
- { name: customer_id, type: string, source_name: cust_id }
Omitting source_name (the common case) reads the input field whose key equals
name, unchanged from before. A channel schema patch’s rename op sets this
alias automatically (see Channels).
Transport vs format
A source declaration has two independent layers:
- Transport (
transport:) selects where the records come from. Two transports exist:file— read bytes from the filesystem, resolved through one of the file matchers (path/glob/regex/paths) — andrest— pull records from a paginated HTTP endpoint under a hard page/record cap (see Network Sources (REST)).transport:is optional and defaults tofile, so a source that omits it reads from disk exactly as before.restneeds therestcapability, which the released binary has; a build compiled without it refuses arestsource at validation withE223rather than running the pipeline short one input (see Optional capabilities). - Format (
type:) selects how the bytes decode into records:csv,json,xml,fixed_width,edifact,x12,hl7,swift.
- type: source
name: orders
config:
name: orders
transport: file # optional; this is the default
type: csv # the on-disk format
path: "./data/orders.csv"
schema:
- { name: order_id, type: int }
A file transport requires exactly one file matcher (path, glob, regex, or paths). Declaring none fails validation with E211; declaring more than one fails with E210. Both are reported at config-load time, before any file is opened.
Choosing files
Interactive companion: the file discovery explainer runs these steps on a sample folder as you change the settings.
Before reading, a file Source builds its list of files in a fixed order:
- Matcher. Paths are relative to the pipeline file’s folder.
path:names one file andpaths:a list. A file that does not exist stops the run withE216, whateveron_no_matchsays.glob:matches names inside one folder:*does not cross a/. Use**to include subfolders (./data/**/*.csv). A leading dot is not special, so*.csvalso matches.draft.csv. An invalid pattern isE212.regex:searches the pipeline’s folder and, by default, every subfolder, and matches anywhere in each file’s path. That path begins with the folder part of the pipeline path as you typed it (clinker run pipelines/orders.yamlgivespipelines/data/orders_2024-01.csv), so anchor the end of the path, for exampledata/orders_\d{4}-\d{2}\.csv$. An unanchored pattern can pick up earlier outputs too. An invalid pattern isE213.
exclude:– a list of glob patterns; a file whose name or full path matches any of them is dropped.- Anything that is not a regular file is dropped.
min_size:/max_size:– decimal units:1KBis 1000 bytes, alsoB,MB,GB.modified_after:/modified_before:– a duration back from the time the pipeline is loaded (30s,15m,2h,3d) or an RFC 3339 timestamp (2024-03-01T00:00:00Z).files.sort_by:name(default; the full path),createdormodified, withfiles.sort_order:asc(default) ordesc. The files of apaths:list are sorted too: the order they are written in is not the reading order.files.take_first:orfiles.take_last:keeps that many files from the sorted list (setting both isE218). With the default ascending name sort,take_first: 5keeps the five earliest names.files.on_no_match:– when nothing is left:error(default,E216),warn(log a warning and produce no rows), orskip(produce no rows quietly).
files.recursive: controls whether regex: searches subfolders (true by default). It has no effect on glob:, which searches subfolders only where the pattern has **.
- type: source
name: orders
config:
name: orders
type: csv
glob: ./data/orders_*.csv
exclude: ["*_partial.csv"]
modified_after: 30d
files:
sort_by: name
sort_order: desc
take_first: 3 # the three latest names
on_no_match: warn
schema:
- { name: order_id, type: string }
With glob:, regex: or paths:, each file is read as its own document: a declared sort_order is checked per file, and an Aggregate rolls up per file (see Sort order and One document per file).
Format types
The type: field inside config: selects the on-disk format. Each format
has its own reference page covering its options and decoding model:
type: | Format | Reference |
|---|---|---|
csv | Delimited text (RFC 4180) | CSV Format |
json | Array / NDJSON / wrapper object | JSON Format |
xml | Element-path-selected record elements | XML Format |
fixed_width | Column-positioned legacy extracts | Fixed-Width Format |
edifact | UN/EDIFACT interchanges | EDIFACT Format |
x12 | ANSI ASC X12 interchanges | X12 Format |
hl7 | HL7 v2.x pipe-and-hat messages | HL7 v2 Format |
swift | SWIFT MT (FIN) messages | SWIFT MT Format |
The same schema: rules apply regardless of format: the reader maps each
decoded record onto the declared schema, and undeclared input fields fall
under the on_unmapped policy
below.
Declared-type failures
Source types are enforced before a record reaches buffering, sorting, or any downstream node. A value that cannot satisfy its authored type rejects the whole row exactly once; Clinker never substitutes the raw string, a sentinel, or an error-derived null.
Empty input has three distinct outcomes:
- an empty value declared as
stringremains the empty string; - an empty non-string value declared
nullable(T)becomes null; - an empty non-string value declared as non-nullable is a type error.
Parsing does not trim whitespace, guess locale conventions, or recognize
case-insensitive null sentinels. Integer overflow, invalid/out-of-range dates,
decimal precision overflow, and decimal values that would need rounding to
the declared scale are type errors. Accepted decimal and date values retain
their declared precision.
The E126 diagnostic identifies source, file, one-based row and column,
field, and declared type. Its value preview is sanitized to one line and
limited to 256 rendered UTF-8 bytes; controls, bidi characters, diagnostic
delimiters, backslashes, and invalid UTF-8 are explicit indivisible escape
tokens. When truncated, the preview ends in one … without splitting a token
or Unicode scalar and reports the original byte length. The complete original
record/value is retained only in the configured DLQ. See Error Handling &
DLQ for strategy and threshold
behavior.
Already-decoded values obey the same declarations without hidden coercions:
string admits only a string, null admits only null, and any admits every
supported native value while preserving its represented value. Numeric
conversions must be exact; non-finite floating-point values and decimal values
that would require rounding are rejected. With multiple: true, the scalar
declaration is applied to every array element. If any element fails, the whole
source record is rejected and the complete original array remains available
through the DLQ.
Error strategies and complete populations
fail_fast aborts on the first declared source failure. continue and
best_effort require a DLQ and use the configured threshold over the complete
population:
rejected records / attempted records
The run aborts only when that ratio is strictly greater than the threshold.
Equality is accepted, so a threshold of 0.1 admits exactly one rejected
record in a population of ten but not two. A zero threshold aborts on the first
rejection; a threshold of 1.0 admits an all-rejected population for the
continuing strategies.
For ordered file sources, Clinker establishes the complete attempted and rejected population before any accepted record, punctuation, downstream side effect, or output byte is released. That population is applied exactly once whether execution remains resident, spills, or fuses a downstream node. A threshold violation therefore cannot leave a committed prefix of output.
on_unmapped — undeclared input fields
The per-source on_unmapped policy decides what to do with input fields the source’s schema: block does not name. Three modes — auto_widen (default), drop, reject:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
on_unmapped:
mode: auto_widen # default; other values: drop, reject
schema:
- { name: order_id, type: string }
- { name: amount, type: float }
See Auto-Widen & Schema Drift for the full
specification: how undeclared columns flow through each downstream node
type, the include_unmapped Sink flag, E315 merge-policy
mismatch, and fixed-width behavior.
Sort order
If each physical input file is pre-sorted, declare the record order so the planner can admit order-dependent strategies such as streaming aggregation:
- type: source
name: sorted_transactions
config:
name: sorted_transactions
type: csv
path: "./data/transactions_sorted.csv"
schema:
- { name: account_id, type: string }
- { name: txn_date, type: date }
- { name: amount, type: float }
sort_order:
- { field: "account_id", order: asc }
- { field: "txn_date", order: asc }
Clinker binds these fields to the declared source schema, compares the typed values by the same rule every sort uses (see How values are ordered), and verifies each physical file independently before any record from that file reaches an order-dependent consumer. The declaration never means that a multi-file source is globally sorted: the last key in one file is not compared with the first key in the next file.
on_unsorted controls the result of the first adjacent inversion:
| Value | Behavior |
|---|---|
warn (default) | Stably repair the complete physical file with the shared bounded-memory sort, emit one W307 warning for that file, then release it. A file already in order emits no warning. |
error | Reject the physical file without releasing an unverified prefix. The diagnostic identifies the source, file, adjacent rows, and keys. |
sort_order:
- { field: "account_id", order: asc, null_order: last }
- { field: "txn_date", order: desc, null_order: first }
on_unsorted: warn
Source ordering accepts null_order: first or last; drop is rejected
when the pipeline is planned because verifying order must not discard source
records. The error gives one fix: delete null_order: drop and add a
Transform after the Source whose whole config is the line it prints,
config: { cxl: "filter not <field>.is_null()" }. With the line deleted, the
Source declares its null keys last; if a file’s null keys arrive first,
write null_order: first instead of deleting the line. That filter needs a
field CXL can name as it is: one identifier of ASCII letters, digits and _,
not starting with a digit and not a CXL keyword. For any other key, such as
order id, filter or a flattened Address.City, the error prints no CXL
and prints a source_name
line for the column instead, such as source_name: "order id". Set the
column’s name to a new identifier, add that line, use the new name wherever
the pipeline names the column, and plan again: the error then prints the
filter on the new name. Equal authored keys
retain arrival order within the selected execution path. Clinker does not add
a source identity, physical filename, or canonical-row tie-breaker.
Verification stages the complete sortable file event sequence behind the
run’s memory arbitrator. It uses the existing stable resident/spill sort and
bounded-fan-in merge machinery when repair spills, while preserving row
identity, source/file provenance, and document context. Only flat sources and
sources with one sortable frame per physical file are admitted; a format whose
nested or repeated framing cannot be reordered losslessly fails during
planning with a correction to remove sort_order or normalize the input.
If a downstream consumer needs one global order across all files, declare
sort_order on the terminal Sink. Use enough output
fields to define a total business order when byte-identical output matters.
The shorthand form is also accepted – a bare string defaults to ascending:
sort_order:
- "account_id"
- { field: "txn_date", order: desc }
Watermarks
An event-time watermark declares which column on the source carries
each record’s event time — the wall-clock instant the event
happened, distinct from when Clinker read the row. When set, Clinker
takes the column on every record, subtracts the source’s delay, and
uses the result to track event-time progress so downstream time
windows know when to close. The delay-corrected value is also stamped
on every record as
$source.event_time, the column a
downstream time-windowed aggregate
uses to assign records to windows.
- type: source
name: clicks
config:
name: clicks
type: csv
path: "./data/clicks.csv"
options:
has_header: true
watermark:
column: event_ts # must be date_time or date
delay: 5s # bounded out-of-order tolerance
idle_timeout: 30s # flip partitions to idle if quiet
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: amount, type: int }
Fields:
-
column(required) — the schema column whose value is each record’s event time. The column’s declared type must bedate_timeordate. Acolumn:that names a field absent fromschema:raises E154; acolumn:whose declared type is neither raises E155. -
delay(optional duration, default unset) — bounded out-of-order tolerance. Each record’s event time is shifted earlier bydelaybefore being folded into the watermark, so the source’s effective watermark trails its observed max event time by this amount. Mirrors Flink’sBoundedOutOfOrdernessWatermarks. Withoutdelay, the watermark advances strictly to the observed max — a single late record routes to the DLQ. -
idle_timeout(optional duration, default unset) — if a source stays quiet longer than this, it stops holding back downstream window-close progress, so windows keep closing when one source pauses. Unset means the source never goes idle.
Durations use the suffixes ms, s, m, h, d. ms is
matched before the single-character s, so 500ms reads as 500
milliseconds, not 500 seconds with a stray m.
A pipeline whose aggregate declares time_window: must have a
watermark.column on every upstream-reachable source. Without it,
event-time progress can never advance and the window can never close —
the planner rejects this with
E156.
Multi-value fields
A field that holds more than one value is declared on the schema column, not on the pipeline:
schema:
- { name: order_id, type: string }
- { name: tags, type: string, multiple: true }
multiple: true says the column holds zero or more values of its declared
type. Reading collects every occurrence of the field into one array — a single
occurrence is still an array, so downstream code never has to branch on how
many values happened to arrive. A field absent from a record has no column at
all and resolves to null, exactly as any other absent column does. CXL sees the
column as an array; the declared type: describes each element and drives
coercion. The declaration describes the shape of the
data, so it serves both directions: a writer that can encode repetition reads
the same declaration.
The split_to_rows, split_values, and join_values blocks below accept a
compact shorthand (a bare field name, or a mapping that omits defaults). To see
the fully-materialized form the engine actually runs — every default spelled
out — print the canonical config with
clinker config --resolved.
It rewrites only those shorthand blocks and leaves the rest of the file
untouched.
Both ends of the declaration are checked at compile, so a shape the formats cannot carry fails before a run starts rather than mid-stream:
| Format | As a source | As an output |
|---|---|---|
json | native — an array | native — an array |
xml | native — repeated child elements | not yet (issue 916) |
csv | delimited cell via split_values | delimited cell via join_values |
fixed_width | delimited cell via split_values | not yet (issue 918) |
edifact, x12, hl7, swift | no — repetition is positional | no — repetition is positional |
A multiple: true column reaching an output that cannot encode it is E359; one
on a source that cannot produce it is E361. E359 covers an output’s own
schema: block too — the attribute is direction-neutral, but the remaining
writers do not encode repetition yet, so declaring it on such a sink would be
accepted and ignored. Run clinker explain --code E361 for the full remediation
of either.
csv and fixed_width read a multiple: true column through a
split_values entry. Neither wire format repeats a field, but a cell’s text
may hold several values separated by a delimiter. Declare that delimiter with a
split_values entry and the reader
parses the cell into the array the column holds. A multiple: true column no
entry covers is rejected by E361 — the reader would have no delimiter and
deliver the raw cell; either add the entry or leave the column single-valued and
split it in a transform (tags.split(";")). The entry is read only on a
single-schema source: a multi-record source of either format runs a backend that
does not consume it. On the output side, a CSV sink joins a multiple: field
into one delimited cell with join_values
(defaulting to ; / on_conflict: error); fixed_width output is still pending
(#918).
A split_values entry also recovers a CSV cell a sink wrote under join_values
on_conflict: escape or encode_json: add escape: "\\" to un-escape an escaped
delimiter, or json: true to read the whole cell as an embedded JSON array.
The segment formats are a permanent no, not a pending one. Repetition there
is a positional coordinate rather than a list: a repeated composite is written
as two axes interleaved in one element (11:B:1^12:B:2), which a flat array
cannot represent without losing the component axis. The faithful shape is one
column per coordinate — for HL7, that is what options.split_fields produces,
with a writer that reassembles the wire field byte-for-byte.
One record per value: split_to_rows
split_to_rows fans a record out to one record per occurrence of a repeated
field. Each entry is either a bare field name or a full mapping, and the two
forms mix freely in one list:
- type: source
name: invoices
config:
name: invoices
type: json
path: "./data/invoices.json"
schema:
- { name: invoice_id, type: int }
- { name: customer, type: string }
- { name: line_item, type: string }
- { name: line_amount, type: float }
- { name: line_no, type: int }
split_to_rows:
- tags # shorthand: field name, all defaults
- field: line_items # full form
keep_empty: true
mode: extract
position_column: line_no
max_output_rows_per_input: 10000
| Key | Default | Meaning |
|---|---|---|
field | — | The repeated field, as a flattened dotted name |
keep_empty | true | Whether a record whose field is empty or absent survives |
mode | extract | extract — the occurrence becomes the record; split — the record shape is kept |
position_column | none | Column receiving each occurrence’s 1-based position |
The field is named as it appears in the input document, not as the schema
exposes it: a column declared source_name: is addressed by that
source_name. The same rule applies to split_values below.
keep_empty defaults to true. A record whose field holds an empty array,
or carries no such field at all, is emitted with that field unset rather than
disappearing. Several widely used engines drop the record instead; a vanished
row is the costliest failure mode there is, so dropping is opt-in here.
mode: extract (the default) makes the occurrence the record: its own
fields are lifted out from under the field name and every field outside the
group is merged onto each output. {"orders": [{"id": 1}]} yields a top-level
id, and repeated <Item><name> children yield name. When lifting lands an
occurrence’s field on a name an outside field already occupies, the occurrence
wins — it is the record, so its own value is not shadowed by the parent it was
merged with. A position_column wins over both: you named it, so a field of
the same name inside or outside the occurrence gives way to the index.
mode: split preserves the record shape: the occurrence’s fields keep
their dotted path (orders.id, Item.name) and each output carries exactly
one occurrence.
Entries apply in declaration order, so two entries multiply. Declaring the same
field twice is rejected at compile (E358), as is fanning out a field the
schema also declares multiple: true — the attribute collects the occurrences
into one array, the fan-out spends them one per record, and a field cannot be
both. On an XML source, two entries may not name nested element groups
either (Item and Item.part): that reader assigns each element to one
occurrence group by document position, and a nested pair leaves the inner
group’s membership ambiguous.
max_output_rows_per_input is valid only alongside a non-empty
split_to_rows block and bounds that cumulative product for one original
JSON object or XML record element. Omit it, or set it to 0, for no ceiling.
With a positive value N, the reader emits exactly the first N records in
the same declaration order shown above. If an N+1 row is attempted, the
reader stops that input and reports an expansion_limit_exceeded source
failure. Under error_handling.strategy: continue, the complete decoded
original input representation is routed to the DLQ; the first N records
remain valid output. This is not
silent truncation: the run records one explicit rejection naming the field,
the configured ceiling, and the exact first violating count (N+1). Under
fail_fast, the run fails at that boundary.
The check is lazy: the reader never materializes or pre-counts the Cartesian
product. Its cursor holds the original input, the declared occurrence lists,
and one output record. This source setting is separate from a Transform’s
max_expansion; when both surfaces
fan out, each enforces its own per-input boundary.
A JSON source accepts a nested pair to produce a two-level expansion, but
only when the outer entry declares mode: split. mode: extract lifts the
occurrence’s own keys to the top level, which removes the dotted path the inner
entry addresses — the inner entry would then match nothing and fan nothing out,
so the pairing is rejected (E358):
split_to_rows:
- { field: orders, mode: split }
- { field: orders.items, mode: split }
Several values in one cell: split_values
split_values parses a delimited cell into the several values a multiple:
column holds. It takes the same bare-name-or-mapping shorthand:
split_values:
- tags # shorthand: default delimiter `;`
- field: codes # full form
delimiter: "|"
schema:
- { name: tags, type: string, multiple: true }
- { name: codes, type: string, multiple: true }
The delimiter defaults to ;. A split_values field the schema does not
declare multiple: true is rejected at compile (E358): splitting produces
several values, and only a multi-value column can hold them. So is an entry
naming a column’s exposed name when that column reads a differently-named input
field — the split runs against the document’s own field names, so name the
source_name.
The entry is read by the JSON and XML readers, and — on a single-schema
source — by the CSV and fixed-width readers. On a multi-record CSV or
fixed-width source, or on any segment format, declaring it is rejected (E358)
rather than silently ignored: those readers are never handed it, so the cell
would arrive unsplit with nothing to say so.
Migrating from array_paths
array_paths: was the earlier form of these declarations. It is no longer read,
and a source still carrying it is rejected at compile (E360) rather than
running with the fan-out silently dropped. An explode path becomes a
split_to_rows: entry, a delimited cell becomes a split_values: entry, and a
path kept as an array becomes multiple: true on the schema column.
mode: extract (the default) reproduces the old projection — the element’s
own fields lifted to the top level of each output record. It does not reproduce
the old cardinality: explode dropped a record whose array was empty, while
keep_empty defaults to true here and keeps it with the element’s fields
unset. Add keep_empty: false to the entry to migrate row-for-row.
Format notes
split_to_rows is honored by the JSON and XML readers, over a file path
and over a rest response body alike. split_values is honored by those two and
also by the single-schema CSV and fixed-width readers. Declaring a knob on
a reader that is never handed it — split_to_rows on any delimited-cell or
segment format, split_values on a multi-record or segment source — is rejected
at compile (E358) rather than accepted and inert, and a multiple: true column
no split_values entry can cover is rejected by E361.
Composition body files are gated by the same four checks as the pipeline that calls them.
JSON — the field names the key holding the array (line_items, or
order.line_items for an array nested under an object). A field present but
holding a single object or scalar rather than an array counts as one
occurrence, and is projected exactly as a one-element array would be — many
producers unwrap a lone element, and XML cannot express the difference at all,
so the two readers agree on this. A field that is absent, holds an empty array,
or is explicitly null has no occurrence and is governed by keep_empty: an
explicit null is how many producers write “no value”, and it counts as none
rather than as one. For the same reason a multiple: true column holding an
explicit null stays null rather than becoming [null], so size() over it
reads the same as it does for a field the document omits.
XML — the field is the repeated child element’s dotted path relative to the
record element (Item, or Items.Item when nested). Repetition and absence
are indistinguishable in XML, so a record with no occurrence of the element is
governed by keep_empty exactly as an empty array is. See
XML Format for the full rules.
Transform Nodes
Transform nodes apply CXL expressions to each record, producing new fields, filtering records, or both. They process one record at a time in streaming fashion with constant memory overhead.
Basic structure
- type: transform
name: enrich
input: customers
config:
cxl: |
emit full_name = first_name.concat(" ", last_name)
emit tier = if lifetime_value >= 10000 then "gold" else "standard"
filter status == "active"
The cxl: field is required and contains a CXL program. The three core CXL statements for transforms are:
emit– adds a field to the record, or replaces the field of the same name. The record’s other input fields are carried through to downstream nodes without being emitted; a Sink withinclude_unmapped: falseis the one place that narrows them (see Sink Nodes).filter– drops records that do not match the boolean condition.let– binds a local variable for use in subsequent expressions (not emitted).
cxl: |
let margin = revenue - cost
emit product_id = product_id
emit margin = margin
emit margin_pct = if revenue > 0 then margin / revenue * 100 else 0
filter margin > 0
Analytic window
The analytic_window field enables cross-source lookups by joining a secondary dataset into the transform. The secondary source is loaded into memory and indexed by the join key.
- type: transform
name: enrich_orders
input: orders
config:
analytic_window:
source: products
on: product_id
group_by: [product_id]
cxl: |
emit order_id = order_id
emit product_name = $window.first().product_name
emit quantity = quantity
emit line_total = quantity * price
The $window.* namespace provides access to the windowed data. Functions like $window.first().<field>, $window.last().<field>, and $window.count() operate over the matched group. first() and last() return a whole record, so name the field to read; without one they return null. See Window Functions.
Validations
Declarative validation checks can be attached to a transform. They run against each record and either route failures to the DLQ (severity error) or log a warning and continue (severity warn).
- type: transform
name: validate_orders
input: raw_orders
config:
cxl: |
emit order_id = order_id
emit amount = amount
emit email = email
validations:
- field: email
check: "not_empty"
severity: error
message: "Email is required"
- check: "amount > 0"
severity: warn
message: "Non-positive amount"
- field: order_id
check: "not_empty"
severity: error
Validation fields
| Field | Required | Description |
|---|---|---|
field | No | Restrict the check to a single field |
check | Yes | Validation name (e.g. "not_empty") or CXL boolean expression |
severity | No | error (default) routes to DLQ; warn logs and continues |
message | No | Custom error message for DLQ entries |
name | No | Validation name for DLQ reporting. Auto-derived from field + check if omitted |
args | No | Additional arguments as key-value pairs |
Expansion cap (max_expansion)
When a transform body contains an emit each statement, every input record can fan out into multiple output records. The max_expansion field caps how many output records a single input record may produce – a safety bound against unexpectedly large arrays.
- type: transform
name: explode_items
input: orders
config:
max_expansion: 5000 # default: 10000
cxl: |
emit each it in items {
emit order_id = order_id
emit sku = it["sku"]
emit price = it["price"]
}
| Field | Type | Default | Description |
|---|---|---|---|
max_expansion | u64 | 10000 | Maximum cumulative output records per input record. |
If a single input record’s emit each block produces more than max_expansion output records, the originating record routes to the DLQ with category expansion_limit_exceeded instead of producing a truncated or unbounded result. No partial output is emitted for that record – the cap is enforced eagerly so the writer never sees records from a runaway expansion.
When to tune
- Lower (e.g.
100,1000) when input arrays are bounded by a known business rule and you want hostile or malformed input to surface as a DLQ entry rather than as a flood of downstream records. - Higher (e.g.
100000,1000000) when legitimate input carries large arrays – for example, an order with a long line-item list or an event carrying a per-second pricing curve.
The DLQ category expansion_limit_exceeded is distinct from generic CXL evaluation failures, so DLQ-side filters and metrics can target expansion runaway specifically. See Error Handling & DLQ for the wider DLQ contract.
Batch size (batch_size)
A streaming-eligible transform hands its output downstream in bounded batches rather than accumulating the whole stage before the next stage runs. batch_size sets how many events (records plus document-boundary punctuations) a batch holds. A per-transform batch_size overrides the pipeline-level pipeline.batch_size for this one stage; omit it to inherit the pipeline value (or the built-in default of 2048).
- type: transform
name: enrich
input: orders
config:
batch_size: 512 # override pipeline.batch_size for this stage
cxl: |
emit order_id = order_id
emit total = quantity * unit_price
| Field | Type | Default | Description |
|---|---|---|---|
batch_size | usize | inherits pipeline.batch_size (else 2048) | Events per streaming batch for this transform. Must be >= 1. |
A batch_size of 0 is rejected at config load (a zero-event batch never flushes). Smaller batches lower the in-flight memory of a streaming stage at the cost of more per-batch bookkeeping; larger batches amortize the bookkeeping at the cost of a larger live working set. The default suits typical record widths — tune it only when a profiling run shows a streaming stage’s per-batch footprint matters. See Streaming vs. Blocking Stages for which stages stream and which fully materialize.
Log directives
Log directives declare bounded structured diagnostic events during transform execution:
- type: transform
name: process
input: validated
config:
cxl: |
emit id = id
emit result = compute(value)
log:
- name: transform.record_processed
level: info
when: per_record
every: 1000
message: "Processed record"
fields: [id]
- name: transform.record_failed
level: warn
when: on_error
message: "Record failed processing"
- name: transform.started
level: debug
when: before_transform
message: "Starting transform"
Log directive fields
| Field | Required | Description |
|---|---|---|
name | Yes | Stable event name: a bounded dotted identifier using ASCII letters, digits, or underscores |
level | Yes | trace, debug, info, warn, or error |
when | Yes | before_transform, after_transform, per_record, or on_error |
message | Yes | Static event message, at most 1024 UTF-8 bytes. Interpolation is rejected; request record values with fields instead |
every | For per_record | Positive record interval. It is required for every per_record event, including explicit every: 1, and rejected for other timings |
fields | No | Up to 256 unique record field names requested as structured attributes. Available only for per_record and on_error events |
condition | No | CXL boolean expression; the event fires only for records where it is true. Available only for when: per_record, and at most 512 UTF-8 bytes |
A transform may declare at most 32 events and request at most 256 fields in aggregate across them. Event names and field selectors use the same grammar as deployment field policy: dot-separated segments beginning with an ASCII letter or underscore, followed by ASCII letters, digits, or underscores.
fields is the only channel by which record data reaches an event — message
is static text. A selector naming a field the incoming record does not carry
contributes nothing, so a directive whose selectors all miss would publish an
event with no attributes at all, which reads exactly like a run whose records
were empty.
Clinker refuses that when the pipeline compiles (E374). The rejection names the
selector, lists the columns the input row does carry, and — when your spelling
is close to one of them — gives you the corrected fields: line to paste:
[E374] transform `enrich` log[0].fields requests `orderId`, which the input
record does not carry; the upstream row has `order_id`, `amount`, `region` —
write `fields: [order_id]`
Selectors bind against the transform’s input row, for per_record and
on_error alike: dispatch fires before this transform’s own cxl: block, so a
column the transform produces cannot be requested. Request the columns it reads
instead.
The check decides what the declared schema can decide. A column that reaches the transform through an open composition port is not visible to it, and a selector naming one of those is still checked only at run time — counted in the run’s admission accounting under the missing-field total.
One event name, one set of fields
An event name is the identity a collector groups records by, and it carries no
node identity of its own. Two transforms may emit the same event — that is how
one thing that happens in several places is reported as one thing — but they
have to declare the same fields, or the collector receives two record shapes
under one name and nothing downstream can separate them.
Clinker checks this across the whole pipeline, composition bodies included, and refuses a disagreement with E375:
[E375] transform "shape" in composition "customer_enrichment": `log` event
"transform.customer_seen" is also declared by transform "seen" with different
fields; one event name carries one set of fields — give both declarations the
same `fields`, or give one of them its own event name
A composition used twice is not a conflict: both instances declare the same event with the same fields, and agreement is what the rule asks for.
Logging only the records you care about
every thins a per-record event by count. condition selects it by content —
use it when the interesting records are rare and you want all of them rather
than every thousandth record:
log:
- name: transform.large_order
level: info
when: per_record
every: 1
condition: "amount > 1000"
message: "large order"
fields: [order_id, amount]
The two compose: every is applied first, then condition, so every: 100
with a condition logs every hundredth record that also matches.
A condition is CXL, checked when the pipeline compiles. It must resolve to a
boolean, and it is evaluated against the transform’s input record — the one
that arrived, before this transform’s own cxl: block runs. A field the
transform only produces is therefore not in scope; write the condition in terms
of the fields the transform reads.
A condition decides only whether an event fires. It cannot add anything to
one: the values that leave the process are still exactly the fields you
requested, each still subject to deployment policy. Narrowing a condition can
never widen what is exported.
Transform declarations name events, request fields, and may gate a per-record event on its own input; they do not choose a destination, credentials, routing, redaction, or sampling policy. Each requested event-field pair is denied unless deployment observability policy explicitly allows, hashes, or replaces it. Telemetry delivery is bounded and best effort and cannot change transform results or published output — including a condition that fails to evaluate, which drops its event rather than failing the run.
Every transform event also carries the fixed correlation fields
execution_id, batch_id, and pipeline_name. Unlike requested record
fields, these are not default-deny and are not gated by
field_policy: they are engine-supplied identity that never derives from a
source, and they are what makes an exported event joinable to the machine
stream and to the lineage events. A deployment that allows an event without
also writing three correlation rules still gets telemetry it can correlate.
Because they are exported verbatim, choose a --batch-id that is safe to send
to your collector — an identifier, not a tenant name or anything else you
would not want retained there. A field_policy rule naming one of these three
fields does not redact it. Source paths, records, secrets, and raw error text
are never implicit attributes.
The former log_rule directive key and pipeline-level log_rules block are
rejected. Move event identity and safe field requests into the transform’s
log: entries as shown above; keep routing, privacy, credentials, and sampling
in deployment policy.
Complete example
- type: source
name: employees
config:
name: employees
type: csv
path: "./data/employees.csv"
schema:
- { name: employee_id, type: string }
- { name: first_name, type: string }
- { name: last_name, type: string }
- { name: department, type: string }
- { name: salary, type: int }
- { name: hire_date, type: date }
- type: transform
name: enrich_employees
description: "Compute display name and tenure"
input: employees
config:
cxl: |
emit employee_id = employee_id
emit display_name = last_name.concat(", ", first_name)
emit department = department.upper()
emit salary = salary
emit annual_bonus = if salary >= 80000 then salary * 0.15
else salary * 0.10
validations:
- field: employee_id
check: "not_empty"
severity: error
message: "Employee ID is required"
- check: "salary > 0"
severity: warn
message: "Salary should be positive"
log:
- name: transform.employee_processed
level: info
when: per_record
every: 5000
message: "Processing employees"
fields: [employee_id]
Route Nodes
Route nodes split a stream of records into named branches based on CXL boolean conditions. Each branch becomes an independent output port that downstream nodes can wire to using port syntax.
Interactive companion: the Route and Merge explainer shows, record by record, which conditions are checked and where each record goes, and how a Merge rejoins the branches.
Basic structure
- type: route
name: split_by_value
input: orders
config:
mode: exclusive
conditions:
high: "amount.to_int() > 1000"
medium: "amount.to_int() > 100"
default: low
This creates three output ports: split_by_value.high, split_by_value.medium, and split_by_value.low.
Conditions
The conditions: field is an ordered map of branch names to CXL boolean expressions. Each expression is evaluated against the incoming record.
conditions:
priority: "urgency == \"high\" and amount > 500"
standard: "urgency == \"medium\""
bulk: "quantity > 100"
default: other
Condition keys become the port names used in downstream input: wiring.
Compile-time checking
Branch conditions are typechecked when the pipeline is compiled – the plan-building pass you trigger with clinker run pipeline.yaml --explain or a bare --dry-run (see Validation & Dry Run). Each condition is checked against the concrete column types of the route’s input, so a condition that references an unknown column or compares incompatible types fails at compile time rather than partway through a run.
Compile failures surface as a CXL diagnostic keyed to the failure class – E202 for a branch condition that does not parse, E203 for one that references an unknown column, and E200 for one that compares incompatible types – and name the offending branch, for example split_by_value (branch high), so in a multi-branch route the error points at the specific branch rather than just the route node.
Default branch
The default: field is required. Records for which no condition is true are routed to the default branch. A condition whose result is null counts as not true.
A condition that fails to evaluate is not “no match”. Under error_handling.strategy: continue the record is dead-lettered and takes no branch, not even one whose condition held, and not the default; under fail_fast the run stops. In exclusive mode the conditions after the first true one are never evaluated, so a condition further down cannot fail for that record. See An evaluation error is never false.
Routing modes
Exclusive (default)
In exclusive mode, conditions are evaluated in declaration order and the first matching condition wins. A record appears in exactly one branch. Order matters – put more specific conditions first.
mode: exclusive
conditions:
vip: "lifetime_value > 100000"
high: "lifetime_value > 10000"
medium: "lifetime_value > 1000"
default: standard
A customer with lifetime_value = 50000 is not over 100000, so vip is not true; high is the first true condition and wins. A customer with lifetime_value = 150000 is true for all three conditions, but goes to vip alone, because vip is checked first. Listing medium first would send both customers to medium.
Inclusive
In inclusive mode, all matching conditions route the record. A single record can appear in multiple branches simultaneously.
mode: inclusive
conditions:
needs_review: "amount > 10000"
flagged: "status == \"flagged\""
international: "country != \"US\""
default: standard
A flagged international order over 10000 would appear in needs_review, flagged, and international – three copies routed to three branches.
Downstream wiring
Downstream nodes reference route branches using port syntax: route_name.branch_name. The default branch is reached the same way, by its name. Several nodes can read the same branch; each receives every record on it.
A node whose own name equals a branch name can also reference the Route by its bare name and receives that branch: a Sink named high with input: classify reads classify.high.
- type: route
name: classify
input: transactions
config:
mode: exclusive
conditions:
high: "amount > 1000"
medium: "amount > 100"
default: low
- type: transform
name: high_value_processing
input: classify.high
config:
cxl: |
emit txn_id = txn_id
emit amount = amount
emit review_flag = true
- type: transform
name: standard_processing
input: classify.medium
config:
cxl: |
emit txn_id = txn_id
emit amount = amount
- type: sink
name: low_value_out
input: classify.low
config:
name: low_value_out
type: csv
path: "./output/low_value.csv"
Constraints
- Give every condition a distinct name; each name is a separate port.
- Give
defaulta name that no condition uses. This is not currently checked when the pipeline is planned: adefaultwith the same name as a condition is folded into that condition’s branch, so its records cannot be told apart from the matches. - Declare at least one condition. A Route with none is not rejected today; every record takes the default branch.
Complete example
pipeline:
name: order_routing
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
schema:
- { name: order_id, type: int }
- { name: region, type: string }
- { name: amount, type: float }
- { name: priority, type: string }
- type: route
name: by_region
input: orders
config:
mode: exclusive
conditions:
domestic: "region == \"US\" or region == \"CA\""
emea: "region == \"UK\" or region == \"DE\" or region == \"FR\""
apac: "region == \"JP\" or region == \"AU\" or region == \"SG\""
default: other
- type: sink
name: domestic_orders
input: by_region.domestic
config:
name: domestic_orders
type: csv
path: "./output/domestic.csv"
- type: sink
name: emea_orders
input: by_region.emea
config:
name: emea_orders
type: csv
path: "./output/emea.csv"
- type: sink
name: apac_orders
input: by_region.apac
config:
name: apac_orders
type: csv
path: "./output/apac.csv"
- type: sink
name: other_orders
input: by_region.other
config:
name: other_orders
type: csv
path: "./output/other_regions.csv"
Merge Nodes
Merge nodes concatenate multiple upstream branches into a single stream. They are the counterpart to route nodes – where a route splits one stream into many, a merge joins many streams back into one.
Merge is for streamwise concatenation of inputs that share a schema. For record-level joining across inputs that have different schemas, see Combine Nodes.
Interactive companion: the Route and Merge explainer shows how each mode mixes its inputs, and what an inclusive Route’s copies look like after a Merge.
Basic structure
- type: merge
name: combined
inputs:
- east_data
- west_data
config: {}
Note the key differences from other node types:
- Uses
inputs:(plural), notinput:(singular). - The
config:block is empty – all wiring is on the node header. - Using
input:(singular) on a merge node is a parse error.
Wiring
The inputs: field is a list of upstream node references. These can be bare node names or port references from route nodes:
- type: merge
name: rejoin
inputs:
- process_high
- process_medium
- classify.low # Port syntax for a route branch
config: {}
Downstream nodes wire to the merge as a normal single-input reference:
- type: sink
name: final_output
input: rejoin
config:
name: final_output
type: csv
path: "./output/combined.csv"
Modes
Merge’s cross-input ordering discipline is selected by config.mode. Two modes exist; concat is the default.
concat (default)
Predecessor records drain in declaration order: inputs[0] flows to output first, then inputs[1], then inputs[2], and so on. Within a single predecessor, its arrival order is preserved. Output is reproducible run-to-run for the same predecessor paths.
- type: merge
name: combined
inputs: [east, west]
config:
mode: concat
interleave
Records flow to output as they become available from any predecessor. Each input’s arrival order is preserved; cross-input order follows wall-clock arrival and is not promised.
- type: merge
name: combined
inputs: [east, west]
config:
mode: interleave
Seeded interleave — interleave_seed:
Snapshot tests and benchmarks that need reproducible cross-input ordering can opt into a deterministic schedule:
- type: merge
name: combined
inputs: [east, west]
config:
mode: interleave
interleave_seed: 42
With a seed, the cross-input order is reproducible from run to run regardless of upstream timing. To get there, the Merge reads all of its inputs into memory before emitting, so a seeded interleave buffers more than the other modes — use it for tests and benchmarks, not high-volume production merges.
Choosing a mode
| Mode | Order | When to use |
|---|---|---|
concat | Inputs emitted in declaration order, each fully drained before the next. | Downstream depends on a stable, declaration-ordered sequence (byte-identical output, contiguous time partitions). |
interleave (unseeded) | Records emitted as they arrive; per-input order preserved, cross-input order varies. | Lowest latency and the consumer is order-insensitive (e.g. an aggregator grouping by key, or a writer that doesn’t assert on row order). |
interleave (seeded) | Reproducible cross-input order. | Tests and benchmarks that assert on exact row sequence. Buffers all inputs in memory. |
For high-volume merges, prefer concat or unseeded interleave — both stream their inputs and let a slow downstream consumer naturally throttle the upstream readers, so memory stays bounded. A seeded interleave does not, because it buffers everything first.
Record ordering
Records arrive in the order described by the mode in use — see Modes
and Choosing a mode above. Merge does not promote matching
per-input sort_order declarations into one global sort: concatenating two
independently sorted files can still put a high key before a lower key at the
input boundary.
For concat and seeded interleave, exact sequence is a supported oracle for
the same input paths and configuration. For unseeded interleave, compare the
decoded record multiset and aggregate values; snapshotting incidental
cross-input arrival order would assert behavior Clinker does not promise. If
you need one sorted sequence regardless of Merge mode, declare sort_order on
the downstream Sink. That terminal sort uses only the
authored keys and does not invent a hidden identity tie-breaker.
Use cases
Reuniting route branches
The most common pattern is routing records through different processing paths and then merging them back together:
- type: route
name: classify
input: orders
config:
mode: exclusive
conditions:
high: "amount > 1000"
default: standard
- type: transform
name: process_high
input: classify.high
config:
cxl: |
emit order_id = order_id
emit amount = amount
emit surcharge = amount * 0.02
emit tier = "premium"
- type: transform
name: process_standard
input: classify.standard
config:
cxl: |
emit order_id = order_id
emit amount = amount
emit surcharge = 0
emit tier = "standard"
- type: merge
name: all_orders
inputs:
- process_high
- process_standard
config: {}
- type: sink
name: result
input: all_orders
config:
name: result
type: csv
path: "./output/all_orders.csv"
A Merge passes on every record from every input and does not remove duplicates. Branches of an inclusive Route can carry the same record, and rejoining them puts that record in the output once per branch it took. Use an exclusive Route when each record should come out once.
Unioning multiple sources
Merge nodes can combine records from multiple source files that share the same schema:
- type: source
name: jan_sales
config:
name: jan_sales
type: csv
path: "./data/sales_jan.csv"
schema:
- { name: sale_id, type: int }
- { name: amount, type: float }
- { name: region, type: string }
- type: source
name: feb_sales
config:
name: feb_sales
type: csv
path: "./data/sales_feb.csv"
schema:
- { name: sale_id, type: int }
- { name: amount, type: float }
- { name: region, type: string }
- type: merge
name: all_sales
inputs:
- jan_sales
- feb_sales
config: {}
- type: aggregate
name: totals
input: all_sales
config:
group_by: [region]
cxl: |
emit total = sum(amount)
emit count = count(*)
Schema constraints across inputs
Merge concatenates streams positionally against the merge node’s output_schema (taken from the first input). Every input must therefore agree on column shape — same column names, same on_unmapped policy, same correlation_key set.
Disagreement on the $widened auto_widen sidecar (one source uses auto_widen, another uses drop / reject) fails compile with E315. See Auto-Widen & Schema Drift → E315 for the full diagnostic shape and remediation.
Combine Nodes
Combine nodes are the N-ary record-combining operator. Every input is declared up front and bound to a qualifier; the where: expression matches records across inputs using qualified field references (e.g. orders.product_id == products.product_id); the cxl: body shapes the output row.
Combine is distinct from merge: merge concatenates upstream branches that share a schema, while combine joins records across inputs that have different schemas.
Interactive companion: the Combine playground shows, driver row by driver row, which build rows where: matches and what match: and on_miss: do with them, including range joins.
Basic structure
- type: combine
name: enrich
input:
orders: orders # qualifier: upstream node name
products: products
config:
where: "orders.product_id == products.product_id"
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_name = products.product_name
emit amount = orders.amount
propagate_ck: driver
Note the differences from other node types:
- Uses
input:as a map, binding qualifier names to upstream node references. Other nodes useinput:as a single string orinputs:as a list of strings. - Every field reference inside
where:andcxl:must be qualified (<qualifier>.<field>). Bare field names are a compile error. - Using
inputs:(plural list) on a combine node is a parse error.
Wiring
Each entry in the input: map binds a qualifier to an upstream node:
input:
orders: orders # qualifier "orders" -> source node "orders"
products: products
high_priority: classify.high # qualifier "high_priority" -> route port
Qualifiers are local names used inside where: and cxl:; they do not need to match the upstream node name. Upstream references can be bare node names or port references from a route node.
Iteration order in the input: map is preserved and used as the default driver-selection order (see Choosing the driving input below).
Configuration fields
| Field | Required | Default | Description |
|---|---|---|---|
where | Yes | – | CXL boolean expression matching records across inputs. Must contain at least one cross-input equality or range conjunct (a predicate with neither is rejected at plan time — see Predicate requirements). |
match | No | first | Match cardinality: first, all, or collect. |
on_miss | No | null_fields | Driver-record handling on zero predicate matches: null_fields, skip, or error. |
cxl | Yes (except under match: collect) | – | Emit statements defining the output row. Empty under match: collect. |
drive | No | first input | Explicit driver-input qualifier. Overrides the iteration-order default. |
strategy | No | auto | Execution strategy hint: auto or grace_hash. |
propagate_ck | Yes | – | Selects which correlation-key columns ride onto the output. driver keeps the driver’s CK only; all unions every input’s CK columns; { named: [<field>, ...] } carries an explicit subset. See Correlation-key propagation below. |
max_output_rows | No | unlimited | Opt-in cap on the number of rows this combine may emit. When set, the run fails loud (diagnostic E325) the moment the output would exceed the cap — it never truncates to a partial result. See Output-size cap. |
The where: predicate
The where: expression is a CXL boolean expression evaluated for every candidate record pair across inputs. It must contain at least one cross-input equality – an equality with field references from two different inputs:
where: "orders.product_id == products.product_id"
Compound predicates combine multiple conjuncts with and. Each conjunct is classified by the planner:
- Equi conjunct – a cross-input equality (
a.x == b.y). Drives the hash lookup or sort-merge join. - Range conjunct – a cross-input ordered comparison (
a.start <= b.ts and b.ts <= a.end). Handled by a range join (IEJoin), whether or not an equi conjunct also links the same two inputs. - Residual conjunct – any other CXL predicate (intra-input filter, function call, etc.). Applied as a post-filter after the equi/range match.
where: |
orders.product_id == products.product_id
and orders.amount >= 100
and products.region == "us-east"
Above: the equi conjunct drives the join; orders.amount >= 100 and products.region == "us-east" are applied as residuals.
Every combine predicate must carry at least one cross-input equality or range conjunct. A predicate with neither — a pure residual with no decomposable cross-input comparison — is rejected at plan time with diagnostic E313; there is no supported execution strategy for it. Pure-range predicates without an equi conjunct are fully supported via IEJoin.
Non-orderable range keys
A range conjunct compares values that must be orderable at runtime: integers, finite floats, exact decimals, dates, and datetimes. When a record’s range key evaluates to a non-orderable value — SQL NULL, a non-finite float (NaN/infinity), or any other type — that record can never satisfy the range comparison, so it is routed out of the range match rather than joined:
Decimal and mixed-numeric range keys. Exact fixed-point
decimalrange keys are fully supported on every join strategy — a monetary band join (amount >= tier.floor and amount < tier.ceiling) matches correctly. A range conjunct that mixes anintegerand afloatoperand across the two inputs is also supported: the integer is compared as a float, exactly as the>=/<operators compare it elsewhere.Decimal range keys are placed on a shared fixed-point grid with up to 18 fractional digits and an integer magnitude up to roughly 1.7 × 10²⁰ — well beyond any realistic monetary value. A decimal range value outside that grid (more than 18 fractional digits, which would truncate, or a magnitude that would overflow the grid) is not silently dropped: the run stops with diagnostic
E326naming the combine, so a wrong or empty result is never emitted. Rescale or narrow the compared values if you hit it.Datetime range keys.
dateanddatetimerange keys compare at their native resolution — adatetimeto the nanosecond, across the full representable calendar (no microsecond rounding; instants before 1677 or after 2262 stay exact). Two timestamps that differ only below the microsecond therefore match, sort, and group as distinct instants on every join strategy, so a sub-microsecond as-of or band lookup neither drops a boundary match nor merges two near-simultaneous events.
abs/min/max/clampand rejected range keys (E327).abs,min,max, andclampreturn thenumericsupertype (int | float). When both operands of the range conjunct recover the same concrete type —abs(int) >= abs(int), or amin/max/clampwhose result can only be one type (all-intor all-float, recovered through nested calls and arithmetic such asabs(a.x + 1)orabs(min(a.i, b.i))) — the axis is exactly as safe as a plain matching-typed key and the join runs normally. Otherwise the conjunct cannot be reduced to one exact numeric axis and is rejected at plan time withE327, rather than routed to the join where it could silently drop rows or mismatch.E327fires when such an operand stays genuinely ambiguous — a mixed pair likeabs(int) >= abs(float), or a per-row-ambiguous result likemin(int, float)that may return either the integer or the float operand — or on a non-orderable pairing such as a string comparison or adecimalcompared against afloat/numeric. Compare a matching-typed range key instead. (A supported int/float/decimal/date/datetime range key, including a mixed int/float field pair, never triggersE327.)
- A driver record with a non-orderable range key is treated as a zero-match driver and handled by
on_miss(null_fields/skip/error). - A build record with a non-orderable range key is dropped (it can match nothing).
This is a runtime routing decision on the data, distinct from the plan-time E313 rejection above (which is about the predicate shape, not the values).
Match modes
match: first
Emit one output row per driver record, using the first matching build-side record. Standard 1:1 enrichment. Default.
“First” means the earliest in the build input’s arrival order: of the build records that match a driver, the one the build input delivered first. The same order sets the row order of match: all within a driver and the element order of a match: collect array. It holds identically for every join strategy the planner may pick and whatever shape the where: predicate has. One exception remains. When the build input exceeds the memory limit, the Combine sets part of it aside on disk and splits that part into smaller pieces until each fits. If a piece still does not fit after splitting, the Combine processes it in chunks and currently decides first, collect and on_miss once per chunk rather than once per driver, so a driver can receive more than one first row or more than one collect row. A piece fails to fit in three cases: the build rows sharing one key value do not fit on their own; the build input has already been split into the maximum of 4,096 pieces; or splitting a piece again leaves most of its rows together, as happens when many rows share a key value or their key values happen to land in the same piece. So the exception can reach a driver even when no single key value is over the limit, for example with a large build input under a tight memory limit. This is a known defect.
- To pick a different record, order the build input upstream. For example, to enrich each driver with the latest price, deliver the build input sorted descending on its date, such as with a descending
sort_orderon the build Source. A Combine has no ordering option of its own. - A correlation key orders its Source. Declaring
correlation_keyon the build Source makes the planner sort that Source’s rows by the key before they reach the Combine, unless its declaredsort_orderalready begins with the key. That order is what “first” follows. See Correlation keys. - An unseeded
interleaveMerge has no fixed order. When the build input is such a Merge, its cross-input order follows arrival at run time, so “first” follows that unfixed order too. Useconcat, or a seededinterleave, for a build input whose order must be the same on every run.
config:
where: "orders.product_id == products.product_id"
match: first
cxl: |
emit order_id = orders.order_id
emit product_name = products.product_name
The where: predicate selects the match; the cxl: body is a post-match projection that runs once on the chosen build record. Selection and projection are separate steps: if the body filters the row out (a filter that fails, or a body that emits nothing), that one output row is dropped. The combine does not fall back to a later matching build, and the driver is not treated as unmatched — it matched the predicate, the body just produced no row. on_miss (below) never fires for such a driver; it fires only when the predicate matched nothing at all. This holds identically for every join strategy the planner may pick.
Behavior change. This is a change in observable output for existing pipelines whose
where:predicate carries a range, equi+range, or single-inequality comparison — the shapes the planner runs as a sort-merge join or an IEJoin (both the pure-range block-band path and the equi+range hash-partitioned path). On any of those three strategies, a driver that matched the predicate but whose body skipped every candidate was previously routed toon_miss— firingnull_fields,skip, orerror. It now silently produces no row, matching the pure-equality strategies (in-memory hash and grace hash), which already behaved this way and are unchanged. A pipeline that relied on the old routing (for example,on_miss: errortripping on a body-skipped driver) no longer sees it.
When where: fails to evaluate. A where: that raises an error for a candidate, such as a division by zero, has not said whether that candidate matches. It is neither a match nor a non-match. Under match: first the earliest candidate that is not a non-match decides. If it matched, the driver is enriched with it. If its where: failed, the driver is dead-lettered with that build row and writes no output row, even when a later candidate would match: taking the later one would publish a row that depends on an evaluation that failed. A candidate after the deciding one is not part of the driver’s result, so its failure is never written. This holds on every join strategy.
match: all
Emit one output row for every matching build-side record. 1:N fan-out – if a driver record matches three build records, three rows are emitted, in the build input’s arrival order (see match: first).
config:
where: "employees.department == benefits.department"
match: all
cxl: |
emit employee_id = employees.employee_id
emit benefit = benefits.benefit_name
Each output row depends only on its own pair, so a candidate whose where: fails to evaluate is dead-lettered with its build row while the driver’s other matches are still emitted.
match: collect
Gather every matching build-side record into a single Array-typed field on the output row. The driver record appears once; the build matches are aggregated into an array, in the build input’s arrival order (see match: first). The cxl: body must be empty under collect – the combine node synthesizes the output as { driver fields..., <build_qualifier>: Array }.
config:
where: "orders.product_id == products.product_id"
match: collect
cxl: ""
A per-group entry limit of 10,000 prevents unbounded growth.
A collect row states the complete set of a driver’s matches. If any candidate’s where: fails to evaluate, that set is unknown, so the driver writes no row, neither a partial array nor an empty one; each failing candidate is dead-lettered with its build row. Candidates past the 10,000-entry limit are still checked, so every failure among them is dead-lettered too.
Use collect when you need the set of matches as a single structured value; use all when you need a flat row per match.
Unmatched records (on_miss)
on_miss controls what happens to driver records with zero predicate matches — drivers for which no build-side record satisfied where:, because every candidate evaluated it to false or null, or because there was no candidate at all. Two kinds of driver are not misses and never reach on_miss, under any match mode and on every join strategy:
- A driver that matched the predicate but whose
cxl:body skipped or failed the row (seematch: first). It simply produces no output row for that match, and a body failure is dead-lettered. - A driver any of whose candidates failed to evaluate
where:. Each failure is dead-lettered, and whateveron_misssays, it does not fire:on_miss: errordoes not stop the run andon_miss: null_fieldswrites no null-filled row, because the driver was never shown to have no match.
| Value | Semantics |
|---|---|
null_fields (default) | Build-side fields resolve to null. Driver record is still emitted. Equivalent to left-join. |
skip | Driver record is dropped. Equivalent to inner-join. |
error | Pipeline fails on the first unmatched driver record. |
config:
where: "orders.product_id == products.product_id"
on_miss: skip
on_miss: error is useful for strict referential integrity where any miss should halt processing. on_miss: skip is the inner-join shape. on_miss: null_fields is the left-join shape and the default.
Composite keys
Chain multiple cross-input equalities with and:
config:
where: |
sales.department == targets.department
and sales.region == targets.region
cxl: |
emit department = sales.department
emit region = sales.region
emit actual = sales.amount
emit goal = targets.goal
All conjuncts must hold for a record pair to match.
Multi-input combine (three or more)
Combine accepts any number of inputs. Each pair of inputs that should be related needs an explicit cross-input equality:
- type: combine
name: fully_enriched
input:
orders: orders
products: products
categories: categories
config:
where: |
orders.product_id == products.product_id
and products.category_id == categories.category_id
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_name = products.product_name
emit category_name = categories.name
emit amount = orders.amount
propagate_ck: driver
The planner builds a join tree by walking equalities pairwise: starting from the driver, it adds one input at a time, always one linked by an equality to the inputs already joined. It is designed to prefer the input with the fewest rows at each step, but those row counts are not available yet when it runs, so in practice it takes the linked inputs in the order they are declared.
Choosing the driving input
The driver is the input whose records flow through one at a time during execution; the other inputs are materialized as build-side hash tables (or IEJoin index structures). By default the first input in the input: map is the driver.
Use drive: to override:
config:
where: "orders.product_id == products.product_id"
drive: products
cxl: |
emit product_id = products.product_id
emit product_name = products.product_name
emit sample_order_id = orders.order_id
With drive: products, the pipeline emits one row per product enriched with a matching order, instead of one row per order enriched with its product. Pick the driver based on which side you want to iterate over (typically the larger stream, or the one whose ordering you want to preserve).
Strategy hint
| Value | Behavior |
|---|---|
auto (default) | Planner picks a strategy from the predicate shape. Equalities only: an in-memory hash join, or grace hash when the build side’s estimated size is close to the memory limit. Equalities plus ranges: a range join (IEJoin) that also groups by the equal values. Ranges only: a sort-merge join when there is a single comparison between two int, float, date or datetime fields of the same type (not decimal) and both inputs already arrive sorted ascending on those fields, each from a single sorted file or stream; otherwise a range join (IEJoin). Both give the same result. |
grace_hash | Force grace hash join (disk-spilling partitioned hash). Applies only to pure-equi predicates; ignored on predicates with range conjuncts. |
The choice is made when the pipeline is planned, from an estimate of the inputs’ size. Under auto, an equal-ids join runs in memory unless that estimate says the build side is too large. If the estimate is low or missing (for example, a glob: source whose files are not known in advance) and the build side turns out larger than the memory budget, the run stops with E310 MemoryBudgetExceeded; it does not switch to disk partway through (#1337). Set strategy: grace_hash when the build side may be larger than the memory budget: it partitions the build side to disk from the start.
Correlation-key propagation
Combine declares which correlation-key columns its output rows carry via the required propagate_ck field.
- type: combine
name: enriched
input:
orders: orders
products: products
config:
where: "orders.product_id == products.product_id"
cxl: |
emit order_id = orders.order_id
emit product_name = products.name
propagate_ck: driver # driver-only (today's behavior)
propagate_ck: all # union of every input's $ck.* columns
propagate_ck:
named: [order_id] # explicit subset (intersected with upstream)
driver– output carries only the driver input’s correlation-key columns. Build-side records contribute body fields, but their group identity is consumed by the match; when a match fails, the build record’s dead letter follows the failing driver’s correlation group.all– output carries every input’s correlation-key columns. Use when the build side carries keys that downstream operators need to read.named: [<field>, ...]– an explicit subset. Use to project a multi-field key down to a single field after a join.
Driver wins on a name collision: if both the driver and a build input declare the same key field, the output keeps the driver’s value. See the Correlation-key combine interaction reference for how each match mode fills the propagated key (especially match: collect).
propagate_ck is required on every combine; pipelines without an explicit value fail to compile. Existing pipelines migrate by adding propagate_ck: driver, which is bit-for-bit equivalent to today’s behavior.
Output-size cap (max_output_rows)
max_output_rows is an opt-in ceiling on how many rows a combine may emit. It defaults to unlimited; set it to guard against a permissive or mis-specified predicate that would explode a small pair of inputs into a huge result (for example, a range join over an unexpectedly hot key producing a near cross product):
config:
where: "orders.ts >= prices.effective_from"
cxl: |
emit order_id = orders.order_id
emit price = prices.amount
propagate_ck: driver
match: all
max_output_rows: 1000000
Semantics:
- Fail-loud, never truncate. The moment the combine would emit more than the cap, the run stops with diagnostic
E325naming the node and the cap. It does not produce a capped or partial result — a silently truncated join would corrupt downstream data. - Independent of the memory budget. This is a result-size guard, not memory pressure. A runaway join can be perfectly bounded in memory (its output spills to disk) yet still produce far more rows than intended;
max_output_rowscaps the row count regardless of bytes. - Covers the whole output, on every strategy. The cap counts every emitted output row across all match modes and any
on_missrows, and is enforced identically whichever join strategy the planner picks (hash build-probe, grace-hash, sort-merge, or the IEJoin block-band). collectcounts driver rows. Undermatch: collecta combine emits one output row per driver row (each carrying an array of up to 10 000 collected matches), somax_output_rowsbounds the driver-row count, not the number of collected array elements.- Dead-lettered rows are not counted. A matched pair whose
cxl:body or residual raises a recoverable eval failure is routed to the DLQ, not the output, so it does not count toward the cap. Only that pair is dead-lettered: undermatch: allthe driver’s other matches are still evaluated and emitted, whichever join strategy runs. A failing residual is neither a match nor a miss: undermatch: firsta failure on the deciding candidate is the driver’s only result, and undermatch: collectit leaves the driver with no row (seematch: firstandmatch: collect). The driver row and the matched build row are dead-lettered, each with its own Source and row number; each failure writes its own pair, so a driver that fails against several build rows is dead-lettered once per failure, and a build row that several failing drivers matched is dead-lettered once for each of them. - N-ary combines cap the final output. A combine whose
where:spans three or more inputs is decomposed into a chain of binary steps;max_output_rowsguards the final combined output, not the intermediate chain steps.
If the large result is expected, raise the cap (or omit the field). If it is not, tighten the where: predicate. Run clinker explain --code E325 for the full remediation guide.
Memory considerations
Each non-driving (build-side) input is held in memory while the join runs, so plan for roughly 1.5–2× its file size — a 50 MB lookup table needs about 75–100 MB. An equal-ids join spills its build side to disk only when it runs as a grace hash join, which the planner chooses from its size estimate or which strategy: grace_hash forces. An in-memory join whose build side outgrows the budget stops the run with E310 instead of spilling (#1337). Set the budget with pipeline.memory.limit; see Memory Tuning.
Range and equi+range predicates (the IEJoin block-band strategy) are bounded on both input axes and the output: each side is drained into disk-backed, key-sorted blocks, and the emitted rows accumulate in a spillable sort buffer rather than a resident vector. So a range join whose inputs — or whose result — exceed the memory budget spills automatically and completes, rather than failing. Even a single hot key whose block-pair is a near cross product streams through a bounded nested loop instead of materializing the whole candidate set. Use max_output_rows if you want such a runaway result to stop rather than spill.
Document boundaries
A Combine passes document boundaries through to its output, so a per-document Aggregate after a join still rolls up per document. A driver source that carries several documents (a glob: over monthly files, say) produces one roll-up per driver document after the join, not one fold spanning all of them. A document carried on both the driver and the build side opens and closes exactly once downstream. See Document Context & Envelopes for the per-document aggregation model.
Complete example
pipeline:
name: order_enrichment
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
schema:
- { name: order_id, type: string }
- { name: product_id, type: string }
- { name: amount, type: float }
- type: source
name: products
config:
name: products
type: csv
path: "./data/products.csv"
schema:
- { name: product_id, type: string }
- { name: product_name, type: string }
- { name: category, type: string }
- type: combine
name: enrich
input:
orders: orders
products: products
config:
where: "orders.product_id == products.product_id"
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_id = orders.product_id
emit product_name = products.product_name
emit category = products.category
emit amount = orders.amount
propagate_ck: driver
- type: sink
name: result
input: enrich
config:
name: result
type: csv
path: "./output/enriched_orders.csv"
See also
- Multi-Input Combine – recipe-style walkthrough with input data and expected output.
- Merge Nodes – streamwise concatenation; the right operator when inputs share a schema and no per-record matching is needed.
- Memory Tuning – memory budget, spill thresholds, and strategy overrides.
Aggregate Nodes
Aggregate nodes group records by one or more fields and compute summary values using CXL aggregate functions. They consume all input records in a group before emitting a single summary record per group.
Basic structure
- type: aggregate
name: dept_totals
input: employees
config:
group_by: [department]
cxl: |
emit total_salary = sum(salary)
emit headcount = count(*)
emit avg_salary = avg(salary)
Group-by fields pass through automatically – you do not need to emit them. In this example, the output records contain department, total_salary, headcount, and avg_salary.
Group-by fields
The group_by: field is a list of field names from the input schema. Records sharing the same values for all group-by fields are placed in the same group.
group_by: [region, department]
cxl: |
emit total_salary = sum(salary)
emit max_salary = max(salary)
This produces one output record per unique (region, department) combination.
Values are the same group key when they are equal under the rule sorting uses (see How values are ordered):
- Numbers group by exact value, whatever their type:
1,1.0and the decimal1are one group, and distinct integers are never merged, however large. -0.0and0.0are one group.- Every
NaN, of either sign, is one group, separate from the group of null values. - A group reports the value of its first-arriving row. An integer column is
written as integers; a column holding both
5and5.0writes5if the integer arrived first and5.0if the float did. The value is the same whether or not the Aggregate spilled to disk.
The same rule groups Cull and Reshape partition_by, a window’s group_by,
correlation keys, distinct and output splitting.
Global aggregation
An empty group_by list treats the entire input as a single group, producing exactly one output record:
- type: aggregate
name: grand_totals
input: orders
config:
group_by: []
cxl: |
emit grand_total = sum(amount)
emit record_count = count(*)
emit avg_order = avg(amount)
Aggregate functions
The following aggregate functions are available in CXL:
| Function | Description |
|---|---|
sum(field) | Sum of all values in the group |
count(*) | Number of records in the group |
avg(field) | Arithmetic mean |
min(field) | Minimum value |
max(field) | Maximum value |
collect(field) | Collect all values into an array |
weighted_avg(value, weight) | Weighted average |
Strategy hint
The strategy: field controls how aggregation is executed:
- type: aggregate
name: totals
input: sorted_data
config:
group_by: [account_id]
strategy: streaming
cxl: |
emit total = sum(amount)
| Strategy | Behavior |
|---|---|
auto | Default. The optimizer chooses based on whether the input is provably sorted for the group-by keys. |
hash | Force hash aggregation. Works on any input ordering. Holds all groups in memory (with disk spill if memory budget is exceeded). |
streaming | Require streaming aggregation. Processes one group at a time with O(1) memory per group. Compile-time error if the input is not provably sorted for the group-by keys. |
When to use streaming
The optimizer chooses streaming aggregation automatically when the Aggregate’s input is still sorted and the first fields of that order are exactly the group_by fields, in any order: take as many fields from the front of the order as there are group_by fields, and they must be the same set. Order [department, day] streams group_by: [department] and group_by: [day, department], but not group_by: [day]. An Aggregate with no group_by always streams. Use strategy: streaming as an explicit assertion – it turns a silent fallback to hash aggregation into a compile error (CXL0419), which is useful for catching sort-order regressions.
The order has to survive every stage between the Source and the Aggregate. A Merge, a Combine, a Reshape or Cull, distinct, and a Transform that writes one of the sort fields all drop it. Writing a field back unchanged counts: emit department = department drops an order on department. A Transform carries every input field through, so there is no need to emit sort fields. The sort order explainer lets you build a chain and see where the order holds.
When to use hash
Hash aggregation works on unsorted input and is the safe default. It uses more memory but handles any data ordering. Memory-aware disk spill kicks in when RSS approaches the pipeline’s memory.limit.
Correlation-key interaction
When a pipeline’s sources declare correlation_key: fields, an aggregate behaves one of two ways depending on its group_by:
group_bycovers the correlation key — if any record in a group fails, the whole group (including the aggregate output row) is sent to the DLQ.group_byomits a correlation-key field — only the failing records are dropped and the affected groups are recomputed, so the surviving rows still produce a correct aggregate.
You do not configure this; the engine picks the behavior from your group_by. The one restriction is that the second case cannot be combined with strategy: streaming. See Correlation Keys for the full rules.
Time-windowed aggregates
When time_window: is set on the aggregate body, records are grouped
not just by group_by but also by event-time window. Each record is
placed into one or more windows by its event time (the
$source.event_time value derived from
the source’s watermark), and each window emits one row per group once
it closes.
Every upstream-reachable source must declare a
watermark: so the engine knows when a window
is complete. Without one, no window ever closes and the planner rejects
the pipeline with
E156.
The engine emits user-declared columns only — window bounds do not
appear in the output unless you compute and emit them yourself. The
emit order is ascending window_start (deterministic), so output
rows naturally group by window.
Tumbling windows
Non-overlapping fixed-size buckets. Each record lands in exactly one
window [floor(t / size) * size, floor(t / size) * size + size).
time_window:
tumbling: { size: 1h }
Input (tumbling_demo.csv):
user_id,event_ts,kind
u1,2026-05-14T10:05:00,click
u2,2026-05-14T10:30:00,click
u1,2026-05-14T10:42:00,click
u1,2026-05-14T11:03:00,click
u2,2026-05-14T11:15:00,click
u2,2026-05-14T11:50:00,click
Output with tumbling: { size: 1h }, group_by: [user_id],
emit n = count(*):
user_id,n
u1,2
u2,1
u1,1
u2,2
Reading top-to-bottom: the first two rows are the [10:00, 11:00)
bucket (u1’s 10:05 and 10:42, then u2’s 10:30); the next two are
the [11:00, 12:00) bucket (u1’s 11:03, then u2’s 11:15 and 11:50).
Each input record contributes to exactly one window.
Hopping windows
Overlapping fixed-size buckets advanced by slide. Each record
lands in ceil(size / slide) windows: slide < size produces
overlap, slide == size degenerates to tumbling, slide > size
produces gaps where some records fall in zero windows.
time_window:
hopping: { size: 1h, slide: 30m }
Input (hopping_demo.csv):
user_id,event_ts,amount
u1,2026-05-14T10:05:00,10
u1,2026-05-14T10:42:00,20
u1,2026-05-14T11:10:00,15
Output with group_by: [user_id], emit total = sum(amount),
emit n = count(*):
user_id,total,n
u1,10,1
u1,30,2
u1,35,2
u1,15,1
Three input records, four output rows — each record fans into two
overlapping size: 1h, slide: 30m windows:
[09:30, 10:30)— just 10:05 →total=10, n=1[10:00, 11:00)— 10:05 + 10:42 →total=30, n=2[10:30, 11:30)— 10:42 + 11:10 →total=35, n=2[11:00, 12:00)— just 11:10 →total=15, n=1
Session windows
Per-key gap-bounded sessions. A new record extends its key’s current
session if its event time is within gap of the session’s last
event time; otherwise it starts a new session. The boundary is
data-driven, not clock-aligned.
time_window:
session: { gap: 10m }
Input (session_demo.csv):
user_id,event_ts,action
u1,2026-05-14T10:00:00,login
u1,2026-05-14T10:07:00,click
u1,2026-05-14T10:13:00,click
u1,2026-05-14T10:50:00,login
u1,2026-05-14T10:55:00,click
Output with group_by: [user_id], emit n = count(*):
user_id,n
u1,3
u1,2
u1’s first three rows form one session (10:00 → 10:07 → 10:13,
consecutive gaps ≤ 10m). The 37-minute idle stretch exceeds gap,
so 10:50 starts a fresh session that runs through 10:55. Two
sessions, two output rows.
Allowed lateness
allowed_lateness extends how long a window stays open past its end
before it emits, giving late-arriving records a grace period to still
be counted. It is distinct from the source-side watermark.delay.
Records that arrive after a window’s end + allowed_lateness route to
the DLQ as LateRecord with stage label time_window:<aggregate-name>.
See DLQ category: LateRecord
for the DLQ row layout.
- type: aggregate
name: hourly
input: clicks
config:
group_by: [user_id]
time_window:
tumbling: { size: 1h }
allowed_lateness: 30s
cxl: |
emit n = count(*)
Default (unset) means no grace beyond the watermark — windows close
the instant min_across_sources crosses window_end. Set
allowed_lateness when the source’s watermark.delay alone is too
small to absorb the observed out-of-order tail.
Worked example: multi-source session window
This pipeline merges two independent login feeds and groups per-user events into gap-bounded sessions. When several sources feed one time-windowed aggregate, a window cannot close until every source has advanced past the window’s end — the slowest source paces the others, so no window emits before all of its records have arrived.
pipeline:
name: multi_source_session
nodes:
- type: source
name: src_web
description: Web login events.
config:
name: src_web
type: csv
path: ./data/session_logins.csv
options:
has_header: true
watermark:
column: event_ts
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: source, type: string }
- type: source
name: src_mobile
description: Mobile login events.
config:
name: src_mobile
type: csv
path: ./data/session_mobile.csv
options:
has_header: true
watermark:
column: event_ts
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: source, type: string }
- type: merge
name: all_logins
inputs: [src_web, src_mobile]
- type: aggregate
name: user_sessions
input: all_logins
config:
group_by: [user_id]
time_window:
session: { gap: 5m }
allowed_lateness: 30s
cxl: |
emit user_id = user_id
emit logins = count(*)
- type: sink
name: results
input: user_sessions
config:
name: results
type: csv
path: ./output/multi_source_session.csv
Both sources declare their own watermark.column independently, and
each source’s records carry an event time regardless of which feed
delivered them — so the aggregate does not care which column name a
given source used. A session cannot emit until both src_web and
src_mobile have advanced past the session’s end plus its
allowed_lateness. Drop the watermark: block on either source and
the planner rejects the pipeline with
E156.
Run it from the repo:
cargo run -p clinker -- run examples/pipelines/multi_source_session.yaml
Complete example
- type: source
name: transactions
config:
name: transactions
type: csv
path: "./data/transactions.csv"
schema:
- { name: account_id, type: string }
- { name: txn_date, type: date }
- { name: amount, type: float }
- { name: category, type: string }
sort_order:
- { field: "account_id", order: asc }
- type: aggregate
name: account_summary
input: transactions
config:
group_by: [account_id]
strategy: streaming
cxl: |
emit total_amount = sum(amount)
emit txn_count = count(*)
emit avg_amount = avg(amount)
emit max_amount = max(amount)
emit categories = collect(category)
- type: sink
name: summary_output
input: account_summary
config:
name: summary_output
type: json
path: "./output/account_summary.json"
The output is JSON because collect(category) emits an array. JSON serializes
it as a native array; XML can instead write it as repeated child elements, and
CSV can join it into a delimited cell. For a scalar-only output such as
fixed-width, coerce it first with a downstream Transform (for example
emit categories = categories.join(";")).
Reshape Nodes
Reshape nodes observe a whole correlation group and, per group, mutate the rows whose state caused a rule to fire while synthesizing new rows derived from those trigger rows. They are the node for “look at everything an entity did, then fix one record and insert the record that should have been there” — work no other node can do:
- Aggregate reduces a group to one summary row.
- Transform emits 0 or 1 record per input record.
- Combine joins records across sources.
None of those produce new records derived from a group’s observed state while preserving the originals. Reshape does.
Reshape is a blocking grouping operator: it buffers every record of a group before any output row leaves, because a rule cannot decide what to synthesize until it has seen the whole group. It has a single output.
Basic structure
- type: reshape
name: backfill_plans
input: plans
config:
partition_by: [employee_id]
order_by:
- { field: plan_start, order: asc }
rules:
- name: fix_long_plan_years
when: "plan_start - plan_end > 365"
mutate:
set:
plan_end: "plan_start"
synthesize:
copy_from: trigger
overrides:
status: "'synthesized'"
For each employee_id group, every row where plan_start - plan_end > 365 (the trigger) has its plan_end rewritten and a new row synthesized from it with status overridden to synthesized.
partition_by
A list of field names. Records sharing the same values for all partition_by fields form one group, and every rule observes and acts within a single group. This is the correlation key the operator reasons over.
Whole-input grouping (partition_by: [])
An empty list is the degenerate case of that key: every record shares it, so the entire input forms one group and the rules apply across the whole dataset rather than per entity.
- type: reshape
name: relabel
input: rows
config:
partition_by: [] # one group: the whole input
order_by:
- { field: amount, order: asc }
rules:
- name: flag_large
when: "amount > 100"
mutate:
set:
label: "'large'"
order_by then sorts the whole input as a unit, and the no-cascade contract applies across every record at once. Two consequences follow from there being only one group:
- The whole input must fit
memory.limitat finalize, since a single group has to be resident when its rules fire (see Limits). Whole-input grouping is for datasets that fit the budget, not for bulk row-at-a-time work — a per-record mutation with no cross-record dependency belongs in a Transform, which streams. - A mutation conflict rolls back the whole group, so one conflict rolls back the entire run’s records rather than one entity’s.
Values Reshape cannot key
Reshape groups a record under a single null group whenever it cannot build a key from the partition value. Today that covers all of:
- the column being absent from the record
- an explicit null
- an empty string (
"") - an array- or map-valued cell (a multi-value column)
A NaN float, of either sign, is one group of its own, not part of the null group. Numbers group by exact value, as in Aggregate: 1, 1.0 and the decimal 1 are one group, and distinct integers never merge, however large.
So account="" and account=null land in the same Reshape group. Note that Cull does not fold empty strings — there, account="" and account=null are two groups. The difference is unintentional and tracked in #1022; until it is resolved, do not assume one node’s grouping matches the other’s for blank values.
If a blank-heavy column is your partition key, expect one large null group. Normalize blanks upstream (a Transform that maps "" to a real sentinel) when you want them grouped separately.
order_by
Optional. A list of sort fields applied within each group before rules run, so order-dependent synthesis is deterministic. Arrival order breaks ties.
Each entry is either a field name, which sorts ascending, or a map { field, order, null_order }:
| Key | Values | Default |
|---|---|---|
field | a column of the input | required |
order | asc or desc | asc |
null_order | first or last | last |
order_by:
- plan_start # same as { field: plan_start }
- { field: plan_end, order: desc, null_order: first }
Each group is ordered this way before its rules run, exactly as a Sink sort_order orders rows (see How values are ordered): a null goes where null_order puts it, last by default for asc and desc alike; numbers compare by their exact value across integers, floats and decimals; every NaN is one value after every number, and so comes first under desc; -0.0 equals 0.0; and rows the order calls equal keep their arrival order.
null_order: drop is rejected when the pipeline is planned. order_by only arranges the rows of a group; it never removes any. To leave out rows whose field is null, filter them in a Transform before the Reshape:
- type: transform
name: started_only
input: plans
config:
cxl: |
filter not plan_start.is_null()
Rules
Each entry in rules: is a declarative rule with a name, a trigger predicate, and optional mutation and synthesis actions. Rules are evaluated in declaration order — but every rule observes the same original group snapshot (see No cascade below).
when — the trigger predicate
when is a CXL boolean expression evaluated against each row in the group. A row for which when is true is a trigger row for that rule: its mutate rewrites it, and its synthesize derives new rows from it. CXL boolean operators are and / or / not (Clinker’s expression language is not SQL).
mutate — in-place trigger-row mutation
mutate:
set:
plan_end: "plan_start"
note: "concat(note, ' (corrected)')"
Each set: entry is field: <CXL expression>. The expression evaluates against the original trigger row and overwrites that field’s value on the row.
Two restrictions are enforced at compile time:
- A
set:target must already exist in the upstream schema. Reshape mutates existing columns; it does not add new ones. Emit the column from an upstream Transform first if you need it. - A
set:may not write apartition_byfield — group identity must survive Reshape.
synthesize — deriving new rows
synthesize:
copy_from: trigger
overrides:
plan_date: "'2024-01-01'"
status: "'synthesized'"
For each trigger row, synthesize emits one new row:
copy_from: trigger— the new row starts as a copy of the trigger row’s values, thenoverridesare applied on top.copy_from: none— the new row’s columns start null;overridesmust supply every column (enforced at compile time, so a synthesized row is never silently empty).
Each overrides: entry is field: <CXL expression>, evaluated against the trigger row.
Whichever copy_from a rule uses, a synthesized row keeps its trigger row’s $source.* identity — $source.name, $source.file, and $source.event_time. Downstream, $source.name on a synthesized row names the Source its trigger came from, and if a later stage dead-letters the row, the entry is attributed to that Source: it counts toward that Source’s per_source rate limit and is written to that Source’s per_source dead-letter file when one is configured. See Error Handling & DLQ.
No cascade
Every rule observes the original group state. A row mutated by rule A is not re-observed by rule B, and rule B’s when predicate sees the row’s original values, not rule A’s edits. This is a deliberate guarantee:
- Determinism — cascade would make rule order silently change output.
- Single observation — the group is observed once.
To sequence dependent transformations, chain two Reshape nodes in the DAG so the second observes the first’s output.
Mutation conflicts
If two rules write the same field on the same row, that is a mutation conflict. Some conflicts are caught at compile time when the rules’ selectors statically overlap; content-dependent collisions that cannot be proven at compile time are caught at runtime.
A runtime conflict routes a dead-letter-queue entry under the mutation_conflict category, and the whole correlation group rolls back — none of that group’s mutated or synthesized rows reach the output. The DLQ entry’s stage label is reshape:<node>:<rule_a>+<rule_b>, naming the colliding rule pair. See Error Handling & DLQ.
Known issue: today only the row where the conflict happened gets a DLQ entry. The group’s other rows are dropped without one, so they appear in neither the output nor the DLQ (#1349).
Audit stamps
Reshape stamps three engine-written columns on its output records so the provenance of a synthesized or mutated row is queryable downstream:
| Column | Meaning |
|---|---|
$meta.synthetic | true on a synthesized row, false on an original (including a mutated trigger row) |
$meta.synthesized_by | the <node>:<rule> that synthesized the row (empty on originals) |
$meta.mutated_by | comma-separated <node>:<rule> labels of every rule that mutated the row (empty if none) |
Like the $ck.* correlation columns, these $meta.* columns stay out of the default writer output — they are available for downstream CXL and audit, not silently dumped into your output files.
Memory model
Reshape is a blocking, grouped operator: it groups every input record by partition_by before any rule fires, because each rule must observe its whole correlation group (the no-cascade contract forbids folding a group incrementally). It therefore cannot stream — the full group set materializes before the first output row leaves.
That per-group buffer is governed by the same central memory arbitrator every other blocking operator polls (see Memory & Spill). As records are grouped, Reshape tracks the live in-memory footprint and, whenever the run crosses the soft spill threshold (80% of memory.limit by default), it spills buffered groups to disk:
- What spills: the raw input records, never the post-processed output rows. On reload at finalize, mutation and synthesis re-run in memory exactly as they would have without spilling, so the output is identical whether a group stayed resident or round-tripped through disk — including its within-group row order, which is restored to arrival order after a reload. (Two caveats apply; see Limits below.) Spilling input records (rather than output rows) is also what keeps a
copy_from: nonesynthesized row — built against the wider output schema — from being reconstructed against the wrong schema; input records all share one uniform schema. - Spill priority:
15, between grace-hash Combine (10) and external sort (20). A grouped record buffer costs more to evict than grace partitions (reload pays the re-synthesis CPU) but less than an external-sort merge. Reshape cannot back-pressure — once its predecessor has drained there is no upstream channel to pause — so under memory pressure it always spills its own buffer in-thread rather than pausing a producer. - Largest-first, stop at the threshold: when the budget trips, Reshape evicts the largest resident groups first and stops as soon as the resident footprint drops back under the soft threshold — it does not drain every group. The cross-group resident peak is bounded near the soft limit.
- Skew (one giant group): a single correlation group whose resident tail alone exceeds the budget is spilled incrementally — sliced by the upper bits of each record’s admission sequence so successive spill waves evict fractions of the one group — while smaller groups stay resident. The ingest-time resident peak therefore stays bounded even under one giant skewed group.
Reshape’s buffer is byte-budget-only: it does not honor the error_handling.max_group_buffer record-count cap. That cap bounds the correlation-commit group buffer used by the retraction machinery, a different buffer; Reshape’s footprint is bounded by memory.limit and the spill path, not by a per-group record count.
The on-disk spill volume Reshape produces is surfaced per stage in clinker run --explain (the Estimated spill volume and Spill compression sections) and, after a run, in the actual per-stage spill totals.
Limits
Two current limitations qualify the “identical whether spilled or resident” guarantee above:
-
A single correlation group must fit the memory budget at finalize. The no-cascade contract requires the whole group to be resident when its rules fire, so even though cross-group and ingest-time peaks spill to disk, the finalize reload of one group needs that group to fit. Skew slicing bounds the ingest peak, but a single correlation group larger than
memory.limithas no in-budget representation. Rather than risk an out-of-memory crash, the run fails loud withE310 MemoryBudgetExceeded, naming the Reshape node, the offendingpartition_bygroup, and its footprint against the budget:E310 backfill: arena exceeded budget (128000/8192) [one Reshape correlation group [employee_id="employee-00000"] does not fit; the reported use is that group's reload footprint alone. ...]Raising
memory.limitis the only fix that leaves your output unchanged. Raise it clear of the reported figure — finalize also holds the run’s remaining groups, so that figure is a floor, not a target.The other two levers both change what you get, and are worth knowing only so you can weigh them deliberately:
- Dropping columns this node does not read, in an upstream Transform, shrinks each buffered record. But Reshape’s output row is the upstream columns plus its
$meta.*audit columns — it never drops a column itself — so anything you strip upstream is also missing from the written output. - Narrowing
partition_byshrinks the group too, but that key defines the group the rules evaluate against, so a narrower key changes which records each rule sees and therefore changes your results. Treat the grouping key as a modelling decision, never as a memory knob.
(A future two-pass finalize could lift this limit.)
Run
clinker explain --code E310for remediation keyed to whichever memory surface overran. - Dropping columns this node does not read, in an upstream Transform, shrinks each buffered record. But Reshape’s output row is the upstream columns plus its
-
Reshape rules cannot reference
$docdocument context. Because the spill round-trip does not yet preserve document envelope context, a$doc.*reference in a rule’swhen,mutate.set, orsynthesizeexpression would resolve to the real envelope for a resident group but to null for a spilled one — output that depends on the memory budget. A pipeline whose Reshape rules reference$docis rejected at compile time. Move the$doclookup into an upstream Transform that copies the value into a record column, then reference that column in the Reshape rule.
Cull Nodes
Cull nodes observe a whole correlation group and remove the entire group when a group-level predicate holds, routing the removed records to a side-output port instead of discarding them. They are the node for “look at everything an entity did, then set the whole entity aside for review or reprocessing” — work that operates on an aggregate property of a group, not on individual rows:
- Route fans out per record on a row predicate.
- Reshape mutates rows and synthesizes new ones within a group.
- Aggregate reduces a group to one summary row.
None of those remove a whole correlation group based on an aggregate property of the group and emit the removed rows on a second stream. Cull does.
Cull is a blocking grouping operator: it buffers every record of a group before any output leaves, because a group-level predicate (e.g. “this group has more than 100 rows”) cannot be decided until the whole group is seen. It has two output ports: the main port (kept groups) and the removed_to side-output port (removed groups).
Basic structure
- type: cull
name: flag_large_histories
input: backfill
config:
partition_by: [employee_id]
removed_to: review
rules:
- name: too_many_plans
drop_group_when: "count(*) > 3"
For each employee_id group, if the group holds more than three rows the whole group is routed to the review side output; every other group flows to the main output.
partition_by
A list of field names. Records sharing the same values for all partition_by fields form one group, and the removal predicate observes one whole group at a time. This is the correlation key the operator reasons over. partition_by must cover every visible correlation-key field so group identity is preserved on both output ports.
Whole-input grouping (partition_by: [])
An empty list is the degenerate case of that key: every record shares it, so the entire input forms one group and drop_group_when decides the whole dataset at once — every record is kept, or every record is routed to removed_to.
- type: cull
name: drop_bad
input: events
config:
partition_by: [] # one group: the whole input
removed_to: removed
rules:
- name: drop_any_error
drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"
That pipeline routes all records to removed if any single record has status: error. Keyed by [account] the same rule would remove only the offending account’s records — so reach for whole-input grouping when the decision genuinely concerns the batch (an all-or-nothing gate on a delivery), not when you meant a per-entity rule.
The whole input must then fit memory.limit at finalize, since a single group has to be resident when its predicate runs (see the limit below). The sibling group-count bound does not apply: one group is as few as the decision state can be.
Values Cull cannot key
Cull groups a record under a single null group in two cases: the column is absent from the record, or its value is an explicit null.
Everything else is either its own group or a hard error:
- An empty string (
"") is its own group, distinct fromaccount=null. - Numbers group by exact value, as in Aggregate:
1,1.0and the decimal1are one group, and distinct integers never merge, however large. - A
NaNfloat, of either sign, is one group of its own, distinct from the null group. - An array- or map-valued cell aborts the run rather than grouping — a partition key must be a single scalar value. This abort currently presents as an internal error, but it is a data condition, not an engine defect: fix the offending column rather than treating the message as an engine invariant failure.
Reshape treats empty strings and multi-value cells differently: it folds both into its null group instead. The blank-versus-null divergence is tracked in #1022. The multi-value behavior is separate; until both rules are deliberately aligned or retained, do not assume one node’s grouping matches the other’s.
order_by
Optional. A list of sort fields that orders the rows of each group as they are written. It does not change which groups are removed: the removal rule is evaluated over the group in arrival order (#1264). Arrival order breaks ties.
Each entry is either a field name, which sorts ascending, or a map { field, order, null_order }:
| Key | Values | Default |
|---|---|---|
field | a column of the input | required |
order | asc or desc | asc |
null_order | first or last | last |
order_by:
- txn_date # same as { field: txn_date }
- { field: amount, order: desc, null_order: first }
A group’s rows are ordered exactly as a Sink sort_order orders rows (see How values are ordered): a null goes where null_order puts it, last by default for asc and desc alike; numbers compare by their exact value across integers, floats and decimals; every NaN is one value after every number, and so comes first under desc; -0.0 equals 0.0; and rows the order calls equal keep their arrival order.
null_order: drop is rejected when the pipeline is planned. order_by only arranges the rows of a group; it never removes any, and every record of a group is kept or removed together. To leave out rows whose field is null, filter them in a Transform before the Cull:
- type: transform
name: dated_only
input: transactions
config:
cxl: |
filter not txn_date.is_null()
Rules
Each entry in rules: is a declarative removal rule with a name and a drop_group_when predicate. A group is removed when any rule’s predicate holds (the rules are OR-combined).
drop_group_when — the group-level removal predicate
drop_group_when is a CXL boolean expression evaluated in aggregate context over the whole group (group-by = partition_by). Because it is an aggregate expression, it uses CXL’s aggregate functions:
| Aggregate | Meaning |
|---|---|
count(*) | number of rows in the group |
sum(<expr>) | sum of an expression over the group |
min(<expr>) / max(<expr>) | minimum / maximum over the group |
avg(<expr>) | mean over the group |
rules:
- name: too_many_plans
drop_group_when: "count(*) > 3"
- name: high_total
drop_group_when: "sum(amount) > 10000"
CXL’s bare aggregate vocabulary is sum / count / min / max / avg / collect / weighted_avg — there is no bare any() aggregate. To express “remove the group if any row matches a condition”, sum an indicator and compare to zero:
rules:
- name: drop_error_groups
# Remove any account group containing at least one `error` row.
drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"
Ordered comparisons (>, <, >=, <=) work over every comparable aggregate type, not just numbers — the predicate uses the same comparison rules as a Transform. Numbers, strings, and dates all order:
rules:
- name: late_alphabet
drop_group_when: "max(name) > 'M'" # string ordering
- name: recent_hire
drop_group_when: "max(hired) >= #2020-01-01#" # date ordering (`#YYYY-MM-DD#` literal)
A group whose aggregate operand is null (for example max(...) over an all-null column) compares as false — a null operand never removes the group and never errors.
Comments in a predicate
A drop_group_when predicate may carry a # line-comment to explain the rule inline. Each rule’s predicate is parsed on its own, so a trailing comment applies only to that rule:
rules:
- name: drop_error_groups
drop_group_when: "sum(if status == 'error' then 1 else 0) > 0 # any error row removes the group"
- name: high_total
drop_group_when: "sum(amount) > 10000 # large accounts"
The comment is source text only — it never changes the compiled decision. (A #YYYY-MM-DD# date literal is unaffected: it lexes as a date, not a comment.)
Output ports: main and removed_to
Cull has two producer-side output ports, the same mechanism a Route uses for its branches — not the dead-letter queue. Removed records are valid rows the operator deliberately partitions onto a second stream, not errors.
Downstream nodes draw from the two ports by reference:
- The main output (kept groups) is referenced by the Cull node’s bare name:
input: flag_large_histories. - The side output (removed groups) is referenced as
<cull>.<removed_to>:input: flag_large_histories.review.
- type: sink
name: kept
input: flag_large_histories # main port — kept groups
config: { name: kept, type: csv, path: kept.csv }
- type: sink
name: review
input: flag_large_histories.review # side-output port — removed groups
config: { name: review, type: csv, path: review.csv }
removed_to must be a non-empty name distinct from the Cull node’s own name (enforced at compile time, so the two ports are always distinguishable).
A single downstream node may draw from both ports — for example a Merge recombining the kept and removed streams (inputs: [flag_large_histories, flag_large_histories.review]) — and receives the union of both ports’ records.
removed_to is not the DLQ
The removed_to port carries the unchanged upstream schema — Cull does not widen, and both ports emit exactly the input columns. Removed records are not DlqEntrys and never appear in the dead-letter queue or its counters; they flow down a normal data edge to whatever node draws the removed_to port (an audit sink, a reprocessing branch, another transform). Use the DLQ for errors; use a Cull side output for valid records you want to handle separately. See Error Handling & DLQ for the error path.
Memory model
Cull is a blocking, grouped operator: it groups every input record by partition_by before any record leaves, because the group-level drop_group_when predicate is an aggregate property of the whole group and cannot be folded into a per-record keep/remove decision. It therefore cannot stream — the full group set materializes before the first output row leaves.
That per-group buffer is governed by the same central memory arbitrator every other blocking operator polls (see Memory & Spill). As records are grouped, Cull tracks the live in-memory footprint and, whenever the run crosses the soft spill threshold (80% of memory.limit by default), it spills buffered groups to disk:
- What spills: the raw input records. On reload at finalize, each group is re-split onto its output port exactly as it would have been without spilling, so the output is identical whether a group stayed resident or round-tripped through disk — including within-group row order, which is restored to arrival order after a reload. (The per-group removal decision is computed from an in-memory aggregate over the same records; that aggregate state is
O(distinct groups)and is never spilled — only the raw records spill. It cannot spill because Cull has no upstream channel to pause, so instead it is bounded by a hard check: if that state plus the run’s other live charged memory would exceedmemory.limit, the run fails loud with a memory-budget error rather than growing it unbounded. See the limit below.) - Spill priority:
15, between grace-hash Combine (10) and external sort (20), matching Reshape. Cull cannot back-pressure — once its predecessor has drained there is no upstream channel to pause — so under memory pressure it always spills its own buffer in-thread rather than pausing a producer. - Largest-first, stop at the threshold: when the budget trips, Cull evicts the largest resident groups first and stops as soon as the resident footprint drops back under the soft threshold — it does not drain every group.
- Skew (one giant group): a single correlation group whose resident tail alone exceeds the budget is spilled incrementally — sliced by the upper bits of each record’s admission sequence — while smaller groups stay resident, so the ingest-time resident peak stays bounded even under one giant skewed group.
The on-disk spill volume Cull produces is surfaced per stage in clinker run --explain (the Estimated spill volume and Spill compression sections) and, after a run, in the actual per-stage spill totals.
Limit: a single group must fit the finalize budget
Cull evaluates its group-level predicate against the whole group at once, so even though cross-group and ingest-time peaks spill to disk, the finalize reload of one group needs that group to fit the memory budget. Skew slicing bounds the ingest peak, but a single correlation group larger than memory.limit has no in-budget representation. Rather than risk an out-of-memory crash, the run fails loud with E310 MemoryBudgetExceeded, naming the Cull node, the offending partition_by group, and its footprint against the budget:
E310 drop_big: arena exceeded budget (512000/8192) [one Cull correlation
group [account="BIG"] does not fit; the reported use is that group's reload
footprint alone. ...]
Raising memory.limit is the only fix that leaves your output unchanged. Raise it clear of the reported figure — finalize also holds the run’s remaining groups and the per-group decision map, so that figure is a floor, not a target.
The other two levers both change what you get, and are worth knowing only so you can weigh them deliberately:
- Dropping columns this node does not read, in an upstream Transform, shrinks each buffered record. But Cull filters rows, never columns — both its ports carry the unchanged upstream schema — so anything you strip upstream is also missing from the main and
removed_tooutputs. - Narrowing
partition_byshrinks the group too, but that key defines the groupdrop_group_whenevaluates over. With a rule likecount(*) > 100, splitting one account across a finer key drops each resulting group below the threshold, so an account that should have been removed is emitted on the main port instead — the run “works” and quietly returns a different result set. Treat the grouping key as a modelling decision, never as a memory knob.
Run clinker explain --code E310 for remediation keyed to whichever memory surface overran, or see the memory guide.
There is a second, symmetric bound on the number of groups. The per-group removal decision is held in an in-memory aggregate that is O(distinct groups) and — unlike the raw records — cannot spill. If a partition key is so high-cardinality that the decision state plus the run’s other live charged memory would exceed memory.limit (many small groups rather than one giant group), the run likewise fails loud with E310, rather than growing that state unbounded. Here, coarsening partition_by is a legitimate fix only if the coarser key is the grouping you actually meant — the same caveat as above applies. Otherwise, raise memory.limit.
Several distinct memory surfaces can raise E310 against a Cull node, so read the [...] detail to see which one overran. The ones documented here are Cull correlation group [...] (the giant-group case above), Cull drop-decision aggregate state (the group-count bound), Cull cross-region tee admission (a downstream stage in a different deferred region forces this node’s output to be parked in memory), and node-buffer materialization overlap (a consumer must collect one of Cull’s sequential port scans into a resident vector). Cull’s main and removed_to handoff buffers themselves are spill-eligible, including when a port fans out. Other surfaces the shared runtime charges may name this node too — the detail string is the authority, not this list.
Envelope Nodes
Envelope nodes frame a body stream into per-document documents. An Envelope is a discrete, composable stage you can place after any operator — a Transform, a Merge, a Combine, or an Aggregate — to declare “from here on, treat the records as belonging to framed documents.” It mirrors the message/EDI/XML envelope-wrapper pattern (the Enterprise Integration Patterns Envelope Wrapper, XProc’s p:wrap-sequence): the body is the payload, and the envelope is the document boundary around it.
This page documents the preserve and concat framing strategies and the orthogonal header: / footer: synthesis that layers on top of either.
Interactive companion: How many documents? shows the documents a Sink writes under preserve and concat, with and without a synthesized footer, how an Aggregate after the Envelope rolls up, and which shapes E347 and E355 reject.
Basic structure
- type: envelope
name: framed
body: merged
config:
strategy: preserve
The node reads its body: input and emits the same records, framed per document. A downstream Sink with reconstruct_envelope: true then writes one framed document per body grain.
Inputs: body, the wired header, and the not-yet-wired trailer
| Input | Required | Status |
|---|---|---|
body | yes | The records to frame into documents. |
header | no | A 1-row-per-grain header stream. A wired value replaces each body document’s header with the matching header record, attached by document grain. |
trailer | no | A stream whose records append to each framed document. Accepted in config but not yet wired — a wired value is rejected at plan validation this release. |
When you omit header:, an Envelope frames each body record using the body’s own ambient envelope — the document context every record already carries from its source.
Wiring a header: port replaces the document’s header
A wired header: input is a second stream carrying one header record per document, each on the same document grain as the body it frames. The node attaches a header to a body strictly by grain, so the header record reaching it must carry the body’s grain — and the node then replaces that document’s ambient header with the wired header record (the framing grain is preserved; only the header changes). This is transform-in-place header replacement: rewrite a header’s values upstream — for example, override a batch id or stamp a run date — and frame the body with the rewritten header instead of the source’s original.
nodes:
# `rewrite_header` rewrites the source header's values while keeping each
# record's grain, so the rewritten header still grounds to its body document.
- type: envelope
name: framed
body: payments
header: rewrite_header
config: { strategy: preserve }
The header record must carry a body document’s grain. A grain-preserving Transform of the source’s promoted header keeps it; a replacement from a different source establishes it via a business-key join against the body. A header record whose grain matches no in-flight body document (or carries the synthetic, ungrounded grain a Transform stamps onto a record it builds from scratch) cannot be placed, so the run fails with E351 (run clinker explain --code E351 for the full write-up):
envelope "framed": a wired header record carries document grain <grain>, which
matches no in-flight body grain (or is a synthetic / ungrounded grain). The node
attaches a header to a body strictly by grain, so it cannot place a header that
grounds to no body document.
Exactly one header record may carry each body document’s grain. When the wired header stream carries two or more records on the same grain, the node has no rule to fold a second header onto an already-framed document, so it refuses to silently keep one and drop the rest. The run fails with E352 (run clinker explain --code E352 for the full write-up):
envelope "framed": the wired header input carries two or more records for
document grain <grain> — exactly one header record per document grain is
required.
Deduplicate the header stream to one record per grain upstream — an aggregate or distinct on the grain’s business key, or a Transform that emits a single rewritten header per source document.
A wired trailer: input is still rejected this release with a clear “not yet supported” message:
envelope node "framed": explicit `trailer` input wiring is not yet supported —
omit it to frame with the body's own envelope
strategy: preserve
preserve emits one framed document per body grain. It is a transparent framing stage: body records pass through with their document context and grain unchanged, and the document-boundary signals are forwarded verbatim. Inserting a preserve Envelope between a body stage and a Sink is byte-identical to today’s per-document framing — its value is being the explicit, composable stage that later strategies extend, not a change in output.
preserve is the default, so config: { strategy: preserve } and an empty config: {} are equivalent.
Framing is keyed on the document grain, never the source file
The grain is the level at which one logical document is reconstructed — and it is not always one-per-file:
- A nested X12 interchange frames once per interchange. The
GSfunctional-group andSTtransaction-set levels inherit the interchange grain, so anISA … IEAinterchange is one framed document regardless of how many groups or transaction sets it nests. - An HL7 multi-message file frames once per message. Each
MSHmessage opens its own grain, so a single file containing several messages produces several framed documents.
Because framing keys on the grain rather than the source file, splitting or combining files never silently changes the document count.
strategy: concat
concat does the opposite of preserve: it collapses a multi-document body into one framed document. Every body record is re-stamped onto a single consolidated document context, so the body opens and closes exactly once regardless of how many documents fed in. This is the strategy to use when several source documents — say two files joined by a Merge — should write as one consolidated document with a single header and footer.
nodes:
- type: merge
name: both
inputs: [file_a, file_b]
- type: envelope
name: framed
body: both
config: { strategy: concat }
- type: sink
name: out
input: framed
config:
name: out
type: csv
path: out.csv
reconstruct_envelope: true
Re-stamping changes only the framing (the grain) and the ambient $doc.* view a record sees — it does not disturb per-record fields. In particular $source.file is a real column stamped when each record is read, so it still reports the record’s own originating file after a concat. Concat is lossless on per-record provenance; it changes only which document the record is framed inside.
The consolidated header, and the two-headers conflict
One consolidated document can carry only one envelope header. concat derives it from the headers of the documents that contribute body records, taking one header per document:
- Every header agrees (or there is only one) → the consolidated document carries that common header.
- No document carries a header → the consolidated document is headerless.
- A headed document and a headerless document → the single header wins; the headerless document coexists with it (no conflict).
Only documents that contribute body records take part: a document that carries a header but no body records frames nothing once consolidated, so it never enters the comparison. Header identity is structural — two documents share a header when they declare the same sections, in the same order, with the same field values. Engine-added fields whose names start with $ (such as $raw, the raw segment text) are ignored, so two files whose headers differ only in raw content fold to one header, and the consolidated document keeps the first document’s full header, $raw included. A difference in a field you declare (for example an extracted control number) makes the headers distinct.
When the body carries two or more distinct non-empty headers, concat refuses to silently keep one and drop the rest. The run fails with E350 (run clinker explain --code E350 for the full write-up):
envelope "framed": concat collapses the body into one framed document, but the
body carried 2 distinct non-empty envelope headers — one document can frame only
one header, so concat will not silently drop the rest. Make the headers identical
upstream, or add a header-folding strategy that declares which header the
consolidated document keeps.
To resolve a conflict, either keep the documents separate with preserve, or make the headers identical upstream (project them to the same sections and values). header: synthesis (below) does not resolve it: concat compares the input headers and raises E350 before any synthesis runs (#1385).
Synthesizing a header and footer
header: and footer: are orthogonal to the framing strategy. The strategy decides how many output documents there are (preserve = one per body grain, concat = one consolidated); synthesis decides what header and footer each of those documents carries. The node computes a fresh header (declarative scalar expressions) and footer (streaming aggregates over the framed body) per output document, stamps them as named sections into the document’s envelope, and the same header_from_doc / footer_from_doc writer path renders them.
Both maps are keyed section -> field -> CXL expression. The inner field map preserves declaration order, which is the rendered cell order. A downstream Sink renders a section through header_from_doc / footer_from_doc, and those may name only a section that a feeding Source declares (E346), so in practice a synthesized section reuses a declared section name.
A synthesized section replaces the whole same-named section on the document; it does not add fields to it, so the section’s original fields (for example a source’s interchange.tag) are gone from that document. Other sections ride through untouched. Give header: and footer: different section names: if both name the same section, the header is applied last and replaces the footer, whose fields are lost without an error (#1388).
- type: envelope
name: framed
body: merged
config:
strategy: concat # or preserve — synthesis works the same on either
header: # section -> field -> scalar CXL
group: # a different section from the footer's
sender: $vars.sender_id
created: $pipeline.run_date
footer: # section -> field -> aggregate CXL
interchange:
record_count: count()
total: sum(amount)
Header fields: evaluated once at document open
A header: field is a scalar expression evaluated once per output document, before the body streams. It may read only inputs known at document open — pipeline configuration ($vars), per-document provenance ($source), pipeline-scope state ($pipeline), and the ambient envelope ($doc.*). It may not read a body column, because there is no “current body record” when the header is emitted; a body-column reference is rejected at compile time with E353 (run clinker explain --code E353). Put body-derived values in a footer aggregate instead.
Footer fields: streaming aggregates over the body
A footer: field folds the document’s body records into a footer value at the document’s close. The fold is an O(1) accumulator per open document, so it supports exactly the streaming distributive/algebraic aggregates over a bare field or *:
| Aggregate | Example |
|---|---|
count | count(), count(*) |
sum | sum(amount) |
avg | avg(amount) |
min | min(amount) |
max | max(amount) |
Anything else is rejected at compile time. A holistic or unbounded aggregate (collect, any, weighted_avg) or a composed/multi-argument aggregate (sum(amount * 1.1), weighted_avg(value, weight)) raises E354 (run clinker explain --code E354); a non-aggregate function (median, mode) fails earlier as an unknown function. Project a composed value in an upstream Transform, or compute a holistic value in an upstream Aggregate, then aggregate the bare column here.
Synthesis is orthogonal: the same footer differs across strategies
Because the strategy sets the grain cardinality, the same footer: { interchange: { record_count: count() } } produces a different result on each strategy over the same body:
- Under
strategy: preserve, each body document frames its own output document, so each footer’srecord_countis that document’s body count. - Under
strategy: concat, the whole body collapses to one document, so the single footer’srecord_countis the merged body count.
Placement
An Envelope is a normal single-input, single-output node — put it anywhere a record stream flows:
nodes:
# … sources, a Combine that joins two streams into `merged` …
- type: envelope
name: framed
body: merged
config: { strategy: concat } # preserve here is rejected with E347
- type: sink
name: out
input: framed
config:
name: out
type: csv
path: out.csv
reconstruct_envelope: true
After a Combine, Aggregate or Composition, a Sink with reconstruct_envelope: true needs a concat Envelope in between: those nodes’ rows carry no document of their own, and concat frames them as one. A preserve Envelope there is rejected with E347. For an Aggregate that is right, since its rows would pass through unframed; for a Combine the check is stricter than it needs to be, because the joined rows do keep their driver row’s document (#1384).
After a Combine this works: the joined rows keep their driver row’s document, and concat writes them as one framed document. After an Aggregate it plans without error but does not frame yet: Aggregate rows carry no source file, so the consolidated document has none either, and the Sink writes the rows with no header or footer (#603). Reshape output loses its document the same way (#1317).
Memory model
Both strategies re-park the body into the node’s own buffer slot, which the engine’s memory arbitrator governs and spills to disk under pressure — so neither strategy is bounded by total input size held in RAM.
preserve is a transparent framing pass-through: it forwards records and their document boundaries unchanged. concat additionally re-stamps each record onto the one consolidated document context and replaces the per-document boundaries with a single open/close pair; the header consolidation it does first groups the body records by document — one document’s worth of body records shares one grain and one header — so it does one pass over the records to collect the distinct headers, comparing only one envelope per document (the work is bounded by the number of documents, not the number of body rows).
header: / footer: synthesis adds, on top of the materialized body, one O(1) accumulator per footer field per open document — every allowed footer aggregate (count / sum / avg / min / max) holds a fixed-size state regardless of how many body rows it folds. So a document’s footer state is independent of its body-row count, and synthesis stays within the node’s bounded-memory model.
Sink Nodes
Sink nodes write processed records to files. They are the terminal nodes of a pipeline – every pipeline path must end at a sink (or records are silently dropped).
Terminal-node migration:
type: outputis retired and rejected withE376. Replace only the terminal discriminator with the paste-ready correction:- type: sinkComposition and node output ports, produced artifacts, files and paths, serialization formats, stdout, command or machine output, writer results, and OpenLineage output datasets keep the word “output.” See Production Contracts for the compatibility boundary.
Use the same name for the Sink node and its config.name. At the current
implementation boundary, mismatched names can produce a writer-mode error or
empty published files instead of a planning diagnostic. The examples use
matching names; planning success alone does not establish correct output for
a mismatched pair.
Basic structure
- type: sink
name: result
input: transform_node
config:
name: result
type: csv
path: "./output/result.csv"
The type: field selects the output format: csv, json, xml, fixed_width, edifact, x12, hl7, or swift. The edifact, x12, and swift writers reconstruct one interchange/message envelope around emitted records; the hl7 writer re-emits HL7 v2 segments and optionally wraps them in batch/file envelopes. See EDIFACT Format, X12 Format, HL7 v2 Format, and SWIFT MT Format.
Structured single-writer outputs (edifact, x12, hl7, and swift) accept one concrete document grain per output file. A multi-file source or multi-input merge feeding one of these outputs is rejected instead of being silently written as one merged envelope. To write multiple structured documents, consolidate them deliberately with an Envelope node first or route each document to a separate output path.
Local and network-share destinations
An output path may be on a local filesystem or a mounted NFS/SMB share. Clinker detects the filesystem behind the actual destination and applies its contained-create and same-filesystem promotion rules there; users do not label their production paths with a CI profile. The committed filesystem matrix qualifies Clinker’s semantics against specific loopback NFSv4.1 and SMB3.1.1 mounts, but it cannot certify every vendor appliance, mount option, outage mode, or corporate network. Qualify representative production mounts before depending on atomic promotion during an outage or failover.
Clinker creates Unix output files with owner-only mode 0600. This prevents a
new file from accidentally inheriting broad access in a shared drop zone. If a
different service account or group must consume the result, arrange that access
explicitly with the destination’s ACL/ownership policy; Clinker does not
currently expose an output-mode setting.
For performance, keep spill files and optional staged input copies on a local disk when one is available. Blocking operators can create substantial random I/O, and performing that work directly on a network share adds latency and network traffic. The final output is still written as a hidden file on the destination filesystem and promoted there, so the completed file never relies on a cross-filesystem rename from local storage.
This commit lifecycle applies to single-file, per-source-file fan-out, and
split: outputs. Clinker does not open or truncate an existing final while a
replacement is running. Before publication, Clinker synchronizes and validates
the complete output set, then promotes each hidden file directly to its final
name. An overwrite is one atomic replacement rename: Clinker never moves the
previous final out of the way first. It also never claims that a multi-file set
can be rolled back after some replacements are already visible.
Publication is not one atomic filesystem operation for the whole set. Each
individual rename is atomic, but a reader may briefly observe a mixture while
the finite commit walks several destinations. If a promotion or directory sync
fails, Clinker stops, exits 4, and reports three exact groups: finals that are
visible and synchronized, finals that are visible but whose parent sync failed,
and unpublished hidden partials. Already-visible finals stay visible; remaining
old finals stay untouched. A process or machine crash in that window can leave
the same mixed set plus .partial or .reservation siblings. Reconcile those
named paths before retrying or consuming the set. Clinker does not create or
use .backup files for output publication.
Every collision policy uses a hidden sibling reservation, including overwrite.
This ensures only one live publisher may mutate a final destination. if_exists: error uses a no-replace promotion; if_exists: overwrite and clinker run --force replace only at successful promotion; if_exists: unique_suffix
reserves candidate names until one wins. Reservations never expose zero-byte
final placeholders. A reservation holds an operating-system lock and records
its owner process. A later run reclaims it only after a short creation grace
period and only when the lock is acquirable, proving that no live publisher owns
it. If reservation cleanup fails after successful publication, Clinker exits
4 and names both the visible final and stale reservation as cleanup debt.
When unique_suffix can find no name at all because the destination itself
refuses every candidate — the directory is not writable by this run — the
diagnostic names the path you wrote, not the numbered candidate the search
happened to stop on, and says that the destination rather than the name is what
refused. Fix the directory’s permissions, or point path: somewhere this run
may write.
Rendered fan-out paths are validated as new output paths. Directory traversal, an absolute result produced from a relative template, symbolic-link/reparse ancestors, and cross-filesystem promotion fail before a final is touched. Create the intended destination directories ahead of the run; Clinker does not follow rendered paths while creating missing fan-out parents.
{source_file} and {source_path} create one output route per discovered
source file. Two source files that render to the same destination are rejected
before any output is staged, with both source paths in the diagnostic. Escape a
token as {{source_file}} or {{source_path}} when the braces are intended as
literal filename text. Runtime source names and paths are inserted as opaque
text: braces inside an actual filename are never interpreted as another token.
When fan-out is combined with split:, every source has its own segment
sequence: each starts at sequence 1 and rolls over independently.
With write_meta: true, Clinker writes a .meta.json sidecar for every actual
committed final. A split output therefore gets one sidecar per segment, and a
fan-out output gets one per rendered destination; no sidecar is written for the
unrendered base template. Main outputs and sidecars share the same publication
ledger, so a path collision between any two of them fails before publication
and names both producers. Counters that are not known at sidecar-preparation
time are omitted from the JSON rather than written as misleading zeroes.
When two paths are the same destination
Two Sink nodes — or a Sink node and a DLQ path — that resolve to one file
are rejected at plan time with E322, before any record is read. Deciding that
means deciding when two differently-spelled paths name one file, and that
depends on the volume you are writing to, not on the text. Clinker measures the
volume rather than guessing from its type, by creating and removing a probe file
in the destination directory.
Two paths are the same destination when they differ only in:
.components, or relative versus absolute spelling../out/errors.csvand/data/out/errors.csvfrom/dataare one file everywhere.- A symlinked parent directory. The existing part of each path is resolved, so a link and its target are one place.
- Letter case — only on a volume that ignores case. The default macOS
(APFS) and Windows (NTFS) volumes do; ext4, xfs, and btrfs do not. Where it
applies, it covers the whole of Unicode, not just ASCII:
Ärger.csvandärger.csvare one file, as areΣ.csvandσ.csv.straße.csvandstrasse.csvare always two files — no filesystem treats them as one. - Unicode normal form — only on a volume that ignores it. APFS and HFS+ do,
in both their case-sensitive and case-insensitive variants; ext4 and NTFS do
not. Where it applies,
café.csvwritten with a precomposedéand the same name written aseplus a combining accent are one file. This is independent of the case rule: a case-sensitive APFS volume ignores normal form while still tellingCafé.csvandcafé.csvapart.
Because the last two depend on the volume, the same pipeline can be accepted on Linux and rejected on macOS. That is not an inconsistency — the two disks really do behave differently, and the rejection is the one that prevented a file from being written twice.
Two limits are worth knowing:
- Clinker may report a collision on a volume that would in fact have kept the two files apart — for instance on an older Windows volume whose case table predates a character you used. The run stops with both paths named, and renaming either one clears it.
- The reverse is possible on Windows, which folds some letters according to
Turkish and Azeri rules that no locale-independent table reproduces.
İ.csvandi.csvmay be one file on such a volume while Clinker still sees two. If you write output paths that differ only in dotted or dotlessI, give them distinct names.
Direct broadcast to several outputs
Several Sink nodes may name the same input. This is a broadcast: every
Sink receives every upstream record, regardless of node declaration order.
The run report counts one write per sink, so five input records feeding a CSV
and a JSON Sink produce records_written: 10.
Use a Route node when outputs should receive different subsets.
Writing a field such as _route does not select a destination; it is an
ordinary output column unless a Route condition explicitly reads it.
Field control
Sink nodes can either pass every upstream field through to the writer or restrict output to the fields the upstream transform explicitly emitted. Several options control which fields appear and how they are named.
Unmapped input field passthrough
include_unmapped: false # Default: true
When true (the default), every field on an input record that the upstream transform did not explicitly emit still passes through to the output unchanged. This includes fields the source’s on_unmapped: auto_widen policy absorbed into the per-record $widened sidecar map – their contents expand back to top-level columns at the sink.
When false, only fields named by an emit statement in the upstream transform appear in the output. The $widened sidecar slot is stripped and undeclared input fields are dropped.
When true, how a carried-along column reaches the writer depends on the output format. Self-describing formats (JSON / NDJSON / XML) write each record’s own keys. A CSV output widens its header to the union of every record’s columns when it can materialize the batch, and otherwise — on a bounded-memory streaming path (a Merge, a fused Transform, a single-branch Route, a streaming-strategy Aggregate, or the probe side of a hash-build-probe Combine feeding the output), or an envelope-reconstructing path — fails loudly with a SchemaDrift error rather than dropping a column it cannot fit under its already-committed header. A fixed-width output has no room for an undeclared column and likewise raises SchemaDrift. See Auto-Widen & Schema Drift → Schema drift across records.
Migration notice
The default flipped from false to true in a recent release (see issue #90). Pipelines that relied on the previous behavior – where output records contained only the fields explicitly emitted upstream – must now set include_unmapped: false explicitly to restore that shape.
The flag composes independently with include_correlation_keys: true – see below. See Auto-Widen & Schema Drift -> Output controls for the full specification and cross-format flow examples.
Worked example
Suppose the upstream source emits records with order_id, customer_id, amount, and region, and a transform that emits only one derived field:
- type: transform
name: classify
input: orders
config:
cxl: |
emit amount_bucket = if amount >= 1000 then "high" else "low"
With include_unmapped: true (the default), each output record carries order_id, customer_id, amount, region, and amount_bucket. With include_unmapped: false, each output record carries only amount_bucket. The transform’s CXL is unchanged in both cases – the Sink node decides the field set.
Include correlation-key shadow columns
include_correlation_keys: true # Default: false
When a source declares a correlation_key:, the engine tracks correlation-group identity on hidden columns that are stripped from output by default. Set include_correlation_keys: true to surface them in the writer output — typically for debugging correlation-group routing or auditing DLQ behavior. See Correlation Keys.
include_correlation_keys does not surface auto-widened columns – include_unmapped is the separate flag for that. The two are independent: each, both, or neither can be set.
Nested columns and writer capabilities
JSON writes map and array values recursively. XML maps ordinary keys to child
elements, reserves unescaped @.../#text keys for attributes/text, and writes
an array at a top-level column only when that output-facing column comes from a
multiple: true declaration; arrays nested inside a map remain native XML
structure. CSV likewise joins an array only for a compiled multiple: true
column. An undeclared array at either writer is a routing error, not an implicit
declaration. CSV, fixed-width, EDIFACT, X12, and HL7 reject a map that reaches a
column slot. See JSON,
XML, and
Auto-Widen & Schema Drift.
Field mapping
mapping: declares the columns the file carries – which columns, under what
names, in what order – without changing upstream CXL. It is a sequence, one
item per output column:
mapping:
- order_id # carried through under its own name
- sold_to: customer_id # written as `sold_to`, read from `customer_id`
- contact_email: customer_email
- channel
- sku
Two item shapes:
- A bare column name emits that column unchanged. This is the common case, and it costs one line naming the column once.
- A single-key pair renames. The output name is on the left, the source column on the right – the same side the bare form names. Reading an item left to right always tells you what appears in the file first.
The renames are the only items carrying a colon, so in a wide output they are found by scanning for structure rather than by comparing two names per line.
Order and selection
Declaration order is the output column order. Listed columns are written first, in the order the block declares them, whatever order they arrive in.
include_unmapped governs everything the block does not list. With
include_unmapped: true (the default) unlisted columns are appended after the
declared ones, in their existing relative order. With include_unmapped: false
they are dropped, so the block becomes the complete statement of the output:
include_unmapped: false
mapping:
- department
- surname: last_name
- first_name
Given upstream columns first_name, last_name, department, that writes exactly
department,surname,first_name.
Every record carries every declared column. When a record does not supply an
item’s source column, that column is still written, empty. The file’s shape
follows the block, not the data — so a stream whose records differ in shape (a
multi-record-type source, a column arriving through auto_widen, a composition
body’s open row) still produces one stable column set in declaration order,
rather than one that depends on which record happened to arrive first.
One upstream column may feed two output columns – - sku and
- item_code: sku – because names must be unique on the output side, not the
source side. Declaring the same output name twice is rejected (E364): a
file cannot carry two columns under one header.
For the same reason, an output name that include_unmapped: true would also
carry through is rejected. If upstream already has a sold_to column, writing
- sold_to: customer_id under include_unmapped: true would put two sold_to
columns in the file and readers would resolve the wrong one. Rename the mapped
column, exclude the upstream one, or set include_unmapped: false.
Where the compiler cannot enumerate the upstream columns, the same collision reaches the run. The mapped value wins – the block is your explicit statement of what the file carries – and the displaced upstream column is named in a W366 warning at the end of the run. Applying one of the three fixes above silences it.
Diagnostics
A mapping: item naming a column that does not exist at that point in the
pipeline is rejected at compile time (E365), with the available column list
and a did you mean when the name is a near miss. Nothing is renamed silently.
The compiler cannot always see the column set. Inside a composition body the
rows are open by construction, and under on_unmapped: auto_widen a column can
reach the sink through the sidecar without being declared anywhere. There an
item naming an unknown column compiles even when its name resembles a declared
column: spelling similarity cannot prove that a dynamic field is absent.
W365 reports it after the run if no written record supplied it.
What catches the rest is the end of the run: if no record supplied an item’s source column, that item wrote an empty column in every row, and the run reports it as W365, naming the column to correct. An item some records supply and others do not is a sparse column, not a mistake, and is not reported.
Both W365 and W366 are advisory. They print to standard error when the run finishes and do not change the exit code – the file is written and readable either way, and by the time a stream ends the run’s other outputs have already been flushed.
A column absent from the source’s schema: reaches the sink only through the
auto_widen sidecar, which is expanded to top-level columns only under
include_unmapped: true. A mapping: item may name such a column when that
flag is set; under include_unmapped: false it cannot resolve and is rejected
at compile time.
An empty block – mapping: {} or mapping: [] – is rejected (E364): it
declares an output with no columns. To write every upstream column, remove the
mapping: key rather than emptying it.
Writing the block as a YAML map instead of a sequence is rejected (E364);
the message prints your own block already rewritten. Run clinker explain --code E364
for the migration, and read the direction note there before pasting: releases
before this one documented output_name: source_field but executed the
reverse, so the rewrite swaps each pair’s two sides to preserve what the
pipeline was actually writing.
Excluding fields
Remove specific fields from output:
exclude: [internal_id, _debug_flag, temp_calc]
exclude: matches incoming column names, and runs before mapping:. Two
consequences:
- The columns that survive keep their relative order. Upstream
a, b, c, dwithexclude: [b]writesa, c, d. - Naming a column that a
mapping:item also produces is not a conflict – the exclusion removes the upstream column of that name and leaves the mapped one standing. That is the fix for the two-columns-under-one-header collision above:- sold_to: customer_idwithexclude: [sold_to]writes onesold_tocolumn, carryingcustomer_id’s value.
Excluding a column a mapping: item reads is a different matter, and is
rejected (E364): the exclusion removes the column before the item can read
it, so the item could never resolve.
Header control (CSV)
include_header: true # Default: true
Set to false to omit the CSV header row.
Null handling
preserve_nulls: false # Default: false
When false, null values are written as empty strings. When true, nulls are preserved in the output format’s native null representation (e.g., null in JSON).
Rounding decimals to a declared scale
A Sink node’s optional schema: may declare a column type: decimal with a
scale. A decimal value landing in that column is rounded to the declared
number of fractional places on write, using banker’s rounding — the same
boundary contract a decimal source column applies on read.
schema:
- { name: dept, type: string }
- { name: total, type: decimal, scale: 2 }
- { name: average, type: decimal, scale: 2 }
Decimals compute at full precision inside the pipeline (division and avg
keep every digit), so a declared output scale is how you pin a computed result
to fixed places at the sink: avg(amount) over 1.00, 1.00, 2.00 writes 1.33
into a scale: 2 column, while sum(amount) — already at scale 2 — stays
4.00. This works for every format (CSV, JSON, fixed-width); an output column
with no declared scale, or an output with no schema: block at all, keeps the
full-precision value. Only decimal values in decimal-declared columns are
affected — no other type is coerced. See Decimal — arithmetic
rules for the full boundary-contract model.
The same rounding applies to a Sink node declared inside a
composition body. When its schema: names an
external .schema.yaml file, the path resolves relative to the composition
file’s own directory (not the invoking pipeline’s).
Output format options
CSV
- type: sink
name: csv_out
input: processed
config:
name: csv_out
type: csv
path: "./output/result.csv"
options:
delimiter: "|"
delimiter is a single byte on the wire, so it must be exactly one ASCII
character (for example ,, |, or \t). An empty, multi-character, or
non-ASCII value is rejected at plan validation rather than silently truncated
to its first byte.
JSON
- type: sink
name: json_out
input: processed
config:
name: json_out
type: json
path: "./output/result.json"
options:
format: ndjson # array | ndjson
pretty: true # Pretty-print JSON
array(default) – writes a single JSON array containing all records.ndjson– writes one JSON object per line.
JSON numbers cannot represent non-finite floats; a record carrying NaN or
an infinity fails the write with a JSON error instead of silently becoming
null. See JSON Format.
XML
- type: sink
name: xml_out
input: processed
config:
name: xml_out
type: xml
path: "./output/result.xml"
options:
root_element: "data"
record_element: "row"
attribute_prefix: "@" # emit @-prefixed fields as XML attributes
Fields whose final path segment carries the attribute_prefix (default
@, matching the XML source option) are emitted as XML attributes of
their enclosing element, so attribute fields read from an XML source
round-trip. See XML Format for details.
Fixed-width
- type: sink
name: fw_out
input: processed
config:
name: fw_out
type: fixed_width
path: "./output/result.dat"
schema: "./schemas/output.schema.yaml"
options:
line_separator: crlf
Fixed-width output requires a format schema defining field positions and widths. Fields land at their declared byte ranges with gaps space-filled — see Fixed-Width Format for the layout semantics.
EDIFACT
- type: sink
name: edi_out
input: messages
config:
name: edi_out
type: edifact
path: "./out/result.edi"
options:
interchange: ["UNOA:1", "SENDER", "RECEIVER", "240101:1200", "REF1"]
message_type: "ORDERS:D:96A:UN"
write_una: false
segment_newline: true
The EDIFACT writer reconstructs the interchange envelope around emitted
records, recomputing the UNT/UNZ control counts and echoing the
control references, and release-escapes any element data that carries a
service character. The UNB header comes from interchange (literal
elements) or interchange_from_doc (echoed from a $doc section). An
interchange is a single envelope, so an edifact output cannot be
combined with a split: block — the combination is rejected at
config-validation time (E323). See EDIFACT Format for the
full option reference, the record schema, and the round-trip semantics.
HL7 v2
- type: sink
name: hl7_out
input: messages
config:
name: hl7_out
type: hl7
path: "./out/result.hl7"
options:
file_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "FILE7"]
batch_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "BATCH3"]
segment_newline: true
The HL7 writer re-emits the MSH and body segments from the record
stream, escaping any field data that carries a delimiter character (| →
\F\, ^ → \S\, and so on). When a file_header (or
file_header_from_doc) or batch_header is configured the writer wraps the
messages in an FHS..FTS file or BHS..BTS batch and recomputes the
closing BTS/FTS counts. A batch/file envelope is a single structure, so
an hl7 output cannot be combined with a split: block — the combination
is rejected at config-validation time (E339). See
HL7 v2 Format for the full option reference, the record schema,
the MSH off-by-one, and the round-trip semantics.
Sort order
Sort records before writing:
sort_order:
- { field: "name", order: asc }
- { field: "amount", order: desc, null_order: last }
| Sort option | Values | Default |
|---|---|---|
order | asc, desc | asc |
null_order | first, last, drop | last |
first– nulls sort before all non-null values.last– nulls sort after all non-null values.drop– records with null sort keys are excluded from output.
drop is available only on a Sink sort_order, because only a Sink’s
ordering decides which records are written. A Source sort_order, a Cull or
Reshape order_by and a Transform analytic_window.sort_by only order
records, so they accept first and last and reject drop when the
pipeline is planned, pointing at a filter not <field>.is_null() Transform
instead. When CXL cannot name the field as it is (a name with a space, a CXL
keyword such as filter, or a flattened Address.City), the error prints no
CXL and points at the Source schema’s
source_name
rename, which gives the column a name the filter can use.
drop removes records, so a run using it writes fewer records than it read
and that is not a fault. A missing column counts as a null key: a record that
never carried the sort field is dropped the same as one carrying an explicit
null. With several dropping fields, a record is excluded if any of its keys
is null, and counts once however many of them are.
The excluded records are counted, separately from records_dlq and from
filter losses, so a short output can be attributed rather than guessed at. A
run that dropped any reports the number on completion:
1234 record(s) excluded by null_order: drop
and the same number is written as records_null_dropped in the metrics spool
when one is configured (see Metrics).
Under fan-out the count is per exclusion, not per source record: two Sinks
that each declare a dropping sort_order each drop their own copy, so one
source record excluded at both counts twice — the same multiplicity
records_written carries. Subtracting this from records_total is therefore
only sound on a pipeline with a single dropping Sink.
Nothing else records these records. Unlike a DLQ entry, a dropped record
leaves no artifact to inspect afterwards – if you need to see which records
were removed rather than only how many, route them out with a filter before
the sort instead of declaring drop.
Shorthand: a bare string defaults to ascending with nulls last:
sort_order:
- "name"
- { field: "amount", order: desc }
A Sink sort_order materializes all records that reach that terminal and
re-establishes one order across them, including records from several physical
files or Merge inputs. The guarantee is exactly the authored field sequence,
direction, and null placement. drop is also part of the authored contract:
records with a null in a sort key do not reach the writer.
For a split Sink, that order is global across the complete numbered split set,
not restarted independently inside each file. Clinker sorts and applies
null_order: drop before it rotates the writer. Each numbered file is therefore
a contiguous slice of the one ordered sequence, and concatenating the files in
sequence-number order recovers that sequence. Dropped rows do not count toward
max_records, max_bytes, or the resulting number of split files.
The sort is stable. Equal authored keys retain their upstream arrival order within a given execution path, and the same path produces the same bytes in resident and forced-spill operation. Clinker does not add a source-row, filename, or canonical-record tie-breaker. If upstream strategies can produce different arrival orders, equal-key rows have no cross-strategy relative-order promise. Author enough fields for a total business order before using an exact byte comparison; otherwise validate the decoded record multiset and aggregate values instead.
How values are ordered
Every sort uses one rule for comparing two values: a Sink or Source
sort_order, a Cull or Reshape order_by, a window’s sort_by, and the
check that verifies a Source’s declared order. The rule does not depend on the memory limit, so a sort that
spills to disk writes the same records in the same order as one that fits in
memory.
- Nulls are placed only by
null_order: first, last or dropped. A missing column counts as a null. - Values of one type order naturally: numbers by value, strings by UTF-8
code point (no locale collation, so
"Z"sorts before"a"),falsebeforetrue, and dates and datetimes chronologically. A leap-second datetime sorts with the instant one second later that has the same fraction. - Integers, floats and decimals compare by their exact value, not through
a rounded floating-point copy. The integer
1, the float1.0and the decimal1.00are equal. The integer9007199254740993sorts after the float9007199254740992.0, although the two round to the same float. The decimal0.1sorts before the float0.1, whose exact binary value is slightly larger. - Zero has one position:
-0.0and0.0are equal. - NaN is one value, whatever its sign. It sorts after every number,
infincluded, in ascending order, and so comes first in descending order. - Values of different types, which a column can hold when an expression’s branches produce different types, order by type: booleans, then numbers, then strings, then dates, then datetimes, then arrays, then maps.
Values the rule calls equal keep their arrival order, as described above, at every memory limit.
Physical writer boundaries
Planning derives the writer boundary from the finalized graph, not from how many Sink nodes appear in the YAML. The same ordering promise is therefore enforced at every physical byte-emission path:
- ordinary single-file and split-file record output;
- one output per physical source file;
- reconstructed envelope output per document;
- document DLQ output after the whole document is known to be clean;
- deferred output per correlation group; and
- incremental streaming output.
Complete-population modes apply the exact authored key at their population
boundary using the same bounded-memory spill path. Incremental streaming
cannot truthfully promise a terminal whole-population sort. If a finalized
output mode is incompatible with an authored sort_order, planning rejects
the pipeline instead of weakening the promise. The diagnostic names the
Sink, mode, authored keys, and last reordering stage, and includes a corrected
sort_order form that can be pasted into the source or upstream node.
File splitting
Split output into multiple files based on record count, byte size, or group boundaries:
- type: sink
name: split_output
input: processed
config:
name: split_output
type: csv
path: "./output/result.csv"
split:
max_records: 10000
max_bytes: 10485760 # 10 MB
group_key: "department" # Never split mid-group
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true # Repeat CSV header in each file
oversize_group: warn # warn | error | allow
Split configuration fields
| Field | Required | Default | Description |
|---|---|---|---|
max_records | No | – | Soft record count limit per file |
max_bytes | No | – | Soft byte size limit per file |
group_key | No | – | Field name – never split within a group sharing this key value |
naming | No | "{stem}_{seq:04}.{ext}" | File naming pattern. It must contain exactly one {seq:NN} token, where NN is a decimal width from 1 through 20. {stem} is the base name and {ext} is the file extension. |
repeat_header | No | true | Repeat CSV header row in each split file |
oversize_group | No | warn | What to do when a single key group exceeds file limits |
At least one of max_records or max_bytes should be specified for splitting to have any effect.
The naming grammar is strict: {stem}, {ext}, and the one required
{seq:NN} token are the only placeholders. Unknown placeholders, a bare
{seq}, non-numeric or out-of-range widths, and duplicate or missing sequence
tokens are rejected during configuration validation. For example,
{stem}_{seq:03}.{ext} renders sequence 7 as 007.
For formats whose output wraps the whole file in framing – a JSON array or an XML root element – each split file is a complete, independently valid document: the framing is closed at rotation and reopened for the next file.
When the Sink also declares sort_order, splitting happens after the complete
Sink population has been ordered and null-key drops have been applied. Segment
1 receives the first surviving records, segment 2 the next records, and so on.
The files are individually ordered and together form one ordered sequence when
read by sequence number; split rotation never starts a new independent sort.
Oversize group policies
warn(default) – log a warning and allow the oversized file.error– stop the pipeline.allow– silently allow the oversized file.
When group_key is set, the split point is the first group boundary after the threshold is reached (greedy). Without group_key, files are split at the exact limit.
Streaming writes after an interleave Merge
When a single Sink sits directly after a Merge with mode: interleave whose inputs are all Sources, records are written to disk as they arrive rather than being buffered until the merge finishes. This keeps memory flat and lets a slow writer naturally pace the upstream readers.
- type: source
name: src_a
config: { type: csv, path: a.csv, schema: ... }
- type: source
name: src_b
config: { type: csv, path: b.csv, schema: ... }
- type: merge
name: merged
inputs: [src_a, src_b]
config:
mode: interleave # required
- type: sink
name: out
input: merged
config:
name: out
type: csv
path: out.csv
This is automatic — there is no setting to enable it. It applies only to this
exact shape: one interleave Merge of Sources feeding one non-splitting Sink,
in a pipeline without correlation keys. Any other topology buffers as usual.
Both paths preserve the same record multiset and writer semantics, but an
unseeded interleave does not promise one exact cross-input row sequence. Add a
Sink sort_order with a total business key when exact bytes are required.
Complete example
- type: sink
name: department_reports
input: enriched_employees
config:
name: department_reports
type: csv
path: "./output/employees.csv"
# `include_unmapped: false` makes the mapping the whole output: these four
# columns, in this order, and nothing else. Without it every unlisted
# upstream column would still be appended after them, and an `exclude:`
# would be needed to keep any of them out.
include_unmapped: false
mapping:
- "Employee ID": employee_id
- "Full Name": display_name
- department
- "Annual Salary": salary
include_header: true
sort_order:
- { field: "department", order: asc }
- { field: "display_name", order: asc }
split:
max_records: 5000
group_key: "department"
naming: "employees_{seq:03}.csv"
repeat_header: true
CSV Format
CSV is the default file format. The reader decodes each CSV row (including
quoted multiline cells) into a record whose fields are matched
positionally (or by header name) against the source’s declared
schema:; the writer reverses the process. CSV pairs with the file
transport — see Source Nodes for the transport /
format split and the schema rules every source shares.
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
schema:
- { name: order_id, type: int }
- { name: customer_id, type: int }
- { name: amount, type: float }
- { name: order_date, type: date }
options:
delimiter: "," # default ","
quote_char: "\"" # default "\""
has_header: true # default true
encoding: "utf-8" # default "utf-8"
Declared types
CSV has no native numeric or date types: decoding first produces text cells.
Source ingestion then parses and validates each declared column against
schema: before buffering, sorting, or downstream CXL evaluation. With
type: int, the cell 42 reaches a Transform as an integer; with
type: string, the cell 0042 remains text, including its leading zeroes.
No explicit CXL cast is needed for a column already declared with its intended
type. Casts such as .to_int() are useful when a pipeline deliberately keeps
text in its schema and converts it later.
A value that fails its declared type rejects the row under the configured error policy; it never silently falls back to text. See Declared-type failures.
Options
All CSV options are optional. With no options: block, Clinker uses
standard RFC 4180 defaults.
| Option | Default | Description |
|---|---|---|
delimiter | , | Field separator, exactly one ASCII byte. Set to \t for TSV, ; for semicolon-delimited exports. |
quote_char | " | Quote character that escapes delimiters and newlines inside a field, exactly one ASCII byte. |
has_header | true | When true, the first line names the columns and is consumed, not emitted. When false, fields bind to schema: positionally. |
encoding | utf-8 | Character set each field — including the header row — is decoded through. Supported values are utf-8 (the default) and iso-8859-1 (aliases latin-1, latin1). See Encoding. |
delimiter and quote_char are each a single byte on the wire, so each must
be exactly one ASCII character. An empty, multi-character, or non-ASCII
value (for example "||" or "→") is rejected at plan validation — it is
never silently truncated to its first byte.
Encoding
The reader decodes every field through the source’s declared encoding:
utf-8(the default) is strict — a byte sequence that is not valid UTF-8 fails the run loudly rather than substituting replacement characters, so a mis-declared encoding is caught instead of silently corrupting data.iso-8859-1(Latin-1; also spelledlatin-1orlatin1) maps each byte0xNNto codepointU+00NN, so high bytes such as0xE9(é) from legacy exports decode correctly.
The same options.encoding setting is available on a CSV Sink. It applies to
headers, body fields, joined values, embedded JSON cells and reconstructed
envelope rows. UTF-8 is the default on both sides. Names are case-insensitive;
hyphens, underscores and spaces are ignored. UTF8, ISO8859_1, latin1,
Latin-1 and l1 resolve to the corresponding canonical spelling.
Latin-1 is true ISO-8859-1: bytes 0x80..0x9F are the corresponding control
codepoints, not punctuation from another code page. Characters above U+00FF
(for example €) cannot be written in Latin-1 and cause an error; there is no
replacement character or fallback. UTF-8 input strips its leading UTF-8 BOM;
Latin-1 preserves an initial EF BB BF sequence as the three ordinary characters
. CSV grammar is parsed from bytes before either header or body text is
decoded.
An unsupported encoding is rejected during configuration admission with a
precise error naming the value and a supported correction. Charset is part of
the semantic plan identity; changing input or output charset changes that
identity. Aliases and an omitted or explicit UTF-8 default resolve identically.
Malformed UTF-8 in a header or body cell is an input-data error
(source.data.invalid), including when the header is read to discover columns.
The closed encoding policy across formats is:
| Format | Encoding policy |
|---|---|
| CSV, X12 | Authored options.encoding: UTF-8 or true ISO-8859-1. |
| JSON, XML, fixed-width, SWIFT MT | UTF-8 only; no authored encoding override. |
| EDIFACT | In-band repertoire: UNOA/UNOB ASCII, UNOC ISO-8859-1, UNOY UTF-8; no authored override. |
| HL7 | In-band repertoire: blank/ASCII means ASCII, UNICODE UTF-8 means UTF-8; no authored override. |
Single-schema and multi-record CSV sources support the same two encodings, including textual column headers, discriminator fields, body cells and declared envelope sections. Each file applies its own leading-BOM rule. A UTF-8 BOM is recognized only at the start of that file; the same bytes inside a field are data.
Header handling
With has_header: true, the header row’s names bind input columns to the
schema: entries — column order in the file may differ from the schema.
With has_header: false, binding is strictly positional, so the schema
order must match the file’s column order.
Input columns the schema does not name are governed by the source’s
on_unmapped policy, the same as every other
format.
Multi-value cells (split_values)
A CSV cell holds one string, but that string may pack several values behind a
delimiter (1,a;b;c). Declare the column multiple: true and add a
split_values entry naming the field and its delimiter, and the reader parses
the cell into an array:
- type: source
name: orders
config:
name: orders
type: csv
path: ./orders.csv
split_values:
- { field: tags, delimiter: ";" }
schema:
- { name: order_id, type: string }
- { name: tags, type: string, multiple: true }
tags reads as ["a", "b", "c"]. An empty cell is an empty array; a cell with
no delimiter is a one-element array; each element is coerced to the column’s
declared type:. A quoted cell is unquoted first, so a delimiter inside the
quotes is not a boundary. A multiple: true column with no covering
split_values entry is rejected at compile
(E361).
Multi-record CSV sources reject split_values and multiple: true columns
with E358
and E361; these input options require a single-schema source.
See split_values
in the Source reference for the full grammar.
Writing CSV
On output, the writer emits one row per record with cells in the
output schema’s column order — the same order as the header row —
regardless of how an upstream node ordered the record’s fields. An
output-schema column the record does not carry emits an empty cell,
the same as an explicit null; with include_unmapped: false a record
field the output schema does not name is not written. See
Sink Nodes for header control, field mapping,
and null handling.
- type: sink
name: export
input: orders
config:
name: export
type: csv
path: ./out/orders.csv
options:
encoding: iso-8859-1
Each output operation is prepared completely before any of its bytes reach the destination. The first body row and its automatic header are one operation. An unrepresentable cell therefore writes neither a partial row nor a stray header, and leaves previously accepted bytes and format state intact. Explicit document start and end operations are prepared separately. An I/O failure during delivery can leave a destination prefix and prevents further writes; successful preparation alone is not a delivered row. File publication is a separate contract described in Storage & Spill Location.
CSV quoting needs each complete cell. Its encoding workspace is admitted against the run’s resource budget before allocation and released after that cell; there is no authored field-size or record-size ceiling. A cell that cannot fit fails as a resource error rather than being truncated or dropped. Joined or embedded JSON cells account for the rendered text and encoded bytes while both are live. The complete operation also needs storage for its prepared bytes: it stays in memory unless an explicit spill location is available. Spill does not eliminate the minimum memory needed for a cell, policy, or schema mapping. See Memory Tuning for the accounting boundary and output preparation for storage and cleanup behavior.
Split output captures only the header actually delivered by a successful
operation. With repeat_header: true, later files replay those same names;
with repeat_header: false, only the first file emits the automatic header.
include_header: false suppresses that header throughout. Reconstructed
envelope rows do not become an automatic column header.
An empty stream emits no automatic header. A single empty or null cell is written
as "" followed by a newline; adjacent empty cells are separated by the delimiter.
Reading an ordinary empty cell yields an empty string, so CSV does not distinguish
an empty string from a null unless the pipeline applies its own schema policy.
Writing multi-value cells (join_values)
A multiple: field is joined into one delimited cell on write — the write-side
inverse of split_values. The default needs
no configuration: values join with ;, and a value that itself contains the
delimiter is a hard error rather than a cell that would split back wrongly.
The planner carries the exact output-facing multiple: true column set through
mapping and exclusion into the writer. An array reaching any other CSV column
is rejected as a routing/type-contract error rather than joined implicitly.
- type: sink
name: report
input: orders
config:
name: report
type: csv
path: ./out/report.csv
join_values:
- tags # delimiter ";", on_conflict: error
- { field: notes, delimiter: "|", on_conflict: escape, escape: "\\" }
A field with no join_values entry still joins, with the defaults. An entry
overrides, per field:
delimiter— the separator written between values (default;).on_conflict— what to do when a value contains the delimiter:error(default) — reject the record with the field and element position, preserving its original value for the DLQ, rather than emit a cell that splits back wrongly. This is what makes a defaulted delimiter safe. Undererror_handling.strategy: continue, the offending record goes to the dead-letter queue (categorymulti_value_join_collision) and the run continues; underfail_fastit aborts. The exception is a pipeline where any Source declaresdlq_granularity: document: there the collision fails the run (see Not covered, #933).escape— prefix each delimiter (and each escape character) inside a value withescape(default\), so a matchingsplit_valuesescape:recovers the original. Lossless.delimiterandescapemust each be a single character.encode_json— encode the whole field as an embedded JSON array, recovered by a matchingsplit_valuesjson: true. Preserves every value’s text exactly, including ones carrying the delimiter, quotes, or newlines — nothing is lost or mis-split. (A decimal/date/datetime element serializes as its JSON string form and reads back as a string, re-typed by the column’s declaredtype:, the same round trip every CSV cell takes.)
An empty field emits an empty cell — and, under the delimited policies (error,
escape), a single empty-string value [""] emits an empty cell too, which
reads back as zero values: the delimited encoding cannot tell an empty field from
one empty value. Use encode_json when that distinction matters. A single
non-empty value emits that value with no delimiter. The joined cell is quoted by
the normal CSV rules when it contains the field delimiter, a quote, or a newline.
Declaring join_values on a non-CSV output is rejected at compile
(E362).
Round trip. on_conflict: escape and encode_json are recovered exactly by
a matching source split_values entry:
# write side
join_values:
- { field: tags, on_conflict: escape, escape: "\\" }
# read side (a later pipeline)
split_values:
- { field: tags, escape: "\\" }
Header widening under auto-widen
When auto_widen is in effect and the Sink leaves
include_unmapped at its default of true, different records can carry
different carried-along columns. The header must still be shared by every
row, so Clinker widens it to the union of every record’s columns in
first-seen order: a column that first appears on a later record still gets
its own header slot, and the earlier rows write an empty cell for it. This
pre-scan runs on the buffered output path, where the record batch is
materialized.
An output that streams under a bounded-memory budget cannot pre-scan the
whole batch: a CSV output fused directly after a Merge/Transform, a
single-branch Route, a streaming-strategy Aggregate, or the probe side of
a hash-build-probe Combine, or one reconstructing an envelope (which
suppresses the shared header and streams a headerless body), commits its
columns to the first record. A later record
carrying a column that first record lacked then fails the run with a
SchemaDrift error naming the column, rather than silently dropping it.
Declare the column in the source or output schema: so every record carries
it, or route to a self-describing format (JSON / NDJSON / XML). A
reconstruct_envelope CSV output therefore requires a stable body shape —
every record must carry the same columns.
Multi-record files (header / trailer / body)
Some CSV exports interleave multiple record types in one file — a
header row, many body rows, and a trailer row — each distinguished by a
discriminator column. Declare these with a map-form schema: carrying a
discriminator: and a records: list, instead of the single column-list
schema:. Each record type names its tag (the discriminator value that
identifies it) and its own columns:; the discriminator field must sit at the
same column in every type (usually the first). The reader derives the runtime
superset schema (a lead record_type column plus the union of every record
type’s columns) automatically.
- type: source
name: payments
config:
name: payments
type: csv
path: "./data/payments.csv"
schema: # one multi-record schema (map form)
discriminator: { field: rec_type } # the physical column carrying the type tag
records:
- { id: header, tag: H, columns: [ { name: rec_type, type: string }, { name: batch_id, type: string } ] }
- { id: detail, tag: D, columns: [ { name: rec_type, type: string }, { name: id, type: int }, { name: amount, type: int } ] }
- { id: trailer, tag: T, columns: [ { name: rec_type, type: string }, { name: count, type: int } ] }
structure:
- { record: trailer, count: count } # validate T's count against the body count
envelope:
sections:
head:
extract: { record_type: H } # the H record type surfaces as $doc.head.*
fields:
batch_id: string
The reader emits one record per CSV row, including quoted multiline cells,
on a single superset schema
whose lead record_type column carries the matched type’s id. A
downstream Route discriminates on that column.
Rows of different record types may carry different column counts (ragged
rows) — the reader validates the column count per record type, not
file-wide. A textual column-header row is skipped when has_header is
true (the default), so a leading record_type,name,amount line is not
mistaken for a record of an unknown type. Each declared field honors its
own type / trim / pad, the same as a single-record CSV field.
- Header rows declared as an
envelope:section via therecord_typeextract surface as$doc.<section>.*and are excluded from the body stream (see Envelopes & Document Context). - Trailer rows named by a
structure:constraint are validated as they stream — the declaredcountfield is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body row after the trailer is rejected as content past the document close. - Blank lines (empty or whitespace-only, common after concatenation) are skipped rather than parsed.
- An unknown discriminator value (a tag no
records:entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run underfail_fast; undercontinuewith the default record granularity it dead-letters only that physical row and continues with the next row; underdlq_granularity: documentit condemns the whole file. A record-grained DLQ row carries a JSON array of the decoded CSV cells in_cxl_dlq_source_record, preserving empty cells without guessing which declared layout the unknown tag meant.
JSON Format
The JSON reader turns a JSON document into a record stream. It handles
three physical shapes — a single array of objects, newline-delimited
objects (NDJSON), or a wrapper object that nests the records under a
path — and auto-detects the shape when you do not declare it. Each object
is matched against the source’s declared schema:; see
Source Nodes for the shared schema and transport
rules.
- type: source
name: events
config:
name: events
type: json
path: "./data/events.json"
schema:
- { name: event_id, type: string }
- { name: timestamp, type: date_time }
- { name: payload, type: string }
options:
format: object # array | ndjson | object (auto-detect if omitted)
record_path: "data" # dot-separated keys to the records array
max_index_bytes: 64MB # cap on retained envelope sections (optional)
Text encoding
Input is strict UTF-8. One leading UTF-8 BOM is accepted and removed at each
physical file open, including an envelope pre-scan. UTF-16 and UTF-32 BOMs and
malformed UTF-8 are rejected; convert the file to UTF-8 before running it.
There is no JSON encoding option or lossy fallback. A later invalid file does
not erase records already delivered from earlier files. Validation follows the
reader: it does not promise to discover malformed bytes beyond what it reads.
Output is UTF-8 without a BOM. Ordinary format: ndjson always writes one
compact object followed by exactly one LF, including the last record;
pretty: true does not expand ordinary NDJSON across lines. pretty still
controls array output and reconstructed envelope documents. Envelope framing
is documented separately.
Physical shapes
format | Layout |
|---|---|
array | The file is a single JSON array of objects. |
ndjson | One JSON object per line (newline-delimited JSON). |
object | A single top-level object; record_path locates the records array within it. |
If format is omitted, Clinker auto-detects the shape from the file
content. Declare it explicitly when the file is large enough that you want
to skip detection, or when an object wrapper needs a
record_path.
record_path
record_path is a dot-separated path of object keys, descended from the
document root. data.rows selects the array at {"data": {"rows": [ … ]}}, and
each of its elements becomes one record. This is the canonical statement of the
grammar; other pages link here rather than restate it.
The rules, in full:
- No
$.root marker. It is not JSONPath. Writedata.rows, not$.data.rows. Only the exact leading$.is rejected, so a key that merely starts with$($schema.rows) is still addressable. - No leading
/. A leading slash is how a JSON Pointer is anchored;record_pathis already anchored at the document root. - No empty segments — no doubled separator (
data..rows) and no trailing one (data.). - Omitting
record_pathentirely lets the reader auto-detect the document shape. That is not the same asrecord_path: "", which is a path naming a key called “” and is rejected.
A value breaking any of these fails at compile time with E363, before any input is opened. The diagnostic names the corrected path where one can be derived.
record_path takes precedence over format:. When both are declared the
reader navigates the path and streams the array it finds, whatever format:
says — so pair record_path with format: object (or leave format: off).
Declaring format: ndjson alongside a record_path does not read NDJSON.
Because a JSON key may contain any character, the two rejected prefixes give up
a sliver of addressing: a top-level key literally named $ followed by a nested
key, and a top-level key whose name starts with /, are not reachable through
record_path.
Nested arrays
JSON records frequently embed arrays — line items on an invoice, tags on a product. Three source-level declarations decide what happens to them, all documented on the Source Nodes page:
split_to_rowsfans the array out to one record per element.mode: extract(the default) hoists an object element’s keys onto the output record;mode: splitkeeps the record shape, flattening the element back under the field name (orders.id). An array of scalars keeps the value under the field’s own name under both modes.- A schema column declared
multiple: truekeeps the array as an array, and normalizes a lone scalar into a one-element array so the column’s shape never depends on what a particular document happened to carry. split_valuesparses a delimited string cell into several values.
A record whose declared field holds an empty array, is explicitly null, or
carries no such field at all, is preserved by default — keep_empty defaults to
true, and setting it to false drops such a record. An explicit null is how
many producers write “no value”, so it counts as no occurrence rather than one;
for the same reason a multiple: true column holding an explicit null stays
null rather than becoming [null].
A field that IS present but holds a single object or scalar rather than an array
is one occurrence, projected exactly as a one-element array would be. Producers
routinely unwrap a lone element, so a feed where some documents carry
"line_items": [{…}, {…}] and others carry "line_items": {…} fans both out
the same way and every output record ends up with the same columns. The XML
reader, where a document cannot express the difference at all, already behaved
this way.
Two declared fan-out fields apply in declaration order and multiply. A nested
pair (orders then orders.items) produces the two-level expansion when the
outer entry declares mode: split:
split_to_rows:
- { field: orders, mode: split }
- { field: orders.items, mode: split }
Under mode: extract the outer entry lifts the occurrence’s keys to the top
level, which removes the orders.items path the inner entry addresses — so that
pairing is rejected at compile (E358) rather than silently fanning out only
one level. A duplicated field is rejected too.
Set source-level max_output_rows_per_input: N to bound the cumulative product
without materializing it. The reader emits the first N rows in stable order,
then a first attempted row above the ceiling routes the original JSON object to
the DLQ as expansion_limit_exceeded; 0 or omission is unlimited. See
Source Nodes -> split_to_rows
for the complete error and fail_fast behavior.
Flattened-name collisions
The reader dissolves nested objects into dotted keys, so {"a": {"b": 1}}
becomes the field a.b. When two distinct keys flatten to the same name —
for example a nested {"a": {"b": 1}} alongside a literal {"a.b": 2} in the
same record — only one value could survive, and keeping one while dropping the
other is silent data loss. The reader refuses the record instead, naming the
colliding field. This mirrors the XML reader’s treatment of a repeated element:
both formats now fail loud on an undeclared collision rather than one keeping the
first value and the other the last. If the collision is intentional (both values
belong together), declare the column multiple: true to collect them into an
array in document order; otherwise rename one of the source keys so they no
longer collide. As with XML, detection is per document at read time, so the run
aborts under fail_fast and dead-letters the document under continue
with dlq_granularity: document.
Detection covers two distinct source keys that flatten to the same dotted
name — the nested {"a": {"b": 1}} plus literal {"a.b": 2} case above. It does
not cover a key that is literally duplicated within one JSON object
({"tags": "x", "tags": "y"}): the JSON parser collapses such duplicates
last-wins (keeping "y") before the record reaches collision detection, so that
repeat is silently dropped rather than reported. A collision inside an array
element that a split_to_rows: extract fan-out lifts to the top level (one
element key clashing with a parent field or with another element key) is
likewise not yet detected and still resolves last-wins — tracked by
issue 920.
Bounding envelope retention: max_index_bytes
When a source declares an envelope: and a pipeline reads $doc.* paths
from it, the JSON reader runs a streaming pre-scan that walks the document
once and retains only the declared section subtrees — every other key,
including a multi-megabyte body array, is parsed-and-skipped without being
stored. The retained sections live in a bounded document index.
max_index_bytes caps that index. It is charged incrementally as each
section is parsed, so even a single oversized declared section aborts
mid-parse (naming the section and the cap) rather than risking an
out-of-memory failure. It accepts a decimal size string (64MB, 500KB)
or a bare byte count; optional, defaulting to 64MB. Only the declared
sections a program actually reads are retained, so envelope metadata sits
far below this ceiling in practice — the cap exists to convert an unbounded
mistake into a clear error. See
Document Envelope Context for
the full model.
Non-finite floats
JSON numbers cannot represent NaN, +infinity, or -infinity. Writing a
record (or an envelope section field) that holds a non-finite float to a
JSON output fails with a bounded field diagnostic, rather than silently
substituting null — a substituted null would be indistinguishable from
a genuine source null on read-back. Filter such records or replace the
value in a transform before the JSON output.
Writing JSON
A JSON output writes one object per record, in schema-column order, either as a
single array (format: array, the default) or one object per line
(format: ndjson).
- type: sink
name: enriched
input: processed
config:
name: enriched
type: json
path: "./output/enriched.json"
preserve_nulls: false # omit null columns; native map/array nulls remain values
options:
format: ndjson # array | ndjson
pretty: false # indentation for arrays or reconstructed envelopes
Dotted column names become nested objects
A column name containing a . expands back into nesting, the same way the
XML writer expands one into nested elements. Columns
Address.City and Address.State write as one object:
{"Address":{"City":"Boston","State":"MA"},"name":"Ada"}
This is what makes a JSON-in / JSON-out pipeline reproduce its input shape: the
reader flattened {"Address":{"City":…}} into the column Address.City, and
the writer puts it back. It applies to every JSON output, with no option to turn
it off — a flag would mean the same column name meant different things at
different outputs.
Three points follow from the rule, all shared with the XML writer and specified in full on Field Paths:
- Grouping. Columns sharing a prefix collect into one object, positioned where that prefix first appeared, even when the schema interleaves them.
- Absent children. Under
preserve_nulls: falsea null column emits no key, and an object whose every descendant is absent emits no key at all rather than an empty{}— so it reads back as the absent column it stands for. - Values are untouched. A column holding a map or an array still serializes as that map or array. Expansion adds structure above the value, never inside it.
Native map and array values
CXL can construct maps, arrays, and array comprehensions directly; a JSON output writes those values recursively as native objects and arrays. Map key insertion order and array item order are preserved:
emit payload = {
customer: customer_name,
items: [{sku: item.sku, quantity: item.quantity} for item in line_items],
}
The output contains "payload" as an object with an "items" array. Native
maps and arrays are values, so the default preserve_nulls: false does not remove null map
entries or array items inside them; it controls null output columns and null
leaves created by dotted-column expansion.
JSON and XML share one neutral-map key grammar. After CXL has decoded the string
literal, an unescaped key is its ordinary logical spelling. Exactly one leading
backslash marks a literal reserved-looking key only in these three forms:
\@name, \#text, or \\name. The neutral decoder removes that one marker.
In CXL source, where the string literal itself must escape the backslash, write
"\\@name", "\\#text", or "\\\\name". Other leading-backslash
forms are non-canonical and fail. JSON assigns no structural role to @name or
#text, but it still uses this same decoder: "\\@literal" writes the JSON
key "@literal".
Static and computed map keys follow the same rule. Two authored spellings that decode to the same logical key are a duplicate and fail rather than selecting a winner. Before writing any bytes for a record, the writer validates the entire neutral tree. Scalars have depth zero and each map or array adds one container; depth 64 is accepted and depth 65 is rejected. A failed nested value therefore cannot leave a partial JSON record in the output.
This recursive behavior is native to JSON/NDJSON and XML. Flat, positional, and
message formats do not silently turn a map or array into JSON text. Use an
explicit encoding the destination format declares—such as join_values for a
multi-value flat field—or reshape the value before that output. Without one,
the structured value is rejected before bytes for that record are written.
Keeping a literal . in a key
To emit a key that genuinely contains a ., escape the separator in the column
name. The column a\.b writes the single key "a.b":
schema:
- { name: "a\\.b", type: string } # emits {"a.b": …}
- { name: "a.b", type: string } # emits {"a": {"b": …}}
A [ in a column name is currently literal but reserved; write \[ if you want
it to stay literal indefinitely. See
Field Paths for why.
Note that this is a write-side escape. A source key that literally contains
a . still arrives from the reader as an unescaped column name (the
Flattened-name collisions section above covers
what the reader does), so {"a.b": 1} read and written back comes out as
{"a": {"b": 1}}. Closing that is tracked by
issue 920.
Column names that cannot both be written
Two columns can describe places that cannot both exist in one object — a column
a holding a value alongside a column a.b that needs a to be an object.
Rather than keep one and drop the other, the writer refuses the whole column set
before emitting that record, identifying the offending column and the path
rule to correct. Field Paths
lists every clashing shape.
A column name carrying a malformed escape — a \ that is not part of \.,
\[, or \\, as in a column literally named C:\temp — is refused the same
way, with escape guidance; write C:\\temp for a literal backslash.
Preparation and empty output
Each complete output operation is prepared within the run’s finite resources before delivery. An invalid value or resource refusal during preparation writes none of that operation. A destination failure during delivery can leave a prefix; the writer then stops and never retries or finalizes on teardown. Earlier delivered records remain delivered. See output preparation for spill, cancellation and the separate file-publication boundary.
A CLI source with no body records never opens its native writer and produces an
empty file, including when envelope reconstruction is selected. This differs
from explicitly finalizing a library array writer, which emits [] and an LF.
An explicitly opened empty envelope document retains its declared framing and
has a body count of zero.
XML Format
The XML reader selects record elements by a slash-separated path of element
names and maps each one onto the source’s declared schema:. Child elements
bind to fields by name; attributes bind under a configurable prefix. Namespaces
are stripped by default so schema field names stay clean. See
Source Nodes for the shared schema and transport
rules.
- type: source
name: catalog
config:
name: catalog
type: xml
path: "./data/catalog.xml"
schema:
- { name: product_id, type: int }
- { name: name, type: string }
- { name: price, type: float }
options:
record_path: "catalog/product" # slash-separated element path
attribute_prefix: "@" # prefix for XML attribute fields
namespace_handling: strip # strip | qualify
max_index_bytes: 64MB # cap on retained envelope sections (optional)
Text encoding
XML input must contain valid UTF-8. One leading UTF-8 BOM is removed on every
physical file open, including the envelope pre-scan. UTF-16/32 BOMs are rejected.
An XML declaration may omit encoding or declare UTF-8; other encodings and
conflicting declarations are rejected. Convert such input to UTF-8 rather than
adding an encoding option. Names, attributes, text and CDATA are never decoded
with replacement characters.
Validation applies to bytes the reader consumes. A pre-scan may find a late error before any body record is delivered; a streaming body can have already delivered earlier records. Each subsequent file establishes its own BOM and declaration policy. Metadata adjacent to a selected record does not become an extra row, and repeated matching containers preserve body order and empty rows.
Output is UTF-8 without a BOM or XML declaration. See native document boundaries for envelope and empty-output behavior.
Options
| Option | Default | Description |
|---|---|---|
record_path | — | Slash-separated path of element names selecting the elements that each become one record — see record_path. Omitted, every top-level element becomes one record. |
attribute_prefix | @ | Prefix that distinguishes an element’s attributes from its child elements when both map to schema fields. |
namespace_handling | strip | strip removes namespace prefixes from element and attribute names; qualify preserves the namespace-qualified names. |
max_index_bytes | 64MB | Cap on the bytes the envelope pre-scan retains while extracting declared $doc.* sections. |
record_path
record_path is a slash-separated path of XML element names, matched level
by level starting at the document element. catalog/product selects every
<product> that is a child of the document element <catalog>. This is the
canonical statement of the grammar; other pages link here rather than restate
it.
The rules, in full:
- The path is already anchored at the document element, so it carries no
leading
/. WriteOrders/Order, not/Orders/Order. - No
//. It is not XPath: there is no descendant-or-any-depth step. Name every enclosing element. - No empty segments — no doubled separator (
Orders//Order) and no trailing one (Orders/). - No XPath predicates, axes, or wildcards (
product[@id='7'],child::product,*). Select the elements by path and filter the records in a transform. - Every segment must be a legal XML element name. Under
namespace_handling: qualifyelement names keep their prefix, so a qualified segment (ns:Order) is allowed and is what matches; under the defaultstripthe prefix is gone and the segment is the local name. - Omitting
record_pathentirely makes every top-level element one record. That is not the same asrecord_path: "", which is a path naming an element called “” and is rejected.
A value breaking any of these fails at compile time with E363, before any input is opened. The diagnostic names the corrected path where one can be derived.
record_path and xml_path root differently
The envelope option
extract: { xml_path: … }
is also a slash-path over XML, but it tolerates a leading / — /doc/Head
is its documented form. record_path rejects one.
The two are separate grammars addressing separate things: xml_path locates a
single envelope section anywhere in the document, record_path locates the
record elements the body streams. They are deliberately not aligned — writing
record_path: "/catalog/product" is an error, and writing
xml_path: "/doc/Head" is correct.
Truncated input
A truncated XML document — one whose input ends before an open element’s
closing tag — is rejected with a format error rather than yielding the
partial fields read so far. This holds for a record cut off mid-element, a
skipped-over sibling subtree cut off before it closes, and an envelope
section cut off during the pre-scan (which then attaches no $doc
metadata). This matches the general contract that a
truncated stream always aborts
rather than silently dropping data.
Writing XML
The XML writer expands dotted field names to nested elements, by the same rule
the JSON writer expands them into nested objects —
grouping, ordering, absent-child pruning, and the \. escape for a literal dot
are all specified once on Field Paths. What is specific
to XML is layered on top of that decoding, not instead of it.
The attribute_prefix convention applies in reverse: a field whose final path
segment carries the prefix is emitted as an XML attribute of its enclosing
element instead of a child element. A top-level @id attaches to the record
element’s start tag; a nested Address.@type attaches to the <Address>
element. Records read from an XML source therefore round-trip —
<Record id="7"><name>A</name></Record> reads and writes back unchanged, and
the writer never emits an @-named element.
Each decoded segment must also be a well-formed XML Name, so a segment that
begins with a digit or contains a space is rejected. A literal dot survives —
. is a legal XML name character, so a column declared a\.b emits the single
element <a.b> rather than nesting.
Two column names that cannot both be expanded — a column a holding a value
alongside a column a.b needing a to be a container — are refused before any
byte of the record is written, naming both columns. Earlier versions emitted two
sibling <a> elements for that column set, which this reader then refused on
the way back in.
- type: sink
name: xml_out
input: processed
config:
name: xml_out
type: xml
path: "./output/result.xml"
preserve_nulls: false # omit null elements; null attributes always omit
options:
root_element: "Root" # default Root
record_element: "Record" # default Record
attribute_prefix: "@" # matches the source-side prefix
| Option | Default | Description |
|---|---|---|
root_element | Root | Name of the document root element wrapping all records. |
record_element | Record | Name of the element emitted per record. |
attribute_prefix | @ | Prefix marking a field as an attribute of its enclosing element. Set it to the same value as the source-side prefix when round-tripping; an empty string disables attribute classification (every field emits as an element). |
Attribute handling details:
- A null attribute field is dropped even under
preserve_nulls: true— a null element round-trips as a self-closing tag, but an attribute has no form that reads back as null. - A field with children nested under an attribute-prefixed segment
(e.g.
@a.b) is rejected with a format error: an XML attribute is a leaf and cannot contain elements. - The attribute name (the segment after the prefix) must be a well-formed
XML name — a letter,
_, or:followed by letters, digits,_,-,., or:(plus the XML 1.0 Unicode name ranges). A name with a space,=, quote,/,>, or a leading digit (e.g.@foo bar,@1st) is rejected with a format error rather than emitting a malformed start tag. Non-ASCII letters are accepted, so an attribute name read from a source document round-trips unchanged. - An element with only attribute fields and no children self-closes:
Address.@typealone emits<Address type="home"/>.
Native map and array values
An element-valued CXL map is written recursively. Ordinary keys become child
elements, an unescaped key beginning with attribute_prefix becomes an
attribute on the current element, and the unescaped key #text becomes text in
the current element. Map insertion order controls text/child order; attributes
are collected onto the start tag. Arrays held under an ordinary key repeat that
key as the element name.
emit payload = {
"@kind": "event",
"#text": "before",
item: [
{"@id": 1, "#text": "alpha"},
{"@id": 2, "#text": "beta"},
],
tail: "after",
}
writes:
<payload kind="event">before<item id="1">alpha</item><item id="2">beta</item><tail>after</tail></payload>
JSON and XML share one neutral-map key grammar. After CXL string decoding,
ordinary keys are unescaped. Exactly one leading backslash marks a literal key
only as \@name, \#text, or \\name, and the neutral decoder removes that
one marker. Because the CXL string literal must encode the backslash too, the
source spellings are "\\@name", "\\#text", and "\\\\name".
Other leading-backslash forms are non-canonical and fail. An escape disables
XML’s attribute or text classification; it does not make the decoded spelling a
legal XML name. For example, a decoded @literal still cannot be an element
name, while JSON can write it as an ordinary object key.
The rules are deliberately strict:
- Attribute and
#textvalues must be scalar or null; maps and arrays there are rejected. - A direct array inside another array is rejected because XML has no child name to repeat. Put the inner array under a map key to supply that name.
- Every decoded ordinary key and attribute name must be a well-formed XML name.
- Static and computed keys use identical decoding. Two spellings that decode to
the same logical key—including attribute-looking or
#textspellings—are a collision and reject rather than selecting a winner. - Scalars have depth zero and each map or array adds one container. Depth 64 is accepted; depth 65, malformed escapes, duplicate logical keys, and invalid names reject the whole record before its first byte is emitted. XML never silently falls back to JSON text.
With the default preserve_nulls: false, null child elements and null array
items are omitted; with it enabled they emit as self-closing elements. Null
attributes are always omitted.
The XML writer is deliberately two-pass per record. Its first borrowed pass validates the complete schema/value shape, XML names, and scalar roles before writing any bytes for that record. Its second pass encodes from the borrowed original record into a private prepared operation. Delivery begins only when the complete operation is ready. Authored strings remain borrowed; other scalars are formatted in a fixed 128-byte stack scratch buffer. The writer does not clone or materialize a second nested tree, and it retains no record values or rendered scalar capacity between calls. Its only memoized preparation heap state is the admitted schema-derived element plan, whose size is independent of record value widths; recursive calls are capped at 64 containers. This is separate from the XML reader’s optional envelope pre-scan described below.
Native recursion belongs to XML and JSON/NDJSON. Flat, positional, and message
formats do not stringify maps or arrays implicitly. Declare an encoding that
the destination supports—such as join_values for a multi-value flat
field—or reshape the value before that output; otherwise the structured value
is rejected before bytes for that record are written.
Writing multi-value fields (repeated elements)
A multiple: field is written as repeated child elements, one per value, in
order — the XML counterpart to the CSV writer’s delimited
join_values cell, and the
write-side inverse of reading multiple: true.
The default needs no configuration:
<Order><id>1</id><tags>a</tags><tags>b</tags></Order>
The planner carries the exact output-facing multiple: true column set through
mapping and exclusion into the writer. A top-level array repeats only for one of
those columns; an array reaching any other XML column is rejected as a
routing/type-contract error rather than treated as an implicit declaration.
Arrays nested inside a map remain part of XML’s native recursive structure. A
field with one value emits exactly one element, byte-identical to a scalar
field’s output; a field with an empty array emits nothing (no element, and no
container even when one is configured); an empty-string value emits a
self-closing item element (<tags/>).
A multiple: column that maps to an attribute field (a column whose name
maps to an XML attribute, e.g. @tags, declared multiple: true) is rejected at
compile with E359 — an XML attribute holds a single
value and cannot repeat, and the writer emits repetition only as child elements.
A runtime array reaching an attribute field is likewise rejected by the writer.
To rename the elements, add a join_values entry — the same block the CSV writer
reads, sharing the field key. The XML writer reads two keys from it and ignores
the CSV-only delimiter / on_conflict / escape:
- type: sink
name: xml_out
input: processed
config:
name: xml_out
type: xml
path: "./output/result.xml"
join_values:
- field: tags
repeat_as: Tag # per-item element name; defaults to the field name
wrap_in: Tags # optional container; omit for bare repeats
repeat_as— the element name emitted per item. Defaults to the field’s own element name.wrap_in— a container element bracketing the repeated items. Omit it for bare repeats with no container.
A scalar value on a field that carries a join_values entry is treated as a
one-element sequence: it receives the same repeat_as / wrap_in naming an
array of length one would, so the emitted shape does not depend on whether a lone
value arrived wrapped ([a]) or bare (a) — mirroring how the reader normalizes
a lone scalar into a one-element array. A field with no entry emits the plain
<field>value</field> element.
The two combine into the four arrangements, with no other key:
repeat_as | wrap_in | Output for tags = [a, b] |
|---|---|---|
| — | — | <tags>a</tags><tags>b</tags> |
Tag | — | <Tag>a</Tag><Tag>b</Tag> |
| — | Tags | <Tags><tags>a</tags><tags>b</tags></Tags> |
Tag | Tags | <Tags><Tag>a</Tag><Tag>b</Tag></Tags> |
repeat_as and wrap_in must each be a well-formed XML name, validated the same
way as the root_element / record_element names. Declaring join_values on an
output format that is neither csv nor xml is rejected at compile
(E362).
Round trip. A document read into a multiple: true column with the default
naming writes back to the identical repeated elements — reading
<Order><id>1</id><tags>a</tags><tags>b</tags></Order> into a tags column and
writing it to an XML output with record_element: Order reproduces the input
byte-for-byte.
Repeated elements
When a record element contains repeated child elements, two source-level declarations decide what happens to them, and both take the flattened dotted field name — see Source Nodes → Multi-value fields for the shared grammar. The XML-specific matching rules are below.
A declared field is the repeated element’s dotted path relative to the record
element — the same form the flattened field names use. For a record element
<Order> containing repeated <Item> children, the field is Item; for
<Order><Items><Item>…, it is Items.Item.
One record per occurrence: split_to_rows
- type: source
name: orders
config:
name: orders
type: xml
path: "./data/orders.xml"
options:
record_path: "Orders/Order"
schema:
- { name: id, type: int }
- { name: "Item.name", type: string }
- { name: "Item.qty", type: int }
split_to_rows:
- field: "Item"
mode: split # one output record per <Item> occurrence
Each output carries one occurrence’s fields plus every field outside the group, duplicated onto each record.
Under mode: split the occurrence’s fields keep their full dotted names
(Item.name, Item.@sku), including the element’s attributes. Under the
default mode: extract the declared field’s prefix is lifted off, so the same
document yields name and qty; a repeated scalar element (<Tag>a</Tag>)
has no remainder to lift and takes the declared field’s last segment, so
Tags.Tag yields Tag under extract and stays Tags.Tag under split.
Lifting a prefix off can land an occurrence’s field on a name a field outside
the group already occupies — <Order><name> alongside <Item><name>. The
occurrence wins: under extract it is the record, so its own field is not
shadowed by the parent it was merged with. Use mode: split when you need both
values, which keeps them at name and Item.name.
A declared position_column wins over any field of that name, inside the
occurrence or outside it. position_column: line_no against an <Item> that
carries its own <line_no> child yields the occurrence’s index, not the
document’s value — you named the column, so the index is what it holds.
An occurrence with no content (<Item></Item>) still emits a record, one
carrying only the fields outside the group. A record with no occurrence of
the element is governed by keep_empty: XML cannot distinguish an empty
repetition from an absent element, and the default keep_empty: true passes
the record through unchanged.
Entries apply in declaration order, so two declared fields multiply. Fields must
name disjoint element groups — a duplicated field, or one extending another
(Item and Item.part) — which is rejected at compile (E358), before the
source opens. The disjointness rule is this reader’s: it assigns each element to
one occurrence group by document position, which is sound exactly when the
declared groups do not nest. A JSON source has no such constraint.
Set source-level max_output_rows_per_input: N to bound that cumulative
product without constructing it in memory. The reader emits the first N rows
in document/declaration order, then a first attempted row above the ceiling
routes the original ordered XML field occurrences to the DLQ as
expansion_limit_exceeded; 0 or omission is unlimited. See
Source Nodes -> split_to_rows
for the complete error and fail_fast behavior.
All occurrences in one field: multiple: true
Declaring a schema column multiple: true collects every occurrence of that
flattened field into one array, in document order, instead of keeping only the
first:
schema:
- { name: id, type: int }
- { name: "Tag", type: string, multiple: true }
<Tag>a</Tag><Tag>b</Tag> yields ["a", "b"], and a single <Tag> still
yields a one-element array. Declaring the flattened children of a repeated
container (Item.name, Item.qty) collects each of them independently.
An empty occurrence — an empty-body <Tag></Tag> or a self-closing <Tag/> — is
a real array element, collected in position as a null:
<Tag>a</Tag><Tag></Tag><Tag>b</Tag> yields ["a", null, "b"] rather than
squeezing the empty element out, so the array round-trips its per-item shape. (An
empty text value reads as null, the same rule the reader applies elsewhere; the
self-closing and empty-body forms behave identically.)
A field cannot be both collected and fanned out: naming a multiple: true
column in split_to_rows is rejected at compile (E358).
A repeated element named by neither a split_to_rows entry nor a multiple:
column is a loud error, not a silent drop. Keeping the first occurrence and
discarding the rest would lose data without warning, so the reader refuses the
record and names the offending field, pointing at the two ways to handle a
repeat on purpose: declare the column multiple: true to collect every
occurrence into an array, or add a split_to_rows entry to fan each occurrence
out to its own record. Detection is per document at read time — a plan cannot
know in advance that a particular document repeats a field. Under the default
fail_fast strategy the run aborts with the diagnostic; under continue
with dlq_granularity: document the offending document is routed
to the dead-letter queue and the run continues.
Delimited text in one element: split_values
split_values parses <Tag>a;b;c</Tag> into ["a", "b", "c"]. The field must
also be declared multiple: true.
Bounding envelope retention: max_index_bytes
When a source declares an envelope: and a pipeline reads $doc.* paths
from it, the XML reader runs an event-driven streaming pre-scan that walks
the document once and retains only the declared section subtrees — every
other element, including a multi-megabyte body, is event-walked and dropped
without being flattened into memory. The retained sections live in a bounded
document index.
max_index_bytes caps that index. It is charged incrementally as each
section is built, so even a single oversized declared section aborts
mid-parse (naming the section and the cap) rather than risking an
out-of-memory failure. It accepts a decimal size string (64MB, 500KB)
or a bare byte count; optional, defaulting to 64MB. Only the declared
sections a program actually reads are retained, so envelope metadata sits
far below this ceiling in practice — the cap exists to convert an unbounded
mistake into a clear error.
The reader holds no whole-document buffer: the body walks the document element-at-a-time, and the envelope pre-scan opens the source a second time to walk it independently — a file source is read twice, never buffered. Peak memory is the bounded section index plus a single live record, not the input size. See Document Envelope Context for the full model.
Preparation and delivery failures
Invalid XML names, illegal XML characters, unsupported nested shapes and resource refusal during preparation leave that operation’s destination bytes unchanged. Diagnostics identify a bounded offending field and the rule to correct. No replacement character or JSON-string fallback is written.
A destination may accept a prefix before failing. After that failure the writer refuses further work, including finalization, and dropping it never retries the prefix. Earlier delivered records remain delivered. The CLI publishes staged files only after successful execution; this is a separate boundary from writer delivery. See output preparation.
A CLI source with no body records produces an empty file because no writer is
opened. Explicitly finalizing an unused library writer instead produces
<Root></Root> with default names. An explicitly opened empty envelope retains
its declared framing with a body count of zero.
Fixed-Width Format
Fixed-width files carry no delimiters — each field occupies a fixed column
range on every line, the layout common to mainframe extracts and legacy
COBOL exports. Because the byte layout is not self-describing, each column in
a fixed-width source’s schema: carries its byte layout (start +
width) alongside its CXL type — one unified declaration drives both the
physical slice and compile-time type checking. See
Source Nodes for the shared transport rules.
- type: source
name: legacy_data
config:
name: legacy_data
type: fixed_width
path: "./data/mainframe.dat"
schema:
- { name: account_id, type: string, start: 0, width: 12 }
- { name: balance, type: float, start: 12, width: 10 }
- { name: status_code, type: string, start: 22, width: 2 }
options:
line_separator: crlf # line-ending style
The column layout
Each column pins itself to a byte range with start (a 0-based offset) and
width (a byte count); end (exclusive) may be given instead of width.
Optional per-column formatting keys — justify, pad, trim, truncation
— control padding and trimming on read and write. Because the same column
declaration carries both the byte range and the CXL type, the physical layout
and the types can never drift apart. A layout shared across pipelines can live
in an external .schema.yaml file referenced by schema: layout.schema.yaml.
Writing fixed-width output
A fixed-width output node declares the same column layout in its
schema:. The writer places every field at its declared byte range —
start plus width (or end), resolved exactly as the reader slices —
regardless of the order the columns are declared in, so a file written
with a schema reads back under that same schema. Byte ranges the layout
leaves undeclared (a gap between fields) are filled with spaces. A column
that omits start continues at the previous column’s end, so a
width-only schema lays its fields out sequentially. Two columns whose
byte ranges overlap have no consistent layout; the writer rejects such a
schema when the output opens, naming both columns and their ranges.
Widths are byte counts, matching how the reader slices. When a value
is longer than its field, truncation cuts at a UTF-8 character boundary at
or below the width, so a multi-byte character is never split: the emitted
cell is always valid UTF-8 of exactly width bytes (it may hold fewer
characters than the width when a trailing multi-byte character does not
fit, with the freed bytes pad-filled). Because padding fills exact byte
counts on write and is stripped one character at a time on read, pad must
be a single-byte (ASCII) character — a multi-byte character (such as ·)
or a multi-character string (such as "0 ") is rejected when the schema is
resolved, on both the read and write sides. An absent or empty pad defaults
to a space. Under truncation: error an over-long value is still a hard error
before any slicing.
truncation: warn truncates and reports it; silent performs the same
truncation and reports nothing. Numeric columns default to error, and
other columns default to warn.
When a run finishes, each output that truncated under warn prints one
W367 warning to standard error, naming every such column with the exact
number of values cut, the longest original value in bytes, the column width,
and the numbers of the first eight output records that were cut (records
count from 1 across everything that output wrote, including every file of a
split output; an output that writes one file per source file numbers its files
one after another in file-path order, except that a file whose writer fails is
numbered when it fails). The warning does not change the exit code. The report never
copies a value, so it cannot leak data into logs, and its memory is fixed by
the schema when the output opens: recording a truncation cannot fail, so
warn never turns an over-long value into a rejected record. A record
that is rejected or not delivered is not counted. Truncations are also
counted by the clinker.sink.truncations metric. Run
clinker explain --code W367 for the fixes.
A type: decimal output column with a scale rounds its values to that
many fractional places on write (banker’s rounding), the same contract a
decimal source column applies on read. This matters here: a computed
decimal such as avg(amount) is a full-precision quotient that would
overflow a narrow numeric field — a hard error under the default
truncation: error for numeric fields — so declaring the field’s scale
shrinks it to fixed places that fit. For example, avg over 1.00, 1.00, 2.00 written into { name: average, type: decimal, scale: 2, width: 6 }
emits 1.33; without the scale the 28-digit quotient overflows the
6-byte field and fails.
Options
| Option | Default | Description |
|---|---|---|
line_separator | lf | Record separator: lf, crlf, or none for consecutive fixed-length records. |
Under lf or crlf, the reader buffers each physical line only up to the
declared record width plus a line-terminator allowance. A physical line wider
than the declared width — trailing filler beyond the last declared field, or a
schema that maps only a prefix of a wider fixed-length record — reads its
declared-width portion; the remaining bytes are discarded up to the next line
terminator and the reader continues with the following record. Because the
buffered portion is capped, a malformed file (a corrupt or missing newline)
cannot grow a single record until end of input: memory stays bounded regardless
of how long the physical line runs. A final line with no trailing newline reads
normally as long as its declared fields fit within the width.
Strict selected-cell input
Field offsets remain physical byte offsets. Each selected cell must be valid UTF-8 within its own byte range; a boundary that cuts through a multi-byte character fails rather than shifting the layout or inserting a replacement character. Undeclared gaps and discarded trailing bytes are not decoded, so invalid bytes in those ignored ranges do not invalidate a selected cell. A leading UTF-8 BOM is removed before applying the layout. There is no charset conversion for fixed-width input or output.
A typed numeric cell that cannot be parsed terminates the read, including
under strategy: continue; that policy does not make numeric parse failures
recoverable. This differs from the multi-record unknown-discriminator
handling described below.
Multi-value cells (split_values)
A fixed-width field holds one value, but its text may pack several behind a
delimiter within the field’s byte range. Declare the column multiple: true
and add a split_values entry, and the reader splits the (padding-stripped)
field text and coerces each part to the column’s declared type:
split_values:
- { field: tags, delimiter: ";" }
schema:
- { name: order_id, type: string, start: 0, width: 4 }
- { name: tags, type: string, start: 4, width: 20, multiple: true }
Because the fixed-width reader is the sole coercion pass, each element is typed
here — a multiple: true int field over 1;2;3 reads as [1, 2, 3]. A blank
field is an empty array; a field with no delimiter is a one-element array. A
multiple: true column with no covering split_values entry is rejected at
compile (E361). The entry is read only on a
single-schema source, not the multi-record reader below.
Repeating groups
A positional repeating group occupies a bounded sequence of fixed-width
occurrence records. Declare the logical column as type: map with
multiple: true, put the per-occurrence byte layout in fields, and give the
group a finite positive occurs.max:
schema:
- { name: account_id, type: string, start: 0, width: 8 }
- name: transactions
type: map
multiple: true
start: 8
fields:
- { name: kind, type: string, start: 0, width: 1 }
- { name: code, type: string, start: 1, width: 2 }
occurs:
min: 0
max: 3
fill: pad
on_overflow: error
count_field:
name: transaction_count
width: 1
Child start offsets are relative to one occurrence. As with top-level
columns, a child may omit start to continue after the previous child. The
count field, when present, is a leading physical cell inside the group’s byte
range; the occurrence payload follows it. It controls how many occurrence
maps the reader returns, but it is not a logical record column and cannot be
named from CXL. In the example, 2A01B02 means two three-byte occurrences
followed by one unused padded slot.
occurs.min defaults to zero and cannot exceed max. Both values count
logical occurrences. max has no default and cannot be inferred: it is what
bounds the resolved record width, the reader’s one-line buffer, and the
writer’s one-record buffer. Missing, zero, overflowing, overlapping, hybrid,
or recursively repeated layouts fail during normal pipeline compilation,
before an input or destination is opened.
fill chooses the physical treatment of unused occurrences:
pad(the default) always reservesmaxoccurrence slots and fills unused child cells with their declared padding. Without a count field, the reader infers only a trailing run of completely padded slots. A populated slot after an empty one is invalid, and the writer rejects an authored occurrence that itself renders entirely as padding because it could not be read back unambiguously. Addcount_fieldwhen an all-empty occurrence is meaningful.shiftomits unused slots. A count field makes the following byte position explicit. Without one, a shifted group must be the last physical field; the reader otherwise cannot distinguish group bytes from the next field.
Overflow is an error by default. The diagnostic names the group, its declared
maximum, and the supplied count without printing record values. To choose
lossy output deliberately, set on_overflow: truncate and also select the
retained end with keep: first or keep: last; keep is invalid with the
default error policy. The writer validates and encodes the complete bounded
record before the first destination write, so an invalid later occurrence or
overflow cannot leave a partial record behind.
Repeating groups and delimiter-packed scalar cells are separate encodings. A
bare multiple: true fixed-width column is not a positional group, and a
split_values entry cannot stand in for fields plus occurs. A fixed-width
sink that receives an array of records must declare the same named positional
group in its output schema.
Scalar document headers and footers
With reconstruct_envelope: true, options.envelope.header_from_doc and
options.envelope.footer_from_doc select arbitrary document sections. Their
values are concatenated in section field order, without field padding,
delimiters, or the body’s byte layout. Strings
are verbatim, booleans use true/false, numbers use their natural scalar
spelling, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null emits
no text. Arrays and maps have no scalar envelope representation and fail
before any bytes from that header or footer reach the destination.
An absent section emits nothing. A present section with no fields still
emits its separator: LF, CRLF, or no bytes under line_separator: none.
Header, body record, and footer are separate complete operations, so a bad
footer cannot undo an earlier successful body. See
fixed-width document output
for section selection and input-extraction limits.
Library output and failures
Direct callers construct FixedWidthEncoder with their columns,
FixedWidthWriterConfig, and finite WriterResources, then wrap it in
PreparedWriter. MemoryOnlyResources::new requires an explicit nonzero
budget. The resource-free writer constructor is unavailable. Read the
truncation account through FormatWriter::truncation_summary() (or
writer.encoder().truncation_summary()): None when nothing was truncated,
otherwise per-column counts, longest lengths, and the first delivered record
numbers.
Preparation validates the complete operation before destination writes.
Preparation failure leaves committed state unchanged and permits a corrected
retry. Once delivery starts, a destination can accept a prefix before
failing; the writer then refuses continuation, including flush, and drop
does not retry. flush_bytes() only drains the destination; flush()
finalizes once and drains. See output preparation
for the separate storage and publication guarantees.
Schema drift
Fixed-width is inert with respect to
auto-widen: because every byte is accounted
for by the format schema, there are no “unmapped” trailing columns to
absorb. The on_unmapped policy has no effect on a fixed-width source.
Multi-record files (header / trailer / body)
Mainframe and banking extracts often interleave multiple record types
in one file — a header line, many body lines, and a trailer line — each
identified by a discriminator at a fixed byte position (commonly the
first character). Declare these with a map-form schema: carrying a
discriminator: byte range and a records: list, instead of a flat column
list. Each record type names its tag (the discriminator value) and its own
byte-positioned columns:; the reader synthesizes the lead record_type
column automatically.
- type: source
name: payments
config:
name: payments
type: fixed_width
path: "./data/payments.dat"
schema: # one multi-record schema (map form)
discriminator: { start: 0, width: 1 } # the type tag occupies byte 0
records:
- { id: header, tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
- { id: detail, tag: D, columns: [ { name: id, type: int, start: 1, width: 5 }, { name: amount, type: int, start: 6, width: 4 } ] }
- { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 5 } ] }
structure:
- { record: trailer, count: count } # validate T's count against the body count
envelope:
sections:
head:
extract: { record_type: H } # the H line surfaces as $doc.head.*
fields:
batch_id: string
The reader streams one record per line on a single superset schema
whose lead record_type column carries the matched type’s id. A
downstream Route discriminates on that column; the
file is never buffered.
- Header lines declared as an
envelope:section via therecord_typeextract surface as$doc.<section>.*and are excluded from the body stream (see Envelopes & Document Context). - Trailer lines named by a
structure:constraint are validated as they stream — the declaredcountfield is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body line after the trailer is rejected as content past the document close. - Blank lines (empty or whitespace-only, common after concatenation)
are skipped rather than rejected; a line whose declared field range is
cut off mid-value is a truncation error, not a silently-partial read.
Field parsing — type coercion, padding strip, justification — is shared
with the single-record fixed-width reader, so a declared
typeparses identically on both paths. - An unknown discriminator value (a tag no
records:entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run underfail_fast; undercontinuewith the default record granularity it dead-letters only that physical line and continues with the next line; underdlq_granularity: documentit condemns the whole file. A record-grained DLQ row carries the line text in_cxl_dlq_source_recordwithout guessing which declared field layout the unknown tag meant.
EDIFACT Format
Clinker reads and writes UN/EDIFACT interchanges alongside CSV, JSON,
XML, and fixed-width. An interchange is a finite file: it opens with an
optional UNA service-string advice and a mandatory UNB header, wraps
one or more UNH..UNT messages, and closes with a UNZ trailer. The
reader streams one segment at a time and the writer reconstructs the
envelope around emitted records. The reader decodes release-escape
sequences into clean data values and the writer re-escapes them on
output, so a reader → writer → reader round-trip preserves the data
values and the envelope control references.
Delimiters and the UNA service string
Each segment is terminated by the segment terminator; within a segment, data elements split on the element separator and components on the component separator. A release character escapes a delimiter that occurs as literal data.
When the file begins with a 9-byte UNA prefix, its six service
characters override the defaults in this fixed order: component,
element, decimal, release, repetition, terminator. When UNA is absent,
the syntax Level-A defaults apply:
| Role | Level-A default |
|---|---|
| Component separator | : |
| Element separator | + |
| Decimal notation | . |
| Release / escape | ? |
| Repetition | space (inactive) |
| Segment terminator | ' |
UNA is optional — a parser that requires it would fail on the common
no-UNA interchange, so Clinker assumes Level-A when it is absent.
Release character
The release character (default ?) marks the following byte as literal
data rather than a delimiter: ?+ is a literal + inside an element,
?' is a literal apostrophe (not a terminator), and ?? is a literal
?. The reader decodes these sequences into clean data values, so a
downstream CSV/JSON sink, a CXL string comparison, or a $doc field sees
O'BRIEN, never the wire form O?'BRIEN. The writer re-escapes on
output: any element value that carries the element separator, the segment
terminator, or the release character is release-escaped automatically, so
a value computed by a Transform or sourced from CSV — never
EDIFACT-escaped to begin with — does not corrupt the interchange. A
reader → writer → reader round-trip therefore preserves the data values
exactly.
The component separator inside an element (e.g. the : in the composite
UNOA:1) is kept as part of the element’s text and is not escaped — the
positional element model works above component resolution, so a composite
element round-trips unchanged. A literal colon in free-text data is the
one ambiguity this introduces: because components are not split into
separate fields, a : in a value re-reads as a component boundary.
Repeating elements ride inside one element string intact and are likewise
never truncated to their first repetition.
Newlines between segments
Some producers insert CR/LF after each segment terminator for readability. Those bytes are insignificant and are stripped between segments; CR/LF that appears inside an element is preserved.
Record shape
Each non-service segment becomes one record under a fixed positional schema:
| Column | Meaning |
|---|---|
seg_id | The segment tag (BGM, NAD, …) |
msg_ref | The enclosing message reference (the UNH element 1) |
msg_type | The message type (the UNH element 2, full composite) |
e01, e02, … | The segment’s positional data elements (release sequences decoded) |
Service segments (UNB, UNZ, UNH, UNT) are consumed by the reader
to drive envelope state and validation — they are never emitted as body
records. The UNH segment that opens a message is emitted as a body
record (its seg_id is UNH), carrying the message reference in msg_ref
and the message-type composite in msg_type, with its full positional
element list also stamped onto e01, e02, … — so any UNH element past
the message type (a common access reference, a message subset
identification, and so on) is available as e03 onward and is
reconstructed on write.
The number of eNN columns is controlled by the source max_elements
option (default 32). A segment carrying more data elements than that is
rejected with guidance rather than silently truncated. Absent trailing
elements read as null.
nodes:
- type: source
name: orders
config:
name: orders
type: edifact
glob: ./inbox/*.edi
options:
max_elements: 48 # widen the positional schema for exotic segments
schema:
- { name: seg_id, type: string }
- { name: msg_ref, type: string }
- { name: e01, type: string }
Envelope sections over UNB
The interchange header UNB is extractable as a document envelope
section, exposing its positional elements to CXL as
$doc.<section>.<field>. Use the segment extract rule with the section
field names matching the positional keys e01, e02, …:
envelope:
sections:
interchange:
extract: { segment: "UNB" }
fields:
e05: string # interchange control reference (UNB element 5)
A Transform can then read $doc.interchange.e05 on every body record.
Only the UNB header is extractable as an envelope section. Trailer
segments (UNT, UNZ) arrive after the body and cannot become $doc
fields without buffering the whole interchange — their control counts
are instead validated inline by the reader (see below). A segment
extract naming any tag other than UNB, or an xml_path / json_pointer
extract against an EDIFACT source, is rejected at startup.
Control-count validation
The reader validates the structural integrity claims carried in the trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):
UNTsegment count — must equal the actual number of segments in the message, counting theUNHandUNTthemselves.UNTmessage reference — must echo the openingUNHreference.UNZmessage count — must equal the actual number ofUNHmessages in the interchange.UNZcontrol reference — must echo theUNBcontrol reference.
Clinker locates the UNB control reference correctly even when the
header carries an empty optional element or transmits the date/time of
preparation as two separate parts rather than one combined element, so
such interchanges still validate and round-trip with the trailer echoing
the correct reference.
A missing UNZ at end of input is a truncation error; content after the
UNZ trailer is rejected — including a lone stray release character with
no following segment, which is treated as unterminated (truncated) content
rather than silently dropped.
Routing a count mismatch to the DLQ
By default a UNT/UNZ count mismatch aborts the run. A source declaring
dlq_granularity: document instead dead-letters the whole interchange / file
to the DLQ — the file’s records become a structural_validation trigger plus
document_rejected collaterals, and no record of the malformed file reaches
the sink. The count is only known at the trailer, after the body has streamed,
so the rejection lands at the sink boundary (no record is written out), not
literally before the first record. The grain is the whole file. The
control-reference echo mismatches and every other corruption (truncation,
post-trailer content) always abort, even under the opt-in. See Malformed
envelopes.
Writing EDIFACT
An EDIFACT Sink node reconstructs the envelope around emitted records.
Records map by the same positional columns (seg_id, msg_ref,
msg_type, eNN); trailing null/empty elements are trimmed so no
fabricated delimiters appear, and a column the writer does not recognize
is an error (project the record to the EDIFACT columns first).
Engine-internal $-namespaced columns are excluded automatically.
nodes:
- type: sink
name: out
input: messages
config:
name: out
type: edifact
path: ./out/result.edi
options:
interchange: ["UNOA:1", "SENDER", "RECEIVER", "240101:1200", "REF1"]
message_type: "ORDERS:D:96A:UN"
write_una: false
segment_newline: true
Output options:
| Option | Meaning |
|---|---|
interchange | Literal UNB data elements (release-escaped as needed on write). |
interchange_from_doc | Name of a $doc section to echo the UNB elements from (round-trip). |
message_type | Fallback UNH message type when a record carries no msg_type value. |
write_una | Emit a leading UNA segment (default false). |
segment_newline | Write a newline after each segment terminator (default true). |
Consecutive records are grouped into UNH..UNT messages on msg_ref
transitions. The writer recomputes the UNT segment count and UNZ
message count, and echoes the message and interchange control references,
so the output passes its own count validation on re-read.
interchange_from_doc echoes the header from a record’s document
context. That context is populated by a source’s UNB envelope section
(declare a segment: "UNB" envelope section on the source) and travels
with every body record through the pipeline — including to a sink that
sits directly downstream of the source with no intervening Transform. The
reader stashes the complete, ordered UNB element list (empty middle
elements included), so the reconstructed header is faithful even when a
middle element is empty and the user declares only the fields they care
about. Supply interchange literal elements instead when the records
have no source UNB section to echo.
Character set
EDIFACT names its body character repertoire in-band: the UNB
header’s syntax identifier (data element S001, component 1 — the UNOA
in UNOA:1) declares the repertoire for the whole interchange. There is
therefore no encoding option on an EDIFACT source or sink; the
reader discovers the repertoire from the UNB and the writer re-derives
it from the UNB it emits, so a read → write round-trip is byte-faithful
without any configuration.
UNB syntax level | Repertoire |
|---|---|
UNOA, UNOB | ASCII; a byte >= 0x80 is an error |
UNOC | ISO-8859-1 (Latin-1), one byte per character |
UNOY | UTF-8; invalid byte sequences are an error |
Decoding stays streaming and per-segment — the interchange is never
buffered whole. The syntax identifier is itself ASCII, so it is read
straight from the raw UNB bytes before any text is decoded; the UNB
and every body segment are then decoded through the negotiated
repertoire. The UNB’s own sender and recipient identification elements
may legitimately carry non-ASCII text under UNOC or UNOY, and they
decode (and re-encode on output) under that repertoire too — so a UNOC
interchange whose header or body carries Latin-1 high bytes (for example
accented characters in a party name) parses without error, surfaces the
correct text in $doc.UNB.*, and round-trips byte-for-byte.
The repertoire is enforced loudly. A UNOA/UNOB interchange whose body
carries a high byte fails (“outside the ASCII repertoire”) rather than
silently reinterpreting it; a UNOY interchange with invalid UTF-8 fails
(“not valid UTF-8”); and a UNB declaring an unsupported syntax level
(UNOD..UNOX) fails at startup with a precise error naming the level,
never falling back to a guessed encoding or substituting replacement
characters. On output the writer encodes element text through the same
repertoire the UNB declares; a character the repertoire cannot
represent (a non-ASCII character under UNOA/UNOB, or a codepoint above
U+00FF under UNOC) is rejected rather than emitted truncated.
Limitations
- Functional groups. A single
UNB..UNZinterchange is supported;UNG/UNEfunctional-group segments are rejected with a precise error. - Output splitting. An interchange is a single
UNB..UNZenvelope and cannot be divided across files. Anedifactoutput combined with asplit:block is rejected at config-validation time (diagnosticE323) rather than emitting a structurally corrupt interchange. - Rare degenerate headers. A few unusual header shapes — those that combine an empty date/time slot with a date-only date/time where the control reference normally sits — may not round-trip byte-for-byte on re-emit. Conformant headers and ordinary variations are unaffected.
X12 Format
Clinker reads and writes ANSI ASC X12 interchanges alongside CSV, JSON,
XML, fixed-width, and EDIFACT. An X12 interchange is a finite file with a
three-tier envelope: an ISA..IEA interchange wraps one or more GS..GE
functional groups, and each functional group wraps one or more ST..SE
transaction sets. The reader streams one segment at a time and the writer
reconstructs the three envelope tiers around emitted records.
The three tiers surface as nested document-context levels: the ISA
interchange becomes the file-level $doc document, and each GS group and
ST set opens a nested level whose $doc sections layer over the
enclosing tiers. A body record therefore sees every enclosing tier’s
fields through one $doc.<section>.<field> lookup.
Delimiters and the ISA header
Unlike EDIFACT’s optional UNA service-string advice, X12 declares its
delimiters in a fixed-length 106-byte ISA header. Three delimiter bytes
live at structural positions within it:
| Role | Source in the ISA |
|---|---|
| Element (data) separator | The byte immediately after the ISA tag |
| Sub-element (component) sep. | ISA16, the last single-byte ISA element |
| Segment terminator | The byte immediately after ISA16 |
The reader reads these three bytes from the header rather than assuming a
fixed delimiter set, so an interchange that uses */:/~,
|/^/newline, or any other producer-chosen delimiters parses correctly.
The ISA13 interchange control number is located as the 13th element of
the header split on the discovered element separator — structurally, not by
an absolute byte offset — so producer padding quirks do not misalign it.
On output, an X12 sink that echoes the header via interchange_from_doc
also adopts the source’s discovered delimiter set, so a reconstructed
interchange keeps the exact element separator, sub-element separator, and
segment terminator bytes it arrived with (see Writing
X12). The literal interchange option keeps the writer’s
*/:/~ defaults.
No escape character
X12 has no release/escape character (EDIFACT’s ? has no X12
equivalent). A data value that contains a delimiter byte is therefore
unrepresentable. On output the writer rejects any element value carrying
the element separator or the segment terminator with a precise error rather
than silently corrupting the interchange; re-encode the value or choose
delimiters the data does not contain.
The sub-element (component) separator inside an element (e.g. the : in a
composite A:B:C) is kept as part of the element’s text and is not split —
the positional element model works above component resolution, so a
composite element round-trips unchanged.
Newlines between segments
Some producers insert CR/LF after each segment terminator for readability. Those bytes are insignificant and are stripped between segments; CR/LF that appears inside an element is preserved.
Record shape
Each non-service segment becomes one record under a fixed positional schema:
| Column | Meaning |
|---|---|
seg_id | The segment tag (BEG, PO1, …) |
group_ref | The enclosing functional-group control number (GS06) |
set_ref | The enclosing transaction set control number (ST02) |
set_type | The transaction set identifier code (ST01, e.g. 850) |
e01, e02, … | The segment’s positional data elements |
The reader stamps both envelope control numbers on every body record:
group_ref from GS06 and set_ref from ST02. The same GS06 value also
surfaces through $doc.functional_group.e06 for expressions that need the
full functional-group envelope (see Envelope sections over the three
tiers).
Service segments (ISA, IEA, GS, GE, SE) are consumed by the
reader to drive the envelope and validation — they are never emitted as
body records. The ST segment that opens a transaction set is emitted
as a body record (its seg_id is ST), carrying the set reference and
type.
The number of eNN columns is controlled by the source max_elements
option (default 32). A segment carrying more data elements than that is
rejected with guidance rather than silently truncated. Absent trailing
elements read as null.
nodes:
- type: source
name: orders
config:
name: orders
type: x12
glob: ./inbox/*.x12
options:
max_elements: 48 # widen the positional schema for exotic segments
encoding: iso-8859-1 # decode body element text as Latin-1
schema:
- { name: seg_id, type: string }
- { name: group_ref, type: string }
- { name: set_ref, type: string }
- { name: e01, type: string }
Character set
X12 carries no in-band element that names the body character repertoire
(unlike EDIFACT’s UNB syntax identifier). Element text is therefore
decoded through the charset the source declares in its encoding option,
defaulting to UTF-8. The ISA header’s control fields are ASCII, so the
declared charset affects only the body element text.
encoding value | Repertoire |
|---|---|
utf-8 (default) | UTF-8; invalid bytes are an error |
iso-8859-1 (aliases latin-1, latin1) | ISO-8859-1 (Latin-1), one byte per char |
Decoding stays streaming and per-element — the interchange is never
buffered whole. An interchange whose body carries Latin-1 high bytes (for
example accented characters in free-text name or address fields) parses
without error once the source declares encoding: iso-8859-1, and the
decoded element text matches the source bytes under Latin-1.
A source that omits the option and meets non-UTF-8 bytes fails explicitly
(“segment is not valid UTF-8”) rather than corrupting the data silently; a
source that declares an unsupported encoding fails at startup with a
precise error naming the value. On output, set the same encoding on the
X12 sink so the round-trip is byte-faithful; a character the chosen charset
cannot represent (for example a non-Latin-1 codepoint under iso-8859-1)
is rejected rather than emitted truncated.
Envelope sections over the three tiers
The interchange header ISA is extractable as a file-level document
envelope section, exposing its positional elements to CXL as
$doc.<section>.<field>. Use the segment extract rule with the section
field names matching the positional keys e01, e02, …:
envelope:
sections:
interchange:
extract: { segment: "ISA" }
fields:
e13: string # interchange control number (ISA13)
The GS functional group and the ST transaction set surface
automatically as the nested $doc sections functional_group and
transaction_set, each keyed by positional eNN elements — no envelope
declaration is needed for them. A Transform on any body record can read
all three tiers at once:
emit isa13 = $doc.interchange.e13 # interchange control number
emit gs06 = $doc.functional_group.e06 # group control number (GS06)
emit st02 = $doc.transaction_set.e02 # set control number (ST02)
Naming and typing the nested levels
The ISA header is declared through envelope: because a bounded pre-scan
resolves it before any body streams. The GS group and ST set exist
only mid-file, so they cannot be declared the same way — instead, name them
and give them a typed field schema under the X12 source’s options. The
reader applies the declaration each time it crosses a group or set
boundary, so $doc.<your-name>.<field> exposes the level’s elements under
the name you chose, coerced to the types you declared — the same way a
declared ISA field is typed and coerced:
type: x12
options:
group_section:
name: functional_group # your choice — the engine reserves no name
fields:
e01: string # GS01 functional identifier code
e06: int # GS06 group control number
set_section:
name: transaction_set # your choice
fields:
e01: int # ST01 transaction-set identifier code
e02: string # ST02 set control number
emit functional_id = $doc.functional_group.e01 # typed string
emit group_control = $doc.functional_group.e06 # typed int
emit txn_type = $doc.transaction_set.e01 # typed int
The two levels are declared independently — name one, both, or neither. A
declared field schema is the contract: only the elements it lists surface
in the typed section, and an element the wire carries but the schema omits
is absent from $doc. An element that cannot coerce to its declared type
(declaring the alphabetic GS01 code as an int, say) fails the run with
a precise error. Omit a level’s declaration and it keeps its default name
(functional_group / transaction_set) keyed by untyped positional eNN
strings, unchanged from before this option existed.
Only the ISA header is extractable through the envelope: block.
Trailer segments (SE, GE, IEA) arrive after the body they close and
cannot become $doc fields without buffering the whole interchange — their
control counts are instead validated inline by the reader (see below). A
segment extract naming any tag other than ISA, or an xml_path /
json_pointer extract against an X12 source, is rejected at startup.
Control-count validation
The reader validates the structural integrity claims carried in the trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):
SEsegment count (SE01) — must equal the number of segments in the transaction set, counting theSTandSEthemselves.SEset control number (SE02) — must echo the openingST02.GEtransaction-set count (GE01) — must equal the number ofSTsets in the functional group.GEgroup control number (GE02) — must echo theGS06.IEAfunctional-group count (IEA01) — must equal the number ofGSgroups in the interchange.IEAcontrol number (IEA02) — must echo theISA13.
A missing IEA at end of input is a truncation error; content after the
IEA trailer is rejected.
Routing a count mismatch to the DLQ
By default a SE/GE/IEA count mismatch aborts the run. A source
declaring dlq_granularity: document instead dead-letters the whole
interchange / file to the DLQ — the file’s records become a
structural_validation trigger plus document_rejected collaterals, and no
record of the malformed file reaches the sink. The count is only known at the
trailer, after the body has streamed, so the rejection lands at the sink
boundary (no record is written out), not literally before the first record.
The grain is the whole file: an SE-level mismatch rejects the entire
interchange, not just that transaction set. The control-number echo
mismatches (SE02/GE02/IEA02) and every other corruption (truncation,
post-trailer content) always abort, even under the opt-in. See Malformed
envelopes.
Writing X12
An X12 Sink node reconstructs the three-tier envelope around emitted
records. Records map by the same positional columns (seg_id, group_ref,
set_ref, set_type, and eNN); trailing null/empty
elements are trimmed so no fabricated delimiters appear, and a column the
writer does not recognize is an error (project the record to the X12
columns first). Engine-internal $-namespaced columns are excluded
automatically.
nodes:
- type: sink
name: out
input: messages
config:
name: out
type: x12
path: ./out/result.x12
options:
interchange:
["00", " ", "00", " ", "ZZ", "SENDER ",
"ZZ", "RECEIVER ", "240101", "1200", "U", "00401",
"000000001", "0", "P", ":"]
group_header: ["PO", "SENDER", "RECEIVER", "20240101", "1200", "1", "X", "004010"]
set_type: "850"
segment_newline: true
Output options:
| Option | Meaning |
|---|---|
interchange | Literal ISA data elements (the 16 fixed-width ISA fields). |
interchange_from_doc | Name of a $doc section to echo the ISA elements from (round-trip). |
group_header | Literal GS01..GS08 elements (GS06 control number recomputed per group). |
set_type | Fallback ST01 set type when a record carries no set_type value. |
segment_newline | Write a newline after each segment terminator (default true). |
encoding | Character set element text is encoded through (default utf-8). |
Consecutive records are grouped into ST..SE transaction sets on set_ref
transitions and into GS..GE functional groups on group_ref transitions
— the group discriminator is the outer-tier analog of set_ref. The writer
recomputes the SE segment count, the per-group GE transaction-set count,
and the IEA functional-group count, and echoes the set, group, and
interchange control numbers, so the output passes its own count validation
on re-read.
Multiple functional groups
A real interchange can carry several functional groups inside one
ISA..IEA. The reader-stamped group_ref makes a direct X12-to-X12
pipeline preserve those boundaries: the writer opens a fresh GS..GE
group every time the value changes, echoes it as the group control number
(GS06/GE02), and recomputes that group’s GE01 transaction-set count.
The configured group_header continues to supply the other GS fields for
each reconstructed group. Records from a source without group_ref still
collapse into one functional group.
nodes:
- type: sink
name: out
input: orders
config:
name: out
type: x12
path: ./out/result.x12
options:
interchange_from_doc: interchange
group_header: ["PO", "SENDER", "RECEIVER", "20240101", "1200", "1", "X", "004010"]
Records that share a group_ref value must arrive consecutively, exactly as
records sharing a set_ref must: the writer streams and closes a group the
moment the discriminator changes, so an interleaved stream would reopen a
group it already closed. Sort upstream by group_ref (then set_ref) when
the record order does not already guarantee it.
interchange_from_doc echoes the header from a record’s document context.
That context is populated by a source’s ISA envelope section (declare a
segment: "ISA" envelope section on the source) and travels with every
body record through the pipeline — including to a sink that sits directly
downstream of the source with no intervening Transform. The reader stashes
the complete, ordered ISA element list together with the delimiter set it
discovered from the header, so the reconstructed header is faithful and the
whole output interchange — header, envelopes, and body — is emitted with the
original element separator, sub-element separator, and segment terminator
rather than the writer’s */:/~ defaults. Supply interchange literal
elements instead when the records have no source ISA section to echo;
that path keeps the default delimiters.
Limitations
- Charset. Element text is decoded through the source’s
encodingoption (UTF-8 by default, or ISO-8859-1). An unsupported or unconfigured-but-non-UTF-8 repertoire is rejected explicitly rather than silently corrupted (see Character set above). - No escape character. X12 has no release mechanism, so a data value that contains a delimiter byte is rejected on output rather than corrupting the interchange.
- Consecutive grouping. Both
GS..GEfunctional groups (group_ref) andST..SEtransaction sets (set_ref) close the moment their discriminator changes, so records sharing a value must arrive together; sort upstream when the source order does not already guarantee it. - Output splitting. An interchange is a single
ISA..IEAenvelope and cannot be divided across files. Anx12output combined with asplit:block is rejected at config-validation time (diagnosticE338) rather than emitting a structurally corrupt interchange.
HL7 v2 Format
Clinker reads and writes HL7 v2.x pipe-and-hat messages alongside CSV,
JSON, XML, fixed-width, EDIFACT, and X12. An HL7 v2 file is a finite stream
of carriage-return-terminated segments. The smallest unit is one message,
which always begins with an MSH (message header) segment; messages may
optionally be wrapped in a BHS..BTS batch and an FHS..FTS file
envelope. The reader streams one segment at a time and the writer re-emits
the segments, optionally reconstructing the batch/file envelopes.
The optional envelope tiers surface as nested document-context levels: an
FHS file header becomes the file-level $doc document, and each BHS
batch and MSH message opens a nested level whose $doc sections layer
over the enclosing tiers. A body record therefore sees every enclosing
tier’s fields through one $doc.<section>.<field> lookup. All tiers are
optional — a bare stream of MSH messages with no batch or file wrapping
is a valid HL7 v2 file.
Delimiters and the MSH header
HL7 declares its delimiters in the MSH header rather than assuming a
fixed set. The byte immediately after the MSH tag is the field separator
(MSH-1), and the four bytes that follow are the encoding characters
(MSH-2), in this fixed order:
| Role | Source in MSH-2 | Conventional byte |
|---|---|---|
| Component separator | first encoding character | ^ |
| Repetition separator | second encoding char | ~ |
| Escape character | third encoding char | \ |
| Sub-component sep. | fourth encoding char | & |
The reader reads these bytes from the header, so a message that uses the
conventional |^~\& or any other producer-chosen delimiters parses
correctly. The segment terminator is always a carriage return
(0x0D) — unlike the field and encoding delimiters, it is never
producer-chosen.
When a file opens with an FHS or BHS header instead of MSH, the
delimiters are read from that header; HL7 requires the file’s FHS/BHS
encoding characters to match its messages’ MSH.
The discovered delimiter set travels with each message through the
pipeline, and an HL7 Output re-emits every message with the set its header
declared — a custom-delimiter file round-trips byte-faithfully, never
silently rewritten to the conventional |^~\& (see
Writing HL7).
The MSH off-by-one
MSH-1 is the field separator, so it is implicit — it never appears as a
data field. When the MSH segment is split on the field separator, the
encoding-characters field (MSH-2) is the first data field. As a result,
for any MSH field number N ≥ 2, the positional column is f<N-1>:
MSH-2 is f01, the sending application MSH-3 is f02, the message
type MSH-9 is f08, and the message control id MSH-10 is f09. The
same off-by-one applies to FHS and BHS.
Escape sequences
Field data escapes a literal delimiter character with an escape sequence
\X\, where X names the delimiter: \F\ field separator, \S\
component separator, \T\ sub-component separator, \R\ repetition
separator, \E\ the escape character itself. The reader decodes these into
their literal data byte, so downstream consumers — CSV/JSON output, CXL
string predicates, $doc fields — see clean data, never the wire escapes.
The writer re-escapes any delimiter byte in field data on output, so the
reader → writer → reader round-trip is byte-faithful.
An application escape the positional reader does not decode (e.g. the
formatting escape \.br\) is kept verbatim rather than dropped, so no data
is lost.
The component, repetition, and sub-component separators inside a field
(e.g. the ^ in a composite PATID^^^HOSP^MR) are kept as part of the
field’s text and are not split — the positional field model works above
component resolution, so a composite field round-trips unchanged.
Newlines between segments
Some producers (and CRLF-normalizing transports) add a line feed after each carriage-return terminator. Those bytes are insignificant and are stripped between segments. A producer that omits the trailing carriage return on the final segment is accepted — that shape is common in practice.
Record shape
Each segment becomes one record under a fixed positional schema:
| Column | Meaning |
|---|---|
seg_id | The segment tag (MSH, PID, OBX, …) |
set_ref | The enclosing message’s control id (MSH-10) |
set_type | The enclosing message’s type (MSH-9, e.g. ADT^A01) |
f01, f02, … | The segment’s positional data fields |
The MSH header segment is emitted as a body record (its seg_id is
MSH), carrying the message’s fields positionally. Batch/file envelope
segments (FHS, FTS, BHS, BTS) are consumed by the reader to drive
the document levels and validate counts — they are never emitted as body
records.
The number of fNN columns is controlled by the source max_fields
option (default 64). A segment carrying more data fields than that is
rejected with guidance rather than silently truncated. Absent trailing
fields read as null.
nodes:
- type: source
name: messages
config:
name: messages
type: hl7
glob: ./inbox/*.hl7
options:
max_fields: 128 # widen the positional schema for large OBX segments
schema:
- { name: seg_id, type: string }
- { name: set_ref, type: string }
- { name: f01, type: string }
Component splitting (optional)
By default a composite field rides inside one fNN column with its
component (^), repetition (~), and sub-component (&) separators intact
— the positional model deliberately works above component resolution. When
you want component-level access (the message code MSH-9.1 vs the trigger
event MSH-9.2) without writing CXL string-splitting downstream, opt one or
more fields into splitting with split_fields. The reader explodes the
named field into structured columns, and an HL7 Output re-assembles the
exact wire field from them, so an HL7→HL7 round-trip stays byte-identical.
options:
split_fields:
- { field: f08, components: 2 } # MSH-9 → message code + trigger
- { field: f03, components: 5 } # PID-3 (CX) → its components
- { field: f04, components: 2, subcomponents: 3 } # also expose sub-components
- { field: f13, components: 1, repetitions: 4 } # repeating field → per-repetition
Each split fixes the column width on three structural axes: components
(required, the ^ axis), subcomponents (default 1, the & axis), and
repetitions (default 1, the ~ axis). The schema stays static — it never
varies with per-record data. A field whose data carries more structure on
any axis than the declaration reserves is rejected with guidance, the same
posture as a max_fields overflow; raise the axis count or leave the field
unsplit.
The exploded columns name the path from the field down to a leaf with the
axis letters r, c, s, all 1-based, eliding the default index (1) on
the repetition and sub-component axes so the common component-only case
stays clean:
| Declaration | Columns for f08 |
|---|---|
components: 2 | f08_c1, f08_c2 |
components: 1, subcomponents: 2 | f08_c1_s1, f08_c1_s2 |
components: 1, repetitions: 2 | f08_r1_c1, f08_r2_c1 |
The verbatim fNN column is replaced by the structured columns. Declare the
exploded column names in the source schema: block (or rely on
on_unmapped) the same way you would any other column:
options:
split_fields:
- { field: f08, components: 2 }
schema:
- { name: seg_id, type: string }
- { name: f08_c1, type: string } # MSH-9.1 message code
- { name: f08_c2, type: string } # MSH-9.2 trigger event
emit code = f08_c1
emit trigger = f08_c2
Splitting respects the escape rules: an escaped separator (e.g. \S\, a
literal ^ in data) is not treated as a component boundary — the split
runs on the raw bytes before the escape decodes, so the literal stays inside
one component. On output the writer re-joins the leaves on the separators
verbatim (never escaping ^/~/&) and still escapes any field-separator,
escape, or carriage-return byte inside a leaf, so the round-trip is
byte-faithful.
Envelope sections over the tiers
The file header FHS is extractable as a file-level document envelope
section, exposing its positional fields to CXL as $doc.<section>.<field>.
Use the segment extract rule with the field names matching the positional
keys f01, f02, … :
envelope:
sections:
file:
extract: { segment: "FHS" }
fields:
f07: string # file name / id (FHS-8 under the off-by-one)
The BHS batch and the MSH message surface automatically as the nested
$doc sections batch and transaction_set, each keyed by positional
fNN fields — no envelope declaration is needed for them. A Transform on
any body record can read all available tiers at once:
emit file_id = $doc.file.f07 # FHS file id (declared section)
emit batch_id = $doc.batch.f07 # BHS batch id (auto section)
emit mtype = $doc.transaction_set.f08 # MSH-9 message type (auto section)
emit ctrl = $doc.transaction_set.f09 # MSH-10 control id (auto section)
Only the FHS header is extractable as a declared envelope section, and
only when the file actually opens with one; a bare MSH-led file has no
file-level envelope, so declaring an FHS section against it is rejected at
startup. Trailer segments (BTS, FTS) arrive after the body they close
and cannot become $doc fields without buffering the whole file — their
counts are instead validated inline by the reader (see below). A segment
extract naming any tag other than FHS, or an xml_path / json_pointer
extract against an HL7 source, is rejected at startup.
Control-count validation
The reader validates the structural integrity claims carried in the batch/file trailers as they arrive, failing the run on a mismatch (a truncation or corruption signal):
BTSbatch message count (BTS-1) — must equal the number ofMSHmessages in the batch. An emptyBTS-1disables the check (the count is optional in practice).FTSfile batch count (FTS-1) — must equal the number ofBHSbatches in the file. An emptyFTS-1disables the check.
Content after the FTS file trailer is rejected. A bare MSH-led file
needs no trailers and validates with no count checks.
Routing a count mismatch to the DLQ
By default a BTS/FTS count mismatch aborts the run. A source declaring
dlq_granularity: document instead dead-letters the whole file to the DLQ —
the file’s records become a structural_validation trigger plus
document_rejected collaterals, and no record of the malformed file reaches
the sink. The count is only known at the trailer, after the body has streamed,
so the rejection lands at the sink boundary (no record is written out), not
literally before the first record. The grain is the whole file. Other
corruption (truncation, post-trailer content) always aborts, even under the
opt-in. An empty BTS-1/FTS-1 still disables the check entirely. See
Malformed envelopes.
Writing HL7
An HL7 Sink node re-emits the MSH and body segments from the record
stream, escaping any field data that carries a delimiter byte. Records map
by the same positional columns (seg_id, fNN); trailing null/empty
fields are trimmed so no fabricated delimiters appear, and a column the
writer does not recognize is an error (project the record to the HL7
columns first). Split-leaf columns (f08_c1, f03_r2_c1_s3) are recognized
too — the writer groups them by field and re-assembles the wire value from
the column names alone, so no output option is needed to round-trip a split
source. The reader-stamped set_ref/set_type echoes and engine-internal
$-namespaced columns are excluded automatically.
The writer re-emits each message with the delimiter set its source header
declared: the reader stamps the discovered field separator and encoding
characters into the message’s document context, and the writer adopts that
set at every MSH — field joining, escaping, and split re-assembly all use
it, so a custom-delimiter message round-trips byte-faithfully. A record
that did not come from an HL7 source (or whose document context was
dropped upstream) is written with the conventional |^~\& set.
nodes:
- type: sink
name: out
input: messages
config:
name: out
type: hl7
path: ./out/result.hl7
options:
file_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "FILE7"]
batch_header: ["^~\\&", "LAB", "HOSP", "EHR", "HOSP", "20240102", "BATCH3"]
segment_newline: true
Output options:
| Option | Meaning |
|---|---|
file_header | Literal FHS fields; opens an FHS..FTS file envelope. |
file_header_from_doc | Name of a $doc section to echo the FHS fields from (round-trip). |
batch_header | Literal BHS fields; wraps the messages in a BHS..BTS batch. |
segment_newline | Write a newline after each segment terminator (default true). |
When a file or batch header is configured, the writer recomputes the BTS
batch message count and the FTS file batch count and emits the trailers
at end of stream, so the output passes its own count validation on re-read.
The MSH-2 encoding-characters field of each header is written verbatim so
the delimiter declaration round-trips. With no header options set, the
writer emits a bare stream of messages with no batch/file envelope.
file_header_from_doc echoes the FHS header from a record’s document
context. That context is populated by a source’s FHS envelope section
(declare a segment: "FHS" envelope section on the source) and travels
with every body record through the pipeline. The reader stashes the
complete, ordered FHS field list, so the reconstructed header is
faithful. Supply file_header literal fields instead when the records have
no source FHS section to echo.
Limitations
- Charset. Field text is decoded as UTF-8. Non-UTF-8 messages are rejected explicitly rather than silently corrupted.
- Positional fields by default; opt-in component columns. Components,
repetitions, and sub-components ride inside one positional
fNNfield verbatim unless that field is named insplit_fields(see Component splitting above), in which case the reader explodes it into structured columns and the writer re-assembles them. The reader always decodes the\X\delimiter escapes regardless. - HL7 v3 and FHIR are out of scope. HL7 v3 is XML (use the
xmlformat) and FHIR is JSON/REST (use thejsonformat); this format handles HL7 v2.x pipe-and-hat encoding only. - Output splitting. A batch/file envelope is a single
FHS..FTSstructure and cannot be divided across files. Anhl7output combined with asplit:block is rejected at config-validation time (diagnosticE339) rather than emitting a structurally corrupt file.
SWIFT MT Format
Clinker reads and writes SWIFT MT (FIN) messages alongside CSV, JSON, XML, fixed-width, EDIFACT, X12, and HL7 v2. A SWIFT MT message is a finite file built from brace-balanced blocks. The reader scans the message into retained fields, then emits one body field at a time and surfaces service blocks as document-envelope sections. The writer preserves field-value bytes while reconstructing the message with CRLF structural separators; it does not preserve the original file’s inter-block whitespace or field separators.
Block structure
A SWIFT MT message is a sequence of top-level blocks, each {n:...} where
n is a numeric block id:
| Block | Role | Contents |
|---|---|---|
1 | Basic header | A fixed header string (application id, BIC, …) |
2 | Application header | Input/output direction, message type, recipient |
3 | User header | Optional; may carry nested {tag:value} sub-blocks |
4 | Message text | The message body — a run of :tag:value fields |
5 | Trailer | Optional; may carry nested {tag:value} sub-blocks |
Unlike the flat delimiter-structured EDI formats (HL7, X12, EDIFACT), SWIFT
framing is brace-balanced rather than terminator-delimited. Because of
that, the nested sub-blocks of blocks 3 and 5 (for example
{3:{108:MSGREF}}) are kept intact inside their parent block rather than
mistaken for top-level blocks.
The -} text-block trailer
Block 4 is special. Its body is opaque line-structured free text — a field
value (a :77E: envelope, a :79: narrative, an :86: information line)
legitimately contains {, }, and even -} as data. So braces inside
block 4 are treated as data, not framing: the block closes only on a
line-anchored -} trailer — at the start of the block-4 body or after
either LF or CR. An interior {,
}, or a -} in the middle of a value is data, not a frame boundary. Both
the framing braces and the closing -} trailer are stripped from the stored
values, so a record carries clean tag/value data.
Whitespace between blocks
Producers insert CR/LF (and the \r\n that separates block-4 fields) for
readability. Inter-block whitespace is insignificant and is skipped; the
line breaks inside block 4 delimit the :tag:value fields.
Record shape
Each :tag:value line of block 4 becomes one record under a fixed
positional schema — the same one-line-one-record model the X12 and HL7
readers use:
| Column | Meaning |
|---|---|
block | The block id the line came from (always 4 for body fields) |
tag | The SWIFT field tag without its surrounding colons (20, 32A, 61) |
value | The field value, with continuation lines folded in verbatim |
A multi-line field (a :50K: ordering-customer block, a :77E: / :86:
narrative) keeps its continuation lines: any line of block 4 that does not
begin a new :tag: is folded into the current field’s value with its line
break preserved — including a blank line inside the value, so a narrative
with an internal blank line round-trips faithfully. LF and CRLF are retained
exactly inside values, including mixed separators, leading/trailing spaces,
and trailing blank continuation lines. Only the separator before the next
field or trailer is structural: one CRLF, LF, or (at the trailer) lone CR is
removed. A data CR immediately followed by the structural LF is necessarily
read as one CRLF separator; these bytes cannot express a separate data CR.
A repeated tag (the :61: / :86: statement lines of an MT940, for instance)
streams as one
record per occurrence, in order.
The service blocks (1, 2, 3, 5) are consumed by the reader to serve
envelope sections and drive the message-level document context — they are
never emitted as body records.
nodes:
- type: source
name: payments
config:
name: payments
type: swift
glob: ./inbox/*.swift
options:
max_fields: 20000 # raise the block-4 field ceiling for large messages
schema:
- { name: block, type: string }
- { name: tag, type: string }
- { name: value, type: string }
The max_fields option caps the number of block-4 field lines a single
message may carry (default 10000). A message exceeding it is rejected with
guidance. This is a field-count guard, not a byte budget for retained reader
data; it does not establish constant-memory input parsing.
Envelope sections over the service blocks
A SWIFT MT message is a single envelope: the four service blocks surface as
file-level $doc sections, exposing each block’s text to CXL as
$doc.<section>.body. Every body record can read the enclosing message’s
headers through a $doc.<section>.body lookup.
Declare the sections on the source with the segment extract rule naming the
block id. The whole block body surfaces under the field name body, because
a SWIFT service block carries free-form text (a header string, nested
{sub:tag} blocks) rather than positional elements:
envelope:
sections:
basic:
extract: { segment: "1" } # block 1, the basic header
app:
extract: { segment: "2" } # block 2, the application header
user:
extract: { segment: "3" } # block 3, the user header (nested sub-blocks kept verbatim)
trailer:
extract: { segment: "5" } # block 5, the trailer
A Transform on any body record can read every enclosing block at once:
emit tag = tag
emit value = value
emit basic_hdr = $doc.basic.body # block 1 header string
emit app_hdr = $doc.app.body # block 2 header string
emit user_hdr = $doc.user.body # block 3 body, nested {108:...} kept verbatim
The section names are entirely your choice — the engine reserves none. A
segment extract may name a block either by its numeric id ("1", "3") or
by the stable default label ("basic_header", "app_header",
"user_header", "trailer"); both resolve the same block.
Block 4 is the message-text body streamed as records, not an envelope
section — a segment: "4" extract is rejected at startup. An xml_path or
json_pointer extract against a SWIFT source is likewise rejected, because
those rules belong to the tree formats.
Malformed-message handling
A structurally broken message fails the run with a precise SWIFT error
rather than producing garbled records:
- Unbalanced brace — a block that never closes (or whose brace depth never returns to zero) is a truncation error naming the offending block.
- Missing
-}trailer — a block 4 that runs to end of input without its-}trailer is a truncation error. - Missing or non-numeric block id — a block without a numeric id after
the
{(or with no:separating the id from the body) is rejected. - Malformed
:tag:valueline — a block-4 line with no second colon closing the tag, or an empty tag, is rejected. - Repeated service block — a second
{1:...}(or any repeated service block) in one message is rejected. - Invalid UTF-8 — block ids and bodies are validated before parsing;
invalid bytes in headers, body text, or trailers are never replaced with
substitute characters. A leading UTF-8 BOM is rejected: after optional
ASCII whitespace, the next byte must be an opening
{.
Initialization succeeds only after the complete message has parsed. On
failure, partial fields and service blocks are discarded; later reads are
terminal and cannot reveal partial records or sections. Malformed messages
terminate execution under both fail_fast and continue.
A header-only message (no block 4, or an empty block 4) is valid: it produces no body records and drains cleanly.
Writing SWIFT MT
A SWIFT Sink node re-emits each block-4
record as a :tag:value line and re-frames the single message envelope
around them: the service blocks 1/2/3 first, then block 4 ({4: … -}),
then the optional block-5 trailer. Block-4 free text is opaque, so values are
written verbatim with no escaping — an interior {, }, a mid-line -},
a folded continuation break, and an interior blank line all reproduce as data.
The writer opens block 4 with {4:\r\n, appends \r\n after each
:tag:value, and closes with -} before any block-5 trailer. Existing LF
and CRLF inside values remain unchanged. Re-reading returns the same field
values even when the original input used LF structural separators.
Records map by the tag and value columns. The block column is the
constant 4 discriminator (an empty block is treated as block 4, so a
Transform that projects only tag/value writes fine); a record carrying a
block other than 4 is rejected, because service blocks are never emitted
as records — they ride the document context.
Strings are written verbatim; other scalar values use their natural display spelling, and null renders empty text. Arrays and maps are rejected. A tag must be nonempty and contain no colon, CR, or LF.
nodes:
- type: sink
name: out
input: messages
config:
name: out
type: swift
path: ./out/message.swift
options:
basic_header_from_doc: basic
app_header_from_doc: app
user_header_from_doc: user
trailer_from_doc: trailer
Each service block is written from a literal body or echoed from a
user-declared $doc section:
| Option | Meaning |
|---|---|
basic_header | Literal block-1 body, written verbatim as {1:<body>}. |
basic_header_from_doc | Name of a $doc section to echo the block-1 body from. |
app_header | Literal block-2 body. |
app_header_from_doc | Name of a $doc section to echo the block-2 body from. |
user_header | Literal block-3 body (nested {sub:tag} content kept verbatim). |
user_header_from_doc | Name of a $doc section to echo the block-3 body from. |
trailer | Literal block-5 body, written after block 4 closes. |
trailer_from_doc | Name of a $doc section to echo the block-5 body from. |
The *_from_doc options name the section the user declared on the source
— the engine reserves no section name. A literal *_header wins over its
*_from_doc companion when both are set (the same rule applies to trailer);
a service block with neither is omitted. The _from_doc echo reads the block
body verbatim from the section’s
body field — the same single-field shape the reader writes — so a SWIFT
source’s service blocks (declared as segment envelope sections) round-trip
unchanged when their section names are passed back to the writer here.
The first successfully delivered record supplies the document service bodies;
its trailer is retained until finalization, even if later records carry a
different context. A selected section must contain body. Missing sections
or fields, structured values, and unbalanced braces in service bodies fail
before delivering that operation, including the first message header.
The document context that carries the $doc sections rides on each body
record, so the *_from_doc echoes require at least one block-4 record to
read from. Explicitly finalizing a library writer with zero records emits
{4:\r\n-}, wrapped only by literal-configured service blocks; document
echoes are skipped because no record supplies context. The CLI opens its
writer lazily: a zero-record run publishes an empty file, even when literal
service options are configured.
Block-4 free text has no escape mechanism, so values are written verbatim.
Almost any value round-trips faithfully, but two shapes are unrepresentable
when the value is built from arbitrary records (CSV/JSON → Transform →
SWIFT): a value whose continuation line — a line after a folded line break —
begins with the block terminator -} after LF or CR (which would re-read as
an early block close), or with a : tag marker after LF (which would
re-read as a spurious field). A bare CR before : is literal data.
The writer rejects such a value with a clear error rather than emitting
silently-corrupt output. Values read from a SWIFT source can never take these
shapes, so a read → write → read round-trip is always safe.
A SWIFT MT message is a single indivisible envelope, so a swift output
cannot be combined with a split: block, including max_records splitting — the pairing is rejected
at config-validation time (diagnostic E342).
Library output and failures
Direct callers use SwiftEncoder::new with a schema, SwiftWriterConfig,
and finite WriterResources, then PreparedWriter::new. A standalone
MemoryOnlyResources provider requires an explicit nonzero budget. The
resource-free writer constructor is unavailable.
The first record’s headers and body are prepared together; finalization
prepares the closing block and trailer together. Preparation failure leaves
the destination and committed state unchanged and permits a corrected retry.
Delivery failure can leave an accepted prefix and poisons continuation.
flush_bytes() drains without closing the message; flush() finalizes once
and drains, with no duplicate trailer on repeated calls. Drop does not finalize
or retry writes. See output preparation
for resource and publication boundaries.
Limitations
- UTF-8 only. SWIFT MT messages are decoded as UTF-8; a non-UTF-8 block id or body is rejected explicitly rather than corrupted silently.
- One message per input. Adjacent complete messages are rejected. Reader materialization keeps its existing allocation behavior; the finite writer budget does not account for the reader’s retained fields or service blocks.
- Field-content parsing. The reader exposes each
:tag:valueline as atag/valuepair verbatim. Parsing a field’s internal structure (the sub-fields of a:32A:value-date/currency/amount, say) is a CXL concern downstream of the source, not a reader responsibility.
Network Sources (REST)
A Source reads from the filesystem by default. To pull records from a
network endpoint instead, declare a transport: block on the Source. The
transport selects where records come from; it sits above the on-disk
type: (the format), which for a REST source still selects how the
response bodies decode.
A network transport is a finite-pull source: it runs on its own
thread, drives a synchronous client to cursor exhaustion, then exits.
There is no daemon, no event loop, and no async runtime — the same single-
process, run-to-drain model as a file pipeline. Finiteness is a hard
property of the reader: a REST source enforces explicit page and record
limits, so an unbounded endpoint cannot keep it running forever. If the
server offers another continuation after max_pages, the source fails
closed instead of reporting a truncated pull as successful completion.
A network source still requires a schema: block. That authored
schema is the row-to-record target: the reader maps each decoded object
onto it, coercing values leniently. A per-row value that cannot coerce is
left unchanged at the reader and routed to the dead-letter queue at the
Transform stage — identical to file-source semantics. A network source
declares no file matcher (path / glob / regex / paths);
declaring one is a configuration error (E219).
Because a network source has no file path, its $source.file provenance
column and the {source_file} output template both resolve to a stable
synthetic identifier, <source:NAME>, where NAME is the Source node’s
name.
REST sources
A rest source issues paginated HTTP GETs against a base URL, decoding
each response body through the declared json or xml format. (Other
formats are rejected with E220 — a REST body is a multi-record
document, not a flat CSV/fixed-width stream.)
nodes:
- type: source
name: orders_api
config:
name: orders_api
type: json
options:
format: array # each page body is a JSON array of objects
transport:
kind: rest
url: https://api.example.com/v1/orders
max_pages: 50 # HARD page cap — required
pagination:
strategy: link_header
auth:
scheme: bearer
token: "${ORDERS_TOKEN}"
schema:
- { name: order_id, type: int }
- { name: total, type: float }
- { name: placed_at, type: date_time }
Pagination strategies
The pagination.strategy selects how the reader advances pages and
detects the last one. max_records bounds emitted records. max_pages
bounds requests and requires the server to reach an actual terminal page;
an offered continuation beyond that bound is an error.
-
none(default) — a singleGET; the body is the whole result. -
offset—?offset=N&limit=L, advancing the offset by the page size each request. The last page is the one that returns fewer rows thanlimit. A page containing exactlylimitrows is not proof of end-of-input, so Clinker must issue one more bounded request to observe a short or empty terminal page. Sizemax_pagesto leave room for that probe; reaching the cap first fails withpage_limit_reachedinstead of accepting a possibly truncated result.pagination: strategy: offset limit: 200 offset_param: offset # optional, defaults shown limit_param: limit -
cursor_token— the reader reads a continuation token from a JSON pointer in each response and sends it back on the next request. Paging stops when the token field is absent or null.pagination: strategy: cursor_token cursor_param: page_token next_token_pointer: /meta/next_page # RFC 6901 JSON pointer -
link_header— the reader follows the URL in the response’s RFC 8288Link: <…>; rel="next"header until no such link is present. Registered relation tokens are case-insensitive, sonext,Next, andNEXThave the same meaning.pagination: strategy: link_header
Continuation and redirect safety
Every server-directed continuation or redirect is resolved against the
effective response URL and normalized before another request is built. Only
the original normalized origin is allowed. Cross-origin targets, HTTPS-to-HTTP
downgrades, malformed or conflicting rel="next" metadata, redirect or
continuation cycles, and traversal beyond the configured bounds fail before a
foreign or repeated request is sent.
Normalization resolves . and .. exactly as RFC 3986 does, and changes
nothing else about the path. An empty segment is a segment: /v1//items/../p
names /v1//p, not /v1/p, because a doubled slash is a different resource on
any server that does not collapse it. A path ending in .. resolves to the
directory above, trailing slash included. Two targets that differ only in an
empty segment therefore stay two pages, and a pull that visits both is not a
continuation cycle.
Link continuation metadata is parsed and authorized only when
pagination.strategy is link_header. Other strategies ignore it because
their continuation authority comes from the configured offset, cursor, or
single-request contract.
A Link header is read as bytes, so a parameter this reader never consults —
a title or a type carrying an accented character, an emoji, or anything
else outside ASCII — does not affect the pull. Only the target inside <…> is
decoded, because only the target has to become a URL: a target that is not
valid UTF-8 is reported as malformed metadata. When a header carries several
comma-separated links and one of them cannot be parsed, the rest are still
read, so a reply naming two different next pages is reported as the conflict
it is rather than as unreadable metadata.
Authentication
auth.scheme selects the credential sent on every request:
-
none(default) — no auth header. -
bearer— sendsAuthorization: Bearer <token>. -
header— sends an arbitrary static header, e.g. an API key.auth: scheme: header name: X-API-Key value: "${API_KEY}"
Corporate proxies
REST sources use the process proxy environment: ALL_PROXY, HTTPS_PROXY, or
HTTP_PROXY (including their lowercase forms), with NO_PROXY/no_proxy for
bypass rules. This lets the standalone clinker run CLI reach a third-party
vendor through a corporate forward proxy without a central orchestrator or a
pipeline-specific proxy key. Keep proxy credentials out of pipeline YAML and
avoid printing credential-bearing proxy URLs.
The current Rust TLS configuration trusts the bundled public Web PKI roots. A proxy that tunnels HTTPS works when the vendor certificate remains visible and chains to those roots. A TLS-inspecting proxy that substitutes a certificate from a private corporate CA is not currently supported by a Clinker trust-store setting, even if that CA is installed in the operating-system store. In that case Clinker fails closed with a TLS/proxy classification; do not disable certificate verification as a workaround.
Reliability and finiteness knobs
| Key | Default | Meaning |
|---|---|---|
max_pages | — | Required. Hard ceiling on pages fetched, regardless of the server. |
max_records | none | Optional hard ceiling on records emitted. |
retries | 3 | Bounded retries on a transient failure (5xx, connect/timeout error, or a transient body-delivery timeout/reset). A 4xx is fatal — retrying cannot help. |
timeout_secs | 30 | Per-request timeout. Bounds in-flight time so an interrupt lands within the shutdown window. |
Request diagnostics report a failure class, attempt number, page number, HTTP status when available, and the query-free target path. Authorization headers, request bodies, response bodies, and URL query values are never included. A proxy or vendor error therefore remains actionable without copying credentials or signed query parameters into logs.
A retryable body-delivery failure discards the partial body and retries the
whole page within the same bounded retry budget. What counts as retryable is
one rule for the whole request, whichever phase observed the failure: a
dropped, reset, or timed-out exchange is retried, and a failure that would
arrive identically on every attempt is not. Body-size violations, TLS
failures, an unroutable URL, a host that does not resolve, and local material
this process cannot read are therefore fatal at once rather than retried —
reported against attempt 1, and without spending timeout_secs once per
remaining attempt on a condition that cannot change. retries is a ceiling on
attempts worth making, not a number of attempts every failure receives.
A partial-page decode failure routes that page’s offending rows to the DLQ per-row, exactly like a file source; it does not abort the pull.
Shutdown
On SIGINT/SIGTERM the reader polls its cancellation handle at each
page boundary and stops cleanly with a normal end-of-input — the same
graceful drain a file source performs. The timeout_secs per-request
bound caps how long a single in-flight request can delay that stop.
A request already in flight when the signal arrives is reported as a cancellation too, not as a failing endpoint — including a page whose body was being read when the connection dropped, which the reader would otherwise have retried. What the signal took away was that retry, so the run’s outcome is the cancellation. An endpoint that would have failed identically however many attempts remained is still reported as that endpoint failure, because no retry was lost: a supervisor re-queuing the batch would only repeat it.
Auto-Widen & Schema Drift
When an input file carries columns the source’s declared schema: block does not name, Clinker decides what to do with them via the per-source on_unmapped policy. The default is auto_widen, which preserves the extra columns end-to-end so schema drift never silently breaks a pipeline. This page covers the three modes, how undeclared columns flow downstream, the output controls, and the related diagnostics.
The three modes
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
on_unmapped:
mode: auto_widen # default; other values: drop, reject
schema:
- { name: order_id, type: string }
- { name: amount, type: float }
auto_widen(default) — undeclared input fields are carried along with each record and re-expanded to top-level columns at the output (when the Sink node’sinclude_unmappedis left at its default oftrue). Nothing is silently lost, and you don’t have to declare every column up front. How a widened column reaches the output depends on the output format — self-describing formats (JSON / NDJSON / XML) carry it per record, tabular CSV widens its header to the union of every record’s columns, and fixed-width (a positional layout with no room for undeclared columns) fails loudly rather than dropping it. See Output controls below.drop— undeclared input fields are silently stripped at read time. The source carries only its declaredschema:.reject— any record carrying a field not in the declared schema fails the source with a diagnostic naming the offending field. The strict choice when unexpected columns should be treated as errors.
CXL expressions can only read fields you declared in schema: — carried-along undeclared fields are not visible to CXL, only to the output. To use an undeclared field in an expression, add it to the source schema:.
How undeclared columns flow downstream
Carried-along columns follow these rules through each node type:
| Node type | Behavior |
|---|---|
| Transform | Passed through unchanged (transforms are row-preserving). |
| Aggregate | Dropped — per-row extra columns have no meaning on a grouped row. To keep one, add it to group_by or emit it explicitly. |
| Combine | The driver’s carried columns ride through; build-side ones are dropped. To keep a build-side field, emit it explicitly in the combine body via <build_qualifier>.<field>. |
| Route / Merge | Passed through. Merge requires every input to share the same on_unmapped policy — mixing fails with E315 (see below). |
| Composition | The body inherits the parent’s carried columns and whatever the body’s last node carries flows back out. |
| Output | Expanded to top-level columns when include_unmapped: true (the default); stripped when false. How a widened column reaches a CSV / XML / fixed-width writer depends on the format — see Schema drift across records. |
Output controls
- type: sink
name: out
input: src
config:
name: out
type: json
path: out.json
include_unmapped: true # default: true
When true (the default), undeclared fields the source carried along are expanded back to top-level columns at the sink — useful for pass-through pipelines where every original column should reach the output. Set include_unmapped: false to write only the columns explicitly emitted upstream.
include_unmapped is independent of include_correlation_keys: each can be set on its own, and include_correlation_keys never surfaces auto-widened columns.
Cross-format flow
Expansion happens before the writer runs, so a CSV source with auto_widen feeding a JSON output with include_unmapped: true produces JSON objects whose keys include both the declared columns and the absorbed ones:
input.csv: id,extra,city
1,foo,Paris
output.json: {"id": "1", "extra": "foo", "city": "Paris"}
Schema drift across records (tabular formats)
Different records can carry different auto-widened columns — for example a Merge of two sources where one carries region and the other category, so region appears only on the first source’s rows and category only on the second’s. Each output format handles that heterogeneity differently:
- JSON / NDJSON / XML are self-describing: each record writes its own keys/elements, so a column present on only some records is simply absent from the others. Nothing is lost.
- CSV needs one header shared by every row. When the output can be materialized (the common buffered path), Clinker pre-scans the batch and writes a header that is the union of every record’s columns, in first-seen order; rows that lack a later-appearing column write an empty cell for it. Nothing is lost.
- On a bounded-memory CSV path — a streaming output fused directly after a
Merge/Transform, a single-branchRoute, a streaming-strategyAggregate, or the probe side of a hash-build-probeCombine, or an envelope-reconstructing output — the writer commits its header to the first record before it has seen the rest, so a union is impossible. A later record carrying a column the header lacks then fails the run loudly with aSchemaDrifterror naming the format and column, rather than silently writing a narrower row. Declare the column in the source (or output)schema:so every record carries it, or route to a self-describing format. - Fixed-width is positional — every column occupies a declared byte range, and there is no room for an undeclared one — so any carried-along column reaching a fixed-width output is a
SchemaDrifterror. Fixed-width sources never auto-widen (see below), so this only arises when a fixed-width output sits downstream of a source that does.
Writer errors on unexpanded columns
CSV and fixed-width writers can only write flat scalar columns. If an
unexpanded map reaches one of those writers, the write fails with an
UnserializableMapValue error naming the format and column. JSON and XML can
write a user-visible nested value natively, although the engine-stamped
$widened sidecar is still expanded or stripped by the Output projection and
is never an author-facing XML structure.
The fix is to either leave include_unmapped at its default of true, so the columns are expanded to top-level before writing, or to convert the value to a scalar in CXL before emitting it. The error message lists both routes.
E315 — Merge inputs must agree on policy
Merge concatenates its inputs positionally, so every input must agree on column shape — same column names, same on_unmapped policy, same correlation_key set. If two upstream sources disagree on whether they carry auto-widened columns (one uses auto_widen, another uses drop / reject), compilation fails:
E315: merge "merged": input schemas disagree on the `$widened` auto_widen sidecar column.
The fix is to set every merge upstream source to the same on_unmapped policy.
Fixed-width sources
Fixed-width sources are positional — the reader only sees the byte ranges your schema defines, so there are never any “extra” columns to absorb. auto_widen has no effect on a fixed-width source; use on_unmapped: drop (or reject) to make that explicit and silence the informational log the engine emits otherwise.
Clinker Expression Language Overview
Clinker Expression Language (CXL) is a per-record ETL expression language. Every program operates on one record at a time, producing output fields, filtering records, or computing derived values.
Programs are sequences of emit, let, filter, and distinct statements
that execute from top to bottom against the current record.
Core expression forms
| Intent | CXL form |
|---|---|
| Produce or rename a field | emit alias = col |
| Keep only matching records | filter condition |
| Combine conditions | and / or / not (keywords) |
| Choose the first non-null value | a ?? b |
| Choose between values | if ... then ... else ... or match { } |
Boolean operators are keywords
CXL uses English keywords for boolean logic, not symbols:
$ cxl eval -e 'emit result = true and false' --field dummy=1
{
"result": false
}
The operators &&, ||, and ! are syntax errors in CXL. Always use and, or, and not.
System namespaces use $ prefix
CXL provides built-in namespaces for accessing pipeline state, metadata, and window functions. All system namespaces are prefixed with $:
$pipeline.*– pipeline execution context (name, counters, provenance) and pipeline-scope declared state$source.*– per-source context and source-scope declared state$record.*– per-record scoped state (travels with the record, never an output column)$window.*– window function calls$vars.*– static, channel-overridable configuration$config.*– a composition’s config parameters, read inside its body (constant-folded per instantiation)
$ cxl eval -e 'emit name = $pipeline.name'
{
"name": "cxl-eval"
}
Compile-time type checking
CXL is statically type-checked, so type errors are caught before any data is processed. Run cxl check to validate a transform before a run. Errors come with source locations and fix suggestions.
$ cxl check transform.cxl
ok: transform.cxl is valid
If there are type errors, the checker reports them with spans:
error[typecheck]: cannot apply '+' to String and Int (at transform.cxl:12)
help: convert one operand — use .to_int() or .to_string()
A minimal CXL program
emit greeting = "hello"
emit doubled = amount * 2
filter amount > 0
This program:
- Emits a constant string field
greeting - Emits
doubledas twice the inputamount - Filters out records where
amountis not positive
Try it:
$ cxl eval -e 'emit greeting = "hello"' -e 'emit doubled = amount * 2' \
--field amount=5
{
"greeting": "hello",
"doubled": 10
}
Statement order matters
CXL statements execute sequentially. Later statements can reference fields produced by earlier emit or let statements:
$ cxl eval -e 'let tax_rate = 0.21' -e 'emit tax = price * tax_rate' \
--field price=100
{
"tax": 21.0
}
A filter statement short-circuits execution – if the condition is false, remaining statements do not run and the record is excluded from output.
Types & Literals
CXL has 10 value types. Every field value, literal, and expression result is one of these types.
Value types
| Type | Rust backing | Description |
|---|---|---|
| Null | Value::Null | Missing or absent value |
| Bool | bool | true or false |
| Integer | i64 | 64-bit signed integer |
| Float | f64 | 64-bit double-precision float |
| Decimal | Decimal | Exact base-10 fixed-point number for money/financials |
| String | FieldStr | UTF-8 text |
| Date | NaiveDate | Calendar date without timezone |
| DateTime | NaiveDateTime | Date and time without timezone |
| Array | OwnedValues | Ordered collection of values |
| Map | OwnedMap | Key-value pairs |
Literal syntax
Integers
Standard decimal notation. Negative values use the unary minus operator.
$ cxl eval -e 'emit a = 42' -e 'emit b = -5' -e 'emit c = 0'
{
"a": 42,
"b": -5,
"c": 0
}
Floats
Decimal notation with a dot. Must have digits on both sides of the decimal point.
$ cxl eval -e 'emit a = 3.14' -e 'emit b = -0.5'
{
"a": 3.14,
"b": -0.5
}
Strings
Double-quoted or single-quoted. Supports escape sequences: \\, \", \', \n, \t, \r.
$ cxl eval -e 'emit greeting = "hello world"'
{
"greeting": "hello world"
}
Booleans
The keywords true and false.
$ cxl eval -e 'emit flag = true' -e 'emit neg = not flag'
{
"flag": true,
"neg": false
}
Dates
Hash-delimited ISO 8601 format: #YYYY-MM-DD#.
$ cxl eval -e 'emit d = #2024-01-15#'
{
"d": "2024-01-15"
}
Null
The keyword null.
$ cxl eval -e 'emit nothing = null'
{
"nothing": null
}
Arrays and comprehensions
Array literals accept full expressions and preserve their written order:
emit values = [order_id, amount * 2, null]
Use one for clause and an optional trailing if to construct an array from
another array:
emit positive_doubles = [item * 2 for item in values if item > 0]
The source must be an array; null and scalar sources are errors. The binding
is local to the item expression and predicate, cannot destructure, and cannot
shadow an input field or surrounding let binding.
Maps
Map literals preserve key insertion order. Bare identifiers and quoted strings are static keys; brackets hold a computed expression whose result must be a non-null string:
emit payload = {
customer: customer_name,
items: [{sku: item.sku, quantity: item.quantity} for item in line_items],
[dynamic_key]: dynamic_value,
}
Duplicate keys are errors, including two differently escaped spellings that decode to the same logical key. Nested maps and arrays are limited to 64 container levels, and CXL construction is limited to 10 MiB per input record; both limits fail the record instead of growing without bound.
Nested keys use one canonical escape grammar. After CXL string decoding, one
leading backslash makes a reserved-looking key literal: \@name, \#text,
or \\name. In CXL source each backslash in a quoted string is itself escaped,
so write "\\@name", "\\#text", or "\\\\name". Other leading-backslash
forms are rejected. Output formats decide how the neutral nested value is
encoded. JSON removes the structural escape when writing the key; XML assigns
roles to the unescaped forms as described in
Writing XML.
The runnable examples/pipelines/nested_values.yaml pipeline sends one
constructed value to both JSON and XML so the two native encodings can be
compared directly.
Schema types
When declaring column types in YAML pipeline schemas, use these type names:
| Schema type | CXL type | Description |
|---|---|---|
string | String | Text values |
int | Integer | 64-bit integers |
float | Float | 64-bit floats |
decimal | Decimal | Exact base-10 fixed-point (money) — see below |
bool | Bool | Boolean values |
date | Date | Calendar dates |
date_time | DateTime | Date and time |
array | Array | Ordered collections |
numeric | Int or Float | Union type – accepts either |
any | Any | Unknown type – no type constraints |
nullable(T) | Nullable(T) | Wrapper – value may be null |
Example YAML schema declaration:
schema:
employee_id: int
name: string
salary: nullable(float)
start_date: date
Type promotion
CXL automatically promotes types in mixed expressions:
Int + Float promotes to Float:
$ cxl eval -e 'emit result = 2 + 3.5'
{
"result": 5.5
}
Null + T produces Nullable(T): Any operation involving null produces a nullable result.
$ cxl eval -e 'emit result = null + 5'
{
"result": null
}
Nullable(A) + B unifies to Nullable(unified): When a nullable value meets a non-nullable value, the result type wraps the unified inner type in Nullable.
The decimal type
float is an IEEE-754 binary float: it cannot represent most base-10 fractions
exactly, so 0.1 + 0.2 is 0.30000000000000004, not 0.3. That rounding is
unacceptable for money. The decimal type is an exact base-10 fixed-point
number — 0.10 + 0.20 is exactly 0.30 — and is the correct type for
monetary amounts, prices, tax, and any figure that must round like decimal
arithmetic on paper.
Declare a decimal column with type: decimal and a scale (the number of
fractional digits). precision (total significant digits) is optional
validation metadata:
schema:
- { name: amount, type: decimal, scale: 2 }
- { name: tax_rate, type: decimal, scale: 4 }
A decimal column parses its raw text into an exact value and rounds off any
excess precision to the column scale (round-half-to-even, the unbiased
“banker’s rounding” used in accounting), so a scale: 2 column stores 2.567
as 2.57. This is one edge of a boundary contract: a declared scale pins
a value to that many places at the boundary it is declared on — a source
column’s scale on read, an output column’s scale on write (see Aggregating
decimals) — while decimals keep full precision inside
the pipeline.
Arithmetic rules
decimal ⊗ decimal → decimal— exact.decimal ⊗ int → decimal— the integer widens exactly, soamount + 1andprice * quantitystay exact decimals.decimal ⊗ floatis a type error. Mixing an exact decimal with a binary float would silently lose precision, so CXL rejects it and asks for an explicit cast. Choose the trade-off deliberately:amount.to_float() * rate— opt into binary float precision.rate.to_decimal() * amount— bring the float into exact decimal math (the float→decimal step is the one acknowledged lossy conversion).
- Division and
avgcompute at full precision — the exact quotient, not a binary-float approximation. Inside the pipeline a computed decimal keeps every digit; it is pinned to a fixed number of places only at a boundary that declares ascale. Declaring the output columntype: decimalwith ascalerounds the value to that many places on write (banker’s rounding), exactly as adecimalsource column rounds on read — soavg(amount)emitted into ascale: 2output column writes1.33, not the full quotient. With no declared output scale the full precision is preserved; use an explicit round in CXL when you need fixed places mid-pipeline.
Comparisons follow the same rule: decimal < int is fine, decimal < float
requires a cast.
The branches of a conditional follow it too. An if, a match or a ?? whose
branches are a decimal and a float does not compile, because its result would
be a decimal on some rows and a float on others:
cannot mix decimal and float without an explicit cast: the branches of this `if` are a decimal (`amount`) and a float (`price`); declare `price` a decimal in its Source schema, `type: decimal` in place of `type: float`, so the branches have one numeric type
The message gives one fix. When the float is a Source column, declare it
type: decimal in its Source schema: the reader parses the column’s text
exactly, so if flag then amount else price is a decimal holding the values
the file holds. (A JSON number read into a decimal column is still parsed
through a float first; see #1299.) When the float is computed
rather than read from a Source column, convert the decimal side instead:
if flag then amount.to_float() else price * 2.0 is a float. Converting a
float with .to_decimal() does not make it exact: the decimal keeps the
float’s binary digits.
Casting
x.to_decimal() converts an int, string, or float into a decimal (try_decimal
is the lenient form that yields null on failure). d.to_int(),
d.to_float(), and d.to_string() convert a decimal back out.
Worked example — an exact invoice total
$ cxl eval -e 'emit total = ("19.99".to_decimal() * 3) + "4.80".to_decimal()'
{
"total": "64.77"
}
19.99 * 3 = 59.97, + 4.80 = 64.77 — exact, with no binary-float drift.
(In a pipeline, declare the source columns type: decimal instead of casting;
JSON output renders a decimal as a scale-preserving string.)
Aggregating decimals
sum, avg, min, max, count, and distinct all work over a decimal
column and stay exact — no binary float ever touches a running total:
sum(amount)returns adecimal: the exact total of the group’s values at the largest scale among them, rounded once (half to even) only when it does not fit a decimal at that scale. The sum of1.00,-1.00and2is2.00in any order, and a total of amounts with two decimal places is exact to the cent.avg(amount)issum(amount) / count(amount), adecimalat full division precision, andweighted_avg(v, w)issum(v * w) / sum(w): the same digits and scale as those expressions give.- Only a group with no non-null value gives null. A decimal total outside the
decimal range, a
weighted_avgwhose weights total zero or whose row product is out of range, and a group holding both decimals and floats each fail the group with anaggregate_finalizeerror that names the fix (see Aggregate functions). min/maxreturn the exact extremum, andcountreturns an integer.- Group-by and
distinctkeys are scale-normalized: two decimals that are numerically equal group together regardless of scale, so2.50and2.5fall in one group. This holds even when the aggregation spills to disk.
To pin an aggregate result to fixed places on the way out, declare the output
column type: decimal with a scale: the value is rounded to that scale on
write (banker’s rounding), so avg(amount) over 1.00, 1.00, 2.00 writes 1.33
into a scale: 2 output column while sum(amount) stays 4.00. An output
column with no declared scale keeps the full-precision quotient. This applies to
every format — CSV, JSON, and fixed-width — because the rounding happens as the
record is projected onto the output. For fixed-width output it is often
required: a full-precision quotient overflows a narrow numeric field, which is a
hard error, whereas the rounded value fits.
weighted_avg also stays exact over decimals: a decimal value or weight (or
both) gives sum(value * weight) / sum(weight) over exact totals, at full
division precision. A zero total weight is an error, as x / 0 is. A decimal
in one position mixed with a binary float in the other is a type error,
matching the decimal ⊗ float arithmetic rule. Declare a float Source column
type: decimal so the value and weight share one numeric domain, or, when the
float is computed, convert the decimal argument with .to_float().
Type unification rules
When two types meet in an expression, CXL coerces them automatically:
- Numbers combine: mixing an integer and a float gives a float (
2 + 3.5is5.5). - A decimal and a float never combine, in an operator or in the branches of an
if,matchor??: declare a float Source columntype: decimal, or convert the decimal side with.to_float()when the float is computed. - Arithmetic and ordering comparisons with
nullgivenull.==and!=never do (null == nullistrue), andand/orgive a definite answer when the other side settles it. See Null Handling. - Mismatched types are an error:
String + Intfails. Convert first with.to_int()or.to_string()so both sides are the same type.
Operators & Expressions
CXL provides arithmetic, comparison, boolean, null coalescing, and string operators. Boolean logic uses keywords (and, or, not), not symbols.
Arithmetic operators
| Operator | Description | Example |
|---|---|---|
+ | Addition (or string concatenation) | 2 + 3 |
- | Subtraction | 10 - 4 |
* | Multiplication | 3 * 5 |
/ | Division | 10 / 3 |
% | Modulo (remainder) | 10 % 3 |
$ cxl eval -e 'emit result = 2 + 3 * 4'
{
"result": 14
}
Multiplication binds tighter than addition, so 2 + 3 * 4 is 2 + (3 * 4) = 14, not (2 + 3) * 4 = 20.
$ cxl eval -e 'emit result = 10 % 3'
{
"result": 1
}
Comparison operators
| Operator | Description | Example |
|---|---|---|
== | Equal | x == 0 |
!= | Not equal | x != 0 |
> | Greater than | x > 10 |
< | Less than | x < 10 |
>= | Greater than or equal | x >= 10 |
<= | Less than or equal | x <= 10 |
$ cxl eval -e 'emit result = 5 > 3' --field dummy=1
{
"result": true
}
Boolean operators
CXL uses keywords for boolean logic. The symbols &&, ||, and ! are not valid CXL syntax.
| Operator | Description | Example |
|---|---|---|
and | Logical AND | a and b |
or | Logical OR | a or b |
not | Logical NOT (unary) | not a |
$ cxl eval -e 'emit result = true and not false'
{
"result": true
}
$ cxl eval -e 'emit result = 5 > 3 or 10 < 2'
{
"result": true
}
Null coalesce operator
The ?? operator returns its left operand if non-null, otherwise its right operand.
$ cxl eval -e 'emit result = null ?? "default"'
{
"result": "default"
}
$ cxl eval -e 'emit result = "present" ?? "default"'
{
"result": "present"
}
Like the branches of an if, the two sides of ?? must not be a decimal and a
float: amount ?? price does not compile. As for if, the fix is to declare
price with type: decimal in its Source schema, or, when the float side is
computed rather than a Source column, to convert the decimal side with
.to_float().
String concatenation
Use .concat() to join strings in compiled pipelines. Numeric + is not a
substitute for explicit text concatenation at the planner’s type boundary.
$ cxl eval -e 'emit result = "hello".concat(" ", "world")'
{
"result": "hello world"
}
Unary operators
| Operator | Description | Example |
|---|---|---|
- | Numeric negation | -x |
not | Boolean negation | not done |
$ cxl eval -e 'emit result = -42'
{
"result": -42
}
Method calls
Methods are called on a receiver using dot notation:
$ cxl eval -e 'emit result = "hello".upper()'
{
"result": "HELLO"
}
Methods can be chained:
$ cxl eval -e 'emit result = " hello ".trim().upper()'
{
"result": "HELLO"
}
Field references
Bare identifiers reference fields from the input record:
$ cxl eval -e 'emit result = price * qty' \
--field price=10 \
--field qty=3
{
"result": 30
}
Qualified field references use dot notation for multi-source pipelines: source.field.
Operator precedence
From highest (binds tightest) to lowest:
| Precedence | Operators | Associativity |
|---|---|---|
| 1 (highest) | . (method calls, field access) | Left |
| 2 | - (unary), not | Prefix |
| 3 | * / % | Left |
| 4 | + - | Left |
| 5 | == != > < >= <= | Left |
| 6 | and | Left |
| 7 | or | Left |
| 8 (lowest) | ?? | Right |
Use parentheses to override precedence:
$ cxl eval -e 'emit result = (2 + 3) * 4'
{
"result": 20
}
Comments
Line comments start with # (when not followed by a digit – digit-prefixed # starts a date literal):
# This is a comment
emit total = price * qty # inline comment
emit deadline = #2024-12-31# # this is a date literal, not a comment
Statements
CXL programs are sequences of statements that execute top-to-bottom against each input record. Statement order matters – later statements can reference values produced by earlier ones.
emit
The emit statement produces an output field. Each emit becomes a column in the output record.
emit name = expression
$ cxl eval -e 'emit greeting = "hello"' -e 'emit doubled = 21 * 2'
{
"greeting": "hello",
"doubled": 42
}
Multiple emit statements build up the output record field by field:
$ cxl eval -e 'emit first = "Alice"' -e 'emit last = "Smith"' \
-e 'emit full = first.concat(" ", last)'
{
"first": "Alice",
"last": "Smith",
"full": "Alice Smith"
}
let
The let statement creates a local variable binding. The variable is available to subsequent statements but is NOT included in the output record.
let name = expression
$ cxl eval -e 'let tax_rate = 0.21' -e 'emit tax = 100 * tax_rate'
{
"tax": 21.0
}
Note that tax_rate does not appear in the output – only emit statements produce output fields.
filter
The filter statement keeps a record only when its condition is exactly true; a condition that is false or null excludes it (see Null in conditions). When a filter excludes a record, remaining statements do not execute (short-circuit).
filter condition
$ cxl eval -e 'filter amount > 0' -e 'emit result = amount * 2' \
--field amount=5
{
"result": 10
}
When the filter condition is false or null, the entire record is dropped and no output is produced.
Filters can appear anywhere in the statement sequence. Place them early to skip unnecessary computation:
filter status == "active"
let discount = if tier == "gold" then 0.2 else 0.1
emit final_price = price * (1 - discount)
distinct
The distinct statement deduplicates records. The bare form deduplicates on all emitted fields. The by form deduplicates on a specific field.
distinct
distinct by field_name
In a pipeline, distinct tracks values seen so far and drops records that have already been emitted with the same key.
emit to a scoped namespace
An emit whose target is a $pipeline.*, $source.*, or $record.* name writes a producer-declared scoped variable instead of an output column. The variable must be listed in the writing Transform’s config.declares: block. $record.* is the per-record store that travels with the record but never serializes as an output column.
emit $record.quality_flag = if amount < 0 then "suspect" else "ok"
Read it downstream via the same namespace:
filter $record.quality_flag == "ok"
See Scoped Variables for the declaration model and the three scopes’ lifetimes.
trace
The trace statement emits debug logging. It has no effect on the output record. Trace messages are only visible when tracing is enabled at the appropriate level.
trace "processing record"
trace warn "unusual value detected"
trace info if amount > 10000 then "high value transaction"
Trace levels: trace (default), debug, info, warn, error. An optional guard condition (via if) limits when the trace fires.
Statement ordering
Statements execute sequentially. A statement can reference any field or variable defined by a preceding emit or let:
$ cxl eval -e 'let base = 100' -e 'let rate = 0.15' \
-e 'emit subtotal = base * rate' \
-e 'emit total = base + subtotal'
{
"subtotal": 15.0,
"total": 115.0
}
Referencing a name before it is defined is a resolve-time error:
emit total = base + tax # error: 'base' is not defined yet
let base = 100
let tax = base * 0.21
use
The use statement imports a CXL module for reuse. See Modules & use for details.
use shared.dates as d
emit fy = d::fiscal_year(invoice_date)
Conditionals
CXL provides two conditional expression forms: if/then/else and match. Both are expressions – they return values and can be used anywhere an expression is expected.
If / then / else
The basic conditional expression:
if condition then value else alternative
$ cxl eval -e 'emit label = if amount > 100 then "high" else "low"' \
--field amount=250
{
"label": "high"
}
The else branch is optional. When omitted, records where the condition is false or null produce null. A null condition takes the else branch when there is one:
$ cxl eval -e 'emit bonus = if score > 90 then score * 0.1' \
--field score=80
{
"bonus": null
}
Branches of one numeric type
The two branches of an if must not be a decimal and a float, because the
result would be a decimal on some rows and a float on others. Such an if does
not compile:
cannot mix decimal and float without an explicit cast: the branches of this `if` are a decimal (`amount`) and a float (`price`); declare `price` a decimal in its Source schema, `type: decimal` in place of `type: float`, so the branches have one numeric type
The message gives one fix. When the float branch is a Source column, it is
the column’s schema type: declare price with type: decimal in place of
type: float, and the reader parses its text exactly, so both branches are
decimals holding the values the file holds. (A JSON number read into a
decimal column is still parsed through a float first; see #1299.)
When the float branch is computed rather than read from a Source column, the
fix converts the decimal branch with .to_float() instead, accepting binary
float precision. An integer branch is fine beside either: it widens exactly
into the decimal or float.
Chained conditionals
Chain multiple conditions with else if:
$ cxl eval -e 'emit tier = if amount > 1000 then "platinum"
else if amount > 500 then "gold"
else if amount > 100 then "silver"
else "bronze"' \
--field amount=750
{
"tier": "gold"
}
Nested usage
Since if/then/else is an expression, it can be used inside other expressions:
$ cxl eval -e 'emit price = base * (if member then 0.8 else 1.0)' \
--field base=100 \
--field member=true
{
"price": 80.0
}
Match
The match expression provides pattern matching. It comes in two forms: value matching (with a subject) and condition matching (without a subject).
Value form (with subject)
Match a subject expression against literal patterns:
match subject {
pattern1 => result1,
pattern2 => result2,
_ => default
}
$ cxl eval -e 'emit label = match status {
"A" => "Active",
"I" => "Inactive",
"P" => "Pending",
_ => "Unknown"
}' \
--field status=A
{
"label": "Active"
}
The wildcard _ is the catch-all arm. It matches any value not covered by preceding arms.
Condition form (without subject)
When no subject is provided, each arm’s pattern is evaluated as a boolean condition, and the first arm whose condition is exactly true wins. An arm whose condition is null is skipped:
match {
condition1 => result1,
condition2 => result2,
_ => default
}
$ cxl eval -e 'emit tier = match {
amount > 1000 => "high",
amount > 100 => "medium",
_ => "low"
}' \
--field amount=500
{
"tier": "medium"
}
Practical examples
Tiered pricing:
emit discount = match {
qty >= 1000 => 0.25,
qty >= 100 => 0.15,
qty >= 10 => 0.05,
_ => 0.0
}
Status code mapping:
emit status_text = match http_code {
200 => "OK",
201 => "Created",
400 => "Bad Request",
404 => "Not Found",
500 => "Internal Server Error",
_ => "HTTP ".concat(http_code.to_string())
}
Region classification:
emit region = match country {
"US" => "North America",
"CA" => "North America",
"MX" => "North America",
"GB" => "Europe",
"DE" => "Europe",
"FR" => "Europe",
_ => "Other"
}
Arms of one numeric type
As with if, the arms of a match must not include both a decimal and a
float. The error names the first decimal arm and the first float arm, counted
from 1, or the field when an arm is a bare field, and gives the same one fix:
the Source schema type when the float arm is a float column, otherwise
.to_float() on the decimal arm.
Match arms are evaluated in order
The first matching arm wins. Place more specific conditions before general ones:
# Correct: specific before general
emit category = match {
amount > 10000 => "enterprise",
amount > 1000 => "business",
_ => "personal"
}
# Wrong: first arm always matches
emit category = match {
amount > 0 => "personal", # catches everything positive
amount > 1000 => "business", # never reached
amount > 10000 => "enterprise", # never reached
_ => "unknown"
}
Built-in Methods
CXL provides built-in scalar methods organized into categories. Methods are called on a receiver value using dot notation: receiver.method(args).
Null propagation
Most methods return null when the receiver is null. This means null values flow through method chains without causing errors. The exceptions are documented in Introspection & Debug.
Method categories
String Methods (24 methods)
Text manipulation: case conversion, trimming, padding, searching, splitting, regex matching.
| Method | Description |
|---|---|
upper, lower | Case conversion |
trim, trim_start, trim_end | Whitespace removal |
starts_with, ends_with, contains | Substring testing |
replace | Find and replace |
substring, left, right | Extraction |
pad_left, pad_right | Padding |
repeat, reverse | Repetition and reversal |
length | Character count |
split, join | Splitting and joining |
matches, find, capture | Regex operations |
format, concat | Formatting and concatenation |
Numeric Methods (8 methods)
Rounding, clamping, and comparison for integers and floats.
| Method | Description |
|---|---|
abs | Absolute value |
ceil, floor | Ceiling and floor |
round, round_to | Rounding to decimal places |
clamp | Constrain to range |
min, max | Pairwise minimum/maximum |
Date & Time Methods (13 methods)
Date component extraction, arithmetic, and formatting.
| Method | Description |
|---|---|
year, month, day | Date component extraction |
hour, minute, second | Time component extraction (DateTime only) |
add_days, add_months, add_years | Date arithmetic |
diff_days, diff_months, diff_years | Date difference |
format_date | Custom date formatting |
Conversion Methods (11 methods)
Type conversion in strict (error on failure) and lenient (null on failure) variants.
| Method | Description |
|---|---|
to_int, to_float, to_string, to_bool | Strict conversion |
to_date, to_datetime | Strict date parsing |
try_int, try_float, try_bool | Lenient conversion |
try_date, try_datetime | Lenient date parsing |
Introspection & Debug (5 methods)
Type inspection, null checking, and debugging. These are the only methods that accept null receivers without propagating null.
| Method | Description |
|---|---|
type_of | Returns the type name as a string |
is_null | Tests for null |
is_empty | Tests for empty string, empty array, or null |
catch | Null fallback (equivalent to ??) |
debug | Passthrough with tracing side effect |
Path Methods (5 methods)
File path component extraction.
| Method | Description |
|---|---|
file_name | Full filename with extension |
file_stem | Filename without extension |
extension | File extension |
parent | Parent directory path |
parent_name | Parent directory name |
Array Methods
Traversal and transformation over nested arrays. Closure-bearing methods take an arrow-syntax closure and evaluate it per element.
| Method | Description |
|---|---|
filter, map, find, any, flat_map | Closure-bearing traversal |
remove | Drop the element at a given index |
length, join | Cross-listed on arrays (also defined on strings) |
Map Methods
Builders and accessors for Value::Map payloads. All map methods return new maps – they never mutate the receiver.
| Method | Description |
|---|---|
keys, values | List map keys / values as arrays |
merge | Union of two maps (right wins on conflict) |
set | Insert / replace an entry, by single key or by a nested a.b[0].c path |
remove_field | Drop a single entry by top-level key |
unset | Delete an entry by single key or by a nested a.b[0].c path (array index removes-and-shifts; missing path is a no-op) |
String Methods
CXL provides 24 built-in methods for string manipulation. All string methods return null when the receiver is null (null propagation).
Case conversion
upper()
Converts all characters to uppercase.
$ cxl eval -e 'emit result = "hello world".upper()'
{
"result": "HELLO WORLD"
}
lower()
Converts all characters to lowercase.
$ cxl eval -e 'emit result = "Hello World".lower()'
{
"result": "hello world"
}
Whitespace trimming
trim()
Removes leading and trailing whitespace.
$ cxl eval -e 'emit result = " hello ".trim()'
{
"result": "hello"
}
trim_start()
Removes leading whitespace only.
$ cxl eval -e 'emit result = " hello ".trim_start()'
{
"result": "hello "
}
trim_end()
Removes trailing whitespace only.
$ cxl eval -e 'emit result = " hello ".trim_end()'
{
"result": " hello"
}
Substring testing
starts_with(prefix: String) -> Bool
Tests whether the string starts with the given prefix.
$ cxl eval -e 'emit result = "hello world".starts_with("hello")'
{
"result": true
}
ends_with(suffix: String) -> Bool
Tests whether the string ends with the given suffix.
$ cxl eval -e 'emit result = "report.csv".ends_with(".csv")'
{
"result": true
}
contains(substring: String) -> Bool
Tests whether the string contains the given substring.
$ cxl eval -e 'emit result = "hello world".contains("lo wo")'
{
"result": true
}
Find and replace
replace(find: String, replacement: String) -> String
Replaces all occurrences of find with replacement.
$ cxl eval -e 'emit result = "foo-bar-baz".replace("-", "_")'
{
"result": "foo_bar_baz"
}
Extraction
substring(start: Int [, length: Int]) -> String
Extracts a substring starting at start (0-based character index). If length is provided, takes at most that many characters. If omitted, takes all remaining characters.
$ cxl eval -e 'emit result = "hello world".substring(6)'
{
"result": "world"
}
$ cxl eval -e 'emit result = "hello world".substring(0, 5)'
{
"result": "hello"
}
left(n: Int) -> String
Returns the first n characters.
$ cxl eval -e 'emit result = "hello world".left(5)'
{
"result": "hello"
}
right(n: Int) -> String
Returns the last n characters.
$ cxl eval -e 'emit result = "hello world".right(5)'
{
"result": "world"
}
Padding
pad_left(width: Int [, char: String]) -> String
Left-pads the string to the given width. Default pad character is a space.
$ cxl eval -e 'emit result = "42".pad_left(5, "0")'
{
"result": "00042"
}
$ cxl eval -e 'emit result = "hi".pad_left(6)'
{
"result": " hi"
}
pad_right(width: Int [, char: String]) -> String
Right-pads the string to the given width. Default pad character is a space.
$ cxl eval -e 'emit result = "hi".pad_right(6, ".")'
{
"result": "hi...."
}
Repetition and reversal
repeat(n: Int) -> String
Repeats the string n times.
$ cxl eval -e 'emit result = "ab".repeat(3)'
{
"result": "ababab"
}
reverse() -> String
Reverses the characters in the string.
$ cxl eval -e 'emit result = "hello".reverse()'
{
"result": "olleh"
}
Length
length() -> Int
Returns the number of characters in the string. Also works on arrays, returning the number of elements.
$ cxl eval -e 'emit result = "hello".length()'
{
"result": 5
}
Splitting and joining
split(delimiter: String) -> Array
Splits the string by the delimiter, returning an array of strings.
$ cxl eval -e 'emit result = "a,b,c".split(",")'
{
"result": ["a", "b", "c"]
}
join(delimiter: String) -> String
Joins an array of values into a string with the given delimiter. The receiver must be an array.
$ cxl eval -e 'emit result = "a,b,c".split(",").join(" - ")'
{
"result": "a - b - c"
}
Regex operations
matches(pattern: String) -> Bool
Tests whether the string fully matches the given regex pattern.
$ cxl eval -e 'emit result = "abc123".matches("^[a-z]+[0-9]+$")'
{
"result": true
}
find(pattern: String) -> Bool
Tests whether the string contains a substring matching the given regex pattern (partial match).
$ cxl eval -e 'emit result = "hello world 42".find("[0-9]+")'
{
"result": true
}
capture(pattern: String [, group: Int]) -> String
Extracts a capture group from the first regex match. Default group is 0 (the full match).
$ cxl eval -e 'emit result = "order-12345".capture("order-([0-9]+)", 1)'
{
"result": "12345"
}
Formatting and concatenation
format(fmt: String) -> String
Formats the receiver value as a string.
$ cxl eval -e 'emit result = 42.format("")'
{
"result": "42"
}
concat(args: String…) -> String
Concatenates the receiver with one or more string arguments. Null arguments are treated as empty strings.
$ cxl eval -e 'emit result = "hello".concat(" ", "world")'
{
"result": "hello world"
}
This is variadic – it accepts any number of string arguments:
$ cxl eval -e 'emit result = "a".concat("b", "c", "d")'
{
"result": "abcd"
}
Numeric Methods
CXL provides 8 built-in methods for numeric operations. These methods work on both Integer and Float values (the Numeric receiver type). All return null when the receiver is null.
abs, ceil, floor, round, and round_to also accept an exact
decimal receiver and return an exact decimal
result — d.round(2) rounds d to two fractional digits using banker’s
rounding, staying in the decimal domain rather than converting to a float.
The type checker infers decimal for these calls as well, so the result
composes with other decimals without a cast:
amount.round_to(2) + fee typechecks when both columns are decimal.
abs() -> Numeric
Returns the absolute value. Preserves the original type (Int stays Int, Float stays Float).
$ cxl eval -e 'emit result = (-42).abs()'
{
"result": 42
}
$ cxl eval -e 'emit result = (-3.14).abs()'
{
"result": 3.14
}
ceil() -> Int
Rounds up to the nearest integer. Returns the value unchanged for integers.
$ cxl eval -e 'emit result = 3.2.ceil()'
{
"result": 4
}
$ cxl eval -e 'emit result = (-3.2).ceil()'
{
"result": -3
}
floor() -> Int
Rounds down to the nearest integer. Returns the value unchanged for integers.
$ cxl eval -e 'emit result = 3.8.floor()'
{
"result": 3
}
$ cxl eval -e 'emit result = (-3.2).floor()'
{
"result": -4
}
round([decimals: Int]) -> Float
Rounds to the specified number of decimal places. Default is 0 decimal places.
On a decimal receiver the result is a decimal (exact banker’s rounding),
not a float.
$ cxl eval -e 'emit result = 3.456.round()'
{
"result": 3.0
}
$ cxl eval -e 'emit result = 3.456.round(2)'
{
"result": 3.46
}
round_to(decimals: Int) -> Float
Rounds to the specified number of decimal places. Unlike round(), the decimals argument is required.
On a decimal receiver the result is a decimal (exact banker’s rounding),
not a float.
$ cxl eval -e 'emit result = 3.14159.round_to(3)'
{
"result": 3.142
}
Use round_to when you want to be explicit about precision in financial or scientific calculations:
$ cxl eval -e 'emit price = 19.995.round_to(2)'
{
"price": 20.0
}
clamp(min: Numeric, max: Numeric) -> Numeric
Constrains the value to the given range. Returns min if the value is below it, max if above, or the value itself if within range.
$ cxl eval -e 'emit result = 150.clamp(0, 100)'
{
"result": 100
}
$ cxl eval -e 'emit result = (-5).clamp(0, 100)'
{
"result": 0
}
$ cxl eval -e 'emit result = 50.clamp(0, 100)'
{
"result": 50
}
min(other: Numeric) -> Numeric
Returns the smaller of the receiver and the argument.
$ cxl eval -e 'emit result = 10.min(20)'
{
"result": 10
}
$ cxl eval -e 'emit result = 10.min(5)'
{
"result": 5
}
max(other: Numeric) -> Numeric
Returns the larger of the receiver and the argument.
$ cxl eval -e 'emit result = 10.max(20)'
{
"result": 20
}
$ cxl eval -e 'emit result = 10.max(5)'
{
"result": 10
}
Practical examples
Clamp a percentage:
emit pct = (completed / total * 100).clamp(0, 100).round_to(1)
Absolute difference:
emit diff = (actual - expected).abs()
Floor division for batch numbering:
emit batch = (row_number / 1000).floor()
Date & Time Methods
CXL provides 13 built-in methods for date and time manipulation. These methods work on Date and DateTime values. All return null when the receiver is null.
Component extraction
year() -> Int
Returns the year component.
$ cxl eval -e 'emit result = #2024-03-15#.year()'
{
"result": 2024
}
month() -> Int
Returns the month component (1-12).
$ cxl eval -e 'emit result = #2024-03-15#.month()'
{
"result": 3
}
day() -> Int
Returns the day-of-month component (1-31).
$ cxl eval -e 'emit result = #2024-03-15#.day()'
{
"result": 15
}
hour() -> Int
Returns the hour component (0-23). DateTime only – returns null for Date values.
$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime().hour()'
{
"result": 14
}
minute() -> Int
Returns the minute component (0-59). DateTime only – returns null for Date values.
$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime().minute()'
{
"result": 30
}
second() -> Int
Returns the second component (0-59). DateTime only – returns null for Date values.
$ cxl eval -e 'emit result = "2024-03-15T14:30:45".to_datetime().second()'
{
"result": 45
}
Date arithmetic
add_days(n: Int) -> Date
Adds n days to the date. Use negative values to subtract. Works on both Date and DateTime.
$ cxl eval -e 'emit result = #2024-01-15#.add_days(10)'
{
"result": "2024-01-25"
}
$ cxl eval -e 'emit result = #2024-01-15#.add_days(-5)'
{
"result": "2024-01-10"
}
add_months(n: Int) -> Date
Adds n months to the date. Day is clamped to the last day of the target month if necessary.
$ cxl eval -e 'emit result = #2024-01-31#.add_months(1)'
{
"result": "2024-02-29"
}
$ cxl eval -e 'emit result = #2024-03-15#.add_months(-2)'
{
"result": "2024-01-15"
}
add_years(n: Int) -> Date
Adds n years to the date. Leap day (Feb 29) is clamped to Feb 28 in non-leap years.
$ cxl eval -e 'emit result = #2024-02-29#.add_years(1)'
{
"result": "2025-02-28"
}
Date difference
diff_days(other: Date) -> Int
Returns the number of days between the receiver and the argument (receiver - other). Positive when the receiver is later.
$ cxl eval -e 'emit result = #2024-03-15#.diff_days(#2024-03-01#)'
{
"result": 14
}
$ cxl eval -e 'emit result = #2024-01-01#.diff_days(#2024-03-15#)'
{
"result": -74
}
diff_months(other: Date) -> Int
Returns the difference in months between two dates.
Note: This method currently returns
null(unimplemented). Usediff_daysand divide by 30 as an approximation.
diff_years(other: Date) -> Int
Returns the difference in years between two dates.
Note: This method currently returns
null(unimplemented). Usediff_daysand divide by 365 as an approximation.
Formatting
format_date(format: String) -> String
Formats the date/datetime using a chrono format string. See chrono format syntax.
Common format specifiers:
| Specifier | Description | Example |
|---|---|---|
%Y | 4-digit year | 2024 |
%m | 2-digit month | 03 |
%d | 2-digit day | 15 |
%H | Hour (24h) | 14 |
%M | Minute | 30 |
%S | Second | 00 |
%B | Full month name | March |
%b | Abbreviated month | Mar |
%A | Full weekday | Friday |
$ cxl eval -e 'emit result = #2024-03-15#.format_date("%B %d, %Y")'
{
"result": "March 15, 2024"
}
$ cxl eval -e 'emit result = #2024-03-15#.format_date("%Y/%m/%d")'
{
"result": "2024/03/15"
}
Practical examples
Fiscal year calculation (April start):
let d = invoice_date
emit fiscal_year = if d.month() < 4 then d.year() - 1 else d.year()
Age in days:
emit days_since = now.diff_days(created_date)
Quarter:
emit quarter = match {
invoice_date.month() <= 3 => "Q1",
invoice_date.month() <= 6 => "Q2",
invoice_date.month() <= 9 => "Q3",
_ => "Q4"
}
ISO week format:
emit formatted = order_date.format_date("%Y-W%V")
Conversion Methods
CXL provides two families of conversion methods: strict (7 methods) and lenient (6 methods). Strict conversions raise an error on failure, halting pipeline execution. Lenient conversions return null on failure, allowing graceful handling of dirty data.
All conversion methods accept any receiver type (Any).
Strict conversions
Use strict conversions for required fields where invalid data should halt processing.
to_int() -> Int
Converts the receiver to an integer. Errors on failure.
- Float: truncates toward zero
- String: parses as integer
- Bool:
truebecomes1,falsebecomes0
$ cxl eval -e 'emit result = "42".to_int()'
{
"result": 42
}
$ cxl eval -e 'emit result = 3.9.to_int()'
{
"result": 3
}
to_float() -> Float
Converts the receiver to a float. Errors on failure.
- Integer: promotes to float
- String: parses as float
$ cxl eval -e 'emit result = "3.14".to_float()'
{
"result": 3.14
}
$ cxl eval -e 'emit result = 42.to_float()'
{
"result": 42.0
}
to_decimal() -> Decimal
Converts the receiver to an exact decimal. Errors on failure.
- Integer: converts exactly
- String: parses base-10 exactly (
"19.99"→19.99, never via a binary float) - Float: converts via the binary value — this is the one lossy direction, made
explicit precisely because
decimal * floatis otherwise a type error
Use to_decimal() to bring a value into exact decimal arithmetic. JSON output
renders a decimal as a scale-preserving string.
$ cxl eval -e 'emit result = "0.10".to_decimal() + "0.20".to_decimal()'
{
"result": "0.30"
}
to_string() -> String
Converts any value to its string representation. Never fails.
$ cxl eval -e 'emit result = 42.to_string()'
{
"result": "42"
}
$ cxl eval -e 'emit result = true.to_string()'
{
"result": "true"
}
to_bool() -> Bool
Converts the receiver to a boolean. Errors on failure.
- String:
"true","1","yes"becometrue;"false","0","no"becomefalse(case-insensitive) - Integer:
0isfalse, everything else istrue
$ cxl eval -e 'emit result = "yes".to_bool()'
{
"result": true
}
$ cxl eval -e 'emit result = 0.to_bool()'
{
"result": false
}
to_date([format: String]) -> Date
Parses a string to a Date. Without a format argument, expects ISO 8601 (YYYY-MM-DD). With a format, uses chrono strftime syntax.
$ cxl eval -e 'emit result = "2024-03-15".to_date()'
{
"result": "2024-03-15"
}
$ cxl eval -e 'emit result = "15/03/2024".to_date("%d/%m/%Y")'
{
"result": "2024-03-15"
}
to_datetime([format: String]) -> DateTime
Parses a string to a DateTime. Without a format argument, expects ISO 8601 (YYYY-MM-DDTHH:MM:SS). With a format, uses chrono strftime syntax.
$ cxl eval -e 'emit result = "2024-03-15T14:30:00".to_datetime()'
{
"result": "2024-03-15T14:30:00"
}
Lenient conversions
Use lenient conversions for optional or dirty data fields. They return null instead of raising errors, making them safe to combine with ?? for fallback values.
try_int() -> Int
Attempts to convert to integer. Returns null on failure.
$ cxl eval -e 'emit a = "42".try_int()' -e 'emit b = "abc".try_int()'
{
"a": 42,
"b": null
}
try_float() -> Float
Attempts to convert to float. Returns null on failure.
$ cxl eval -e 'emit a = "3.14".try_float()' -e 'emit b = "N/A".try_float()'
{
"a": 3.14,
"b": null
}
try_decimal() -> Decimal
Attempts to convert to an exact decimal. Returns null on failure.
$ cxl eval -e 'emit a = "19.99".try_decimal()' -e 'emit b = "N/A".try_decimal()'
{
"a": "19.99",
"b": null
}
try_bool() -> Bool
Attempts to convert to boolean. Returns null on failure.
$ cxl eval -e 'emit a = "yes".try_bool()' -e 'emit b = "maybe".try_bool()'
{
"a": true,
"b": null
}
try_date([format: String]) -> Date
Attempts to parse a string as a Date. Returns null on failure.
$ cxl eval -e 'emit a = "2024-03-15".try_date()' \
-e 'emit b = "not a date".try_date()'
{
"a": "2024-03-15",
"b": null
}
try_datetime([format: String]) -> DateTime
Attempts to parse a string as a DateTime. Returns null on failure.
$ cxl eval -e 'emit a = "2024-03-15T14:30:00".try_datetime()' \
-e 'emit b = "invalid".try_datetime()'
{
"a": "2024-03-15T14:30:00",
"b": null
}
When to use each
Strict conversions (to_*) for:
- Required fields that must be valid
- Schema-enforced data where bad input should halt the pipeline
- Fields already validated upstream
Lenient conversions (try_*) for:
- Optional fields that may be missing or malformed
- Dirty data with mixed formats
- Fields where a fallback value is acceptable
Practical patterns
Safe numeric parsing with fallback:
emit amount = raw_amount.try_float() ?? 0.0
Parse dates from multiple formats:
emit parsed = raw_date.try_date("%Y-%m-%d")
?? raw_date.try_date("%m/%d/%Y")
?? raw_date.try_date("%d-%b-%Y")
Strict conversion for required fields:
emit employee_id = raw_id.to_int() # halts on bad data -- correct behavior
emit salary = raw_salary.to_float() # must be numeric
Lenient conversion for optional fields:
emit bonus = raw_bonus.try_float() # null if missing or non-numeric
emit total = salary + (bonus ?? 0.0) # safe arithmetic
Introspection & Debug
CXL provides 4 introspection methods and 1 debug method. The four introspection methods are the only methods that accept null receivers without propagating null – they are designed specifically for inspecting and handling null values. debug follows ordinary null propagation.
type_of() -> String
Returns the type name of the receiver as a string. Works on any value, including null.
Type name strings: "string", "int", "float", "decimal", "bool", "date", "datetime", "null", "array", "map".
$ cxl eval -e 'emit a = 42.type_of()' -e 'emit b = "hello".type_of()' \
-e 'emit c = null.type_of()'
{
"a": "int",
"b": "string",
"c": "null"
}
Useful for branching on dynamic types:
emit formatted = match value.type_of() {
"int" => value.to_string().concat(" (integer)"),
"float" => value.round_to(2).to_string().concat(" (decimal)"),
_ => value.to_string()
}
is_null() -> Bool
Returns true if the receiver is null, false otherwise. This is the primary way to test for null values – it is NOT subject to null propagation.
$ cxl eval -e 'emit a = null.is_null()' -e 'emit b = 42.is_null()'
{
"a": true,
"b": false
}
Use in filter statements:
filter not field.is_null()
is_empty() -> Bool
Returns true for empty strings, empty arrays, or null values. Returns false for all other values.
$ cxl eval -e 'emit a = "".is_empty()' -e 'emit b = "hello".is_empty()' \
-e 'emit c = null.is_empty()'
{
"a": true,
"b": false,
"c": true
}
Useful for filtering out blank or missing records:
filter not name.is_empty()
catch(fallback: Any) -> Any
Returns the receiver if it is non-null, otherwise returns the fallback value. This is the method equivalent of the ?? operator.
$ cxl eval -e 'emit a = null.catch("default")' \
-e 'emit b = "present".catch("default")'
{
"a": "default",
"b": "present"
}
catch and ?? are interchangeable:
# These two are equivalent:
emit name = raw_name.catch("Unknown")
emit name = raw_name ?? "Unknown"
debug(label: String) -> Any
Passes the receiver through unchanged while emitting a trace log with the given label. Zero overhead when tracing is disabled. The return value is always the receiver, making it safe to insert into any expression chain. A null receiver returns null without logging.
$ cxl eval -e 'emit result = 42.debug("check value")'
{
"result": 42
}
Insert debug anywhere in a method chain for inspection without affecting the output:
emit total = price.debug("price")
* qty.debug("qty")
When tracing is enabled, this produces log lines like:
TRACE source_row=1 source_file=input.csv: price: Integer(100)
TRACE source_row=1 source_file=input.csv: qty: Integer(5)
Null-safe summary
| Method | Null receiver behavior |
|---|---|
type_of() | Returns "null" |
is_null() | Returns true |
is_empty() | Returns true |
catch(x) | Returns x |
debug(l) | Returns null, logs nothing (propagation) |
| All other methods | Return null (propagation) |
Path Methods
CXL provides 5 built-in methods for extracting components from file path strings. All path methods take a string receiver and return a string. They return null when the receiver is null or when the requested component does not exist.
file_name() -> String
Returns the full filename (with extension) from the path.
$ cxl eval -e 'emit result = "/data/reports/sales.csv".file_name()'
{
"result": "sales.csv"
}
file_stem() -> String
Returns the filename without the extension.
$ cxl eval -e 'emit result = "/data/reports/sales.csv".file_stem()'
{
"result": "sales"
}
extension() -> String
Returns the file extension (without the leading dot).
$ cxl eval -e 'emit result = "/data/reports/sales.csv".extension()'
{
"result": "csv"
}
Returns null when no extension is present:
$ cxl eval -e 'emit result = "/data/reports/README".extension()'
{
"result": null
}
parent() -> String
Returns the parent directory path.
$ cxl eval -e 'emit result = "/data/reports/sales.csv".parent()'
{
"result": "/data/reports"
}
parent_name() -> String
Returns just the name of the parent directory (not the full path).
$ cxl eval -e 'emit result = "/data/reports/sales.csv".parent_name()'
{
"result": "reports"
}
Practical examples
Organize output by source directory:
emit source_dir = $pipeline.source_file.parent_name()
emit source_type = $pipeline.source_file.extension()
Extract file identifiers:
emit file_id = $pipeline.source_file.file_stem()
emit is_csv = $pipeline.source_file.extension() == "csv"
Route by file type:
let ext = input_path.extension()
emit format = match ext {
"csv" => "delimited",
"json" => "structured",
"xml" => "markup",
_ => "unknown"
}
Array Methods
CXL provides closure-bearing and non-closure array builtins for traversing and transforming nested arrays carried on a single record. The closure-bearing methods take an arrow-syntax closure and evaluate it once per element.
Null propagation
Every array method returns null when the receiver is null. The closure body is not invoked on a null receiver.
Closure-bearing methods
filter(it => Bool) -> Array
Returns a new array containing the elements for which the closure body evaluates to true.
- type: transform
name: filter_items
input: orders
config:
cxl: |
emit kept = items.filter(it => it["price"] > 5)
For an input record where items is [{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}], kept is [{"sku":"a","price":10},{"sku":"b","price":20}].
map(it => T) -> Array
Returns a new array whose elements are the closure body’s value for each input element. The element type need not match the input element type.
cxl: |
emit skus = items.map(it => it["sku"])
emit doubled_prices = items.map(it => it["price"] * 2)
skus is ["a", "b", "c"]; doubled_prices is [20, 40, 10].
find(it => Bool) -> Element | Null
Returns the first element for which the closure body evaluates to true. Returns null if no element matches.
cxl: |
emit first_premium = items.find(it => it["price"] > 15)
first_premium is {"sku":"b","price":20} for the running example.
any(it => Bool) -> Bool
Returns true if the closure body evaluates to true for at least one element. Returns false if no element matches (including on an empty array).
cxl: |
emit has_cheap = items.any(it => it["price"] < 10)
has_cheap is true.
flat_map(it => Array) -> Array
Like map, but the closure body returns an array per input element; the results are concatenated into a single flat array. A null body result contributes no elements; a non-array body result contributes a single element.
cxl: |
emit all_tags = items.flat_map(it => it["tags"])
For input items carrying tags arrays (e.g. [{"sku":"a","tags":["new"]},{"sku":"b","tags":["sale","new"]}]), all_tags is ["new","sale","new"].
Non-closure methods
remove(index: Int) -> Array
Returns a new array with the element at the given 0-based index removed. The original array is unchanged.
cxl: |
emit shifted = items.remove(1)
shifted is [{"sku":"a","price":10},{"sku":"c","price":5}] – index 0 is preserved, index 2 shifts down to index 1.
If the index is negative or out of range, remove returns the receiver array unchanged.
length() -> Int
Returns the number of elements in the array. length is also defined on strings (see String Methods).
cxl: |
emit item_count = items.length()
item_count is 3.
join(separator: String) -> String
Joins an array of values into a single string with the given separator between elements. Defined as a string method (see String Methods) but accepts array receivers.
cxl: |
emit sku_list = items.map(it => it["sku"]).join(", ")
sku_list is "a, b, c".
Bracket indexing vs .remove
Bracket indexing (items[0]) reads an element by position and returns null when out of range. .remove(idx) returns a new array with the element dropped; out-of-range indices leave the array unchanged. See Nested Paths for the index-access surface.
See also
- Closures – the
it => bodyform used by closure-bearing array methods. - Map Methods – builtins that operate on the map elements typically iterated by these array methods.
- Nested Paths – bracket-index and dotted-path access through nested arrays and maps.
- Emit Each – fan one input record into many output records, one per array element.
Map Methods
CXL provides six built-in methods for working with map values (key-value pairs). Maps arise naturally from JSON object inputs, from the set builder below, and from upstream emits that produce nested structures.
All map methods return new values – they never mutate the receiver. This is copy-on-write semantics: chaining .set then .remove_field produces a fresh map at each step, leaving the upstream binding untouched.
Null propagation
Every map method returns null when the receiver is null or is not a Value::Map.
Method reference
keys() -> Array
Returns the map’s keys as an array of strings, preserving insertion order.
- type: transform
name: list_keys
input: rows
config:
cxl: |
emit field_names = profile.keys()
For an input record where profile is {"name":"Alice","tier":"gold","since":"2021-04"}, field_names is ["name","tier","since"].
values() -> Array
Returns the map’s values as an array, preserving insertion order. Value types are heterogeneous – the array carries each value as-is.
cxl: |
emit field_values = profile.values()
field_values is ["Alice","gold","2021-04"].
merge(other: Map) -> Map
Returns a new map containing every key from the receiver and from other. On conflicting keys, other’s value wins.
cxl: |
emit enriched = profile.merge(overrides)
For profile = {"name":"Alice","tier":"gold"} and overrides = {"tier":"platinum","since":"2021-04"}, enriched is {"name":"Alice","tier":"platinum","since":"2021-04"}.
set(key: String, value: Any) -> Map
Returns a new map with key set to value. If the key was already present, its value is replaced; insertion order is preserved.
cxl: |
emit stamped = profile.set("region", "us-east")
stamped is {"name":"Alice","tier":"gold","since":"2021-04","region":"us-east"}.
Nested paths
key may be a dotted/indexed path that descends into nested maps and arrays, so a single set writes into a deep document. Dots separate map keys; a [n] suffix indexes an array.
cxl: |
emit moved = profile.set("address.city", "NYC")
emit relabel = order.set("items[0].sku", "A-100")
- Auto-create. Missing intermediate map segments are created as empty maps, so a path can build structure that does not yet exist.
{}.set("a.b.c", 7)returns{"a":{"b":{"c":7}}}. This is what letssetassemble a nested document from scratch (matching jqsetpathand Bloblang assignment). - Type conflict -> null. If an intermediate segment already exists but is the wrong kind for the next step – descending into a key whose value is a scalar, indexing a map with
[n], or naming a field on an array – the whole operation returnsnull. Nothing is partially written. - Array index past the end -> null. Indexing past the last element returns
nullfor the whole operation; arrays are never silently grown. The path can only overwrite an array slot that already exists. - A bare key is a single key, not a path.
"region"writes the top-levelregion. Only.and[n]introduce nesting; a key with neither behaves exactly as before.
For profile = {"name":"Alice","address":{"city":"LA"}}, profile.set("address.city", "NYC") is {"name":"Alice","address":{"city":"NYC"}} – the sibling name and any other address keys are preserved.
Known limitation. Because
.and[are path syntax,setcannot yet target a key whose name literally contains a.or[(for example a JSON field literally named"a.b"). Column names solve this with a backslash escape —a\.bis one segment nameda.b, per Field Paths — and the decided direction is forset/unsetkeys to read that same grammar, since both are flat strings addressing a path. Until they do: to write such a key, build it withmergeand a map literal; to remove it, useremove_field, which matches the exact key string.
remove_field(key: String) -> Map
Returns a new map without key. If the key was absent, the receiver is returned unchanged.
cxl: |
emit slim = profile.remove_field("since")
slim is {"name":"Alice","tier":"gold"}.
unset(key: String) -> Map
Returns a new map with the entry addressed by key removed. unset is the deletion counterpart to set and reuses the same dotted/indexed path grammar: dots separate map keys, a [n] suffix indexes an array. A bare key (no . or [n]) drops a top-level entry, exactly like remove_field.
cxl: |
emit pruned = profile.unset("address.city")
emit dropped = order.unset("items[0]")
- Array element removes and shifts.
unset("items[0]")deletes element 0 and shifts the remaining elements down, so the array shrinks by one (matching jqdel). This is deliberately distinct fromset("items[0]", null), which leaves anullhole —unsetmeans delete. - Missing or conflicting path is a no-op. A path that does not resolve — a missing intermediate, a missing final key, an array index past the end, or a type conflict (a field segment against an array, an index segment against a map) — returns the receiver unchanged. This mirrors
remove_fieldon an absent key, and is the opposite ofset, which returnsnullon a conflicting path. - Copy-on-write. Like every map method,
unsetnever mutates the receiver; the upstream binding is untouched.
For profile = {"name":"Alice","address":{"city":"LA","zip":"90001"}}, profile.unset("address.city") is {"name":"Alice","address":{"zip":"90001"}} – the sibling zip and the top-level name are preserved.
Worked example: chained set + remove_field
Map methods compose naturally because each returns a new map.
- type: transform
name: rewrite_profile
input: rows
config:
cxl: |
emit profile =
profile.set("region", "us-east").remove_field("internal_id")
For profile = {"name":"Alice","internal_id":"ix-77","tier":"gold"}, the emitted profile is {"name":"Alice","tier":"gold","region":"us-east"}. The internal_id slot is removed and the region slot is appended; both happen on a fresh map so the upstream record’s profile is unaffected for any other downstream branch.
Parentheses are required
All map methods are method calls and must be written with parentheses, even the zero-argument ones:
profile.keys() -- ok
profile.keys -- parses as a field lookup, not a method call
profile.keys parses as a dotted path – a lookup for a field literally named keys inside profile. That path almost certainly returns null. Always include the parentheses when invoking a map method.
Using map methods inside array closures
Map methods compose with closure-bearing array builtins when the array elements are themselves maps.
cxl: |
emit enriched_items = items.map(it => it.set("region", "us-east"))
emit item_keys = items.map(it => it.keys())
Each it is a map; the closure body invokes a map method on it. enriched_items is an array where every element gained a region field. item_keys is an array of key-name arrays, one per element.
See also
- Closures – arrow-syntax closures often invoke map methods on their
itbinding. - Array Methods – closure-bearing array methods commonly carry maps as their elements.
- Nested Paths – bracket-index access (
profile["name"]) reads a single key without producing a new map.
Window Functions
Window functions allow CXL expressions to access aggregated values across a set of records within an analytic window. Unlike aggregate functions (which collapse groups into single rows), window functions attach computed values to each individual record.
Window functions are accessed via the $window.* namespace and require an analytic_window: configuration on the transform node.
Interactive companion: the window functions explainer shows, for any row, which rows of its partition each function reads.
Configuring an analytic window
Window functions are only available in transform nodes that declare an analytic_window: section in YAML:
nodes:
- name: ranked_sales
type: transform
input: raw_sales
config:
analytic_window:
group_by: [region]
sort_by:
- field: amount
order: desc
cxl: |
emit region = region
emit amount = amount
emit region_total = $window.sum(amount)
emit running_total = $window.cumulative_sum(amount)
emit rank_position = $window.row_number()
Window configuration fields
| Field | Description |
|---|---|
group_by | List of fields to partition the window by (the SQL PARTITION BY axis). |
sort_by | List of { field, order, null_order } ordering specifications: order is asc (default) or desc, and null_order is first or last (default last), placing null keys before or after every value. Values compare by the rule every sort uses; see How values are ordered. null_order: drop is rejected; see Nulls in sort_by. |
source | Optional explicit source-name reference for cross-source windows. |
on | Optional cross-source partition-lookup field. |
Which rows a function reads
There is no frame option. A window function reads one of three things:
- The whole partition.
sum,avg,min,max,count,first_value,last_value,first(),last(),any,every,exists,not_exists,collectanddistinctread every row of the record’s partition, whatever the record’s position. Every record in a partition gets the same$window.sum(amount). - The partition up to the current record.
cumulative_sumis the running total, from the partition’s first record (insort_byorder) through the current one. - A position.
row_number,rankanddense_rankgive the current record’s place insort_byorder;lag(n)andlead(n)read the recordnplaces before or after it.
Nulls in sort_by
sort_by only orders the rows of a partition; every row of the partition is
still seen by the window functions and written by the Transform. A row whose
key is null is placed by null_order: before every value with first, after
every value with last, in either direction.
null_order: drop is rejected when the pipeline is planned:
transform "running": `null_order: drop` is not allowed on `analytic_window.sort_by` for field "amount": `sort_by` only orders the rows of a window partition, placing nulls `first` or `last`, and cannot remove a row. To remove the rows whose "amount" is null, delete `null_order: drop` and add a Transform before this node with `config: { cxl: "filter not amount.is_null()" }`.
The one fix is the filter the error prints: delete null_order: drop and
add a Transform before the windowed one whose whole config is the printed
line. That leaves rows with a null key out of every partition, and also
removes them from the windowed Transform’s output:
- type: transform
name: with_amount
input: orders
config: { cxl: "filter not amount.is_null()" }
For a field CXL cannot write as a bare name, the error prints a
source_name: line instead; see
source_name.
Aggregate window functions
These compute values over the record’s whole partition, except cumulative_sum, which stops at the current record.
$window.sum(field)
Sum of the field values across the whole partition. Null and non-numeric values are skipped; a sum of integers returns a Float.
emit running_total = $window.sum(amount)
$window.cumulative_sum(field)
Running total of the field values from the partition’s first record (in sort_by order) through the current record. Like $window.sum, a sum of integers returns a Float.
emit running_total = $window.cumulative_sum(amount)
$window.avg(field)
Average of the field values across the whole partition. Returns Float.
emit moving_avg = $window.avg(amount)
$window.min(field)
Minimum value in the partition.
emit window_min = $window.min(amount)
$window.max(field)
Maximum value in the partition.
emit window_max = $window.max(amount)
$window.count()
Number of records in the partition, the same for every record in it. Takes no arguments. For a record’s position, use $window.row_number().
emit window_size = $window.count()
$window.first_value(field)
Returns the value of field at the first record of the partition
(ordered by sort_by). Equivalent to SQL FIRST_VALUE(field).
emit opening_amount = $window.first_value(amount)
$window.last_value(field)
Returns the value of field at the last record of the partition
(ordered by sort_by), the same for every record in it.
emit closing_amount = $window.last_value(amount)
Ranking window functions
Zero-argument integer functions that return the current row’s rank within its partition.
$window.row_number()
1-indexed position of the current record within its partition.
emit row_idx = $window.row_number()
$window.rank()
SQL RANK(): rows that share the same sort_by tuple receive the same
rank, and the next distinct row jumps by the size of the tie group.
emit sales_rank = $window.rank()
$window.dense_rank()
SQL DENSE_RANK(): ties share a rank with no gaps between distinct
ranks.
emit sales_dense_rank = $window.dense_rank()
Positional window functions
These return a whole record by position within the partition. Name the field to read after the call, as in $window.lag(1).amount. Without a field name the call returns null on every record, and no error is raised.
$window.first().field
The first record of the partition, in sort_by order.
emit first_amount = $window.first().amount
$window.last().field
The last record of the partition, in sort_by order.
emit last_amount = $window.last().amount
$window.lag(n).field
The record n places before the current record. Returns null if there is no record at that offset.
emit prev_amount = $window.lag(1).amount
emit two_back = $window.lag(2).amount
$window.lead(n).field
The record n places after the current record. Returns null if there is no record at that offset.
emit next_amount = $window.lead(1).amount
Iterable window functions
These evaluate predicates or collect values across the window.
$window.any(predicate)
Returns true if the predicate is true for any record in the window.
emit has_high = $window.any(amount > 1000)
$window.every(predicate)
Returns true if the predicate is true for every record in the window.
emit all_positive = $window.every(amount > 0)
$window.exists(predicate)
Returns true if the predicate is true for at least one record in the
window — a SQL-fluency alias of $window.any.
emit any_high = $window.exists(amount > 1000)
$window.not_exists(predicate)
Returns true if no record in the window satisfies the predicate.
Equivalent to not $window.exists(predicate) and to
$window.every(not predicate).
emit none_negative = $window.not_exists(amount < 0)
$window.collect(field)
Collects all values of the field in the window into an array.
emit all_amounts = $window.collect(amount)
$window.distinct(field)
Collects distinct values of the field in the window into an array.
emit unique_regions = $window.distinct(region)
$window.collect and $window.distinct emit arrays. JSON writes them as native
arrays, XML as repeated child elements, and CSV as a delimited cell. Before a
scalar-only sink such as fixed-width, coerce the value in a downstream
Transform (for example emit regions = unique_regions.join(";")).
Complete example
nodes:
- name: sales_analysis
type: transform
input: daily_sales
config:
analytic_window:
group_by: [store_id]
sort_by:
- field: sale_date
order: asc
cxl: |
emit store_id = store_id
emit sale_date = sale_date
emit daily_revenue = revenue
emit store_avg = $window.avg(revenue)
emit store_total = $window.sum(revenue)
emit revenue_to_date = $window.cumulative_sum(revenue)
emit prev_day_revenue = $window.lag(1).revenue
emit day_over_day = revenue - ($window.lag(1).revenue ?? revenue)
For each sale this adds the store’s average and total over all of its days, its revenue to date, and the change from the previous day.
Correlation-key error handling
Window functions work correctly when a pipeline uses correlation keys for group-atomic error handling: if records are retracted from an upstream group, the window recomputes the affected partitions so its output stays consistent. There is nothing to configure.
Aggregate Functions
Aggregate functions operate across grouped record sets in aggregate nodes, collapsing multiple input records into summary rows. They are distinct from window functions, which attach computed values to each individual record.
Aggregate functions
CXL provides 7 aggregate functions. These are called as free-standing function calls (not method calls) within the CXL block of an aggregate node.
| Function | Signature | Returns | Description |
|---|---|---|---|
sum(expr) | Numeric | Int, Float or Decimal (the input’s type) | Sum of values |
count(*) | – | Int | Count of records in the group |
avg(expr) | Numeric | Float, or Decimal for a decimal input | Arithmetic mean |
min(expr) | Any | Any | Minimum value |
max(expr) | Any | Any | Maximum value |
collect(expr) | Any | Array | All values collected into an array |
weighted_avg(value, weight) | Numeric, Numeric | Float / Decimal | Weighted arithmetic mean |
YAML aggregate node
Aggregate functions are used inside the cxl: block of a node with type: aggregate. The node must declare group_by: fields.
nodes:
- name: dept_summary
type: aggregate
input: employees
config:
group_by: [department]
cxl: |
emit total_salary = sum(salary)
emit headcount = count(*)
emit avg_salary = avg(salary)
emit max_salary = max(salary)
emit min_salary = min(salary)
Group-by fields pass through automatically
Fields listed in group_by: are automatically included in the output. You do NOT need to emit them – they are carried through as group keys.
In the example above, department is automatically present in every output record without an explicit emit department = department statement.
Function details
sum(expr) -> Int, Float or Decimal
Computes the sum of the expression across all records in the group. Null values are skipped.
cxl: |
emit total_revenue = sum(price * quantity)
The result has the type of the values summed: integers give an integer, floats a float, decimals a decimal. Integers summed with floats give a float, and integers summed with decimals a decimal. An integer sum outside the 64-bit integer range is an error.
A float sum is the exact total of the group’s values, rounded once to the
nearest float. It does not depend on the order rows arrive in or on
memory.limit: the same group gives the same bytes whether the Aggregate holds
every group in memory or spills and merges partial sums. A group that holds
1e16, 1.0 and -1e16 sums to 1, in any order, where adding the floats
left to right gives 0 or 1 depending on the order. Integers mixed with floats
are added exactly too, so an integer larger than 2^53 is not rounded before it
is added. A NaN in the group makes the sum NaN, and +inf with -inf makes it
NaN. Results can differ in the last bit from a version that rounded after each
addition.
A decimal sum is the exact total of the group’s values, rounded once (half to
even) only when it does not fit a decimal at its scale. Its scale is the
largest scale among the group’s values, zeros and integers included, so the
sum of 1.00, -1.00 and 2 is 2.00 whatever order the rows arrive in. It
is an error only when the whole group’s exact total is outside the decimal
range, ±79,228,162,514,264,337,593,543,950,335; a group whose running total
passes outside the range and comes back is fine. The error’s fix aggregates the
column as floats, sum(amount.to_float()) with your column in place of
amount, when a binary float’s range and precision will do.
A group whose values are all null gives null. Null is never a substitute for a
failure: a group that fails is an aggregate_finalize error (see Error
categories), which under
strategy: continue goes to the dead-letter output.
Decimal and float in one group
A decimal is never added to a float without an explicit conversion, in an
aggregate as in amount + price. A sum, avg or weighted_avg whose values
in one group include both a decimal and a float fails that group with:
decimal and float in one group: a decimal is never added to a float without an explicit conversion; declare the column that holds the floats `type: decimal` in its Source schema, so every value in the group is a decimal
When the typechecker can see the mix, for example
sum(if flag then amount else price), the pipeline does not compile (E200; see
Conditionals). The run-time error covers what it cannot see:
a value whose type is only known at run time, such as an untyped column or a
numeric result like amount.clamp(0, 100). Declaring the float column
type: decimal in its Source schema keeps the total exact: the reader parses
the column’s text as a decimal, so every value in the group is a decimal. (A
JSON number read into a decimal column is still parsed through a float first;
see #1299.) When the floats are computed upstream rather than read
from a Source column, there is no column to retype: convert the decimal values
with .to_float() instead, accepting binary float precision.
count(*) -> Int
Counts the number of records in the group. The argument is the wildcard *.
cxl: |
emit num_orders = count(*)
avg(expr) -> Float or Decimal
Computes the arithmetic mean. Null values are skipped.
cxl: |
emit avg_order_value = avg(order_total)
avg(x) is sum(x) / count(x): the group’s exact sum, rounded once as sum
rounds it, divided by the number of non-null values, so it is exact over floats
and decimals alike and does not depend on row order or memory.limit. Over
decimals the result
is a decimal, the quotient at full precision, so avg(amount) and
sum(amount) / count(amount) give the same digits and the same scale. Over
floats, and over integers mixed with floats, the result is a float. Over
integers alone it is a float: the exact integer total, converted once to a
float, divided by the count.
A decimal total outside the decimal range is an error, as for sum, and so is
a group mixing decimals and floats. A group whose values are all null gives
null.
min(expr) -> Any
Returns the minimum value in the group. Works on numeric, string, and date types. Null values are skipped; a group whose values are all null gives null.
cxl: |
emit earliest_order = min(order_date)
emit lowest_price = min(unit_price)
Values compare by the rule sorting uses (see How values are ordered):
- Integers, floats and decimals compare by their exact value, so a column that holds both integers and floats (for example one built by an
ifwhose branches return an integer and a float) is compared value by value. - NaN is the largest value, above
inf. - Strings, dates and datetimes order as they do in a Sink sort.
The result does not depend on the order rows arrive in, or on the memory limit.
When several values in the group are equal under that rule, min returns the same one every time: an integer before a decimal before a float, the decimal with fewer fractional digits, and the float with a negative sign before one with a positive sign. So min of 1 and 1.0 is 1, min of the decimals 1.0 and 1.00 is 1.0, and min of -0.0 and 0.0 is -0.0.
max(expr) -> Any
Returns the maximum value in the group. Works on numeric, string, and date types. Null values are skipped; a group whose values are all null gives null.
cxl: |
emit latest_order = max(order_date)
emit highest_price = max(unit_price)
Values compare as they do for min: numbers by their exact value across integer, float and decimal, and NaN is the largest value, so a group that holds a NaN has NaN as its maximum. The result does not depend on the order rows arrive in, or on the memory limit.
When several values in the group are equal, max picks in the reverse order to min: a float before a decimal before an integer, the decimal with more fractional digits, and the float with a positive sign. So max of 1 and 1.0 is 1.0, max of the decimals 1.0 and 1.00 is 1.00, and max of -0.0 and 0.0 is 0.0.
collect(expr) -> Array
Collects all values of the expression into an array. Useful for building lists of values per group.
cxl: |
emit all_order_ids = collect(order_id)
Because collect emits an array, JSON writes it as a native array, XML as
repeated child elements, and CSV as a delimited cell. Coerce it to a scalar for
a format such as fixed-width (for example
emit ids = all_order_ids.join(";")).
weighted_avg(value, weight) -> Float or Decimal
Computes a weighted average: sum(value * weight) / sum(weight). Takes two arguments.
cxl: |
emit weighted_price = weighted_avg(unit_price, quantity)
weighted_avg(v, w) is sum(v * w) / sum(w), with each row’s v * w
computed as it is in any expression and both sums exact, rounded once as sum
rounds them. When either the value or the weight is a decimal, the result is
a decimal at full division precision, and it has the same digits and scale as
sum(v * w) / sum(w). Otherwise it is a float. Over floats each row’s
v * w is a float product, and the products and the weights are summed
exactly, so the average does not depend on row order or memory.limit. Over
integers alone the two exact totals are each converted once to a float and
divided.
These groups fail with an aggregate_finalize error rather than writing a
value:
- Zero total weight. The group’s weights add up to exactly zero, so the
average divides by zero, as
x / 0does in any expression. Rows whose weight is zero add nothing to the average, so the error’s fix drops them with a Transform before the Aggregate,config: { cxl: "filter qty != 0" }with your weight column in place ofqty. A group whose non-zero weights cancel, such as a sale and its return, still totals zero after that filter and still fails. - A row’s product out of range. A row’s decimal
value * weightis outside the decimal range. Retracting that row clears the error. The error’s fix computes the average in floats,weighted_avg(price.to_float(), qty.to_float())with your columns in place ofpriceandqty. - A total or the quotient out of range. A decimal total is outside the decimal range, or the quotient is because the weights nearly cancel.
- Decimal and float in one group, in one row or across rows (see above).
Mixing a decimal with a binary float across the two arguments is a type
error when the typechecker can see it. Declare the float column
type: decimal in its Source schema so both arguments are decimals, or, when
the float is computed rather than read from a Source column, convert the
decimal argument with .to_float(). A group with no row whose value and
weight are both non-null gives null.
Aggregates vs. windows
| Feature | Aggregate node | Window function |
|---|---|---|
| Record output | One row per group | One row per input record |
| Syntax | sum(field) (free-standing) | $window.sum(field) (namespace) |
| Configuration | type: aggregate + group_by: | type: transform + analytic_window: |
| Use case | Summarize groups | Enrich records with group context |
An Aggregate’s sum, avg and weighted_avg are exact (see
sum). A window function’s $window.sum and
$window.avg are not: they still add in the order the rows of the partition
are held, so a window sum over floats can differ in its last bits from the
Aggregate’s sum of the same values.
Combining aggregates with expressions
Aggregate function calls can be mixed with regular CXL expressions in emit statements:
nodes:
- name: category_stats
type: aggregate
input: products
config:
group_by: [category]
cxl: |
emit total_revenue = sum(price * quantity)
emit avg_price = avg(price)
emit margin_pct = (sum(revenue) - sum(cost)) / sum(revenue) * 100
emit product_count = count(*)
emit has_premium = max(price) > 100
Restrictions
letbindings in aggregate transforms are restricted to row-pure expressions (no aggregate function calls inlet).filterin aggregate transforms runs pre-aggregation – it filters input records before grouping.distinctis not permitted inside aggregate transforms. Place a separate distinct transform upstream.
Complete example
pipeline:
name: sales_summary
nodes:
- name: raw_sales
type: source
format: csv
path: sales.csv
- name: monthly_summary
type: aggregate
input: raw_sales
group_by: [region, month]
cxl: |
emit total_sales = sum(amount)
emit order_count = count(*)
emit avg_order = avg(amount)
emit top_sale = max(amount)
emit all_reps = collect(sales_rep)
- name: output
type: sink
input: monthly_summary
format: json
path: summary.json
This pipeline outputs JSON because all_reps = collect(sales_rep)
emits an array, which the tabular writers (CSV/XML/fixed-width)
reject; drop the collect binding or coerce it with a downstream
Transform to keep a CSV sink.
Closures
CXL supports arrow-syntax closures as arguments to closure-bearing array builtins like filter, map, find, any, and flat_map. They give CXL a way to express element-by-element predicates and projections over nested arrays carried inside a single record – without writing a separate transform node per element.
Syntax
it => expression
A closure has one parameter, named it, and a single expression body. The arrow => separates them.
- type: transform
name: filter_items
input: orders
config:
cxl: |
emit kept = items.filter(it => it["price"] > 5)
The body is an expression, not a block of statements. Use if/then/else or match if you need branching inside a closure.
cxl: |
emit price_buckets = items.map(it =>
if it["price"] >= 100 then "premium"
else if it["price"] >= 10 then "standard"
else "value")
Parameter name
The parameter is always it. Other identifiers are not accepted as the closure binding:
items.filter(item => item["price"] > 5) -- parse error
items.filter(it => it["price"] > 5) -- ok
it is recognized in expression position only inside a closure body. Outside of one, it has no special meaning.
Lexical capture
Inside the closure body, the outer record’s fields and let bindings remain visible. For each iteration the closure parameter it is bound to the current element, the body evaluates, then it is removed before the next iteration.
cxl: |
let threshold = 10
emit kept = items.filter(it => it["price"] > threshold)
Here the closure body reads both it (the current array element) and threshold (an outer let binding). The record’s fields are also reachable by name – a closure over items can still read customer_id, region, or any other field on the same record.
Where closures appear
Closures are valid only as method-call arguments to closure-bearing builtins. They cannot be assigned to variables, stored in fields, or passed to non-closure builtins:
let f = it => it * 2 -- rejected at resolve time
emit doubler = it => it * 2 -- rejected at resolve time
If you need to share a closure across multiple call sites, repeat the literal closure expression. CXL has no first-class function values.
Null propagation
Closure-bearing builtins applied to a null receiver return null without evaluating the body. The body is also never called on records where the array is null:
cxl: |
emit kept = items.filter(it => it["price"] > 5)
-- when `items` is null, `kept` is null; the body never runs
This matches the null-propagation policy on every other builtin – see Null Handling for the wider rules.
Worked example: filter and map over a nested array
Suppose each input record carries an items array of objects, each with sku and price:
{"order_id":"O-1","items":[{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}]}
A transform that drops cheap items and projects the remaining SKUs:
- type: transform
name: filter_items
input: orders
config:
cxl: |
emit order_id = order_id
emit kept = items.filter(it => it["price"] > 5)
emit kept_skus = items.filter(it => it["price"] > 5).map(it => it["sku"])
For the input above, the transform produces:
{
"order_id": "O-1",
"kept": [{"sku": "a", "price": 10}, {"sku": "b", "price": 20}],
"kept_skus": ["a", "b"]
}
Bracket-index access (it["price"]) reaches into each map element. See Nested Paths for the full traversal surface.
See also
- Array Methods – the closure-bearing builtins (
filter,map,find,any,flat_map). - Map Methods – callable on map elements inside a closure body.
- Nested Paths – bracket-index and dotted-path navigation through nested arrays and maps.
- Emit Each – statement that fans one input record into many output records, using a binding similar to the closure parameter.
Nested Paths
CXL records can carry nested arrays and maps as field values (for example, a JSON input where each record has an items array of objects). Reaching into that structure uses two complementary forms: dotted paths and bracket indices.
Dotted paths
A dotted identifier path reads a static field name from a map.
doc.metadata.tenant
Each segment must be a valid identifier. Dotted paths are resolved at compile time – the typechecker walks the structure declared in the source schema and reports a missing-field error if any segment doesn’t exist.
- type: transform
name: project_tenant
input: events
config:
cxl: |
emit tenant = doc.metadata.tenant
emit user_id = doc.user.id
Use dotted paths for structures whose shape is fixed and known at authoring time.
Bracket indices
A bracket index reads a runtime-computed key. The receiver may be an array (integer index) or a map (string index).
items[0]
profile["name"]
items.map(it => it["sku"])
Bracket indices are dynamic – the index expression evaluates per record. The typechecker treats the result as Any and does not assert that the key is present.
Integer index on an array
- type: transform
name: first_item
input: orders
config:
cxl: |
emit head = items[0]
emit second = items[1]
For items = [{"sku":"a"},{"sku":"b"},{"sku":"c"}], head is {"sku":"a"} and second is {"sku":"b"}.
Out-of-range indices return null. Negative indices also return null (CXL does not support negative indexing).
String index on a map
cxl: |
emit name = profile["name"]
emit tier = profile["tier"]
Missing keys return null – the lookup never raises an error. This is the same null-propagation policy closure builtins use on their receivers.
Mixing forms
The two forms compose in either order:
cxl: |
emit first_sku = items[0]["sku"]
emit profile_email = users.profile["email"]
items[0]["sku"] is two bracket indices chained – an integer index against the array, then a string index against the resulting map. users.profile["email"] walks a dotted path to reach profile (a map field on users), then bracket-indexes into it for a runtime key.
Null propagation
Every nested-access form propagates null end-to-end. If the receiver is null, the result is null without evaluating the index expression:
cxl: |
emit sku = items[0]["sku"]
-- when `items` is null, `sku` is null
-- when `items[0]` is null, `sku` is also null
This matches the null behavior on dotted paths and on method-call receivers. Records with missing intermediate structure produce nulls in their derived fields rather than aborting the transform.
Method calls on indexed values
A bracket-indexed expression is a regular value, so it composes with any method or further index:
cxl: |
emit head_sku_upper = items[0]["sku"].upper()
emit cheap_skus = items.filter(it => it["price"] < 10).map(it => it["sku"])
The first chain reads a string out of nested structure and uppercases it. The second filters an array of maps by a numeric field and projects the SKU strings out.
See also
- Field Paths – the other surface for the same idea: how a flat column-name string spells a path, with a backslash escape instead of brackets.
- Closures – closures over arrays of maps typically use bracket-index on the
itbinding. - Array Methods – traversal builtins that consume nested arrays.
- Map Methods – builders and accessors for map values.
- Null Handling – the wider null-propagation rules.
Field Paths
A column name is not opaque. An unescaped . inside it separates path
segments, so the column Address.City addresses the path Address → City.
Readers produce such names when they flatten nested input, writers expand them
back into nested output, and the rule for reading them is the same everywhere in
Clinker — one grammar, not one per format.
This page is the reference for that grammar. The places it applies:
- Column names declared in a source or output
schema:. - Column names a self-describing reader infers (the JSON reader flattens
{"a":{"b":1}}to the columna.b; the XML reader flattens<a><b>1</b></a>the same way). - Column names a writer expands back into nesting — see Writing JSON and Writing XML.
It does not cover CXL expressions. Reaching into a value — a map or array
held in one column — uses the expression forms on
Nested Paths (profile.city, profile["a.b"], items[0]).
The two are different surfaces over the same idea: a path is an ordered list of
segments either way, but a column name spells it with . and a backslash
escape, while an expression spells it with dots and brackets.
The rule
Reading a name left to right:
| Input | Meaning |
|---|---|
. | Ends the current segment, starts the next. |
\. | A literal . inside the current segment. |
\\ | A literal \. |
\[ | A literal [. |
\ before anything else | An error. \], \t, and a name ending in \ are all rejected. |
| Any other character | A literal character of the current segment — including a bare [, and ], @, $, /, and whitespace. |
So Address.City is two segments, and a\.b is one segment named a.b.
Three consequences worth stating outright:
- An unrecognized escape is an error, never a literal. A
\must be followed by.,[, or\— nothing else, and not the end of the name.C:\tempis rejected, because silently treating\tas the two characters\andtwould make the encoding ambiguous, and silently dropping the\would rename the column without saying so. WriteC:\\temp. The same applies to\], which is covered below. - Empty segments are real.
a..bis three segments —a, the empty name, andb. An empty key is a value a document can genuinely carry, so the grammar does not reject it. - A name may nest at most 64 levels deep. This matches the depth at which the JSON reader stops flattening, so any name that reader produces is a name a writer can expand. The XML reader flattens with no depth bound, so an extraordinarily deep XML document can produce a name past the cap; writing it back out fails with a clear error rather than recursing without limit.
Writing a literal bracket
[ is currently a literal character, so a column named a[0] works. It is
nonetheless reserved: bracket indexing may later be given meaning inside a
flat name, matching the [n] form CXL expressions already use. Writing \[
means “a literal [” today and will keep meaning exactly that. A bare [ is
not guaranteed to.
So the column a[0] future-proofs as a\[0] — escaping only the opening
bracket:
| Spelling | Result |
|---|---|
a\[0] | One segment named a[0]. Correct, and stable across the reserved-[ change. |
a\[0\] | Rejected — \] is not an escape. |
a[0] | One segment named a[0] today, but a bare [ is the form that is not guaranteed to keep that meaning. |
Only the opening bracket is ever escaped. ] is never escaped and never
needs to be: it would carry meaning only as the close of an unescaped [, so a
literal ] standing on its own is already unambiguous. Escaping it would add
noise without removing any ambiguity.
Writing \] is an error rather than a silently-accepted no-op, for the same
reason C:\temp is: an escape that quietly meant nothing would let two
different names decode to the same path, and one of them would be dropped. The
error names the offending escape and the rule.
If a column name of yours contains [, escaping it now costs nothing and makes
it future-proof.
Expansion on write
A writer that can express nesting rebuilds it from the decoded paths.
Grouping. Columns sharing a prefix collect into one container, positioned
where that prefix first appeared — even when the schema interleaves them.
Columns Address.City, name, Address.State write as:
{"Address":{"City":"Boston","State":"MA"},"name":"Ada"}
Absent children. A column omitted for a record (a null under
preserve_nulls: false) contributes nothing, and a container whose every
descendant is omitted emits no key at all rather than an empty object. That
keeps the round trip honest: reading {"a":{}} produces no column, so writing
no column must produce no "a".
Values are untouched. Expansion adds structure above a column’s value; a
Value::Map or array held in that column still serializes as itself. A column
Items.Item holding [1,2] writes as {"Items":{"Item":[1,2]}}.
No carve-outs. The rule reads the column-name string and nothing else, so it
applies identically to engine-stamped columns. Under
include_correlation_keys: true the column $ck.customer_id writes as
{"$ck":{"customer_id":…}}.
When two names clash
Two columns can describe places that cannot both exist. The writer refuses the whole column set before writing a single byte, naming both columns — it never silently keeps one.
| Columns | Why |
|---|---|
a and a.b | a holds a value and is also the container a.b sits inside. |
a.b and a.b.c | The same clash, one level down. |
a[b and a\[b | Two spellings of the identical path. |
a.b and a\.b do not clash: they address a → b and the single
segment a.b. That is what the escape is for.
When a clash comes from a . you meant literally, the diagnostic offers the
escaped spelling:
XML writer cannot expand this output's column names into nested output: field
names `a.b` and `a.b.c` cannot both be written: `a.b` holds a value and is also
the container `a.b.c` nests inside. Rename one of them, or — if the `.` in `a.b`
is part of the name rather than a nesting separator — declare it as `a\.b`.
Where the grammar does not reach yet
Two surfaces read column names without this grammar. Both are tracked, and both are stated here so the current behavior is a known position rather than a surprise:
- Flat writers emit the raw name. CSV and fixed-width have no nesting to
expand into, so they write the column name verbatim — a column declared
a\.bappears as the literal CSV headera\.b, backslash included. Whether flat writers should emit the decoded name instead is a separate decision. - Readers join without escaping. The JSON and XML readers join flattened
path segments with a plain
., so a source key that literally contains a.({"a.b": 1}) arrives as the columna.b— indistinguishable from a nested{"a":{"b":1}}, and it writes back nested. Closing that means escaping each key as the reader joins it, tracked by issue 920.
set and unset in CXL also use a path grammar of their own that does not yet
support escaping — see the known limitation on
Map Methods.
See also
- Nested Paths — the expression-side forms for reaching into a value.
- Writing JSON — expansion on the JSON write side.
- Writing XML — expansion on the XML write side, plus the attribute convention layered over it.
Emit Each
The emit each statement fans one input record into multiple output records – one per element of an array on the input. The body emits the fields each output record carries. A trailing outer modifier preserves the trigger row when the array is empty or null.
Syntax
emit each <binding> in <source> {
<statements>
}
<binding>is the identifier the body uses to refer to the current array element. The conventional name isit(same as the closure parameter), but any identifier is accepted.<source>is any expression producing an array. Typically a field reference on the input record.- The body is a block of
letandemitstatements that produce one output record per iteration.
Worked example
Suppose each input record carries an items array of objects, each with sku and price:
{"order_id":"O-1","items":[{"sku":"a","price":10},{"sku":"b","price":20},{"sku":"c","price":5}]}
A transform that fans each input into one record per item:
- type: transform
name: explode
input: orders
config:
cxl: |
emit each it in items {
emit order_id = order_id
emit sku = it["sku"]
emit price = it["price"]
}
For the input above, the transform produces three output records:
{"order_id":"O-1","sku":"a","price":10}
{"order_id":"O-1","sku":"b","price":20}
{"order_id":"O-1","sku":"c","price":5}
The body reads both it (the current element) and order_id (an outer record field). Outer-record fields remain visible inside the body for every iteration.
Cardinality
If the source array has N elements, emit each produces exactly N output records. Empty array sources produce zero records. A null source also produces zero records – no DLQ entry, no error – mirroring the explode-on-null convention used elsewhere in CXL.
Known issue: a top-level plain
emit eachover an empty array ornullcurrently passes the input row through as one record instead of producing zero records (#1350). This comes from reading the engine’s code; it has not been confirmed with a pipeline run. Until it is fixed, putfilter not items.is_empty()(with your own field foritems) before the block if you rely on the row being dropped:is_empty()is true for both an empty array andnull.
When fan-out nests, the cardinalities multiply: an outer array of M elements whose inner arrays have N elements each produces up to M×N records. The cumulative max_expansion cap bounds that product.
A non-array, non-null source raises a runtime type-mismatch error and routes the originating record to the DLQ.
Preserving the trigger row: outer
A trailing outer modifier switches emit each to its outer-join variant. The grammar is identical except for the keyword after the source:
emit each <binding> in <source> outer {
<statements>
}
The only behavioral difference is what happens when the source is null or an empty array. Plain emit each drops the trigger row entirely (zero output records). The outer variant instead emits the trigger row once, with <binding> bound to null:
| Source | emit each ... | emit each ... outer |
|---|---|---|
| 3-element | 3 records | 3 records (identical) |
| empty array | 0 records (see the known issue above) | 1 record, binding = null |
null | 0 records (see the known issue above) | 1 record, binding = null |
This is the shape SQL engines spell LATERAL VIEW OUTER EXPLODE (Spark, Hive) or an outer UNNEST (DuckDB): “for each tag on this article emit a tagged row, but keep articles that have no tags.”
Using the worked example above with an order that carries no items:
{"order_id":"O-2","items":[]}
- type: transform
name: explode_outer
input: orders
config:
cxl: |
emit each it in items outer {
emit order_id = order_id
emit sku = it["sku"]
emit price = it["price"]
}
produces a single record that keeps order_id while the per-item fields read through the null binding:
{"order_id":"O-2","sku":null,"price":null}
Outer-record fields (like order_id) and any emit statements preceding the block still apply to the preserved trigger row, so an outer row is never bare.
The source type rule is slightly wider than plain emit each: a statically-null source is accepted (it is the case the variant exists to handle), alongside arrays and Any. Everything else in this page — the cumulative max_expansion cap, the nesting rules, the body-statement restrictions — applies unchanged to the outer variant. The two variants compose freely: an outer block may nest inside a plain emit each block and vice versa.
Output schema
The body’s emit statements define the output record’s field set, the same way emit does in a regular transform body. Fields the body does not emit fall under the Sink node’s include_unmapped policy (see Sink Nodes).
Fields written by the body shadow same-named fields on the originating input record.
Nested fan-out: fan-out within fan-out
An emit each body may itself contain emit each blocks — fan-out within fan-out for one trigger row. This is the canonical “for each article, for each section, for each tag, emit a row” shape:
emit each section in article["sections"] {
emit each tag in section["tags"] {
emit article_id = article_id
emit section = section["name"]
emit tag = tag
}
}
For one input article, this produces one output record per (section, tag) pair. The inner binding (tag) reads the current inner element; the outer binding (section) and any outer-record field (article_id) stay visible inside the inner body. A field name reused as both an outer and inner binding shadows lexically — the inner binding wins inside the inner body, and the outer value is restored when the inner block finishes.
Emits are positional: an emit placed in the outer body before a nested block applies to every leaf record that block produces, but an emit placed after a nested block does not retroactively reach the records that block already emitted. Put the fields shared across leaves above the nested block.
Plain and outer blocks compose in any order. An inner plain emit each over an empty or null array contributes no records for that branch, while an inner emit each ... outer preserves one trigger row (inner binding bound to null) — exactly the per-level semantics from the single-level table, applied at each level.
Nesting is bounded to 32 levels so that adversarially deep input cannot exhaust the parser stack; legitimate document fan-out is only a few levels deep. Beyond that bound, parsing fails with a “nesting too deep” diagnostic.
The flat-array workaround (precompute a flattened array with .flat_map and use a single emit each) is still available and may be clearer for a simple two-level cartesian product, but is no longer required.
Body-statement restrictions
Within the body, let, emit, trace, and nested emit each / emit each ... outer are accepted. filter and distinct are rejected at evaluation time – a body filter would split work between branches the engine can’t represent. Move filter/distinct logic into a downstream transform, or pre-filter the source array with .filter before the emit each block.
Safety cap: max_expansion
By default, emit each can fan one input record into at most 10,000 output records. This limit is counted cumulatively across all nesting levels, so nested fan-out cannot multiply past it. A record that exceeds the limit routes to the DLQ with category expansion_limit_exceeded instead of producing an unbounded result. Override the limit with the max_expansion field in the transform config.
See Transform Nodes -> Expansion Cap for the YAML field and tuning guidance.
See also
- Closures – closures bind a similar
itparameter inside method calls. - Array Methods –
flat_mapis the in-expression cousin ofemit each. - Nested Paths – bracket-index access on the body binding.
- Transform Nodes – the
max_expansioncap and DLQ routing. - Error Handling & DLQ – DLQ category semantics.
System Variables
CXL provides several system variable namespaces prefixed with $. These give CXL expressions access to pipeline execution context, user-defined variables, per-record metadata, and the current time.
$pipeline.* – Pipeline context
Pipeline variables are accessed via $pipeline.member_name. Some are frozen at pipeline start; others update per record.
Stable (frozen at pipeline start)
| Variable | Type | Description |
|---|---|---|
$pipeline.name | String | Pipeline name from YAML config |
$pipeline.execution_id | String | UUID v7, unique per pipeline run |
$pipeline.batch_id | String | From --batch-id CLI flag, or auto-generated UUID v7 |
$pipeline.start_time | DateTime | Frozen at pipeline start, deterministic within a run |
$ cxl eval -e 'emit name = $pipeline.name' \
-e 'emit exec = $pipeline.execution_id'
{
"name": "cxl-eval",
"exec": "00000000-0000-0000-0000-000000000000"
}
Counters
These members exist, but nothing updates them while a pipeline runs, so each
one reads 0 in every expression (#1346).
Do not use them to count progress: a condition such as
$pipeline.total_count % 10000 == 0 is true for every record. The run’s real
counts are reported when it ends, on the summary line and in the metrics file;
see Where did my rows go? and
Metrics & Monitoring.
| Variable | Type | Names the count of |
|---|---|---|
$pipeline.total_count | Int | Records read |
$pipeline.ok_count | Int | Records that reached an output |
$pipeline.dlq_count | Int | Records sent to the dead-letter queue |
$pipeline.filtered_count | Int | Records excluded by filter statements |
$pipeline.distinct_count | Int | Records excluded by distinct statements |
$source.* – Per-record source lineage
$source.* exposes engine-stamped columns that travel with every
record from its origin Source node downstream through merges,
combines, and transforms. They identify where the record came
from and when in event-time it happened. All three columns are
filtered out of default Output projections — reference them
explicitly with emit if you need them in your output schema.
| Variable | Type | Description |
|---|---|---|
$source.file | String | Path of the input file the current record was read from. |
$source.name | String | Name of the Source node that produced the current record. Survives through merge / combine so downstream nodes can branch on origin. |
$source.event_time | DateTime | Engine-stamped event time, delay-corrected by the source’s watermark.delay. Null when the source has no watermark: block, or when the per-record value did not parse. |
filter $source.name == "src_web"
emit origin = $source.name
emit ingest_file = $source.file
emit ts = $source.event_time
$source.event_time is the column a
time-windowed aggregate
reads to assign records to windows. It is only populated for
records from a source that declares
watermark: — otherwise it
holds Null.
$vars.* – User-defined variables
User-defined variables are declared in the YAML pipeline config under pipeline.vars: and accessed via $vars.name in CXL expressions.
YAML declaration
pipeline:
name: invoice_processing
vars:
high_value_threshold: 10000
tax_rate: 0.21
output_currency: "USD"
fiscal_year_start_month: 4
CXL usage
filter amount > $vars.high_value_threshold
emit tax = amount * $vars.tax_rate
emit currency = $vars.output_currency
Variables provide a clean way to externalize configuration from CXL logic. Combined with channels, different variable sets can parameterize the same pipeline for different environments or clients.
$config.* – Composition config parameters
$config.<param> reads a composition’s declared config parameter from inside that composition’s body. It is only available in a composition body — a top-level pipeline declares no config schema, so $config.* there is a compile error.
Each parameter is declared in the composition’s _compose.config_schema: block, then read from the body’s CXL:
# in fraud_check.comp.yaml
_compose:
name: fraud_check
config_schema:
threshold: { type: float, default: 0.8 }
nodes:
- type: transform
name: flag
input: inp
config:
cxl: |
emit order_id = order_id
emit flagged = score >= $config.threshold
Unlike $vars.* (which flows to the executor as a runtime value), $config.<param> is constant-folded at compile time: each reference is replaced by the value resolved for that instantiation, so two call sites of the same composition with different config: compile to different bodies. The resolution precedence, highest first, is a channel/group config: clobber, then the call site’s config:, then the signature default.
Because the value is resolved per instantiation, overriding a config knob via a channel or group config: value clobber changes what the composition body computes — the override is applied to execution, and the winning layer is still recorded in the provenance side-table for channels resolve / explain --field.
$record.* – Per-record scoped state
$record.* is a per-record key-value store that travels with the record through the pipeline but never serializes as an output column. It is the mechanism for tagging records with quality flags, routing hints, or audit information that should not appear in the final output unless explicitly re-emitted as a regular column.
Each $record variable is declared in the writing Transform’s config.declares: block (scope: record) and written from that Transform’s CXL:
Writing record state
- type: transform
name: classify
input: orders
config:
declares:
- { name: quality, scope: record, type: string }
cxl: |
emit order_id = order_id
emit $record.quality = if amount < 0 then "suspect" else "ok"
Reading record state
Any downstream node reads it via $record.<key>:
filter $record.quality == "ok"
emit audit_quality = $record.quality
See Scoped Variables for the full declaration model and the pipeline / source / record lifetimes.
now – Current time
The now keyword returns the current wall-clock time as a DateTime value. It is evaluated fresh per record, so each record gets the actual time of its processing.
$ cxl eval -e 'emit timestamp = now'
{
"timestamp": "2026-04-11T15:30:00"
}
now is useful for timestamping records:
emit processed_at = now
emit days_old = now.diff_days(created_date)
Note:
nowis a keyword, not a function call. Writenow, notnow().
Complete example
pipeline:
name: order_enrichment
vars:
discount_threshold: 500
tax_rate: 0.08
nodes:
- name: orders
type: source
format: csv
path: orders.csv
- name: enrich
type: transform
input: orders
cxl: |
emit order_id = order_id
emit amount = amount
emit discount = if amount > $vars.discount_threshold then 0.1 else 0.0
emit tax = amount * $vars.tax_rate
emit total = amount * (1 - discount) + tax
emit processed_at = now
emit source_file = $source.file
emit pipeline_run = $pipeline.execution_id
- name: output
type: sink
input: enrich
format: csv
path: enriched_orders.csv
Null Handling
Null values in CXL represent missing or absent data. CXL uses null propagation – most operations on null produce null – with specific tools for detecting and handling nulls.
Interactive companion: the null explainer shows, step by step, how an expression is worked out when a field is null, and whether a filter keeps the record.
Null propagation
When a method receives a null receiver, it returns null without executing. This is called null propagation and applies to all methods except the introspection methods.
$ cxl eval -e 'emit result = null.upper()'
{
"result": null
}
Propagation flows through method chains:
$ cxl eval -e 'emit result = null.trim().upper().length()'
{
"result": null
}
Null propagation exceptions
Four methods are exempt from null propagation and actively handle null receivers:
| Method | Null behavior |
|---|---|
is_null() | Returns true |
type_of() | Returns "null" |
is_empty() | Returns true |
catch(x) | Returns x |
debug(label) is not one of them: on a null receiver it returns null without logging, like any other method.
$ cxl eval -e 'emit a = null.is_null()
emit b = null.type_of()
emit c = null.catch("fallback")'
{
"a": true,
"b": "null",
"c": "fallback"
}
Null coalesce operator (??)
The ?? operator returns its left operand if non-null, otherwise its right operand. It is the primary tool for providing default values.
$ cxl eval -e 'emit a = null ?? "default"
emit b = "present" ?? "default"'
{
"a": "default",
"b": "present"
}
Chain multiple ?? operators for fallback chains:
$ cxl eval -e 'emit result = null ?? null ?? "last resort"'
{
"result": "last resort"
}
Three-valued logic
Boolean operations with null follow three-valued logic: read null as “unknown”, and a result is decided only when the known side settles it.
and
| Left | Right | Result |
|---|---|---|
true | null | null |
false | null | false |
null | true | null |
null | false | false |
null | null | null |
The key insight: false and null is false because the result is false regardless of the unknown value.
or
| Left | Right | Result |
|---|---|---|
true | null | true |
false | null | null |
null | true | true |
null | false | null |
null | null | null |
The key insight: true or null is true because the result is true regardless of the unknown value.
not
| Operand | Result |
|---|---|
true | false |
false | true |
null | null |
Arithmetic with null
Any arithmetic operation involving null produces null:
$ cxl eval -e 'emit result = 5 + null'
{
"result": null
}
Comparison with null
== and != never produce null. Two nulls are equal, and a null is not equal to any other value:
$ cxl eval -e 'emit a = null == null
emit b = null != 100'
{
"a": true,
"b": true
}
Every other comparison (<, >, <=, >=) involving null produces null, as arithmetic does:
$ cxl eval -e 'emit result = null > 100'
{
"result": null
}
Null in conditions
A condition that comes out null is treated as “not true”:
filterkeeps a record only when its condition is exactlytrue. A null condition drops the record, just asfalsedoes.iftakes theelsebranch when its condition is null, or produces null when there is noelse.- A
matchwithout a subject skips an arm whose condition is null. Amatchwith a subject compares with==, so anull => …arm matches a null subject.
Because not null is also null, a record whose amount is null fails both filter amount > 100 and filter not (amount > 100). And because != never produces null, filter amount != 100 keeps it. Test for null explicitly with is_null(), or give a default first with ??, when it matters which way an empty value goes.
To test for null, use is_null():
$ cxl eval -e 'emit result = null.is_null()'
{
"result": true
}
Practical patterns
Fallback values with ??
emit name = raw_name ?? "Unknown"
emit amount = raw_amount ?? 0
emit active = is_active ?? false
Safe conversion with try_* and ??
emit price = raw_price.try_float() ?? 0.0
emit qty = raw_qty.try_int() ?? 1
Explicit null testing
filter not amount.is_null()
emit has_email = not email.is_null()
Catch method (equivalent to ??)
emit name = raw_name.catch("Unknown")
Conditional null handling
emit status = if amount.is_null() then "missing"
else if amount < 0 then "invalid"
else "ok"
Filter blank or null
# Filter out records where name is null or empty string
filter not name.is_empty()
Null-safe chaining
When working with fields that may be null, place the null check early or use ??:
# Safe: coalesce first, then transform
emit normalized = (raw_name ?? "").trim().upper()
# Safe: test before use
emit name = if raw_name.is_null() then "N/A" else raw_name.trim()
Modules and use
CXL modules organize reusable constants and pure, single-expression functions.
Module files use the .cxl extension and are admitted while Clinker plans the
pipeline. Execution uses the admitted declarations stored in the compiled plan;
it does not read module files again.
Module files
# rules/shared/finance.cxl
let tax_rate = 0.21
let default_currency = "USD"
fn tax(amount) = amount * tax_rate
fn normalize_currency(value) = value.trim().upper()
A module may contain:
letconstants whose expressions depend only on other constants and pure CXL operations;fndeclarations with named parameters and one expression body; andusedeclarations for other modules.
Functions cannot contain statements such as emit, filter, or distinct.
Recursive function calls and cyclic module imports are rejected during
planning.
Where use is recognized
Planning resolves module imports from every field that carries executable CXL, including:
- a Transform’s primary expression, validation checks, and per-record log conditions;
- an Aggregate’s expression;
- every Route condition;
- a Combine predicate and body;
- Envelope header and footer expressions;
- Reshape rule conditions, mutations, and synthesized overrides;
- Cull group-drop conditions; and
- the same fields inside reachable composition bodies.
Ordinary strings do not participate in module resolution. Node names,
validation and log messages, output paths, and other descriptive text cannot
introduce an import merely by containing text that resembles use.
Importing and using a module
Module identities and member access both use dot notation:
use shared.finance as finance
emit tax = finance.tax(amount)
emit currency = finance.default_currency
The alias is optional. Without as, the last identity segment is the alias:
use shared.finance
emit tax = finance.tax(amount)
There is no :: member syntax and no wildcard import. A missing member, calling
a constant, or reading a function without parentheses is a planning error with
the offending module and member named in the diagnostic.
Direct imports and private dependencies
Pipeline CXL can access only modules it imports directly. A module may import another module by its absolute logical identity:
# rules/app/invoice.cxl
use shared.finance as finance
let standard_rate = finance.tax_rate
fn invoice_tax(amount) = finance.tax(amount)
# pipeline transform
use app.invoice as invoice
emit tax = invoice.invoice_tax(amount)
shared.finance is included in the admitted transitive closure, but it is
private to app.invoice. The pipeline must add its own use shared.finance if
it needs to address that module directly. Dependencies are never re-exported.
Rules-root selection
Clinker selects exactly one rules root for non-catalog module identities. The precedence is:
- explicit
clinker run --rules-path <DIR>; pipeline.rules_pathin the pipeline YAML;[catalog].rules_rootinclinker.toml; then- the workspace-relative
rules/default.
There is no search path and no first-match shadowing. Every relative candidate is anchored to the selected workspace, not the process working directory or the pipeline file’s directory. See the CLI reference and typed workspace catalog.
An explicit [catalog.rules] entry maps a logical rule identity to a particular
workspace-contained file and takes priority over the derived
<rules-root>/<identity segments>.cxl path for that identity.
Planning bounds and diagnostics
Planning loads only the direct imports and their transitive dependencies. Each canonical module is parsed once. The default closure limits are:
| Limit | Default |
|---|---|
| One module file | 1 MiB |
| Unique modules | 64 |
| Import depth | 32 |
| Total closure source | 16 MiB |
Planning fails before execution for a missing or unreadable module, invalid UTF-8 or CXL, duplicate declarations or aliases, an import/function cycle, or a closure that exceeds a bound. Cycle diagnostics show the complete discovered chain so the import edge to remove is visible.
After loading the complete reachable closure, planning validates both declaration graphs:
- constant dependencies must be acyclic; and
- function calls must be acyclic, including direct, mutual, and cross-module recursion.
Cycle diagnostics report the complete chain with the relevant call or
declaration locations. Imported calls are also checked at the authored call
site. The diagnostic names the logical module and member when the member is not
a function, the argument count is wrong, or the expanded function body is
ill-typed. For example, if shared.numbers.add takes two arguments, the
corrected call is:
use shared.numbers as numbers
emit total = numbers.add(left, right)
Source-file lifetime
Module files are an input to planning, not a runtime dependency. Once planning succeeds, the compiled plan owns the immutable parsed declarations for every admitted direct and transitive module. The same plan can execute repeatedly if those source files are renamed, changed, or removed after planning. Changes take effect only after compiling a new plan.
Removing or changing a required file before planning still fails admission. This boundary prevents a checked plan from silently executing different module code and keeps execution independent of filesystem path authority.
Complete example
# rules/etl/clean.cxl
let max_amount = 999999.99
fn normalize_name(name) = name.trim().upper()
fn safe_amount(raw) = raw.try_float() ?? 0.0
fn flag_suspicious(amount, threshold) =
if amount > threshold then "review" else "ok"
# pipeline CXL
use etl.clean as clean
emit customer = clean.normalize_name(raw_customer)
emit amount = clean.safe_amount(raw_amount)
filter amount <= clean.max_amount
emit review_flag = clean.flag_suspicious(amount, 10000)
The cxl CLI Tool
The cxl command-line tool validates, evaluates, and formats CXL source files. It is the standalone companion to the Clinker pipeline engine, useful for testing expressions, validating transforms, and debugging CXL logic.
Commands
cxl check
Parse, resolve, and type-check a .cxl file. Reports errors with source locations and fix suggestions.
$ cxl check transform.cxl
ok: transform.cxl is valid
On errors:
error[parse]: expected expression, found '}' (at transform.cxl:12)
help: check for missing operand or extra closing brace
error[resolve]: unknown field 'amoutn' (at transform.cxl:5)
help: did you mean 'amount'?
error[typecheck]: cannot apply '+' to String and Int (at transform.cxl:8)
help: convert one operand — use .to_int() or .to_string()
cxl eval
Evaluate CXL expressions against provided data and print the result as JSON.
Inline expression:
$ cxl eval -e 'emit result = 1 + 2'
{
"result": 3
}
From a file with field values:
$ cxl eval transform.cxl \
--field Price=10.5 \
--field Qty=3
From a file with JSON input:
$ cxl eval transform.cxl --record '{"price": 10.5, "qty": 3}'
Multiple inline statements:
$ cxl eval -e 'let tax = 0.21
emit net = price * (1 - tax)' --field price=100
{
"net": 79.0
}
cxl fmt
Parse and pretty-print a .cxl file in canonical format with normalized whitespace and consistent styling.
$ cxl fmt transform.cxl
Output is printed to stdout. Redirect to overwrite:
$ cxl fmt transform.cxl > transform.cxl.tmp && mv transform.cxl.tmp transform.cxl
Input data
–field name=value
Provide individual field values as key-value pairs. Values are automatically type-inferred:
| Input | Inferred type | Example |
|---|---|---|
| Integer pattern | Int | --field count=42 |
| Decimal pattern | Float | --field price=10.5 |
true / false | Bool | --field active=true |
null | Null | --field value=null |
| Anything else | String | --field name=Alice |
$ cxl eval -e 'emit t = amount.type_of()' --field amount=42
{
"t": "int"
}
$ cxl eval -e 'emit t = name.type_of()' --field name=Alice
{
"t": "string"
}
–record JSON
Provide a full JSON object as input. Mutually exclusive with --field.
$ cxl eval -e 'emit total = price * qty' \
--record '{"price": 10.5, "qty": 3}'
{
"total": 31.5
}
JSON types map directly:
| JSON type | CXL type |
|---|---|
null | Null |
true / false | Bool |
| integer number | Int |
| decimal number | Float |
"string" | String |
[array] | Array |
{object} | Map |
Output format
Output is always JSON. Each emit statement produces a key-value pair:
$ cxl eval -e 'emit a = 1
emit b = "two"
emit c = true'
{
"a": 1,
"b": "two",
"c": true
}
Date and DateTime values are serialized as ISO 8601 strings:
$ cxl eval -e 'emit d = #2024-03-15#'
{
"d": "2024-03-15"
}
Exit codes
| Code | Meaning |
|---|---|
| 0 | Success (or warnings only) |
| 1 | Parse, resolve, type-check, or evaluation errors |
| 2 | I/O error (file not found, invalid JSON, etc.) |
Pipeline context in eval mode
When running cxl eval, a minimal pipeline context is provided:
| Variable | Value |
|---|---|
$pipeline.name | "cxl-eval" |
$pipeline.execution_id | Zeroed UUID |
$pipeline.batch_id | Zeroed UUID |
$pipeline.start_time | Current wall-clock time |
$pipeline.source_file | Filename or "<inline>" |
$pipeline.source_row | 1 |
now | Current wall-clock time (live) |
Practical usage
Quick expression testing:
$ cxl eval -e 'emit result = "hello world".upper().split(" ").length()'
{
"result": 2
}
Validate a transform file:
$ cxl check transforms/enrich_orders.cxl && echo "Valid"
Test conditional logic:
$ cxl eval -e 'emit tier = match {
amount > 1000 => "high",
amount > 100 => "med",
_ => "low"
}' \
--field amount=500
{
"tier": "med"
}
Test date operations:
$ cxl eval -e 'emit year = d.year()
emit month = d.month()
emit next_week = d.add_days(7)' \
--record '{"d": "2024-03-15"}'
Test null handling:
$ cxl eval -e 'emit safe = raw.try_int() ?? 0' --field raw=abc
{
"safe": 0
}
CLI Reference
Clinker ships two command-line tools: clinker (the pipeline runner) and cxl (the expression checker/evaluator/formatter, covered in the CXL CLI chapter). This page is the complete reference for clinker.
clinker run
Execute a pipeline.
clinker run [OPTIONS] <CONFIG>
Positional arguments
| Argument | Description |
|---|---|
<CONFIG> | Path to the pipeline YAML configuration file (required) |
Options
| Flag | Default | Description |
|---|---|---|
--memory-limit <SIZE> | YAML memory.limit, else 512M | Memory budget for the execution. Uses the same grammar as the YAML memory.limit: a byte count with an optional binary (1024-based) K/M/G suffix (K = 1024 bytes, M = 1024², G = 1024³), where a bare integer is bytes. Other forms — a decimal GB, an explicit GiB, or a fractional value such as 1.5G — are rejected. When the limit is approached, aggregation operators spill to disk rather than crashing. When passed, this value overrides any memory.limit set in the pipeline YAML; when omitted, the YAML value applies (or the 512M default when the YAML is also silent). An empty or whitespace-only value — as an ops wrapper produces when it forwards an unset variable, e.g. --memory-limit "$CLINKER_MEM" with CLINKER_MEM unset — is treated the same as omitting the flag. A non-empty malformed value (for example the decimal 4GB rather than the binary 4G) is rejected at the CLI boundary with an error naming --memory-limit and echoing the value, so a typo fails loudly instead of silently falling back to the default and shrinking a larger YAML budget. Because the flag simply populates pipeline.memory.limit, a startup budget error (E312) for a value you passed via --memory-limit refers to that same limit. |
--threads <N> | YAML pipeline.concurrency.threads, else number of CPUs | Positive capacity applied independently to the Rayon CPU-kernel pool and to concurrent Source schema/read work across top-level and composition-body Sources. It is not a total operating-system thread limit: Source workers and the Rayon pool remain distinct. The selected value is recorded in execution metrics. Zero is rejected before the config is opened. |
--batch-id <ID> | UUID v7 | Logical-batch correlation available as pipeline.batch_id, in {batch_id} output-path templates, machine events, and opt-in output provenance sidecars. Supplying it does not override the fresh UUIDv7 execution ID and does not provide deduplication, resume, or exactly-once behavior. It is not currently a field in the metrics-spool payload. |
--machine ndjson-v1 | – | Opt into the clinker.run schema-1 lifecycle on stdout. Requires a non-empty --batch-id; conflicts with plan/dry-run output and with --lineage - or --lineage-events -. File-based lineage remains compatible: a plan-only --lineage <FILE> export shares this stream’s identity and closes it with an explicit empty publication inventory, since it runs no attempt. Every line is one compact JSON object; human diagnostics move to stderr. Consumers must concurrently drain both pipes, reject unsupported schema majors, accept only additive schema-1 fields, and reconcile exactly one supported terminal with the actual process status and current-attempt artifact evidence. EOF, malformed output, forced termination, or a missing/duplicate terminal is incomplete, never success. See Running Clinker Directly or Under a Supervisor. |
--explain [FORMAT] | text | Print the execution plan and exit without processing data. Accepted formats: text, json, dot. With json or dot, standard output carries only the document and human diagnostics move to stderr, so a consumer can redirect stdout straight into a parser; with text they stay together on stdout. See Explain Plans. |
--lineage <PATH> | – | Preflight the workspace lineage identity policy, build column lineage, and write it as OpenLineage NDJSON, then exit without processing data. Give a file path, or - for stdout. The export is the whole invocation, so one that cannot be delivered exits non-zero rather than reporting success: a destination the exporter cannot write exits 4, and an event the [observability.lineage] byte caps reject exits 1. Each diagnostic names the destination, states which of the two failed, and prints the configuration change where one applies. A failed export leaves no partial file behind, so a following upload step cannot pick up a stale one. Both this flag and --lineage-events need the lineage capability, which the released binary has; a build compiled without it refuses the flag at validation rather than exiting zero having emitted nothing (see Optional capabilities). See Column Lineage. |
--lineage-events <PATH> | – | Preflight the workspace lineage identity policy, run the pipeline, and emit live OpenLineage run events (a START at run begin, then a terminal COMPLETE / FAIL / ABORT with real timing and row counts) as NDJSON to a file path, or - for stdout. Cannot be combined with --lineage, --explain, --dry-run, or -n. With -, normal run output can interleave with the event stream; use a file for clean NDJSON. See Live run events. |
--dry-run | – | With no -n, performs complete config, overlay, CXL, schema, DAG, resource, and publication-configuration validation, prints resolved outputs, and exits without opening or reading a Source and without opening or publishing a Sink. |
-n, --dry-run-n <N> | – | Bounded preview. Requires --dry-run and a positive N. Clinker checks the limit before every read and reads at most N records from each declared Source, including Sources inside composition bodies. Records drain in stable plan order to the explicit preview stream; configured Sink paths are never opened or published. Preview emits no live run telemetry or lineage lifecycle. |
--dry-run-output <FILE> | stdout | Destination for bounded-preview bytes. Requires --dry-run-n; without it, the option is rejected before config access. All preview Sinks write through this one explicit destination using their configured formats. |
--rules-path <DIR> | selected workspace’s rules/ | Select the CXL module rules root for this run. Precedence is explicit CLI value, then pipeline.rules_path, then [catalog].rules_root, then the workspace-relative rules/ default. A relative value is anchored to the workspace selected by --base-dir or workspace discovery, not the process working directory. One root is selected; Clinker does not search multiple roots. See Modules and use and the typed workspace catalog. |
--base-dir <DIR> | – | Base directory for resolving relative paths in the YAML config. Defaults to the directory containing the config file. |
--allow-absolute-paths | – | Permit absolute file paths in the pipeline YAML. By default, absolute paths are rejected to encourage portable configs. |
--env <NAME> | – | Sets CLINKER_ENV in the current process before the pipeline loads. The current run path does not otherwise consume that value for channel selection; select a channel explicitly with --channel. |
--quiet | – | Suppresses the “applied overlay” summary. Other stdout, tracing, warnings, and errors are not uniformly silenced. |
--force | – | Overrides an output’s if_exists: error policy and permits overwrite. Outputs using the default if_exists: overwrite already overwrite without this flag; unique_suffix keeps its own collision behavior. |
--log-level <LEVEL> | info | Closed logging level: error, warn, info, debug, or trace. Any other spelling is rejected. |
--metrics-spool-dir <DIR> | – | Directory for per-execution metrics files. See Metrics & Monitoring. |
--channel <ID> | – | Apply a logical id from [catalog.channels]. The selected file must also have a [catalog.pipelines] id listed in the channel manifest. Matching groups are target-bounded before labels narrow them. |
--group <NAME> | – | Force-include a group overlay by name (repeatable). The selected pipeline or one of its admitted compositions must appear in the group’s explicit targets: set. Use clinker channels resolve to preview the effective plan. |
--no-auto-groups | – | Suppress selector-derived group membership; only groups named with --group apply. |
--error-threshold is retired and rejected. Configure the typed pipeline
policy instead; the CLI diagnostic prints this paste-ready replacement:
error_handling:
type_error_threshold: 0.05
Credential profile foundation
The current binary does not yet accept a credential-profile option or a credential-profile configuration table. Referenced credentials therefore do not activate a source, destination, or observability exporter through this surface. Do not pass an environment, channel, or group as a substitute: those selectors never choose credentials, and there is no default or sentinel profile.
The run-local foundation that later preflight wiring will call is already bounded. Its default ceilings are 64 named profiles, 256 provider registrations across those profiles, 1 MiB of decoded profile/provider definition state, and 256 simultaneously retained handles. Admission checks all definition counts and bytes before a profile can resolve a requirement. Each live lease reports its exact retained bytes to the run memory arbitrator before provider allocation. The registry has no inbound producer and is not a backpressure target. An arbitrator spill callback queues a request and reports zero bytes freed synchronously; the next registry-owned checkpoint revokes and releases the partial set in reverse acquisition order and unregisters the registry. Cap, memory, and provider failures follow the same fail-closed cleanup path; an explicit run coordinator may pause acquisition until resume.
These are foundation limits, not newly available command-line behavior. A later complete preflight surface must add the one explicit profile selector, credential-required omission checks, and consumer activation together before the option can appear in the options table above.
Examples
# Basic execution
clinker run pipeline.yaml
# Production run with memory budget and forced overwrite
clinker run pipeline.yaml --memory-limit 512M --force --log-level warn
# Validate without processing
clinker run pipeline.yaml --dry-run
# Preview at most 25 records from each declared Source without publishing Sinks
clinker run pipeline.yaml --dry-run -n 25 --dry-run-output preview.csv
# Compile and explain without reading data
clinker run pipeline.yaml --explain text
# Show execution plan as Graphviz
clinker run pipeline.yaml --explain dot | dot -Tpng -o plan.png
# Run with a batch ID available to templates and provenance sidecars
clinker run pipeline.yaml --batch-id "daily-2026-04-11"
# Emit the bounded schema-1 lifecycle for a supervisor
clinker run pipeline.yaml --machine ndjson-v1 --batch-id "daily-2026-04-11"
The standalone command remains the common case and requires no supervisor. Machine mode adds a child-process control stream; it does not add scheduling, retry, heartbeat, or process-tree management to Clinker. A supervising parent must heartbeat independently of advisory progress and start a fresh process with a new execution ID for every retry.
Typed resource failures carry registry-owned messages and policy_required
retry advice. A supervisor must decide whether a fresh attempt is safe;
temporary-storage or delivery failure can follow bytes already accepted by a
destination and does not establish rollback.
| Resource condition | Failure code | Category |
|---|---|---|
| Memory budget refused | runtime.resource.memory_budget_exceeded | infrastructure |
| Allocation or allocation layout unsatisfied | runtime.resource.allocation_failed | infrastructure |
| Spill quota refused | runtime.resource.spill_cap_exceeded | infrastructure |
| Descriptor quota refused | runtime.resource.descriptor_cap_exceeded | infrastructure |
| Temporary storage or readback failed | runtime.resource.storage_failed | infrastructure |
| Continuation refused after delivery failed | runtime.resource.delivery_poisoned | infrastructure |
| Resource finalized or ownership authority violated | runtime.invariant.unknown | internal_invariant |
Explicit resource cancellation is an aborted run, reported as cancelled in
machine mode with exit 130 and no failure classification. It does not depend
on telemetry delivery or a pending shutdown signal. A genuine resource or data
failure retains its classification even when a shutdown signal is pending;
malformed source data remains source.data.invalid with do_not_retry advice.
clinker guess
Preview, exhaustively check, or safely write concrete int or float
replacements for inference-only numeric columns.
clinker guess [OPTIONS] <CONFIG>
With no selector, guess reads the base pipeline. --channel <ID> selects one
cataloged channel plus its target-admitted derived groups; --group <NAME>
selects one explicit, target-admitted group without a channel. The two
selectors conflict. A missing or ambiguous selector is an error rather than a
fallback to the base pipeline, so the report always describes exactly one
effective configuration.
| Flag | Description |
|---|---|
<CONFIG> | Pipeline YAML containing the source-schema numeric leaves to inspect. |
--channel <ID> | Select one cataloged channel and its derived, target-admitted groups. |
--group <NAME> | Select one explicit, target-admitted group without a channel. |
--field <NODE.COLUMN> | Narrow the preview to one numeric source field. Repeatable; repeated selectors are deduplicated in request order. In a multi-record schema, one selector covers every same-named literal numeric owner across records: while leaving same-named concrete declarations unchanged. Unknown, malformed, or entirely concrete fields are rejected. |
--check | Exhaust the frozen, capped manifest and exit 3 if any selected owner remains unresolved. |
--write | Exhaust evidence and edit exactly one inline, literal, single-owner numeric leaf after guarded compare-and-swap revalidation. Mutually exclusive with --check. |
--base-dir <DIR> | Workspace root holding clinker.toml and the channel/group roots. Defaults to the pipeline file’s directory. |
The command constructs readers through the same CSV, JSON, and XML option and
schema-coercion path as runtime ingest. It emits one deterministic JSON document
containing the selected configuration, bounded coverage, parser-owned numeric
evidence, unresolved reasons, proposed types, and an exact semantic YAML patch.
Preview and check never edit. Write reports written only after guarded
publication completes; otherwise it reports not_written and leaves the patch
for manual application. It never edits an overlay.
Each field report carries an owners array. Every owner has its own evidence,
votes, proposed type, unresolved reasons, and exact address. A single-record
field has one canonical /v1/schema/sources/.../columns/... owner.
Multi-record owners use
/v1/schema/sources/.../records/.../columns/..., in authored record order, so
same-named leaves never collapse into a fake single-record address or share a
proposal derived from another record type.
Discovery freezes the complete deterministic manifest up to 4,096 files and
reports its fixed-size path/order/size identity without listing every path.
Preview admits at most four files and 8 MiB of discovered file sizes globally,
round-robin across sources while preserving each source’s stable prefix. Every
selected reader enforces its discovered length and fails if the file is opened
at another size, truncates, or grows; it never hands a format reader bytes past
the admitted boundary. Preview then reads at most 1,024 records globally.
Coverage retains at most four file details per source and reports sampled,
truncated, uncovered, and unreported aggregate counts for the rest. Multi-pass
formats can report physical bytes_read above the admitted input size because
each pass is counted. --check reads every file and record in that same frozen
manifest instead of applying the preview budgets.
Candidate storage is capped at 100,000 source-schema leaves, matching the
canonical YAML parser’s 100,000-node limit; each YAML document is also capped
at 32 MiB. Each parser-owned NumericObservation retains at most 128 numeric
lexeme bytes, and each exact schema owner retains at most eight evidence items.
These limits appear in the JSON report. Candidate, observation, evidence,
coverage, and write-snapshot collections are fixed-bounded independently of
input size. Write retains one fixed-size digest per file in the 4,096-file
manifest and hashes through a fixed 1 MiB buffer; it does not register a
runtime memory consumer.
Write is eligible only for one resolved owner in the base pipeline whose exact
authored bytes and direct YAML provenance identify a literal inline numeric
leaf. It preserves all sibling bytes and proves by canonical reparse that only
that type leaf changed in the resolved staged configuration. Overlays,
external/generated or synthetic ownership, aliases, interpolation, symlinks,
non-local inputs, multiple owners, unresolved evidence, and input/config drift
are patch-only outcomes.
The CLI holds an advisory fs4 lock on the stable sibling
<CONFIG>.clinker-guess.lock file through sibling-temp flush/fsync, final
byte/semantic/input revalidation, atomic rename, and directory fsync. The lock
file remains beside the config so every cooperating replacement locks the same
inode. It must remain a regular non-symlink file and is owner-only on Unix. This
is a fail-closed compare-and-swap for cooperating writers using the same lock,
not a kernel-enforced conditional rename against a writer that ignores advisory
locking; such a writer can race after the final comparison.
Exit 0 means a complete preview was written, including an unresolved preview,
an exhaustive check resolved every owner, or one safe write was published.
Selection, configuration, and field errors exit 1; unresolved checks and
patch-only/no-edit writes exit 3; source discovery, input I/O, reader,
publication, signal-handler, and stdout failures exit 4; interruption exits
130 before any report is emitted or before publication begins.
# Preview every inference-only numeric leaf in the base pipeline
clinker guess pipeline.yaml
# Preview one field under a cataloged channel
clinker guess pipeline.yaml --channel acme --field orders.amount
# Preview one explicit group without a channel
clinker guess pipeline.yaml --group enterprise
# Exhaust the frozen manifest and make unresolved owners fail the command
clinker guess pipeline.yaml --check
# Exhaust evidence and edit one safe directly-authored owner
clinker guess pipeline.yaml --field orders.amount --write
clinker explain
Inspect one compiled field’s provenance or discover registry-owned diagnostic descriptors.
clinker explain <CONFIG> --field <PATH> [OPTIONS]
clinker explain --list [--status <STATUS>] [--category <CATEGORY>]
clinker explain --code <CODE>
Exactly one of --field, --list, or --code is required. A pipeline path is
required only for --field and is rejected for the two static discovery modes.
| Flag | Description |
|---|---|
--field <PATH> | Explain one exact or unambiguous shorthand field address in the compiled pipeline. |
--list | Print every registered descriptor in stable code order. |
--status <STATUS> | With --list, require active or retired-reserved. |
--category <CATEGORY> | With --list, require one of configuration, composition, source-and-expression, execution-and-format, terminal-authoring, security, or advisory. |
--code <CODE> | Print one registered descriptor and its optional longer detail page. |
--channel <ID> | With --field, apply the selected channel before compiling provenance. |
--group <NAME> | With --field, force-include a group overlay (repeatable). |
--no-auto-groups | With --field, suppress selector-derived groups. |
--base-dir <DIR> | Workspace root used by the field-provenance compile path; defaults to .. |
All seven descriptor fields—code, severity, status, category, retryability,
meaning, and correction—come from the same leaf registry in both discovery
views. Closed enum values in the descriptor use lowercase kebab-case; filter
spellings come from the same enum tables used to parse those filters. A detail
page can add examples but cannot define whether a code exists.
Unknown or empty filters, no-match combinations, unknown codes, and conflicting
modes exit nonzero. See Explain Plans for examples and for the
separate clinker run --explain plan display.
clinker metrics collect
Sweep per-execution metrics files from a spool directory into a single NDJSON archive.
clinker metrics collect [OPTIONS]
Options
| Flag | Description |
|---|---|
--spool-dir <DIR> | Spool directory to sweep (required). |
--output-file <FILE> | NDJSON archive destination (required). If the file exists, new entries are appended. |
--delete-after-collect | Remove spool files after they have been successfully written to the archive. |
--dry-run | Preview which files would be collected without writing anything. |
Examples
# Collect and archive, then clean up spool
clinker metrics collect \
--spool-dir /var/spool/clinker/ \
--output-file /var/log/clinker/metrics.ndjson \
--delete-after-collect
# Preview what would be collected
clinker metrics collect \
--spool-dir ./metrics/ \
--output-file ./archive.ndjson \
--dry-run
clinker channels
Inspect and validate the channel/group multi-tenant overlay system.
clinker channels resolve <TARGET> [OPTIONS]
clinker channels lint [OPTIONS]
clinker channels group members <GROUP> [OPTIONS]
clinker channels label set <KEY>=<VALUE> <CHANNEL_ID>... [OPTIONS]
clinker channels resolve
Renders the effective post-overlay plan for one target — the DAG plus per-value provenance (which layer supplied each value, and which group injected which node). This answers “what does tenant X actually run?”.
| Flag | Default | Description |
|---|---|---|
<TARGET> | – | Path to the base pipeline (or composition) YAML to resolve (required). |
--channel <ID> | – | Logical channel id from [catalog.channels]. Matching groups are derived only after explicit target admission. |
--group <NAME> | – | Force-include a group overlay by name (repeatable), subject to the same target set as automatic selection. |
--no-auto-groups | – | Suppress selector-derived group membership. |
--base-dir <DIR> | . | Workspace root holding clinker.toml and the channel/group roots. |
Exits non-zero when the overlay raises an error (e.g. a config key matching no
parameter), so resolve doubles as a targeted check for one tenant.
clinker channels lint
Compiles every target declared by every [catalog.channels] entry and reports
failures — the CI safety net for base-change blast radius. It uses the same
logical identity and target-scope checks as run and explain; a basename or
current working directory never supplies target identity.
| Flag | Default | Description |
|---|---|---|
--base-dir <DIR> | . | Workspace root to lint. |
Exits non-zero if any combination fails to compile or apply. Dangling splice anchors (an op referencing a missing node) and config keys matching no parameter are reported per combination.
clinker channels group members
Lists the channels whose labels currently satisfy a group’s selector — “who is
in this group right now?”. Because membership is derived from labels, this
evaluates the group’s match: selector against each channel’s manifest labels
through the same derivation the overlay resolver uses.
| Flag | Default | Description |
|---|---|---|
<GROUP> | – | Group name (the group.name of a *.group.yaml). |
--base-dir <DIR> | . | Workspace root holding clinker.toml and the channel/group roots. |
A group with no match: selector is explicit-only and reports no derived
members. A channel whose labels make the selector ill-typed or reference an
undeclared label is reported as a selector error (never a silent non-match), and
the command exits non-zero when any such error occurs.
clinker channels label set
Stamps (or overwrites) one label across the named channels by editing each
channel’s channel.cfg.yaml manifest in place. Idempotent: re-running with the
same value writes nothing. Only the manifest’s labels: block is rewritten;
other keys and comments are preserved. The channel manifest must already exist
with a non-empty channel.targets list; label set will not create a
targetless manifest.
| Flag | Default | Description |
|---|---|---|
<KEY>=<VALUE> | – | Label assignment. KEY must be an identifier (letters, digits, _) so a selector can reference it. VALUE is typed by YAML scalar inference (true/false → bool, integers → int, decimals → float, otherwise string). |
<CHANNEL_ID>... | – | One or more channel ids (tenant folder names) to stamp. |
--base-dir <DIR> | . | Workspace root holding clinker.toml and the channel root. |
Because group membership is attribute-derived, label set is the maintenance
operation for group membership: set a label once and every group whose selector
matches gains the channel — no membership list to hand-edit.
Examples
# What does tenant `globex` actually run for this pipeline?
clinker channels resolve pipeline/order_fulfillment.yaml --channel globex
# Preview a group overlay standalone (no channel)
clinker channels resolve pipeline/order_fulfillment.yaml --group enterprise
# Compile every channel/group overlay in the workspace and report failures
clinker channels lint
# Which channels are currently in the `enterprise` group?
clinker channels group members enterprise
# Onboard two tenants into the enterprise tier in one shot
clinker channels label set tier=enterprise globex acme-corp
clinker refactor
Structural refactors that span a base pipeline and every channel/group overlay that references it.
clinker refactor rename-node <TARGET> <OLD> <NEW> [OPTIONS]
clinker refactor rename-node
Renames a base node and propagates the rename to every overlay reference. The overlay op model addresses base nodes by name, so renaming a node otherwise breaks every overlay that referenced it. This command rewrites, in one operation:
- the base node’s
nameand every consumer’sinput:/inputs:/body:/header:/trailer:reference; - a Combine’s named-input map (qualifier key and/or upstream value) and — when
the Combine draws from the renamed node under a same-named qualifier — its
where:/cxl:bodies, rewritten via the CXL parser so only true source qualifiers are touched (a method receiver likeregion.contains(...)is left alone); - across every target-admitted group / channel-manifest / per-target overlay
file: op
target,after,before, injectedalias, explicitinput,rewirekeys and values, an inlinenode, aset config.cxlvalue’s CXL, and top-levelconfigdotted-path prefixes (old.param→new.param).
| Flag | Default | Description |
|---|---|---|
<TARGET> | – | Path to the base pipeline (or composition) YAML that declares the node. |
<OLD> | – | Current node name (must exist in the target). |
<NEW> | – | New node name — identifier only (letters, digits, _); must not already exist in the target. |
--dry-run | – | Print the diff of every file that would change without writing anything. |
--base-dir <DIR> | . | Workspace root holding clinker.toml and the channel/group roots. |
Ambiguity is guarded: renaming to a name that already exists in the target, or
renaming a node that does not exist, is a hard error. A Combine where:/cxl:
body that must be rewritten but does not parse aborts the whole operation before
anything is written. After a real (non-dry-run) run the command re-runs
channels lint so an incomplete rename fails loudly.
Scope is catalog- and target-bounded. Per-target files are matched by their
logical channel.target; channel manifests are admitted only when
channel.targets contains the selected pipeline; and groups are admitted only
when group.targets names that pipeline or a composition in its resolved
closure. Filenames and basenames never establish identity, and a selector match
cannot widen the refactor beyond the group’s declared target set.
Files are rewritten by re-serializing their YAML: key order is preserved, but
comments and incidental scalar styling are normalized. Use --dry-run to review
the exact on-disk diff first.
Examples
# Preview a rename across the base pipeline and every overlay that references it
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases --dry-run
# Apply it, then re-lint the workspace
clinker refactor rename-node pipeline/order_fulfillment.yaml orders purchases
clinker config
Inspect a pipeline configuration file.
clinker config --resolved <CONFIG>
clinker config –resolved
Prints the config with the multi-value shorthand expanded to canonical form.
The bare-field forms of split_to_rows:, split_values:, and join_values:
are rewritten to full mappings with every default spelled out — so you can see
exactly what the engine runs:
- a bare
- line_itemsundersplit_to_rows:becomes- { field: line_items, keep_empty: true, mode: extract }; - a bare
- tagsundersplit_values:becomes- { field: tags, delimiter: ";" }; - a bare
- tagsunderjoin_values:becomes- { field: tags, delimiter: ";", on_conflict: error, escape: "\\" }.
The rewrite is surgical: only those shorthand blocks change. Comments, key
order, indentation, and every other surface are preserved byte-for-byte, so the
output parses to a plan semantically identical to the input, and running
config --resolved on the result is a no-op. Schema columns are already
canonical (multiple: true is always written explicitly), so the schema block is
left untouched.
This is config canonicalization for the pipeline file itself. It is distinct
from clinker channels resolve, which renders the
effective post-overlay plan for a specific tenant.
A few surfaces are deliberately left as written rather than expanded, since
regenerating them would lose information: a shorthand block that carries an
interior comment or blank line between its items is passed through unchanged
(so the comment is never dropped), and a value written as a YAML alias
(*anchor) is left in place — the anchor it points to is expanded at its
definition, so the alias still resolves to the expanded value. The output uses
the input file’s line endings (LF or CRLF).
| Flag | Default | Description |
|---|---|---|
<CONFIG> | – | Path to the pipeline YAML config file (required). The file is validated before it is rewritten, so a malformed config fails with a config error rather than emitting a half-expanded document. |
--resolved | – | Print the fully-expanded canonical form to stdout. Currently the only mode; required. |
Examples
# Show the fully-expanded canonical form
clinker config --resolved pipeline.yaml
# Materialize the shorthand into a new file
clinker config --resolved pipeline.yaml > pipeline.canonical.yaml
See Source Nodes → Multi-value fields for the shorthand these forms expand from.
clinker attempts
Inspect and clean up retained publication attempts owned by a pipeline:
clinker attempts list <PIPELINE> [--path-execution-id <ID>] [--continuation <TOKEN>] [--show-paths] [--format text|json]
clinker attempts inspect <PIPELINE> --execution-id <ID> [--path-execution-id <ID>] [--show-paths] [--format text|json]
clinker attempts purge <PIPELINE> (--execution-id <ID> | --expired) [--path-execution-id <ID>] [--execute] [--continuation <TOKEN>] [--show-paths] [--format text|json]
<PIPELINE> is required and must be a traversal-free, workspace-relative
.yaml or .yml path. Every invocation reloads and compiles that pipeline,
then derives its finite destination-parent roots from the compiled config.
There is no option for supplying a storage root, deletion path, or safety
override.
When the original run used path- or overlay-affecting options, repeat them on
the attempt command: --base-dir, --allow-absolute-paths, --rules-path,
--channel, repeatable --group, and --no-auto-groups. Output templates that
use run identity also require the matching --path-execution-id, --batch-id,
or --timestamp. The path identity is deliberately distinct from inspect and
purge’s --execution-id selector, so purge --expired can reconstruct an
execution-scoped destination without changing its selector. Attempt operations
replay file-source discovery for {source_file} and {source_path} fan-out and
anchor a pipeline without --base-dir at the pipeline’s own directory, matching
run. These values recompile typed ValidatedPath roots and never grant
authority to a caller-supplied deletion path.
list and inspect never mutate retained state. purge is also non-mutating
by default: it reports the attempts that the current retention policy admits.
Only --execute performs bounded cleanup. Live locks, invalid ownership,
unsupported filesystem entries, ambiguous clocks, and unreadable manifests
remain keep decisions even with --execute.
| Command or flag | Behavior |
|---|---|
list | Lists retained attempts across all existing roots owned by the freshly compiled pipeline. |
inspect --execution-id <ID> | Reports one canonical execution ID across those roots. |
purge --execution-id <ID> | Previews one logical execution; add --execute to remove only positively owned, eligible files. |
purge --expired | Previews all policy-expired attempts admitted by the bounded page; add --execute to clean them. |
--continuation <TOKEN> | Resumes the exact plan-, root-, and selector-bound page emitted by a partial result. JSON resume_argv is authoritative; the text command applies platform quoting to the raw opaque token. |
--show-paths | Adds sanitized workspace-relative attempt paths. Machine-local prefixes and sensitive-looking components remain redacted. |
--format json | Emits one compact JSON object with stable field order and logical identifiers. The default is deterministic human-readable text. |
Default output is path-free. It contains the logical root ID, execution ID,
lifecycle state, eligibility, artifact IDs, cleanup debt, and exact bounds.
The compact JSON form carries the same fields plus shell-independent
recovery_argv and resume_argv arrays. Neither form includes record values,
credentials, secrets, or raw debug data.
Safety refusals and incomplete cleanup exit with status 4 and use the stable E371 or E372 data. The report includes its logical failure code, registry-owned retry advice, and a pasteable workspace-relative recovery command, for example:
diagnostic: E371
failure: attempt.retention.manifest_invalid
retry: policy_required
recover: clinker attempts inspect pipelines/orders.yaml --execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159
Examples:
# Path-free, non-mutating inventory
clinker attempts list pipelines/orders.yaml
# Inspect one retained execution as compact JSON
clinker attempts inspect pipelines/orders.yaml \
--execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159 \
--format json
# Preview expired cleanup, then perform the same bounded selection
clinker attempts purge pipelines/orders.yaml --expired
clinker attempts purge pipelines/orders.yaml --expired --execute
See Storage & Spill Location for retention, bounds, destination qualification, and cleanup ordering.
Environment Variables
| Variable | Description |
|---|---|
CLINKER_ENV | Active environment name. Equivalent to --env. Used by when: conditions in channel overrides to select environment-specific configuration. |
CLINKER_METRICS_SPOOL_DIR | Default metrics spool directory. Overridden by --metrics-spool-dir. |
Precedence (highest to lowest): CLI flag, environment variable, YAML config value.
Validation and Admission
clinker-plan is the authority that admits a pipeline to execution. A pipeline
is executable only after the planner has parsed canonical YAML, bound schemas
and compositions, type-checked Clinker Expression Language (CXL), and produced
a CompiledPlan.
Canonical planner validation
clinker run pipeline.yaml --explain text
This compiles the complete pipeline through clinker-plan, the sole authority
that can admit it for execution. The command checks:
- YAML structure and required fields
- CXL syntax and compile-time type checking
- Schema compatibility between connected nodes
- DAG wiring (no cycles, dangling inputs, or missing nodes)
- Plan-time source and output configuration gates
No runtime readers are opened and no output files are created. Planning may
inspect available file metadata or evaluate matchers for cost estimates. The
command exits with code 0 only after the planner produces a CompiledPlan, and
with code 1 for a configuration, schema, or plan diagnostic. Admission does not
prove that later input decoding or I/O will succeed, and the rendered plan does
not prove output correctness.
Composition resource descriptors and bindings are part of this admission. The
planner checks the bounded [catalog.resources] table, declared
_compose.resources_schema slots, call-site and overlay logical identities,
kind/capability compatibility, required slots, fixed locks, and recursive
composition bodies without resolving credentials or opening handles.
An ordinary composition call containing alias: or outputs: fails during
strict YAML parsing with E377 at the authored location. Replace alias: with
the composition node’s name:. Declare ports under _compose.outputs and
refer to them downstream as <composition-node-name>.<port>. The separate
add.alias field remains valid only within an overlay add operation.
Bare --dry-run performs the same planner compilation without rendering the
plan:
clinker run pipeline.yaml --dry-run
Prefer --explain text when reviewing a schema change because the resulting
plan is visible evidence of what the planner admitted. See Explain
Plans for text, JSON, and DOT plan output.
Guessing numeric types and repeated source fields
numeric is an authoring-only placeholder. Runtime planning still rejects it
with E158; use clinker guess to collect the real readers’ parser evidence and
produce an exact patch, then review and apply that patch before compilation.
clinker guess pipeline.yaml
clinker guess pipeline.yaml --field orders.amount --field orders.tax
clinker guess pipeline.yaml --channel production --check
clinker guess pipeline.yaml --field orders.amount --write
With no selector, the base pipeline is inspected. Exactly one --channel ID
or --group NAME selects an effective configuration; the two options conflict.
Repeatable --field node.column selectors narrow literal numeric leaves or
select one concrete column from a single-record CSV, JSON, or XML source for
multiplicity review. Numeric selectors retain their existing meaning and can
represent more than one authored multi-record leaf; the report gives every
exact owner address separately. Concrete multiplicity candidates are limited
to directly authored columns that do not already declare multiple: true.
The default preview is deterministic, bounded, and read-only. It freezes the
configured stable file order (name ascending by default) and reports the
fixed-size identity of its normalized paths, order, and sizes, then allocates
four file opens, 1,024 records, and 8 MiB of admitted file sizes globally in
round-robin source/file order. Each selected reader is pinned to its discovered
length: a pre-open mismatch, truncation, or growth fails instead of expanding
the preview, and no format reader receives bytes beyond that admitted length.
Multi-pass formats may report more physical bytes_read because each bounded
pass rereads the same admitted input. The manifest itself is capped at 4,096 files;
narrow a matcher or use files.take_first /
files.take_last if the selected set is larger. The YAML/configuration cap
limits candidates to 100,000 source-schema leaves. Per owner, at most eight
representative observations are retained, each with at most 128 bytes of
numeric lexeme evidence. Coverage retains at most four file details per source
and reports aggregate sampled, truncated, uncovered, and unreported counts for
the rest. Multiplicity inference retains counters and a fixed interpretation
set, never field values or a raw sample corpus. More than 16 distinct CSV
delimiter candidates is review-only. These fixed bounds are also printed in
the JSON report when they apply.
The manifest identity covers normalized path, configured order, and discovered size; it is not a content hash or a compare-and-swap proof. Preview and check perform no edit. Write additionally streams an exact BLAKE3 snapshot of every file in the capped manifest before evidence collection and compares it again after collection and immediately before publication.
--check uses the same frozen, capped manifest but reads every selected file
and record. It is exhaustive over that manifest rather than subject to the
preview’s open/record/byte sampling budgets. --write is equally exhaustive
and edits only when exactly one resolved owner is directly authored in the
base pipeline. Numeric evidence may replace one literal numeric leaf.
Multiplicity evidence may set one column’s multiple: true and, for CSV, add
one complete split_values entry using the proven delimiter and activated
escape. Both are one owner mutation. An overlay, external/generated owner,
alias, interpolation, existing conflicting split declaration, no-op
already-multiple column, symlink, non-local input, multiple owners, unresolved
evidence, or changed snapshot leaves the pipeline untouched and reports the
patch with exit 3. The edit is reparsed and compared with a typed expected
configuration, so comments, ordering, spans, and every unrelated scalar remain
unchanged.
Publication holds an advisory fs4 lock on the stable sibling
<CONFIG>.clinker-guess.lock file through a sibling-temp flush/fsync, final
exact byte/semantic/input revalidation, atomic replacement, and parent directory
fsync. The lock file remains beside the config so cooperating replacements keep
one lock inode across renames; it must remain regular, non-symlinked, and
owner-only on Unix. This catches any change visible at the final comparison,
but is not a kernel-enforced content-conditional rename: a writer that ignores
the advisory lock can still race after the final check. Use one cooperating
configuration writer per file.
Numeric votes come only from the parser-owned observations used by the shared
runtime reader construction. Exact integers vote int; finite, representation-
safe values vote float. Mixed integer/float evidence resolves to float only
when every integer is exactly representable there. Numeric defaults vote
through the schema parser. Accepted missing/null/empty states abstain but remain
reported, forbidden absence is a conflict, and all-no-value evidence remains
unresolved. No confidence threshold or statistical guess is used.
Repeated-value evidence
Multiplicity is proved per logical record; counts from separate records are never added together. The production reader runs against a temporary schema clone so it can retain ordered repeated values for observation without changing the effective pipeline:
- XML becomes conclusive when one record contains two or more sibling elements at the selected path. A sibling in each of two records is still unconfirmed.
- JSON becomes conclusive when one record contains an array longer than one. Null, empty, and one-element arrays remain unconfirmed.
- CSV becomes conclusive only when exactly one delimiter/activated-escape interpretation parses and re-encodes every non-null cell to the original bytes in the source’s declared character set, and at least one cell produces multiple values. Two surviving interpretations are review-only.
For example, each selected tags field below uses the same existing schema
surface:
schema:
- name: tags
type: string
Conclusive XML has two siblings in one row:
<root><row><tags>a</tags><tags>b</tags></row></root>
Conclusive JSON has an array longer than one:
[{"tags": []}, {"tags": ["a"]}, {"tags": ["a", "b"]}]
Conclusive CSV has one lossless interpretation:
tags
a|b
plain
The corresponding safe CSV edit reuses the normal multi-value syntax:
split_values:
- field: tags
delimiter: "|"
schema:
- name: tags
type: string
multiple: true
These inputs remain review-only or unconfirmed and cannot write:
[{"tags": []}, {"tags": ["a"]}, {"tags": ["b"]}]
tags
a|b;c
d|e;f
Run the exhaustive gate before requesting a write:
clinker guess pipeline.yaml --field values.tags --check
clinker guess pipeline.yaml --field values.tags --write
| Exit | Meaning |
|---|---|
| 0 | Preview completed, including a preview with unresolved owners; exhaustive check resolved every owner; or write published its one safe edit. |
| 1 | Configuration or selection error. |
| 3 | Exhaustive check is unresolved, or write emitted a patch but did not safely edit. |
| 4 | Source discovery, reader, I/O, signal-handler, or report-output failure. |
| 130 | Interrupted before a complete report could be emitted. |
Inspect outcome, every owner-level unresolved_reasons entry, coverage, and
the emitted patch. A preview exit of 0 is not proof that every selected field
resolved; use --check when the exit status must enforce that condition.
Bounded execution preview
clinker run pipeline.yaml --dry-run -n N executes a bounded sample after
planning. N must be positive; the runtime checks the limit before each read
and reads at most N records from each declared Source, including Sources
inside composition bodies. This is a source-record limit, not an output-row
limit: filtering, aggregation, joins, and fan-out can change the output count.
Configured Sink paths are never opened or published. Instead, all preview
Sinks serialize in stable plan order to stdout, or to the one explicitly
selected --dry-run-output PATH. That explicit destination is written; choose
it deliberately. Preview emits no live run telemetry or lineage lifecycle.
Without -n, --dry-run performs planning only and reads no input records.
--dry-run-output requires -n, and -n requires --dry-run. A successful
sample validates only the sampled data; it does not prove that the rest of
the input will pass. See CLI options.
Known limitation: At revision
3b343a4e, some bounded previews fail during cleanup withcompleted node-buffer scope retained ..., even when the same pipeline completes in an ordinary run. Treat that nonzero exit as a failed preview. Planning-only validation remains available; use an ordinary run with disposable output paths to verify actual results.
Advisory workspace schema analysis
clinker-schema is a separate advisory authoring library. It reuses the
canonical typed YAML parser for pipeline structure and schema references, but
it is not called by clinker run. Its warnings and coverage status cannot
admit or reject execution, and an advisory result never overrides the planner.
| Status | Exact meaning |
|---|---|
analyzed | Every applicable facet represented by the advisory model was inspected. This is not planner acceptance. |
partial | Some applicable content was inspected, but an unsupported shape or bounded limit left a gap. Read the attached reasons, then run the canonical planner check. |
skipped | The artifact had no applicable advisory content, such as no external schema reference or no fields to inspect. This is not success. |
failed | The artifact could not be inspected safely or structurally, such as a read, parse, or reference-resolution failure. Diagnose the reason and still use the planner as the execution authority. |
Known advisory limits are explicit rather than silently treated as valid:
- The advisory schema model covers linked external
.schema.yamlmetadata. Inline column lists, generated schemas, multi-record schemas, and planner-owned external schema shapes remain planner concerns; an applicable unsupported external shape reportspartial. - Transform field scanning is a conservative heuristic. It does not reproduce
CXL parsing, name resolution, schema flow, composition binding, or type
checking; when linked schema content is otherwise analyzable, a transform
therefore makes that advisory coverage
partial. - Array element types and object shapes without declared child fields are not modeled completely. Format matching is also unavailable for fixed-width and SWIFT sources, and the advisory Parquet token is not a supported pipeline source format.
- Analysis retains at most 10,000 field descriptors to depth 64, 4,096 schema
references, and 1,024 reasons. Reaching a bound reports
partialrather than dropping the gap.
After reading an advisory report, make the authoring decision against the canonical result:
clinker run pipeline.yaml --explain text
Recommended runtime-validation workflow
- Run
clinker run pipeline.yaml --explain textfor planner admission and inspect the compiled DAG. - Use bare
--dry-runonly when a quiet canonical compile check is preferable. - Run representative data against an isolated destination and inspect it.
- Run the full job only after checking the representative result and destination policy.
Explain Plans
The --explain flag prints the execution plan – the DAG of nodes, their connections, and the parallelism strategy the optimizer has chosen – without reading any data.
Text format
clinker run pipeline.yaml --explain
# or explicitly:
clinker run pipeline.yaml --explain text
The text format shows a human-readable summary of the execution plan:
Execution Plan: customer_etl
============================
Node 0: customers (Source, parallel: file-chunked)
-> transform_1
Node 1: transform_1 (Transform, parallel: record)
-> route_1
Node 2: route_1 (Route, parallel: record)
-> [high] output_high
-> [default] output_standard
Node 3: output_high (Sink, parallel: serial)
Node 4: output_standard (Sink, parallel: serial)
Key information shown:
- Node index and name – the topological position in the DAG. Under
dlq_granularity: documentevery Sink is listed after every other node, which is the order the run dispatches them in (see Document-level DLQ). - Node type – Source, Transform, Aggregate, Route, Merge, Sink, Composition
- Parallelism strategy – how the optimizer plans to execute the node
- Connections – downstream nodes, with port labels for route branches
- Buffer class (Physical Properties section) –
buffer: streamingfor a node that hands its output straight to a single downstream consumer, orbuffer: materializedfor one that holds a whole stage’s output in an inter-stage buffer. See Streaming vs. Blocking Stages for the distinction.
The buffer class is a pre-runtime signal for memory pressure: a materialized node holds its rows against pipeline.memory.limit and may spill to disk once the budget is tight, while a streaming node holds only a small in-flight slice. Use the annotation alongside --memory-limit / pipeline.memory.limit to predict which stages will dominate memory before running the pipeline.
When a Sink declares sort_order, text output also includes a terminal writer
decision:
=== Sink Writer Ordering ===
sink.export:
terminal_order: customer_id asc, created_at desc
disposition: deferred_sort
boundary_mode: records_only
partition_scope: global_split_sequence
disposition: proven_terminal_sort means the final upstream Sort already
establishes the exact authored order; its name appears as proven_by. A
deferred_sort is enforced over the complete writer population and can use the
bounded spill path. boundary_mode names the physical write path.
partition_scope says how wide the promise is: for example,
global_split_sequence is one order across all numbered split files, while
per_source_file is an independent order for each fan-out destination. The
section is absent when no terminal order was authored.
Dead-letter output
When the pipeline has an error_handling.dlq block, the text output ends with
a === Dead-Letter Output === section. It lists every DLQ file the run can
write, with the header the compiled plan fixed for it, so you can check the
columns before any data is read:
=== Dead-Letter Output ===
rejects.csv
sources: (pipeline-wide fallback)
columns: _cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file, _cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field, _cxl_dlq_triggering_value, _cxl_dlq_error_category, _cxl_dlq_error_detail, _cxl_dlq_stage, _cxl_dlq_route, _cxl_dlq_trigger, order_id, order_total, _cxl_dlq_source_record
refunds_rejects.csv
sources: refunds
columns: _cxl_dlq_id, _cxl_dlq_trigger_id, _cxl_dlq_timestamp, _cxl_dlq_source_file, _cxl_dlq_source_name, _cxl_dlq_source_row, _cxl_dlq_triggering_field, _cxl_dlq_triggering_value, _cxl_dlq_error_category, _cxl_dlq_error_detail, _cxl_dlq_stage, _cxl_dlq_route, _cxl_dlq_trigger, refund_id, refund_amount, reason, _cxl_dlq_source_record
Each entry gives:
- The path that names the file, as written in the YAML.
sources: the Sources whoseper_source.<name>.pathroutes to this file. The pipeline-widepathis labelled(pipeline-wide fallback): it takes the rows of every Source without aper_sourcepath, and rows that carry no Source identity.columns: the file’s complete header, in order. A row that lacks one of these columns writes an empty cell.
The files are listed with the pipeline-wide file first, then each
per_source file in Source-name order. The section is absent when the
pipeline has no dlq block. How the columns are chosen is described in
Error Handling.
JSON format
clinker run pipeline.yaml --explain json
Standard output carries only the JSON document: plan warnings and other human
diagnostics are written to stderr in this format and in dot, so redirecting
stdout into a parser is safe. Read stderr as well if you want to see them –
under --explain text they remain on stdout alongside the plan.
Produces a machine-readable JSON object for programmatic consumption. Useful for:
- CI pipelines that need to assert plan properties
- Custom dashboards that visualize execution plans
- Diffing plans between config versions
An authored Sink order adds a writer_boundaries array. Each entry carries the
Sink name, structured terminal_order fields and directions, the
terminal_order_label, disposition, optional proven_by, boundary_mode,
and partition_scope. The key is omitted when no Sink declares an order.
A pipeline with an error_handling.dlq block adds a dead_letter object with
the same information as the text dead-letter section.
Its buckets array has one entry per DLQ file, in the same order:
"dead_letter": {
"buckets": [
{
"path": "rejects.csv",
"sources": [],
"fallback": true,
"header": ["_cxl_dlq_id", "_cxl_dlq_trigger_id", "_cxl_dlq_timestamp", "...", "order_id", "order_total", "_cxl_dlq_source_record"]
},
{
"path": "refunds_rejects.csv",
"sources": ["refunds"],
"fallback": false,
"header": ["_cxl_dlq_id", "_cxl_dlq_trigger_id", "_cxl_dlq_timestamp", "...", "refund_id", "refund_amount", "reason", "_cxl_dlq_source_record"]
}
]
}
sources lists the per_source names routed to the file. fallback is
true for the pipeline-wide file. header is the complete header, including
the _cxl_dlq_* columns (shortened to "..." above). The key is omitted when
the pipeline has no dlq block.
# Compare plans before and after a config change
clinker run old.yaml --explain json > plan_old.json
clinker run new.yaml --explain json > plan_new.json
diff plan_old.json plan_new.json
Graphviz DOT format
clinker run pipeline.yaml --explain dot
Produces a Graphviz DOT graph. Pipe it to dot to render an image:
# PNG
clinker run pipeline.yaml --explain dot | dot -Tpng -o pipeline.png
# SVG (scalable, good for documentation)
clinker run pipeline.yaml --explain dot | dot -Tsvg -o pipeline.svg
# PDF
clinker run pipeline.yaml --explain dot | dot -Tpdf -o pipeline.pdf
This requires the graphviz package to be installed on the system.
The resulting diagram shows:
- Nodes as labeled boxes with type and parallelism annotations
- Edges as arrows with port labels where applicable
- Branch/merge fan-out and fan-in structure
- Terminal order, writer disposition, boundary mode, and partition scope on a
Sink that declares
sort_order
When to use explain
- During development – verify the DAG shape matches your mental model before writing test data.
- After adding route or merge nodes – confirm branch wiring is correct.
- When tuning parallelism – check which strategy the optimizer selected for each node.
- In code review – generate a DOT diagram and include it in the PR for visual confirmation.
Explain parses the YAML and builds the plan without opening runtime readers or processing records. Planning may inspect source metadata or matchers for cost estimates, but it does not create pipeline outputs.
clinker run pipeline.yaml --explain # parse, compile, print the plan
clinker run pipeline.yaml --dry-run # parse and compile without printing the plan
Both commands perform the same compile-time checks: schema binding, CXL type
checking, DAG wiring, and plan-time source and output gates. --explain also
renders the compiled plan; bare --dry-run is the quieter validation form.
Neither command opens runtime readers, processes records, or creates pipeline
outputs.
Retraction section
If at least one Aggregate has a group_by that omits a correlation-key field, the output includes a === Retraction === block. It lists which aggregates and windows use group-atomic retraction (see Correlation Keys) and a rough per-row memory estimate for each, so you can gauge the memory cost before a production run. The block is absent on pipelines that don’t use this mode.
Exact group sizes are unknown until the pipeline runs, so treat the estimates as a planning aid and confirm the live shape with clinker metrics collect after the first run.
Statistics
When the plan carries column statistics, the output ends with a === Statistics === section. Each figure is tagged with where it came from:
- Row counts — an estimate per source. A
[file metadata]figure is estimated from the input file’s size before any record is read; a[exec sketch]figure is an exact count measured during an actual run. These row counts are what the optimizer uses to pick a Combine’s join strategy. - Column sketches — distinct-value counts and frequent-value hints that a Combine gathers over its join keys while records flow, used to speed up matching.
A statistic that was never gathered renders as null rather than a fabricated zero — for example, a multi-file glob source or a network source whose size cannot be read adds no Statistics section at all.
Field provenance
clinker explain <pipeline> --field <path> traces where a single resolved value
comes from across every configuration layer, printing the winning layer plus
each shadowed layer and its source span. The path arity selects what is traced:
<node>.<param>(two parts) — a composition config parameter, resolved across composition defaults and channel/group overlays.<source>.<column>.<attribute>(three parts) — a source-schema attribute (type,scale,precision,format,width,required, …), resolved across the schema-provenance layersBase < Pipeline < Group < Channel.Baseis the source’s own declaredschema:; the higher layers are thepatch_schemaoverlay ops each channel/group applies.
# Where does the `scale` on the orders source's `amount` column come from?
clinker explain pipeline.yaml --field orders.amount.scale
# Resolve the same attribute with a channel overlay applied first.
clinker explain pipeline.yaml --field orders.amount.scale --channel acme_prod
Field: orders.amount.scale
Resolved value: 2
Provenance chain (outermost to innermost):
[WON] Channel → 2 (line 12)
Pipeline → 0 (shadowed) (line 5)
Base → 0 (shadowed)
The [WON] marker names the layer whose value survives; shadowed layers show
what they proposed. An unknown source, column, or attribute is rejected with a
hint listing the valid names at that level.
Reading a plan-time failure
A pipeline that fails a plan-time check never reads any input. The failure is printed before the run starts, and it carries four things:
E363
× source "src": `record_path` "$.rows" starts with the JSONPath root marker
│ `$.`, which is not part of the grammar; `record_path` is a dot-separated
│ path of object keys, descended from the document root (for example
│ `data.rows`). Write "rows" instead
╭─[pipeline.yaml:4:1]
3 │ nodes:
4 │ - type: source
· ────────┬───────
· ╰── declared here
5 │ name: src
╰────
help: `record_path` on a `json` source is a dot-separated path of object
keys descended from the document root: no `$.` JSONPath root marker,
no leading `/`, and no empty segments. It takes precedence over
`format:`, so pair it with `format: object` or leave `format:` off.
Omit `record_path` entirely and the reader auto-detects the document
shape. Run `clinker explain --code E363` for the full grammar.
- The code (
E363) heads the report. Where a page exists for it, hand it toclinker explain --codefor the worked example. - The message names the offending input and the rule it broke.
- The source line is quoted from your YAML, with the offending node underlined.
- The
help:paragraph names the fix. When the gate does not already say so, aSee: clinker explain --code <CODE>line is appended.
Warnings are reported the same way but marked ⚠ rather than ×, so an
advisory is distinguishable from the diagnostic that stopped the run.
The same report is printed under --explain, which compiles the plan before
printing it.
Two notes on where the snippet comes from:
- A pipeline that pulls in a composition body is reported without the quoted source line. A plan-time diagnostic carries a line number but not which file it belongs to, so rather than risk underlining an unrelated line, the report gives the code, message and help alone.
- A channel/group overlay suppresses the snippet only when it rewrites the
compiled config through structural ops, source patches, or composition
config:values. A selection that contributes only runtime vars leaves the pipeline document unchanged, so its snippet remains safe and is retained. - Bare
--dry-runcompiles the plan and prints the same report without reading source data.
Looking up diagnostic codes
clinker explain --list enumerates every registered diagnostic in stable code
order. Each entry includes its code, severity, status, category, retryability,
meaning, and correction. Closed enum values in this descriptor use lowercase
kebab-case. Narrow the list with exact filters:
clinker explain --list
clinker explain --list --status retired-reserved
clinker explain --list --category source-and-expression
The status vocabulary is closed:
| Status | Meaning |
|---|---|
active | The code describes a condition in the current authoring surface. |
retired-reserved | The old condition is no longer accepted, but its identifier remains permanently reserved and still explains the paste-ready correction. |
Categories are configuration, composition, source-and-expression,
execution-and-format, terminal-authoring, security, and advisory.
Unknown or empty filter values fail; a valid combination that matches no code
also fails instead of printing an ambiguous empty result.
clinker explain --code <CODE> prints the same registry-owned descriptor as
the list view, followed by a longer detail page when one exists:
clinker explain --code E15Y # retraction-mode aggregate incompatible with strategy: streaming
clinker explain --code E376 # retired type: output spelling; use type: sink
Not every registered code has a longer page. A registered code without one is
still valid and prints its complete descriptor plus Detail page: none; only a
code absent from the registry is unknown. The See: clinker explain --code <CODE> line is appended to a diagnostic only when the longer page exists.
List and code discovery are static authoring metadata. They do not compile a
pipeline, inspect records, or render runtime values or secrets. Use only the
clinker explain spelling: there is no separate diagnostic command.
Column Lineage
The --lineage flag builds the pipeline’s column-level lineage – which source columns each output column is derived from, and which source columns influence the output as a whole – and writes it as OpenLineage events. Like --explain, it compiles the plan and exits without reading any data, so the lineage is derived statically from the pipeline definition.
# Write to a file
clinker run pipeline.yaml --lineage lineage.ndjson
# Write to stdout (pipe into other tooling)
clinker run pipeline.yaml --lineage -
There are two emission modes:
--lineage– a static, plan-derived export. It compiles the plan and exits without reading data, so it runs instantly and describes the pipeline’s lineage rather than a specific execution.--lineage-events– live run-lifecycle emission. It runs the pipeline and emits aSTARTwhen the run begins and a terminalCOMPLETE/FAIL/ABORTwhen it ends, carrying real timing and row counts. See Live run events below.
Both modes share the same column-lineage facet and the same on-the-wire OpenLineage shape; the live mode wraps it in real run-lifecycle events. In external identity mode, complete events can also cross the independently bounded delivery worker described below. That worker owns only the selected file or stdout sink; it does not share the OTLP Collector worker or its memory arena.
Dataset identity preflight
Both flags require an explicit [observability.lineage] identity policy in the
workspace clinker.toml. The default identity_mode = "external" requires one
exact binding for every emitted Source and Sink node. A binding uses either a
canonical datasource or a complete catalog namespace/name pair:
[observability.lineage]
identity_mode = "external"
[[observability.lineage.dataset]]
node = "source_customers"
canonical_datasource = "s3://warehouse/customers"
[[observability.lineage.dataset]]
node = "output_customers"
catalog_namespace = "analytics"
catalog_name = "customers_clean"
A source or output declared inside a composition body needs its own binding, keyed by the call site it belongs to:
[[observability.lineage.dataset]]
node = "enrich_orders.reference_prices"
canonical_datasource = "s3://warehouse/prices"
Body node names live in their own scope and may legally repeat a top-level
name, so the key is <composition node>.<body source> rather than the bare
name — two call sites of one body can be pointed at different files, and each
gets its own identity.
. joins a call site to a body node, and a key never has to disambiguate that
join from a node’s own name: a . in a node name is refused at plan time with
E010, for every node kind and inside composition bodies too. node = "enrich.ref" therefore always addresses the source ref inside composition
node enrich.
A \ that belongs to a node’s own name is written \\ in the key, so the key
format stays unambiguous on its own rather than by relying on the naming rule.
The same escape covers . — node = "enrich\\.ref" (in TOML, \\ is a
literal backslash; the literal string 'enrich\.ref' says the same thing)
would address a node whose own name is enrich.ref — but no pipeline the
planner accepts can produce that key. Node names without \ — nearly all of
them — are unaffected.
A node whose key cannot be written as a binding — over 128 bytes once the call site is joined to it — is refused by name, naming the limit. The correction is to rename the pipeline nodes the key is built from; there is no binding that can carry an over-long key.
# is reserved in an authored dataset name (catalog_name, or the name
half of a canonical datasource) because it separates a multi-record source’s
record types from their base dataset. Without the restriction a name like
payments#detail would collide with record type detail of a source bound to
payments, and the two would merge in the catalogue — attributing one
dataset’s columns to the other. Namespaces are unaffected.
Clinker validates all required bindings before opening the lineage sink or,
for --lineage-events, discovering sources and creating output attempts.
Missing, duplicate, partial, ambiguous, or invalid bindings fail as
observability.configuration.invalid; rejected values and physical paths are
not copied into the diagnostic. The complete observability policy, including
the required OTLP and authentication tables, is documented under
lineage identity.
The destination file is emptied once the exporter has started, before any
event is written. Point each run at a path you are willing to overwrite, and
copy a record you want to keep before re-running against it. This applies to
both --lineage and --lineage-events.
What that means for a run that produces no events:
- Refused before the exporter starts — an invalid pipeline, a rejected configuration, a lineage binding that does not resolve — the file is untouched and still holds the previous run’s events.
- Refused after the exporter starts, or fails before its first event, the file is empty. An empty file means this run wrote nothing; it never means the previous run’s result still stands.
- A plan-only
--lineageexport that wrote nothing removes the destination, so no zero-byte artifact is left for a later step to publish. This applies only to a regular file: a destination that always reports zero length, such as/dev/nullor a FIFO, is a successful export and is never removed. It is also skipped when the export ran out of flush time, because the exporter may still be writing.
If a consumer must distinguish “this run produced no lineage” from “an older run’s lineage is still here”, give each run its own destination path rather than relying on the state of a shared one.
Path-derived dataset names remain available only through the exact local compatibility spelling below. This mode is visibly labeled on stderr and is for local diagnostics, not external delivery:
[observability.lineage]
identity_mode = "local_diagnostic_paths"
Output format
The output is NDJSON (one JSON object per line) conforming to the OpenLineage 2-0-2 core spec. A run is described by a START event followed by a COMPLETE event that share one runId:
{"eventType":"START","run":{"runId":"019f030d-0b3e-7ee1-86ec-1bb5b4a2776b","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"namespace":"clinker","name":"audit_join","facets":{"clinker_pipeline":{"sourceHash":"7fd096a9..."},"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, ...}
{"eventType":"COMPLETE","run":{"runId":"019f030d-0b3e-7ee1-86ec-1bb5b4a2776b","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"namespace":"clinker","name":"audit_join","facets":{"clinker_pipeline":{"sourceHash":"7fd096a9..."},"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "inputs":[...], "outputs":[{"namespace":"analytics","name":"audit_report","facets":{"columnLineage":{ ... }}}]}
runIdis a UUID v7 minted for this export and shared by both events. Because--lineageis a static, plan-derived export, theSTART/COMPLETEpair describes the pipeline’s lineage, not an executed data run — no rows are processed and the two events share one timestamp. A separateclinker runmints its ownrunId. (For real timing and row counts tied to an actual execution, use--lineage-events.)- Correlation is copied from one immutable CLI lifecycle snapshot:
runIdis the generated execution ID, the clinker-definedclinker_batchrun facet carries the caller/generatedbatchId, and theclinker_semanticPlanjob facet carries the effective fingerprint algorithm, schema version, and digest. Static and live events do not independently mint or parse these identities. job.namespaceisclinker;job.nameis the pipeline name. The pipeline’s content hash rides in theclinker_pipelinejob facet (sourceHash), not the job name – so the name stays stable across edits while runs of the same definition remain correlatable.inputsare the source datasets;outputsare the sink datasets. External mode uses the exact configured canonical or catalog identities, so relocating a pipeline does not change its lineage graph. Explicitlocal_diagnostic_pathscompatibility mode instead uses thefilenamespace with resolved paths (and falls back to theclinkernamespace plus the node name for a network source).- The dataset namespace/name identifies the stable collection. A concrete
logical partition or location would be emitted as the standard role-specific
input/output subset facet — which rides under
inputFacetson an input andoutputFacetson an output, the positions its schema names, not under the dataset-levelfacets— and an explicitly authorized alias as the standard symlinks facet, which is a plain dataset facet and does ride underfacets; neither is ever inferred from worker paths, attempt paths, hashes, or process context. No pipeline emits either facet today – the workspace config exposes no subset or symlink fields, so nothing can authorize one. - The
columnLineagefacet is attached to each output dataset on theCOMPLETEevent.
Reading the columnLineage facet
The facet has two parts, mirroring the OpenLineage ColumnLineageDatasetFacet:
"columnLineage": {
"fields": {
"amount": { "inputFields": [
{ "namespace":"file", "name":".../audit_orders.csv", "field":"amount",
"transformations":[{"type":"DIRECT","subtype":"IDENTITY"}] }
]}
},
"dataset": [
{ "namespace":"file", "name":".../audit_orders.csv", "field":"order_id",
"transformations":[{"type":"INDIRECT","subtype":"JOIN"}] }
]
}
fields– DIRECT (value-derivation) lineage, keyed per output column: the source columns each output column’s value is computed from. A rename (emit full = name), a multi-hop chain, or a path through a composition body (including nested compositions) collapses to the originating source column. A column whose value derives from an envelope read ($doc.<section>.<field>, bare / indexed / inside a larger expression) gets a DIRECT input field on the originating source dataset whosefieldis the rendered$doc.…path – so envelope-derived columns trace back to the document section they came from.dataset– INDIRECT (influence) lineage for the dataset as a whole: source columns that shaped which rows exist, via filtering, joining, grouping, or sorting – collected once rather than duplicated across every column.
Each transformation carries a type (DIRECT / INDIRECT) and a subtype (IDENTITY, TRANSFORMATION, AGGREGATION, JOIN, GROUP_BY, FILTER, SORT, CONDITIONAL).
Multi-record sources
A multi-record flat file carries several record shapes in one physical file, discriminated by a lead record_type column. Record types differ in their columns, not in which rows they select, so each is treated as its own logical dataset rather than as a subset of one flat superset dataset:
- Each record type is a dataset named
<dataset>#<id>– the source’s bound dataset identity with the record type’sidas a#fragment. Underidentity_mode = "external"that is the configured canonical or catalog identity (namespaces3://payments-lake, nameraw/payments#detail), so no filesystem path enters the name; underlocal_diagnostic_pathsit is the resolved file path (.../payments.txt#detail). Its columns are exactly that record type’s declared columns, so an output column that derives from a detail-record field traces to…#detail, and one from a header field traces to…#header. - A column declared by several record types (unified into one superset column) lists each owning
#<id>dataset as an input field, so a derived output column traces to every record type it could have come from. - The engine-stamped
record_typediscriminator lead column belongs to the container rather than to any one record type, so it stays on the base dataset (no fragment) – aRoutethat branches onrecord_typestill references{<base>, record_type}. - The run’s
inputslist the base dataset followed by each#<id>record-type dataset, in record-type declaration order. Declaring them is load-bearing, not cosmetic: a lineage consumer resolves acolumnLineageinput field only against datasets the run declared as inputs, so a record-type dataset left out ofinputswould have its column edges silently dropped on ingest.
A record type’s parent / join_key – the intra-file hierarchy linking a child record type to its parent – is not emitted as a lineage edge, since no plan operation performs that join.
Live run events
--lineage-events <PATH> runs the pipeline and emits OpenLineage run events tied to that actual execution, as NDJSON to a file path (or - for stdout):
clinker run pipeline.yaml --lineage-events events.ndjson
Unlike --lineage (which exits before reading data), this processes data, so it cannot be combined with --lineage, --explain, --dry-run, or -n.
Prefer a file path for a clean stream. With
-(stdout), the run’s own stdout output — for example the per-stage spill-volume summary — interleaves with the event lines, so stdout is not pure NDJSON. Writing to a file keeps the events unmixed.
A run emits a START when it begins, then exactly one terminal event when it ends:
START– offered to the lineage path before the run body executes. In local diagnostic mode it is written synchronously; in external mode it is admitted non-blockingly to the bounded lineage queue and can be dropped under the configured policy. It carries the input and output datasets by identity plus the shared batch and semantic-plan correlation facets; no completed dataset facets exist yet.COMPLETE– the run finished. It carries the input datasets and the output datasets with theircolumnLineagefacets, exactly like the static export.FAIL– the run errored. It carries the standard OpenLineageerrorMessagerun facet and the clinker-definedclinker_failurefacet. Both are derived from the same bounded, sanitized classification used by machine supervision; the latter adds stablecode,category, andretryAdvicefields.ABORT– the run was interrupted (e.g. aSIGINT/SIGTERMshutdown) and drained what it could before unwinding.
{"eventType":"START","eventTime":"2026-07-03T17:00:00Z","run":{"runId":"019f...","facets":{"clinker_batch":{"batchId":"batch-42"}}},"job":{"facets":{"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "inputs":[...], "outputs":[...]}
{"eventType":"COMPLETE","eventTime":"2026-07-03T17:00:04Z","run":{"runId":"019f...","facets":{"clinker_batch":{"batchId":"batch-42"},"clinker_runStats":{"recordsRead":1000,"recordsWritten":970,"recordsDlq":30,"durationMs":4210}}},"job":{"facets":{"clinker_semanticPlan":{"algorithm":"blake3","semanticSchemaVersion":1,"digest":"4e8c..."}}}, "outputs":[{"...":"...","facets":{"columnLineage":{ ... }}}]}
Key differences from the static export:
runIdis the run’sexecution_id(a UUID v7) — the same identity used across clinker’s provenance sidecars and metrics spool, so an orchestrator can correlate the lineage events with the run’s other artifacts.- Both events carry the same
clinker_batchrun facet andclinker_semanticPlanjob facet. Together withrunId, their batch ID, execution ID, and semantic fingerprint tuple come from the run’s single CLI-owned lifecycle source and exactly match the optional machine stream when both are enabled. - The
STARTand terminal events carry distincteventTimes (run begin and run end), not one shared timestamp. - The terminal event carries a
clinker_runStatsrun facet — a clinker-defined facet withrecordsRead,recordsWritten,recordsDlq, anddurationMs. Counts are pipeline-wide run totals, not per-output. - On
FAIL, the run also carries the standarderrorMessagerun facet (ErrorMessageRunFacet1-0-0) plusclinker_failure, with the shared sanitized message and stable failure code/category/retry advice.
Every started run that reaches a handled executor or publication boundary records one terminal snapshot; output-commit errors close as FAIL from that same source. Terminal emission is best-effort after the run’s authoritative publication decision. A lineage admission drop or sink write, flush, or deadline failure is reported on standard error and does not fail a run whose outputs already landed. A process crash can still leave only a delivered START event.
External delivery boundary
External identity mode is also the boundary for the independently bounded
lineage worker. Each complete event is serialized under
lineage.max_event_bytes and offered without blocking to a queue capped by
lineage.queue_bytes; a full queue or oversized event drops the newest event.
One synchronous worker owns the selected file or stdout sink, and shutdown
waits no longer than lineage.flush_timeout_ms. Its typed outcome distinguishes
normal shutdown, write failure (including permission errors), flush failure,
and deadline expiry, with accepted, dropped, and full counters reported
separately.
Within that deadline the worker stops taking new events off the queue
halfway through, keeping the rest of the budget to finish the event it is
already writing. A destination too slow to keep up therefore receives a file
that is short — missing its last events — rather than one that ends inside
a half-written record, which an NDJSON reader cannot parse. Nothing is added to
the deadline: the whole flush still ends within lineage.flush_timeout_ms. A
destination that stops accepting bytes altogether cannot be waited on, so in
that one case the file may end mid-record, and the delivery outcome reports
that separately from its counters. It has no access to the telemetry arena or Collector worker.
What the run prints
A delivery that lost events, ended on anything other than a normal shutdown, or left its destination inside a record prints one line on standard error:
clinker: lineage delivery outcome: status=deadline-exceeded error_kind=none accepted=2 dropped=0 full=0 records_complete=false
records_complete is the completeness of the file, where the counters
beside it are the completeness of the export. It is true on every normal
shutdown and on the slow-destination path above — a short file is still valid
NDJSON — and false only when the worker was abandoned inside a write. That is
the one state a consumer cannot determine for itself: a truncated NDJSON file
simply ends, with nothing in it to say more was coming, and the counters cannot
tell you either, because a run that gave up on a slow destination reports the
same accepted total whether or not the last record made it out whole.
Being a condition rather than a count, it breaks the clean-run silence on its own: a run that dropped no events still prints this line if it left the destination unreadable. A normal shutdown that dropped nothing and ended on a record boundary prints nothing at all.
The plan-only --lineage export, whose whole invocation is the export, says
the same thing in prose when it misses its flush deadline, and its correction
follows from it. A short export ends “on a record boundary, short by the events
that never got out” and is simply re-run. An export that “ends inside a record
and is not readable as NDJSON” must be discarded rather than published as this
run’s lineage, and only then re-run.
A sink write or flush failure leaves the same two files, and says so the same way. Where the deadline path ran out of time on an export that was otherwise going fine, here the destination itself refused, so two separate facts are reported: what the destination was left holding, and where the retry should point. An export that “ends on a record boundary and is readable as NDJSON” is reported without a disposal instruction; one that “ends inside a record and is not readable as NDJSON” must be discarded first, and the correction says so ahead of the retry advice. The retry advice itself is unchanged — a permanent refusal (permission denied, read-only filesystem, a directory) asks for a different destination, anything else asks for a re-run — because re-running against a destination that has just refused a write may refuse it again.
Nothing is said about a destination that is not there: a failure that wrote no bytes at all leaves an empty file, which this path removes so a publish step cannot upload it as the run’s lineage, and the diagnostic then describes no file.
This external worker is not a second identity mode and does not make local
paths suitable as catalog identity. The explicit local_diagnostic_paths
mode remains a synchronous compatibility path for local file or console
inspection and cannot enter the external delivery worker.
Lineage uses logical dataset bindings, not working-directory or attempt paths.
It copies the batch ID, execution ID, semantic fingerprint, and terminal facts
from the same immutable lifecycle snapshot used by machine supervision and
OTLP, while keeping its delivery result independent. A lineage fault cannot
change final or DLQ bytes, exit status, machine terminal payload, publication
inventory, visible final set, or retained failed-attempt evidence. Event values
also remain outside this identity-only payload; telemetry field policy is
enforced before Collector queue admission. If [observability] is absent,
neither the lineage worker nor the Collector worker exists.
When to use
- Impact analysis – before changing a source schema, see which outputs and columns depend on it.
- Auditing & governance – feed the OpenLineage events into a catalog (e.g. Marquez) to track data provenance.
- Review – attach the lineage of a new pipeline to a PR to confirm the intended derivations.
Because --lineage reads no data, it runs instantly and works on a pipeline whose inputs do not yet exist.
Limitations
Lineage is derived from the compiled plan, so a few constructs are approximated:
- A column-grain
$docread is traced as DIRECT lineage (seefieldsabove) in a transform projection, a combine body, a composition body, and an aggregate emit, attributed only to a source whose envelope declares the section. A$docread in an influence predicate – a route condition, a culldrop_group_when, or a combinewhere– is surfaced as INDIRECT influence (FILTERfor route and cull,JOINfor combine). Two$doccases remain uncovered: a whole-section envelope echo (an output header/footer regenerated from a source document section, with no output column or expression); and any$docreference in a Reshape rule, which the compiler rejects outright (Reshape re-runs its rules after a per-group spill that drops envelope context), so there is no Reshape envelope lineage to produce. - A match: collect combine declared without a projection body produces coarse column lineage: each collected column derives (as
TRANSFORMATION) from every build-side column, because there is no body expression to pin the exact source column. - INDIRECT influence covers route/cull predicates, join keys, aggregate grouping, and correlation sort over record columns (and
$docenvelope terms in route/cull/combine predicates, as above). An aggregate’s pre-aggregation rowfilter, a transform-inlinefilter, and Reshapeorder_by/partition_byare not (yet) attributed as influence. - Constant and
count(*)columns (which have no source input) are omitted fromfields; engine-stamped columns ($ck.*,$meta.*,$source.*) are skipped, mirroring the default writer.
Memory Tuning
Clinker is designed to be a good neighbor on shared servers. Rather than consuming all available memory, it works within a configurable budget and reaches for back-pressure or disk spill before it runs out.
What the budget measures
Memory attribution and process RSS measure different things. Clinker samples RSS to detect pressure, while individual operators report the data they retain. Allocator overhead, thread stacks, native I/O workspace and startup allocations also contribute to RSS. Setting a budget therefore does not promise that the process’s resident size will always equal its accounted data size.
CSV input decoding and output preparation use the run’s finite resource budget. Growing decoded cells, owned record values and retained document metadata carry their accounting with them until the last owner releases the allocation. Output policy, schema mappings, captured headers and prepared operation storage are also admitted before allocation. Replacing a buffer accounts for both old and new storage while they overlap; sharing a value does not release its charge.
An operation that cannot obtain its working memory fails as a resource error. CSV does not silently truncate a large cell, impose a separate authored cell-size limit, or treat resource refusal as a record eligible for the DLQ. An explicitly configured spill location can hold prepared output bytes, but cell rendering and metadata still require memory. See output preparation.
This is not a whole-process allocation or constant-memory guarantee. The CSV parser’s raw buffers and the intermediate JSON tree used to parse JSON-encoded cells remain outside this admission boundary. Unchanged format readers and later legacy record copies also have separate or incomplete accounting. Source and Combine paths can still retain whole inputs; input-sized residency remains tracked in #1183. A successful small fixture does not establish that a larger input fits. The existing YAML and CLI tuning controls below remain the controls for this budget.
The memory: block
All pipeline-level memory tuning lives under a single optional block:
pipeline:
name: my_pipeline
memory:
limit: "1G" # optional — defaults to 512M
backpressure: pause # optional — defaults to pause
The entire block is optional. A pipeline with no opinions about memory writes nothing:
pipeline:
name: my_pipeline
…and gets the runtime defaults (512 MB hard limit, backpressure: pause).
Individual fields are also optional. Setting just one is fine:
pipeline:
name: my_pipeline
memory:
limit: "2G"
Setting the memory limit
CLI flag (highest priority):
clinker run pipeline.yaml --memory-limit 512M
YAML config:
pipeline:
memory:
limit: "512M"
When --memory-limit is passed it overrides pipeline.memory.limit for that run; omit the flag and the YAML value applies unchanged, falling back to the 512 MB default only when neither is set. An empty or whitespace-only flag value — as an ops wrapper produces when it forwards an unset variable, so --memory-limit "$CLINKER_MEM" expands to --memory-limit "" — is treated exactly like omitting the flag: the YAML value (or default) applies, rather than the run aborting. Suffixes are binary (1024-based): K = 1024 bytes, M = 1024², G = 1024³; a bare integer is bytes. (This differs from the decimal KB/MB/GB used by min_size/max_size, which are 1000-based.)
Default: 512 MB.
Invalid values: the two entry points treat a malformed limit differently.
An empty or unparseable memory.limit in the YAML (for example a stray
non-numeric value) falls back to the 512 MB default. A non-empty malformed
--memory-limit flag, by contrast, is rejected up front with a config error
that names --memory-limit and echoes the value you passed — so a typo such as
the decimal 4GB (the binary suffix is 4G) fails loudly instead of silently
collapsing to the default and shrinking a larger budget set in your YAML. (An
empty or whitespace-only flag value is not malformed: it is treated as if the
flag were omitted, as noted above.)
Either way, a value whose size is well-formed but too large to represent — its
scaled byte count exceeds the maximum a 64-bit counter can hold — is rejected
rather than wrapping to a small budget (the YAML overflow error names
memory.limit; the flag overflow error names --memory-limit). Pick a limit
that fits your host’s real memory.
A well-formed but undersized value is a different case: it is not a malformed
flag, so it clears the boundary check, and whether it aborts the run depends on
the backpressure policy. Under a producer-pausing policy (pause, the
default, or both) a value below the process’s baseline resident memory is
rejected at startup by the budget gate as E312. Under the non-pausing spill
policy that startup gate does not fire: the run proceeds and relies on spilling
to stay within the budget rather than aborting. Because --memory-limit simply
populates pipeline.memory.limit, that E312 — which names the limit and
echoes the offending byte value — refers to the same limit you passed via the
flag.
Choosing a backpressure policy
When memory use approaches the limit (the soft threshold is 80 % of limit), something has to give up memory. The backpressure knob chooses what:
| Value | Behavior |
|---|---|
pause (default) | Where possible, pause an upstream reader so it stops producing until pressure eases; when a paused reader is about to be needed, first spill downstream state and then proceed, so a pause never stalls the run. |
spill | Never pause a producer — always free memory by spilling a stage to disk. |
both | Pause where possible, otherwise spill whichever stage is holding the most memory. |
pause is the right default for most pipelines: pausing a fast Source feeding a slow downstream stage is cheaper than writing its buffered records to disk. Reach for spill or both only when you have a specific reason to prefer a different posture — for example, both when one large stage dominates the budget and you want it spilled first.
How pause and resume work
Under pause (and both), a producer paused because memory crossed the soft threshold is resumed automatically once memory recedes — it is never left parked. Pause and resume use two watermarks to avoid flapping (a hysteresis band):
- Pause when live memory rises above the soft threshold,
0.80 × limit. - Resume when live memory falls back below the lower resume watermark,
resume_threshold × limit(default0.70 × limit).
Between the two watermarks nothing changes, so a normal batch-to-batch swing in memory cannot make a producer flap between paused and resumed on every poll.
A paused reader also never blocks the run. When the engine reaches a stage that needs a paused reader’s records, it first sheds reclaimable downstream state to disk and then resumes the reader and proceeds — so pause throttles producers under pressure but degrades to spill-and-continue at the point it would otherwise wait, rather than stalling.
resume_threshold
resume_threshold tunes the low watermark of that band, as a fraction of the hard limit:
pipeline:
memory:
limit: "1G"
resume_threshold: 0.65 # optional — defaults to 0.70
- Default:
0.70(omit the field to take it). - Valid range: strictly greater than
0and strictly less than the0.80soft threshold, so the resume point always sits below the pause point. A value outside(0, 0.80)— including0.0, a negative, or anything≥ 0.80— is rejected at plan time withE324. A misspelled key is rejected as an unknown field. - Lower widens the band: a paused producer stays paused longer (smoother, but slower to re-open).
- Higher narrows the band: faster to re-open, but more prone to flapping.
Only the pausing policies (pause, both) use this watermark; under spill no producer is ever paused, so it has no effect.
Streaming batch size (batch_size)
pipeline.batch_size sets how many events (records plus document-boundary punctuations) a streaming-eligible stage hands off to its downstream consumer at a time over a back-pressured channel. For a fused stage (Source → Transform → Sink, Merge.interleave of Sources) it bounds the in-flight working set to one batch rather than the whole stage, because the stage pulls records off a live upstream channel without ever building a full result. The other streaming stages build their full result first and stream it in batches; there the knob sizes only the inter-stage slice, not the producer’s footprint. The knob is optional; omit it to use the built-in default of 2048 events. See Streaming vs. Blocking Stages for the distinction.
pipeline:
name: orders_rollup
batch_size: 1024 # optional; default 2048
A per-transform override is available on a Transform’s config.batch_size (see Transform Nodes); it takes precedence over the pipeline value for that one stage. A batch_size of 0 is rejected at config load. The knob affects only the memory profile of streaming stages, never their output — blocking stages (sort, hash Aggregate, Combine build side) ignore it and continue to fully materialize. See Streaming vs. Blocking Stages for the full model.
Behavior under memory pressure
You don’t manage memory by hand — the engine does it within the budget you set. What this means in practice:
-
Spillable stages always complete if disk space is available, regardless of input size. When a blocking stage (sort, hash Aggregate, a grace hash or range Combine) outgrows the budget, it spills to disk instead of failing.
- An in-memory equal-ids Combine does not spill. The planner picks the in-memory hash join or the disk-spilling grace hash join from its size estimate before the run starts. If an in-memory join’s build side outgrows the budget, the run stops with
E310 MemoryBudgetExceeded(#1337). Setstrategy: grace_hashon a Combine whose build side may not fit; see Combine Nodes. - Range and equality+range Combine also spill. A Combine whose
where:joins the two inputs on an inequality (<,<=,>,>=) — a pure band join such asorders.amount >= bands.lo and orders.amount < bands.hi, or an equality-plus-range join such asorders.region == bands.region and orders.amount >= bands.lo— runs the block-band strategy, which is spill-bounded on both axes: it external-sorts each input side to disk and accumulates its matched output in a spillable sort as well, so it completes on inputs and result sizes larger than the budget. When an equality key is present, records are grouped by that key (via its hash) before the range walk and only same-key pairs are joined; a single very common key is spread across many disk-backed blocks and joined a bounded pair at a time, so even a heavily skewed key stays within the budget. The join still aborts withE310 MemoryBudgetExceededonly as a last resort — when a single indivisible unit of work (one pair of input blocks plus the join’s scratch arrays) cannot fit the hard limit even on its own. If you hit that, give the pipeline more headroom or narrow the predicate.
- An in-memory equal-ids Combine does not spill. The planner picks the in-memory hash join or the disk-spilling grace hash join from its size estimate before the run starts. If an in-memory join’s build side outgrows the budget, the run stops with
-
Performance degrades gracefully. Under pressure you’ll see slower execution and possibly disk I/O — not a crash.
-
The limit is a soft ceiling, not a hard wall. Momentary spikes may briefly exceed it before the engine reacts. Only if memory blows past the limit outright does the run abort with
E310 MemoryBudgetExceeded, which names the stage that overran. -
Shared handoff buffers spill and re-scan sequentially. A stage that fans out to several consumers, or feeds a composition input port, records the exact reader count for each output port and may spill that shared slot like any other materialized buffer. Readers run one at a time over immutable backing; spilled and mixed buffers open one file at a time for each scan, and the final reader takes the authoritative slot. Clinker never pre-forks one copy per destination. A consumer that must collect the scan into a full resident vector reserves that materialization before allocating it. If the projected overlap exceeds the hard limit, it aborts with
E310 MemoryBudgetExceeded, names that consumer, and reportsNodeBufferas the budget category. More spill space can keep the shared backing off-heap, but it cannot eliminate an individual operator’s required resident working set. -
One oversized correlation group trips
E310in Reshape and Cull. Both apply their rules to a whole group at once, so a group must fit the budget at finalize even though cross-group and ingest-time peaks spill. A group larger thanmemory.limithas no in-budget representation, so the run aborts withE310 MemoryBudgetExceedednaming the node and the offendingpartition_bygroup. Raisingmemory.limitclear of the reported figure is the only fix that leaves your output unchanged (finalize also holds the run’s other groups, so that figure is a floor, not a target). Dropping unread columns in an upstream Transform also shrinks the group, but both nodes write every input column through, so those columns leave the output too. Do not narrowpartition_byto clear it — that key defines the group the rules evaluate over, so a narrower key makes the run succeed by changing your results. Both nodes also hold per-group bookkeeping that isO(distinct groups)and cannot spill — the spill path evicts buffered records, never group entries — but they handle an extreme partition cardinality differently, and the difference is the diagnostic you get:- Cull fails loud. Its in-memory drop-decision state is combined with the run’s other live charged memory and checked against the budget on every admission, so a cardinality that would breach
memory.limitaborts withE310 MemoryBudgetExceedednaming the decision state instead of a group. - Reshape has no such gate. Its group map and group-order list grow with the distinct-group count unchecked, so a high-cardinality Reshape can be OOM-killed rather than reporting
E310; the resulting status is platform-dependent. Keep the group count well within the budget on a Reshape; there is no engine check to catch it for you. This gap is tracked in #1027.
- Cull fails loud. Its in-memory drop-decision state is combined with the run’s other live charged memory and checked against the budget on every admission, so a cardinality that would breach
-
clinker explain --code E310covers every surface. The rendered diagnostic names which one overran in its[...]detail; the explain page keys its remediation to that. -
A Reshape partition value it cannot key is folded into the null group, silently. Where Cull aborts on an array or map
partition_byvalue, Reshape groups that record undernull— alongside every record whose partition column is missing, explicitly null, or an empty string — and the run exits 0. ANaNis not among them: in both nodes everyNaNis one group of its own. Nothing in the output marks the merge, so if any of those shapes can occur in your partition column, normalize or filter it in an upstream Transform. See Values Reshape cannot key. -
These are limits, not bugs. Every case above is an ordinary operational condition with a documented code, so it reports through
E310rather than as an internal error. Four conditions do still present as aninternal errortoday despite being about your data or your host, not an engine defect:- an array or map value in a Cull
partition_bycolumn (Reshape folds these into the null group instead, per the bullet above); - a runtime data error inside a Cull
drop_group_whenexpression; - a runtime data error inside a Reshape rule’s
when,mutate.set, orsynthesizeexpression — the direct analogue of the Cull case above. Like it, this aborts the run rather than routing the record to the dead-letter queue, even understrategy: continue; - a spill read/write failure in a Reshape or Cull group buffer, including the spill volume filling up; #1021 tracks preserving its spill diagnostic and exit classification.
Any other
internal erroris a Clinker defect worth reporting. - an array or map value in a Cull
Some stages stream (they hold only a small in-flight slice of records) and some materialize (they hold a whole stage’s worth before emitting). clinker run --explain annotates each node with buffer: streaming or buffer: materialized so you can see which stages will dominate the budget before you run. See Streaming vs. Blocking Stages for which is which.
Dead-letter output
Dead-lettered rows are not collected in memory until the run ends. Each row is formatted under its DLQ file’s header, which is fixed when the pipeline compiles, and written straight into a staged copy of that file. Each open DLQ file costs one fixed 64 KiB write buffer, and the number of DLQ files is fixed by the pipeline. Apart from the held failures described below, a run whose every row fails therefore uses about as much memory for dead letters as a run where none do. That is why the DLQ writers are not charged to the memory budget: they do not grow with input.
What does grow with failures is the DLQ files themselves, and they are
bounded by disk: free space at the staging location, and the publication
attempt’s byte ceiling (storage.publication.max_attempt_bytes). To stop a
run before it produces a large DLQ, set a breaker: dlq.max_rate and
dlq.per_source.<name>.max_rate (E315/E316), or type_error_threshold
(E368). See
How DLQ output is written,
Bounding how much can dead-letter
and How the DLQ columns are chosen.
Some failures are held in memory, uncharged, until the stage that found them
finishes, and are written then: join_values collisions at a Sink that writes
on its own thread, Aggregate add_record failures found on the Aggregate’s
input thread, and Combine output-row failures found on the Combine’s driver
thread or inside a grace-hash, sort-merge or IEJoin join. Records a correlation key
holds until their group is decided are that feature’s own state, not DLQ
output.
Under dlq_granularity: document three things are held, all charged to the
memory budget:
- Each Sink holds every open document’s records in a buffer that spills to disk when the budget needs the memory, until the document’s verdict is final.
- A failed document’s failing records are held as their dead-letter rows
until the document is rejected. They move to one file in the spill
directory when the budget needs the memory, and count toward
storage.spill.disk_cap_bytes(E320). If one more held row would not fit even with every held row on disk, the run fails with E310. - For each rejected document, a compressed record of the rows already written, so a row several Sinks held is written once. It never spills; if it would pass the limit once every held row is on disk, the run fails with E310.
See Document-level DLQ and, for what each of these costs, How DLQ output is written.
Sizing guidelines
| Workload | Recommended limit | Notes |
|---|---|---|
| Small files (<10 MB) | 128M | Minimal memory pressure |
| Medium files (10–50 MB) | 256M | Covers most ETL jobs |
| Large files or complex aggregations | 512M (default) – 1G | Multiple group-by keys, large cardinality |
| Multiple large group-by keys | 1G+ | High-cardinality distinct values |
Target workload: Clinker is optimized for 1–5 input files of up to 100 MB each, processing 10K–2M records per run.
Aggregation strategy interaction
Memory consumption depends heavily on the aggregation strategy the optimizer selects:
-
Hash aggregation accumulates state in a hash map. Memory usage is proportional to the number of distinct group-by values. With high-cardinality keys, this can consume significant memory before spill triggers.
-
Streaming aggregation processes groups in order and emits results as each group completes. Memory usage is minimal (proportional to a single group’s state) but requires the input to be sorted by the group-by keys.
-
strategy: auto(the default) lets the optimizer choose based on the declared sort order of the input. If the data arrives sorted by the group-by keys, streaming aggregation is selected automatically.
To influence strategy selection:
- type: aggregate
name: rollup
input: sorted_data
config:
group_by: [department]
strategy: streaming # force streaming (input MUST be sorted)
cxl: |
emit total = sum(amount)
Only force streaming when you are certain the input is sorted by the group-by keys. If the data is not sorted, results will be incorrect. Use auto when in doubt.
Oversized single rows
An aggregate that keeps min, max, or another value-buffering binding holds each contributing row’s raw values until the group finalizes. If a single input row’s buffered footprint is larger than the entire memory.limit, no amount of spilling can hold it — spill would only re-read the same oversized row. The engine surfaces this per-row overflow rather than absorbing it:
- With
error_handling.strategy: fail_fast(the default), the run aborts withE310 MemoryBudgetExceeded, naming the aggregate stage and reporting the offending row’s byte footprint against the budget. - With
strategy: continue, the offending record is routed to the dead-letter queue under theaggregate_finalizecategory, and the run proceeds.
This is almost always a sign the budget is set far too low for the record shape — raise memory.limit so a typical row fits comfortably.
Compositions
A composition (a reusable sub-pipeline included via use:) does not get its own memory budget — its operators share the parent pipeline’s budget and spill to the same temporary directory. A spilled composition input is scanned sequentially, and its materialization stays charged continuously as ownership moves into the body; it is not briefly dropped from the accounting or charged twice. During body-input schema conversion, the old input and new records are both charged for the short interval when both allocations exist. If that materialization would exceed the hard limit, E310 names the composition call-site directly. If a later budget overrun happens inside the composition, the error names that same call-site (e.g. enrich_call) so you can locate it, prefixing the message with in composition "enrich_call": ... when the overrun is internal to the body.
Monitoring memory usage
Use the metrics system to track peak_rss_bytes across runs:
clinker run pipeline.yaml --metrics-spool-dir ./metrics/
The metrics file includes peak_rss_bytes, which shows the maximum resident memory during execution. If this consistently approaches your memory limit, consider increasing the budget or restructuring the pipeline to reduce intermediate state.
Shared server considerations
On servers running JVM applications, memory is often at a premium. Recommendations:
- Set
--memory-limitormemory.limitexplicitly rather than relying on the default. Know your budget. - Use
--threadsto limit CPU contention alongside memory limits. - Monitor
peak_rss_bytesin production metrics to right-size the limit over time. - Schedule large pipelines during off-peak hours when JVM heap pressure is lower.
Storage & Spill Location
Blocking operators — Aggregate, sort, and grace-hash Combine — accumulate
state in memory up to the configured budget, then spill to disk when a
soft or hard memory threshold trips, rather than running the process out of
memory. By default those spill files land in the operating system’s temporary
directory. The [storage] block in clinker.toml lets you redirect them.
Output preparation
CSV, JSON, XML, fixed-width and SWIFT output in the CLI and executor prepares each complete output operation before delivering its bytes. For CSV, the first body row and its automatic header share one operation; explicit document start and end are separate operations. The same finite-resource preparation API is available to library integrations. Fixed-width prepares document headers, body records and footers separately. SWIFT prepares its service headers together with the first body record, then prepares the closing block and trailer at finalization.
Prepared output bytes stay in memory unless storage.spill.dir supplies an explicit
spill location. This differs from the operator spill default described below:
output preparation does not silently use the operating system’s temporary
directory. Configured spill uses the run’s disk budget and a finite descriptor
allowance. It does not remove the memory required for a rendered CSV cell,
retained header, format configuration and schema plans, fixed-width warning
history, or a SWIFT document trailer. Fixed-width retains every committed
warning; resource refusal fails the next operation rather than discarding
history. See
Memory Tuning.
Failure before delivery writes none of that operation’s bytes. Once delivery starts, ordinary I/O can accept a prefix before failing. The writer then refuses further operations, including flush, so the original failure is not hidden by later calls. Prepared bytes and accepted destination bytes are different counts; a partly delivered operation does not become a committed record.
Temporary files remain charged until their removal is confirmed, including when cleanup must be retried. Dropping a handle or requesting cancellation does not turn failed cleanup into free disk capacity. Resource telemetry can be dropped when its fixed arena is full; admission, output and cleanup do not depend on those signals being retained.
These rules add no storage setting and make no atomic-publication promise for an arbitrary destination. Output publication governs the separate file-publication boundary. EDIFACT, X12 and HL7 retain their existing writer paths. These five codecs do not establish all-format migration or admission of existing reader/parser allocations.
The [storage] block
Storage settings are a property of the workspace, not of an individual
pipeline, so they live in clinker.toml at the workspace root rather than in
the per-pipeline YAML:
[storage.spill]
dir = "/var/clinker/spill" # optional; operator default = OS temp dir
disk_cap_bytes = "10GB" # optional; default = unlimited
compress = "auto" # optional; auto | off | on (default = auto)
[storage.staging]
enabled = false # opt-in; default off
dir = "/var/clinker/staging" # required when enabled
patterns = ["/mnt/nfs/data/**"] # which sources to stage
[storage.publication]
mode = "direct" # direct | local_then_publish
destination_profile = "local" # local | nfs_v4_1 | smb_3_1_1
failed_retention_seconds = 86400 # 24 hours; zero is allowed
The whole block is optional. With no clinker.toml, or a clinker.toml that
omits [storage], blocking operators spill to the OS temp directory. Prepared
output preparation stays in memory unless storage.spill.dir is set.
Table names are checked
Clinker reads five top-level tables — [catalog], [storage],
[observability], [channel], and [group] — and passes over any other
top-level table, so a clinker.toml may carry tables meant for other tooling.
A table name that is a misspelling of one of those five is refused instead,
naming the table you wrote and the one it was mistaken for. Absence is how a
workspace says “off” — omitting [observability] disables telemetry — so a
name that misses by a letter would otherwise turn a policy off silently:
clinker.toml table [observabilty] is not a table clinker reads, and is a
misspelling of [observability]; nothing under it would have been applied —
write `[observability]`, or rename the table so it is not mistakable for one
Keys inside a recognized table are strict already: an unknown key there is refused outright.
Output publication and retained attempts
[storage.publication] is the only author-facing block for output publication,
destination qualification, retained attempts, and bounded cleanup. It is
optional; the defaults use destination-local quarantine (mode = "direct"),
the local filesystem profile, and 24-hour retention for failed attempts.
NFS and SMB profiles are accepted only when the operating system identifies
the matching filesystem family; an unidentifiable remote drive is rejected
rather than guessed.
[storage.publication]
mode = "direct"
destination_profile = "local"
failed_retention_seconds = 86400
creation_grace_seconds = 300
max_attempt_bytes = "4GB"
retained_byte_limit = "8GB"
retained_attempt_limit = 8
min_free_bytes = "2GB"
sweep_entry_limit = 1000
sweep_byte_limit = "8GB"
sweep_time_limit_ms = 2000
| Setting | Default | Hard limit | Meaning |
|---|---|---|---|
failed_retention_seconds | 86,400 (24 hours) | 604,800 (7 days) | How long incomplete, abandoned, or otherwise failed terminal attempts remain. Zero is valid and makes them immediately eligible after the live lock is released. |
creation_grace_seconds | 300 | 3,600 | Grace period before a non-terminal attempt can be considered abandoned. |
max_attempt_bytes | 4 GB | 16 GB | Maximum admitted estimate for one publication attempt. |
retained_byte_limit | 8 GB | 64 GB | Aggregate retained-attempt byte ceiling used by policy. |
retained_attempt_limit | 8 | 128 | Aggregate retained-attempt count ceiling. |
min_free_bytes | 2 GB | 64 GB | Additional free-space headroom required by admission. |
sweep_entry_limit | 1,000 | 10,000 | Maximum directory entries considered by one cleanup page. |
sweep_byte_limit | 8 GB | 64 GB | Maximum regular-file bytes considered by one cleanup page. Must be at least max_attempt_bytes + 4,194,304B so one maximum attempt and its bounded manifest can always make progress. |
sweep_time_limit_ms | 2,000 | 30,000 | Maximum monotonic elapsed time for one cleanup page. |
The capacity observation is advisory. It is a one-time comparison of the
attempt estimate plus min_free_bytes against observed free space; it reserves
no blocks or quota. A later write or synchronization can still fail with
ENOSPC or EDQUOT, and Clinker retains exact attempt state rather than
claiming publication succeeded.
Configuration resolution rejects a sweep budget that cannot inspect one
maximum-sized attempt plus the bounded 4 MiB manifest. The diagnostic reports
the exact minimum and a paste-ready sweep_byte_limit setting; this prevents a
valid admitted attempt from becoming permanently too large for cleanup.
Aggregate admission counts physical manifest-owned staging and quarantine
files across every destination and local-spool root. Temporary local and
destination copies both count while both exist; an uninspectable size fails
admission instead of being treated as zero. Count and byte inventory, expiry
cleanup, the limit check, and attempt-root creation are serialized across the
same root set, so concurrent runs cannot both consume the final retained slot.
Before releasing that serialization boundary, Clinker records the admitted
estimate in every owned attempt root. A later process charges at least one
copy of that reservation for the execution until exact artifact sizes replace
it. Admission locks live only inside the internal .clinker-attempts
namespace, so an authored output leaf cannot replace the mutex.
Lowering retained_attempt_limit does not hide attempts that were admitted by
an earlier configuration. Listing and expired purging continue to page over
the bounded physical namespace and report policy debt until the retained count
is back within the new limit.
Publication modes and destination profiles
direct writes each artifact into restrictive quarantine on its destination
filesystem, synchronizes it, and only then promotes it to the final leaf.
local_then_publish requires local_spool_dir; Clinker writes and verifies the
local copy, copies and verifies it into destination-local quarantine, and
promotes only that destination copy. There is no copy fallback that writes
directly to a visible final, and no automatic mode fallback. A destination
profile mismatch fails closed.
The supported destination profile is a qualification claim, not a spelling
that makes an arbitrary mount safe. NFSv4.1 and SMB3.1.1 support requires an
actual mounted destination qualified with the selected profile, including real
late-failure behavior. Hosted qualification also runs the ordinary publication
admission API from independent processes against each mounted profile, in
opposite multi-root order, and requires exactly one process to enter the final
retained-count and retained-byte slot without deadlock. This production lock
proof is separate from the filesystem’s byte-range/OFD lock probe. Release evidence for low space must observe a real
ENOSPC. Injected EDQUOT is useful seam coverage but is non-qualifying unless
a real quota is provisioned and the filesystem reports EDQUOT during the
mounted test.
Publication truth is per artifact. Earlier artifacts can be synchronized and visible while a later promotion fails, so an output set is never described as atomically published across artifacts. Quarantine ownership is scoped to one execution ID; attempts never share writable staging ownership.
Retention and metadata-last cleanup
Every retained attempt has a restrictive owner manifest and live lock. Cleanup
accepts only roots derived from a freshly compiled workspace-relative pipeline,
canonical execution IDs or the typed expired selector, and bound opaque
continuations. It never accepts a raw deletion path.
Cleanup keeps data on any ambiguity. It opens entries through retained, no-follow handles, verifies the supported manifest and allowed children, acquires the live lock, checks wall-clock eligibility, and stays within the configured entry, byte, and monotonic-time limits. Owned artifact bytes are removed first. Clinker then revalidates the directory, removes the owner manifest and liveness metadata last, and removes the now-empty attempt root. A crash or refusal before that sequence finishes remains explicit cleanup debt for a bounded retry.
Use the non-mutating operator surface to inspect or preview retained state:
clinker attempts list pipelines/orders.yaml
clinker attempts inspect pipelines/orders.yaml \
--execution-id 018f47a2-9a41-7a27-b4d6-4f7137e3c159
clinker attempts purge pipelines/orders.yaml --expired
clinker attempts purge pipelines/orders.yaml --expired --execute
Repeat any path or overlay identity used by the run (--base-dir, absolute-path
permission, rules root, channel/groups, --path-execution-id, batch identity,
or timestamp) so the command recompiles the same typed owned roots. File-backed
fan-out roots normally replay source discovery. A retained failure also keeps
a bounded, plan-bound receipt containing the logical source identities used by
{source_file} and {source_path} plus path-free identifiers for every owned
output or spool root. If a source is later removed or its directory is renamed,
attempt commands re-render those authored templates and require the resulting
validated roots to match the receipt exactly before listing, inspection, or
purge. Successful publication removes the receipt with the rest of the attempt.
Attempt commands do not accept a raw cleanup path. Continuation tokens are opaque raw values;
JSON output provides authoritative structured recovery/resume argument arrays
for shell-independent automation.
Output is path-free by default and reports logical root, execution, and artifact
IDs. --show-paths adds only sanitized workspace-relative paths; machine-local
prefixes, sensitive-looking components, credentials, secrets, record values,
and raw staging detail remain redacted. Successful operations exit 0. Invalid
selectors, pipelines, and continuations exit 1. Bounded partial work or any
cleanup debt exits 4 with E371/E372 retry and workspace-relative recovery
guidance. See the attempt command reference
and exit-code contract.
storage.spill.dir — where spill files go
When dir is set, the per-run spill directory (clinker-spill-<random>/) is
created under that path, and every blocking operator writes its spill files
there. When dir is omitted, the per-run directory is created under the OS
temp directory (std::env::temp_dir, typically $TMPDIR or /tmp).
The directory is validated once at startup, before any input is read. If the path does not exist, is a file, or is not writable, the run fails immediately with a diagnostic naming the setting:
storage.spill.dir /var/clinker/spill does not exist; create it or point at an existing volume
Validating up front — rather than at the first spill — means a misconfigured spill volume fails fast, while the run is cheap to abandon, instead of after minutes of work. (This is the trap DuckDB fell into when its temp-directory setting was honored only lazily, duckdb/duckdb#9401.)
Why redirect spill off /tmp
On many Linux hosts — especially systemd-managed ones — /tmp is mounted as
tmpfs, which is backed by RAM (and swap), not disk. Spilling there does
not actually free physical memory: the spill bytes stay resident, defeating
the whole point of the memory budget. If df -T /tmp reports a tmpfs
filesystem, point storage.spill.dir at a path on a real block device so
spilling moves pressure off RAM and onto disk.
Prefer local spill for network-share pipelines
When sources and outputs live on NFS or SMB, point storage.spill.dir at a
real local disk if possible. Spill workloads include repeated reads, writes,
merges, and synchronization; running them on the share usually multiplies
latency and network I/O without improving the durability of the final output.
Likewise, optional source staging can copy matched share inputs to local disk
before execution. Output commit remains destination-local: Clinker writes and
flushes a hidden file on the output share, then promotes it on that same
filesystem so the final rename does not cross devices.
Inspecting the resolved spill root
clinker run --explain prints the resolved spill root and where it came from,
so you can confirm the setting took effect before committing to a run:
Spill root: /var/clinker/spill [storage.spill.dir]
…or, with no configuration:
Spill root: /tmp [OS temp dir (default)]
The same --explain output reports the resolved disk cap on the next line:
Spill disk cap: 10737418240 bytes [storage.spill.disk_cap_bytes]
…or, with no cap configured:
Spill disk cap: unlimited (default)
Finally, --explain reports the resolved compression decision per
spill-writing operator, so you can see which spills will be LZ4-framed (lz4)
and which will be written raw (off) before the run starts. Under auto the
choice varies by operator width:
Spill compression: Auto [storage.spill.compress]
Aggregate 'totals' → lz4
Sort 'by_amount' → off
Only operators that actually write spill files appear here: the external sort, the hash Aggregate, the grace-hash / sort-merge Combine, and the pure-range (block-band) IEJoin Combine, which external-sorts each side and writes its min/max-tagged blocks to disk and spills its matched-output sort runs the same way. The remaining in-memory join strategies — the inline hash build/probe and the equi+range IEJoin (hash-partitioned range join) — run their kernel entirely in RAM and never open a spill file, so spill compression does not apply to them and they are omitted from this list, even though they carry a spill priority for memory arbitration.
storage.spill.disk_cap_bytes — cap concurrent spill
By default a run will spill as much as it needs, limited only by the physical
space on the spill volume. disk_cap_bytes sets a budget on the spill the run
holds at once: the on-disk size of the spill files live at any moment. When
that footprint would cross the cap, the run aborts with a dedicated diagnostic
instead of continuing to fill the volume. Because the cap tracks what is
concurrently on disk, an operator that deletes intermediate spill files as it
consumes them (such as the merge that folds a heavily fragmented external sort
back together) does not count those transient files twice — only the disk a run
actually occupies at once is charged against the cap.
[storage.spill]
dir = "/mnt/fast-ssd/clinker-spill"
disk_cap_bytes = "50GB"
The value accepts the same human-readable byte-size grammar as the source
size filters — a bare integer is bytes, and KB/MB/GB suffixes use
decimal units (1GB = 1,000,000,000 bytes), matching du, df, and the AWS
CLI. Omitting the key leaves spill unlimited, exactly as before.
The cap is a policy ceiling, deliberately independent of both the memory
budget and the physical volume size. A run can sit well inside its
memory.limit and still exhaust local disk through an unbounded stream of
spill files; the cap lets an operator bound that on a shared volume. It is the
guard DataFusion shipped without (apache/datafusion#15358) until production
runs filled volumes.
storage.spill.compress — LZ4 compression policy
Spill files are postcard-encoded record streams. By default each stream is wrapped in an LZ4 frame, which shrinks large spilled runs. But LZ4 carries a per-frame fixed cost — clearing the compressor’s internal state on every frame reset — and on small spills that cost can outweigh the byte savings. The LZ4 v1.8.2 release notes call this out directly, and Pentaho Kettle ships explicit guidance to turn spill compression off for small rows.
compress controls the policy:
[storage.spill]
compress = "auto" # auto | off | on (default = auto)
| Mode | Behavior |
|---|---|
auto (default) | Compress only when a spilled batch is projected large enough to amortize LZ4’s per-frame cost — both ≥ 4 KiB and ≥ 1024 rows. Below either threshold the batch is written raw. The projection comes from the operator’s schema width and the run’s batch_size, so the decision is made per blocking operator. |
off | Never compress. Postcard records are written straight to disk with no LZ4 frame. Cheapest for small spills; largest on-disk size. |
on | Always compress with an LZ4 frame. The pre-knob behavior, best for spills of large, compressible rows. |
Each spill file records its compression choice in a one-byte header tag, so the read path always dispatches to the right decoder regardless of the mode the file was written with — changing the knob between runs never breaks re-reading an earlier run’s files.
The 4 KiB / 1024-row thresholds mark the empirical crossover: below them the
LZ4 frame’s fixed cost dominates the small amount of compressible payload, and
writing raw is faster end-to-end (the spill_compression benchmark sweeps
batch sizes from 256 B to 64 KiB and confirms auto tracks the faster of
on / off across the range). Most pipelines should leave compress at
auto; set on when spilling wide, highly compressible rows to a
space-constrained volume, and off when spills are dominated by many small
batches.
Observability — what the planner will do before you run
clinker run --explain is plan-only (it reads no input and spills nothing), so
it is the safe place to see what a run would do to the spill volume and to the
staging dir before committing to it. On top of the resolved spill root, disk
cap, and compression decision documented above, --explain surfaces three
storage-observability sections, and a real clinker run reports the matching
actuals at end-of-run so you can calibrate the estimate.
A note on byte units. Three different unit conventions appear across the storage surface, and it helps to know which is which before comparing figures:
- Config values you write (
disk_cap_bytes = "10GB") use decimal units —1GB= 1,000,000,000 bytes — matchingdu,df, and the AWS CLI (see the disk-cap grammar). - The
=== Estimated Spill Volume ===section humanizes with binary suffixes —K/M/G= KiB/MiB/GiB — so it lines up with thepredicted_peakfigure on each stage’s Physical Properties line, which uses the same humanizer. - The cap-headroom line and the post-run actuals print raw bytes with no suffix, so the cap-minus-estimate subtraction and the estimate-vs-actual comparison are exact rather than rounded.
When you calibrate the estimate against the post-run actual, convert the binary
estimate suffix to bytes first (1K = 1024 bytes, 1M = 1,048,576 bytes) so you
are comparing the same unit the actuals report.
Estimated spill volume per stage
The === Estimated Spill Volume === section lists one line per spill-writing
stage (hash Aggregate, external sort, grace-hash / sort-merge Combine, and the
pure-range block-band IEJoin Combine) with its plan-time spill-volume estimate,
followed by a total. The remaining in-memory join strategies (inline hash
build/probe, equi+range IEJoin) never write spill files, so they do not appear
here and do not inflate the total:
=== Estimated Spill Volume ===
Estimated spill volume (per blocking stage):
[aggregation:hash] dept_totals → 1K
[sort] by_amount → 4K
Total: 5K
Each figure is the operator’s coarse predicted peak live state — the same
predicted_peak the Physical Properties arbitration line shows — and bytes
render in binary units (K/M/G = KiB/MiB/GiB). Summing rather than maxing is
the conservative choice for a preflight: two blocking operators can be live and
spilled at the same time, so their footprints add.
A streaming-only pipeline (no blocking operator) has nothing that spills, so the section is omitted entirely.
Unknown stages. The estimate is seeded from input file sizes resolved at
plan time. A stage whose volume cannot be known before the run renders
unknown instead of a misleading 0B, and the total notes that unknown stages
are excluded:
[aggregation:hash] dept_totals → unknown
Total (known stages): 0B (excludes stages whose volume is unknown at plan time
— a network source, a missing or unreadable input, or a glob/regex matcher
whose discovery fails)
The seed is known for every file-backed matcher whose files can be sized at
plan time: a single-file path: source, an explicit paths: list, and a
glob: or regex: matcher. A glob/regex seed runs the same discovery resolver
the run uses — applying its exclude, min_size/max_size,
modified_after/before, take, and sort filters — and sums the matched
files’ sizes, so the estimate names exactly the bytes the run will read with no
second implementation to drift. A glob/regex that matches nothing seeds zero
(rendered as unknown, since there is no spill volume to preview). The seed is
genuinely unknown for a network source, for a missing or unreadable input
file, and for a glob/regex matcher whose discovery itself fails (an invalid
pattern, or no match under on_no_match: error) — the run surfaces the same
error at startup. Check the post-run actuals below to calibrate any estimate.
Staging plan per source
When storage.staging is enabled, the === Staging Plan === section reports,
for each source (and each discovered file under a multi-file matcher): whether
it would be staged, the resolved content-addressed staged path, and — under
on_existing = reuse — the reuse-if-fresh cache decision (hit if a committed
prior copy still matches the live source, miss if it would be re-staged):
=== Staging Plan ===
Source 'orders':
/data/in/orders-2024.csv → staged: yes, path: /mnt/local/staging/3f2a…b1.staged, reuse: hit
/data/in/orders-2025.csv → staged: yes, path: /mnt/local/staging/9c4e…07.staged, reuse: miss
The reuse prediction runs the exact freshness check (mtime + size against the
committed manifest) the real run makes, read-only — --explain copies nothing.
A source that matches no staging pattern reports
staged: no (no pattern match, reads in place); a network source reports
not stagable (network source reads in place). When staging is disabled the
section states that every source reads in place.
Cap headroom
When a spill cap is configured, --explain reports the headroom (cap minus
estimate) with the same per-invocation disclaimer the startup
cap-headroom preflight carries, and the same 80%
warning:
Cap headroom: 5000000000 bytes free (5000000000 estimated of 10000000000 cap, 50%)
[per invocation — does NOT account for sibling invocations sharing the spill
volume under partition-and-run]
Machine-readable form — --explain json
clinker run --explain json emits the whole plan as JSON for tooling (the
canvas, dashboards, CI gates). The same storage observability the text form
prints lives under a structured storage_summary object, so a consumer reads
per-stage spill estimates and the cap / staging summary without re-parsing
prose:
{
"schema_version": "1",
"nodes": [ ... ],
"node_properties": { ... },
"storage_summary": {
"spill_root": { "path": "/mnt/fast-ssd/clinker-spill", "source": "storage.spill.dir" },
"spill_disk_cap_bytes": 1000000000,
"estimated_spill": {
"per_stage": [
{ "node_name": "dept_totals", "display_name": "[aggregation:hash] dept_totals", "estimate_bytes": 1024 },
{ "node_name": "by_amount", "display_name": "[sort] by_amount", "estimate_bytes": 4096 }
],
"total_known_bytes": 5120,
"any_unknown": false
},
"spill_compression": {
"mode": "auto",
"per_operator": [
{ "node_name": "dept_totals", "display_name": "[aggregation:hash] dept_totals", "compression": "lz4" },
{ "node_name": "by_amount", "display_name": "[sort] by_amount", "compression": "off" }
]
},
"cap_headroom": {
"headroom_bytes": 999994880,
"estimated_bytes": 5120,
"cap_bytes": 1000000000,
"pct_of_cap": 0.000512,
"over_threshold": false
},
"staging": { "enabled": false, "sources": [] }
}
}
The fields mirror the text sections one-for-one: estimated_spill is the
=== Estimated Spill Volume === section (a stage whose volume is unknown at
plan time carries estimate_bytes: null and sets any_unknown: true),
spill_compression is the Spill compression: projection, cap_headroom is
the cap-headroom line (omitted when no cap is configured or the estimate is
zero), and staging is the === Staging Plan === section. The JSON and DOT
formats emit only their machine payload — the human-readable
=== Resolved Outputs === preamble the text form prints is suppressed so the
output parses cleanly.
Post-run actuals — calibrating the estimate
A real clinker run that spills prints a per-stage actual spill-volume
section at end-of-run, so you can compare it against the --explain estimate
for the same stage — the calibration loop that turns a coarse pre-run estimate
into a trustworthy one over repeated runs:
=== Spill Volume (actual, per stage) ===
dept_totals → 1048576 bytes
by_amount → 4194304 bytes
Total: 5242880 bytes (compare against the --explain estimate)
The per-stage breakdown sums to the pipeline-wide cumulative spill total. A run that stayed within memory spilled nothing and prints no section. A large estimate-vs-actual delta is the single highest-leverage signal when a pipeline starts spilling unexpectedly (the failure mode behind Polars’ documented 13.5× spill amplification, where an optimizer interaction turned 30 GB of input into 400 GB of spill with no per-stage visibility).
Note on the
--explaincompression projection. The per-operator spill-compression decision shown underSpill compression:is projected from the same column count the operator’s runtime spill writer sees, so the projectedautoverdict matches the file the run actually writes. A hash Aggregate and a grace-hash / sort-merge Combine project against their output schema (engine-stamped identity columns included), exactly the width their dispatch arms resolve compression against; an enforcer sort projects against the width of the records flowing into it — its upstream’s emitted schema — which is the width its sort buffer reads at runtime. The read path also dispatches on each spill file’s own one-byte header tag, so re-reading is robust regardless.
Distinguishing the runtime storage-abort conditions
A run that fails while spilling or staging emits one of several distinct
diagnostics so a single glance at the error tells you exactly what to fix —
instead of every disk and memory problem rendering as one ambiguous “out of
memory” message (the trap DuckDB hit in duckdb/duckdb#14142, where a temp-dir
cap was reported as “Out of Memory Error … 187.3 GiB/187.3 GiB used” and users
inspected df only to find free space). The aborts split along two axes: the
spill side (in-memory operator state landing on disk) and the staging
side (matched source files copied to local disk before they are read).
Spill aborts
| Condition | Code | What happened | What to do |
|---|---|---|---|
| Out of memory | E310 | An operator’s in-RAM state crossed the hard memory.limit (a true RSS overrun). | Raise memory.limit, reduce input, or let the operator spill. |
| Spill cap exceeded | E320 | Cumulative spill bytes crossed storage.spill.disk_cap_bytes. The volume may still have free space — you hit the configured budget. | Raise disk_cap_bytes, point storage.spill.dir at a larger volume, or reduce the spill footprint. |
| Spill volume full | E321 | The OS reported the spill volume out of space (ENOSPC). The physical disk filled. | Free space on the volume, or move storage.spill.dir to a larger mount. |
| Spill directory unavailable | (Spill) | The spill directory went bad mid-run — unmounted, remounted read-only, deleted by a cleaner, or permissions revoked. | Remount/restore the volume; stop the over-eager cleaner. |
The key separations:
- E310 vs E320 — an OOM is an in-RAM overrun; a cap-exceeded is a disk-budget stop. A run can hit E320 while comfortably inside its memory envelope, so conflating the two would point you at the wrong knob.
- E320 vs E321 — E320 is the budget you set; E321 is the disk itself
running dry. If you removed
disk_cap_bytes, an over-large run would no longer trip E320 and would instead spill until the volume filled (E321).
(A future per-operator memory-reservation surface will add a fifth, reservation-exhausted condition; it is not part of the engine yet.)
Staging-copy aborts
When storage.staging is enabled, copying a matched source to local disk can
fail in three distinct ways. Like the spill split, each has its own code so a
content-corruption problem never renders as a budget problem and vice versa.
Staging runs before any record flows, so these surface as startup-style
validation failures.
| Condition | Code | What happened | What to do |
|---|---|---|---|
| Staged copy corrupt | E335 | The local copy’s BLAKE3 digest did not match the source — the transport (e.g. a soft-mount NFS share) delivered different bytes than the source holds. | Re-run over a healthy transport, harden the mount, or stage from a stable snapshot. Do not set verify = "none" to silence it — that hides corruption, not fixes it. |
| Staging cap exceeded | E336 | The cumulative bytes staged this run would cross storage.staging.disk_cap_bytes. The volume may still have free space — you hit the configured budget, not a full disk. | Raise disk_cap_bytes, point storage.staging.dir at a larger volume, narrow storage.staging.patterns, or remove the cap. |
| Staged copy already exists | E337 | A staged copy of this source already exists and on_existing = error refuses to touch it. | Remove the existing copy, or switch on_existing to overwrite (re-stage) or reuse (reuse a fresh copy). |
The same cap-vs-full-disk separation applies here as on the spill side: E336 is the budget you set (mirroring E320), so it must not render as an out-of-space message — a physically full staging volume instead surfaces as a staging I/O error (mirroring E321). E335 is distinct from a generic staging I/O error: an I/O error means the OS reported a fault, whereas E335 means the copy completed cleanly yet still does not match the source.
Startup storage validation
Before a run spawns its first source-ingest thread — after the plan compiles
but before any input is read or any byte is spilled or staged — Clinker runs a
single comprehensive validation pass over the resolved [storage]
configuration. It rejects configurations that are physically wrong for the job,
each with a stable diagnostic code, the offending clinker.toml field, and a
clinker explain --code <CODE> pointer. Validating up front fails a
misconfigured volume while the run is still cheap to abandon, rather than after
minutes of work when the first spill or staged copy hits the bad volume.
| Code | Rejected configuration | Why |
|---|---|---|
| E330 | storage.spill.dir on an in-memory filesystem (Linux tmpfs / ramfs, Windows RAM disk). | Spilling there keeps the bytes in RAM, so it frees no physical memory and defeats the memory budget. |
| E331 | storage.spill.dir on a network filesystem (NFS / SMB / CIFS / FUSE). | A spill target on a soft-mounted share risks silent truncation and mmap data loss — the failure modes spill exists to avoid. |
| E332 | storage.staging.dir on a network filesystem. | Staging copies inputs off a flaky share; a staging dir that is itself on a share reintroduces the fragility staging exists to escape. |
| E333 | storage.staging.dir on the same physical device as a matched (staged) source. | The copy moves no I/O off the source volume, so it buys nothing while still spending time and space. Applies only to matched sources. |
| E334 | storage.spill.dir equal to storage.staging.dir. | Spill files and staged source copies are sized and cleaned up differently; sharing one directory makes accounting and cleanup ambiguous. |
The filesystem-class checks (E330–E332) read the volume type through one
cross-platform detection layer, so they behave identically on Linux, macOS,
and Windows: Linux matches the statfs f_type magic, macOS matches the
f_fstypename string, and Windows maps GetDriveTypeW. (macOS has no native
tmpfs, so E330 only ever fires on Linux and Windows.) The same-device check
(E333) compares the device id on Linux/macOS and the volume serial number on
Windows — the very same probe the staging same-volume rule uses, so there is
one consistent notion of “same device” across the whole run.
Free-space preflight
Separately from the runtime disk cap (E320) and the full-volume surface
(E321), the startup pass runs a free-space preflight: it queries the bytes
available on the spill volume and compares them to the run’s estimated spill
footprint (the sum of every blocking operator’s predicted peak state, the same
estimate --explain surfaces). When the spill volume looks too small, the run
prints a warning and continues:
W330: spill volume /var/clinker/spill has 2000000000 bytes free but the run is
estimated to spill up to 8000000000 bytes; the run may abort with a full-volume
error (E321) at the final spill — point storage.spill.dir at a larger volume or
reduce the spill footprint (raise memory.limit, partition the input)
This is advisory, not fatal: the estimate is a coarse upper bound (it
ignores spill compression and the streaming drain), so the run may well finish
within the available space. The warning exists so a long pipeline that would
die at its final spill surfaces that risk before it runs for an hour, rather
than after. The free-space query uses a cross-platform probe (statvfs on
Unix, GetDiskFreeSpaceExW on Windows) that returns a 64-bit byte count, so
the historical 32-bit f_bavail truncation never affects the comparison.
Cap-headroom preflight
When storage.spill.disk_cap_bytes is configured, the same startup pass also
runs a cap-headroom preflight: it compares the run’s estimated spill volume
to the configured cap and warns when the estimate reaches 80% of the cap.
Unlike the free-space preflight (which probes the physical volume), this checks
the run against the policy ceiling you set, so it fires even on a volume with
plenty of free space:
W331: this run is estimated to spill up to 9000000000 bytes, which is 90% of the
configured spill cap storage.spill.disk_cap_bytes (10000000000 bytes); the run
may abort with a spill-cap error (E320) before it finishes — raise disk_cap_bytes
or reduce the spill footprint (raise memory.limit, partition the input). This
headroom is per invocation: if you partition the input and run several clinker
invocations against the same spill volume and cap, they share the cap, so the
real headroom is smaller than this figure
Like W330, this is advisory, not fatal — the estimate is a coarse upper
bound, so a run that compresses well or never trips its memory budget may finish
comfortably under the cap. It fires on a normal clinker run (before ingestion,
at startup), not only under --explain, so an operator sees the signal on the
real run even when they did not explicitly inspect the plan first.
Per-invocation accounting. The cap and the headroom figure are scoped to a
single clinker invocation. Under the partition-and-run model — where you
split a large input by file or key and launch several clinker processes that
share one spill volume and one disk_cap_bytes — the physical spill volume is
shared by every sibling, so the real headroom is smaller than any one
invocation’s figure. The warning text states this explicitly rather than
silently presenting a per-invocation number as a whole-volume guarantee. Clinker
is single-process by design (one invocation = one OS process), so the engine
cannot see its siblings; the disclaimer is the honest stance.
Mid-run spill failures
The startup check guarantees the spill directory is writable when the run begins, but it can still go bad mid-run — an NFS share remounts read-only, a volume unmounts, an over-eager temp-file cleaner deletes the directory, or permissions are revoked. When a spill write fails because the directory has vanished or become read-only, the run aborts cleanly with a distinct diagnostic rather than a generic I/O error or a panic:
spill directory /var/clinker/spill became unavailable mid-run: No such file or directory
(the directory may have been unmounted, remounted read-only, deleted by an
external cleaner, or had its permissions revoked)
This surfaces the directory-level cause directly, so the fix (remount the volume, stop the cleaner, restore permissions) is obvious from the message.
Crash purge of orphaned spill directories
A run’s spill directory (clinker-spill-<random>/) is normally removed when the
run ends — a clean exit, a run that aborts with a fatal error, or even a panic
all delete it. But a SIGKILL, the Linux OOM-killer, or a power loss kills the
process before that cleanup runs, leaking the directory and every spill file
inside it. Over many crashed runs that fills the spill volume.
To prevent that, a run cleans up orphaned spill directories at startup — but
only when a spill directory is explicitly configured (storage.spill.dir),
before it creates its own. It removes only directories left by dead runs and
never touches one a concurrent run is still using.
When storage.spill.dir is not set, the spill root defaults to the OS temp
directory (std::env::temp_dir, typically $TMPDIR or /tmp), and no
startup purge runs there. In the default case a run still cleans up its own
spill directory on every exit short of a hard kill; a directory leaked into the
OS temp directory by a hard kill is the operating system’s temp-reaper’s
responsibility, not Clinker’s. The purge is confined to a configured spill root
because Clinker owns that volume but does not own the shared OS temp directory.
storage.staging — opt-in source staging
Reading source files directly from a network share (NFS, SMB) couples every run to the share’s availability and quirks: a soft-mount can silently truncate a read, and latency multiplies across many small files. Source staging copies matched source files to a local volume before the pipeline reads them, so the run works from stable local copies. It is off by default and activated per workspace by pattern match — pipelines that don’t opt in behave exactly as before.
[storage.staging]
enabled = true
dir = "/var/clinker/staging" # required when enabled
patterns = [
"/mnt/nfs/data/**",
"//fileserver/share/**",
]
disk_cap_bytes = "50GB" # optional; cap on bytes copied per run (default unlimited)
verify = "blake3" # optional; blake3 | none (default blake3)
on_existing = "overwrite" # optional; overwrite | reuse | error (default overwrite)
cleanup = "on_success" # optional; on_success | always | never (default on_success)
| Key | Default | Meaning |
|---|---|---|
enabled | false | Master switch. When false, patterns is ignored and every source reads in place. |
dir | — | Local directory the copies are written under. Required when enabled. |
patterns | [] | Glob patterns selecting which source paths to stage. A source is staged only when enabled and its path matches at least one pattern. Empty ⇒ nothing is staged. |
disk_cap_bytes | unlimited | Cumulative cap on bytes copied per run. Same byte-size grammar as the spill cap ("50GB", bare integers are bytes). |
verify | blake3 | Post-copy integrity check. blake3 hashes source and copy and requires a match — the only check that catches a soft-mount’s silent truncation. none skips the check. |
on_existing | overwrite | What to do when a staged copy of this source already exists from a prior run: overwrite re-copies unconditionally; reuse reuses the existing copy only when it is still fresh (the source’s modification time and size match what was recorded when it was staged), otherwise re-copies; error fails the run rather than touch the existing copy. See The staging cache below. |
cleanup | on_success | When staged copies are deleted relative to the run’s outcome: on_success removes them after a clean exit but keeps them after a failure so the operator can inspect the exact inputs the failed run saw; always removes them regardless; never keeps them as a persistent reuse cache for a later reuse run. See Cleanup. |
Pattern matching
patterns uses the same glob grammar as a source’s exclude: list. Each
pattern is tested against both the full path and the basename, so
/mnt/nfs/** matches a deep path by its full path while *.csv matches any
CSV by basename. ** crosses directory boundaries; * does not.
Startup validation
When enabled, staging is validated once at startup, before any input is
opened, so a misconfiguration fails the run immediately rather than at the
first copy. The run is refused when:
diris unset.dirdoes not exist, is a file, or is not writable (probed with a real create-and-delete, so a read-only mount or restrictive ACL is caught).- a
patternsentry is not a valid glob. dirsits on the same volume as a matched source. Staging within one volume copies bytes without moving I/O off the slow share — a well-documented anti-pattern — so it is refused up front rather than left to surface as a confusingly slow pipeline. The check compares the source’s and the staging dir’s storage volume (the device id on Linux/macOS, the volume mount root on Windows); pointdirat a local disk on a different volume.
The same-volume rule applies only to matched sources: a source the patterns don’t select reads in place, so its volume is irrelevant.
How a file is staged
Staging copies the matched source to your local staging directory once, then
verifies the copy against the source (with verify = blake3, the default, a
content mismatch fails the run with E335). From then on
the pipeline reads from the local copy. The same source always resolves to the
same staged file, so a later run can find and reuse a prior copy.
The staging cache (on_existing)
Because staged copies live at stable paths, a copy from a prior run is still on
disk when the next run starts (unless cleanup removed it). on_existing
decides what happens when that prior copy is found:
| Mode | Behavior |
|---|---|
overwrite (default) | Always re-stage. The prior copy is removed and the source is copied fresh. The safe default: a copy from a crashed run must not be trusted. |
reuse | Reuse the prior copy only when it is still fresh — the source’s current modification time and size both match what was recorded when it was staged. A fresh match skips the copy entirely (no bytes read off the share, nothing charged against the disk cap). A changed mtime or size means the source was rewritten, so the copy is stale and is re-staged. |
error | Fail the run with a clear diagnostic if a staged copy already exists, rather than overwrite or reuse it. For workflows that want an explicit “the cache is already populated” stop. |
reuse is the mode that turns staging into a cache: re-running the same
pipeline over an unchanged network share copies nothing on the second run. The
freshness check is mtime + size, not a re-hash, so it is cheap.
Staging is safe to run from several clinker invocations at once over a shared
staging volume: a source is copied exactly once no matter how many runs race for
it, a run always reads a complete copy, and no run fails because a sibling was
reading, cleaning up, or re-staging the same source.
Cleanup (on_success | always | never)
cleanup decides when a run’s staged copies are removed, keyed on the run’s
outcome:
| Mode | Behavior |
|---|---|
on_success (default) | Remove the copies after a clean exit; keep them after a failure (or an interrupted / DLQ-producing run) so the operator can inspect the exact inputs the run saw and re-run without re-fetching. |
always | Remove the copies when the run ends, success or failure. |
never | Keep the copies indefinitely as a persistent reuse cache. Combine with on_existing = reuse to make repeated runs over a stable source copy-free. The operator reclaims the staging dir manually (or lets the next run’s crash purge eventually reap stale entries). |
Each staged file’s manifest is removed alongside it, so cleanup never leaves a manifest pointing at a staged file that is gone.
Crash purge of orphaned artifacts
A SIGKILL, the Linux OOM-killer, or a power loss can kill a run before its
cleanup runs, leaving half-finished staging artifacts behind. To stop those
from accumulating, every run cleans up leftover artifacts from dead runs at
startup, before it stages anything. A complete staged copy is the reuse cache
and is always kept; only incomplete leftovers are reclaimed.
File permissions
Staged copies hold verbatim source records — potentially PII, credentials, or financial data — so on Unix they are created with owner-only permissions. On Windows staged files inherit the staging directory’s permissions, so restrict the directory if the volume is shared with other users.
Crash durability and the parent-directory fsync
Staged copies survive a crash: a later run finds a complete file or nothing at all, never a half-written one.
Streaming vs. Blocking Stages
Every node in a pipeline is one of two kinds at runtime, and the difference is what keeps Clinker’s memory bounded:
- Streaming stages pass records through without holding the whole input. Their memory footprint stays small no matter how large the input is.
- Blocking stages must see their entire input before they can produce any output, so they accumulate state. They stay within the memory budget and spill to disk when it gets tight, rather than holding everything in RAM.
Peak memory includes all concurrently live operator state, source queues, writer buffers, and retained intermediate records. The shared budget and spill policies govern that combined working set; the largest blocking stage alone is not a peak-memory bound.
Interactive companion: the streaming vs. blocking explainer classifies every stage of a few pipeline shapes as you change their settings, and shows the --explain lines.
Which stages stream
A stage streams when two things hold: it is one of the stages listed below, and its output goes to exactly one consumer that can take a stream, which is a Sink, the input of an Aggregate, or the driver side of a hash Combine. Any other stage, a stage that feeds two consumers, and a stage that roots an analytic window keep their output in a buffer instead. For example, in Source → Transform → Transform → Sink only the first Transform streams.
Two shapes stream and hold only one batch at a time, however large the input:
- Source → Transform → Sink chains, where the Transform has no window and its Source feeds only it. Records flow straight from the reader through the transform to the writer.
Mergeininterleavemode without aninterleave_seed, whose inputs are all Sources, each feeding only the Merge.
These hand their output straight to their one consumer, but still build their own result first:
Routewith only one branch wired to a downstream stage. Thedefault:branch counts: a Route with one condition and a wireddefaulthas two consumers, and gives each its own buffer.Mergeinconcatmode, in seededinterleavemode, or ininterleavemode fed by other stages.Aggregatewithstrategy: streaming— when the input is pre-sorted on the group key, each group is emitted as soon as the key advances. (See Aggregate Nodes.)- A hash
Combine’s output, and its driver side, which streams in against the already-built lookup table. - A range
Combine’s output, once it has sorted both sides. Sink— a Sink writes each record to its writer as it arrives. A Sink withsort_order,splitor a per-source-file path takes no stream, so the stage before it keeps a buffer.
Document boundaries (the signals behind $doc.*) flow inline with records through streaming stages, so a document’s close always trails its last record.
Which stages block
A stage blocks when its result depends on records it has not seen yet:
sort— the full input must be present before the first sorted record is known.- Hash
Aggregate— a group’s final value depends on every member, so the group table retains aggregate state for every live group. (Astreaming-strategy Aggregate over pre-sorted input is the exception above.) - A
Combine’s build side — the lookup table is built in full before any driver record is matched. The probe side streams; the build side materializes. - Time-windowed and correlation-key Aggregates — these hold their group state for windowing or for the correlation commit, so they materialize.
A blocking stage keeps its accumulated state inside pipeline.memory.limit and spills to disk when the budget gets tight.
Seeing the classification
clinker run <pipeline>.yaml --explain annotates every node with its class in the Physical Properties section:
sink.report:
buffer: streaming
aggregation.dept_totals:
buffer: materialized
buffer: streaming marks a stage that holds only a small in-flight slice; buffer: materialized marks one that holds a whole stage’s output and may spill it. The annotation follows the rules the executor applies at runtime, with two known gaps:
- A Sink with
reconstruct_envelope: trueturns streaming output off at runtime, but--explainstill reportsbuffer: streamingfor that Sink and for a Transform feeding it. - Under a correlation key,
--explainreports each Sink asbuffer: streaming, though the correlation commit writes its rows.
See Explain Plans and Memory Tuning.
Under dlq_granularity: document, as under a correlation key, streaming handoffs are off for the whole pipeline: no Transform, Merge, Route, Aggregate or Combine hands its output to a streaming consumer, so none of them reports buffer: streaming. Each Sink reports buffer: materialized, because it holds every open document’s records until the document’s verdict is final. A Source read by a single Transform still hands its records straight to that Transform, and keeps buffer: streaming.
Tuning the batch size
The number of records a streaming stage hands downstream at a time is set by pipeline.batch_size (default 2048), with an optional per-transform override. Smaller batches lower in-flight memory at the cost of more per-batch overhead; larger batches do the reverse. The batch size changes only the memory profile of streaming handoffs — never their output, and never the behavior of blocking stages.
Optimizing Pipelines
Clinker keeps memory bounded and spills to disk automatically, so most pipelines run fine with no tuning at all. When you do need a pipeline to run faster or in less memory, a handful of authoring choices do nearly all the work. This page is the practical checklist; the engine mechanics behind each tip live in the separate Engine Internals book.
Let stages stream instead of buffer
The cheapest pipeline is one where records flow straight through without being held in memory. A Source → Transform → Sink chain streams end to end — no intermediate stage is materialized. You get this automatically; the things that break it are fan-out (a Route with several branches, an output that forks) and blocking operators (sort, hash aggregation, the build side of a Combine).
Practical implication: keep the hot path simple. A filter-and-reshape job that’s just Source → Transform → Sink already runs at minimal memory. See Streaming vs. Blocking Stages for which operators stream and which block.
Make aggregation stream with sort_order
A hash Aggregate holds one entry per distinct group key in memory — fine for low-cardinality keys, expensive for high-cardinality ones. If your input is already sorted on the group-by keys, declare it:
- type: source
name: txns
config:
type: csv
path: ./data/transactions_sorted.csv
sort_order:
- { field: account_id, order: asc }
schema:
- { name: account_id, type: string }
- { name: amount, type: float }
- type: aggregate
name: per_account
input: txns
config:
group_by: [account_id]
cxl: |
emit total = sum(amount)
With a matching sort_order, the optimizer switches the aggregate to streaming — it emits each group as the key advances and holds only one group at a time, regardless of cardinality. To make the requirement explicit (and turn a silent fallback to hash aggregation into a compile error), set strategy: streaming. See Aggregate Nodes → Strategy hint.
sort_orderis checked for each file as it is read. Withon_unsorted: warn(the default) a file that is out of order is sorted before it is released and aW307warning names it; withon_unsorted: errorthe file is rejected. A wrong declaration costs time rather than correctness, but declare it only when the data really is sorted. See Source Nodes → Sort order.The order only reaches the Aggregate if every stage in between keeps it: a Merge, a Combine,
distinct, or a Transform that writes one of the sort fields (evenemit account_id = account_id) drops it. The sort order explainer shows which stages keep it.
Choose the Combine driver side deliberately
A Combine holds each non-driving (build-side) input in memory as a lookup table, then streams the driver against it. So:
- Put the smaller relation on the build side and drive with the larger stream — you iterate the big input once and keep only the small one resident. Plan for roughly 1.5–2× the build file’s size in memory.
- The driver also sets output order and which side’s correlation identity propagates, so pick it for those reasons too. See Combine Nodes and Correlation Keys → Combine interaction.
A large build side isn’t a failure — the join spills to disk automatically — but spilling is slower than staying in memory, so sizing the driver right is the main lever.
Size the memory budget
The default budget is 512 MB. Raise it when a pipeline does high-cardinality aggregation or large joins and you have the RAM; lower it to be a good neighbor on a shared box. The budget is a target, not a hard wall — stages spill rather than fail when they exceed it.
pipeline:
name: my_pipeline
memory:
limit: "1G"
Full sizing guidance and the backpressure knob are in Memory Tuning.
Reduce intermediate state
Less data in flight means less to buffer and spill:
- Filter early. Drop records you don’t need in the first Transform, before they reach a blocking stage.
- Project narrowly. Emit only the fields downstream stages actually use; carrying wide records through a sort or aggregate costs memory per row.
- Aggregate before joining when you can — feeding a small rolled-up relation into a Combine is cheaper than joining raw rows and aggregating after.
Confirm with --explain and metrics
Before running, clinker run pipeline.yaml --explain annotates each node with buffer: streaming or buffer: materialized, so you can see which stages will dominate memory. After a run, the metrics spool reports peak_rss_bytes — if it consistently approaches your limit, raise the budget or cut intermediate state. See Explain Plans.
Metrics & Monitoring
Clinker writes per-execution metrics as JSON files to a spool directory. These files can be collected into an NDJSON archive for ingestion into monitoring systems.
Interactive companion: Where did my rows go? follows every row of ten small pipelines into the end-of-run counters, and shows which rows are in no counter at all.
Enabling metrics
There are three ways to enable metrics collection, listed from highest to lowest priority:
CLI flag:
clinker run pipeline.yaml --metrics-spool-dir ./metrics/
Environment variable:
export CLINKER_METRICS_SPOOL_DIR=./metrics/
clinker run pipeline.yaml
YAML config:
pipeline:
metrics:
spool_dir: "./metrics/"
When metrics are enabled, each execution writes one JSON file to the spool directory, named <execution_id>.json.
Metrics schema
Each metrics file follows schema version 3. The collector rejects spool files written under an older schema version, so upgrading clinker across a schema bump means draining the spool first.
{
"execution_id": "01912345-6789-7abc-def0-123456789abc",
"schema_version": 3,
"pipeline_name": "customer_etl",
"config_path": "/opt/clinker/pipelines/daily_etl.yaml",
"hostname": "prod-etl-01",
"started_at": "2026-04-11T10:00:00Z",
"finished_at": "2026-04-11T10:00:05Z",
"duration_ms": 5000,
"exit_code": 0,
"records_total": 50000,
"records_ok": 49950,
"records_written": 49950,
"records_dlq": 50,
"records_null_dropped": 0,
"execution_mode": "Streaming",
"peak_rss_bytes": 134217728,
"thread_count": 4,
"input_files": ["./data/customers.csv"],
"output_files": ["./output/enriched.csv"],
"dlq_path": "./output/errors.csv",
"error": null,
"retraction": {
"groups_recomputed": 0,
"partitions_dispatched": 0,
"iterations": 0,
"degrade_fallback_count": 0,
"synthetic_ck_columns_emitted_total": 0,
"synthetic_ck_fanout_lookups_total": 0,
"synthetic_ck_fanout_rows_expanded_total": 0
},
"per_source_record_counts": { "customers": 50000 },
"per_source_dlq_counts": { "customers": 50 }
}
Field reference
| Field | Type | Description |
|---|---|---|
execution_id | string | UUID v7 or custom --batch-id value |
schema_version | integer | Schema version of this payload; currently 3 |
pipeline_name | string | The name from the pipeline YAML |
config_path | string | Absolute path to the config file |
hostname | string | Machine hostname |
started_at | string | ISO 8601 UTC timestamp |
finished_at | string | ISO 8601 UTC timestamp |
duration_ms | integer | Wall-clock duration in milliseconds |
exit_code | integer | Process exit code (see Exit Codes) |
records_total | integer | Records read from every Source, including records rejected while being read (for example a value that does not fit its declared type) and the lookup inputs of a Combine. per_source_record_counts splits it by Source |
records_ok | integer | Distinct source records that reached at least one output. Under inclusive Route fan-out one input matching N branches counts once. An Aggregate output row counts as one record of its group (the first one read), so the other records of the group are not counted; a Combine output row counts as its driver record, so lookup records are not counted; a whole-input Aggregate (group_by: []) over no records still writes one row and counts 1 |
records_written | integer | Total writes across all sinks. Equals records_ok for single-output exclusive pipelines; exceeds it under inclusive Route fan-out or multiple Sinks |
records_dlq | integer | Rows written to the dead-letter queue, collateral rows included. A source row counts once for each failure it took part in, with or without a correlation key |
records_null_dropped | integer | Records excluded by a Sink null_order: drop sort field (see Sort order). Counts exclusions, not distinct source records: like records_written, one source record dropped at two Sinks counts twice. Absent from spool files written before this counter existed, where it reads as 0 |
execution_mode | string | DAG-derived execution summary: Streaming (no full-stage materialization required) or TwoPass (a blocking stage forces an accumulation pass) |
peak_rss_bytes | integer/null | Peak resident set size in bytes, sampled across chunk boundaries on Linux, macOS, and Windows. null on platforms where RSS sampling is unavailable |
thread_count | integer | Thread pool size used |
input_files | array | Paths to all source files |
output_files | array | Paths to all output files written |
dlq_path | string/null | Path to the DLQ file, or null if none |
error | string/null | Error message on exit 1/3/4, or null on success (exit 0) and partial success (exit 2) |
retraction | object | Correlation-key retraction counters (see below). All-zero on strict pipelines, which never enter the relaxed loop |
per_source_record_counts | object | Ingest record count per Source node, keyed by node name. A source that read zero records is present with a count of 0 |
per_source_dlq_counts | object | DLQ entry count per Source node; sources with zero DLQ entries are absent. The values sum to at most records_dlq — see the note below |
The sum of per_source_dlq_counts values is at most records_dlq, and can
be less: a failure in a Combine emit or a post-aggregate row is not traceable
to a single declared source, so it is counted in records_dlq but not in this
per-source breakdown. For pipelines whose dead-letters all originate at a
declared source, the two match exactly.
The retraction object carries the relaxed correlation-key retraction
orchestrator’s counters: groups_recomputed, partitions_dispatched,
iterations, degrade_fallback_count,
synthetic_ck_columns_emitted_total, synthetic_ck_fanout_lookups_total,
and synthetic_ck_fanout_rows_expanded_total. Every field is 0 on
strict pipelines and on relaxed pipelines that never trigger a retraction.
See Correlation Keys for the underlying
mechanism.
Collecting metrics
The spool directory accumulates one file per execution. Use clinker metrics collect to sweep them into an NDJSON archive:
clinker metrics collect \
--spool-dir ./metrics/ \
--output-file ./metrics/archive.ndjson \
--delete-after-collect
This appends all spool files to the archive (one JSON object per line) and removes the originals. The NDJSON format is compatible with most log aggregation and monitoring tools.
Preview without writing:
clinker metrics collect \
--spool-dir ./metrics/ \
--output-file ./metrics/archive.ndjson \
--dry-run
Integration with monitoring systems
Grafana / Prometheus
Parse the NDJSON archive with a log shipper (Promtail, Filebeat, Vector) and create dashboards tracking:
duration_ms– execution time trendsrecords_dlq– data quality over timepeak_rss_bytes– memory utilization
Datadog
Ship NDJSON to Datadog Logs, then create metrics from log attributes:
# Example: tail the archive and ship to Datadog
tail -f ./metrics/archive.ndjson | datadog-agent log-stream
ELK Stack
Filebeat can ingest NDJSON directly:
# filebeat.yml
filebeat.inputs:
- type: log
paths:
- /var/log/clinker/metrics.ndjson
json.keys_under_root: true
Simple alerting with jq
For environments without a full monitoring stack, use jq to query the archive directly:
# Find all runs with DLQ entries in the last 24 hours
jq 'select(.records_dlq > 0)' metrics/archive.ndjson
# Find runs that exceeded 400MB RSS
jq 'select(.peak_rss_bytes > 419430400)' metrics/archive.ndjson
# Average duration by pipeline
jq -s 'group_by(.pipeline_name) | map({
pipeline: .[0].pipeline_name,
avg_ms: (map(.duration_ms) | add / length)
})' metrics/archive.ndjson
Workspace OTLP and lineage policy
Deployment observability is optional and disabled when clinker.toml has no
[observability] table. It is workspace policy, not pipeline YAML, and does
not participate in the compiled plan’s semantic fingerprint. A present table
is one complete policy: callers may supply a complete resolved replacement
only when the workspace table is absent; individual fields are never merged.
The workspace loader validates this policy without opening a source, output,
attempt directory, worker, credential provider, or network connection. It
keeps the Collector endpoint as length-bounded raw text exactly as authored.
The one shape it requires of that text is the shape it requires of every other
authored string in the table: non-empty, within its byte cap, with no
surrounding whitespace and no embedded control character. Padding and a
carriage return are not part of an endpoint under any parse, and refusing them
here names observability.otlp.endpoint and hands you a pasteable correction,
where the later network boundary can only report that some endpoint was
unusable.
The network admission boundary parses the text itself later, before any delivery effect; scheme, authority, credentials, paths, query strings, fragments, normalization, and the fixed OTLP signal routes are deliberately not decided by the workspace parser. Collector reachability is not a configuration admission check.
A complete fixed-capacity example is:
[observability]
arena_bytes = "4MB"
ordinary_lane_bytes = "3MB"
high_severity_lane_bytes = "1MB"
max_batch_bytes = "256KB"
max_attributes_per_event = 32
max_attribute_bytes = "4KB"
drop_policy = "drop_newest"
sample_every = 1
rate_limit_per_second = 1000
rate_limit_burst = 1000
flush_timeout_ms = 15000
[observability.otlp]
endpoint = "https://collector.example.com"
connect_timeout_ms = 1000
request_timeout_ms = 5000
retry_max_attempts = 3
retry_total_timeout_ms = 10000
max_response_bytes = "64KB"
[observability.otlp.auth]
mode = "none"
[observability.lineage]
queue_bytes = "1MB"
max_event_bytes = "64KB"
drop_policy = "drop_newest"
flush_timeout_ms = 5000
identity_mode = "external"
[[observability.lineage.dataset]]
node = "source_customers"
canonical_datasource = "s3://warehouse/customers"
[[observability.lineage.dataset]]
node = "output_customers"
catalog_namespace = "analytics"
catalog_name = "customers_clean"
[[observability.field_policy]]
event = "run.completed"
field = "records_written"
action = "allow"
[[observability.field_policy]]
event = "transform.customer_seen"
field = "customer_id"
action = "hash"
[[observability.field_policy]]
event = "transform.customer_seen"
field = "email"
action = "replace"
replacement = "[redacted]"
Byte-size strings use decimal units (1KB = 1,000 bytes and
1MB = 1,000,000 bytes). The fixed defaults and hard ceilings are:
| Key | Default | Hard ceiling or relationship |
|---|---|---|
arena_bytes | "4MB" | "64MB"; equals the exact sum of both lane caps |
ordinary_lane_bytes | three quarters of the arena | "64MB" and disjoint from the high-severity lane |
high_severity_lane_bytes | one quarter of the arena | "64MB" and disjoint from the ordinary lane |
max_batch_bytes | "256KB" | "1MB" and no larger than either lane |
max_attributes_per_event | 32 | 256 |
max_attribute_bytes | "4KB" | "64KB" |
sample_every | 1 | 1,000,000 |
rate_limit_per_second | 1,000 | 1,000,000 |
rate_limit_burst | 1,000 | 1,000,000 |
flush_timeout_ms | 15,000 | 60,000 |
otlp.connect_timeout_ms | 1,000 | 60,000 and no greater than request timeout |
otlp.request_timeout_ms | 5,000 | 60,000 and no greater than retry total |
otlp.retry_max_attempts | 3 | 10 |
otlp.retry_total_timeout_ms | 10,000 | 60,000 and no greater than flush timeout |
otlp.max_response_bytes | "64KB" | "1MB" |
lineage.queue_bytes | "1MB" | "64MB", reserved independently of the telemetry arena |
lineage.max_event_bytes | "64KB" | "1MB" and no larger than its lineage queue |
lineage.flush_timeout_ms | 5,000 | 60,000 |
Every byte default above is the quantity its own spelling parses to, so writing a default out in full changes nothing.
Sizing the arena
The two lanes partition the arena exactly: no telemetry byte is charged twice, and none of the arena is unreachable. You may write any of the three and leave the rest to be worked out from what you wrote:
arena_bytesalone — the lanes split it three-to-one, whatever its size.arena_bytes = "8MB"gives a 6 MB ordinary lane and a 2 MB high-severity one.- One lane alone — the arena stays at its default and the other lane takes the remainder.
- Both lanes — the arena is their sum.
- All three — the equality is checked rather than adjusted, and a disagreement is refused before the run starts.
The arena is the budget in every case: a lane is never allowed to grow it. A lane that does not fit inside the arena is refused, naming both keys.
An exported attribute longer than max_attribute_bytes is cut to fit and
marked with a trailing …, and the marker is charged against the same cap, so
a marked value is never longer than an unmarked one would have been. The mark
matters because a bare prefix is not obviously a prefix: an amount of
123456789 cut to four bytes reads as 1234, and a timestamp cut short is
still a well-formed timestamp. A dashboard reading 1… fails visibly; one
reading 1234 charts a wrong number. Note that the mark is a signal, not a
guarantee — a free-text value that genuinely ends in … is indistinguishable
from a truncated one.
Node names are exported verbatim under the same treatment. A name is never dropped for the characters it contains, so a Transform named with a space or a non-ASCII character still produces a span.
Both delivery paths admit with drop_policy = "drop_newest"; there is no
blocking, unbounded, or disk-spool spelling. The telemetry arena contains two
disjoint lanes: trace, debug, and info signals occupy the ordinary lane,
while warn and error occupy the high-severity lane. The lineage queue is a
separate reservation and cannot be expressed as an alias of either telemetry
lane or the arena.
sample_every = N keeps one in every N signals, counted within each lane
separately. The lanes are disjoint precisely so ordinary volume cannot crowd
out problems, and sampling honours that: with sample_every = 10, a Transform
emitting nine per-record info events for every error still keeps one in ten
of its errors, and that fraction does not change when the info volume does.
A run that raises its per-record logging therefore does not quietly thin out
its error reporting.
What the machine terminal reports about export
Under --machine ndjson-v1 the terminal event carries an observability
object summarising what the exporter did. It holds one counter group per
signal — logs, metrics, and traces, each with accepted, rejected,
attempts, and failures — plus flush_complete.
Read flush_complete before reading the counters. When it is true the
counts are the run’s final accounting. When it is false the exporter did not
get to the end of the flush: either it ran past flush_timeout_ms, or it could
not take the signal arena from the pipeline before giving up on it. Either way
the counts are what had been recorded at that point, deliveries may still have
been in flight, signals may remain that were never sent, and a low accepted
means “we stopped counting” rather than “the collector refused them”. Those call
for different responses, and only one of them is a collector problem.
Any 2xx answer counts as delivered. A collector’s own success status is
200, and its body is where a partial success declares the records it refused
— those are what rejected counts. A gateway in front of a collector may
answer 202 Accepted or 204 No Content instead, and there is no such body to
read: the whole chunk counts under accepted, because the answer says it was
taken. A rejection is a 4xx or a 5xx.
A delivery cut short after its request was fully sent is not sent again, even
with attempts left in the budget. The collector may already hold that batch —
a reply that is only slow cannot be told apart from one that was lost — and
repeating it would ingest the same log records twice and count the same
monotonic sums twice, which is wrong rather than merely unconfirmed. Such a
batch is counted under failures, having spent fewer attempts than
retry_max_attempts allowed: its delivery is unconfirmed, not known to have
failed.
A collector that habitually answers slowly needs a larger
otlp.request_timeout_ms, not more attempts.
A retryable status (429, 502, 503, 504) waits before the next attempt.
The wait starts at a fixed 100 ms and doubles before each retry after the
first, so a collector recovering from load is not handed the whole attempt
budget inside one of its recovery windows.
A Retry-After on that answer overrides the backoff whenever it asks for
longer: the collector said how long “later” is, and the export waits at least
that. Only the delay-in-seconds spelling is read. If the field asks for longer
than otlp.retry_total_timeout_ms leaves — or arrives in the HTTP-date
spelling, which this exporter does not parse — the delivery stops there rather
than re-sending sooner than the collector asked or sleeping out the whole
deadline. It is counted under failures with the throttle as its cause and
retry_with_backoff as its advice; retrying belongs to whatever supervises the
run. otlp.retry_total_timeout_ms remains the ceiling on all of it, so no
requested wait can extend an export.
Delivery outcomes never change execution, publication, or the process exit status; the summary is an observation about the export, not about the run.
What the machine terminal reports about admission
The per-signal groups above are export-side, and an exporter can only count
what reached it. A run that discarded most of its signals at the arena would
otherwise report accepted = N, rejected = 0, flush_complete = true — a clean,
complete-looking export of a silently truncated dataset. The admission object
beside them says what the arena took and what it refused, before any export:
"admission": {
"counts_complete": true,
"accepted": 9,
"dropped": {
"sampled": 8, "rate_limited": 0, "queue_full": 0, "contended": 0,
"oversize": 0, "invalid_identity": 0, "undecodable": 0
},
"lanes": {
"ordinary": { "sampled": 4, "queue_full": 0, "retained_bytes": 0, "capacity_bytes": 32000 },
"high_severity": { "sampled": 4, "queue_full": 0, "retained_bytes": 0, "capacity_bytes": 32000 }
},
"fields": { "denied": 0, "truncated": 0, "limit_dropped": 0, "missing": 0 },
"arena_recoveries": 0,
"retained_bytes": 0, "peak_retained_bytes": 2477, "capacity_bytes": 64000
}
counts_complete says whether the rest of the object is a final accounting.
These counters are read from the producer’s arena, and the arena keeps changing
while the exporter drains it — undecodable is credited at drain. A flush that
ran to completion joined the exporter first, so nothing was left to credit and
counts_complete is true. A flush that expired on flush_timeout_ms
detached an exporter that is still draining, so the read landed mid-drain and
counts_complete is false: every number below it is whatever had been
reached, and a low one means “we could not finish counting” rather than
“nothing was lost”. The flush stays bounded either way — a finishing run does
not wait on an unresponsive collector — so this flag, not a longer wait, is
what keeps a truncated view from looking complete.
accepted counts the logs and spans the arena took. Metric points are
coalesced into fixed counters rather than admitted as signals, so none of them
appear here.
Each key under dropped is one reason a signal never became exportable. They
are named as OpenTelemetry error.type values — a full arena is queue_full,
not full — so these map onto SDK self-observability metrics without a
rename. undecodable is the one member of the set that is also counted in
accepted: those signals were admitted, then could not be read back at drain.
fields is a different kind of number and must not be added into a loss
total. Those counts describe what became of the fields of records that were
accepted — values denied, values truncated, attributes dropped at the per-event
cap, and values a directive requested that the record did not carry. They
reduce what a record says; they never discard one.
The first three are policy doing what you configured it to do. missing is
not: a transform’s log directive asked for a column and the record did not
have one. Most such requests are refused when the pipeline compiles (E374).
What reaches this counter is the case the planner cannot decide — a selector
naming a column that arrives through an open composition port. It is credited
where the signal is built, so under a sampling policy it sees one miss per
sampled event rather than one per record: read it as “this is happening”, not
as how often.
arena_recoveries is neither a drop nor a quantity of anything lost. It counts
the times the arena resumed from a poisoned lock — telemetry panicked while
holding its own guard, and the arena carried on rather than taking the run down
with it. A non-zero value says every counter beside it was produced by a
subsystem that faulted mid-run. Treat it as a defect report against Clinker,
and read the run’s other telemetry numbers with that in mind; the pipeline’s
own results are unaffected, because telemetry never changes them.
Reconciling admission against export
Every signal counted under dropped is one the collector never saw. A
delivery accounts for every item in its chunk — a chunk travels whole, so
accepted + rejected equals the items handed to the exporter whatever the
outcome was. That gives an exact identity for a run where flush_complete is
true:
admission.accepted
= (logs.accepted + logs.rejected)
+ (traces.accepted + traces.rejected)
+ admission.dropped.undecodable
- 1
The - 1 is the run-lifecycle span, which the exporter synthesizes at the
final flush rather than drawing from the arena. metrics has no term because
metric points are not admitted signals. When flush_complete is false the
export counters are a partial accounting by definition and the identity does
not apply; admission.counts_complete is false on that same run, because the
arena side was read while the detached exporter was still draining it.
Any shortfall against that is arena loss, and dropped says which kind.
Why the lanes are split
sampled and queue_full are the two refusals that can cost an error while
the ordinary lane is the thing under pressure, so both are attributed per lane.
That is what makes the sampling guarantee above checkable from a run’s own
accounting: with sample_every = 10, lanes.high_severity.sampled holds at
one in ten of the high-severity signals produced however far
lanes.ordinary.sampled climbs beside it. A single total cannot show that, and
an author reading one would have no way to tell what share of their errors
survived.
rate_limited and contended have no per-lane spelling. The rate limiter and
the arena lock are properties of the shared arena rather than of a lane.
retained_bytes is what the lane still held when the run finished, and
capacity_bytes is its reservation. Note that max_batch_bytes is the
per-slot bound, so a lane holds lane_bytes / max_batch_bytes signals between
drains — that ratio, not the byte size alone, is what queue_full is measured
against.
Telemetry loss without --machine
A run without --machine ndjson-v1 discards the terminal object entirely, so
when anything was dropped Clinker writes one line to standard error:
clinker: telemetry admission outcome: accepted=9 dropped=8 sampled=8 rate_limited=0 queue_full=0 contended=0 oversize=0 invalid_identity=0 undecodable=0 ordinary_sampled=4 ordinary_queue_full=0 high_sampled=4 high_queue_full=0 missing_fields=0 arena_recoveries=0 counts_complete=true
It mirrors the lineage delivery line, including its suppression rule: a run that dropped nothing prints nothing. A line that appeared on every run reading all zeroes is noise an operator learns to skip, and the one run that did lose signals would be skipped with it.
missing_fields and arena_recoveries break that silence on their own. Both
sit outside dropped, and neither is anything an operator asked for: an
attribute the collector never received, and telemetry having panicked under its
own guard. The denied and truncated field counters stay silent by contrast,
because they are policy doing exactly what it was configured to do — and they
are not on this line at all.
The suppression is on the counters being final and clean, not on their reading
zero. A run whose flush expired prints the line whatever the numbers say, with
counts_complete=false: all-zero counts taken mid-drain are not evidence of a
clean run, and staying silent on them would report a run that may well have
lost signals as one that certainly did not.
Like the lineage line, this is an observation. Telemetry loss does not change execution, publication, the machine terminal result, or the exit status.
Runtime ownership and failure isolation
The deployment path keeps capability ownership narrow:
| Boundary | Owned capability |
|---|---|
| Workspace plan/config | Secret-free raw endpoint text plus numeric, capacity, retry, and deadline bounds. |
| Network | The sole endpoint admission, a private admitted-endpoint proof, fixed OTLP signal routes, and transport. |
| Executor | The real log, metric, and trace producers plus the fixed-memory telemetry arena. |
| Lineage | Canonical/catalog dataset identity, authorized subset and symlink facts, and independently bounded event delivery. |
| CLI | Pre-effect composition, one immutable lifecycle-fact source, worker lifecycle, and separate typed delivery outcomes. |
The workspace policy loader owns only the secret-free raw endpoint string.
For an enabled run, the CLI’s first capability transition calls the network
crate’s sole endpoint-admission API. That API accepts one HTTPS origin and
derives exactly three routes: /v1/logs, /v1/metrics, and /v1/traces.
Relative, malformed, HTTP, credential-bearing, path-bearing, query-bearing,
fragment-bearing, and already signal-specific endpoint text is rejected — and
so is text that is not exactly the origin it names: surrounding whitespace, or
any embedded control character such as a carriage return, is refused rather
than trimmed or ignored, so no endpoint value can smuggle a header into a
request. Each is rejected as
observability.otlp.endpoint with a pasteable HTTPS-origin correction before
source discovery, output attempts, arena reservation, worker construction, or
network effects. Rejected text is not echoed.
After admission, the CLI combines the admitted origin with the configured
request, retry, response, arena, and flush bounds in one immutable run-local
bundle. auth.mode = "none" is the supported production capability today and
sends no credential headers. auth.mode = "reference" remains a logical,
secret-free policy name, but the run fails before exporter effects until the
AUTH-01 credential applicator supplies that capability; the applicator will
not be allowed to change the admitted origin or fixed routes.
Logs, metrics, and traces share one finite telemetry arena and exporter worker,
but retain distinct typed per-signal delivery outcomes and fixed aggregate
counters. Those producers are the executor’s actual lifecycle, runtime, and
terminal producers; the transport does not invent equivalent events. The
OpenLineage worker has its own queue, byte cap, sink, deadline, counters, and
typed outcome. The arena allocation and both worker spawns complete before
source discovery, staging, publication-attempt creation, sink writes, or a
lineage START; inability to create either worker fails admission without
those effects as observability.delivery.failed, with exit 4 and
retry_with_backoff. Invalid endpoint, authentication, identity, and bounds
policy remains observability.configuration.invalid with do_not_retry.
Both paths copy the same batch ID, execution ID, semantic-plan
algorithm/version/digest, and terminal facts from one immutable lifecycle
snapshot; neither path reconstructs or owns those facts. Collector partial
acceptance, rejection, transport failure, shutdown, or flush expiry, and
lineage drop, sink failure, or deadline expiry remain optional observations:
they do not change final or DLQ bytes, process status, the machine terminal
result, publication inventory, visible finals, or retained failed-attempt
evidence. The machine terminal exposes aggregate per-signal counters only; it
does not flatten or replace either typed delivery outcome.
Field policy is applied before telemetry enters the arena. Denied values never
reach Collector request bodies, OpenLineage events, counters, diagnostics, or
machine records. With no [observability] table, Clinker performs no endpoint
admission, arena reservation, worker creation, or exporter I/O.
External --lineage and --lineage-events exports serialize each complete
OpenLineage event within lineage.max_event_bytes before attempting immediate
admission to the byte-bounded lineage queue. A full queue drops the newest
event instead of delaying the finite job. One lineage-only synchronous worker
owns the destination and receives no output, DLQ, publication, or machine-mode
authority. At completion Clinker waits no longer than
lineage.flush_timeout_ms; dropped events, write or flush failure, and a
deadline-exceeded worker are reported separately on standard error and do not
change the authoritative ETL/publication result. The explicit
local_diagnostic_paths compatibility mode remains a local synchronous file
or console export and cannot use this external delivery path.
Exported signal shapes
Clinker speaks OTLP/JSON over the three fixed routes. What it puts on the wire:
Traces. One run is one trace. The lifecycle span clinker.run is the trace
root and every Transform and Sink span is a child of it, so a collector can
reconstruct the whole run from any single span. Each span carries a trace id, a
span id, and both startTimeUnixNano and endTimeUnixNano.
A Transform is one span, emitted when the Transform finishes and covering the
interval it ran for. It is not a start record followed by an end record.
That shape was not exportable: a span requires both timestamps, so a
start-only record is not a valid span, and two independently admitted halves
can be sampled or dropped separately, leaving a collector holding one half of a
pair it cannot use. If you want to observe that a Transform has begun while it
is still running, that signal is the clinker.transform.started metric below,
which is recorded before the work runs and exported on the normal metric
cadence — the span’s startTimeUnixNano then tells you exactly when it began.
Metrics. The Transform counters — clinker.transform.started,
clinker.transform.completed, clinker.transform.records, and
clinker.transform.errors — are exported as monotonic sums with delta
aggregation temporality. Each exported point is the count accumulated since the
previous export, not a running total, and carries the startTimeUnixNano and
timeUnixNano bounding that interval. Sum the deltas to get the run total;
reading any single point as an absolute value will understate the run, often by
a large factor on a long one.
Each real Sink writer work unit likewise emits one clinker.sink.started and
exactly one of clinker.sink.completed, clinker.sink.failed, or
clinker.sink.interrupted. clinker.sink.records, clinker.sink.errors, and
clinker.sink.bytes report the rows handled, errors observed, and serialized
bytes accepted by that writer boundary. clinker.sink.truncations counts the
values a fixed-width writer cut to fit a truncation: warn column — the same
count the end-of-run W367 warning reports. Synchronous, fused streaming, and
correlation-deferred writers all use the same counter names and one closed
clinker.sink span. A failed flush can therefore report bytes accepted before
the destination rejected the flush; the failed terminal counter remains the
authoritative outcome.
Each dead-letter (DLQ) file a run writes is a work unit of its own. It emits
one clinker.dead_letter.started when the file receives its first row, and
exactly one of clinker.dead_letter.completed (the file was written in full
and handed to publication), clinker.dead_letter.failed (a write or flush
failed, or the run failed), or clinker.dead_letter.interrupted (the run was
interrupted). clinker.dead_letter.records and clinker.dead_letter.bytes
report the rows written to the file and the bytes its writer accepted, header
included. Each unit closes one clinker.dead_letter span whose
clinker.logical_node attribute is dead_letter[<n>], where <n> is the
file’s position, counting from 0, in the === Dead-Letter Output === section
of clinker run --explain; the path never appears. A DLQ file that receives
no rows emits nothing, and neither do dead letters that have no destination:
count those with records_dlq. Preview runs write no DLQ file and emit no
dead-letter signals.
clinker.correlation.group_overflows counts the correlation-key groups that
went over error_handling.max_group_buffer, one per group when the run
commits it. It counts a group whose rows all failed on their own too, which
writes no group_size_exceeded row, so it can exceed the number of
group_size_exceeded rows in the DLQ. See
Correlation Keys.
An instrument that recorded nothing in an interval carries no points for it. That is an ordinary interval, not a malformed export: the batch is delivered and the points its other instruments did record arrive intact.
Transforms that a correlated commit re-runs. A pipeline with a relaxed
correlation-key aggregate converges at commit time: the engine re-runs the
transforms downstream of that aggregate, retracting the rows that turned out to
fail, until the result stops changing. Only the converged result is published,
and the exported signals describe it the same way — one span covering the whole
convergence, one clinker.transform.started, one clinker.transform.completed,
and record and error counts taken from the converged pass rather than added up
across the discarded ones. An every: cadence continues across the passes
rather than restarting on each. So a transform inside a convergence reports the
rows the run actually carried, and its counters stay summable alongside every
other transform’s.
A convergence that does not finish — a failure in one of the re-run transforms,
or an interrupting signal — still reports every transform it had passed over.
Each gets one span with an ERROR status covering the interval it ran for, one
clinker.transform.started, and the record and error counts the interrupted
pass had reached. It gets no clinker.transform.completed: nothing completed,
and that counter is what tells you whether everything did.
Logs. Each emission of an authored log: directive becomes one OTLP log
record: severityText from the directive’s level, the directive’s message
as the body, the event name as the clinker.event attribute, the three
run-correlation attributes described under
Authentication and privacy, and whichever
requested record fields the field policy allowed, hashed, or replaced.
Request size. Delivery is bounded per request, not per record. A drained
batch that would exceed one request’s byte budget is split across as many
requests as it needs rather than discarded. max_batch_bytes bounds one stored
record inside the arena; it is not the request size, and the two are not the
same number.
Authentication and privacy
Authentication is always explicit. Credential-free delivery uses exactly:
[observability.otlp.auth]
mode = "none"
Referenced delivery retains one provider-neutral logical name:
[observability.otlp.auth]
mode = "reference"
reference = "telemetry/production"
The reference is not a credential. A later run-local authentication provider must resolve it before effects. Inline headers, bearer/basic values, environment-variable names, and mixed auth variants are rejected; omission does not mean anonymous delivery. Diagnostics name the authored field and show a safe corrected table without echoing an endpoint, credential value, record value, or physical path.
Event fields are denied by default. Each [[observability.field_policy]]
entry selects exactly one dotted event/field pair and one allow, hash, or
replace action. replacement is required only for replace, and duplicate
rules for the same pair are invalid. A replacement is written verbatim into
the exported record, so it is held to the same shape as the other authored
strings here: non-empty, bounded, with no surrounding whitespace and no
embedded control character.
Field policy governs record fields — values a Transform selected out of the
data being processed. It does not govern run correlation. Every exported log
record carries clinker.execution_id, clinker.batch_id, and
clinker.pipeline_name unconditionally, with no [[observability.field_policy]]
entry required for them.
The reason is worth stating plainly, because the opposite behaviour was a
defect rather than a policy: those three values are identifiers Clinker
generates for the run itself, not data read from a source. A record-field
privacy policy has nothing to decide about them. Subjecting them to it meant a
workspace that declared an event without also writing three correlation rules
exported log records with no correlation at all — telemetry that could not be
joined to the machine stream’s execution_id, to the lineage events, or to
another run of the same pipeline — and it counted all three as privacy denials
on every event, inflating the arena’s denied-field accounting by three per
event. Correlation is now carried outside the policy entirely, so neither
happens.
A [[observability.field_policy]] rule whose field happens to be named
execution_id, batch_id, or pipeline_name still parses. It governs a
record column of that name, if a directive requests one; it has no effect on
the correlation attributes above.
Lineage identity
identity_mode = "external" is the default. Every externally emitted source
or Sink node needs exactly one binding: either canonical_datasource, or
the complete catalog_namespace/catalog_name pair. Missing, duplicate,
partial, and mixed bindings fail validation; Clinker does not synthesize an
external identity from a working directory, worker path, temporary root, URL,
attempt identifier, or path hash.
The stable collection identity remains the dataset namespace/name. When the runtime has explicitly authorized a concrete logical partition or location, lineage represents it with the standard input/output subset facet rather than changing that collection name. Explicitly authorized aliases use the standard symlinks facet. The current resolved workspace config has no author-facing subset or symlink fields, so Clinker does not infer either fact from local paths, attempt directories, hashes, or process context.
The only path-derived compatibility mode is the exact, explicit value below. It accepts no external dataset bindings and is for labeled local diagnostics, not external delivery:
[observability.lineage]
identity_mode = "local_diagnostic_paths"
Operational recommendations
- Always enable metrics in production. The overhead is negligible (one small JSON write at the end of each run).
- Run
metrics collect --delete-after-collecton a schedule (e.g., hourly) to prevent spool directory growth. - Use
--batch-idwith meaningful identifiers to correlate metrics across retries and environments. - Alert on
records_dlq > 0to catch data quality regressions early. - Track
peak_rss_bytestrends to anticipate when memory limits need adjustment.
Exit Codes & Error Diagnosis
Clinker uses structured exit codes to communicate the outcome of a pipeline run. These codes are designed for integration with schedulers, cron, CI systems, and monitoring tools.
Exit code reference
| Code | Meaning | Description |
|---|---|---|
| 0 | Success | Pipeline completed successfully, or an attempt operation completed without cleanup debt. A purge preview that safely selects nothing is also successful. |
| 1 | Configuration or argument error | Invalid YAML, CXL syntax error, type mismatch, DAG wiring problem, invalid attempt selector, invalid continuation, a --lineage export rejected by the configured observability caps, a --lineage destination that will refuse every identical retry — one the process may not write, that does not exist, that is read-only, or that is not a file. Fix the pipeline configuration or command arguments. |
| 2 | Partial success | Pipeline ran to completion, but some records were routed to the dead-letter queue. Check the DLQ file. |
| 3 | Evaluation error | CXL runtime error during record processing (e.g., division by zero, type coercion failure). |
| 4 | Infrastructure or retained cleanup debt | File/format failure, disk full, a --lineage export that failed in a way a retry may resolve — a reader that went away, a write that timed out, a volume that was out of space, or a flush that exceeded its deadline — an environment that refused the SIGINT/SIGTERM handler the run requires, in which case the run stops before reading or writing anything and the environment is what must change, or an attempt operation stopped with bounded, ambiguous, live, or otherwise retryable cleanup debt. This status never means completed-with-DLQ. |
| 130 | Cancelled | Graceful SIGINT or SIGTERM cancellation won before publication, or a required --machine lifecycle record could not be written before publication. Final paths for the current attempt remain unchanged. |
Known issue: a command-line usage error (an unknown flag or a missing argument) currently exits 2, the partial-success code, instead of 1, except under
clinker attempts(#1372). A scheduler that treats 2 as “completed with dead letters” cannot tell the two apart yet; check stderr for a usage message.
For an ordinary standalone run, these statuses are the complete process
result. A --machine ndjson-v1 consumer must additionally require exactly one
supported terminal event whose result and embedded exit, where present, match
the actual child status and current-attempt artifact evidence. EOF, malformed
or unsupported output, a duplicate or missing terminal, forced termination,
or any mismatch is an incomplete attempt even if an older final already
exists. See Running Clinker Directly or Under a Supervisor.
Understanding exit code 2
Exit code 2 is not a crash. It means:
- The pipeline started and ran to completion.
- All viable records were processed and written to output files.
- Some records could not be processed and were diverted to the dead-letter queue.
Your scheduler should distinguish completed-with-DLQ from an aborted run. Whether downstream work may continue is a data-quality policy decision. The DLQ contains the problematic records and their rejection diagnostics.
To bound the tolerated type-error ratio, configure the YAML policy:
error_handling:
strategy: continue
type_error_threshold: 0.05
dlq:
path: ./output/errors.csv
format: csv
The threshold is a ratio from 0 to 1, not a record count. Exceeding the configured
type-error ratio produces exit 3. It does not count every possible DLQ reason.
The retired --error-threshold flag is rejected; see
error handling for the policy’s population.
Diagnosing failures
Exit code 1: Configuration error
The error message includes a span-annotated diagnostic pointing to the exact location of the problem:
Error: CXL type error in node 'transform_1'
--> pipeline.yaml:25:15
|
25 | emit total = amount + name
| ^^^^^^^^^^^^^ cannot add Int and String
Action: Fix the YAML or CXL expression indicated in the diagnostic, then re-run with --dry-run to confirm the fix.
Exit code 2: Partial success (DLQ entries)
Check the DLQ file for details:
# The DLQ path is shown in the run output and in metrics
cat output/errors.csv
Common causes:
- Null values in fields that a CXL expression does not handle
- Data that does not match the declared schema (e.g., non-numeric value in an integer column)
- Coercion failures between types
Action: Review the DLQ records, fix the data or add null handling to CXL expressions, and re-run.
Exit code 3: Evaluation error
A CXL expression failed at runtime. The error message includes the failing expression and the record that triggered it:
Error: division by zero in node 'compute_ratio'
expression: emit ratio = total / count
record: {total: 500, count: 0}
Action: Add guard conditions to the CXL expression:
emit ratio = if count == 0 then 0 else total / count
Exit code 4: Infrastructure or retained cleanup debt
File system or format errors:
Error: file not found: ./data/customers.csv
--> pipeline.yaml:8:12
Common causes:
- Input file does not exist or path is wrong
- Permission denied on input or output directories
- Output file already exists (use
--forceto overwrite) - Disk full during output writing
- Input file format does not match the declared type (e.g., invalid CSV)
- A retained-attempt query reached its entry, byte, or monotonic-time bound
- Cleanup kept an attempt because ownership, liveness, manifest, clock, or filesystem evidence was ambiguous
For clinker attempts, exit 4 includes path-free cleanup debt plus E371
(unsafe or invalid attempt refused) or E372 (cleanup incomplete or budget
exhausted) when the debt belongs to a logical execution. Follow the emitted
retry advice and pasteable clinker attempts inspect <workspace-relative.yaml> --execution-id <id> command. If a continuation is present, paste its exact
resume command to continue the bounded page. Do not treat status 4 as authority
to delete a directory manually.
Action: Fix file paths, permissions, or disk space for run failures. For attempt operations, inspect the named logical execution and resolve the stated debt before retrying or executing another bounded purge.
Exit code 130: Cancelled
Exit 130 means the attempt stopped before publication and the current attempt’s final paths are unchanged. Two things produce it:
- A SIGINT or SIGTERM that won the cancellation gate before the first final rename.
- Under
--machine ndjson-v1, a required lifecycle record that could not be written. The run refuses to publish an outcome it cannot report, so a broken control pipe stops the attempt rather than promoting silently.
A discardable machine record — a periodic progress observation — is not
in that set. Losing one is reported on stderr as machine progress channel failed and the run continues to its real outcome; it never converts a
completed run into a cancellation or discards computed output. A supervisor
should therefore read 130 as “nothing was published”, never as “an advisory
record went missing”.
Cancellation is also recorded as a cancellation everywhere else it is
reported: the OpenLineage terminal is ABORT and the --machine terminal is
cancelled, regardless of which source observed the signal first. A source
that notices cancellation while a request is in flight — a REST page read, for
instance — produces the same terminals as a file source that drains and stops
at a chunk boundary.
Action: Re-run the attempt from the beginning of the input with the same
--batch-id. There is no resume or checkpoint state to recover; see
Retry and identity boundaries.
If the run was not cancelled by an operator, check stderr for a machine
control-channel write failure and confirm the consumer is draining stdout.
Plan-time diagnostic codes
The process exit codes above tell a scheduler whether the run
succeeded. The E### codes below appear inside the structured
Error: messages a configuration error (exit code 1) prints, and
identify the specific compile-time check that rejected the
pipeline. The codes below cover the event-time watermark and
time-windowed aggregate surface
(issue #61);
related code sets live in
Pipeline Variables,
Channels, and
Correlation Keys.
| Code | Trigger | Remediation |
|---|---|---|
| E154 | A source declares watermark.column: <col> but <col> is not present in that source’s schema: block. | Add the column to schema:, or remove the watermark: block. |
| E155 | A source declares watermark.column: <col> and the column exists, but its declared CXL type is not date_time or date. | Change the column’s type: to date_time or date, or point watermark.column at a column that already has one of those types. |
| E156 | An aggregate declares time_window: but at least one upstream-reachable source does not declare watermark.column. | Add watermark: { column: <event-time-column> } to each listed source, or remove time_window: from the aggregate. Without a watermark on every upstream source, min_across_sources never advances past None and the window can never close. |
| E157 | A source declares an external schema: file (schema: path.schema.yaml) that could not be read or parsed as a SourceSchema. | Fix the file path or its contents. A schema file is a bare column list or a multi-record discriminator:/records: map — it may not itself point at another schema file. |
| E158 | A source column’s declared type is (or wraps) the inference-only numeric union. | Declare a concrete int or float. numeric is int | float resolved during type unification and never carries into a compiled source schema. |
| E159 | A source pairs a generated schema with a non-EDI format. | generated (engine-synthesized positional columns) is valid only for the EDI-family formats (edifact, x12, hl7, swift). Declare an explicit column list for any other format. |
See Source Nodes → Watermarks and Aggregate Nodes → Time-windowed aggregates for the field semantics each code is enforcing.
DLQ category: LateRecord
When a time-windowed aggregate sees a record whose event time falls
inside an already-closed window
(window_end + allowed_lateness < min_across_sources), the engine
routes the record to the DLQ instead of attempting to fold it into
a finalized accumulator. Mirrors Flink’s
sideOutputLateData
and Spark Structured Streaming’s late-data drop.
The DLQ row carries:
_cxl_dlq_error_category=late_record_cxl_dlq_stage=time_window:<aggregate-name>_cxl_dlq_error_detail— the closed window’s[start, end)bounds as i64 nanoseconds since the Unix epoch
Tune watermark.delay (source-side, applies before any aggregate)
or allowed_lateness (operator-side, applies per aggregate) to
absorb expected out-of-order tails before they reach this path.
Scheduler integration
For running Clinker under a workflow orchestrator (Temporal, Airflow, Dagster) — mapping these exit codes onto a retry policy, plus the cancellation and output-atomicity guarantees — see Running Under a Workflow Orchestrator.
Cron script
#!/bin/bash
set -euo pipefail
PIPELINE=/opt/clinker/pipelines/daily_etl.yaml
METRICS_DIR=/var/spool/clinker/
# `set -e` would end the script as soon as clinker exits non-zero, before the
# `case` below runs. `|| EXIT=$?` captures the status without stopping.
EXIT=0
clinker run "$PIPELINE" \
--memory-limit 512M \
--log-level warn \
--metrics-spool-dir "$METRICS_DIR" \
--force || EXIT=$?
case $EXIT in
0)
echo "$(date): Success" >> /var/log/clinker/daily_etl.log
;;
2)
echo "$(date): Warning - DLQ entries produced" >> /var/log/clinker/daily_etl.log
mail -s "Clinker ETL Warning: DLQ entries" ops@company.com < /dev/null
;;
*)
echo "$(date): FAILURE (exit code $EXIT)" >> /var/log/clinker/daily_etl.log
mail -s "Clinker ETL FAILURE (exit $EXIT)" ops@company.com < /dev/null
;;
esac
exit $EXIT
CI pipeline (GitHub Actions)
- name: Run ETL pipeline
run: clinker run pipeline.yaml --dry-run
# Exit code 1 fails the build on config errors
- name: Smoke test with real data
run: clinker run pipeline.yaml --dry-run -n 100
# Catches runtime evaluation errors
Systemd
Systemd Type=oneshot services interpret non-zero exit codes as failures. To allow exit code 2 (partial success) without triggering service failure:
[Service]
Type=oneshot
SuccessExitStatus=2
ExecStart=/opt/clinker/bin/clinker run /opt/clinker/pipelines/daily_etl.yaml --force
Known issue: a command-line usage error, such as a mistyped flag or a missing argument, currently exits 2 instead of 1 (except under
clinker attempts), soSuccessExitStatus=2would also count a typo inExecStartas a success (#1372). After editing the unit, run itsExecStartcommand once by hand and check that it does not print a usage error.
Production Deployment
Clinker is a single statically-linked binary with no runtime dependencies. Deployment is straightforward: copy the binary to the server.
Installation
# Copy the binary
scp target/release/clinker user@server:/opt/clinker/bin/
# Verify it runs
ssh user@server /opt/clinker/bin/clinker --version
No JVM, no Python, no container runtime required.
Recommended directory structure
/opt/clinker/
bin/
clinker # The binary
pipelines/
daily_etl.yaml # Pipeline configs
weekly_report.yaml
data/ # Input data (or symlinks to data locations)
output/ # Output files
rules/ # CXL module files (for use statements)
metrics/ # Metrics spool directory
Create a dedicated user:
sudo useradd --system --home-dir /opt/clinker --shell /usr/sbin/nologin clinker
sudo chown -R clinker:clinker /opt/clinker
Systemd service
For scheduled one-shot execution:
[Unit]
Description=Clinker ETL - Daily Customer Processing
After=network.target
[Service]
Type=oneshot
ExecStart=/opt/clinker/bin/clinker run /opt/clinker/pipelines/daily_etl.yaml \
--memory-limit 512M \
--log-level warn \
--metrics-spool-dir /var/spool/clinker/ \
--force
WorkingDirectory=/opt/clinker
User=clinker
Group=clinker
SuccessExitStatus=2
# Resource limits
MemoryMax=1G
CPUQuota=200%
# Logging
StandardOutput=journal
StandardError=journal
SyslogIdentifier=clinker-daily
[Install]
WantedBy=multi-user.target
Pair with a systemd timer for scheduling:
[Unit]
Description=Run Clinker daily ETL at 2 AM
[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
[Install]
WantedBy=timers.target
sudo systemctl enable --now clinker-daily.timer
Note: SuccessExitStatus=2 tells systemd that exit code 2 (partial success with DLQ entries) is not a service failure. See Exit Codes for the full reference.
Known issue: a command-line usage error, such as a mistyped flag or a missing argument, currently exits 2 instead of 1 (except under
clinker attempts), soSuccessExitStatus=2would also count a typo inExecStartas a success (#1372). After editing the unit, run itsExecStartcommand once by hand and check that it does not print a usage error.
Cron scheduling
# Run daily at 2 AM, log to syslog
0 2 * * * /opt/clinker/bin/clinker run \
/opt/clinker/pipelines/daily_etl.yaml \
--log-level warn --force \
2>&1 | logger -t clinker
# Collect metrics hourly
0 * * * * /opt/clinker/bin/clinker metrics collect \
--spool-dir /var/spool/clinker/ \
--output-file /var/log/clinker/metrics.ndjson \
--delete-after-collect
Environment-based configuration
Use the CLINKER_ENV variable or --env flag to activate environment-specific overrides:
# Production
CLINKER_ENV=production clinker run pipeline.yaml
# Staging
CLINKER_ENV=staging clinker run pipeline.yaml
Combined with channel overrides in the pipeline YAML, this allows a single pipeline definition to target different file paths, connection strings, or thresholds per environment.
Logging
Log levels for production
| Level | Use case |
|---|---|
warn | Recommended for production cron jobs. Prints warnings and errors only. |
info | Default. Includes progress messages. Useful during initial deployment. |
error | Minimal output. Only prints when something fails. |
debug | Troubleshooting. Generates significant output. |
trace | Development only. Extremely verbose. |
Directing logs
To syslog via logger:
clinker run pipeline.yaml --log-level warn 2>&1 | logger -t clinker
To a log file:
clinker run pipeline.yaml --log-level warn 2>> /var/log/clinker/etl.log
Systemd journal captures stdout and stderr automatically when running as a service.
DLQ monitoring
When a pipeline exits with code 2, records that could not be processed are written to the dead-letter queue file. Set up a daily check:
#!/bin/bash
# Check for DLQ files produced today
DLQ_DIR=/opt/clinker/output/
DLQ_FILES=$(find "$DLQ_DIR" -name "*_errors.csv" -mtime 0 -size +0c)
if [ -n "$DLQ_FILES" ]; then
echo "DLQ entries found:" | mail -s "Clinker DLQ Alert" ops@company.com <<EOF
The following DLQ files were produced today:
$DLQ_FILES
Review the files and address data quality issues.
EOF
fi
Batch ID for tracing
Use --batch-id with a meaningful, consistent naming scheme:
# Date-based
clinker run pipeline.yaml --batch-id "daily-$(date +%Y-%m-%d)"
# Include environment
clinker run pipeline.yaml --batch-id "prod-daily-$(date +%Y-%m-%d)"
The batch ID appears in metrics output and log lines, making it easy to correlate a specific run across logs, metrics, and DLQ files. On retries, use a different batch ID (e.g., append -retry-1) to distinguish attempts.
Upgrades
To upgrade Clinker:
- Validate the new version against your pipelines:
/opt/clinker/bin/clinker-new run pipeline.yaml --dry-run - Replace the binary:
cp clinker-new /opt/clinker/bin/clinker - Verify:
/opt/clinker/bin/clinker --version
There is no configuration migration. Pipeline YAML files are forward-compatible within the same major version.
Running Clinker Directly or Under a Supervisor
Clinker is a finite synchronous CLI. The ordinary command is the primary path and needs no worker, service, scheduler, or orchestration component:
clinker run pipeline.yaml
An external scheduler or workflow runner can opt into a versioned child-process protocol when it needs machine-readable lifecycle evidence:
clinker run pipeline.yaml --machine ndjson-v1 --batch-id logical-batch
Machine mode does not change who owns the run. Clinker still validates, executes, cancels, and publishes one finite attempt. The parent process owns scheduling, retry and backoff, deadlines, heartbeats, process lifetime, and direct-child reaping. Clinker embeds no orchestrator SDK, worker, daemon, or service runtime.
Stream ownership and compatibility
In machine mode, stdout contains only compact UTF-8 NDJSON and is flushed after each event. Human diagnostics and tracing use non-ANSI stderr. The parent must drain stdout and stderr concurrently; reading one pipe to completion before the other can deadlock when an OS pipe fills.
Every event carries these fields:
| Field | Meaning |
|---|---|
protocol | Always clinker.run. |
schema | Protocol major version. ndjson-v1 emits 1. |
event | Lifecycle kind such as started, plan_resolved, progress, publication_artifacts, completed, failed, or cancelled. |
seq | Zero-based sequence, increasing by exactly one within this process. |
batch_id | Caller-supplied logical-batch correlation retained across retries. |
execution_id | Fresh non-overridable UUIDv7 generated for this process. |
plan_identity | pending, resolved with the semantic plan fingerprint, or unavailable after admission failure. |
The resolved fingerprint covers the compiled topology, schemas, composition
and CXL dependencies, winning channel/group config values, and all four
runtime-variable scopes. A composition contributes the body each call site
actually binds, so two channels that patch one shared body differently are
two plans. It excludes deployment-only file locations and layer source
formatting, so relocating equivalent inputs does not invalidate a pinned plan
while a value that can change execution does. Where rejected records are
written is a location and is excluded; whether they are written is not — a
pipeline with no error_handling.dlq.path and no per-source override
discards them, and reads as a different plan from one that keeps them.
The version field carries the fingerprint schema, currently 3. A digest
is comparable only with another digest of the same version: when the schema
changes, the same pipeline yields a different digest, and the version is how
a consumer holding a pinned value tells that apart from a changed plan.
Progress events add a bounded logical phase, kind, elapsed time, counts, and
truncation flags. They never contain records, secrets, source URLs, or
physical paths. Periodic records follow Clinker’s own clock rather than
internal engine activity, so a run inside one long operation keeps producing
them; they stay advisory and bounded and never replace the parent’s own
heartbeat. Progress records below specifies the counts
and what a consumer may conclude from them. Failed terminals add a stable failure code,
broad category, sanitized message, and retry_with_backoff, do_not_retry, or
policy_required advice. Before a publication-aware terminal, bounded
publication_artifacts records carry the path-free inventory in ordered
chunks. Artifact entries contain only artifact_id, kind, and state.
The terminal carries the attempt’s completeness, cleanup-debt count, total
artifact count, and counts by artifact state. Every NDJSON record, including a
maximum-cardinality inventory chunk and its terminal summary, is at most 16
KiB.
One invocation that reaches a terminal without running an attempt is
supported: the plan-only --lineage <FILE> export, which preflights the
identity policy, writes its document, and returns before any data is read. It
shares this stream’s execution_id and batch_id, so the exported document is
correlatable with the invocation that produced it, and it closes with
completed / success / exit 0 carrying an explicit empty publication —
zero artifacts, every state count zero, no cleanup debt. An absent inventory
would be read on that row as publication complete for a run that published
nothing; an empty one says the same thing the reconciliation table already has
vocabulary for. Its document is written and flushed before that terminal is
attempted, so a terminal the pipe refuses there is reconciled exactly as a
published run’s is: exit 4 with
infrastructure.delivery.unreportable_outcome and policy_required, never
retry_with_backoff, because the export already exists on disk. Modes that
write their own document to standard output —
--explain, --dry-run, -n, --lineage -, --lineage-events - — are
refused at admission with exit 1 before any record is written.
A required lifecycle record that cannot be delivered while nothing has yet been
read, written, or staged ends the run at exit 130 with a cancelled
terminal, and the attempt’s final paths are unchanged. That applies to
started, to the planning transition, and to plan_resolved, on a run and
on the plan-only --lineage <FILE> export alike: a supervisor that stops
reading during plan compile learns the same thing about the same condition
whichever it asked for.
After plan resolution, Clinker starts the required machine-progress worker
before source discovery, staging, attempt creation, sink writes, or lifecycle
START. If that worker cannot be created, the stream ends with exactly one
infrastructure.runtime.transient failed terminal and exit 4; no run effect
has started, and a supervisor may retry with backoff.
A consumer must reject an unsupported schema major. Within schema 1 it may
ignore additive fields and unknown nonterminal event kinds, but those additions
carry no completion or failure meaning. Missing required fields, malformed
UTF-8 or JSON, non-monotonic sequence, identity changes, duplicate terminals,
or EOF without a terminal make the attempt incomplete.
Records reach the stream whole or not at all, and a parent that reads slowly
does not change what the stream says. A record the pipe refuses is not written,
takes no seq with it, and is not what the next record is numbered after, so a
lost advisory observation leaves the numbering dense rather than shifting every
record after it. A record the pipe accepts only part of is completed by the
next write of this stream — never restarted — so a momentarily full pipe cannot
produce two copies of one record or two terminals. What was already delivered
is likewise not repeated: a terminal retried after a refusal sends only the
inventory chunks the pipe has not taken, so each chunk index appears exactly
once ahead of the terminal that counts them.
Terminal, exit, and artifact reconciliation
Neither a terminal event nor a process status is sufficient alone. Accept a controlled outcome only when the supported terminal family, its embedded exit where present, the actual child status, and current-attempt artifact evidence agree:
| Terminal evidence | Required process status | Artifact interpretation | Adapter result |
|---|---|---|---|
completed, result success, exit 0 | 0 | Publication is complete; every reported artifact is individually complete. | Success. |
completed, result completed_with_dlq, exit 2 | 2 | Publication is complete and includes the reported complete DLQ artifact. | Completed under the caller’s data-quality policy. |
failed, embedded exit 1, 3, or 4 | The same exit | Use the exact reported publication state. When publication is absent the state could not be reported; infer nothing about the visible set from its absence. A reported visible subset, if any, consists only of individually complete artifacts. | Failure; apply the typed retry advice and caller policy. |
cancelled | 130 | Graceful cancellation won before publication; final paths for this attempt remain unchanged. | Cancellation. |
| No terminal, malformed stream, unsupported major, duplicate terminal, forced termination, or mismatched exit | Any | Do not infer current-attempt success from a pre-existing final or a visible complete subset. | Incomplete attempt. |
Exit 4 is intentionally broad; the typed failed terminal distinguishes retry
advice without requiring the parent to parse rendered diagnostics. EOF alone
never proves success. A control-pipe failure can prevent terminal delivery, so
even an otherwise plausible exit remains incomplete without matching terminal
evidence.
That is also why a failed terminal delivery is never reported as a transient
runtime fault. When a run has published and only the terminal saying so cannot
be written, a retried terminal that does get through reports exit 4 with
infrastructure.delivery.unreportable_outcome and policy_required — never
retry_with_backoff — and carries the publication the refused terminal
carried. The failure is on the reporting channel, not in execution: the finals
are visible, the lineage and OTLP terminals for the same run recorded its
completion, and re-running the batch would duplicate published data. A
supervisor reconciles it as a failure whose artifact evidence is complete and
whose repetition is a policy decision, not an automatic one. When neither
terminal reaches the stream the attempt is incomplete by the table above, which
is the same reconciliation it always was.
The terminal family follows the exit code, and the completed family covers
only exit 0 and exit 2. Any other non-cancellation exit is written as a
failed terminal carrying the run’s own failure classification, including
after an earlier failed emission that could not be encoded and so left the
single terminal slot free. A non-zero exit is never restated as result
success; if no terminal can be encoded at all, the stream ends without one
and the attempt is incomplete by the table above.
Publication is atomic per artifact, not for the artifact set. A failure or
uncontrolled stop during multi-artifact publication can leave an exact subset
of newly promoted, individually complete finals visible. Consumers that need a
complete set must wait for reconciled success and verify the expected artifact
inventory. Reassemble artifact chunks only when every chunk index from zero to
chunk_count - 1 is present in sequence before the terminal and the assembled
count matches the terminal’s artifact_count and state_counts. Never treat a
partial terminal stream as set-wide success.
Language-neutral adapter loop
- Launch one child with a stable caller-owned
batch_id, piped stdout and stderr, and an overall attempt deadline. Retain only bounded sanitized tails for diagnostics. - Drain stdout and stderr concurrently. Parse stdout incrementally as UTF-8 NDJSON and validate the first record’s protocol major, identities, and sequence.
- Heartbeat the external scheduler on an independent cadence. Report only the latest sanitized identity, sequence, logical phase/counts, and snapshot age; do not wait for or translate a Clinker progress event into a heartbeat.
- Enforce a total Start-to-Close-style deadline for the whole process. A no-progress timeout is an additional explicit deployment policy, not a substitute for the overall deadline.
- On cancellation or deadline, deliver the platform’s real graceful signal to the direct child, keep both pipes draining, and start a separately bounded cancellation grace period. If grace expires, force termination exactly once. A forced stop is incomplete, not cancelled or successful.
- Always wait for and reap the direct child before joining both drain tasks. Accept an outcome only through the terminal, process-status, and artifact reconciliation table above.
- On retry, launch a completely fresh process from the beginning of the input
with the same
batch_id. Require a newexecution_id; retain no Clinker progress event as checkpoint or resume state.
The heartbeat interval must be below the scheduler’s heartbeat timeout. The overall attempt deadline must cover ordinary execution and publication; the grace period is a separate bounded interval for cooperative cancellation before forced termination.
Progress records
A progress event carries a progress object and a truncation object:
{"event":"progress","seq":7,
"progress":{"phase":"executing","kind":"periodic","elapsed_ms":1031,
"records_read":200000,
"bytes_read":27000835,"bytes_total":81777788,
"files_done":0,"files_total":1},
"truncation":{"detail":false,"events":false}}
kind is transition for a lifecycle edge (planning, executing,
finalizing, publishing) and periodic for an advisory observation inside
a phase.
The counts
records_read is the number of source records read so far, across every
source. It never decreases within a run, and the last one a run emits is the
number of records that run read. It has no companion total, and no total
will be added. The terminal carries no record count of its own: a progress
record is the only place this stream reports one, so there is nothing to
reconcile it against and no disagreement to arbitrate. A source is read as a stream, so its
record count is not established until its last record has been read: any
“records remaining” Clinker could publish mid-run would be a guess presented
as a measurement. A supervisor answers is this run moving by comparing
records_read between two events, which is the question the record is here
to answer. When will it finish is not a question this stream answers.
bytes_read and bytes_total are the denominator to prefer. bytes_total
is the summed on-disk size of every input, read from file metadata before
anything is opened, so it is measured rather than estimated. bytes_read
advances within a file, not only at file boundaries, which is what makes
it useful on the common single-large-file run where every other count sits
still until the end.
bytes_total is never 0. A run with no bytes to read publishes null
instead, because a zero total is not a denominator — dividing by it yields
NaN, not 0%, and a run with nothing to read is not a run that is nought
per cent through anything.
bytes_total is null in three cases. First, when any source cannot supply a
size — a network source, or a path whose metadata will not read. Second, when
any source is read more than once: json and xml sources re-open their
input to pre-scan the envelope before streaming the body, so their bytes cross
the counter twice and the count becomes IO performed rather than input
consumed. Rather than publish a total the count will overrun, Clinker
withdraws it. Third, as above, when the run’s inputs are empty. bytes_read
is still emitted in every one of those cases and still rises monotonically —
it remains a usable liveness signal, just not a fraction.
bytes_read counts file-backed input only. A source with no bytes on
disk — a network source, for instance — contributes nothing to it, so a run
reading only from such a source reports bytes_read: 0 for its whole life
while records flow normally. A null bytes_total is the signal that this
may be so: when bytes_total is null, judge liveness from records_read,
not from bytes. Reading a still 0 byte count as a stalled run is the one
mistake this field invites.
Like files_total, bytes_total is null on the earliest records of a run,
before source discovery has established it. It is written once and never
changes afterwards — including never reverting to null — so a consumer may
cache it on first sight, and its arrival mid-stream is normal rather than an
identity change.
files_done and files_total are a second, independent denominator, useful
where the byte one is withdrawn. A source’s file set is enumerated at startup,
so the count is known rather than estimated. files_total is null when
any source reads from something other than an enumerated file set.
Neither denominator is a substitute for the other and they are never combined. A denominator covering only part of a run’s work is withdrawn rather than published, because nothing on the wire would say which part it covered.
files_total is also null on the earliest records of a run, before source
discovery has completed. It is written once and never changes afterwards, so
a consumer may cache it on first sight; it becoming non-null mid-stream is
normal and is not an identity change.
Two things a consumer must not do with these counts:
- Do not treat either ratio as a completion percentage. Both measure input consumed, not work finished. They reach 100% while sort merges, aggregate finalization, and output publication are still running, which is exactly the shape of a progress bar that sits at 100% and appears to hang.
- Do not divide by a
nulltotal. Absence means unknown for this run, not zero and not an error. Clinker publishes no percentage of its own for this reason: a ratio against an absent denominator is the one value that turns a missing number into a wrong one.
The record cap and the cadence floor
Periodic records are bounded twice. At most one is emitted per second, and at
most 128 per run. After the 128th, Clinker emits exactly one further
record with truncation.events set to true, and then stops emitting
periodic records for the rest of the run. Transitions are never capped, so
finalizing and publishing still arrive, as does the terminal.
Because the two bounds compose, any run longer than roughly two minutes stops producing periodic progress well before it ends. This is deliberate: it is what keeps the stream bounded regardless of how long a run takes.
truncation.events is the in-band signal for it. A consumer that sees it
should conclude that the periodic stream has ended and the run is
continuing normally. It must not conclude that the run has stalled,
and must not treat the silence that follows as evidence of a hung process.
A supervisor that kills or retries a run on progress silence will kill
healthy long runs, and for a pipeline that writes output, a retry means the
work is done twice.
Liveness is the parent’s own responsibility, on the parent’s own clock — see step 3 of the adapter loop. The run’s real outcome is the terminal record, which always arrives.
Cancellation and process trees
Clinker installs handlers for SIGINT and SIGTERM. If graceful cancellation wins
the atomic gate before publication, it exits 130 and leaves finals unchanged.
If publication wins first, later signals do not relabel or erase already
complete visible artifacts; the bounded promotion finishes with completed or
failed truth.
A cancellation is reported as a cancellation on every surface that reports it,
and which source noticed the signal does not change that. A file source drains
to a chunk boundary and stops; a source that observes cancellation while a
request is already in flight — a REST page read, for instance — unwinds from
inside that read. Both produce exit 130, a cancelled machine terminal, and
an OpenLineage ABORT. Cancellation is never reported as an engine failure
class, so alerting keyed on the lineage or OTLP terminal does not page for an
operator-initiated stop.
Exit 130 also covers one non-signal case: a required machine lifecycle
record that could not be written before publication. Clinker refuses to
publish an outcome it cannot report, so a broken control pipe stops the attempt
with finals unchanged. Discardable records are excluded from that rule — a lost
periodic progress observation is reported on stderr and the run continues to
its real outcome, and never converts a completed run into a cancellation.
Signal-handler installation is admission-critical. If installation fails,
Clinker exits with infrastructure status 4 before opening the machine
protocol stream or touching sources, staging, attempts, outputs, or lineage.
The direct-child contract is exercised on Linux with a real SIGTERM rather
than a closed control pipe or an in-process cancellation shortcut. The
cooperative case verifies exit 130, a matching cancelled terminal,
unchanged final paths, and child reaping while both bounded drains remain live.
The uncooperative case verifies that the grace interval is independent of the
overall attempt deadline, force happens once only after that interval, the
direct child is reaped before drain joining, and the attempt remains
incomplete. Process groups and descendant ownership are deliberately outside
that direct-child proof.
The adapter owns the platform termination domain. On POSIX, an adapter that creates descendants may place them in a process group and signal that group. On Windows, such an adapter may use a Job Object, including kill-on-close policy. These are adapter responsibilities: Clinker does not ship a process-group manager or Job Object owner. The parent must still reap the direct Clinker child after graceful or forced termination.
Retry and identity boundaries
batch_id is correlation only. Each invocation generates a fresh
execution_id and starts input from the beginning. Clinker has no cross-attempt
checkpoint, resume cursor, deduplication state, distributed transaction, or
exactly-once guarantee. Safe application retry depends on stable input and the
destination’s chosen if_exists policy.
Cancellation and forced termination preserve failed-attempt evidence under
the configured storage retention policy.
An incomplete attempt keeps its staging directory and manifest without changing
pre-existing finals; the manifest’s eligible_after timestamp, 24 hours by
default, controls when ordinary cleanup may reclaim it. An immediate retry does
not resume or mutate that directory: it starts a new attempt with a fresh
execution_id and its own staging state.
Progress is advisory liveness evidence, not a durable checkpoint or external
heartbeat. Machine events are control evidence, not a secret-bearing event bus
or compliance log. Ordinary users remain free to run the standalone command
without --machine or any supervisor.
See also
CSV-to-CSV Transform
This recipe reads employee data from a CSV file, computes salary tiers using CXL expressions, and writes the enriched result to a new CSV file.
Input data
employees.csv:
id,name,department,salary
1,Alice Chen,Engineering,95000
2,Bob Martinez,Marketing,62000
3,Carol Johnson,Engineering,88000
4,Dave Williams,Sales,71000
5,Eva Brown,Marketing,58000
6,Frank Lee,Engineering,102000
Pipeline
salary_tiers.yaml:
pipeline:
name: salary_tiers
nodes:
- type: source
name: employees
config:
name: employees
type: csv
path: "./employees.csv"
schema:
- { name: id, type: int }
- { name: name, type: string }
- { name: department, type: string }
- { name: salary, type: int }
- type: transform
name: classify
input: employees
config:
cxl: |
emit id = id
emit name = name
emit department = department
emit salary = salary
emit level = if salary >= 90000 then "senior" else "junior"
emit salary_band = match {
salary >= 100000 => "100k+",
salary >= 90000 => "90-100k",
salary >= 70000 => "70-90k",
_ => "under 70k"
}
- type: sink
name: report
input: classify
config:
name: report
type: csv
path: "./output/salary_report.csv"
error_handling:
strategy: fail_fast
Run it
# Validate first
clinker run salary_tiers.yaml --dry-run
# Preview output
clinker run salary_tiers.yaml --dry-run -n 3
# Full run
mkdir -p output
clinker run salary_tiers.yaml
At revision 3b343a4e, bounded preview of this example can fail with an internal
node-buffer cleanup error; the ordinary run produces the output below. See
the preview limitation.
Expected output
output/salary_report.csv:
id,name,department,salary,level,salary_band
1,Alice Chen,Engineering,95000,senior,90-100k
2,Bob Martinez,Marketing,62000,junior,under 70k
3,Carol Johnson,Engineering,88000,junior,70-90k
4,Dave Williams,Sales,71000,junior,70-90k
5,Eva Brown,Marketing,58000,junior,under 70k
6,Frank Lee,Engineering,102000,senior,100k+
Key points
Schema declaration. The source node declares the schema explicitly with typed columns. This enables compile-time type checking of CXL expressions – if you write salary + name, the type checker catches the error before any data is read.
Emit statements. A Transform adds or replaces the fields it emits. Other input
columns can pass through. To select the public output columns deliberately, use
the Sink’s mapping/exclude settings and include_unmapped: false; an emit list
alone is not a data-leakage boundary.
Match expressions. The match block evaluates conditions top to bottom and returns the value of the first matching arm. The _ wildcard is the default case and must appear last.
Error handling. The fail_fast strategy aborts the pipeline on the first record error. For production pipelines processing dirty data, consider continue instead, which routes the failing record to the dead-letter queue and keeps going – see Error Handling & DLQ.
Variations
Filtering records
Add a filter statement to exclude records:
- type: transform
name: classify
input: employees
config:
cxl: |
filter salary >= 60000
emit id = id
emit name = name
emit salary = salary
Records where salary < 60000 are dropped silently – they do not appear in the output or the DLQ.
Computed columns with type conversion
cxl: |
emit id = id
emit name = name
emit monthly_salary = (salary.to_float() / 12.0).round_to(2)
emit salary_display = "$".concat(salary.to_string())
The .to_float() conversion is required because salary is declared as int and division by a float literal requires matching types.
Multi-Input Combine
This recipe enriches order records with product metadata from a separate catalog stream using a combine node. Combine is a first-class N-ary operator: every input is declared up front, and the where expression uses qualified field references (orders.product_id, products.product_id) to express the join.
Input data
orders.csv:
order_id,product_id,quantity,unit_price
ORD-001,PROD-A,5,29.99
ORD-002,PROD-B,2,149.99
ORD-003,PROD-A,1,29.99
ORD-004,PROD-C,10,9.99
ORD-005,PROD-B,3,149.99
products.csv:
product_id,product_name,category
PROD-A,Widget Pro,Hardware
PROD-B,DataSync License,Software
PROD-C,Cable Kit,Hardware
Pipeline
order_enrichment.yaml:
pipeline:
name: order_enrichment
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./orders.csv"
schema:
- { name: order_id, type: string }
- { name: product_id, type: string }
- { name: quantity, type: int }
- { name: unit_price, type: float }
- type: source
name: products
config:
name: products
type: csv
path: "./products.csv"
schema:
- { name: product_id, type: string }
- { name: product_name, type: string }
- { name: category, type: string }
- type: combine
name: enrich
input:
orders: orders
products: products
config:
where: "orders.product_id == products.product_id"
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_id = orders.product_id
emit product_name = products.product_name
emit category = products.category
emit quantity = orders.quantity
emit unit_price = orders.unit_price
emit line_total = orders.quantity.to_float() * orders.unit_price
propagate_ck: driver
- type: sink
name: result
input: enrich
config:
name: result
type: csv
path: "./output/enriched_orders.csv"
Run it
clinker run order_enrichment.yaml --dry-run
clinker run order_enrichment.yaml --dry-run -n 3
clinker run order_enrichment.yaml
Expected output
output/enriched_orders.csv:
order_id,product_id,product_name,category,quantity,unit_price,line_total
ORD-001,PROD-A,Widget Pro,Hardware,5,29.99,149.95
ORD-002,PROD-B,DataSync License,Software,2,149.99,299.98
ORD-003,PROD-A,Widget Pro,Hardware,1,29.99,29.99
ORD-004,PROD-C,Cable Kit,Hardware,10,9.99,99.90
ORD-005,PROD-B,DataSync License,Software,3,149.99,449.97
How combine works
A combine node declares every input in its input: map, binding each upstream stream to a qualifier used inside expressions:
- type: combine
name: enrich
input:
orders: orders # qualifier: upstream_node
products: products
config:
where: "orders.product_id == products.product_id"
propagate_ck: driver
The config: block carries four fields that shape behavior:
where– a CXL boolean expression. Every field reference must be qualified with its input name. The expression must contain at least one cross-input equality (e.g.orders.product_id == products.product_id); additional range or arbitrary conjuncts can be combined withand.match–first(default),all, orcollect. See below.on_miss–null_fields(default),skip, orerror. Applies only to records on the driving input that find no match.cxl– emit statements that shape the output row. Undermatch: collect, this field must be empty; the combine node auto-derives the output schema.
Match modes
match: first
Emit one output row per driver record, using the first matching build-side record. This is the standard 1:1 enrichment. When no match exists, the behavior is governed by on_miss.
config:
where: "orders.product_id == products.product_id"
match: first
match: all
Emit one output row for every matching build-side record. This is 1:N fan-out – if a driver record matches three build records, three rows are emitted.
- type: combine
name: expand_benefits
input:
employees: employees
benefits: benefits
config:
where: "employees.department == benefits.department"
match: all
cxl: |
emit employee_id = employees.employee_id
emit benefit = benefits.benefit_name
propagate_ck: driver
An employee in a department with three benefits produces three output records.
match: collect
Gather every matching build-side record into a single Array-typed field on the output row. The driver record appears once; the build matches are aggregated into a list. The cxl: body must be empty under match: collect – the combine node synthesizes the output as { driver fields..., <build_qualifier>: Array }.
- type: combine
name: gather
input:
orders: orders
products: products
config:
where: "orders.product_id == products.product_id"
match: collect
cxl: ""
propagate_ck: driver
Use collect when you need the set of matches as a single structured value (e.g. every price history row for an order). Use all when you need one flat row per match.
Unmatched records (on_miss)
on_miss controls what happens to driver records with zero matches:
config:
where: "orders.product_id == products.product_id"
on_miss: null_fields # default: emit with build fields set to null
config:
where: "orders.product_id == products.product_id"
on_miss: skip # inner-join semantics: drop unmatched drivers
config:
where: "orders.product_id == products.product_id"
on_miss: error # fail the pipeline on first unmatched driver
Use skip for inner-join semantics, null_fields for left-join semantics, and error for strict referential integrity where any miss should halt processing.
Composite keys
Chain multiple equalities with and to combine on more than one field. Each conjunct is a separate cross-input equality:
- type: combine
name: match_by_region
input:
sales: sales
targets: targets
config:
where: |
sales.department == targets.department
and sales.region == targets.region
cxl: |
emit department = sales.department
emit region = sales.region
emit actual = sales.amount
emit goal = targets.goal
propagate_ck: driver
Both equalities must hold for a record pair to match.
Equi plus residual filter
The where clause can mix equi predicates with additional filter conjuncts. Non-equality conjuncts are applied as a residual filter after the equi match:
- type: combine
name: high_value_enrichment
input:
orders: orders
products: products
config:
where: |
orders.product_id == products.product_id
and orders.amount >= 100
match: first
on_miss: skip
cxl: |
emit order_id = orders.order_id
emit product_name = products.product_name
emit amount = orders.amount
propagate_ck: driver
The equi conjunct drives the hash lookup; the amount >= 100 conjunct is evaluated as a post-filter. At least one cross-input equality is required in every combine.
Multi-input combine (three or more)
Combine accepts any number of inputs. Each pair of inputs that should be related needs an explicit equality in the where clause:
- type: combine
name: fully_enriched
input:
orders: orders
products: products
categories: categories
config:
where: |
orders.product_id == products.product_id
and products.category_id == categories.category_id
match: first
on_miss: null_fields
cxl: |
emit order_id = orders.order_id
emit product_name = products.product_name
emit category_name = categories.name
emit amount = orders.amount
propagate_ck: driver
Input order in the input: map is preserved, and downstream reasoning treats the first input as the default driving side unless a drive: hint overrides it.
Choosing the driving input
By default the planner picks a driving (probe) input and builds hash tables for the rest. Use drive: to force a specific input to be the driver – typically the larger stream, or the one whose ordering you want to preserve:
- type: combine
name: product_driven
input:
orders: orders
products: products
config:
where: "orders.product_id == products.product_id"
match: first
drive: products
cxl: |
emit product_id = products.product_id
emit product_name = products.product_name
emit sample_order_id = orders.order_id
propagate_ck: driver
With drive: products, the pipeline emits one row per product enriched with a matching order, instead of one row per order enriched with its product.
Memory considerations
Build-side inputs are materialized in memory as hash tables keyed by the equi columns. For each non-driving input, plan for roughly 1.5-2x the raw CSV size in heap. A 50 MB product catalog typically uses 75-100 MB of hash-table memory. Tune with --memory-limit; see Memory Tuning for spill thresholds and strategy overrides.
Document boundaries
When the driver carries document boundaries – a glob: source where each file is its own document – the Combine forwards those boundaries to its output, so a per-document Aggregate after the join rolls up per driver document. See Combine – Document boundaries.
Routing to Multiple Outputs
This recipe splits a stream of order records into separate output files based on business rules. High-value orders go to one file, standard orders to another.
Input data
orders.csv:
order_id,customer,amount,region
ORD-001,Acme Corp,15000,US
ORD-002,Globex,450,EU
ORD-003,Initech,8500,US
ORD-004,Umbrella,22000,APAC
ORD-005,Stark Ind,950,US
ORD-006,Wayne Ent,3200,EU
Pipeline
order_routing.yaml:
pipeline:
name: order_routing
vars:
high_value_threshold: { type: int, default: 5000 }
nodes:
- type: source
name: orders
config:
name: orders
type: csv
path: "./orders.csv"
schema:
- { name: order_id, type: string }
- { name: customer, type: string }
- { name: amount, type: float }
- { name: region, type: string }
- type: route
name: split_by_value
input: orders
config:
mode: exclusive
conditions:
high: "amount >= $vars.high_value_threshold"
default: standard
- type: sink
name: high_value_output
input: split_by_value.high
config:
name: high_value_output
type: csv
path: "./output/high_value.csv"
- type: sink
name: standard_output
input: split_by_value.standard
config:
name: standard_output
type: csv
path: "./output/standard.csv"
Run it
clinker run order_routing.yaml --dry-run
mkdir -p output
clinker run order_routing.yaml
Expected output
output/high_value.csv:
order_id,customer,amount,region
ORD-001,Acme Corp,15000,US
ORD-003,Initech,8500,US
ORD-004,Umbrella,22000,APAC
output/standard.csv:
order_id,customer,amount,region
ORD-002,Globex,450,EU
ORD-005,Stark Ind,950,US
ORD-006,Wayne Ent,3200,EU
How routing works
Port syntax
Route nodes produce named output ports. Downstream nodes reference these ports using dot syntax: split_by_value.high and split_by_value.standard.
The port names come from two places:
- Condition names in the
conditionsmap (here,high) - The
defaultfield (here,standard)
Exclusive mode
With mode: exclusive, each record goes to exactly one branch. Conditions are evaluated top to bottom – the first matching condition wins, and the record is sent to that port. Records that match no condition go to the default port.
Pipeline variables
The threshold is defined in pipeline.vars and referenced in the CXL expression as $vars.high_value_threshold. This makes it easy to adjust the threshold without editing the route condition, and channel overrides can change it per environment.
Variations
Multiple branches
Route nodes can have any number of named branches:
- type: route
name: split_by_region
input: orders
config:
mode: exclusive
conditions:
us: "region == \"US\""
eu: "region == \"EU\""
apac: "region == \"APAC\""
default: other
- type: sink
name: us_output
input: split_by_region.us
config:
name: us_output
type: csv
path: "./output/us_orders.csv"
- type: sink
name: eu_output
input: split_by_region.eu
config:
name: eu_output
type: csv
path: "./output/eu_orders.csv"
# ... additional outputs for apac, other
Transform before output
Insert a transform between the route and output to shape the data differently per branch:
- type: transform
name: enrich_high_value
input: split_by_value.high
config:
cxl: |
emit order_id = order_id
emit customer = customer
emit amount = amount
emit priority = "URGENT"
emit review_required = true
- type: sink
name: high_value_output
input: enrich_high_value
config:
name: high_value_output
type: csv
path: "./output/high_value.csv"
Combining routing with aggregation
Route first, then aggregate each branch independently:
- type: aggregate
name: high_value_summary
input: split_by_value.high
config:
group_by: [region]
cxl: |
emit total = sum(amount)
emit count = count(*)
This produces a per-region summary of high-value orders only.
Aggregation & Rollups
This recipe demonstrates grouping records and computing summary statistics. The pipeline filters active sales records, then rolls them up by department.
Input data
sales.csv:
id,department,amount,status,rep
1,Engineering,5000,active,Alice
2,Marketing,3000,active,Bob
3,Engineering,7000,active,Carol
4,Sales,4000,inactive,Dave
5,Marketing,2000,active,Eva
6,Engineering,9500,active,Frank
7,Sales,6000,active,Grace
8,Marketing,1500,inactive,Hank
Pipeline
dept_rollup.yaml:
pipeline:
name: dept_rollup
nodes:
- type: source
name: sales
config:
name: sales
type: csv
path: "./sales.csv"
schema:
- { name: id, type: int }
- { name: department, type: string }
- { name: amount, type: float }
- { name: status, type: string }
- { name: rep, type: string }
- type: transform
name: active_only
input: sales
config:
cxl: |
filter status == "active"
- type: aggregate
name: rollup
input: active_only
config:
group_by: [department]
cxl: |
emit total = sum(amount)
emit count = count(*)
emit average = avg(amount)
emit maximum = max(amount)
emit minimum = min(amount)
- type: sink
name: report
input: rollup
config:
name: report
type: csv
path: "./output/dept_totals.csv"
sort_order: [{ field: department, order: asc }]
Run it
clinker run dept_rollup.yaml --dry-run
mkdir -p output
clinker run dept_rollup.yaml
Expected output
output/dept_totals.csv:
department,total,count,average,maximum,minimum
Engineering,21500,3,7166.666666666667,9500,5000
Marketing,5000,2,2500,3000,2000
Sales,6000,1,6000,6000,6000
One row per department, ordered by the Sink’s explicit sort_order. The
average column is a binary float and is not implicitly rounded to two decimal
places. The inactive records (Dave’s $4000, Hank’s $1500) are excluded by the filter.
How aggregation works
Group-by keys
The group_by field lists the columns that define each group. Records with the same values for all group-by columns are aggregated together. The group-by columns appear automatically in the output – you do not need to emit them.
Aggregate functions
Available aggregate functions in CXL:
| Function | Description |
|---|---|
sum(expr) | Sum of values |
count(*) | Number of records |
avg(expr) | Arithmetic mean |
min(expr) | Minimum value |
max(expr) | Maximum value |
first(expr) | First value encountered |
last(expr) | Last value encountered |
Per-document aggregation
Per-document aggregation works after any upstream that forwards
document boundaries — a document-aware source (a glob: / paths:
source that treats each file as its own document, or an enveloped format
like XML or EDI), a Merge, or a Combine. The Aggregate produces one
set of grouped rows per document rather than a single aggregate
spanning every file. Each document’s groups finalize and emit at that
document’s close boundary, so a glob over twelve monthly files yields
twelve independent monthly roll-ups. A plain single-file source is one
document and still emits a single aggregate. This holds whether the
Aggregate’s upstream streams or materializes its output — both flush
per document.
A Merge of distinct single-document sources forwards each source’s
per-document close on every mode — concat, seeded interleave, and
the fused unseeded all-Source interleave fast path — so each source
flushes its own roll-up. And per-document aggregation also works after a
Combine on any join
strategy — for example a driver glob: source joined to a lookup table,
then a group-by Aggregate, yields one roll-up per driver document:
nodes:
- type: source # driver: each monthly file is a document
name: orders
config: { name: orders, type: csv, glob: "./orders/*.csv", schema: [ ... ] }
- type: source # small lookup table
name: products
config: { name: products, type: csv, path: "./products.csv", schema: [ ... ] }
- type: combine
name: enrich
input: { orders: orders, products: products }
config: { where: "orders.product_id == products.product_id", match: first, on_miss: skip, propagate_ck: driver, cxl: "..." }
- type: aggregate
name: monthly_totals # one roll-up per driver document (per month)
input: enrich
config: { group_by: [category], cxl: "..." }
See Envelopes & Document Context
for the boundary rules across every Merge mode and Combine strategy.
Strategy selection
Clinker offers two aggregation strategies:
-
Hash aggregation (default): Builds an in-memory hash map keyed by the group-by columns. Works with any input order. Memory usage is proportional to the number of distinct groups.
-
Streaming aggregation: Processes records in order, emitting each group’s result as soon as the next group starts. Requires input sorted by the group-by keys. Uses minimal memory regardless of the number of groups.
The default strategy (auto) selects streaming when the optimizer can prove the input is sorted by the group-by keys, and hash otherwise. You can force a strategy:
config:
group_by: [department]
strategy: streaming # requires sorted input
See Memory Tuning for details on memory implications.
Variations
Multiple group-by keys
config:
group_by: [department, region]
cxl: |
emit total = sum(amount)
emit count = count(*)
Produces one row per unique (department, region) combination.
Pre-aggregation transform
Compute derived fields before aggregating:
- type: transform
name: prepare
input: sales
config:
cxl: |
filter status == "active"
emit department = department
emit amount = amount
emit is_large = amount >= 5000
- type: aggregate
name: rollup
input: prepare
config:
group_by: [department]
cxl: |
emit total = sum(amount)
emit large_count = sum(if is_large then 1 else 0)
emit small_count = sum(if not is_large then 1 else 0)
Aggregation followed by routing
Aggregate first, then route the summary rows:
- type: aggregate
name: rollup
input: active_only
config:
group_by: [department]
cxl: |
emit total = sum(amount)
- type: route
name: split_by_total
input: rollup
config:
mode: exclusive
conditions:
large: "total >= 10000"
default: small
This routes departments with over $10,000 in total sales to one output and the rest to another.
No group-by (grand total)
Omit group_by to aggregate all records into a single output row:
config:
cxl: |
emit grand_total = sum(amount)
emit record_count = count(*)
emit average_amount = avg(amount)
Time-windowed rollups
To group records into event-time buckets, declare a
watermark: on every source
and a time_window:
on the aggregate. Each window emits one rollup per group when it closes.
Three patterns cover the common shapes; all three
ship as runnable pipelines under examples/pipelines/.
Tumbling: hourly click counts
Non-overlapping one-hour buckets per user. Use when each record should contribute to exactly one reporting bucket.
examples/pipelines/tumbling_clicks.yaml:
pipeline:
name: tumbling_clicks
nodes:
- type: source
name: clicks
description: Per-user click stream with an event-time column.
config:
name: clicks
type: csv
path: ./data/tumbling_clicks.csv
options:
has_header: true
watermark:
column: event_ts
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: kind, type: string }
- type: aggregate
name: hourly_clicks
description: Per-user click count, bucketed by event-time hour.
input: clicks
config:
group_by: [user_id]
time_window:
tumbling: { size: 1h }
cxl: |
emit user_id = user_id
emit n = count(*)
- type: sink
name: results
input: hourly_clicks
config:
name: results
type: csv
path: ./output/tumbling_clicks.csv
error_handling:
strategy: fail_fast
Run:
cargo run -p clinker -- run examples/pipelines/tumbling_clicks.yaml
Each hour-aligned bucket emits one row per user_id once its time
window has passed. Records that arrive out-of-order land
in the DLQ as late_record — add delay: on the source or
allowed_lateness: on the aggregate if the input has a known
out-of-order tail.
Hopping: 1-hour sums advanced every 5 minutes
Overlapping one-hour windows that move forward every 5 minutes. Use for moving averages and rolling sums where one record should contribute to multiple overlapping reports.
examples/pipelines/hopping_sliding_5m_1h.yaml:
pipeline:
name: hopping_sliding_5m_1h
nodes:
- type: source
name: clicks
config:
name: clicks
type: csv
path: ./data/hopping_clicks.csv
options:
has_header: true
watermark:
column: event_ts
delay: 5s
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: amount, type: int }
- type: aggregate
name: sliding_amount
input: clicks
config:
group_by: [user_id]
time_window:
hopping:
size: 1h
slide: 5m
allowed_lateness: 30s
cxl: |
emit user_id = user_id
emit total = sum(amount)
emit n = count(*)
- type: sink
name: results
input: sliding_amount
config:
name: results
type: csv
path: ./output/hopping_sliding_5m_1h.csv
error_handling:
strategy: fail_fast
Run:
cargo run -p clinker -- run examples/pipelines/hopping_sliding_5m_1h.yaml
Each record fans into ceil(size / slide) = 12 overlapping
windows, so the output row count is roughly 12× the active-window
record count. The source’s delay: 5s plus the aggregate’s
allowed_lateness: 30s give the pipeline 35 seconds of total grace
beyond strict event-time order before a record drops to the DLQ.
Session: per-user multi-source login sessions
Variable-duration windows bounded by inactivity, computed across two independent sources. Use for activity grouping where the window length is data-driven rather than clock-aligned.
examples/pipelines/multi_source_session.yaml:
pipeline:
name: multi_source_session
nodes:
- type: source
name: src_web
description: Web login events.
config:
name: src_web
type: csv
path: ./data/session_logins.csv
options:
has_header: true
watermark:
column: event_ts
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: source, type: string }
- type: source
name: src_mobile
description: Mobile login events.
config:
name: src_mobile
type: csv
path: ./data/session_mobile.csv
options:
has_header: true
watermark:
column: event_ts
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: source, type: string }
- type: merge
name: all_logins
inputs: [src_web, src_mobile]
- type: aggregate
name: user_sessions
input: all_logins
config:
group_by: [user_id]
time_window:
session: { gap: 5m }
allowed_lateness: 30s
cxl: |
emit user_id = user_id
emit logins = count(*)
- type: sink
name: results
input: user_sessions
config:
name: results
type: csv
path: ./output/multi_source_session.csv
error_handling:
strategy: fail_fast
Run:
cargo run -p clinker -- run examples/pipelines/multi_source_session.yaml
Each source declares its own watermark.column independently. A
session can’t emit until both src_web and src_mobile have caught
up past session_end + allowed_lateness, so the rollup waits for the
slower source before closing. Drop the watermark: block on either
source and the pipeline is rejected at plan time with
E156.
When to pick each
| Kind | Bucket shape | Typical use |
|---|---|---|
tumbling | Disjoint, clock-aligned, fixed width | Hourly metrics, daily rollups, billing periods. |
hopping | Overlapping, clock-aligned, fixed width | Moving averages, sliding sums, anomaly detection where each record should affect multiple reports. |
session | Variable width, gap-bounded, per-key | User sessions, telemetry burst grouping, activity envelopes where the window length is data-driven. |
Slowly-Changing Dimensions (SCD Type 2)
This recipe splits over-long dimension records into closed historical rows plus a continuation row, the shape an SCD Type 2 backfill produces. It uses a Reshape node to both mutate the record that needs splitting and synthesize the continuation row, per correlation group.
It ships as a runnable pipeline at examples/pipelines/scd_type2.yaml.
The problem
Each subject (here, an employee) has a history of dimension records — benefit plans, addresses, price tiers — each with a validity window. A record whose window runs longer than it should (it was never split when the underlying fact changed) needs to be closed at the boundary, with a fresh continuation record carrying the rest of the window forward. That is a per-group transformation: the whole employee’s history is one correlation group, and the split must observe the original group, not a half-mutated one.
Reshape fits exactly: it groups by partition_by, and within each group a rule can both mutate the trigger record and synthesize a derived one — all against the original group snapshot (the no-cascade contract).
Input data
data/scd_plans.csv — each employee’s plan history, with start/end day-numbers:
employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,100,baseline
E002,200,150,baseline
E002,2000,300,baseline
E003,300,250,baseline
E004,400,380,baseline
E004,1500,400,baseline
E005,500,450,baseline
E001, E002, and E004 each have one row whose window exceeds the one-year boundary (plan_start - plan_end > 365); the others are already short enough.
Pipeline
scd_type2.yaml:
pipeline:
name: scd_type2_backfill
# A small budget forces the spill path even on this tiny fixture, so the
# example also demonstrates bounded-memory Reshape. Raise or remove it for
# production volumes.
memory: { limit: "16K", backpressure: spill }
nodes:
- type: source
name: plans
config:
name: plans
type: csv
path: ./data/scd_plans.csv
options:
has_header: true
schema:
- { name: employee_id, type: string }
- { name: plan_start, type: int }
- { name: plan_end, type: int }
- { name: status, type: string }
- type: reshape
name: backfill
input: plans
config:
partition_by: [employee_id]
order_by:
- { field: plan_start, order: asc }
rules:
- name: split_long_plan
when: "plan_start - plan_end > 365"
mutate:
set:
plan_end: "plan_start" # close the over-long window
synthesize:
copy_from: none
overrides:
employee_id: "employee_id"
plan_start: "plan_start"
plan_end: "plan_end" # the rest of the window
status: "'synthesized'"
- type: sink
name: out
input: backfill
config:
name: out
type: csv
path: ./output/scd_type2.csv
error_handling:
strategy: continue
Run it
cargo run -p clinker -- run examples/pipelines/scd_type2.yaml
Expected output
output/scd_type2.csv:
employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,1000,baseline
E001,1000,100,synthesized
E002,200,150,baseline
E002,2000,2000,baseline
E002,2000,300,synthesized
E003,300,250,baseline
E004,400,380,baseline
E004,1500,1500,baseline
E004,1500,400,synthesized
E005,500,450,baseline
Eight input rows produce eleven output rows: the three trigger rows have their plan_end closed at plan_start, and each emits one status=synthesized continuation row. The untriggered rows (E003, E005, the short rows of every employee) pass through unchanged.
How it works
One rule, two actions
The single rule’s when predicate selects the trigger rows. For each trigger row:
mutate.setrewritesplan_endtoplan_start, closing the over-long window at its start boundary. The mutated row keeps its identity and is emitted in place.synthesizederives a brand-new continuation row.copy_from: nonestarts from an all-null base and every column is supplied by anoverridesexpression, so the continuation row is fully constructed from the trigger’s values rather than copied. It is markedstatus=synthesizedso downstream stages can tell originals from engine-derived rows.
Because Reshape applies every rule against the original group snapshot, the mutate and the synthesize both read the trigger row as it arrived — the mutation never feeds back into the synthesis.
Audit provenance
Reshape stamps $meta.synthetic, $meta.synthesized_by, and $meta.mutated_by on every output row (see Audit stamps). These stay out of the default CSV output but are available for downstream CXL — a follow-on Route or Transform can filter on $meta.synthetic to handle generated rows separately.
Bounded memory and spill
The example’s memory.limit: "16K" with backpressure: spill is deliberately tiny so the run exercises Reshape’s disk-spill path on a small fixture. Reshape buffers each employee’s group, and when the budget trips it spills the raw input records to disk and re-runs synthesis on reload — the output is identical whether a group stayed in memory or round-tripped through disk. The per-stage spill volume appears in clinker run --explain and in the post-run spill summary. For real workloads, drop the artificial limit (the default budget is 512 MB) and Reshape stays in memory until it genuinely needs to spill.
Two limits apply: a single correlation group must still fit the memory budget at finalize (the no-cascade contract reloads the whole group to apply its rules — a group larger than the budget fails loud rather than crashing), and Reshape rules cannot reference $doc document context while spill is in play (such a pipeline is rejected at compile time). Each employee’s group in this example is tiny, so neither limit is reached here. See Reshape’s memory model and Memory & Spill for the full picture.
Idempotence
Re-running this pipeline over its own output does not re-trigger: a closed row has plan_start - plan_end == 0, and a synthesized continuation row likewise sits inside the boundary, so the when predicate fires only on genuinely over-long windows. That makes the backfill safe to apply repeatedly.
Backfill, Then Cull for Review
This recipe chains two grouping operators: a Reshape node backfills each subject’s history, then a Cull node sets aside whole subjects that need manual review — routing them to a second output stream instead of dropping them.
It ships as a runnable pipeline at examples/pipelines/employee_plan_backfill.yaml.
The problem
You have run an SCD-style backfill over each employee’s benefit-plan history (closing over-long windows and synthesizing continuation rows). After backfilling, some employees end up with a large or otherwise unusual plan history that an analyst should eyeball before it lands in the clean dataset. You want two outputs:
- a clean stream for the employees whose history looks fine, and
- a review stream for the flagged employees — their records intact, not discarded, not errored.
That is a per-group decision based on an aggregate property of the whole group (“this employee has more than three plan rows”), and the flagged records belong on a second valid data stream, not in the dead-letter queue. Cull fits exactly: it groups by partition_by, evaluates a group-level drop_group_when predicate, and emits removed groups on a first-class removed_to side-output port.
Input data
data/employee_plans.csv — each employee’s plan history, with start/end day-numbers:
employee_id,plan_start,plan_end,status
E001,100,90,baseline
E001,1000,100,baseline
E002,200,150,baseline
E003,50,40,baseline
E003,300,250,baseline
E003,600,550,baseline
E003,900,850,baseline
E001 has one over-long window (1000 - 100 > 365); E003 has four plan rows.
The pipeline
nodes:
- type: source
name: plans
config:
name: plans
type: csv
path: ./data/employee_plans.csv
schema:
- { name: employee_id, type: string }
- { name: plan_start, type: int }
- { name: plan_end, type: int }
- { name: status, type: string }
- type: reshape
name: backfill
input: plans
config:
partition_by: [employee_id]
order_by:
- { field: plan_start, order: asc }
rules:
- name: split_long_plan
when: "plan_start - plan_end > 365"
mutate:
set:
plan_end: "plan_start"
synthesize:
copy_from: trigger
overrides:
status: "'synthesized'"
- type: cull
name: flag_large_histories
input: backfill
config:
partition_by: [employee_id]
removed_to: review
rules:
- name: too_many_plans
drop_group_when: "count(*) > 3"
- type: sink
name: out
input: flag_large_histories # main port — kept employees
config: { name: out, type: csv, path: ./output/employee_plans_clean.csv }
- type: sink
name: review
input: flag_large_histories.review # side-output port — flagged employees
config: { name: review, type: csv, path: ./output/employee_plans_review.csv }
How it works
- Reshape (
backfill) groups byemployee_idand, for the over-long window, closes it (plan_end = plan_start) and synthesizes a continuation row markedstatus=synthesized.E001gains a synthesized row, ending up with three rows. - Cull (
flag_large_histories) groups the backfilled rows byemployee_idagain and evaluatescount(*) > 3over each whole group.E003has four rows, so the wholeE003group is routed to thereviewside-output port;E001(three rows) andE002(one row) flow to the main output. - Two outputs draw the two ports:
outreferences the Cull node by name (the main port, kept employees);reviewreferencesflag_large_histories.review(the side-output port, flagged employees).
The main output (employee_plans_clean.csv) carries E001 (with its synthesized continuation row) and E002; the review output (employee_plans_review.csv) carries all four of E003’s rows. Both streams carry the unchanged input schema — Cull does not widen, and the flagged records are valid rows on a normal data edge, never DLQ entries.
Expressing group-level conditions
drop_group_when is an aggregate predicate over the whole group. CXL’s bare aggregates are sum / count / min / max / avg / collect / weighted_avg — there is no bare any(). To flag a group when any row matches a condition, sum an indicator and compare to zero:
rules:
# Flag the whole employee for review if any plan row is still flagged
# `status == 'error'` after backfill.
- name: any_error
drop_group_when: "sum(if status == 'error' then 1 else 0) > 0"
See the Cull node reference for the full predicate vocabulary, the producer-side port model, and the bounded-memory spill behavior.
File Splitting
This recipe demonstrates splitting large output files into smaller chunks, optionally keeping related records together.
Basic record-count splitting
Split output into files of at most 5,000 records each:
pipeline:
name: monthly_report
nodes:
- type: source
name: transactions
config:
name: transactions
type: csv
path: "./data/transactions.csv"
schema:
- { name: id, type: int }
- { name: date, type: string }
- { name: department, type: string }
- { name: amount, type: float }
- { name: description, type: string }
- type: sink
name: split_output
input: transactions
config:
name: split_output
type: csv
path: "./output/report.csv"
split:
max_records: 5000
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true
Output files
output/report_0001.csv (5000 records + header)
output/report_0002.csv (5000 records + header)
output/report_0003.csv (remaining records + header)
Naming pattern variables
| Variable | Description | Example |
|---|---|---|
{stem} | Base filename without extension | report |
{ext} | File extension | csv |
{seq:04} | Zero-padded sequence number (width 4) | 0001 |
The path field provides the template: ./output/report.csv means stem is report and ext is csv.
Header behavior
When repeat_header: true, each output file includes the CSV header row. This is the recommended setting – each file is self-contained and can be processed independently.
Grouped splitting
Keep all records with the same group key value in the same file:
split:
max_records: 5000
group_key: "department"
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true
oversize_group: warn
With group_key: "department", the splitter ensures that all records for a given department land in the same output file. A new file starts only at a group boundary (when the department value changes), even if the current file has not reached max_records yet.
Oversize group policy
If a single group contains more records than max_records, the oversize_group setting controls behavior:
| Policy | Behavior |
|---|---|
warn (default) | Log a warning and write all records for the group into one file, exceeding the limit |
error | Stop the pipeline with an error |
allow | Silently allow the oversized file |
For example, if max_records is 5,000 but the Engineering department has 7,000 records, the warn policy produces a file with 7,000 records and logs a warning.
Byte-based splitting
Split by file size instead of record count:
split:
max_bytes: 10485760 # 10 MB per file
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true
The splitter estimates the current file size and starts a new file when the limit is approached. The actual file size may slightly exceed the limit because the current record is always completed before splitting.
Byte-based splitting works with every output format. For formats that wrap the whole file in framing – a JSON array or an XML root element – each rotation closes the current file’s framing and reopens it in the next, so every chunk is a complete, independently valid document (its own [ ... ] array or <Root> ... </Root> tree), never a fragment.
Combined limits
Use both max_records and max_bytes together – whichever limit is reached first triggers a new file:
split:
max_records: 10000
max_bytes: 5242880 # 5 MB
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true
This is useful when record sizes vary widely. Short records might produce a tiny file at 10,000 records, while long records might hit the byte limit well before 10,000.
Full pipeline example
A complete pipeline that reads a large transaction file, filters it, and splits the output:
pipeline:
name: split_transactions
nodes:
- type: source
name: transactions
config:
name: transactions
type: csv
path: "./data/all_transactions.csv"
schema:
- { name: id, type: int }
- { name: date, type: string }
- { name: department, type: string }
- { name: category, type: string }
- { name: amount, type: float }
- type: transform
name: current_year
input: transactions
config:
cxl: |
filter date.starts_with("2026")
- type: sink
name: chunked
input: current_year
config:
name: chunked
type: csv
path: "./output/transactions_2026.csv"
split:
max_records: 5000
group_key: "department"
naming: "{stem}_{seq:04}.{ext}"
repeat_header: true
oversize_group: warn
clinker run split_transactions.yaml --force
Practical considerations
-
Downstream consumers. Splitting is useful when the receiving system has file size limits (e.g., an upload API that accepts files up to 10 MB) or when parallel processing of chunks is desired.
-
Record ordering. Records within each output file maintain their original order from the pipeline. Across files, the sequence number (
{seq}) indicates the order. -
Group key sorting. For
group_keyto work correctly, the input should ideally be sorted by the group key. If the input is not sorted, records for the same group may appear in multiple files. Pre-sort with a transform if needed, or accept the split-group behavior. -
Overwrite behavior. Use
--forcewhen re-running a pipeline with splitting enabled. Without it, the pipeline aborts if any of the output chunk files already exist.
Intra-Record Closures
This recipe shows the complete intra-record fan-out shape: an NDJSON source where each record carries an array of line items, a transform that filters items by price and then fans each remaining item into its own output record, and a flat NDJSON sink ready for downstream billing.
The pieces involved:
- Arrow-syntax closures for predicates and projections.
- Array methods (
filter,map) for in-place transformation. - Bracket-index access (
it["sku"]) for reading fields off each map element. emit eachfor fan-out.- The Sink node’s
include_unmappedflag for controlling which fields reach the sink.
Input data
orders.ndjson – one JSON object per line, each carrying a nested items array:
{"order_id":"O-1","customer":"alice@example.com","items":[{"sku":"a","price":10,"qty":2},{"sku":"b","price":20,"qty":1},{"sku":"c","price":3,"qty":5}]}
{"order_id":"O-2","customer":"bob@example.com","items":[{"sku":"a","price":10,"qty":1},{"sku":"d","price":50,"qty":1}]}
Each record has two order-level fields (order_id, customer) and an items array whose elements are maps with sku, price, and qty.
Goal
For each order:
- Drop items priced under $5 (a sub-threshold cutoff).
- Fan the surviving items into one output record each, carrying the order-level identifiers plus the per-item fields.
- Compute the per-line revenue (
unit_price * qty) for each output record.
Pipeline
billing_lines.yaml:
pipeline:
name: billing_lines
nodes:
- type: source
name: orders
config:
name: orders
type: json
options:
format: ndjson
path: "./orders.ndjson"
schema:
- { name: order_id, type: string }
- { name: customer, type: string }
- { name: items, type: any }
- type: transform
name: filter_lines
input: orders
config:
cxl: |
emit order_id = order_id
emit customer = customer
emit item_count = items.length()
emit kept = items.filter(it => it["price"] >= 5)
- type: transform
name: explode
input: filter_lines
config:
max_expansion: 10000
cxl: |
emit each it in kept {
emit order_id = order_id
emit customer = customer
emit sku = it["sku"]
emit unit_price = it["price"]
emit qty = it["qty"]
emit line_total = it["price"] * it["qty"]
}
- type: sink
name: lines_out
input: explode
config:
name: lines_out
type: json
path: "./output/billing_lines.ndjson"
options:
format: ndjson
include_unmapped: false
exclude: [items, kept]
error_handling:
strategy: continue
Run it
# Validate first
clinker run billing_lines.yaml --dry-run
# Preview the first few output records
clinker run billing_lines.yaml --dry-run -n 3
# Full run
clinker run billing_lines.yaml
Expected output
output/billing_lines.ndjson:
{"order_id":"O-1","customer":"alice@example.com","sku":"a","unit_price":10,"qty":2,"line_total":20}
{"order_id":"O-1","customer":"alice@example.com","sku":"b","unit_price":20,"qty":1,"line_total":20}
{"order_id":"O-2","customer":"bob@example.com","sku":"a","unit_price":10,"qty":1,"line_total":10}
{"order_id":"O-2","customer":"bob@example.com","sku":"d","unit_price":50,"qty":1,"line_total":50}
Order O-1’s three input items collapse to two output records (the sku=c line was filtered out because its price was below $5). Order O-2’s two items both survive the filter and produce two output records.
How it works
Filter stage. The filter_lines transform reads each order, runs items.filter(it => it["price"] >= 5) to drop sub-threshold items, and stashes the survivors in a kept field. The closure body uses bracket indexing (it["price"]) because each it is a map; bracket indexing returns null for missing keys without aborting. The same record also carries an item_count projection so downstream nodes could route or audit on the original (pre-filter) item count.
Explode stage. The explode transform contains one emit each block over kept. For each surviving item, the body emits a flat record with the order-level identifiers (order_id, customer) repeated, plus the per-item fields lifted out of it. The body has no filter or nested emit each – those are forbidden inside the block; pre-filter upstream as we did, or post-filter in a downstream transform.
include_unmapped: false. The default Sink policy is to pass every unmapped input field through. Here we set it to false so the order-level items array (carried through from the source), the item_count projection, and the intermediate kept array (used only as the fan-out source) do not leak into the per-line output. The exclude: [items, kept] list provides a belt-and-suspenders defense against future renaming.
max_expansion: 10000. Caps how many output records a single input order may produce. The default is 10000; we set it explicitly here so the value is visible in the YAML. Orders with arrays larger than the cap route to the DLQ with category expansion_limit_exceeded (see Transform Nodes -> Expansion Cap).
Variations
Pass through every input field
Remove include_unmapped: false (or set it to true) and the original order-level fields plus the intermediate kept array will appear on every output record. Useful when downstream consumers expect a complete record context, or when you need to audit what was filtered.
Emit a single record per order with the kept-items array
Drop the explode transform and route filter_lines directly to the Sink. Each output record stays at order grain, with kept carrying the post-filter array. This is the same pipeline minus the fan-out step.
Reach for .flat_map instead of two transforms
When the per-element transformation is simple enough to fit in a single closure body, flat_map collapses the filter + project + explode pattern into one expression. It produces a flat array, which downstream nodes still see as a single field on the input record; the explicit emit each is what produces multiple output records.
Rewrite a nested field in place with .set
When you want to keep the record at order grain but mutate a value buried inside it, the set map method takes a dotted/indexed path and rewrites a single leaf, leaving every sibling untouched:
cxl: |
emit order = order.set("items[0].sku", "A-100").set("ship.region", "us-east")
The first set overwrites the SKU of the first item; the second writes ship.region, auto-creating the ship map if the order had no ship field yet. Because set is copy-on-write, this builds a fresh order document without disturbing the upstream binding. A path that conflicts with the existing shape (descending into a scalar, or an array index past the end) yields null for that set rather than partially writing – guard with catch if a path may not match every record.
See also
- Closures – the
it => bodyform. - Array Methods –
filter,map,find,any,flat_map,remove,length,join. - Map Methods –
keys,values,merge,set,remove_field,unset. - Nested Paths – bracket-index and dotted-path navigation.
- Emit Each – the fan-out statement.
- Transform Nodes –
max_expansionand DLQ routing. - Sink Nodes –
include_unmappedand field control.