Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Interactive Guides

Some parts of Clinker are easier to understand by trying them than by reading about them. Each guide below is a single page that runs in your browser, on a desktop or a phone. You change a setting or tap a record, and the page shows what the engine does with it. The guides don’t run a pipeline; each one reproduces the rules described on the reference page it links back to.

Where does a null go?

How CXL works out an expression when a field is empty: and, or and not with null, why null == null is true, and why a filter drops a record whose condition comes out null. Reference: Null Handling.

Correlation keys

How one failing line takes the rest of its order to the DLQ, and how an Aggregate’s group_by decides between rejecting a whole group and recomputing totals without the failed line. Reference: Correlation Keys.

Document context

Where $doc.* values come from, how each file becomes its own document with its own Aggregate roll-up, and what dlq_granularity: document rejects. Reference: Document Envelope Context.

How many documents?

What an Envelope node’s preserve and concat do to the documents a Sink writes, how a synthesized footer differs between them, how an Aggregate after the Envelope rolls up, and which pipeline shapes E347 and E355 reject. Reference: Envelope Nodes.

Which files get read?

How a Source turns glob:, regex:, paths: and its filters into the list of files it reads, and in what order. Reference: Source Nodes → Choosing files.

Is my data still sorted?

Which stages keep a Source’s declared sort_order and which drop it, and when an Aggregate can stream instead of holding every group. Reference: Aggregate Nodes and Source Nodes.

Route and Merge

Where each record goes when a Route’s conditions are true, not true, or fail, in exclusive and inclusive mode, and how a Merge rejoins the branches. Reference: Route Nodes and Merge Nodes.

Combine playground

Which build rows where: matches for each driver row, and what match:, on_miss: and drive: do with them, including range joins. Reference: Combine Nodes.

Window functions

Which rows of a partition each $window.* function reads, and why $window.sum is a partition total while $window.cumulative_sum is the running total. Reference: Window Functions.

Where did my rows go?

What the end-of-run line N total, N ok, N written, N dlq counts, why the numbers often don’t add up, and which rows are in no number at all: filtered and duplicate rows, rows folded into Aggregate groups, Combine misses and Route branches nothing reads. Reference: Metrics & Monitoring.

Streaming vs. blocking

Which stages stream and which hold their output in a buffer, for several pipeline shapes, with the --explain lines each one produces. Reference: Streaming vs. Blocking Stages.

For engine developers

The Clinker Engine Internals book has three more:

  • a memory system explainer, linked from its Memory Arbitration & Scheduling chapter. It covers the memory budget, back-pressure, spilling to disk and the scheduler, with a simulator.
  • a range-join explainer, linked from its Combine Internals chapter. It runs the block-band IEJoin on small inputs: sorting and slicing into blocks, pruning block pairs, the kernel step by step, and the nested-loop fallback.
  • a retraction-loop explainer, linked from its Retraction Protocol chapter. It replays the commit loop that corrects an Aggregate whose group_by leaves out a correlation-key field, on three of the engine’s test pipelines.