Interactive Guides
Some parts of Clinker are easier to understand by trying them than by reading about them. Each guide below is a single page that runs in your browser, on a desktop or a phone. You change a setting or tap a record, and the page shows what the engine does with it. The guides don’t run a pipeline; each one reproduces the rules described on the reference page it links back to.
Where does a null go?
How CXL works out an expression when a field is empty: and, or and not
with null, why null == null is true, and why a filter drops a record whose
condition comes out null. Reference: Null Handling.
Correlation keys
How one failing line takes the rest of its order to the DLQ, and how an
Aggregate’s group_by decides between rejecting a whole group and recomputing
totals without the failed line. Reference: Correlation Keys.
Document context
Where $doc.* values come from, how each file becomes its own document with
its own Aggregate roll-up, and what dlq_granularity: document rejects.
Reference: Document Envelope Context.
How many documents?
What an Envelope node’s preserve and concat do to the documents a Sink
writes, how a synthesized footer differs between them, how an Aggregate after
the Envelope rolls up, and which pipeline shapes E347 and E355 reject.
Reference: Envelope Nodes.
Which files get read?
How a Source turns glob:, regex:, paths: and its filters into the list of
files it reads, and in what order. Reference:
Source Nodes → Choosing files.
Is my data still sorted?
Which stages keep a Source’s declared sort_order and which drop it, and when
an Aggregate can stream instead of holding every group. Reference:
Aggregate Nodes and Source Nodes.
Route and Merge
Where each record goes when a Route’s conditions are true, not true, or fail, in exclusive and inclusive mode, and how a Merge rejoins the branches. Reference: Route Nodes and Merge Nodes.
Combine playground
Which build rows where: matches for each driver row, and what match:,
on_miss: and drive: do with them, including range joins. Reference:
Combine Nodes.
Window functions
Which rows of a partition each $window.* function reads, and why
$window.sum is a partition total while $window.cumulative_sum is the
running total. Reference: Window Functions.
Where did my rows go?
What the end-of-run line N total, N ok, N written, N dlq counts, why the
numbers often don’t add up, and which rows are in no number at all: filtered
and duplicate rows, rows folded into Aggregate groups, Combine misses and Route
branches nothing reads. Reference: Metrics & Monitoring.
Streaming vs. blocking
Which stages stream and which hold their output in a buffer, for several
pipeline shapes, with the --explain lines each one produces. Reference:
Streaming vs. Blocking Stages.
For engine developers
The Clinker Engine Internals book has three more:
- a memory system explainer, linked from its Memory Arbitration & Scheduling chapter. It covers the memory budget, back-pressure, spilling to disk and the scheduler, with a simulator.
- a range-join explainer, linked from its Combine Internals chapter. It runs the block-band IEJoin on small inputs: sorting and slicing into blocks, pruning block pairs, the kernel step by step, and the nested-loop fallback.
- a retraction-loop explainer, linked from its Retraction Protocol chapter.
It replays the commit loop that corrects an Aggregate whose
group_byleaves out a correlation-key field, on three of the engine’s test pipelines.