Clinker user guide · ← Source Nodes

Which files
get read?

A Source can read one file or many. Before it reads a byte, Clinker builds the list: it finds candidates, filters them, sorts them, keeps some, and decides what to do if none are left. Pick the settings and watch each file go through the funnel to the final reading order.

01

The funnelEight steps, always in this order

The pipeline's folder

1 · matcher (exactly one)
files.recursive
2 · exclude
4 · min_size
5 · modified_after
6 · sort_by
sort_order
7 · take
8 · on_no_match

      
02

SurprisesEach of these is how the engine works today

regex: searches everything

It walks the pipeline's whole folder, subfolders included, and matches anywhere in each file's path. The path starts with the folder you typed to reach the pipeline file (clinker run pipelines/orders.yaml gives pipelines/data/…), so match the end of it. orders_.*\.csv also finds out/orders_report.csv, last run's output; anchor it: data/orders_\d{4}-\d{2}\.csv$.

* stops at folders

In glob:, * stays inside one folder; ** goes into subfolders. files.recursive doesn't change a glob: use **. It does apply to regex:, which is recursive unless you set it to false.

Hidden files match

*.csv matches .orders_draft.csv: a leading dot is not treated specially. Exclude such files if editors or tools leave them behind.

Your list gets sorted

The files in paths: are sorted like any other match, by name unless files.sort_by says otherwise. The order you write them in is not the reading order.

take_first is the first after sorting

With the default ascending name sort, take_first: 2 keeps the two earliest names. For the latest files, sort desc or use take_last.

Times and sizes are fixed at load

modified_after: 30d means 30 days before the pipeline was loaded. Sizes are decimal: 1KB is 1000 bytes. Times can also be an RFC 3339 timestamp such as 2024-03-01T00:00:00Z.

A missing path: is an error

A path: or paths: entry that doesn't exist stops the run (E216, reported as a file-system error), whatever on_no_match says. on_no_match applies when the filters leave nothing.

Many files, one stream per file

With glob:, regex: or paths:, each file is its own document: a sort_order is checked per file and an Aggregate rolls up per file. See is my data still sorted?