Source Nodes
Source nodes read data from files and are the entry points of every pipeline. They have no input: field – they produce records, they do not consume them.
Basic structure
- type: source
name: customers
config:
name: customers
type: csv
path: "./data/customers.csv"
schema:
- { name: customer_id, type: int }
- { name: name, type: string }
- { name: email, type: string }
- { name: status, type: string }
- { name: amount, type: float }
Schema declaration
The schema: field is required on every source node. Runtime ingestion does not guess types from data: declare each column’s name and CXL type explicitly. This schema drives compile-time type checking across the entire pipeline.
Each entry is a { name, type } pair:
schema:
- { name: employee_id, type: string }
- { name: salary, type: int }
- { name: hired_at, type: date_time }
- { name: is_active, type: bool }
- { name: notes, type: { nullable: string } }
Available types
| Type | Description |
|---|---|
null | The null value only |
string | UTF-8 text |
int | 64-bit signed integer |
float | 64-bit IEEE 754 floating point |
decimal | Exact fixed-point number, subject to the declared precision and scale |
bool | Boolean (true / false) |
date | Calendar date |
date_time | Date with time component |
array | Ordered sequence of values |
map | String-keyed object |
any | Unknown type – field used in type-agnostic contexts |
{ nullable: T } | Nullable wrapper around a concrete inner type, for example { nullable: int } |
Check the schema through the planner
The schema that governs execution is the one accepted by clinker-plan while
compiling the complete pipeline. Run this after changing a source schema:
clinker run pipeline.yaml --explain text
Exit code 0 means the planner parsed the canonical YAML, resolved schema
references and overlays, bound the pipeline, type-checked CXL, and produced a
compiled plan. It does not prove that input bytes satisfy the declarations or
that the resulting output is correct; execute representative input against an
isolated destination and compare the result before a production run.
Workspace tooling may also present advisory schema analysis. That analysis can
help find likely authoring mistakes, but it cannot admit or reject a pipeline.
Its analyzed, partial, skipped, and failed statuses describe only how
much the advisory model inspected. See Validation and
Admission for the
status meanings and known limits.
Without a column format:, date_time accepts the default offset-free forms
and RFC 3339 timestamps with Z or an explicit numeric offset, such as
2026-01-31T08:27:00Z and 2026-01-31T10:57:00+02:30. Zoned timestamps are
normalized to UTC before entering Clinker’s timezone-free date_time
representation. Parsing is exact: surrounding whitespace, malformed calendar
values, and out-of-range timestamps are rejected rather than trimmed, guessed,
overflowed, or rounded. A column-level strftime format: is exclusive: only
that authored format is tried.
A source column’s declared type must be concrete: numeric — the
inference-only int | float union CXL resolves during type unification — is
not a valid source column type. Declaring one is rejected at compile with
E158; declare int or
float explicitly.
clinker guess pipeline.yaml provides an authoring-only preview for columns
explicitly marked numeric. It proposes concrete int or float declarations;
it does not infer arbitrary schemas or change the runtime admission rules.
The default preview is bounded and read-only. --check exhausts the admitted
manifest; --write can publish one guarded, unambiguous inline edit. Resolve
all numeric declarations to concrete types before executing the pipeline.
See Numeric schema authoring and the
guess command.
long_unique — storage hint for high-cardinality text
A string column may carry an optional long_unique: true flag. It is an
advisory, opt-in hint, not a type change: it tells Clinker the column’s
values are long and effectively unique — never repeated across records — so
the run uses less memory for that column. Typical candidates are UUIDs
rendered as text, street addresses, and free-text comment or note fields.
schema:
- { name: ticket_id, type: string, long_unique: true } # 36-char UUID
- { name: notes, type: string, long_unique: true } # free text
- { name: department, type: string } # low-cardinality, default
The flag lowers memory use only. A value’s content and its
comparison, grouping, join, sort, and output behavior are all unchanged — a
long_unique value behaves identically to the same text in any other column.
Omitting the flag (the common case) leaves the default behavior untouched. Set
it only when you know a column is genuinely high-cardinality free text; on a
column whose values repeat, leave it off.
source_name — read a differently-named physical column
A column may carry an optional source_name naming the physical input
column it reads from, when that differs from the exposed name. The reader
matches input fields by physical name and re-labels the value under name, so
downstream CXL and the output see name carrying the physical column’s data.
schema:
# read the physical `cust_id` column, expose it downstream as `customer_id`
- { name: customer_id, type: string, source_name: cust_id }
Omitting source_name (the common case) reads the input field whose key equals
name, unchanged from before. A channel schema patch’s rename op sets this
alias automatically (see Channels).
Transport vs format
A source declaration has two independent layers:
- Transport (
transport:) selects where the records come from. Two transports exist:file— read bytes from the filesystem, resolved through one of the file matchers (path/glob/regex/paths) — andrest— pull records from a paginated HTTP endpoint under a hard page/record cap (see Network Sources (REST)).transport:is optional and defaults tofile, so a source that omits it reads from disk exactly as before.restneeds therestcapability, which the released binary has; a build compiled without it refuses arestsource at validation withE223rather than running the pipeline short one input (see Optional capabilities). - Format (
type:) selects how the bytes decode into records:csv,json,xml,fixed_width,edifact,x12,hl7,swift.
- type: source
name: orders
config:
name: orders
transport: file # optional; this is the default
type: csv # the on-disk format
path: "./data/orders.csv"
schema:
- { name: order_id, type: int }
A file transport requires exactly one file matcher (path, glob, regex, or paths). Declaring none fails validation with E211; declaring more than one fails with E210. Both are reported at config-load time, before any file is opened.
Choosing files
Interactive companion: the file discovery explainer runs these steps on a sample folder as you change the settings.
Before reading, a file Source builds its list of files in a fixed order:
- Matcher. Paths are relative to the pipeline file’s folder.
path:names one file andpaths:a list. A file that does not exist stops the run withE216, whateveron_no_matchsays.glob:matches names inside one folder:*does not cross a/. Use**to include subfolders (./data/**/*.csv). A leading dot is not special, so*.csvalso matches.draft.csv. An invalid pattern isE212.regex:searches the pipeline’s folder and, by default, every subfolder, and matches anywhere in each file’s path. That path begins with the folder part of the pipeline path as you typed it (clinker run pipelines/orders.yamlgivespipelines/data/orders_2024-01.csv), so anchor the end of the path, for exampledata/orders_\d{4}-\d{2}\.csv$. An unanchored pattern can pick up earlier outputs too. An invalid pattern isE213.
exclude:– a list of glob patterns; a file whose name or full path matches any of them is dropped.- Anything that is not a regular file is dropped.
min_size:/max_size:– decimal units:1KBis 1000 bytes, alsoB,MB,GB.modified_after:/modified_before:– a duration back from the time the pipeline is loaded (30s,15m,2h,3d) or an RFC 3339 timestamp (2024-03-01T00:00:00Z).files.sort_by:name(default; the full path),createdormodified, withfiles.sort_order:asc(default) ordesc. The files of apaths:list are sorted too: the order they are written in is not the reading order.files.take_first:orfiles.take_last:keeps that many files from the sorted list (setting both isE218). With the default ascending name sort,take_first: 5keeps the five earliest names.files.on_no_match:– when nothing is left:error(default,E216),warn(log a warning and produce no rows), orskip(produce no rows quietly).
files.recursive: controls whether regex: searches subfolders (true by default). It has no effect on glob:, which searches subfolders only where the pattern has **.
- type: source
name: orders
config:
name: orders
type: csv
glob: ./data/orders_*.csv
exclude: ["*_partial.csv"]
modified_after: 30d
files:
sort_by: name
sort_order: desc
take_first: 3 # the three latest names
on_no_match: warn
schema:
- { name: order_id, type: string }
With glob:, regex: or paths:, each file is read as its own document: a declared sort_order is checked per file, and an Aggregate rolls up per file (see Sort order and One document per file).
Format types
The type: field inside config: selects the on-disk format. Each format
has its own reference page covering its options and decoding model:
type: | Format | Reference |
|---|---|---|
csv | Delimited text (RFC 4180) | CSV Format |
json | Array / NDJSON / wrapper object | JSON Format |
xml | Element-path-selected record elements | XML Format |
fixed_width | Column-positioned legacy extracts | Fixed-Width Format |
edifact | UN/EDIFACT interchanges | EDIFACT Format |
x12 | ANSI ASC X12 interchanges | X12 Format |
hl7 | HL7 v2.x pipe-and-hat messages | HL7 v2 Format |
swift | SWIFT MT (FIN) messages | SWIFT MT Format |
The same schema: rules apply regardless of format: the reader maps each
decoded record onto the declared schema, and undeclared input fields fall
under the on_unmapped policy
below.
Declared-type failures
Source types are enforced before a record reaches buffering, sorting, or any downstream node. A value that cannot satisfy its authored type rejects the whole row exactly once; Clinker never substitutes the raw string, a sentinel, or an error-derived null.
Empty input has three distinct outcomes:
- an empty value declared as
stringremains the empty string; - an empty non-string value declared
nullable(T)becomes null; - an empty non-string value declared as non-nullable is a type error.
Parsing does not trim whitespace, guess locale conventions, or recognize
case-insensitive null sentinels. Integer overflow, invalid/out-of-range dates,
decimal precision overflow, and decimal values that would need rounding to
the declared scale are type errors. Accepted decimal and date values retain
their declared precision.
The E126 diagnostic identifies source, file, one-based row and column,
field, and declared type. Its value preview is sanitized to one line and
limited to 256 rendered UTF-8 bytes; controls, bidi characters, diagnostic
delimiters, backslashes, and invalid UTF-8 are explicit indivisible escape
tokens. When truncated, the preview ends in one … without splitting a token
or Unicode scalar and reports the original byte length. The complete original
record/value is retained only in the configured DLQ. See Error Handling &
DLQ for strategy and threshold
behavior.
Already-decoded values obey the same declarations without hidden coercions:
string admits only a string, null admits only null, and any admits every
supported native value while preserving its represented value. Numeric
conversions must be exact; non-finite floating-point values and decimal values
that would require rounding are rejected. With multiple: true, the scalar
declaration is applied to every array element. If any element fails, the whole
source record is rejected and the complete original array remains available
through the DLQ.
Error strategies and complete populations
fail_fast aborts on the first declared source failure. continue and
best_effort require a DLQ and use the configured threshold over the complete
population:
rejected records / attempted records
The run aborts only when that ratio is strictly greater than the threshold.
Equality is accepted, so a threshold of 0.1 admits exactly one rejected
record in a population of ten but not two. A zero threshold aborts on the first
rejection; a threshold of 1.0 admits an all-rejected population for the
continuing strategies.
For ordered file sources, Clinker establishes the complete attempted and rejected population before any accepted record, punctuation, downstream side effect, or output byte is released. That population is applied exactly once whether execution remains resident, spills, or fuses a downstream node. A threshold violation therefore cannot leave a committed prefix of output.
on_unmapped — undeclared input fields
The per-source on_unmapped policy decides what to do with input fields the source’s schema: block does not name. Three modes — auto_widen (default), drop, reject:
- type: source
name: orders
config:
name: orders
type: csv
path: "./data/orders.csv"
on_unmapped:
mode: auto_widen # default; other values: drop, reject
schema:
- { name: order_id, type: string }
- { name: amount, type: float }
See Auto-Widen & Schema Drift for the full
specification: how undeclared columns flow through each downstream node
type, the include_unmapped Sink flag, E315 merge-policy
mismatch, and fixed-width behavior.
Sort order
If each physical input file is pre-sorted, declare the record order so the planner can admit order-dependent strategies such as streaming aggregation:
- type: source
name: sorted_transactions
config:
name: sorted_transactions
type: csv
path: "./data/transactions_sorted.csv"
schema:
- { name: account_id, type: string }
- { name: txn_date, type: date }
- { name: amount, type: float }
sort_order:
- { field: "account_id", order: asc }
- { field: "txn_date", order: asc }
Clinker binds these fields to the declared source schema, compares the typed values by the same rule every sort uses (see How values are ordered), and verifies each physical file independently before any record from that file reaches an order-dependent consumer. The declaration never means that a multi-file source is globally sorted: the last key in one file is not compared with the first key in the next file.
on_unsorted controls the result of the first adjacent inversion:
| Value | Behavior |
|---|---|
warn (default) | Stably repair the complete physical file with the shared bounded-memory sort, emit one W307 warning for that file, then release it. A file already in order emits no warning. |
error | Reject the physical file without releasing an unverified prefix. The diagnostic identifies the source, file, adjacent rows, and keys. |
sort_order:
- { field: "account_id", order: asc, null_order: last }
- { field: "txn_date", order: desc, null_order: first }
on_unsorted: warn
Source ordering accepts null_order: first or last; drop is rejected
when the pipeline is planned because verifying order must not discard source
records. The error gives one fix: delete null_order: drop and add a
Transform after the Source whose whole config is the line it prints,
config: { cxl: "filter not <field>.is_null()" }. With the line deleted, the
Source declares its null keys last; if a file’s null keys arrive first,
write null_order: first instead of deleting the line. That filter needs a
field CXL can name as it is: one identifier of ASCII letters, digits and _,
not starting with a digit and not a CXL keyword. For any other key, such as
order id, filter or a flattened Address.City, the error prints no CXL
and prints a source_name
line for the column instead, such as source_name: "order id". Set the
column’s name to a new identifier, add that line, use the new name wherever
the pipeline names the column, and plan again: the error then prints the
filter on the new name. Equal authored keys
retain arrival order within the selected execution path. Clinker does not add
a source identity, physical filename, or canonical-row tie-breaker.
Verification stages the complete sortable file event sequence behind the
run’s memory arbitrator. It uses the existing stable resident/spill sort and
bounded-fan-in merge machinery when repair spills, while preserving row
identity, source/file provenance, and document context. Only flat sources and
sources with one sortable frame per physical file are admitted; a format whose
nested or repeated framing cannot be reordered losslessly fails during
planning with a correction to remove sort_order or normalize the input.
If a downstream consumer needs one global order across all files, declare
sort_order on the terminal Sink. Use enough output
fields to define a total business order when byte-identical output matters.
The shorthand form is also accepted – a bare string defaults to ascending:
sort_order:
- "account_id"
- { field: "txn_date", order: desc }
Watermarks
An event-time watermark declares which column on the source carries
each record’s event time — the wall-clock instant the event
happened, distinct from when Clinker read the row. When set, Clinker
takes the column on every record, subtracts the source’s delay, and
uses the result to track event-time progress so downstream time
windows know when to close. The delay-corrected value is also stamped
on every record as
$source.event_time, the column a
downstream time-windowed aggregate
uses to assign records to windows.
- type: source
name: clicks
config:
name: clicks
type: csv
path: "./data/clicks.csv"
options:
has_header: true
watermark:
column: event_ts # must be date_time or date
delay: 5s # bounded out-of-order tolerance
idle_timeout: 30s # flip partitions to idle if quiet
schema:
- { name: user_id, type: string }
- { name: event_ts, type: date_time }
- { name: amount, type: int }
Fields:
-
column(required) — the schema column whose value is each record’s event time. The column’s declared type must bedate_timeordate. Acolumn:that names a field absent fromschema:raises E154; acolumn:whose declared type is neither raises E155. -
delay(optional duration, default unset) — bounded out-of-order tolerance. Each record’s event time is shifted earlier bydelaybefore being folded into the watermark, so the source’s effective watermark trails its observed max event time by this amount. Mirrors Flink’sBoundedOutOfOrdernessWatermarks. Withoutdelay, the watermark advances strictly to the observed max — a single late record routes to the DLQ. -
idle_timeout(optional duration, default unset) — if a source stays quiet longer than this, it stops holding back downstream window-close progress, so windows keep closing when one source pauses. Unset means the source never goes idle.
Durations use the suffixes ms, s, m, h, d. ms is
matched before the single-character s, so 500ms reads as 500
milliseconds, not 500 seconds with a stray m.
A pipeline whose aggregate declares time_window: must have a
watermark.column on every upstream-reachable source. Without it,
event-time progress can never advance and the window can never close —
the planner rejects this with
E156.
Multi-value fields
A field that holds more than one value is declared on the schema column, not on the pipeline:
schema:
- { name: order_id, type: string }
- { name: tags, type: string, multiple: true }
multiple: true says the column holds zero or more values of its declared
type. Reading collects every occurrence of the field into one array — a single
occurrence is still an array, so downstream code never has to branch on how
many values happened to arrive. A field absent from a record has no column at
all and resolves to null, exactly as any other absent column does. CXL sees the
column as an array; the declared type: describes each element and drives
coercion. The declaration describes the shape of the
data, so it serves both directions: a writer that can encode repetition reads
the same declaration.
The split_to_rows, split_values, and join_values blocks below accept a
compact shorthand (a bare field name, or a mapping that omits defaults). To see
the fully-materialized form the engine actually runs — every default spelled
out — print the canonical config with
clinker config --resolved.
It rewrites only those shorthand blocks and leaves the rest of the file
untouched.
Both ends of the declaration are checked at compile, so a shape the formats cannot carry fails before a run starts rather than mid-stream:
| Format | As a source | As an output |
|---|---|---|
json | native — an array | native — an array |
xml | native — repeated child elements | not yet (issue 916) |
csv | delimited cell via split_values | delimited cell via join_values |
fixed_width | delimited cell via split_values | not yet (issue 918) |
edifact, x12, hl7, swift | no — repetition is positional | no — repetition is positional |
A multiple: true column reaching an output that cannot encode it is E359; one
on a source that cannot produce it is E361. E359 covers an output’s own
schema: block too — the attribute is direction-neutral, but the remaining
writers do not encode repetition yet, so declaring it on such a sink would be
accepted and ignored. Run clinker explain --code E361 for the full remediation
of either.
csv and fixed_width read a multiple: true column through a
split_values entry. Neither wire format repeats a field, but a cell’s text
may hold several values separated by a delimiter. Declare that delimiter with a
split_values entry and the reader
parses the cell into the array the column holds. A multiple: true column no
entry covers is rejected by E361 — the reader would have no delimiter and
deliver the raw cell; either add the entry or leave the column single-valued and
split it in a transform (tags.split(";")). The entry is read only on a
single-schema source: a multi-record source of either format runs a backend that
does not consume it. On the output side, a CSV sink joins a multiple: field
into one delimited cell with join_values
(defaulting to ; / on_conflict: error); fixed_width output is still pending
(#918).
A split_values entry also recovers a CSV cell a sink wrote under join_values
on_conflict: escape or encode_json: add escape: "\\" to un-escape an escaped
delimiter, or json: true to read the whole cell as an embedded JSON array.
The segment formats are a permanent no, not a pending one. Repetition there
is a positional coordinate rather than a list: a repeated composite is written
as two axes interleaved in one element (11:B:1^12:B:2), which a flat array
cannot represent without losing the component axis. The faithful shape is one
column per coordinate — for HL7, that is what options.split_fields produces,
with a writer that reassembles the wire field byte-for-byte.
One record per value: split_to_rows
split_to_rows fans a record out to one record per occurrence of a repeated
field. Each entry is either a bare field name or a full mapping, and the two
forms mix freely in one list:
- type: source
name: invoices
config:
name: invoices
type: json
path: "./data/invoices.json"
schema:
- { name: invoice_id, type: int }
- { name: customer, type: string }
- { name: line_item, type: string }
- { name: line_amount, type: float }
- { name: line_no, type: int }
split_to_rows:
- tags # shorthand: field name, all defaults
- field: line_items # full form
keep_empty: true
mode: extract
position_column: line_no
max_output_rows_per_input: 10000
| Key | Default | Meaning |
|---|---|---|
field | — | The repeated field, as a flattened dotted name |
keep_empty | true | Whether a record whose field is empty or absent survives |
mode | extract | extract — the occurrence becomes the record; split — the record shape is kept |
position_column | none | Column receiving each occurrence’s 1-based position |
The field is named as it appears in the input document, not as the schema
exposes it: a column declared source_name: is addressed by that
source_name. The same rule applies to split_values below.
keep_empty defaults to true. A record whose field holds an empty array,
or carries no such field at all, is emitted with that field unset rather than
disappearing. Several widely used engines drop the record instead; a vanished
row is the costliest failure mode there is, so dropping is opt-in here.
mode: extract (the default) makes the occurrence the record: its own
fields are lifted out from under the field name and every field outside the
group is merged onto each output. {"orders": [{"id": 1}]} yields a top-level
id, and repeated <Item><name> children yield name. When lifting lands an
occurrence’s field on a name an outside field already occupies, the occurrence
wins — it is the record, so its own value is not shadowed by the parent it was
merged with. A position_column wins over both: you named it, so a field of
the same name inside or outside the occurrence gives way to the index.
mode: split preserves the record shape: the occurrence’s fields keep
their dotted path (orders.id, Item.name) and each output carries exactly
one occurrence.
Entries apply in declaration order, so two entries multiply. Declaring the same
field twice is rejected at compile (E358), as is fanning out a field the
schema also declares multiple: true — the attribute collects the occurrences
into one array, the fan-out spends them one per record, and a field cannot be
both. On an XML source, two entries may not name nested element groups
either (Item and Item.part): that reader assigns each element to one
occurrence group by document position, and a nested pair leaves the inner
group’s membership ambiguous.
max_output_rows_per_input is valid only alongside a non-empty
split_to_rows block and bounds that cumulative product for one original
JSON object or XML record element. Omit it, or set it to 0, for no ceiling.
With a positive value N, the reader emits exactly the first N records in
the same declaration order shown above. If an N+1 row is attempted, the
reader stops that input and reports an expansion_limit_exceeded source
failure. Under error_handling.strategy: continue, the complete decoded
original input representation is routed to the DLQ; the first N records
remain valid output. This is not
silent truncation: the run records one explicit rejection naming the field,
the configured ceiling, and the exact first violating count (N+1). Under
fail_fast, the run fails at that boundary.
The check is lazy: the reader never materializes or pre-counts the Cartesian
product. Its cursor holds the original input, the declared occurrence lists,
and one output record. This source setting is separate from a Transform’s
max_expansion; when both surfaces
fan out, each enforces its own per-input boundary.
A JSON source accepts a nested pair to produce a two-level expansion, but
only when the outer entry declares mode: split. mode: extract lifts the
occurrence’s own keys to the top level, which removes the dotted path the inner
entry addresses — the inner entry would then match nothing and fan nothing out,
so the pairing is rejected (E358):
split_to_rows:
- { field: orders, mode: split }
- { field: orders.items, mode: split }
Several values in one cell: split_values
split_values parses a delimited cell into the several values a multiple:
column holds. It takes the same bare-name-or-mapping shorthand:
split_values:
- tags # shorthand: default delimiter `;`
- field: codes # full form
delimiter: "|"
schema:
- { name: tags, type: string, multiple: true }
- { name: codes, type: string, multiple: true }
The delimiter defaults to ;. A split_values field the schema does not
declare multiple: true is rejected at compile (E358): splitting produces
several values, and only a multi-value column can hold them. So is an entry
naming a column’s exposed name when that column reads a differently-named input
field — the split runs against the document’s own field names, so name the
source_name.
The entry is read by the JSON and XML readers, and — on a single-schema
source — by the CSV and fixed-width readers. On a multi-record CSV or
fixed-width source, or on any segment format, declaring it is rejected (E358)
rather than silently ignored: those readers are never handed it, so the cell
would arrive unsplit with nothing to say so.
Migrating from array_paths
array_paths: was the earlier form of these declarations. It is no longer read,
and a source still carrying it is rejected at compile (E360) rather than
running with the fan-out silently dropped. An explode path becomes a
split_to_rows: entry, a delimited cell becomes a split_values: entry, and a
path kept as an array becomes multiple: true on the schema column.
mode: extract (the default) reproduces the old projection — the element’s
own fields lifted to the top level of each output record. It does not reproduce
the old cardinality: explode dropped a record whose array was empty, while
keep_empty defaults to true here and keeps it with the element’s fields
unset. Add keep_empty: false to the entry to migrate row-for-row.
Format notes
split_to_rows is honored by the JSON and XML readers, over a file path
and over a rest response body alike. split_values is honored by those two and
also by the single-schema CSV and fixed-width readers. Declaring a knob on
a reader that is never handed it — split_to_rows on any delimited-cell or
segment format, split_values on a multi-record or segment source — is rejected
at compile (E358) rather than accepted and inert, and a multiple: true column
no split_values entry can cover is rejected by E361.
Composition body files are gated by the same four checks as the pipeline that calls them.
JSON — the field names the key holding the array (line_items, or
order.line_items for an array nested under an object). A field present but
holding a single object or scalar rather than an array counts as one
occurrence, and is projected exactly as a one-element array would be — many
producers unwrap a lone element, and XML cannot express the difference at all,
so the two readers agree on this. A field that is absent, holds an empty array,
or is explicitly null has no occurrence and is governed by keep_empty: an
explicit null is how many producers write “no value”, and it counts as none
rather than as one. For the same reason a multiple: true column holding an
explicit null stays null rather than becoming [null], so size() over it
reads the same as it does for a field the document omits.
XML — the field is the repeated child element’s dotted path relative to the
record element (Item, or Items.Item when nested). Repetition and absence
are indistinguishable in XML, so a record with no occurrence of the element is
governed by keep_empty exactly as an empty array is. See
XML Format for the full rules.