Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Source Nodes

Source nodes read data from files and are the entry points of every pipeline. They have no input: field – they produce records, they do not consume them.

Basic structure

- type: source
  name: customers
  config:
    name: customers
    type: csv
    path: "./data/customers.csv"
    schema:
      - { name: customer_id, type: int }
      - { name: name, type: string }
      - { name: email, type: string }
      - { name: status, type: string }
      - { name: amount, type: float }

Schema declaration

The schema: field is required on every source node. Runtime ingestion does not guess types from data: declare each column’s name and CXL type explicitly. This schema drives compile-time type checking across the entire pipeline.

Each entry is a { name, type } pair:

schema:
  - { name: employee_id, type: string }
  - { name: salary, type: int }
  - { name: hired_at, type: date_time }
  - { name: is_active, type: bool }
  - { name: notes, type: { nullable: string } }

Available types

TypeDescription
nullThe null value only
stringUTF-8 text
int64-bit signed integer
float64-bit IEEE 754 floating point
decimalExact fixed-point number, subject to the declared precision and scale
boolBoolean (true / false)
dateCalendar date
date_timeDate with time component
arrayOrdered sequence of values
mapString-keyed object
anyUnknown type – field used in type-agnostic contexts
{ nullable: T }Nullable wrapper around a concrete inner type, for example { nullable: int }

Check the schema through the planner

The schema that governs execution is the one accepted by clinker-plan while compiling the complete pipeline. Run this after changing a source schema:

clinker run pipeline.yaml --explain text

Exit code 0 means the planner parsed the canonical YAML, resolved schema references and overlays, bound the pipeline, type-checked CXL, and produced a compiled plan. It does not prove that input bytes satisfy the declarations or that the resulting output is correct; execute representative input against an isolated destination and compare the result before a production run.

Workspace tooling may also present advisory schema analysis. That analysis can help find likely authoring mistakes, but it cannot admit or reject a pipeline. Its analyzed, partial, skipped, and failed statuses describe only how much the advisory model inspected. See Validation and Admission for the status meanings and known limits.

Without a column format:, date_time accepts the default offset-free forms and RFC 3339 timestamps with Z or an explicit numeric offset, such as 2026-01-31T08:27:00Z and 2026-01-31T10:57:00+02:30. Zoned timestamps are normalized to UTC before entering Clinker’s timezone-free date_time representation. Parsing is exact: surrounding whitespace, malformed calendar values, and out-of-range timestamps are rejected rather than trimmed, guessed, overflowed, or rounded. A column-level strftime format: is exclusive: only that authored format is tried.

A source column’s declared type must be concrete: numeric — the inference-only int | float union CXL resolves during type unification — is not a valid source column type. Declaring one is rejected at compile with E158; declare int or float explicitly.

clinker guess pipeline.yaml provides an authoring-only preview for columns explicitly marked numeric. It proposes concrete int or float declarations; it does not infer arbitrary schemas or change the runtime admission rules. The default preview is bounded and read-only. --check exhausts the admitted manifest; --write can publish one guarded, unambiguous inline edit. Resolve all numeric declarations to concrete types before executing the pipeline. See Numeric schema authoring and the guess command.

long_unique — storage hint for high-cardinality text

A string column may carry an optional long_unique: true flag. It is an advisory, opt-in hint, not a type change: it tells Clinker the column’s values are long and effectively unique — never repeated across records — so the run uses less memory for that column. Typical candidates are UUIDs rendered as text, street addresses, and free-text comment or note fields.

schema:
  - { name: ticket_id,  type: string, long_unique: true }   # 36-char UUID
  - { name: notes,      type: string, long_unique: true }   # free text
  - { name: department, type: string }                      # low-cardinality, default

The flag lowers memory use only. A value’s content and its comparison, grouping, join, sort, and output behavior are all unchanged — a long_unique value behaves identically to the same text in any other column. Omitting the flag (the common case) leaves the default behavior untouched. Set it only when you know a column is genuinely high-cardinality free text; on a column whose values repeat, leave it off.

source_name — read a differently-named physical column

A column may carry an optional source_name naming the physical input column it reads from, when that differs from the exposed name. The reader matches input fields by physical name and re-labels the value under name, so downstream CXL and the output see name carrying the physical column’s data.

schema:
  # read the physical `cust_id` column, expose it downstream as `customer_id`
  - { name: customer_id, type: string, source_name: cust_id }

Omitting source_name (the common case) reads the input field whose key equals name, unchanged from before. A channel schema patch’s rename op sets this alias automatically (see Channels).

Transport vs format

A source declaration has two independent layers:

  • Transport (transport:) selects where the records come from. Two transports exist: file — read bytes from the filesystem, resolved through one of the file matchers (path / glob / regex / paths) — and rest — pull records from a paginated HTTP endpoint under a hard page/record cap (see Network Sources (REST)). transport: is optional and defaults to file, so a source that omits it reads from disk exactly as before. rest needs the rest capability, which the released binary has; a build compiled without it refuses a rest source at validation with E223 rather than running the pipeline short one input (see Optional capabilities).
  • Format (type:) selects how the bytes decode into records: csv, json, xml, fixed_width, edifact, x12, hl7, swift.
- type: source
  name: orders
  config:
    name: orders
    transport: file        # optional; this is the default
    type: csv              # the on-disk format
    path: "./data/orders.csv"
    schema:
      - { name: order_id, type: int }

A file transport requires exactly one file matcher (path, glob, regex, or paths). Declaring none fails validation with E211; declaring more than one fails with E210. Both are reported at config-load time, before any file is opened.

Choosing files

Interactive companion: the file discovery explainer runs these steps on a sample folder as you change the settings.

Before reading, a file Source builds its list of files in a fixed order:

  1. Matcher. Paths are relative to the pipeline file’s folder.
    • path: names one file and paths: a list. A file that does not exist stops the run with E216, whatever on_no_match says.
    • glob: matches names inside one folder: * does not cross a /. Use ** to include subfolders (./data/**/*.csv). A leading dot is not special, so *.csv also matches .draft.csv. An invalid pattern is E212.
    • regex: searches the pipeline’s folder and, by default, every subfolder, and matches anywhere in each file’s path. That path begins with the folder part of the pipeline path as you typed it (clinker run pipelines/orders.yaml gives pipelines/data/orders_2024-01.csv), so anchor the end of the path, for example data/orders_\d{4}-\d{2}\.csv$. An unanchored pattern can pick up earlier outputs too. An invalid pattern is E213.
  2. exclude: – a list of glob patterns; a file whose name or full path matches any of them is dropped.
  3. Anything that is not a regular file is dropped.
  4. min_size: / max_size: – decimal units: 1KB is 1000 bytes, also B, MB, GB.
  5. modified_after: / modified_before: – a duration back from the time the pipeline is loaded (30s, 15m, 2h, 3d) or an RFC 3339 timestamp (2024-03-01T00:00:00Z).
  6. files.sort_by: name (default; the full path), created or modified, with files.sort_order: asc (default) or desc. The files of a paths: list are sorted too: the order they are written in is not the reading order.
  7. files.take_first: or files.take_last: keeps that many files from the sorted list (setting both is E218). With the default ascending name sort, take_first: 5 keeps the five earliest names.
  8. files.on_no_match: – when nothing is left: error (default, E216), warn (log a warning and produce no rows), or skip (produce no rows quietly).

files.recursive: controls whether regex: searches subfolders (true by default). It has no effect on glob:, which searches subfolders only where the pattern has **.

- type: source
  name: orders
  config:
    name: orders
    type: csv
    glob: ./data/orders_*.csv
    exclude: ["*_partial.csv"]
    modified_after: 30d
    files:
      sort_by: name
      sort_order: desc
      take_first: 3        # the three latest names
      on_no_match: warn
    schema:
      - { name: order_id, type: string }

With glob:, regex: or paths:, each file is read as its own document: a declared sort_order is checked per file, and an Aggregate rolls up per file (see Sort order and One document per file).

Format types

The type: field inside config: selects the on-disk format. Each format has its own reference page covering its options and decoding model:

type:FormatReference
csvDelimited text (RFC 4180)CSV Format
jsonArray / NDJSON / wrapper objectJSON Format
xmlElement-path-selected record elementsXML Format
fixed_widthColumn-positioned legacy extractsFixed-Width Format
edifactUN/EDIFACT interchangesEDIFACT Format
x12ANSI ASC X12 interchangesX12 Format
hl7HL7 v2.x pipe-and-hat messagesHL7 v2 Format
swiftSWIFT MT (FIN) messagesSWIFT MT Format

The same schema: rules apply regardless of format: the reader maps each decoded record onto the declared schema, and undeclared input fields fall under the on_unmapped policy below.

Declared-type failures

Source types are enforced before a record reaches buffering, sorting, or any downstream node. A value that cannot satisfy its authored type rejects the whole row exactly once; Clinker never substitutes the raw string, a sentinel, or an error-derived null.

Empty input has three distinct outcomes:

  • an empty value declared as string remains the empty string;
  • an empty non-string value declared nullable(T) becomes null;
  • an empty non-string value declared as non-nullable is a type error.

Parsing does not trim whitespace, guess locale conventions, or recognize case-insensitive null sentinels. Integer overflow, invalid/out-of-range dates, decimal precision overflow, and decimal values that would need rounding to the declared scale are type errors. Accepted decimal and date values retain their declared precision.

The E126 diagnostic identifies source, file, one-based row and column, field, and declared type. Its value preview is sanitized to one line and limited to 256 rendered UTF-8 bytes; controls, bidi characters, diagnostic delimiters, backslashes, and invalid UTF-8 are explicit indivisible escape tokens. When truncated, the preview ends in one … without splitting a token or Unicode scalar and reports the original byte length. The complete original record/value is retained only in the configured DLQ. See Error Handling & DLQ for strategy and threshold behavior.

Already-decoded values obey the same declarations without hidden coercions: string admits only a string, null admits only null, and any admits every supported native value while preserving its represented value. Numeric conversions must be exact; non-finite floating-point values and decimal values that would require rounding are rejected. With multiple: true, the scalar declaration is applied to every array element. If any element fails, the whole source record is rejected and the complete original array remains available through the DLQ.

Error strategies and complete populations

fail_fast aborts on the first declared source failure. continue and best_effort require a DLQ and use the configured threshold over the complete population:

rejected records / attempted records

The run aborts only when that ratio is strictly greater than the threshold. Equality is accepted, so a threshold of 0.1 admits exactly one rejected record in a population of ten but not two. A zero threshold aborts on the first rejection; a threshold of 1.0 admits an all-rejected population for the continuing strategies.

For ordered file sources, Clinker establishes the complete attempted and rejected population before any accepted record, punctuation, downstream side effect, or output byte is released. That population is applied exactly once whether execution remains resident, spills, or fuses a downstream node. A threshold violation therefore cannot leave a committed prefix of output.

on_unmapped — undeclared input fields

The per-source on_unmapped policy decides what to do with input fields the source’s schema: block does not name. Three modes — auto_widen (default), drop, reject:

- type: source
  name: orders
  config:
    name: orders
    type: csv
    path: "./data/orders.csv"
    on_unmapped:
      mode: auto_widen     # default; other values: drop, reject
    schema:
      - { name: order_id, type: string }
      - { name: amount, type: float }

See Auto-Widen & Schema Drift for the full specification: how undeclared columns flow through each downstream node type, the include_unmapped Sink flag, E315 merge-policy mismatch, and fixed-width behavior.

Sort order

If each physical input file is pre-sorted, declare the record order so the planner can admit order-dependent strategies such as streaming aggregation:

- type: source
  name: sorted_transactions
  config:
    name: sorted_transactions
    type: csv
    path: "./data/transactions_sorted.csv"
    schema:
      - { name: account_id, type: string }
      - { name: txn_date, type: date }
      - { name: amount, type: float }
    sort_order:
      - { field: "account_id", order: asc }
      - { field: "txn_date", order: asc }

Clinker binds these fields to the declared source schema, compares the typed values by the same rule every sort uses (see How values are ordered), and verifies each physical file independently before any record from that file reaches an order-dependent consumer. The declaration never means that a multi-file source is globally sorted: the last key in one file is not compared with the first key in the next file.

on_unsorted controls the result of the first adjacent inversion:

ValueBehavior
warn (default)Stably repair the complete physical file with the shared bounded-memory sort, emit one W307 warning for that file, then release it. A file already in order emits no warning.
errorReject the physical file without releasing an unverified prefix. The diagnostic identifies the source, file, adjacent rows, and keys.
    sort_order:
      - { field: "account_id", order: asc, null_order: last }
      - { field: "txn_date", order: desc, null_order: first }
    on_unsorted: warn

Source ordering accepts null_order: first or last; drop is rejected when the pipeline is planned because verifying order must not discard source records. The error gives one fix: delete null_order: drop and add a Transform after the Source whose whole config is the line it prints, config: { cxl: "filter not <field>.is_null()" }. With the line deleted, the Source declares its null keys last; if a file’s null keys arrive first, write null_order: first instead of deleting the line. That filter needs a field CXL can name as it is: one identifier of ASCII letters, digits and _, not starting with a digit and not a CXL keyword. For any other key, such as order id, filter or a flattened Address.City, the error prints no CXL and prints a source_name line for the column instead, such as source_name: "order id". Set the column’s name to a new identifier, add that line, use the new name wherever the pipeline names the column, and plan again: the error then prints the filter on the new name. Equal authored keys retain arrival order within the selected execution path. Clinker does not add a source identity, physical filename, or canonical-row tie-breaker.

Verification stages the complete sortable file event sequence behind the run’s memory arbitrator. It uses the existing stable resident/spill sort and bounded-fan-in merge machinery when repair spills, while preserving row identity, source/file provenance, and document context. Only flat sources and sources with one sortable frame per physical file are admitted; a format whose nested or repeated framing cannot be reordered losslessly fails during planning with a correction to remove sort_order or normalize the input.

If a downstream consumer needs one global order across all files, declare sort_order on the terminal Sink. Use enough output fields to define a total business order when byte-identical output matters.

The shorthand form is also accepted – a bare string defaults to ascending:

    sort_order:
      - "account_id"
      - { field: "txn_date", order: desc }

Watermarks

An event-time watermark declares which column on the source carries each record’s event time — the wall-clock instant the event happened, distinct from when Clinker read the row. When set, Clinker takes the column on every record, subtracts the source’s delay, and uses the result to track event-time progress so downstream time windows know when to close. The delay-corrected value is also stamped on every record as $source.event_time, the column a downstream time-windowed aggregate uses to assign records to windows.

- type: source
  name: clicks
  config:
    name: clicks
    type: csv
    path: "./data/clicks.csv"
    options:
      has_header: true
    watermark:
      column: event_ts       # must be date_time or date
      delay: 5s              # bounded out-of-order tolerance
      idle_timeout: 30s      # flip partitions to idle if quiet
    schema:
      - { name: user_id, type: string }
      - { name: event_ts, type: date_time }
      - { name: amount, type: int }

Fields:

  • column (required) — the schema column whose value is each record’s event time. The column’s declared type must be date_time or date. A column: that names a field absent from schema: raises E154; a column: whose declared type is neither raises E155.

  • delay (optional duration, default unset) — bounded out-of-order tolerance. Each record’s event time is shifted earlier by delay before being folded into the watermark, so the source’s effective watermark trails its observed max event time by this amount. Mirrors Flink’s BoundedOutOfOrdernessWatermarks. Without delay, the watermark advances strictly to the observed max — a single late record routes to the DLQ.

  • idle_timeout (optional duration, default unset) — if a source stays quiet longer than this, it stops holding back downstream window-close progress, so windows keep closing when one source pauses. Unset means the source never goes idle.

Durations use the suffixes ms, s, m, h, d. ms is matched before the single-character s, so 500ms reads as 500 milliseconds, not 500 seconds with a stray m.

A pipeline whose aggregate declares time_window: must have a watermark.column on every upstream-reachable source. Without it, event-time progress can never advance and the window can never close — the planner rejects this with E156.

Multi-value fields

A field that holds more than one value is declared on the schema column, not on the pipeline:

schema:
  - { name: order_id, type: string }
  - { name: tags, type: string, multiple: true }

multiple: true says the column holds zero or more values of its declared type. Reading collects every occurrence of the field into one array — a single occurrence is still an array, so downstream code never has to branch on how many values happened to arrive. A field absent from a record has no column at all and resolves to null, exactly as any other absent column does. CXL sees the column as an array; the declared type: describes each element and drives coercion. The declaration describes the shape of the data, so it serves both directions: a writer that can encode repetition reads the same declaration.

The split_to_rows, split_values, and join_values blocks below accept a compact shorthand (a bare field name, or a mapping that omits defaults). To see the fully-materialized form the engine actually runs — every default spelled out — print the canonical config with clinker config --resolved. It rewrites only those shorthand blocks and leaves the rest of the file untouched.

Both ends of the declaration are checked at compile, so a shape the formats cannot carry fails before a run starts rather than mid-stream:

FormatAs a sourceAs an output
jsonnative — an arraynative — an array
xmlnative — repeated child elementsnot yet (issue 916)
csvdelimited cell via split_valuesdelimited cell via join_values
fixed_widthdelimited cell via split_valuesnot yet (issue 918)
edifact, x12, hl7, swiftno — repetition is positionalno — repetition is positional

A multiple: true column reaching an output that cannot encode it is E359; one on a source that cannot produce it is E361. E359 covers an output’s own schema: block too — the attribute is direction-neutral, but the remaining writers do not encode repetition yet, so declaring it on such a sink would be accepted and ignored. Run clinker explain --code E361 for the full remediation of either.

csv and fixed_width read a multiple: true column through a split_values entry. Neither wire format repeats a field, but a cell’s text may hold several values separated by a delimiter. Declare that delimiter with a split_values entry and the reader parses the cell into the array the column holds. A multiple: true column no entry covers is rejected by E361 — the reader would have no delimiter and deliver the raw cell; either add the entry or leave the column single-valued and split it in a transform (tags.split(";")). The entry is read only on a single-schema source: a multi-record source of either format runs a backend that does not consume it. On the output side, a CSV sink joins a multiple: field into one delimited cell with join_values (defaulting to ; / on_conflict: error); fixed_width output is still pending (#918).

A split_values entry also recovers a CSV cell a sink wrote under join_values on_conflict: escape or encode_json: add escape: "\\" to un-escape an escaped delimiter, or json: true to read the whole cell as an embedded JSON array.

The segment formats are a permanent no, not a pending one. Repetition there is a positional coordinate rather than a list: a repeated composite is written as two axes interleaved in one element (11:B:1^12:B:2), which a flat array cannot represent without losing the component axis. The faithful shape is one column per coordinate — for HL7, that is what options.split_fields produces, with a writer that reassembles the wire field byte-for-byte.

One record per value: split_to_rows

split_to_rows fans a record out to one record per occurrence of a repeated field. Each entry is either a bare field name or a full mapping, and the two forms mix freely in one list:

- type: source
  name: invoices
  config:
    name: invoices
    type: json
    path: "./data/invoices.json"
    schema:
      - { name: invoice_id, type: int }
      - { name: customer, type: string }
      - { name: line_item, type: string }
      - { name: line_amount, type: float }
      - { name: line_no, type: int }
    split_to_rows:
      - tags                      # shorthand: field name, all defaults
      - field: line_items         # full form
        keep_empty: true
        mode: extract
        position_column: line_no
    max_output_rows_per_input: 10000
KeyDefaultMeaning
field—The repeated field, as a flattened dotted name
keep_emptytrueWhether a record whose field is empty or absent survives
modeextractextract — the occurrence becomes the record; split — the record shape is kept
position_columnnoneColumn receiving each occurrence’s 1-based position

The field is named as it appears in the input document, not as the schema exposes it: a column declared source_name: is addressed by that source_name. The same rule applies to split_values below.

keep_empty defaults to true. A record whose field holds an empty array, or carries no such field at all, is emitted with that field unset rather than disappearing. Several widely used engines drop the record instead; a vanished row is the costliest failure mode there is, so dropping is opt-in here.

mode: extract (the default) makes the occurrence the record: its own fields are lifted out from under the field name and every field outside the group is merged onto each output. {"orders": [{"id": 1}]} yields a top-level id, and repeated <Item><name> children yield name. When lifting lands an occurrence’s field on a name an outside field already occupies, the occurrence wins — it is the record, so its own value is not shadowed by the parent it was merged with. A position_column wins over both: you named it, so a field of the same name inside or outside the occurrence gives way to the index.

mode: split preserves the record shape: the occurrence’s fields keep their dotted path (orders.id, Item.name) and each output carries exactly one occurrence.

Entries apply in declaration order, so two entries multiply. Declaring the same field twice is rejected at compile (E358), as is fanning out a field the schema also declares multiple: true — the attribute collects the occurrences into one array, the fan-out spends them one per record, and a field cannot be both. On an XML source, two entries may not name nested element groups either (Item and Item.part): that reader assigns each element to one occurrence group by document position, and a nested pair leaves the inner group’s membership ambiguous.

max_output_rows_per_input is valid only alongside a non-empty split_to_rows block and bounds that cumulative product for one original JSON object or XML record element. Omit it, or set it to 0, for no ceiling. With a positive value N, the reader emits exactly the first N records in the same declaration order shown above. If an N+1 row is attempted, the reader stops that input and reports an expansion_limit_exceeded source failure. Under error_handling.strategy: continue, the complete decoded original input representation is routed to the DLQ; the first N records remain valid output. This is not silent truncation: the run records one explicit rejection naming the field, the configured ceiling, and the exact first violating count (N+1). Under fail_fast, the run fails at that boundary.

The check is lazy: the reader never materializes or pre-counts the Cartesian product. Its cursor holds the original input, the declared occurrence lists, and one output record. This source setting is separate from a Transform’s max_expansion; when both surfaces fan out, each enforces its own per-input boundary.

A JSON source accepts a nested pair to produce a two-level expansion, but only when the outer entry declares mode: split. mode: extract lifts the occurrence’s own keys to the top level, which removes the dotted path the inner entry addresses — the inner entry would then match nothing and fan nothing out, so the pairing is rejected (E358):

    split_to_rows:
      - { field: orders, mode: split }
      - { field: orders.items, mode: split }

Several values in one cell: split_values

split_values parses a delimited cell into the several values a multiple: column holds. It takes the same bare-name-or-mapping shorthand:

    split_values:
      - tags                      # shorthand: default delimiter `;`
      - field: codes              # full form
        delimiter: "|"
    schema:
      - { name: tags,  type: string, multiple: true }
      - { name: codes, type: string, multiple: true }

The delimiter defaults to ;. A split_values field the schema does not declare multiple: true is rejected at compile (E358): splitting produces several values, and only a multi-value column can hold them. So is an entry naming a column’s exposed name when that column reads a differently-named input field — the split runs against the document’s own field names, so name the source_name.

The entry is read by the JSON and XML readers, and — on a single-schema source — by the CSV and fixed-width readers. On a multi-record CSV or fixed-width source, or on any segment format, declaring it is rejected (E358) rather than silently ignored: those readers are never handed it, so the cell would arrive unsplit with nothing to say so.

Migrating from array_paths

array_paths: was the earlier form of these declarations. It is no longer read, and a source still carrying it is rejected at compile (E360) rather than running with the fan-out silently dropped. An explode path becomes a split_to_rows: entry, a delimited cell becomes a split_values: entry, and a path kept as an array becomes multiple: true on the schema column.

mode: extract (the default) reproduces the old projection — the element’s own fields lifted to the top level of each output record. It does not reproduce the old cardinality: explode dropped a record whose array was empty, while keep_empty defaults to true here and keeps it with the element’s fields unset. Add keep_empty: false to the entry to migrate row-for-row.

Format notes

split_to_rows is honored by the JSON and XML readers, over a file path and over a rest response body alike. split_values is honored by those two and also by the single-schema CSV and fixed-width readers. Declaring a knob on a reader that is never handed it — split_to_rows on any delimited-cell or segment format, split_values on a multi-record or segment source — is rejected at compile (E358) rather than accepted and inert, and a multiple: true column no split_values entry can cover is rejected by E361.

Composition body files are gated by the same four checks as the pipeline that calls them.

JSON — the field names the key holding the array (line_items, or order.line_items for an array nested under an object). A field present but holding a single object or scalar rather than an array counts as one occurrence, and is projected exactly as a one-element array would be — many producers unwrap a lone element, and XML cannot express the difference at all, so the two readers agree on this. A field that is absent, holds an empty array, or is explicitly null has no occurrence and is governed by keep_empty: an explicit null is how many producers write “no value”, and it counts as none rather than as one. For the same reason a multiple: true column holding an explicit null stays null rather than becoming [null], so size() over it reads the same as it does for a field the document omits.

XML — the field is the repeated child element’s dotted path relative to the record element (Item, or Items.Item when nested). Repetition and absence are indistinguishable in XML, so a record with no occurrence of the element is governed by keep_empty exactly as an empty array is. See XML Format for the full rules.