Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

XML Format

The XML reader selects record elements by a slash-separated path of element names and maps each one onto the source’s declared schema:. Child elements bind to fields by name; attributes bind under a configurable prefix. Namespaces are stripped by default so schema field names stay clean. See Source Nodes for the shared schema and transport rules.

- type: source
  name: catalog
  config:
    name: catalog
    type: xml
    path: "./data/catalog.xml"
    schema:
      - { name: product_id, type: int }
      - { name: name, type: string }
      - { name: price, type: float }
    options:
      record_path: "catalog/product"    # slash-separated element path
      attribute_prefix: "@"             # prefix for XML attribute fields
      namespace_handling: strip         # strip | qualify
      max_index_bytes: 64MB             # cap on retained envelope sections (optional)

Text encoding

XML input must contain valid UTF-8. One leading UTF-8 BOM is removed on every physical file open, including the envelope pre-scan. UTF-16/32 BOMs are rejected. An XML declaration may omit encoding or declare UTF-8; other encodings and conflicting declarations are rejected. Convert such input to UTF-8 rather than adding an encoding option. Names, attributes, text and CDATA are never decoded with replacement characters.

Validation applies to bytes the reader consumes. A pre-scan may find a late error before any body record is delivered; a streaming body can have already delivered earlier records. Each subsequent file establishes its own BOM and declaration policy. Metadata adjacent to a selected record does not become an extra row, and repeated matching containers preserve body order and empty rows.

Output is UTF-8 without a BOM or XML declaration. See native document boundaries for envelope and empty-output behavior.

Options

OptionDefaultDescription
record_path—Slash-separated path of element names selecting the elements that each become one record — see record_path. Omitted, every top-level element becomes one record.
attribute_prefix@Prefix that distinguishes an element’s attributes from its child elements when both map to schema fields.
namespace_handlingstripstrip removes namespace prefixes from element and attribute names; qualify preserves the namespace-qualified names.
max_index_bytes64MBCap on the bytes the envelope pre-scan retains while extracting declared $doc.* sections.

record_path

record_path is a slash-separated path of XML element names, matched level by level starting at the document element. catalog/product selects every <product> that is a child of the document element <catalog>. This is the canonical statement of the grammar; other pages link here rather than restate it.

The rules, in full:

  • The path is already anchored at the document element, so it carries no leading /. Write Orders/Order, not /Orders/Order.
  • No //. It is not XPath: there is no descendant-or-any-depth step. Name every enclosing element.
  • No empty segments — no doubled separator (Orders//Order) and no trailing one (Orders/).
  • No XPath predicates, axes, or wildcards (product[@id='7'], child::product, *). Select the elements by path and filter the records in a transform.
  • Every segment must be a legal XML element name. Under namespace_handling: qualify element names keep their prefix, so a qualified segment (ns:Order) is allowed and is what matches; under the default strip the prefix is gone and the segment is the local name.
  • Omitting record_path entirely makes every top-level element one record. That is not the same as record_path: "", which is a path naming an element called “” and is rejected.

A value breaking any of these fails at compile time with E363, before any input is opened. The diagnostic names the corrected path where one can be derived.

record_path and xml_path root differently

The envelope option extract: { xml_path: … } is also a slash-path over XML, but it tolerates a leading / — /doc/Head is its documented form. record_path rejects one.

The two are separate grammars addressing separate things: xml_path locates a single envelope section anywhere in the document, record_path locates the record elements the body streams. They are deliberately not aligned — writing record_path: "/catalog/product" is an error, and writing xml_path: "/doc/Head" is correct.

Truncated input

A truncated XML document — one whose input ends before an open element’s closing tag — is rejected with a format error rather than yielding the partial fields read so far. This holds for a record cut off mid-element, a skipped-over sibling subtree cut off before it closes, and an envelope section cut off during the pre-scan (which then attaches no $doc metadata). This matches the general contract that a truncated stream always aborts rather than silently dropping data.

Writing XML

The XML writer expands dotted field names to nested elements, by the same rule the JSON writer expands them into nested objects — grouping, ordering, absent-child pruning, and the \. escape for a literal dot are all specified once on Field Paths. What is specific to XML is layered on top of that decoding, not instead of it.

The attribute_prefix convention applies in reverse: a field whose final path segment carries the prefix is emitted as an XML attribute of its enclosing element instead of a child element. A top-level @id attaches to the record element’s start tag; a nested Address.@type attaches to the <Address> element. Records read from an XML source therefore round-trip — <Record id="7"><name>A</name></Record> reads and writes back unchanged, and the writer never emits an @-named element.

Each decoded segment must also be a well-formed XML Name, so a segment that begins with a digit or contains a space is rejected. A literal dot survives — . is a legal XML name character, so a column declared a\.b emits the single element <a.b> rather than nesting.

Two column names that cannot both be expanded — a column a holding a value alongside a column a.b needing a to be a container — are refused before any byte of the record is written, naming both columns. Earlier versions emitted two sibling <a> elements for that column set, which this reader then refused on the way back in.

- type: sink
  name: xml_out
  input: processed
  config:
    name: xml_out
    type: xml
    path: "./output/result.xml"
    preserve_nulls: false              # omit null elements; null attributes always omit
    options:
      root_element: "Root"              # default Root
      record_element: "Record"          # default Record
      attribute_prefix: "@"             # matches the source-side prefix
OptionDefaultDescription
root_elementRootName of the document root element wrapping all records.
record_elementRecordName of the element emitted per record.
attribute_prefix@Prefix marking a field as an attribute of its enclosing element. Set it to the same value as the source-side prefix when round-tripping; an empty string disables attribute classification (every field emits as an element).

Attribute handling details:

  • A null attribute field is dropped even under preserve_nulls: true — a null element round-trips as a self-closing tag, but an attribute has no form that reads back as null.
  • A field with children nested under an attribute-prefixed segment (e.g. @a.b) is rejected with a format error: an XML attribute is a leaf and cannot contain elements.
  • The attribute name (the segment after the prefix) must be a well-formed XML name — a letter, _, or : followed by letters, digits, _, -, ., or : (plus the XML 1.0 Unicode name ranges). A name with a space, =, quote, /, >, or a leading digit (e.g. @foo bar, @1st) is rejected with a format error rather than emitting a malformed start tag. Non-ASCII letters are accepted, so an attribute name read from a source document round-trips unchanged.
  • An element with only attribute fields and no children self-closes: Address.@type alone emits <Address type="home"/>.

Native map and array values

An element-valued CXL map is written recursively. Ordinary keys become child elements, an unescaped key beginning with attribute_prefix becomes an attribute on the current element, and the unescaped key #text becomes text in the current element. Map insertion order controls text/child order; attributes are collected onto the start tag. Arrays held under an ordinary key repeat that key as the element name.

emit payload = {
  "@kind": "event",
  "#text": "before",
  item: [
    {"@id": 1, "#text": "alpha"},
    {"@id": 2, "#text": "beta"},
  ],
  tail: "after",
}

writes:

<payload kind="event">before<item id="1">alpha</item><item id="2">beta</item><tail>after</tail></payload>

JSON and XML share one neutral-map key grammar. After CXL string decoding, ordinary keys are unescaped. Exactly one leading backslash marks a literal key only as \@name, \#text, or \\name, and the neutral decoder removes that one marker. Because the CXL string literal must encode the backslash too, the source spellings are "\\@name", "\\#text", and "\\\\name". Other leading-backslash forms are non-canonical and fail. An escape disables XML’s attribute or text classification; it does not make the decoded spelling a legal XML name. For example, a decoded @literal still cannot be an element name, while JSON can write it as an ordinary object key.

The rules are deliberately strict:

  • Attribute and #text values must be scalar or null; maps and arrays there are rejected.
  • A direct array inside another array is rejected because XML has no child name to repeat. Put the inner array under a map key to supply that name.
  • Every decoded ordinary key and attribute name must be a well-formed XML name.
  • Static and computed keys use identical decoding. Two spellings that decode to the same logical key—including attribute-looking or #text spellings—are a collision and reject rather than selecting a winner.
  • Scalars have depth zero and each map or array adds one container. Depth 64 is accepted; depth 65, malformed escapes, duplicate logical keys, and invalid names reject the whole record before its first byte is emitted. XML never silently falls back to JSON text.

With the default preserve_nulls: false, null child elements and null array items are omitted; with it enabled they emit as self-closing elements. Null attributes are always omitted.

The XML writer is deliberately two-pass per record. Its first borrowed pass validates the complete schema/value shape, XML names, and scalar roles before writing any bytes for that record. Its second pass encodes from the borrowed original record into a private prepared operation. Delivery begins only when the complete operation is ready. Authored strings remain borrowed; other scalars are formatted in a fixed 128-byte stack scratch buffer. The writer does not clone or materialize a second nested tree, and it retains no record values or rendered scalar capacity between calls. Its only memoized preparation heap state is the admitted schema-derived element plan, whose size is independent of record value widths; recursive calls are capped at 64 containers. This is separate from the XML reader’s optional envelope pre-scan described below.

Native recursion belongs to XML and JSON/NDJSON. Flat, positional, and message formats do not stringify maps or arrays implicitly. Declare an encoding that the destination supports—such as join_values for a multi-value flat field—or reshape the value before that output; otherwise the structured value is rejected before bytes for that record are written.

Writing multi-value fields (repeated elements)

A multiple: field is written as repeated child elements, one per value, in order — the XML counterpart to the CSV writer’s delimited join_values cell, and the write-side inverse of reading multiple: true. The default needs no configuration:

<Order><id>1</id><tags>a</tags><tags>b</tags></Order>

The planner carries the exact output-facing multiple: true column set through mapping and exclusion into the writer. A top-level array repeats only for one of those columns; an array reaching any other XML column is rejected as a routing/type-contract error rather than treated as an implicit declaration. Arrays nested inside a map remain part of XML’s native recursive structure. A field with one value emits exactly one element, byte-identical to a scalar field’s output; a field with an empty array emits nothing (no element, and no container even when one is configured); an empty-string value emits a self-closing item element (<tags/>).

A multiple: column that maps to an attribute field (a column whose name maps to an XML attribute, e.g. @tags, declared multiple: true) is rejected at compile with E359 — an XML attribute holds a single value and cannot repeat, and the writer emits repetition only as child elements. A runtime array reaching an attribute field is likewise rejected by the writer.

To rename the elements, add a join_values entry — the same block the CSV writer reads, sharing the field key. The XML writer reads two keys from it and ignores the CSV-only delimiter / on_conflict / escape:

- type: sink
  name: xml_out
  input: processed
  config:
    name: xml_out
    type: xml
    path: "./output/result.xml"
    join_values:
      - field: tags
        repeat_as: Tag      # per-item element name; defaults to the field name
        wrap_in: Tags       # optional container; omit for bare repeats
  • repeat_as — the element name emitted per item. Defaults to the field’s own element name.
  • wrap_in — a container element bracketing the repeated items. Omit it for bare repeats with no container.

A scalar value on a field that carries a join_values entry is treated as a one-element sequence: it receives the same repeat_as / wrap_in naming an array of length one would, so the emitted shape does not depend on whether a lone value arrived wrapped ([a]) or bare (a) — mirroring how the reader normalizes a lone scalar into a one-element array. A field with no entry emits the plain <field>value</field> element.

The two combine into the four arrangements, with no other key:

repeat_aswrap_inOutput for tags = [a, b]
——<tags>a</tags><tags>b</tags>
Tag—<Tag>a</Tag><Tag>b</Tag>
—Tags<Tags><tags>a</tags><tags>b</tags></Tags>
TagTags<Tags><Tag>a</Tag><Tag>b</Tag></Tags>

repeat_as and wrap_in must each be a well-formed XML name, validated the same way as the root_element / record_element names. Declaring join_values on an output format that is neither csv nor xml is rejected at compile (E362).

Round trip. A document read into a multiple: true column with the default naming writes back to the identical repeated elements — reading <Order><id>1</id><tags>a</tags><tags>b</tags></Order> into a tags column and writing it to an XML output with record_element: Order reproduces the input byte-for-byte.

Repeated elements

When a record element contains repeated child elements, two source-level declarations decide what happens to them, and both take the flattened dotted field name — see Source Nodes → Multi-value fields for the shared grammar. The XML-specific matching rules are below.

A declared field is the repeated element’s dotted path relative to the record element — the same form the flattened field names use. For a record element <Order> containing repeated <Item> children, the field is Item; for <Order><Items><Item>…, it is Items.Item.

One record per occurrence: split_to_rows

- type: source
  name: orders
  config:
    name: orders
    type: xml
    path: "./data/orders.xml"
    options:
      record_path: "Orders/Order"
    schema:
      - { name: id, type: int }
      - { name: "Item.name", type: string }
      - { name: "Item.qty", type: int }
    split_to_rows:
      - field: "Item"
        mode: split            # one output record per <Item> occurrence

Each output carries one occurrence’s fields plus every field outside the group, duplicated onto each record.

Under mode: split the occurrence’s fields keep their full dotted names (Item.name, Item.@sku), including the element’s attributes. Under the default mode: extract the declared field’s prefix is lifted off, so the same document yields name and qty; a repeated scalar element (<Tag>a</Tag>) has no remainder to lift and takes the declared field’s last segment, so Tags.Tag yields Tag under extract and stays Tags.Tag under split.

Lifting a prefix off can land an occurrence’s field on a name a field outside the group already occupies — <Order><name> alongside <Item><name>. The occurrence wins: under extract it is the record, so its own field is not shadowed by the parent it was merged with. Use mode: split when you need both values, which keeps them at name and Item.name.

A declared position_column wins over any field of that name, inside the occurrence or outside it. position_column: line_no against an <Item> that carries its own <line_no> child yields the occurrence’s index, not the document’s value — you named the column, so the index is what it holds.

An occurrence with no content (<Item></Item>) still emits a record, one carrying only the fields outside the group. A record with no occurrence of the element is governed by keep_empty: XML cannot distinguish an empty repetition from an absent element, and the default keep_empty: true passes the record through unchanged.

Entries apply in declaration order, so two declared fields multiply. Fields must name disjoint element groups — a duplicated field, or one extending another (Item and Item.part) — which is rejected at compile (E358), before the source opens. The disjointness rule is this reader’s: it assigns each element to one occurrence group by document position, which is sound exactly when the declared groups do not nest. A JSON source has no such constraint.

Set source-level max_output_rows_per_input: N to bound that cumulative product without constructing it in memory. The reader emits the first N rows in document/declaration order, then a first attempted row above the ceiling routes the original ordered XML field occurrences to the DLQ as expansion_limit_exceeded; 0 or omission is unlimited. See Source Nodes -> split_to_rows for the complete error and fail_fast behavior.

All occurrences in one field: multiple: true

Declaring a schema column multiple: true collects every occurrence of that flattened field into one array, in document order, instead of keeping only the first:

    schema:
      - { name: id, type: int }
      - { name: "Tag", type: string, multiple: true }

<Tag>a</Tag><Tag>b</Tag> yields ["a", "b"], and a single <Tag> still yields a one-element array. Declaring the flattened children of a repeated container (Item.name, Item.qty) collects each of them independently.

An empty occurrence — an empty-body <Tag></Tag> or a self-closing <Tag/> — is a real array element, collected in position as a null: <Tag>a</Tag><Tag></Tag><Tag>b</Tag> yields ["a", null, "b"] rather than squeezing the empty element out, so the array round-trips its per-item shape. (An empty text value reads as null, the same rule the reader applies elsewhere; the self-closing and empty-body forms behave identically.)

A field cannot be both collected and fanned out: naming a multiple: true column in split_to_rows is rejected at compile (E358).

A repeated element named by neither a split_to_rows entry nor a multiple: column is a loud error, not a silent drop. Keeping the first occurrence and discarding the rest would lose data without warning, so the reader refuses the record and names the offending field, pointing at the two ways to handle a repeat on purpose: declare the column multiple: true to collect every occurrence into an array, or add a split_to_rows entry to fan each occurrence out to its own record. Detection is per document at read time — a plan cannot know in advance that a particular document repeats a field. Under the default fail_fast strategy the run aborts with the diagnostic; under continue with dlq_granularity: document the offending document is routed to the dead-letter queue and the run continues.

Delimited text in one element: split_values

split_values parses <Tag>a;b;c</Tag> into ["a", "b", "c"]. The field must also be declared multiple: true.

Bounding envelope retention: max_index_bytes

When a source declares an envelope: and a pipeline reads $doc.* paths from it, the XML reader runs an event-driven streaming pre-scan that walks the document once and retains only the declared section subtrees — every other element, including a multi-megabyte body, is event-walked and dropped without being flattened into memory. The retained sections live in a bounded document index.

max_index_bytes caps that index. It is charged incrementally as each section is built, so even a single oversized declared section aborts mid-parse (naming the section and the cap) rather than risking an out-of-memory failure. It accepts a decimal size string (64MB, 500KB) or a bare byte count; optional, defaulting to 64MB. Only the declared sections a program actually reads are retained, so envelope metadata sits far below this ceiling in practice — the cap exists to convert an unbounded mistake into a clear error.

The reader holds no whole-document buffer: the body walks the document element-at-a-time, and the envelope pre-scan opens the source a second time to walk it independently — a file source is read twice, never buffered. Peak memory is the bounded section index plus a single live record, not the input size. See Document Envelope Context for the full model.

Preparation and delivery failures

Invalid XML names, illegal XML characters, unsupported nested shapes and resource refusal during preparation leave that operation’s destination bytes unchanged. Diagnostics identify a bounded offending field and the rule to correct. No replacement character or JSON-string fallback is written.

A destination may accept a prefix before failing. After that failure the writer refuses further work, including finalization, and dropping it never retries the prefix. Earlier delivered records remain delivered. The CLI publishes staged files only after successful execution; this is a separate boundary from writer delivery. See output preparation.

A CLI source with no body records produces an empty file because no writer is opened. Explicitly finalizing an unused library writer instead produces <Root></Root> with default names. An explicitly opened empty envelope retains its declared framing with a body count of zero.