Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Fixed-Width Format

Fixed-width files carry no delimiters — each field occupies a fixed column range on every line, the layout common to mainframe extracts and legacy COBOL exports. Because the byte layout is not self-describing, each column in a fixed-width source’s schema: carries its byte layout (start + width) alongside its CXL type — one unified declaration drives both the physical slice and compile-time type checking. See Source Nodes for the shared transport rules.

- type: source
  name: legacy_data
  config:
    name: legacy_data
    type: fixed_width
    path: "./data/mainframe.dat"
    schema:
      - { name: account_id,  type: string, start: 0,  width: 12 }
      - { name: balance,     type: float,  start: 12, width: 10 }
      - { name: status_code, type: string, start: 22, width: 2 }
    options:
      line_separator: crlf    # line-ending style

The column layout

Each column pins itself to a byte range with start (a 0-based offset) and width (a byte count); end (exclusive) may be given instead of width. Optional per-column formatting keys — justify, pad, trim, truncation — control padding and trimming on read and write. Because the same column declaration carries both the byte range and the CXL type, the physical layout and the types can never drift apart. A layout shared across pipelines can live in an external .schema.yaml file referenced by schema: layout.schema.yaml.

Writing fixed-width output

A fixed-width output node declares the same column layout in its schema:. The writer places every field at its declared byte range — start plus width (or end), resolved exactly as the reader slices — regardless of the order the columns are declared in, so a file written with a schema reads back under that same schema. Byte ranges the layout leaves undeclared (a gap between fields) are filled with spaces. A column that omits start continues at the previous column’s end, so a width-only schema lays its fields out sequentially. Two columns whose byte ranges overlap have no consistent layout; the writer rejects such a schema when the output opens, naming both columns and their ranges.

Widths are byte counts, matching how the reader slices. When a value is longer than its field, truncation cuts at a UTF-8 character boundary at or below the width, so a multi-byte character is never split: the emitted cell is always valid UTF-8 of exactly width bytes (it may hold fewer characters than the width when a trailing multi-byte character does not fit, with the freed bytes pad-filled). Because padding fills exact byte counts on write and is stripped one character at a time on read, pad must be a single-byte (ASCII) character — a multi-byte character (such as ·) or a multi-character string (such as "0 ") is rejected when the schema is resolved, on both the read and write sides. An absent or empty pad defaults to a space. Under truncation: error an over-long value is still a hard error before any slicing.

truncation: warn truncates and reports it; silent performs the same truncation and reports nothing. Numeric columns default to error, and other columns default to warn.

When a run finishes, each output that truncated under warn prints one W367 warning to standard error, naming every such column with the exact number of values cut, the longest original value in bytes, the column width, and the numbers of the first eight output records that were cut (records count from 1 across everything that output wrote, including every file of a split output; an output that writes one file per source file numbers its files one after another in file-path order, except that a file whose writer fails is numbered when it fails). The warning does not change the exit code. The report never copies a value, so it cannot leak data into logs, and its memory is fixed by the schema when the output opens: recording a truncation cannot fail, so warn never turns an over-long value into a rejected record. A record that is rejected or not delivered is not counted. Truncations are also counted by the clinker.sink.truncations metric. Run clinker explain --code W367 for the fixes.

A type: decimal output column with a scale rounds its values to that many fractional places on write (banker’s rounding), the same contract a decimal source column applies on read. This matters here: a computed decimal such as avg(amount) is a full-precision quotient that would overflow a narrow numeric field — a hard error under the default truncation: error for numeric fields — so declaring the field’s scale shrinks it to fixed places that fit. For example, avg over 1.00, 1.00, 2.00 written into { name: average, type: decimal, scale: 2, width: 6 } emits 1.33; without the scale the 28-digit quotient overflows the 6-byte field and fails.

Options

OptionDefaultDescription
line_separatorlfRecord separator: lf, crlf, or none for consecutive fixed-length records.

Under lf or crlf, the reader buffers each physical line only up to the declared record width plus a line-terminator allowance. A physical line wider than the declared width — trailing filler beyond the last declared field, or a schema that maps only a prefix of a wider fixed-length record — reads its declared-width portion; the remaining bytes are discarded up to the next line terminator and the reader continues with the following record. Because the buffered portion is capped, a malformed file (a corrupt or missing newline) cannot grow a single record until end of input: memory stays bounded regardless of how long the physical line runs. A final line with no trailing newline reads normally as long as its declared fields fit within the width.

Strict selected-cell input

Field offsets remain physical byte offsets. Each selected cell must be valid UTF-8 within its own byte range; a boundary that cuts through a multi-byte character fails rather than shifting the layout or inserting a replacement character. Undeclared gaps and discarded trailing bytes are not decoded, so invalid bytes in those ignored ranges do not invalidate a selected cell. A leading UTF-8 BOM is removed before applying the layout. There is no charset conversion for fixed-width input or output.

A typed numeric cell that cannot be parsed terminates the read, including under strategy: continue; that policy does not make numeric parse failures recoverable. This differs from the multi-record unknown-discriminator handling described below.

Multi-value cells (split_values)

A fixed-width field holds one value, but its text may pack several behind a delimiter within the field’s byte range. Declare the column multiple: true and add a split_values entry, and the reader splits the (padding-stripped) field text and coerces each part to the column’s declared type:

    split_values:
      - { field: tags, delimiter: ";" }
    schema:
      - { name: order_id, type: string, start: 0, width: 4 }
      - { name: tags, type: string, start: 4, width: 20, multiple: true }

Because the fixed-width reader is the sole coercion pass, each element is typed here — a multiple: true int field over 1;2;3 reads as [1, 2, 3]. A blank field is an empty array; a field with no delimiter is a one-element array. A multiple: true column with no covering split_values entry is rejected at compile (E361). The entry is read only on a single-schema source, not the multi-record reader below.

Repeating groups

A positional repeating group occupies a bounded sequence of fixed-width occurrence records. Declare the logical column as type: map with multiple: true, put the per-occurrence byte layout in fields, and give the group a finite positive occurs.max:

    schema:
      - { name: account_id, type: string, start: 0, width: 8 }
      - name: transactions
        type: map
        multiple: true
        start: 8
        fields:
          - { name: kind, type: string, start: 0, width: 1 }
          - { name: code, type: string, start: 1, width: 2 }
        occurs:
          min: 0
          max: 3
          fill: pad
          on_overflow: error
        count_field:
          name: transaction_count
          width: 1

Child start offsets are relative to one occurrence. As with top-level columns, a child may omit start to continue after the previous child. The count field, when present, is a leading physical cell inside the group’s byte range; the occurrence payload follows it. It controls how many occurrence maps the reader returns, but it is not a logical record column and cannot be named from CXL. In the example, 2A01B02 means two three-byte occurrences followed by one unused padded slot.

occurs.min defaults to zero and cannot exceed max. Both values count logical occurrences. max has no default and cannot be inferred: it is what bounds the resolved record width, the reader’s one-line buffer, and the writer’s one-record buffer. Missing, zero, overflowing, overlapping, hybrid, or recursively repeated layouts fail during normal pipeline compilation, before an input or destination is opened.

fill chooses the physical treatment of unused occurrences:

  • pad (the default) always reserves max occurrence slots and fills unused child cells with their declared padding. Without a count field, the reader infers only a trailing run of completely padded slots. A populated slot after an empty one is invalid, and the writer rejects an authored occurrence that itself renders entirely as padding because it could not be read back unambiguously. Add count_field when an all-empty occurrence is meaningful.
  • shift omits unused slots. A count field makes the following byte position explicit. Without one, a shifted group must be the last physical field; the reader otherwise cannot distinguish group bytes from the next field.

Overflow is an error by default. The diagnostic names the group, its declared maximum, and the supplied count without printing record values. To choose lossy output deliberately, set on_overflow: truncate and also select the retained end with keep: first or keep: last; keep is invalid with the default error policy. The writer validates and encodes the complete bounded record before the first destination write, so an invalid later occurrence or overflow cannot leave a partial record behind.

Repeating groups and delimiter-packed scalar cells are separate encodings. A bare multiple: true fixed-width column is not a positional group, and a split_values entry cannot stand in for fields plus occurs. A fixed-width sink that receives an array of records must declare the same named positional group in its output schema.

Scalar document headers and footers

With reconstruct_envelope: true, options.envelope.header_from_doc and options.envelope.footer_from_doc select arbitrary document sections. Their values are concatenated in section field order, without field padding, delimiters, or the body’s byte layout. Strings are verbatim, booleans use true/false, numbers use their natural scalar spelling, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null emits no text. Arrays and maps have no scalar envelope representation and fail before any bytes from that header or footer reach the destination.

An absent section emits nothing. A present section with no fields still emits its separator: LF, CRLF, or no bytes under line_separator: none. Header, body record, and footer are separate complete operations, so a bad footer cannot undo an earlier successful body. See fixed-width document output for section selection and input-extraction limits.

Library output and failures

Direct callers construct FixedWidthEncoder with their columns, FixedWidthWriterConfig, and finite WriterResources, then wrap it in PreparedWriter. MemoryOnlyResources::new requires an explicit nonzero budget. The resource-free writer constructor is unavailable. Read the truncation account through FormatWriter::truncation_summary() (or writer.encoder().truncation_summary()): None when nothing was truncated, otherwise per-column counts, longest lengths, and the first delivered record numbers.

Preparation validates the complete operation before destination writes. Preparation failure leaves committed state unchanged and permits a corrected retry. Once delivery starts, a destination can accept a prefix before failing; the writer then refuses continuation, including flush, and drop does not retry. flush_bytes() only drains the destination; flush() finalizes once and drains. See output preparation for the separate storage and publication guarantees.

Schema drift

Fixed-width is inert with respect to auto-widen: because every byte is accounted for by the format schema, there are no “unmapped” trailing columns to absorb. The on_unmapped policy has no effect on a fixed-width source.

Multi-record files (header / trailer / body)

Mainframe and banking extracts often interleave multiple record types in one file — a header line, many body lines, and a trailer line — each identified by a discriminator at a fixed byte position (commonly the first character). Declare these with a map-form schema: carrying a discriminator: byte range and a records: list, instead of a flat column list. Each record type names its tag (the discriminator value) and its own byte-positioned columns:; the reader synthesizes the lead record_type column automatically.

- type: source
  name: payments
  config:
    name: payments
    type: fixed_width
    path: "./data/payments.dat"
    schema:                                    # one multi-record schema (map form)
      discriminator: { start: 0, width: 1 }    # the type tag occupies byte 0
      records:
        - { id: header,  tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
        - { id: detail,  tag: D, columns: [ { name: id, type: int, start: 1, width: 5 }, { name: amount, type: int, start: 6, width: 4 } ] }
        - { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 5 } ] }
      structure:
        - { record: trailer, count: count }     # validate T's count against the body count
    envelope:
      sections:
        head:
          extract: { record_type: H }          # the H line surfaces as $doc.head.*
          fields:
            batch_id: string

The reader streams one record per line on a single superset schema whose lead record_type column carries the matched type’s id. A downstream Route discriminates on that column; the file is never buffered.

  • Header lines declared as an envelope: section via the record_type extract surface as $doc.<section>.* and are excluded from the body stream (see Envelopes & Document Context).
  • Trailer lines named by a structure: constraint are validated as they stream — the declared count field is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body line after the trailer is rejected as content past the document close.
  • Blank lines (empty or whitespace-only, common after concatenation) are skipped rather than rejected; a line whose declared field range is cut off mid-value is a truncation error, not a silently-partial read. Field parsing — type coercion, padding strip, justification — is shared with the single-record fixed-width reader, so a declared type parses identically on both paths.
  • An unknown discriminator value (a tag no records: entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run under fail_fast; under continue with the default record granularity it dead-letters only that physical line and continues with the next line; under dlq_granularity: document it condemns the whole file. A record-grained DLQ row carries the line text in _cxl_dlq_source_record without guessing which declared field layout the unknown tag meant.