Fixed-Width Format
Fixed-width files carry no delimiters — each field occupies a fixed column
range on every line, the layout common to mainframe extracts and legacy
COBOL exports. Because the byte layout is not self-describing, each column in
a fixed-width source’s schema: carries its byte layout (start +
width) alongside its CXL type — one unified declaration drives both the
physical slice and compile-time type checking. See
Source Nodes for the shared transport rules.
- type: source
name: legacy_data
config:
name: legacy_data
type: fixed_width
path: "./data/mainframe.dat"
schema:
- { name: account_id, type: string, start: 0, width: 12 }
- { name: balance, type: float, start: 12, width: 10 }
- { name: status_code, type: string, start: 22, width: 2 }
options:
line_separator: crlf # line-ending style
The column layout
Each column pins itself to a byte range with start (a 0-based offset) and
width (a byte count); end (exclusive) may be given instead of width.
Optional per-column formatting keys — justify, pad, trim, truncation
— control padding and trimming on read and write. Because the same column
declaration carries both the byte range and the CXL type, the physical layout
and the types can never drift apart. A layout shared across pipelines can live
in an external .schema.yaml file referenced by schema: layout.schema.yaml.
Writing fixed-width output
A fixed-width output node declares the same column layout in its
schema:. The writer places every field at its declared byte range —
start plus width (or end), resolved exactly as the reader slices —
regardless of the order the columns are declared in, so a file written
with a schema reads back under that same schema. Byte ranges the layout
leaves undeclared (a gap between fields) are filled with spaces. A column
that omits start continues at the previous column’s end, so a
width-only schema lays its fields out sequentially. Two columns whose
byte ranges overlap have no consistent layout; the writer rejects such a
schema when the output opens, naming both columns and their ranges.
Widths are byte counts, matching how the reader slices. When a value
is longer than its field, truncation cuts at a UTF-8 character boundary at
or below the width, so a multi-byte character is never split: the emitted
cell is always valid UTF-8 of exactly width bytes (it may hold fewer
characters than the width when a trailing multi-byte character does not
fit, with the freed bytes pad-filled). Because padding fills exact byte
counts on write and is stripped one character at a time on read, pad must
be a single-byte (ASCII) character — a multi-byte character (such as ·)
or a multi-character string (such as "0 ") is rejected when the schema is
resolved, on both the read and write sides. An absent or empty pad defaults
to a space. Under truncation: error an over-long value is still a hard error
before any slicing.
A type: decimal output column with a scale rounds its values to that
many fractional places on write (banker’s rounding), the same contract a
decimal source column applies on read. This matters here: a computed
decimal such as avg(amount) is a full-precision quotient that would
overflow a narrow numeric field — a hard error under the default
truncation: error for numeric fields — so declaring the field’s scale
shrinks it to fixed places that fit. For example, avg over 1.00, 1.00, 2.00 written into { name: average, type: decimal, scale: 2, width: 6 }
emits 1.33; without the scale the 28-digit quotient overflows the
6-byte field and fails.
Options
| Option | Default | Description |
|---|---|---|
line_separator | platform | Line-ending style (lf / crlf) used to split the file into records. |
Under lf or crlf, the reader buffers each physical line only up to the
declared record width plus a line-terminator allowance. A physical line wider
than the declared width — trailing filler beyond the last declared field, or a
schema that maps only a prefix of a wider fixed-length record — reads its
declared-width portion; the remaining bytes are discarded up to the next line
terminator and the reader continues with the following record. Because the
buffered portion is capped, a malformed file (a corrupt or missing newline)
cannot grow a single record until end of input: memory stays bounded regardless
of how long the physical line runs. A final line with no trailing newline reads
normally as long as its declared fields fit within the width.
Multi-value cells (split_values)
A fixed-width field holds one value, but its text may pack several behind a
delimiter within the field’s byte range. Declare the column multiple: true
and add a split_values entry, and the reader splits the (padding-stripped)
field text and coerces each part to the column’s declared type:
split_values:
- { field: tags, delimiter: ";" }
schema:
- { name: order_id, type: string, start: 0, width: 4 }
- { name: tags, type: string, start: 4, width: 20, multiple: true }
Because the fixed-width reader is the sole coercion pass, each element is typed
here — a multiple: true int field over 1;2;3 reads as [1, 2, 3]. A blank
field is an empty array; a field with no delimiter is a one-element array. A
multiple: true column with no covering split_values entry is rejected at
compile (E361). The entry is read only on a
single-schema source, not the multi-record reader below.
Schema drift
Fixed-width is inert with respect to
auto-widen: because every byte is accounted
for by the format schema, there are no “unmapped” trailing columns to
absorb. The on_unmapped policy has no effect on a fixed-width source.
Multi-record files (header / trailer / body)
Mainframe and banking extracts often interleave multiple record types
in one file — a header line, many body lines, and a trailer line — each
identified by a discriminator at a fixed byte position (commonly the
first character). Declare these with a map-form schema: carrying a
discriminator: byte range and a records: list, instead of a flat column
list. Each record type names its tag (the discriminator value) and its own
byte-positioned columns:; the reader synthesizes the lead record_type
column automatically.
- type: source
name: payments
config:
name: payments
type: fixed_width
path: "./data/payments.dat"
schema: # one multi-record schema (map form)
discriminator: { start: 0, width: 1 } # the type tag occupies byte 0
records:
- { id: header, tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
- { id: detail, tag: D, columns: [ { name: id, type: int, start: 1, width: 5 }, { name: amount, type: int, start: 6, width: 4 } ] }
- { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 5 } ] }
structure:
- { record: trailer, count: count } # validate T's count against the body count
envelope:
sections:
head:
extract: { record_type: H } # the H line surfaces as $doc.head.*
fields:
batch_id: string
The reader streams one record per line on a single superset schema
whose lead record_type column carries the matched type’s id. A
downstream Route discriminates on that column; the
file is never buffered.
- Header lines declared as an
envelope:section via therecord_typeextract surface as$doc.<section>.*and are excluded from the body stream (see Envelopes & Document Context). - Trailer lines named by a
structure:constraint are validated as they stream — the declaredcountfield is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body line after the trailer is rejected as content past the document close. - Blank lines (empty or whitespace-only, common after concatenation)
are skipped rather than rejected; a line whose declared field range is
cut off mid-value is a truncation error, not a silently-partial read.
Field parsing — type coercion, padding strip, justification — is shared
with the single-record fixed-width reader, so a declared
typeparses identically on both paths. - An unknown discriminator value (a tag no
records:entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch: it aborts the run by default, or underdlq_granularity: documentcondemns the whole file to the dead-letter sink and the run continues.