Fixed-Width Format
Fixed-width files carry no delimiters — each field occupies a fixed column
range on every line, the layout common to mainframe extracts and legacy
COBOL exports. Because the byte layout is not self-describing, each column in
a fixed-width source’s schema: carries its byte layout (start +
width) alongside its CXL type — one unified declaration drives both the
physical slice and compile-time type checking. See
Source Nodes for the shared transport rules.
- type: source
name: legacy_data
config:
name: legacy_data
type: fixed_width
path: "./data/mainframe.dat"
schema:
- { name: account_id, type: string, start: 0, width: 12 }
- { name: balance, type: float, start: 12, width: 10 }
- { name: status_code, type: string, start: 22, width: 2 }
options:
line_separator: crlf # line-ending style
The column layout
Each column pins itself to a byte range with start (a 0-based offset) and
width (a byte count); end (exclusive) may be given instead of width.
Optional per-column formatting keys — justify, pad, trim, truncation
— control padding and trimming on read and write. Because the same column
declaration carries both the byte range and the CXL type, the physical layout
and the types can never drift apart. A layout shared across pipelines can live
in an external .schema.yaml file referenced by schema: layout.schema.yaml.
Writing fixed-width output
A fixed-width output node declares the same column layout in its
schema:. The writer places every field at its declared byte range —
start plus width (or end), resolved exactly as the reader slices —
regardless of the order the columns are declared in, so a file written
with a schema reads back under that same schema. Byte ranges the layout
leaves undeclared (a gap between fields) are filled with spaces. A column
that omits start continues at the previous column’s end, so a
width-only schema lays its fields out sequentially. Two columns whose
byte ranges overlap have no consistent layout; the writer rejects such a
schema when the output opens, naming both columns and their ranges.
Widths are byte counts, matching how the reader slices. When a value
is longer than its field, truncation cuts at a UTF-8 character boundary at
or below the width, so a multi-byte character is never split: the emitted
cell is always valid UTF-8 of exactly width bytes (it may hold fewer
characters than the width when a trailing multi-byte character does not
fit, with the freed bytes pad-filled). Because padding fills exact byte
counts on write and is stripped one character at a time on read, pad must
be a single-byte (ASCII) character — a multi-byte character (such as ·)
or a multi-character string (such as "0 ") is rejected when the schema is
resolved, on both the read and write sides. An absent or empty pad defaults
to a space. Under truncation: error an over-long value is still a hard error
before any slicing.
truncation: warn truncates and reports it; silent performs the same
truncation and reports nothing. Numeric columns default to error, and
other columns default to warn.
When a run finishes, each output that truncated under warn prints one
W367 warning to standard error, naming every such column with the exact
number of values cut, the longest original value in bytes, the column width,
and the numbers of the first eight output records that were cut (records
count from 1 across everything that output wrote, including every file of a
split output; an output that writes one file per source file numbers its files
one after another in file-path order, except that a file whose writer fails is
numbered when it fails). The warning does not change the exit code. The report never
copies a value, so it cannot leak data into logs, and its memory is fixed by
the schema when the output opens: recording a truncation cannot fail, so
warn never turns an over-long value into a rejected record. A record
that is rejected or not delivered is not counted. Truncations are also
counted by the clinker.sink.truncations metric. Run
clinker explain --code W367 for the fixes.
A type: decimal output column with a scale rounds its values to that
many fractional places on write (banker’s rounding), the same contract a
decimal source column applies on read. This matters here: a computed
decimal such as avg(amount) is a full-precision quotient that would
overflow a narrow numeric field — a hard error under the default
truncation: error for numeric fields — so declaring the field’s scale
shrinks it to fixed places that fit. For example, avg over 1.00, 1.00, 2.00 written into { name: average, type: decimal, scale: 2, width: 6 }
emits 1.33; without the scale the 28-digit quotient overflows the
6-byte field and fails.
Options
| Option | Default | Description |
|---|---|---|
line_separator | lf | Record separator: lf, crlf, or none for consecutive fixed-length records. |
Under lf or crlf, the reader buffers each physical line only up to the
declared record width plus a line-terminator allowance. A physical line wider
than the declared width — trailing filler beyond the last declared field, or a
schema that maps only a prefix of a wider fixed-length record — reads its
declared-width portion; the remaining bytes are discarded up to the next line
terminator and the reader continues with the following record. Because the
buffered portion is capped, a malformed file (a corrupt or missing newline)
cannot grow a single record until end of input: memory stays bounded regardless
of how long the physical line runs. A final line with no trailing newline reads
normally as long as its declared fields fit within the width.
Strict selected-cell input
Field offsets remain physical byte offsets. Each selected cell must be valid UTF-8 within its own byte range; a boundary that cuts through a multi-byte character fails rather than shifting the layout or inserting a replacement character. Undeclared gaps and discarded trailing bytes are not decoded, so invalid bytes in those ignored ranges do not invalidate a selected cell. A leading UTF-8 BOM is removed before applying the layout. There is no charset conversion for fixed-width input or output.
A typed numeric cell that cannot be parsed terminates the read, including
under strategy: continue; that policy does not make numeric parse failures
recoverable. This differs from the multi-record unknown-discriminator
handling described below.
Multi-value cells (split_values)
A fixed-width field holds one value, but its text may pack several behind a
delimiter within the field’s byte range. Declare the column multiple: true
and add a split_values entry, and the reader splits the (padding-stripped)
field text and coerces each part to the column’s declared type:
split_values:
- { field: tags, delimiter: ";" }
schema:
- { name: order_id, type: string, start: 0, width: 4 }
- { name: tags, type: string, start: 4, width: 20, multiple: true }
Because the fixed-width reader is the sole coercion pass, each element is typed
here — a multiple: true int field over 1;2;3 reads as [1, 2, 3]. A blank
field is an empty array; a field with no delimiter is a one-element array. A
multiple: true column with no covering split_values entry is rejected at
compile (E361). The entry is read only on a
single-schema source, not the multi-record reader below.
Repeating groups
A positional repeating group occupies a bounded sequence of fixed-width
occurrence records. Declare the logical column as type: map with
multiple: true, put the per-occurrence byte layout in fields, and give the
group a finite positive occurs.max:
schema:
- { name: account_id, type: string, start: 0, width: 8 }
- name: transactions
type: map
multiple: true
start: 8
fields:
- { name: kind, type: string, start: 0, width: 1 }
- { name: code, type: string, start: 1, width: 2 }
occurs:
min: 0
max: 3
fill: pad
on_overflow: error
count_field:
name: transaction_count
width: 1
Child start offsets are relative to one occurrence. As with top-level
columns, a child may omit start to continue after the previous child. The
count field, when present, is a leading physical cell inside the group’s byte
range; the occurrence payload follows it. It controls how many occurrence
maps the reader returns, but it is not a logical record column and cannot be
named from CXL. In the example, 2A01B02 means two three-byte occurrences
followed by one unused padded slot.
occurs.min defaults to zero and cannot exceed max. Both values count
logical occurrences. max has no default and cannot be inferred: it is what
bounds the resolved record width, the reader’s one-line buffer, and the
writer’s one-record buffer. Missing, zero, overflowing, overlapping, hybrid,
or recursively repeated layouts fail during normal pipeline compilation,
before an input or destination is opened.
fill chooses the physical treatment of unused occurrences:
pad(the default) always reservesmaxoccurrence slots and fills unused child cells with their declared padding. Without a count field, the reader infers only a trailing run of completely padded slots. A populated slot after an empty one is invalid, and the writer rejects an authored occurrence that itself renders entirely as padding because it could not be read back unambiguously. Addcount_fieldwhen an all-empty occurrence is meaningful.shiftomits unused slots. A count field makes the following byte position explicit. Without one, a shifted group must be the last physical field; the reader otherwise cannot distinguish group bytes from the next field.
Overflow is an error by default. The diagnostic names the group, its declared
maximum, and the supplied count without printing record values. To choose
lossy output deliberately, set on_overflow: truncate and also select the
retained end with keep: first or keep: last; keep is invalid with the
default error policy. The writer validates and encodes the complete bounded
record before the first destination write, so an invalid later occurrence or
overflow cannot leave a partial record behind.
Repeating groups and delimiter-packed scalar cells are separate encodings. A
bare multiple: true fixed-width column is not a positional group, and a
split_values entry cannot stand in for fields plus occurs. A fixed-width
sink that receives an array of records must declare the same named positional
group in its output schema.
Scalar document headers and footers
With reconstruct_envelope: true, options.envelope.header_from_doc and
options.envelope.footer_from_doc select arbitrary document sections. Their
values are concatenated in section field order, without field padding,
delimiters, or the body’s byte layout. Strings
are verbatim, booleans use true/false, numbers use their natural scalar
spelling, dates use YYYYMMDD, datetimes use YYYYMMDDhhmmss, and null emits
no text. Arrays and maps have no scalar envelope representation and fail
before any bytes from that header or footer reach the destination.
An absent section emits nothing. A present section with no fields still
emits its separator: LF, CRLF, or no bytes under line_separator: none.
Header, body record, and footer are separate complete operations, so a bad
footer cannot undo an earlier successful body. See
fixed-width document output
for section selection and input-extraction limits.
Library output and failures
Direct callers construct FixedWidthEncoder with their columns,
FixedWidthWriterConfig, and finite WriterResources, then wrap it in
PreparedWriter. MemoryOnlyResources::new requires an explicit nonzero
budget. The resource-free writer constructor is unavailable. Read the
truncation account through FormatWriter::truncation_summary() (or
writer.encoder().truncation_summary()): None when nothing was truncated,
otherwise per-column counts, longest lengths, and the first delivered record
numbers.
Preparation validates the complete operation before destination writes.
Preparation failure leaves committed state unchanged and permits a corrected
retry. Once delivery starts, a destination can accept a prefix before
failing; the writer then refuses continuation, including flush, and drop
does not retry. flush_bytes() only drains the destination; flush()
finalizes once and drains. See output preparation
for the separate storage and publication guarantees.
Schema drift
Fixed-width is inert with respect to
auto-widen: because every byte is accounted
for by the format schema, there are no “unmapped” trailing columns to
absorb. The on_unmapped policy has no effect on a fixed-width source.
Multi-record files (header / trailer / body)
Mainframe and banking extracts often interleave multiple record types
in one file — a header line, many body lines, and a trailer line — each
identified by a discriminator at a fixed byte position (commonly the
first character). Declare these with a map-form schema: carrying a
discriminator: byte range and a records: list, instead of a flat column
list. Each record type names its tag (the discriminator value) and its own
byte-positioned columns:; the reader synthesizes the lead record_type
column automatically.
- type: source
name: payments
config:
name: payments
type: fixed_width
path: "./data/payments.dat"
schema: # one multi-record schema (map form)
discriminator: { start: 0, width: 1 } # the type tag occupies byte 0
records:
- { id: header, tag: H, columns: [ { name: batch_id, type: string, start: 1, width: 9 } ] }
- { id: detail, tag: D, columns: [ { name: id, type: int, start: 1, width: 5 }, { name: amount, type: int, start: 6, width: 4 } ] }
- { id: trailer, tag: T, columns: [ { name: count, type: int, start: 1, width: 5 } ] }
structure:
- { record: trailer, count: count } # validate T's count against the body count
envelope:
sections:
head:
extract: { record_type: H } # the H line surfaces as $doc.head.*
fields:
batch_id: string
The reader streams one record per line on a single superset schema
whose lead record_type column carries the matched type’s id. A
downstream Route discriminates on that column; the
file is never buffered.
- Header lines declared as an
envelope:section via therecord_typeextract surface as$doc.<section>.*and are excluded from the body stream (see Envelopes & Document Context). - Trailer lines named by a
structure:constraint are validated as they stream — the declaredcountfield is checked against the actual body-record count at document close — and excluded from the body stream. A declared trailer that never appears is an incomplete-document error; a body line after the trailer is rejected as content past the document close. - Blank lines (empty or whitespace-only, common after concatenation)
are skipped rather than rejected; a line whose declared field range is
cut off mid-value is a truncation error, not a silently-partial read.
Field parsing — type coercion, padding strip, justification — is shared
with the single-record fixed-width reader, so a declared
typeparses identically on both paths. - An unknown discriminator value (a tag no
records:entry declares) is a structural-integrity failure, classified separately from a trailer count mismatch. It aborts the run underfail_fast; undercontinuewith the default record granularity it dead-letters only that physical line and continues with the next line; underdlq_granularity: documentit condemns the whole file. A record-grained DLQ row carries the line text in_cxl_dlq_source_recordwithout guessing which declared field layout the unknown tag meant.