Many files wrap their records in a header and a trailer: a batch id at the top, a record count at the bottom. Clinker lets every record read those values as $doc.<section>.<field>. This page shows where the values come from, how each file becomes its own document, and what happens when one record in a document fails.
The Source's envelope: block lists the sections to pull out of each file, where each one is, and which of its fields to keep, with their types.
- type: source
name: payments
config:
name: payments
type: json
glob: ./payments-*.json
options:
record_path: records
envelope:
sections:
:
extract: { json_pointer: "/BatchInfo" }
fields:
batch_id: string
run_date: date
operator: string
:
extract: { json_pointer: "/Summary" }
fields:
total: int
schema:
- { name: id, type: string }
- { name: region, type: string }
- { name: amount, type: int }
envelope: block doesn't declare is rejected when the pipeline is checked (E341), instead of quietly giving null.When a Source reads several files (with glob: or paths:), each file is its own document with its own section values. A record's $doc.* always comes from the file it was read from.
Document boundaries travel with the records. A grouped or global Aggregate gives one set of results per document: it finishes and emits a document's groups when that document ends, then starts afresh for the next. Twelve monthly files through one Aggregate give twelve monthly roll-ups, not one yearly total.
- type: aggregate
name: by_region
input: payments
config:
group_by: [region]
cxl: |
emit region = region
emit total = sum(amount)
The reference page covers how an Aggregate treats documents arriving through a Merge of several Sources, or after a Combine.
By default only the failing record is dead-lettered. When a half-processed file is worse than none, for example an interchange or a batch with a control total, set dlq_granularity: document on the Source to reject the whole file instead.
- type: source
name: payments
config:
…
dlq_granularity:
error_handling:
strategy: continue # required
_cxl_dlq_trigger_id.strategy: continue (E344); it can't be combined with a correlation_key anywhere in the pipeline (E370) or with a per-file Sink path (E343). Failures inside an Aggregate, Combine or Reshape stay per record and don't reject the document yet (#1232).| format | extract: | notes |
|---|---|---|
| JSON | json_pointer: "/Head" | Header and trailer sections, read before the records. |
| XML | xml_path: "/doc/Head" | Same as JSON. |
| Multi-record CSV / fixed-width | record_type: H | Header record types only. A trailer record's count is checked against the body instead of becoming a section. |
| X12, EDIFACT, HL7 | segment: ISA / UNB / FHS | The file header. Inner levels (functional groups, transaction sets, batches, messages) appear as extra sections the reader supplies. |
| SWIFT MT | segment: "1", "2", "3" or "5" | Service blocks. Block 4 is the message body, not a section. |
| Plain CSV / fixed-width | none | No envelope to read; declaring one is rejected (E356). |
| REST sources | none | No document; $doc is rejected (E349). |
A field holding an array or map can be indexed with a fixed value: $doc.Head.items[0], $doc.Head.meta["run_date"]. An index computed from the record is rejected, because the reader must know before reading which parts of the file to keep.
JSON and XML keep only the declared sections, capped by options.max_index_bytes (default 64MB). A section over the cap stops the run with an error naming it.
Each field is converted to its declared type (string, int, float, bool, date, date_time) when the file is opened. A value that can't convert fails the Source, naming the section, field and value.