Clinker user guide · ← back to Document Envelope Context

The envelope
around your records

Many files wrap their records in a header and a trailer: a batch id at the top, a record count at the bottom. Clinker lets every record read those values as $doc.<section>.<field>. This page shows where the values come from, how each file becomes its own document, and what happens when one record in a document fails.

01

SectionsYou name them; every record can read them

The Source's envelope: block lists the sections to pull out of each file, where each one is, and which of its fields to keep, with their types.

- type: source
  name: payments
  config:
    name: payments
    type: json
    glob: ./payments-*.json
    options:
      record_path: records
    envelope:
      sections:
        :
          extract: { json_pointer: "/BatchInfo" }
          fields:
            batch_id: string
            run_date: date
            operator: string
        :
          extract: { json_pointer: "/Summary" }
          fields:
            total: int
    schema:
      - { name: id, type: string }
      - { name: region, type: string }
      - { name: amount, type: int }
  • The names are yours. The engine reserves none. Rename a section on the right and watch the CXL change with it.
  • The trailer is readable from the first record. JSON and XML sources read the declared sections before any record is streamed, so a total at the bottom of the file is already known on row one.
  • Typos fail early. Reading a section or field the envelope: block doesn't declare is rejected when the pipeline is checked (E341), instead of quietly giving null.
  • Missing data is null. A declared field or section that a particular file doesn't contain reads as null.
a Transform emits:
02

One document per fileAnd an Aggregate rolls up each one separately

When a Source reads several files (with glob: or paths:), each file is its own document with its own section values. A record's $doc.* always comes from the file it was read from.

Document boundaries travel with the records. A grouped or global Aggregate gives one set of results per document: it finishes and emits a document's groups when that document ends, then starts afresh for the next. Twelve monthly files through one Aggregate give twelve monthly roll-ups, not one yearly total.

- type: aggregate
  name: by_region
  input: payments
  config:
    group_by: [region]
    cxl: |
      emit region = region
      emit total  = sum(amount)
Only one document's groups are held at a time, so memory follows the largest file rather than all of them together. A single-file Source is one document, so it gives one roll-up as usual. A Combine keeps the boundaries too, so an Aggregate after a join rolls up per driver document.

The reference page covers how an Aggregate treats documents arriving through a Merge of several Sources, or after a Combine.

Aggregate output one roll-up per document, labelled by file here
Not what you get one roll-up across all files
03

Rejecting a documentTap a record to make it fail in a Transform

By default only the failing record is dead-lettered. When a half-processed file is worse than none, for example an interchange or a batch with a control total, set dlq_granularity: document on the Source to reject the whole file instead.

- type: source
  name: payments
  config:
    …
    dlq_granularity: 
error_handling:
  strategy: continue      # required
  • The first failing record is the root cause, with its real error category.
  • Every other record of that file that reached a Sink, or failed too, is written as document_rejected, pointing back at the root cause through _cxl_dlq_trigger_id.
  • No record read from a rejected file is written by any Sink. Other files carry on untouched.
  • The document is the whole file: for X12 or EDIFACT, the whole interchange, not one transaction set.
Rules that come with it: it needs strategy: continue (E344); it can't be combined with a correlation_key anywhere in the pipeline (E370) or with a per-file Sink path (E343). Failures inside an Aggregate, Combine or Reshape stay per record and don't reject the document yet (#1232).
dlq_granularity
Sink output
DLQ
04

Formats & limitsWhere sections can come from

formatextract:notes
JSONjson_pointer: "/Head"Header and trailer sections, read before the records.
XMLxml_path: "/doc/Head"Same as JSON.
Multi-record CSV / fixed-widthrecord_type: HHeader record types only. A trailer record's count is checked against the body instead of becoming a section.
X12, EDIFACT, HL7segment: ISA / UNB / FHSThe file header. Inner levels (functional groups, transaction sets, batches, messages) appear as extra sections the reader supplies.
SWIFT MTsegment: "1", "2", "3" or "5"Service blocks. Block 4 is the message body, not a section.
Plain CSV / fixed-widthnoneNo envelope to read; declaring one is rejected (E356).
REST sourcesnoneNo document; $doc is rejected (E349).

Indexing a section

A field holding an array or map can be indexed with a fixed value: $doc.Head.items[0], $doc.Head.meta["run_date"]. An index computed from the record is rejected, because the reader must know before reading which parts of the file to keep.

Big sections

JSON and XML keep only the declared sections, capped by options.max_index_bytes (default 64MB). A section over the cap stops the run with an error naming it.

Typed fields

Each field is converted to its declared type (string, int, float, bool, date, date_time) when the file is opened. A value that can't convert fails the Source, naming the section, field and value.