Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Network Sources (REST)

A Source reads from the filesystem by default. To pull records from a network endpoint instead, declare a transport: block on the Source. The transport selects where records come from; it sits above the on-disk type: (the format), which for a REST source still selects how the response bodies decode.

A network transport is a finite-pull source: it runs on its own thread, drives a synchronous client to cursor exhaustion, then exits. There is no daemon, no event loop, and no async runtime — the same single- process, run-to-drain model as a file pipeline. Finiteness is a hard property of the reader: a REST source enforces explicit page and record limits, so an unbounded endpoint cannot keep it running forever. If the server offers another continuation after max_pages, the source fails closed instead of reporting a truncated pull as successful completion.

A network source still requires a schema: block. That authored schema is the row-to-record target: the reader maps each decoded object onto it, coercing values leniently. A per-row value that cannot coerce is left unchanged at the reader and routed to the dead-letter queue at the Transform stage — identical to file-source semantics. A network source declares no file matcher (path / glob / regex / paths); declaring one is a configuration error (E219).

Because a network source has no file path, its $source.file provenance column and the {source_file} output template both resolve to a stable synthetic identifier, <source:NAME>, where NAME is the Source node’s name.

REST sources

A rest source issues paginated HTTP GETs against a base URL, decoding each response body through the declared json or xml format. (Other formats are rejected with E220 — a REST body is a multi-record document, not a flat CSV/fixed-width stream.)

nodes:
  - type: source
    name: orders_api
    config:
      name: orders_api
      type: json
      options:
        format: array        # each page body is a JSON array of objects
      transport:
        kind: rest
        url: https://api.example.com/v1/orders
        max_pages: 50         # HARD page cap — required
        pagination:
          strategy: link_header
        auth:
          scheme: bearer
          token: "${ORDERS_TOKEN}"
      schema:
        - { name: order_id, type: int }
        - { name: total,    type: float }
        - { name: placed_at, type: date_time }

Pagination strategies

The pagination.strategy selects how the reader advances pages and detects the last one. max_records bounds emitted records. max_pages bounds requests and requires the server to reach an actual terminal page; an offered continuation beyond that bound is an error.

  • none (default) — a single GET; the body is the whole result.

  • offset — ?offset=N&limit=L, advancing the offset by the page size each request. The last page is the one that returns fewer rows than limit. A page containing exactly limit rows is not proof of end-of-input, so Clinker must issue one more bounded request to observe a short or empty terminal page. Size max_pages to leave room for that probe; reaching the cap first fails with page_limit_reached instead of accepting a possibly truncated result.

    pagination:
      strategy: offset
      limit: 200
      offset_param: offset     # optional, defaults shown
      limit_param: limit
    
  • cursor_token — the reader reads a continuation token from a JSON pointer in each response and sends it back on the next request. Paging stops when the token field is absent or null.

    pagination:
      strategy: cursor_token
      cursor_param: page_token
      next_token_pointer: /meta/next_page   # RFC 6901 JSON pointer
    
  • link_header — the reader follows the URL in the response’s RFC 8288 Link: <…>; rel="next" header until no such link is present. Registered relation tokens are case-insensitive, so next, Next, and NEXT have the same meaning.

    pagination:
      strategy: link_header
    

Continuation and redirect safety

Every server-directed continuation or redirect is resolved against the effective response URL and normalized before another request is built. Only the original normalized origin is allowed. Cross-origin targets, HTTPS-to-HTTP downgrades, malformed or conflicting rel="next" metadata, redirect or continuation cycles, and traversal beyond the configured bounds fail before a foreign or repeated request is sent.

Normalization resolves . and .. exactly as RFC 3986 does, and changes nothing else about the path. An empty segment is a segment: /v1//items/../p names /v1//p, not /v1/p, because a doubled slash is a different resource on any server that does not collapse it. A path ending in .. resolves to the directory above, trailing slash included. Two targets that differ only in an empty segment therefore stay two pages, and a pull that visits both is not a continuation cycle.

Link continuation metadata is parsed and authorized only when pagination.strategy is link_header. Other strategies ignore it because their continuation authority comes from the configured offset, cursor, or single-request contract.

A Link header is read as bytes, so a parameter this reader never consults — a title or a type carrying an accented character, an emoji, or anything else outside ASCII — does not affect the pull. Only the target inside <…> is decoded, because only the target has to become a URL: a target that is not valid UTF-8 is reported as malformed metadata. When a header carries several comma-separated links and one of them cannot be parsed, the rest are still read, so a reply naming two different next pages is reported as the conflict it is rather than as unreadable metadata.

Authentication

auth.scheme selects the credential sent on every request:

  • none (default) — no auth header.

  • bearer — sends Authorization: Bearer <token>.

  • header — sends an arbitrary static header, e.g. an API key.

    auth:
      scheme: header
      name: X-API-Key
      value: "${API_KEY}"
    

Corporate proxies

REST sources use the process proxy environment: ALL_PROXY, HTTPS_PROXY, or HTTP_PROXY (including their lowercase forms), with NO_PROXY/no_proxy for bypass rules. This lets the standalone clinker run CLI reach a third-party vendor through a corporate forward proxy without a central orchestrator or a pipeline-specific proxy key. Keep proxy credentials out of pipeline YAML and avoid printing credential-bearing proxy URLs.

The current Rust TLS configuration trusts the bundled public Web PKI roots. A proxy that tunnels HTTPS works when the vendor certificate remains visible and chains to those roots. A TLS-inspecting proxy that substitutes a certificate from a private corporate CA is not currently supported by a Clinker trust-store setting, even if that CA is installed in the operating-system store. In that case Clinker fails closed with a TLS/proxy classification; do not disable certificate verification as a workaround.

Reliability and finiteness knobs

KeyDefaultMeaning
max_pages—Required. Hard ceiling on pages fetched, regardless of the server.
max_recordsnoneOptional hard ceiling on records emitted.
retries3Bounded retries on a transient failure (5xx, connect/timeout error, or a transient body-delivery timeout/reset). A 4xx is fatal — retrying cannot help.
timeout_secs30Per-request timeout. Bounds in-flight time so an interrupt lands within the shutdown window.

Request diagnostics report a failure class, attempt number, page number, HTTP status when available, and the query-free target path. Authorization headers, request bodies, response bodies, and URL query values are never included. A proxy or vendor error therefore remains actionable without copying credentials or signed query parameters into logs.

A retryable body-delivery failure discards the partial body and retries the whole page within the same bounded retry budget. What counts as retryable is one rule for the whole request, whichever phase observed the failure: a dropped, reset, or timed-out exchange is retried, and a failure that would arrive identically on every attempt is not. Body-size violations, TLS failures, an unroutable URL, a host that does not resolve, and local material this process cannot read are therefore fatal at once rather than retried — reported against attempt 1, and without spending timeout_secs once per remaining attempt on a condition that cannot change. retries is a ceiling on attempts worth making, not a number of attempts every failure receives.

A partial-page decode failure routes that page’s offending rows to the DLQ per-row, exactly like a file source; it does not abort the pull.

Shutdown

On SIGINT/SIGTERM the reader polls its cancellation handle at each page boundary and stops cleanly with a normal end-of-input — the same graceful drain a file source performs. The timeout_secs per-request bound caps how long a single in-flight request can delay that stop.

A request already in flight when the signal arrives is reported as a cancellation too, not as a failing endpoint — including a page whose body was being read when the connection dropped, which the reader would otherwise have retried. What the signal took away was that retry, so the run’s outcome is the cancellation. An endpoint that would have failed identically however many attempts remained is still reported as that endpoint failure, because no retry was lost: a supervisor re-queuing the batch would only repeat it.