Network Sources (REST)
A Source reads from the filesystem by default. To pull records from a
network endpoint instead, declare a transport: block on the Source. The
transport selects where records come from; it sits above the on-disk
type: (the format), which for a REST source still selects how the
response bodies decode.
A network transport is a finite-pull source: it runs on its own
thread, drives a synchronous client to cursor exhaustion, then exits.
There is no daemon, no event loop, and no async runtime — the same single-
process, run-to-drain model as a file pipeline. Finiteness is a hard
property of the reader: a REST source enforces explicit page and record
limits, so an unbounded endpoint cannot keep it running forever. If the
server offers another continuation after max_pages, the source fails
closed instead of reporting a truncated pull as successful completion.
A network source still requires a schema: block. That authored
schema is the row-to-record target: the reader maps each decoded object
onto it, coercing values leniently. A per-row value that cannot coerce is
left unchanged at the reader and routed to the dead-letter queue at the
Transform stage — identical to file-source semantics. A network source
declares no file matcher (path / glob / regex / paths);
declaring one is a configuration error (E219).
Because a network source has no file path, its $source.file provenance
column and the {source_file} output template both resolve to a stable
synthetic identifier, <source:NAME>, where NAME is the Source node’s
name.
REST sources
A rest source issues paginated HTTP GETs against a base URL, decoding
each response body through the declared json or xml format. (Other
formats are rejected with E220 — a REST body is a multi-record
document, not a flat CSV/fixed-width stream.)
nodes:
- type: source
name: orders_api
config:
name: orders_api
type: json
options:
format: array # each page body is a JSON array of objects
transport:
kind: rest
url: https://api.example.com/v1/orders
max_pages: 50 # HARD page cap — required
pagination:
strategy: link_header
auth:
scheme: bearer
token: "${ORDERS_TOKEN}"
schema:
- { name: order_id, type: int }
- { name: total, type: float }
- { name: placed_at, type: date_time }
Pagination strategies
The pagination.strategy selects how the reader advances pages and
detects the last one. max_records bounds emitted records. max_pages
bounds requests and requires the server to reach an actual terminal page;
an offered continuation beyond that bound is an error.
-
none(default) — a singleGET; the body is the whole result. -
offset—?offset=N&limit=L, advancing the offset by the page size each request. The last page is the one that returns fewer rows thanlimit. A page containing exactlylimitrows is not proof of end-of-input, so Clinker must issue one more bounded request to observe a short or empty terminal page. Sizemax_pagesto leave room for that probe; reaching the cap first fails withpage_limit_reachedinstead of accepting a possibly truncated result.pagination: strategy: offset limit: 200 offset_param: offset # optional, defaults shown limit_param: limit -
cursor_token— the reader reads a continuation token from a JSON pointer in each response and sends it back on the next request. Paging stops when the token field is absent or null.pagination: strategy: cursor_token cursor_param: page_token next_token_pointer: /meta/next_page # RFC 6901 JSON pointer -
link_header— the reader follows the URL in the response’s RFC 8288Link: <…>; rel="next"header until no such link is present. Registered relation tokens are case-insensitive, sonext,Next, andNEXThave the same meaning.pagination: strategy: link_header
Continuation and redirect safety
Every server-directed continuation or redirect is resolved against the
effective response URL and normalized before another request is built. Only
the original normalized origin is allowed. Cross-origin targets, HTTPS-to-HTTP
downgrades, malformed or conflicting rel="next" metadata, redirect or
continuation cycles, and traversal beyond the configured bounds fail before a
foreign or repeated request is sent.
Normalization resolves . and .. exactly as RFC 3986 does, and changes
nothing else about the path. An empty segment is a segment: /v1//items/../p
names /v1//p, not /v1/p, because a doubled slash is a different resource on
any server that does not collapse it. A path ending in .. resolves to the
directory above, trailing slash included. Two targets that differ only in an
empty segment therefore stay two pages, and a pull that visits both is not a
continuation cycle.
Link continuation metadata is parsed and authorized only when
pagination.strategy is link_header. Other strategies ignore it because
their continuation authority comes from the configured offset, cursor, or
single-request contract.
A Link header is read as bytes, so a parameter this reader never consults —
a title or a type carrying an accented character, an emoji, or anything
else outside ASCII — does not affect the pull. Only the target inside <…> is
decoded, because only the target has to become a URL: a target that is not
valid UTF-8 is reported as malformed metadata. When a header carries several
comma-separated links and one of them cannot be parsed, the rest are still
read, so a reply naming two different next pages is reported as the conflict
it is rather than as unreadable metadata.
Authentication
auth.scheme selects the credential sent on every request:
-
none(default) — no auth header. -
bearer— sendsAuthorization: Bearer <token>. -
header— sends an arbitrary static header, e.g. an API key.auth: scheme: header name: X-API-Key value: "${API_KEY}"
Corporate proxies
REST sources use the process proxy environment: ALL_PROXY, HTTPS_PROXY, or
HTTP_PROXY (including their lowercase forms), with NO_PROXY/no_proxy for
bypass rules. This lets the standalone clinker run CLI reach a third-party
vendor through a corporate forward proxy without a central orchestrator or a
pipeline-specific proxy key. Keep proxy credentials out of pipeline YAML and
avoid printing credential-bearing proxy URLs.
The current Rust TLS configuration trusts the bundled public Web PKI roots. A proxy that tunnels HTTPS works when the vendor certificate remains visible and chains to those roots. A TLS-inspecting proxy that substitutes a certificate from a private corporate CA is not currently supported by a Clinker trust-store setting, even if that CA is installed in the operating-system store. In that case Clinker fails closed with a TLS/proxy classification; do not disable certificate verification as a workaround.
Reliability and finiteness knobs
| Key | Default | Meaning |
|---|---|---|
max_pages | — | Required. Hard ceiling on pages fetched, regardless of the server. |
max_records | none | Optional hard ceiling on records emitted. |
retries | 3 | Bounded retries on a transient failure (5xx, connect/timeout error, or a transient body-delivery timeout/reset). A 4xx is fatal — retrying cannot help. |
timeout_secs | 30 | Per-request timeout. Bounds in-flight time so an interrupt lands within the shutdown window. |
Request diagnostics report a failure class, attempt number, page number, HTTP status when available, and the query-free target path. Authorization headers, request bodies, response bodies, and URL query values are never included. A proxy or vendor error therefore remains actionable without copying credentials or signed query parameters into logs.
A retryable body-delivery failure discards the partial body and retries the
whole page within the same bounded retry budget. What counts as retryable is
one rule for the whole request, whichever phase observed the failure: a
dropped, reset, or timed-out exchange is retried, and a failure that would
arrive identically on every attempt is not. Body-size violations, TLS
failures, an unroutable URL, a host that does not resolve, and local material
this process cannot read are therefore fatal at once rather than retried —
reported against attempt 1, and without spending timeout_secs once per
remaining attempt on a condition that cannot change. retries is a ceiling on
attempts worth making, not a number of attempts every failure receives.
A partial-page decode failure routes that page’s offending rows to the DLQ per-row, exactly like a file source; it does not abort the pull.
Shutdown
On SIGINT/SIGTERM the reader polls its cancellation handle at each
page boundary and stops cleanly with a normal end-of-input — the same
graceful drain a file source performs. The timeout_secs per-request
bound caps how long a single in-flight request can delay that stop.
A request already in flight when the signal arrives is reported as a cancellation too, not as a failing endpoint — including a page whose body was being read when the connection dropped, which the reader would otherwise have retried. What the signal took away was that retry, so the run’s outcome is the cancellation. An endpoint that would have failed identically however many attempts remained is still reported as that endpoint failure, because no retry was lost: a supervisor re-queuing the batch would only repeat it.