Clinker user guide · ← back to Combine Nodes

Combine
playground who meets whom

A Combine walks one input, the driver, a row at a time and looks each row up in the other inputs, the build side. Three settings decide what comes out: where: picks the build rows that match a driver row, match: decides what to do with several matches, and on_miss: decides what happens when there are none. Change any of them here and watch every driver row's result.

01

MatchingOrders looked up in a product list that repeats some ids

The product list has two rows for P1 and two for P4, so a lookup can find more than one match. P2 is inactive, and P3's pack isn't a number. Order O3 asks for P9, which doesn't exist.

- type: combine
  name: enrich
  input:
    orders: orders        # first entry drives
    products: products
  config:
    where: 
    match: 
    on_miss: 
    propagate_ck: driver
    cxl: 

match:

  • first (default): one row per driver, from the earliest matching build row in the order the build input delivered them. To pick a different one, sort the build input upstream.
  • all: one row per matching build row.
  • collect: one row per driver with every match gathered into an array. The cxl: body must be empty.

on_miss:

  • null_fields (default): keep the driver row; build fields are null.
  • skip: drop the driver row.
  • error: stop the run at the first driver row with no match.
Two things that are not a miss. A driver whose match the body then filters out produces no row, and the engine does not try the next match. A driver for which where: failed to evaluate is dead-lettered with that build row. In both cases on_miss never fires, because the driver wasn't shown to have no match. Try the “+ pack check” condition, or turn the body filter on.
where:
match:
on_miss:
drive:
Output
DLQ strategy: continue
02

Range joinsMatch on “between”, not on equal ids

A where: built from <, <=, > or >= across the two inputs is a range join. Here each order finds its price tier. Mid and Promo overlap, so an amount between 100 and 130 matches both.

where: "orders.amount >= tiers.lo
        and orders.amount < tiers.hi"
  • “First” still means the earliest tier in the tiers file, not the narrowest or the closest.
  • An order with an empty amount can't be compared, so it is treated as having no match and goes to on_miss. A tier with an empty bound can match nothing.
  • Range keys must be numbers, decimals, dates or datetimes. Comparing, for example, text with a number is rejected when the pipeline is checked (E327).
Range joins are where results can grow quickly: a loose condition can match almost every pair. Set max_output_rows to stop the run (E325) instead of writing a huge result.
match:
on_miss:
strategy: IEJoin (block-band), chosen because the condition is ranges only
Output
03

More to knowSettings and limits around the match

Which side drives

The first entry in input: drives unless drive: names another. The output has one row per driver row (with first or collect), so pick the side you want to iterate over. The other side is held in memory, so plan for about 1.5–2× its file size.

How it runs

The planner picks the join method from the shape of where:, before the run starts: a hash join for equal ids, IEJoin for ranges. For equal ids it uses the disk-spilling grace hash join only if its size estimate says the build side is too big, or if you set strategy: grace_hash. An in-memory join whose build side then outgrows the memory budget stops with E310; it doesn't switch to disk partway (#1337). Every method gives the same rows.

Conditions that can't run

where: needs at least one comparison between the inputs, either equality or a range. A condition that only filters one side is rejected (E313). Extra conditions after the first one are checked on each match.

Three or more inputs

Each pair of inputs that should be related needs its own equality, joined with and. The planner chains the joins; max_output_rows limits the final output.

Correlation keys

propagate_ck: is required and says which correlation keys the output carries. A correlation key also sorts its Source, which changes what “first” picks. See the correlation keys explainer.

A known gap

When the build side is far over the memory limit and can't be split small enough, the Combine works through it in chunks and currently decides first, collect and on_miss per chunk, so a driver can get more than one first row. The Combine Nodes page describes when this can happen.