A Combine walks one input, the driver, a row at a time and looks each row up in the other inputs, the build side. Three settings decide what comes out: where: picks the build rows that match a driver row, match: decides what to do with several matches, and on_miss: decides what happens when there are none. Change any of them here and watch every driver row's result.
The product list has two rows for P1 and two for P4, so a lookup can find more than one match. P2 is inactive, and P3's pack isn't a number. Order O3 asks for P9, which doesn't exist.
- type: combine
name: enrich
input:
orders: orders # first entry drives
products: products
config:
where:
match:
on_miss:
propagate_ck: driver
cxl:
cxl: body must be empty.where: failed to evaluate is dead-lettered with that build row. In both cases on_miss never fires, because the driver wasn't shown to have no match. Try the “+ pack check” condition, or turn the body filter on.A where: built from <, <=, > or >= across the two inputs is a range join. Here each order finds its price tier. Mid and Promo overlap, so an amount between 100 and 130 matches both.
where: "orders.amount >= tiers.lo
and orders.amount < tiers.hi"
on_miss. A tier with an empty bound can match nothing.E327).max_output_rows to stop the run (E325) instead of writing a huge result.The first entry in input: drives unless drive: names another. The output has one row per driver row (with first or collect), so pick the side you want to iterate over. The other side is held in memory, so plan for about 1.5–2× its file size.
The planner picks the join method from the shape of where:, before the run starts: a hash join for equal ids, IEJoin for ranges. For equal ids it uses the disk-spilling grace hash join only if its size estimate says the build side is too big, or if you set strategy: grace_hash. An in-memory join whose build side then outgrows the memory budget stops with E310; it doesn't switch to disk partway (#1337). Every method gives the same rows.
where: needs at least one comparison between the inputs, either equality or a range. A condition that only filters one side is rejected (E313). Extra conditions after the first one are checked on each match.
Each pair of inputs that should be related needs its own equality, joined with and. The planner chains the joins; max_output_rows limits the final output.
propagate_ck: is required and says which correlation keys the output carries. A correlation key also sorts its Source, which changes what “first” picks. See the correlation keys explainer.
When the build side is far over the memory limit and can't be split small enough, the Combine works through it in chunks and currently decides first, collect and on_miss per chunk, so a driver can get more than one first row. The Combine Nodes page describes when this can happen.