Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Types & Literals

CXL has 10 value types. Every field value, literal, and expression result is one of these types.

Value types

TypeRust backingDescription
NullValue::NullMissing or absent value
Boolbooltrue or false
Integeri6464-bit signed integer
Floatf6464-bit double-precision float
DecimalDecimalExact base-10 fixed-point number for money/financials
StringFieldStrUTF-8 text
DateNaiveDateCalendar date without timezone
DateTimeNaiveDateTimeDate and time without timezone
ArrayOwnedValuesOrdered collection of values
MapOwnedMapKey-value pairs

Literal syntax

Integers

Standard decimal notation. Negative values use the unary minus operator.

$ cxl eval -e 'emit a = 42' -e 'emit b = -5' -e 'emit c = 0'
{
  "a": 42,
  "b": -5,
  "c": 0
}

Floats

Decimal notation with a dot. Must have digits on both sides of the decimal point.

$ cxl eval -e 'emit a = 3.14' -e 'emit b = -0.5'
{
  "a": 3.14,
  "b": -0.5
}

Strings

Double-quoted or single-quoted. Supports escape sequences: \\, \", \', \n, \t, \r.

$ cxl eval -e 'emit greeting = "hello world"'
{
  "greeting": "hello world"
}

Booleans

The keywords true and false.

$ cxl eval -e 'emit flag = true' -e 'emit neg = not flag'
{
  "flag": true,
  "neg": false
}

Dates

Hash-delimited ISO 8601 format: #YYYY-MM-DD#.

$ cxl eval -e 'emit d = #2024-01-15#'
{
  "d": "2024-01-15"
}

Null

The keyword null.

$ cxl eval -e 'emit nothing = null'
{
  "nothing": null
}

Arrays and comprehensions

Array literals accept full expressions and preserve their written order:

emit values = [order_id, amount * 2, null]

Use one for clause and an optional trailing if to construct an array from another array:

emit positive_doubles = [item * 2 for item in values if item > 0]

The source must be an array; null and scalar sources are errors. The binding is local to the item expression and predicate, cannot destructure, and cannot shadow an input field or surrounding let binding.

Maps

Map literals preserve key insertion order. Bare identifiers and quoted strings are static keys; brackets hold a computed expression whose result must be a non-null string:

emit payload = {
  customer: customer_name,
  items: [{sku: item.sku, quantity: item.quantity} for item in line_items],
  [dynamic_key]: dynamic_value,
}

Duplicate keys are errors, including two differently escaped spellings that decode to the same logical key. Nested maps and arrays are limited to 64 container levels, and CXL construction is limited to 10 MiB per input record; both limits fail the record instead of growing without bound.

Nested keys use one canonical escape grammar. After CXL string decoding, one leading backslash makes a reserved-looking key literal: \@name, \#text, or \\name. In CXL source each backslash in a quoted string is itself escaped, so write "\\@name", "\\#text", or "\\\\name". Other leading-backslash forms are rejected. Output formats decide how the neutral nested value is encoded. JSON removes the structural escape when writing the key; XML assigns roles to the unescaped forms as described in Writing XML.

The runnable examples/pipelines/nested_values.yaml pipeline sends one constructed value to both JSON and XML so the two native encodings can be compared directly.

Schema types

When declaring column types in YAML pipeline schemas, use these type names:

Schema typeCXL typeDescription
stringStringText values
intInteger64-bit integers
floatFloat64-bit floats
decimalDecimalExact base-10 fixed-point (money) — see below
boolBoolBoolean values
dateDateCalendar dates
date_timeDateTimeDate and time
arrayArrayOrdered collections
numericInt or FloatUnion type – accepts either
anyAnyUnknown type – no type constraints
nullable(T)Nullable(T)Wrapper – value may be null

Example YAML schema declaration:

schema:
  employee_id: int
  name: string
  salary: nullable(float)
  start_date: date

Type promotion

CXL automatically promotes types in mixed expressions:

Int + Float promotes to Float:

$ cxl eval -e 'emit result = 2 + 3.5'
{
  "result": 5.5
}

Null + T produces Nullable(T): Any operation involving null produces a nullable result.

$ cxl eval -e 'emit result = null + 5'
{
  "result": null
}

Nullable(A) + B unifies to Nullable(unified): When a nullable value meets a non-nullable value, the result type wraps the unified inner type in Nullable.

The decimal type

float is an IEEE-754 binary float: it cannot represent most base-10 fractions exactly, so 0.1 + 0.2 is 0.30000000000000004, not 0.3. That rounding is unacceptable for money. The decimal type is an exact base-10 fixed-point number — 0.10 + 0.20 is exactly 0.30 — and is the correct type for monetary amounts, prices, tax, and any figure that must round like decimal arithmetic on paper.

Declare a decimal column with type: decimal and a scale (the number of fractional digits). precision (total significant digits) is optional validation metadata:

schema:
  - { name: amount, type: decimal, scale: 2 }
  - { name: tax_rate, type: decimal, scale: 4 }

A decimal column parses its raw text into an exact value and rounds off any excess precision to the column scale (round-half-to-even, the unbiased “banker’s rounding” used in accounting), so a scale: 2 column stores 2.567 as 2.57. This is one edge of a boundary contract: a declared scale pins a value to that many places at the boundary it is declared on — a source column’s scale on read, an output column’s scale on write (see Aggregating decimals) — while decimals keep full precision inside the pipeline.

Arithmetic rules

  • decimal ⊗ decimal → decimal — exact.
  • decimal ⊗ int → decimal — the integer widens exactly, so amount + 1 and price * quantity stay exact decimals.
  • decimal ⊗ float is a type error. Mixing an exact decimal with a binary float would silently lose precision, so CXL rejects it and asks for an explicit cast. Choose the trade-off deliberately:
    • amount.to_float() * rate — opt into binary float precision.
    • rate.to_decimal() * amount — bring the float into exact decimal math (the float→decimal step is the one acknowledged lossy conversion).
  • Division and avg compute at full precision — the exact quotient, not a binary-float approximation. Inside the pipeline a computed decimal keeps every digit; it is pinned to a fixed number of places only at a boundary that declares a scale. Declaring the output column type: decimal with a scale rounds the value to that many places on write (banker’s rounding), exactly as a decimal source column rounds on read — so avg(amount) emitted into a scale: 2 output column writes 1.33, not the full quotient. With no declared output scale the full precision is preserved; use an explicit round in CXL when you need fixed places mid-pipeline.

Comparisons follow the same rule: decimal < int is fine, decimal < float requires a cast.

The branches of a conditional follow it too. An if, a match or a ?? whose branches are a decimal and a float does not compile, because its result would be a decimal on some rows and a float on others:

cannot mix decimal and float without an explicit cast: the branches of this `if` are a decimal (`amount`) and a float (`price`); declare `price` a decimal in its Source schema, `type: decimal` in place of `type: float`, so the branches have one numeric type

The message gives one fix. When the float is a Source column, declare it type: decimal in its Source schema: the reader parses the column’s text exactly, so if flag then amount else price is a decimal holding the values the file holds. (A JSON number read into a decimal column is still parsed through a float first; see #1299.) When the float is computed rather than read from a Source column, convert the decimal side instead: if flag then amount.to_float() else price * 2.0 is a float. Converting a float with .to_decimal() does not make it exact: the decimal keeps the float’s binary digits.

Casting

x.to_decimal() converts an int, string, or float into a decimal (try_decimal is the lenient form that yields null on failure). d.to_int(), d.to_float(), and d.to_string() convert a decimal back out.

Worked example — an exact invoice total

$ cxl eval -e 'emit total = ("19.99".to_decimal() * 3) + "4.80".to_decimal()'
{
  "total": "64.77"
}

19.99 * 3 = 59.97, + 4.80 = 64.77 — exact, with no binary-float drift. (In a pipeline, declare the source columns type: decimal instead of casting; JSON output renders a decimal as a scale-preserving string.)

Aggregating decimals

sum, avg, min, max, count, and distinct all work over a decimal column and stay exact — no binary float ever touches a running total:

  • sum(amount) returns a decimal: the exact total of the group’s values at the largest scale among them, rounded once (half to even) only when it does not fit a decimal at that scale. The sum of 1.00, -1.00 and 2 is 2.00 in any order, and a total of amounts with two decimal places is exact to the cent.
  • avg(amount) is sum(amount) / count(amount), a decimal at full division precision, and weighted_avg(v, w) is sum(v * w) / sum(w): the same digits and scale as those expressions give.
  • Only a group with no non-null value gives null. A decimal total outside the decimal range, a weighted_avg whose weights total zero or whose row product is out of range, and a group holding both decimals and floats each fail the group with an aggregate_finalize error that names the fix (see Aggregate functions).
  • min / max return the exact extremum, and count returns an integer.
  • Group-by and distinct keys are scale-normalized: two decimals that are numerically equal group together regardless of scale, so 2.50 and 2.5 fall in one group. This holds even when the aggregation spills to disk.

To pin an aggregate result to fixed places on the way out, declare the output column type: decimal with a scale: the value is rounded to that scale on write (banker’s rounding), so avg(amount) over 1.00, 1.00, 2.00 writes 1.33 into a scale: 2 output column while sum(amount) stays 4.00. An output column with no declared scale keeps the full-precision quotient. This applies to every format — CSV, JSON, and fixed-width — because the rounding happens as the record is projected onto the output. For fixed-width output it is often required: a full-precision quotient overflows a narrow numeric field, which is a hard error, whereas the rounded value fits.

weighted_avg also stays exact over decimals: a decimal value or weight (or both) gives sum(value * weight) / sum(weight) over exact totals, at full division precision. A zero total weight is an error, as x / 0 is. A decimal in one position mixed with a binary float in the other is a type error, matching the decimal ⊗ float arithmetic rule. Declare a float Source column type: decimal so the value and weight share one numeric domain, or, when the float is computed, convert the decimal argument with .to_float().

Type unification rules

When two types meet in an expression, CXL coerces them automatically:

  • Numbers combine: mixing an integer and a float gives a float (2 + 3.5 is 5.5).
  • A decimal and a float never combine, in an operator or in the branches of an if, match or ??: declare a float Source column type: decimal, or convert the decimal side with .to_float() when the float is computed.
  • Arithmetic and ordering comparisons with null give null. == and != never do (null == null is true), and and/or give a definite answer when the other side settles it. See Null Handling.
  • Mismatched types are an error: String + Int fails. Convert first with .to_int() or .to_string() so both sides are the same type.