Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 63 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,69 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr

---

## [0.1.0-preview.0.3.0] - 2026-05-25

Adds the per-row arithmetic and date/datetime math operators to the `each*` family established in `0.2.0`. Together with the existing per-row boolean/comparison operators, this closes the gap that previously made per-row computed columns inexpressible (e.g. `unit_price * quantity AS subtotal`, `sum(unit_price * quantity)`, `order_date + 30 days`, `shipped_at - ordered_at > 48h`). Purely additive on top of `0.2.0`.

Tracking issue: #47.

### Added

- **Per-row numeric arithmetic**: `eachAdd`, `eachSubtract`, `eachMultiply`, `eachDivide`. `values` is an array of `numericReturning | numericArrayReturning` items (broadcast scalars, zip arrays), `minItems: 2`. Result is `numericArrayReturning`. Mirrors the existing single-value `arithmetic` family but accepts fields directly. Slots into `select` (computed columns), aggregate `arg`, and the right operand of any `each*` comparison. Added to `numericArrayReturning`.
- **Per-row date math**:
- `eachDateAddDays` — `{ left: date \| dateArray, right: number \| numberArray }` → `dateArrayReturning`. Adds N days per row.
- `eachDateDiffDays` — `{ left: date \| dateArray, right: date \| dateArray }` → `numericArrayReturning`. Difference in days.
- **Per-row datetime math**:
- `eachDatetimeAddSeconds` — `{ left: datetime \| datetimeArray, right: number \| numberArray }` → `dateTimeArrayReturning`. Adds N seconds per row.
- `eachDatetimeDiffSeconds` — `{ left: datetime \| datetimeArray, right: datetime \| datetimeArray }` → `numericArrayReturning`. Difference in seconds.
- **Per-row time math**:
- `eachTimeAddSeconds` — `{ left: time \| timeArray, right: number \| numberArray }` → `timeArrayReturning`. Adds N seconds per row; wrap/saturate/error behaviour around `00:00:00` is interpreter-defined.
- `eachTimeDiffSeconds` — `{ left: time \| timeArray, right: time \| timeArray }` → `numericArrayReturning`. Difference in seconds.
- New schema groups `eachArithmetics`, `eachDateArithmetics`, `eachTimeArithmetics`, `eachDatetimeArithmetics`, plus the `eachArithmetic` union. The three diff operators are referenced from `numericArrayReturning`; `eachDateAddDays` from `dateArrayReturning`; `eachTimeAddSeconds` from `timeArrayReturning`; `eachDatetimeAddSeconds` from `dateTimeArrayReturning`.
- Five new reference samples:
- `17_each_arithmetic_select.json` — `eachMultiply(unit_price, quantity)` as a computed `select` column.
- `18_aggregate_of_each_multiply.json` — `sum(eachMultiply(unit_price, quantity))` grouped by user — the textbook line-item revenue query.
- `19_each_date_add_days.json` — `eachDateAddDays(order_date, 30)` derives a `delivery_eta` column.
- `20_each_datetime_diff_where.json` — `eachDatetimeDiffSeconds` inside `eachGreaterThan` filters orders whose ship time exceeds 48 hours.
- `21_each_time_math.json` — `eachTimeAddSeconds` and `eachTimeDiffSeconds` over `clock_in` / `clock_out` time fields.

### Changed

- **`CLAUDE.md` rewritten around the "two operator families" model.** The old "fields cannot appear in arithmetic / comparison" rules — outdated since `0.2.0` and now wrong for arithmetic too — were replaced with explicit guidance on when to use each family. A new "Where each family fits" table maps clauses to accepted families. The "Field equality uses `arrayEquality`" rule is now "prefer `eachEqual`; `arrayEquality` is reserved for whole-sequence equality".
- **README "Operations" section** extended with per-operator subsections for `eachAdd`/`eachSubtract`/`eachMultiply`/`eachDivide`, `eachDateAddDays`/`eachDateDiffDays`, `eachDatetimeAddSeconds`/`eachDatetimeDiffSeconds`. The "Two families" overview now lists arithmetic alongside booleans/comparisons.

### Interpreter notes

- **Result type of per-row arithmetic**: every `eachX` arithmetic / date / time / datetime operator returns a numeric / date / time / datetime *column* aligned with the current row set. Maps to SQL projection expressions, LINQ row lambdas, or in-memory column transforms — same model as `0.2.0`'s `each*` comparisons.
- **Broadcast vs zip semantics (mixed `*Returning` and `*ArrayReturning` operands)**: every `each*` slot that admits both kinds uses one shared evaluation rule:
1. The surrounding row set fixes a row count `N` — determined by `from` + `joins` + `where` for `where` / `select` / `join.on` expressions, or by the group size for expressions nested inside an aggregate `arg` after `groupBy`.
2. Each `*Returning` operand is **broadcast** — repeated `N` times so it has one value per row. (In SQL backends, this is implicit — a scalar in a projection IS the value for every row.)
3. Each `*ArrayReturning` operand is already aligned with the same `N` rows by construction (it comes from the same row set / group).
4. The operator runs **element-wise across all operands**, producing a length-`N` result vector.

So `eachAdd([fieldA, scalar, fieldB])` over 3 rows with `fieldA = [10, 20, 30]`, `scalar = 5`, `fieldB = [1, 2, 3]` evaluates to `[16, 27, 38]`. This is what makes `eachMultiply(unit_price, 1.05)` (5% per-row markup) and `eachAdd(base_price, tax, shipping)` (sum three columns per row) work in one place.

The return type is always `*ArrayReturning` even if every operand is a `*Returning`: `eachAdd(2, 3)` in a `select` over a 4-row table yields `[5, 5, 5, 5]`, not `5`. Interpreters should not optimize this away as a constant — the row-count alignment is structural. Use single-value `add` for purely scalar work.

Backend implementation hints:
- SQL: scalars become literals in the projection / `WHERE` predicate; `*ArrayReturning` operands become column references. The SQL engine handles broadcast natively. No special case needed.
- LINQ / row-lambda backends: emit `row => operator(arg1(row), arg2(row), …)` where each `arg_i` resolves to either a captured constant (broadcast) or a row-accessor (zip).
- Column-vector backends (in-memory analytics): determine `N` from the parent context; materialize each scalar as a length-`N` vector via broadcast; then run the operator element-wise.
- **Subtract / divide ordering**: same as their single-value siblings — left-to-right fold over the `values` array. For `eachSubtract([a, b, c])` over `N` rows, position `i` evaluates `a[i] - b[i] - c[i]`.
- **Aggregate over per-row arithmetic**: `sum` / `min_*` / `max_*` / `average_*` already accepted `*ArrayReturning` as `arg`; with `eachX` now in those unions, the same dispatch covers expressions like `sum(eachMultiply(field_a, field_b))` without additional cases.
- **Unit choice for date / time / datetime math**: days for `date`; seconds for `time` and `datetime`. Larger units are obtained by composition — e.g. "add N hours to a datetime" is `eachDatetimeAddSeconds(dt, eachMultiply(n, 3600))`. There is intentionally no `interval` type and no per-unit operator family; interpreters should not emulate one. Negative `right` values produce subtraction.
- **Diff operator direction**: `eachDateDiffDays` / `eachTimeDiffSeconds` / `eachDatetimeDiffSeconds` all compute `left - right`, so a positive result means `left` is later. The schema fixes this convention.
- **`eachTimeAddSeconds` overflow**: time-of-day is bounded (`00:00:00`–`23:59:59.…`). Wrap / saturate / error behaviour around the day boundary is interpreter-defined — pick one and document it.
- **Division by zero / numeric overflow / null propagation**: schema doesn't constrain. Define per-backend.
- **No `each*` arithmetic on `string`**: out of scope for this release; addable later without breakage.
- **CLAUDE.md interpreter rule**: any backend that previously rejected fields under `add` / `multiply` / etc. should keep rejecting them there — the *single-value* arithmetic family is unchanged. Fields now flow exclusively through the per-row family.

### Versioning

Minor preview bump (`0.1.0-preview.0.2.0` → `0.1.0-preview.0.3.0`). All schema changes are additive; no existing operator or shape was renamed, removed, or tightened. Both `version` and `$id` updated.

---

## [0.1.0-preview.0.2.0] - 2026-05-25

This release establishes the **per-row predicate family (`each*`)** as a first-class, type-distinct set of expressions. All `each*` operators now return a per-row boolean *column* (`booleanArrayReturning`), not a single boolean — making PureQL's type system express the row-vs-scalar distinction that SQL hides. This is purely additive on top of `0.1.0-preview.0.1.0`: no operator existing in that release was removed or renamed in a user-visible way.
Expand Down
78 changes: 65 additions & 13 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,33 +32,85 @@ print('valid')

## Critical design rules (read before editing samples)

### Fields live in `arrayReturning`, not `singleValueReturning`
### Two predicate / expression families

A field reference (`{ entity, field, type }`) is an **array-returning** expression — it represents a whole column. Consequently:
Every operator in the schema belongs to one of two parallel families:

- Fields go in `select`, `groupBy`, `orderBy`, `join.on`, and as `arg` to aggregate functions.
- Fields **cannot** appear directly as operands of `arithmetic` (`add`, `multiply`, etc.) — use an aggregate like `sum` to reduce them first.
- Fields **cannot** appear as operands of `comparison` (`greaterThan`, etc.) — those operators accept `numericReturning` / `stringReturning` (scalars, aggregates, arithmetic), not field references.
- **Single-value family** (`and`, `or`, `not`, `equal`, `greaterThan`, …, `add`, `subtract`, `multiply`, `divide`) — operands and result are single values. Use when both sides reduce to one value per query (or per group, inside `having`): typically scalars, parameters, aggregates, or arithmetic over them.
- **Per-row (`each*`) family** (`eachAnd`, `eachOr`, `eachNot`, `eachEqual`, `eachGreaterThan`, …, `eachAdd`, `eachMultiply`, `eachDateAddDays`, `eachDatetimeDiffSeconds`, …) — operate per row of the current row set; result is a vector aligned with the input rows. Use when at least one operand is a field, or to build computed per-row columns.

### Field equality uses `arrayEquality`
**Do not mix families inside the same boolean operator.** `and`/`or`/`not` take only single-boolean children; `eachAnd`/`eachOr`/`eachNot` take only per-row boolean children.

Comparing a field to a literal requires the literal to be an **array scalar**:
### Where each family fits

| Clause | Accepts |
|---|---|
| `where` | single-value boolean **or** per-row boolean (per-row is the common case) |
| `join.on` | single-value boolean **or** per-row boolean (per-row equi-join is the common case) |
| `having` | single-value boolean **only** — operands must reduce to one value per group |
| `select` | any value-returning expression, including per-row computed columns |
| `groupBy` / `orderBy` | field references only |

### Fields are `arrayReturning`

A field reference (`{ entity, field, type }`) is an **array-returning** expression — it represents a whole column. Fields:

- Go in `select`, `groupBy`, `orderBy`, as `arg` to aggregate functions, and as operands of any `each*` operator.
- **Cannot** appear directly in single-value `add`/`multiply`/`greaterThan`/`equal`/etc. — use the per-row `eachX` variant for the row context, or reduce with an aggregate (`sum`, `count`, `max_*`, …) for the single-value context.

### Field equality: prefer `eachEqual` over `arrayEquality`

For "field = literal" or "field = field" filtering, use **`eachEqual`** with an unwrapped scalar on the right:

```json
{
"operator": "equal",
"operator": "eachEqual",
"left": { "entity": "users", "field": "status", "type": { "name": "string" } },
"right": { "type": { "name": "stringArray" }, "value": ["active"] }
"right": { "type": { "name": "string" }, "value": "active" }
}
```

The type on the right is `stringArray`, not `string`. This is `string_array_equality` under `arrayEquality`.
`arrayEquality` (the `equal` operator with `*ArrayReturning` on both sides) still validates and is semantically distinct — it asks "are these two **whole sequences** equal as wholes?" and returns one boolean. Reserve it for that intent (e.g. comparing two parameter arrays). The historical idiom of `equal(field, [singleValue])` as a per-row filter has been migrated out of all bundled samples.

### Per-row equality vs single-value equality

| Operator | Operands | Result | Typical placement |
|---|---|---|---|
| `equal` (single-value) | two `*Returning` | one boolean | `having` against aggregates |
| `equal` (whole-array) | two `*ArrayReturning` | one boolean | rare — whole-sequence equality |
| `eachEqual` | `*ArrayReturning` left, `*Returning` or `*ArrayReturning` right | one boolean per row | `where`, `join.on` |

### Range comparisons

Two parallel sets, same convention:

- Single-value: `greaterThan` / `lessThan` / `greaterThanOrEqual` / `lessThanOrEqual` over `numericReturning` / `stringReturning` / `dateReturning` / etc. Use in `having`.
- Per-row: `eachGreaterThan` / `eachLessThan` / `eachGreaterThanOrEqual` / `eachLessThanOrEqual`. `left` is `*ArrayReturning` (typically a field); `right` is `*Returning` (broadcast scalar) **or** `*ArrayReturning` (element-wise other field). Use in `where` / `join.on`.

### Arithmetic and date / time / datetime math

Same convention:

- Single-value `add` / `subtract` / `multiply` / `divide` — `values` items are `numericReturning` only. Use to combine aggregates and constants (e.g. `multiply(sum(total), 0.05)`).
- Per-row `eachAdd` / `eachSubtract` / `eachMultiply` / `eachDivide` — `values` items are `numericReturning | numericArrayReturning`. Use for computed per-row columns (e.g. `eachMultiply(unit_price field, quantity field)`).
- Date math: `eachDateAddDays(date, n_days) → date`, `eachDateDiffDays(date1, date2) → number`.
- Time math: `eachTimeAddSeconds(time, n_seconds) → time`, `eachTimeDiffSeconds(time1, time2) → number`.
- Datetime math: `eachDatetimeAddSeconds(datetime, n_seconds) → datetime`, `eachDatetimeDiffSeconds(dt1, dt2) → number`.

Unit choice: days for `date`, seconds for `time` and `datetime`. Larger units are expressed via composition with `eachMultiply` (e.g. `eachDatetimeAddSeconds(dt, eachMultiply(hours, 3600))`). No `interval` type exists. Diff operators evaluate `left - right`, so a positive result means `left` is later. `eachTimeAddSeconds` overflow / wrap behaviour around `00:00:00` is intentionally interpreter-defined.

### Broadcast vs zip: how mixed `*Returning` / `*ArrayReturning` operands evaluate

Every `each*` operator that admits both kinds on the same slot (`values` items in arithmetic, `left` / `right` in date/time math, `right` in comparisons, etc.) uses the same evaluation model:

`singleValueEquality` (with `stringReturning` / `numericReturning`) is for comparing aggregates or scalars to each other, not for field comparisons.
1. The surrounding row set fixes a row count `N` — determined by `from` + `joins` + `where` for `where` / `select` / `join.on` expressions, or by the group size for expressions inside an aggregate `arg` after `groupBy`.
2. Each `*Returning` (single-value) operand is **broadcast** — conceptually repeated `N` times so it has one value per row.
3. Each `*ArrayReturning` operand is already aligned with the same `N` rows by construction (it comes from the same row set).
4. The operation runs **element-wise across all operands**, producing a length-`N` result vector.

### Range comparisons only on single-value expressions
So `eachAdd([fieldA, scalar, fieldB])` over 3 rows with `fieldA = [10, 20, 30]`, `scalar = 5`, `fieldB = [1, 2, 3]` evaluates to `[16, 27, 38]`. Mixing kinds is intended and is how common patterns are expressed: `eachMultiply(unit_price, 1.05)` adds a 5% per-row markup; `eachAdd(base_price, tax, shipping)` sums three columns per row.

`greaterThan` / `lessThan` etc. operate on `numericReturning` / `stringReturning` / `dateReturning` / etc. — scalars, parameters, and aggregates only. Use them in `having` to filter groups by aggregate results.
This is why `eachX` operators do not collapse to their single-value siblings when handed only scalar operands: the *return type* is still `*ArrayReturning`, aligned with the row set. `eachAdd(2, 3)` in a `select` over a 4-row entity yields the vector `[5, 5, 5, 5]`, not the scalar `5`. Use single-value `add` for purely scalar work; `eachAdd` exists specifically because at least one operand is row-aligned.

### `joinItem` has no alias

Expand Down
Loading
Loading