From 07b74b91120084529e9b2cbab85ddce52d57ff7d Mon Sep 17 00:00:00 2001 From: Dmitry Kurochkin Date: Thu, 24 Sep 2026 13:18:53 +0000 Subject: [PATCH 1/4] docs(spec): define broadcast semantics for mixed select Without groupBy, a select that mixes array-returning items (fields, each* columns) with single-value items (aggregates, scalars) now has a defined result: N rows, with each single-value item repeated on every row. A select of only single-value items yields one row. Aggregates are computed over all N rows before distinct and pagination. Add sample 22 (grand total and share-of-total next to order fields) and document the rule in README, CLAUDE.md and CHANGELOG. Closes #52 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 4 ++++ CLAUDE.md | 2 ++ README.md | 12 ++++++++++-- samples/22_select_broadcast.json | 24 ++++++++++++++++++++++++ 4 files changed, 40 insertions(+), 2 deletions(-) create mode 100644 samples/22_select_broadcast.json diff --git a/CHANGELOG.md b/CHANGELOG.md index 5608f70..d5af685 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,10 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr ## [Unreleased] +### 🔧 Improvements + +- **Defined results for fields mixed with aggregates in `select`**: A query without `groupBy` can list fields next to aggregates. The aggregate's value is repeated on every row, e.g. an order total next to the grand total of all orders. Aggregates are always calculated over every matching row, before pagination. (#52) + ### ⚠️ Breaking Changes - **`having` now requires `groupBy`**: Queries that filter with `having` must also group their rows with `groupBy`. Previously the schema accepted `having` on its own, even though it has no meaning without groups. (#40) diff --git a/CLAUDE.md b/CLAUDE.md index 025e0a7..9001380 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -112,6 +112,8 @@ So `eachAdd([fieldA, scalar, fieldB])` over 3 rows with `fieldA = [10, 20, 30]`, This is why `eachX` operators do not collapse to their single-value siblings when handed only scalar operands: the *return type* is still `*ArrayReturning`, aligned with the row set. `eachAdd(2, 3)` in a `select` over a 4-row entity yields the vector `[5, 5, 5, 5]`, not the scalar `5`. Use single-value `add` for purely scalar work; `eachAdd` exists specifically because at least one operand is row-aligned. +The same model applies to `select` itself when it mixes kinds (without `groupBy`): if any item is `*ArrayReturning`, the result has `N` rows and every single-value item (e.g. `sum(total)`) is broadcast to each row; if all items are single-value, the result is one row. Aggregates see all `N` rows — they are computed before `distinct` and `pagination`. So `[field, sum(field)]` without `groupBy` is valid and well-defined, not an error. + ### `joinItem` has no alias Only the root `from` expression supports an `alias`. Joined entities are always referenced by their `entity` name string. diff --git a/README.md b/README.md index 895bac1..5b5d99c 100644 --- a/README.md +++ b/README.md @@ -91,7 +91,7 @@ Parameters are named placeholders resolved at execution time, analogous to prepa Each item in `select` is a value-returning expression (field, scalar, aggregate, arithmetic, boolean expression) with an optional `alias`. -When `groupBy` is present, `select` accepts **single-value expressions only** (aggregates, scalars, parameters, arithmetic over them). Fields and per-row `each*` columns are rejected by the schema, because a group has no single value for them. The `groupBy` fields are output automatically as the leading result columns, in `groupBy` order and named after the field, followed by the `select` entries: +**With `groupBy`**, `select` accepts **single-value expressions only** (aggregates, scalars, parameters, arithmetic over them). Fields and per-row `each*` columns are rejected by the schema, because a group has no single value for them. The `groupBy` fields are output automatically as the leading result columns, in `groupBy` order and named after the field, followed by the `select` entries: ```json "select": [ @@ -104,7 +104,14 @@ When `groupBy` is present, `select` accepts **single-value expressions only** (a Result columns: `user_id`, `order_count`. -Ungrouped example: +**Without `groupBy`**, `select` may mix single-value and array-returning items. The result shape is: + +- **At least one array-returning item** (field or `each*` column): the result has `N` rows, where `N` is the row count after `from` + `joins` + `where`. Every single-value item (aggregate, scalar, parameter, arithmetic) is **broadcast**: the same value is repeated in every row. If `N = 0`, the result is empty. +- **Only single-value items**: the result has exactly one row. + +Aggregates are computed over all `N` rows before `distinct` and `pagination` apply, so `take: 10` pages the rows but does not change `sum(...)`. See [`22_select_broadcast.json`](samples/22_select_broadcast.json). + +Example: `name` and `email` are fields, so `order_count` is repeated on every row: ```json "select": [ @@ -507,3 +514,4 @@ The [`samples/`](samples/) directory contains query examples ordered by complexi | [`19_each_date_add_days.json`](samples/19_each_date_add_days.json) | `eachDateAddDays` to derive a `delivery_eta` column from `order_date + 30 days` | | [`20_each_datetime_diff_where.json`](samples/20_each_datetime_diff_where.json) | `eachDatetimeDiffSeconds` inside `eachGreaterThan` to filter orders by ship-time | | [`21_each_time_math.json`](samples/21_each_time_math.json) | `eachTimeAddSeconds` (time + offset) and `eachTimeDiffSeconds` (shift duration) | +| [`22_select_broadcast.json`](samples/22_select_broadcast.json) | Fields next to an aggregate in `select` — `sum` broadcast to every row, plus `eachDivide(total, sum(total))` as a share-of-total column | diff --git a/samples/22_select_broadcast.json b/samples/22_select_broadcast.json new file mode 100644 index 0000000..2ccac78 --- /dev/null +++ b/samples/22_select_broadcast.json @@ -0,0 +1,24 @@ +{ + "from": { "entity": "orders" }, + "select": [ + { "entity": "orders", "field": "id", "type": { "name": "uuid" } }, + { "entity": "orders", "field": "total_amount", "type": { "name": "number" } }, + { + "operator": "sum", + "arg": { "entity": "orders", "field": "total_amount", "type": { "name": "number" } }, + "alias": "grand_total" + }, + { + "operator": "eachDivide", + "values": [ + { "entity": "orders", "field": "total_amount", "type": { "name": "number" } }, + { + "operator": "sum", + "arg": { "entity": "orders", "field": "total_amount", "type": { "name": "number" } } + } + ], + "alias": "share_of_total" + } + ], + "pagination": { "skip": 0, "take": 10 } +} From 1583d6e9f736717a3afbeff4a8eeb38687ef5f19 Mon Sep 17 00:00:00 2001 From: Dmitry Kurochkin Date: Thu, 24 Sep 2026 13:24:30 +0000 Subject: [PATCH 2/4] feat(schema)!: sort grouped queries by single-value expressions Add expressionOrderByItem ({ expression: singleValueReturning, direction }). With groupBy, orderBy items must be expressionOrderByItem (root dependentSchemas); without groupBy they must be orderByItem (root anyOf). A group has no single value for a field, so sorting groups by a field is rejected; sort by an aggregate instead, and wrap group keys in an aggregate such as min_string. Migrate sample 12 to sort by sum(total_amount) desc, then min_string(users.name) asc, and document the rule in README, CLAUDE.md and CHANGELOG. BREAKING CHANGE: grouped queries that sort by a field no longer validate; use an aggregate expression instead. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + CLAUDE.md | 12 +++++++-- PureQL-Specification.json | 49 ++++++++++++++++++++++++++++++++++- README.md | 20 ++++++++++---- samples/12_complex_query.json | 12 ++++++++- 5 files changed, 85 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index d5af685..709a0c7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,7 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr - **`having` now requires `groupBy`**: Queries that filter with `having` must also group their rows with `groupBy`. Previously the schema accepted `having` on its own, even though it has no meaning without groups. (#40) - **`groupBy` can no longer be empty**: `groupBy` must list at least one field. To skip grouping, leave the clause out. - **Grouped queries select only single values**: When a query has `groupBy`, `select` accepts only aggregates, scalars, parameters and arithmetic over them. The `groupBy` fields now appear in the result automatically as the first columns, so remove them from `select`. Fields that aren't grouped and per-row `each*` columns are rejected, because a group has no single value for them. (#52) +- **Grouped queries sort by expressions**: When a query has `groupBy`, each `orderBy` item holds an `expression` (usually an aggregate such as `sum(...)`) in place of a `field`. To sort groups by a key, wrap it in an aggregate, e.g. `min_string(users.name)`. Queries without `groupBy` keep sorting by `field`. --- diff --git a/CLAUDE.md b/CLAUDE.md index 9001380..a1dccea 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -49,7 +49,8 @@ Every operator in the schema belongs to one of two parallel families: | `join.on` | single-value boolean **or** per-row boolean (per-row equi-join is the common case) | | `having` | single-value boolean **only** — operands must reduce to one value per group | | `select` | any value-returning expression, including per-row computed columns — **single-value only when `groupBy` is present** | -| `groupBy` / `orderBy` | field references only | +| `groupBy` | field references only | +| `orderBy` | `{ field, direction }` without `groupBy`; `{ expression: , direction }` with `groupBy` | ### Fields are `arrayReturning` @@ -118,10 +119,17 @@ The same model applies to `select` itself when it mixes kinds (without `groupBy` Only the root `from` expression supports an `alias`. Joined entities are always referenced by their `entity` name string. -### `groupBy` and `orderBy` take `field` objects +### `groupBy` takes `field` objects Not select expressions — just plain `{ entity, field, type }` field references. No aliases, no operators. +### `orderBy` shape depends on `groupBy` + +- Without `groupBy`: items are `orderByItem` — `{ field, direction }`. Expression items are rejected (root `anyOf`). +- With `groupBy`: items are `expressionOrderByItem` — `{ expression, direction }` where `expression` is `singleValueReturning` (typically an aggregate). Field items are rejected (root `dependentSchemas`). +- Never reference a `select` alias from `orderBy` — repeat the expression. Aliases are name references the schema cannot check. +- To sort groups by a key, wrap it in an aggregate (e.g. `min_string(users.name)`): every value in a group equals the key. + ### Grouped `select` is single-value only When `groupBy` is present, every `select` item must be `singleValueReturning` (enforced via root `dependentSchemas`). Group keys are emitted automatically as the leading result columns, so **never repeat `groupBy` fields in `select`** — the schema rejects them, along with any non-grouped field or `each*` column. diff --git a/PureQL-Specification.json b/PureQL-Specification.json index 0ebea87..d50aff9 100644 --- a/PureQL-Specification.json +++ b/PureQL-Specification.json @@ -2181,6 +2181,25 @@ } } }, + "expressionOrderByItem": { + "type": "object", + "required": [ + "expression" + ], + "properties": { + "expression": { + "$ref": "#/definitions/singleValueReturning" + }, + "direction": { + "type": "string", + "enum": [ + "asc", + "desc" + ], + "default": "asc" + } + } + }, "equality": { "oneOf": [ { @@ -3249,7 +3268,14 @@ "orderBy": { "type": "array", "items": { - "$ref": "#/definitions/orderByItem" + "oneOf": [ + { + "$ref": "#/definitions/orderByItem" + }, + { + "$ref": "#/definitions/expressionOrderByItem" + } + ] } }, "pagination": { @@ -3272,10 +3298,31 @@ "items": { "$ref": "#/definitions/singleValueReturning" } + }, + "orderBy": { + "items": { + "$ref": "#/definitions/expressionOrderByItem" + } } } } }, + "anyOf": [ + { + "required": [ + "groupBy" + ] + }, + { + "properties": { + "orderBy": { + "items": { + "$ref": "#/definitions/orderByItem" + } + } + } + } + ], "title": "PureQL specification", "type": "object" } diff --git a/README.md b/README.md index 5b5d99c..51cc2c8 100644 --- a/README.md +++ b/README.md @@ -12,7 +12,7 @@ PureQL is a JSON-based declarative query language for relational data. Queries a | `joins` | no | Array of join clauses | | `groupBy` | no | Fields to group rows by (at least one); group keys are output automatically | | `having` | no | Boolean filter applied after grouping; requires `groupBy` | -| `orderBy` | no | Fields to order results by | +| `orderBy` | no | Fields to order results by; single-value expressions when `groupBy` is present | | `pagination` | no | `skip` and `take` for paging | | `distinct` | no | When `true`, deduplicate result rows (default: `false`) | @@ -164,15 +164,25 @@ Each join specifies its type (`inner`, `left`, `right`, `full`), the entity to j ### `groupBy` / `orderBy` -`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` accepts an array of `orderByItem` objects, each pairing a `field` with an optional `direction` (`"asc"` | `"desc"`, default `"asc"`). +`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` items have an optional `direction` (`"asc"` | `"desc"`, default `"asc"`). The sort key is written differently depending on `groupBy`: + +**Without `groupBy`**: each item pairs a `field` with a direction (`orderByItem`). + +```json +"orderBy": [ + { "field": { "entity": "users", "field": "name", "type": { "name": "string" } }, "direction": "asc" }, + { "field": { "entity": "orders", "field": "total", "type": { "name": "number" } }, "direction": "desc" } +] +``` + +**With `groupBy`**: each item holds a single-value `expression` (`expressionOrderByItem`), usually an aggregate. Fields are rejected, because a group has no single value for a field. To sort by an aggregate that is also selected, repeat the expression rather than referencing its alias, so that the schema can validate it. To sort by a group key, wrap it in an aggregate such as `min_string`; within a group every key value is the same, so the aggregate returns the key itself. ```json "groupBy": [ { "entity": "orders", "field": "user_id", "type": { "name": "uuid" } } ], "orderBy": [ - { "field": { "entity": "users", "field": "name", "type": { "name": "string" } }, "direction": "asc" }, - { "field": { "entity": "orders", "field": "total", "type": { "name": "number" } }, "direction": "desc" } + { "expression": { "operator": "sum", "arg": { "entity": "orders", "field": "total", "type": { "name": "number" } } }, "direction": "desc" } ] ``` @@ -504,7 +514,7 @@ The [`samples/`](samples/) directory contains query examples ordered by complexi | [`09_arithmetic.json`](samples/09_arithmetic.json) | `add`, `multiply`, `divide` on aggregate results | | [`10_parameters.json`](samples/10_parameters.json) | Named scalar parameters in per-row predicates | | [`11_distinct.json`](samples/11_distinct.json) | `distinct: true` to deduplicate results | -| [`12_complex_query.json`](samples/12_complex_query.json) | Full query: joins, per-row `where`, groupBy, single-value `having`, arithmetic, parameters, orderBy with direction, pagination | +| [`12_complex_query.json`](samples/12_complex_query.json) | Full query: joins, per-row `where`, groupBy, single-value `having`, arithmetic, parameters, grouped `orderBy` by aggregate expressions, pagination | | [`13_range_filter.json`](samples/13_range_filter.json) | `eachGreaterThan` and `eachLessThan` combined with `eachAnd` | | [`14_each_field_to_field.json`](samples/14_each_field_to_field.json) | Per-row range comparison between two fields (no scalar threshold) | | [`15_each_not_equal.json`](samples/15_each_not_equal.json) | `eachNot` wrapping `eachEqual` — the idiom for "field ≠ literal" | diff --git a/samples/12_complex_query.json b/samples/12_complex_query.json index ee0afc0..b03b865 100644 --- a/samples/12_complex_query.json +++ b/samples/12_complex_query.json @@ -110,7 +110,17 @@ }, "orderBy": [ { - "field": { "entity": "users", "field": "name", "type": { "name": "string" } }, + "expression": { + "operator": "sum", + "arg": { "entity": "o", "field": "total_amount", "type": { "name": "number" } } + }, + "direction": "desc" + }, + { + "expression": { + "operator": "min_string", + "arg": { "entity": "users", "field": "name", "type": { "name": "string" } } + }, "direction": "asc" } ], From aca7d6d1be36213bd5cc3e1c67804a9382a12285 Mon Sep 17 00:00:00 2001 From: Dmitry Kurochkin Date: Thu, 24 Sep 2026 14:17:55 +0000 Subject: [PATCH 3/4] feat(schema)!: unify orderBy items around context-dependent expressions Replace the { field } / { expression } split with a single orderByItem { expression, direction }. The allowed expression kind follows the query context, the same split as where/having and grouped select: array-returning without groupBy (fields and each* computations), single-value with groupBy (aggregates and arithmetic over them). Drop expressionOrderByItem. This makes sorting by computed values possible in both contexts. Add sample 23 (order items sorted by eachMultiply(unit_price, quantity)) and update README, CLAUDE.md and CHANGELOG. BREAKING CHANGE: orderBy items use { "expression": X } in place of { "field": X }. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 3 ++- CLAUDE.md | 10 +++++--- PureQL-Specification.json | 42 ++++++++++++++----------------- README.md | 33 +++++++++++++++++++----- samples/23_order_by_computed.json | 33 ++++++++++++++++++++++++ 5 files changed, 87 insertions(+), 34 deletions(-) create mode 100644 samples/23_order_by_computed.json diff --git a/CHANGELOG.md b/CHANGELOG.md index 709a0c7..7b056c0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,6 +11,7 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr ### 🔧 Improvements +- **Sort by computed values**: `orderBy` can sort by any per-row calculation, e.g. line total (`unit_price × quantity`), and grouped queries can sort by any aggregate or arithmetic over aggregates, e.g. average order value. - **Defined results for fields mixed with aggregates in `select`**: A query without `groupBy` can list fields next to aggregates. The aggregate's value is repeated on every row, e.g. an order total next to the grand total of all orders. Aggregates are always calculated over every matching row, before pagination. (#52) ### ⚠️ Breaking Changes @@ -18,7 +19,7 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr - **`having` now requires `groupBy`**: Queries that filter with `having` must also group their rows with `groupBy`. Previously the schema accepted `having` on its own, even though it has no meaning without groups. (#40) - **`groupBy` can no longer be empty**: `groupBy` must list at least one field. To skip grouping, leave the clause out. - **Grouped queries select only single values**: When a query has `groupBy`, `select` accepts only aggregates, scalars, parameters and arithmetic over them. The `groupBy` fields now appear in the result automatically as the first columns, so remove them from `select`. Fields that aren't grouped and per-row `each*` columns are rejected, because a group has no single value for them. (#52) -- **Grouped queries sort by expressions**: When a query has `groupBy`, each `orderBy` item holds an `expression` (usually an aggregate such as `sum(...)`) in place of a `field`. To sort groups by a key, wrap it in an aggregate, e.g. `min_string(users.name)`. Queries without `groupBy` keep sorting by `field`. +- **`orderBy` items hold an expression in place of a field**: Each item is now `{ "expression": ..., "direction": ... }`. Replace `{ "field": X }` with `{ "expression": X }`. Without `groupBy`, the expression is per-row: a field, or a computed value such as `eachMultiply(unit_price, quantity)`. With `groupBy`, it's a single value per group, such as `sum(total)`. To sort groups by a key, wrap it in an aggregate, e.g. `min_string(users.name)`. --- diff --git a/CLAUDE.md b/CLAUDE.md index a1dccea..c1d3073 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -50,7 +50,7 @@ Every operator in the schema belongs to one of two parallel families: | `having` | single-value boolean **only** — operands must reduce to one value per group | | `select` | any value-returning expression, including per-row computed columns — **single-value only when `groupBy` is present** | | `groupBy` | field references only | -| `orderBy` | `{ field, direction }` without `groupBy`; `{ expression: , direction }` with `groupBy` | +| `orderBy` | `{ expression, direction }` — per-row expression without `groupBy`, single-value expression with `groupBy` | ### Fields are `arrayReturning` @@ -123,10 +123,12 @@ Only the root `from` expression supports an `alias`. Joined entities are always Not select expressions — just plain `{ entity, field, type }` field references. No aliases, no operators. -### `orderBy` shape depends on `groupBy` +### `orderBy` keys follow the query's context -- Without `groupBy`: items are `orderByItem` — `{ field, direction }`. Expression items are rejected (root `anyOf`). -- With `groupBy`: items are `expressionOrderByItem` — `{ expression, direction }` where `expression` is `singleValueReturning` (typically an aggregate). Field items are rejected (root `dependentSchemas`). +Every item is `orderByItem` — `{ expression, direction }`. There is no `{ field }` form: a field is just an array-returning `expression`. + +- Without `groupBy`: `expression` must be `*ArrayReturning` (field, `each*` computation). Single-value keys are constants and are rejected (root `anyOf`). +- With `groupBy`: `expression` must be `singleValueReturning` (aggregate, arithmetic over aggregates). Fields and `each*` are rejected (root `dependentSchemas`). - Never reference a `select` alias from `orderBy` — repeat the expression. Aliases are name references the schema cannot check. - To sort groups by a key, wrap it in an aggregate (e.g. `min_string(users.name)`): every value in a group equals the key. diff --git a/PureQL-Specification.json b/PureQL-Specification.json index d50aff9..f8c7c58 100644 --- a/PureQL-Specification.json +++ b/PureQL-Specification.json @@ -2170,25 +2170,20 @@ } }, "orderByItem": { - "type": "object", - "required": ["field"], - "properties": { - "field": { "$ref": "#/definitions/field" }, - "direction": { - "type": "string", - "enum": ["asc", "desc"], - "default": "asc" - } - } - }, - "expressionOrderByItem": { "type": "object", "required": [ "expression" ], "properties": { "expression": { - "$ref": "#/definitions/singleValueReturning" + "oneOf": [ + { + "$ref": "#/definitions/singleValueReturning" + }, + { + "$ref": "#/definitions/arrayReturning" + } + ] }, "direction": { "type": "string", @@ -3268,14 +3263,7 @@ "orderBy": { "type": "array", "items": { - "oneOf": [ - { - "$ref": "#/definitions/orderByItem" - }, - { - "$ref": "#/definitions/expressionOrderByItem" - } - ] + "$ref": "#/definitions/orderByItem" } }, "pagination": { @@ -3301,7 +3289,11 @@ }, "orderBy": { "items": { - "$ref": "#/definitions/expressionOrderByItem" + "properties": { + "expression": { + "$ref": "#/definitions/singleValueReturning" + } + } } } } @@ -3317,7 +3309,11 @@ "properties": { "orderBy": { "items": { - "$ref": "#/definitions/orderByItem" + "properties": { + "expression": { + "$ref": "#/definitions/arrayReturning" + } + } } } } diff --git a/README.md b/README.md index 51cc2c8..23dfe9b 100644 --- a/README.md +++ b/README.md @@ -12,7 +12,7 @@ PureQL is a JSON-based declarative query language for relational data. Queries a | `joins` | no | Array of join clauses | | `groupBy` | no | Fields to group rows by (at least one); group keys are output automatically | | `having` | no | Boolean filter applied after grouping; requires `groupBy` | -| `orderBy` | no | Fields to order results by; single-value expressions when `groupBy` is present | +| `orderBy` | no | Expressions to order results by: per-row without `groupBy`, single-value with it | | `pagination` | no | `skip` and `take` for paging | | `distinct` | no | When `true`, deduplicate result rows (default: `false`) | @@ -164,18 +164,38 @@ Each join specifies its type (`inner`, `left`, `right`, `full`), the entity to j ### `groupBy` / `orderBy` -`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` items have an optional `direction` (`"asc"` | `"desc"`, default `"asc"`). The sort key is written differently depending on `groupBy`: +`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` is an array of `orderByItem` objects: `{ "expression": , "direction": "asc" | "desc" }` (`direction` defaults to `"asc"`). Items are applied in order, each one breaking ties left by the previous one. -**Without `groupBy`**: each item pairs a `field` with a direction (`orderByItem`). +The sort key has to produce one value per **result row**, so which kind of expression is allowed depends on whether the query groups: + +| Query | A result row is | `expression` must be | Examples | +|---|---|---|---| +| without `groupBy` | a row | array-returning | field, `eachMultiply(unit_price, quantity)` | +| with `groupBy` | a group | single-value | `sum(total)`, `divide(sum(total), count(id))` | + +This is the same split as `where` / `having` and as `select` with and without `groupBy`. The schema rejects the other kind: a single-value key without `groupBy` is a constant and would not sort anything, and a field with `groupBy` has no single value per group. + +To sort by a computed `select` column, repeat the expression rather than referencing its alias. The schema cannot check that an alias exists, but it can validate the expression. To sort groups by a group key, wrap it in an aggregate such as `min_string`; within a group every key value is the same, so the aggregate returns the key itself. + +Without `groupBy` (see [`23_order_by_computed.json`](samples/23_order_by_computed.json)): ```json "orderBy": [ - { "field": { "entity": "users", "field": "name", "type": { "name": "string" } }, "direction": "asc" }, - { "field": { "entity": "orders", "field": "total", "type": { "name": "number" } }, "direction": "desc" } + { + "expression": { + "operator": "eachMultiply", + "values": [ + { "entity": "order_items", "field": "unit_price", "type": { "name": "number" } }, + { "entity": "order_items", "field": "quantity", "type": { "name": "number" } } + ] + }, + "direction": "desc" + }, + { "expression": { "entity": "order_items", "field": "id", "type": { "name": "uuid" } } } ] ``` -**With `groupBy`**: each item holds a single-value `expression` (`expressionOrderByItem`), usually an aggregate. Fields are rejected, because a group has no single value for a field. To sort by an aggregate that is also selected, repeat the expression rather than referencing its alias, so that the schema can validate it. To sort by a group key, wrap it in an aggregate such as `min_string`; within a group every key value is the same, so the aggregate returns the key itself. +With `groupBy`: ```json "groupBy": [ @@ -525,3 +545,4 @@ The [`samples/`](samples/) directory contains query examples ordered by complexi | [`20_each_datetime_diff_where.json`](samples/20_each_datetime_diff_where.json) | `eachDatetimeDiffSeconds` inside `eachGreaterThan` to filter orders by ship-time | | [`21_each_time_math.json`](samples/21_each_time_math.json) | `eachTimeAddSeconds` (time + offset) and `eachTimeDiffSeconds` (shift duration) | | [`22_select_broadcast.json`](samples/22_select_broadcast.json) | Fields next to an aggregate in `select` — `sum` broadcast to every row, plus `eachDivide(total, sum(total))` as a share-of-total column | +| [`23_order_by_computed.json`](samples/23_order_by_computed.json) | `orderBy` on a computed per-row expression (`eachMultiply(unit_price, quantity)` desc), then `id` as a tie-breaker | diff --git a/samples/23_order_by_computed.json b/samples/23_order_by_computed.json new file mode 100644 index 0000000..e5f5a5e --- /dev/null +++ b/samples/23_order_by_computed.json @@ -0,0 +1,33 @@ +{ + "from": { "entity": "order_items" }, + "select": [ + { "entity": "order_items", "field": "id", "type": { "name": "uuid" } }, + { "entity": "order_items", "field": "unit_price", "type": { "name": "number" } }, + { "entity": "order_items", "field": "quantity", "type": { "name": "number" } }, + { + "operator": "eachMultiply", + "values": [ + { "entity": "order_items", "field": "unit_price", "type": { "name": "number" } }, + { "entity": "order_items", "field": "quantity", "type": { "name": "number" } } + ], + "alias": "subtotal" + } + ], + "orderBy": [ + { + "expression": { + "operator": "eachMultiply", + "values": [ + { "entity": "order_items", "field": "unit_price", "type": { "name": "number" } }, + { "entity": "order_items", "field": "quantity", "type": { "name": "number" } } + ] + }, + "direction": "desc" + }, + { + "expression": { "entity": "order_items", "field": "id", "type": { "name": "uuid" } }, + "direction": "asc" + } + ], + "pagination": { "skip": 0, "take": 20 } +} From 723af7573a2a31d4653484aa86f6369185cad0e3 Mon Sep 17 00:00:00 2001 From: Dmitry Kurochkin Date: Thu, 24 Sep 2026 14:31:19 +0000 Subject: [PATCH 4/4] refactor(schema): allow single-value orderBy keys without groupBy Drop the root anyOf that rejected single-value orderBy expressions in ungrouped queries. Such keys are constants and leave the order unchanged, like OrderBy(r => 1) in LINQ, so rejecting them added schema complexity without catching an ambiguous query. Grouped queries still require single-value keys via dependentSchemas. Co-Authored-By: Claude Opus 5.5 --- CLAUDE.md | 2 +- PureQL-Specification.json | 20 -------------------- README.md | 10 +++++----- 3 files changed, 6 insertions(+), 26 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c1d3073..b826688 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -127,7 +127,7 @@ Not select expressions — just plain `{ entity, field, type }` field references Every item is `orderByItem` — `{ expression, direction }`. There is no `{ field }` form: a field is just an array-returning `expression`. -- Without `groupBy`: `expression` must be `*ArrayReturning` (field, `each*` computation). Single-value keys are constants and are rejected (root `anyOf`). +- Without `groupBy`: `expression` is normally `*ArrayReturning` (field, `each*` computation). A single-value key validates but is a constant, so it leaves the order unchanged. - With `groupBy`: `expression` must be `singleValueReturning` (aggregate, arithmetic over aggregates). Fields and `each*` are rejected (root `dependentSchemas`). - Never reference a `select` alias from `orderBy` — repeat the expression. Aliases are name references the schema cannot check. - To sort groups by a key, wrap it in an aggregate (e.g. `min_string(users.name)`): every value in a group equals the key. diff --git a/PureQL-Specification.json b/PureQL-Specification.json index f8c7c58..a212fe9 100644 --- a/PureQL-Specification.json +++ b/PureQL-Specification.json @@ -3299,26 +3299,6 @@ } } }, - "anyOf": [ - { - "required": [ - "groupBy" - ] - }, - { - "properties": { - "orderBy": { - "items": { - "properties": { - "expression": { - "$ref": "#/definitions/arrayReturning" - } - } - } - } - } - } - ], "title": "PureQL specification", "type": "object" } diff --git a/README.md b/README.md index 23dfe9b..177f8d7 100644 --- a/README.md +++ b/README.md @@ -166,14 +166,14 @@ Each join specifies its type (`inner`, `left`, `right`, `full`), the entity to j `groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` is an array of `orderByItem` objects: `{ "expression": , "direction": "asc" | "desc" }` (`direction` defaults to `"asc"`). Items are applied in order, each one breaking ties left by the previous one. -The sort key has to produce one value per **result row**, so which kind of expression is allowed depends on whether the query groups: +The sort key produces one value per **result row**, so the kind of expression to use depends on whether the query groups: -| Query | A result row is | `expression` must be | Examples | +| Query | A result row is | `expression` | Examples | |---|---|---|---| -| without `groupBy` | a row | array-returning | field, `eachMultiply(unit_price, quantity)` | -| with `groupBy` | a group | single-value | `sum(total)`, `divide(sum(total), count(id))` | +| without `groupBy` | a row | array-returning (single-value is allowed but is a constant, so it does not change the order) | field, `eachMultiply(unit_price, quantity)` | +| with `groupBy` | a group | single-value **only** | `sum(total)`, `divide(sum(total), count(id))` | -This is the same split as `where` / `having` and as `select` with and without `groupBy`. The schema rejects the other kind: a single-value key without `groupBy` is a constant and would not sort anything, and a field with `groupBy` has no single value per group. +This is the same split as `where` / `having` and as `select` with and without `groupBy`. With `groupBy`, the schema rejects fields and `each*` keys, because a group has no single value for them. Without `groupBy`, a single-value key (e.g. `sum(total)`) is valid but is the same for every row, so it leaves the order unchanged, just as `OrderBy(r => 1)` does in LINQ. To sort by a computed `select` column, repeat the expression rather than referencing its alias. The schema cannot check that an alias exists, but it can validate the expression. To sort groups by a group key, wrap it in an aggregate such as `min_string`; within a group every key value is the same, so the aggregate returns the key itself.