Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,11 +9,17 @@ Versioning follows Semantic Versioning with preview suffix `major.minor.patch-pr

## [Unreleased]

### 🔧 Improvements

- **Sort by computed values**: `orderBy` can sort by any per-row calculation, e.g. line total (`unit_price × quantity`), and grouped queries can sort by any aggregate or arithmetic over aggregates, e.g. average order value.
- **Defined results for fields mixed with aggregates in `select`**: A query without `groupBy` can list fields next to aggregates. The aggregate's value is repeated on every row, e.g. an order total next to the grand total of all orders. Aggregates are always calculated over every matching row, before pagination. (#52)

### ⚠️ Breaking Changes

- **`having` now requires `groupBy`**: Queries that filter with `having` must also group their rows with `groupBy`. Previously the schema accepted `having` on its own, even though it has no meaning without groups. (#40)
- **`groupBy` can no longer be empty**: `groupBy` must list at least one field. To skip grouping, leave the clause out.
- **Grouped queries select only single values**: When a query has `groupBy`, `select` accepts only aggregates, scalars, parameters and arithmetic over them. The `groupBy` fields now appear in the result automatically as the first columns, so remove them from `select`. Fields that aren't grouped and per-row `each*` columns are rejected, because a group has no single value for them. (#52)
- **`orderBy` items hold an expression in place of a field**: Each item is now `{ "expression": ..., "direction": ... }`. Replace `{ "field": X }` with `{ "expression": X }`. Without `groupBy`, the expression is per-row: a field, or a computed value such as `eachMultiply(unit_price, quantity)`. With `groupBy`, it's a single value per group, such as `sum(total)`. To sort groups by a key, wrap it in an aggregate, e.g. `min_string(users.name)`.

---

Expand Down
16 changes: 14 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,8 @@ Every operator in the schema belongs to one of two parallel families:
| `join.on` | single-value boolean **or** per-row boolean (per-row equi-join is the common case) |
| `having` | single-value boolean **only** — operands must reduce to one value per group |
| `select` | any value-returning expression, including per-row computed columns — **single-value only when `groupBy` is present** |
| `groupBy` / `orderBy` | field references only |
| `groupBy` | field references only |
| `orderBy` | `{ expression, direction }` — per-row expression without `groupBy`, single-value expression with `groupBy` |

### Fields are `arrayReturning`

Expand Down Expand Up @@ -112,14 +113,25 @@ So `eachAdd([fieldA, scalar, fieldB])` over 3 rows with `fieldA = [10, 20, 30]`,

This is why `eachX` operators do not collapse to their single-value siblings when handed only scalar operands: the *return type* is still `*ArrayReturning`, aligned with the row set. `eachAdd(2, 3)` in a `select` over a 4-row entity yields the vector `[5, 5, 5, 5]`, not the scalar `5`. Use single-value `add` for purely scalar work; `eachAdd` exists specifically because at least one operand is row-aligned.

The same model applies to `select` itself when it mixes kinds (without `groupBy`): if any item is `*ArrayReturning`, the result has `N` rows and every single-value item (e.g. `sum(total)`) is broadcast to each row; if all items are single-value, the result is one row. Aggregates see all `N` rows — they are computed before `distinct` and `pagination`. So `[field, sum(field)]` without `groupBy` is valid and well-defined, not an error.

### `joinItem` has no alias

Only the root `from` expression supports an `alias`. Joined entities are always referenced by their `entity` name string.

### `groupBy` and `orderBy` take `field` objects
### `groupBy` takes `field` objects

Not select expressions — just plain `{ entity, field, type }` field references. No aliases, no operators.

### `orderBy` keys follow the query's context

Every item is `orderByItem` — `{ expression, direction }`. There is no `{ field }` form: a field is just an array-returning `expression`.

- Without `groupBy`: `expression` is normally `*ArrayReturning` (field, `each*` computation). A single-value key validates but is a constant, so it leaves the order unchanged.
- With `groupBy`: `expression` must be `singleValueReturning` (aggregate, arithmetic over aggregates). Fields and `each*` are rejected (root `dependentSchemas`).
- Never reference a `select` alias from `orderBy` — repeat the expression. Aliases are name references the schema cannot check.
- To sort groups by a key, wrap it in an aggregate (e.g. `min_string(users.name)`): every value in a group equals the key.

### Grouped `select` is single-value only

When `groupBy` is present, every `select` item must be `singleValueReturning` (enforced via root `dependentSchemas`). Group keys are emitted automatically as the leading result columns, so **never repeat `groupBy` fields in `select`** — the schema rejects them, along with any non-grouped field or `each*` column.
Expand Down
29 changes: 26 additions & 3 deletions PureQL-Specification.json
Original file line number Diff line number Diff line change
Expand Up @@ -2171,12 +2171,26 @@
},
"orderByItem": {
"type": "object",
"required": ["field"],
"required": [
"expression"
],
"properties": {
"field": { "$ref": "#/definitions/field" },
"expression": {
"oneOf": [
{
"$ref": "#/definitions/singleValueReturning"
},
{
"$ref": "#/definitions/arrayReturning"
}
]
},
"direction": {
"type": "string",
"enum": ["asc", "desc"],
"enum": [
"asc",
"desc"
],
"default": "asc"
}
}
Expand Down Expand Up @@ -3272,6 +3286,15 @@
"items": {
"$ref": "#/definitions/singleValueReturning"
}
},
"orderBy": {
"items": {
"properties": {
"expression": {
"$ref": "#/definitions/singleValueReturning"
}
}
}
}
}
}
Expand Down
53 changes: 46 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ PureQL is a JSON-based declarative query language for relational data. Queries a
| `joins` | no | Array of join clauses |
| `groupBy` | no | Fields to group rows by (at least one); group keys are output automatically |
| `having` | no | Boolean filter applied after grouping; requires `groupBy` |
| `orderBy` | no | Fields to order results by |
| `orderBy` | no | Expressions to order results by: per-row without `groupBy`, single-value with it |
| `pagination` | no | `skip` and `take` for paging |
| `distinct` | no | When `true`, deduplicate result rows (default: `false`) |

Expand Down Expand Up @@ -91,7 +91,7 @@ Parameters are named placeholders resolved at execution time, analogous to prepa

Each item in `select` is a value-returning expression (field, scalar, aggregate, arithmetic, boolean expression) with an optional `alias`.

When `groupBy` is present, `select` accepts **single-value expressions only** (aggregates, scalars, parameters, arithmetic over them). Fields and per-row `each*` columns are rejected by the schema, because a group has no single value for them. The `groupBy` fields are output automatically as the leading result columns, in `groupBy` order and named after the field, followed by the `select` entries:
**With `groupBy`**, `select` accepts **single-value expressions only** (aggregates, scalars, parameters, arithmetic over them). Fields and per-row `each*` columns are rejected by the schema, because a group has no single value for them. The `groupBy` fields are output automatically as the leading result columns, in `groupBy` order and named after the field, followed by the `select` entries:

```json
"select": [
Expand All @@ -104,7 +104,14 @@ When `groupBy` is present, `select` accepts **single-value expressions only** (a

Result columns: `user_id`, `order_count`.

Ungrouped example:
**Without `groupBy`**, `select` may mix single-value and array-returning items. The result shape is:

- **At least one array-returning item** (field or `each*` column): the result has `N` rows, where `N` is the row count after `from` + `joins` + `where`. Every single-value item (aggregate, scalar, parameter, arithmetic) is **broadcast**: the same value is repeated in every row. If `N = 0`, the result is empty.
- **Only single-value items**: the result has exactly one row.

Aggregates are computed over all `N` rows before `distinct` and `pagination` apply, so `take: 10` pages the rows but does not change `sum(...)`. See [`22_select_broadcast.json`](samples/22_select_broadcast.json).

Example: `name` and `email` are fields, so `order_count` is repeated on every row:

```json
"select": [
Expand Down Expand Up @@ -157,15 +164,45 @@ Each join specifies its type (`inner`, `left`, `right`, `full`), the entity to j

### `groupBy` / `orderBy`

`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` accepts an array of `orderByItem` objects, each pairing a `field` with an optional `direction` (`"asc"` | `"desc"`, default `"asc"`).
`groupBy` accepts an array of field references. Group keys are added to the result automatically, so they are not repeated in `select`. `orderBy` is an array of `orderByItem` objects: `{ "expression": <expr>, "direction": "asc" | "desc" }` (`direction` defaults to `"asc"`). Items are applied in order, each one breaking ties left by the previous one.

The sort key produces one value per **result row**, so the kind of expression to use depends on whether the query groups:

| Query | A result row is | `expression` | Examples |
|---|---|---|---|
| without `groupBy` | a row | array-returning (single-value is allowed but is a constant, so it does not change the order) | field, `eachMultiply(unit_price, quantity)` |
| with `groupBy` | a group | single-value **only** | `sum(total)`, `divide(sum(total), count(id))` |

This is the same split as `where` / `having` and as `select` with and without `groupBy`. With `groupBy`, the schema rejects fields and `each*` keys, because a group has no single value for them. Without `groupBy`, a single-value key (e.g. `sum(total)`) is valid but is the same for every row, so it leaves the order unchanged, just as `OrderBy(r => 1)` does in LINQ.

To sort by a computed `select` column, repeat the expression rather than referencing its alias. The schema cannot check that an alias exists, but it can validate the expression. To sort groups by a group key, wrap it in an aggregate such as `min_string`; within a group every key value is the same, so the aggregate returns the key itself.

Without `groupBy` (see [`23_order_by_computed.json`](samples/23_order_by_computed.json)):

```json
"orderBy": [
{
"expression": {
"operator": "eachMultiply",
"values": [
{ "entity": "order_items", "field": "unit_price", "type": { "name": "number" } },
{ "entity": "order_items", "field": "quantity", "type": { "name": "number" } }
]
},
"direction": "desc"
},
{ "expression": { "entity": "order_items", "field": "id", "type": { "name": "uuid" } } }
]
```

With `groupBy`:

```json
"groupBy": [
{ "entity": "orders", "field": "user_id", "type": { "name": "uuid" } }
],
"orderBy": [
{ "field": { "entity": "users", "field": "name", "type": { "name": "string" } }, "direction": "asc" },
{ "field": { "entity": "orders", "field": "total", "type": { "name": "number" } }, "direction": "desc" }
{ "expression": { "operator": "sum", "arg": { "entity": "orders", "field": "total", "type": { "name": "number" } } }, "direction": "desc" }
]
```

Expand Down Expand Up @@ -497,7 +534,7 @@ The [`samples/`](samples/) directory contains query examples ordered by complexi
| [`09_arithmetic.json`](samples/09_arithmetic.json) | `add`, `multiply`, `divide` on aggregate results |
| [`10_parameters.json`](samples/10_parameters.json) | Named scalar parameters in per-row predicates |
| [`11_distinct.json`](samples/11_distinct.json) | `distinct: true` to deduplicate results |
| [`12_complex_query.json`](samples/12_complex_query.json) | Full query: joins, per-row `where`, groupBy, single-value `having`, arithmetic, parameters, orderBy with direction, pagination |
| [`12_complex_query.json`](samples/12_complex_query.json) | Full query: joins, per-row `where`, groupBy, single-value `having`, arithmetic, parameters, grouped `orderBy` by aggregate expressions, pagination |
| [`13_range_filter.json`](samples/13_range_filter.json) | `eachGreaterThan` and `eachLessThan` combined with `eachAnd` |
| [`14_each_field_to_field.json`](samples/14_each_field_to_field.json) | Per-row range comparison between two fields (no scalar threshold) |
| [`15_each_not_equal.json`](samples/15_each_not_equal.json) | `eachNot` wrapping `eachEqual` — the idiom for "field ≠ literal" |
Expand All @@ -507,3 +544,5 @@ The [`samples/`](samples/) directory contains query examples ordered by complexi
| [`19_each_date_add_days.json`](samples/19_each_date_add_days.json) | `eachDateAddDays` to derive a `delivery_eta` column from `order_date + 30 days` |
| [`20_each_datetime_diff_where.json`](samples/20_each_datetime_diff_where.json) | `eachDatetimeDiffSeconds` inside `eachGreaterThan` to filter orders by ship-time |
| [`21_each_time_math.json`](samples/21_each_time_math.json) | `eachTimeAddSeconds` (time + offset) and `eachTimeDiffSeconds` (shift duration) |
| [`22_select_broadcast.json`](samples/22_select_broadcast.json) | Fields next to an aggregate in `select` — `sum` broadcast to every row, plus `eachDivide(total, sum(total))` as a share-of-total column |
| [`23_order_by_computed.json`](samples/23_order_by_computed.json) | `orderBy` on a computed per-row expression (`eachMultiply(unit_price, quantity)` desc), then `id` as a tie-breaker |
12 changes: 11 additions & 1 deletion samples/12_complex_query.json
Original file line number Diff line number Diff line change
Expand Up @@ -110,7 +110,17 @@
},
"orderBy": [
{
"field": { "entity": "users", "field": "name", "type": { "name": "string" } },
"expression": {
"operator": "sum",
"arg": { "entity": "o", "field": "total_amount", "type": { "name": "number" } }
},
"direction": "desc"
},
{
"expression": {
"operator": "min_string",
"arg": { "entity": "users", "field": "name", "type": { "name": "string" } }
},
"direction": "asc"
}
],
Expand Down
24 changes: 24 additions & 0 deletions samples/22_select_broadcast.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"from": { "entity": "orders" },
"select": [
{ "entity": "orders", "field": "id", "type": { "name": "uuid" } },
{ "entity": "orders", "field": "total_amount", "type": { "name": "number" } },
{
"operator": "sum",
"arg": { "entity": "orders", "field": "total_amount", "type": { "name": "number" } },
"alias": "grand_total"
},
{
"operator": "eachDivide",
"values": [
{ "entity": "orders", "field": "total_amount", "type": { "name": "number" } },
{
"operator": "sum",
"arg": { "entity": "orders", "field": "total_amount", "type": { "name": "number" } }
}
],
"alias": "share_of_total"
}
],
"pagination": { "skip": 0, "take": 10 }
}
33 changes: 33 additions & 0 deletions samples/23_order_by_computed.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
{
"from": { "entity": "order_items" },
"select": [
{ "entity": "order_items", "field": "id", "type": { "name": "uuid" } },
{ "entity": "order_items", "field": "unit_price", "type": { "name": "number" } },
{ "entity": "order_items", "field": "quantity", "type": { "name": "number" } },
{
"operator": "eachMultiply",
"values": [
{ "entity": "order_items", "field": "unit_price", "type": { "name": "number" } },
{ "entity": "order_items", "field": "quantity", "type": { "name": "number" } }
],
"alias": "subtotal"
}
],
"orderBy": [
{
"expression": {
"operator": "eachMultiply",
"values": [
{ "entity": "order_items", "field": "unit_price", "type": { "name": "number" } },
{ "entity": "order_items", "field": "quantity", "type": { "name": "number" } }
]
},
"direction": "desc"
},
{
"expression": { "entity": "order_items", "field": "id", "type": { "name": "uuid" } },
"direction": "asc"
}
],
"pagination": { "skip": 0, "take": 20 }
}
Loading