diff --git a/README.md b/README.md index e9e465f0..8ee66cf1 100644 --- a/README.md +++ b/README.md @@ -149,7 +149,30 @@ ORDER BY sales DESC LIMIT 100; ``` -**Supported features:** cross-index `JOIN`s, `JOIN UNNEST`, window functions, aggregations, nested fields, geospatial queries, and more. +Subqueries and derived tables are part of DQL too — **since engine `0.24.0`**: + +```sql +-- Subquery in WHERE — IN / NOT IN, EXISTS / NOT EXISTS, a scalar comparison, +-- or a quantified one (= ANY | SOME, <> ALL, > ALL, >= ANY, …) +SELECT name FROM employees +WHERE department_id IN (SELECT id FROM departments WHERE region = 'EU'); + +-- Derived table in FROM (also valid in JOIN, nested to any depth) +SELECT d.category, d.n +FROM (SELECT category, COUNT(*) AS n FROM bi_events GROUP BY category) d +WHERE d.n > 10; + +-- Correlated subquery — the body reads the outer row +SELECT c.name FROM customers c +WHERE NOT EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id); +``` + +- **Uncorrelated `WHERE` subqueries run on Elasticsearch itself**, so they work on every surface — including a plain REPL with no extensions. +- **Derived tables and correlated subqueries run on the relational engine** — since engine `0.24.0` with arrow-extensions `0.3.4` (`softclient4es-arrow-extensions`), the same engine that executes cross-index JOINs: the REPL's default install, the JDBC driver, the ADBC driver, the Arrow Flight SQL server and Federation all carry it. A venue without it refuses the statement with a clear error instead of executing it against the first index named. +- **Non-recursive CTEs run on that same relational engine** — a CTE reference *is* a derived table, so it carries the same venue requirement. A `WITH` clause is accepted at the top of a `SELECT` only; `WITH RECURSIVE` and CTE column lists (`WITH a (x, y) AS …`) are refused by name. +- **Set operators run on that engine too** — `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL`. `UNION ALL` is the exception: Elasticsearch answers it directly with one `_msearch`, at every venue. Branches are matched by column **position**, and the result takes the first branch's names. + +**Supported features:** cross-index `JOIN`s, `JOIN UNNEST`, subqueries, derived tables, CTEs, set operators, window functions, aggregations, nested fields, geospatial queries, and more. 📖 **[DQL Documentation](documentation/sql/dql_statements.md)** @@ -509,6 +532,9 @@ Materialized views with JOINs rely on **Elasticsearch Watcher** to automatically - [x] Arrow Flight SQL server (gRPC, Docker) - [x] ADBC driver (in-process, columnar) - [x] Cross-index JOINs +- [x] Subqueries (`IN` / `EXISTS` / scalar / quantified, correlated or not) and derived tables — `0.24.0` +- [x] Non-recursive CTEs (`WITH name AS (SELECT …)`) — `0.24.0` +- [x] Set operators (`UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL`) — `0.24.0` - [ ] Advanced monitoring dashboard - [ ] Additional SQL functions - [ ] ES|QL bridge diff --git a/documentation/client/adbc_driver.md b/documentation/client/adbc_driver.md index b715649e..03c988de 100644 --- a/documentation/client/adbc_driver.md +++ b/documentation/client/adbc_driver.md @@ -194,7 +194,7 @@ Multi-cluster **federation** — joining across *separate* ES clusters — is ** ## What does NOT work yet -Subqueries (`IN (SELECT …)`, `EXISTS`, scalar, derived tables) and CTEs (`WITH`) are not supported in the current release — they arrive in a later release. Write the JOIN explicitly instead. See the Known Limitations & Roadmap (`../sql/known_limitations.md`) for the full list. +**Since engine `0.24.0`**, subqueries (`IN (SELECT …)` / `NOT IN`, `EXISTS` / `NOT EXISTS`, scalar and quantified comparisons) and derived tables (`FROM (SELECT …)`, `JOIN (SELECT …)`) are supported, correlated or not. **Non-recursive CTEs** (`WITH name AS (SELECT …)`) are supported too, at the top of a `SELECT`; `WITH RECURSIVE` and CTE column lists are refused by name. **Set operators** — `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL` — are supported as well; everything but `UNION ALL` runs on the relational engine this driver ships. See the Known Limitations & Roadmap (`../sql/known_limitations.md#subqueries-and-derived-tables`) for the forms that are still refused. --- diff --git a/documentation/client/arrow_flight_sql.md b/documentation/client/arrow_flight_sql.md index a4d65e24..614637ca 100644 --- a/documentation/client/arrow_flight_sql.md +++ b/documentation/client/arrow_flight_sql.md @@ -197,7 +197,7 @@ The single-cluster sidecar on this page is the free shape. Multi-cluster **feder ## What does NOT work yet -Subqueries (`IN (SELECT …)`, `EXISTS`, scalar, derived tables) and CTEs (`WITH`) are not supported in the current release — they arrive in a later release. Write the JOIN explicitly instead. See the Known Limitations & Roadmap (`../sql/known_limitations.md`) for the full list. +**Since engine `0.24.0`**, subqueries (`IN (SELECT …)` / `NOT IN`, `EXISTS` / `NOT EXISTS`, scalar and quantified comparisons) and derived tables (`FROM (SELECT …)`, `JOIN (SELECT …)`) are supported, correlated or not. **Non-recursive CTEs** (`WITH name AS (SELECT …)`) are supported too, at the top of a `SELECT`; `WITH RECURSIVE` and CTE column lists are refused by name. **Set operators** — `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL` — are supported as well; everything but `UNION ALL` runs on the relational engine this driver ships. See the Known Limitations & Roadmap (`../sql/known_limitations.md#subqueries-and-derived-tables`) for the forms that are still refused. --- @@ -218,7 +218,7 @@ The Arrow Flight SQL sidecar sends one anonymous usage ping per day (no IP, no S ## Known limitations -Subqueries, CTEs (`WITH`), and set operators beyond `UNION ALL` are not in R1 — and some BI tools auto-generate them. See [Known Limitations & Roadmap](../sql/known_limitations.md) for exactly what works today, what's coming in R2a, and the per-tool workaround. +Subqueries, derived tables, non-recursive CTEs — which some BI tools auto-generate — and the `UNION` / `INTERSECT` / `EXCEPT` set operators all work since engine `0.24.0`. See [Known Limitations & Roadmap](../sql/known_limitations.md#subqueries-and-derived-tables) for exactly what works today and what is still refused. --- diff --git a/documentation/client/bi_tools.md b/documentation/client/bi_tools.md index 637aaa71..e01b6623 100644 --- a/documentation/client/bi_tools.md +++ b/documentation/client/bi_tools.md @@ -37,11 +37,14 @@ Tableau's own connector documentation says that when the temp-table capabilities *"Tableau will attempt to generate an alternative query to retrieve the necessary results."* The probe therefore costs one failed round trip per connection and is not itself a problem. What -follows it can be: Tableau's alternative for a source without temporary tables uses **subqueries**, -which this release does not accept, so some interactions fail — with the same kind of clear error -naming the statement, never a hang and never a silently wrong answer. Tableau's own documentation -also warns that the subquery path *"can be poor, particularly with large datasets."* See the -Honest-gap note below for what lands when. +follows it is Tableau's alternative for a source without temporary tables, which uses **subqueries** — +and **since engine `0.24.0`** subqueries and derived tables are accepted. (The quoted, +fully-qualified identifiers Tableau emits alongside them have parsed since `0.23.0`; what `0.24.0` +added is executing the derived-table wrapper.) Tableau's own documentation warns +that the subquery path *"can be poor, particularly with large datasets"*, so it is a performance +characteristic to watch rather than a refusal. Note that a derived table runs on the relational engine — +since engine `0.24.0` with arrow-extensions `0.3.4` — and the JDBC driver ships it, so a Tableau +connection has it. **A Tableau datasource customization file (`.tdc`) cannot suppress the probe.** The capability that would do it, `CAP_SUPPRESS_TEMP_TABLE_CHECKS`, is not among the capabilities Tableau documents for @@ -53,10 +56,16 @@ dead end worth not walking down. ## Honest-gap note -The superpower of this release is a **cross-index JOIN** that Elasticsearch can't do, and it runs best -through explicit `JOIN … ON …` SQL — from any tool where you control the statement that is sent (Superset -SQL Lab, DBeaver, Grafana). Some BI tools compose SQL for you: subqueries and CTEs are not in this release -yet, and neither is the quoted, fully-qualified identifier form Tableau generates. Tableau's Custom SQL is -not a way around that — Tableau wraps a custom query inside a `SELECT … FROM ( … )`, which is a derived -table (Tableau's Custom SQL documentation, checked 2026-09-01). Full BI-tool subquery / CTE support is coming in the next release (Quarter 4 2026). See the -website's Known Limitations page for the full picture. +The superpower of this release is a **cross-index JOIN** that Elasticsearch can't do. It runs from explicit +`JOIN … ON …` SQL and, **since engine `0.24.0`**, from the nested SQL a BI tool composes for you: +**subqueries and derived tables are accepted**, and so is the quoted, fully-qualified identifier form +Tableau generates. +Tableau's Custom SQL wraps your query inside a `SELECT … FROM ( … )` (Tableau's Custom SQL documentation, +checked 2026-09-01) — that wrapper is a derived table, which now runs on the relational engine the JDBC +driver ships. + +Non-recursive **CTEs** (`WITH …`) and the **set operators** (`UNION`, `INTERSECT`, `EXCEPT`, with or +without `ALL`) run on that same engine since `0.24.0` — a CTE reference is a derived table, and a set +operation is its branches executed separately and combined. `UNION ALL` alone still runs on Elasticsearch, +as it always has. See the website's Known Limitations page for the full picture, including the subquery +and set-operator forms that are still refused by name. diff --git a/documentation/client/jdbc.md b/documentation/client/jdbc.md index d9ce5a76..415987dd 100644 --- a/documentation/client/jdbc.md +++ b/documentation/client/jdbc.md @@ -278,7 +278,7 @@ The JDBC driver supports the full SQL Gateway syntax: - **DDL** — CREATE/ALTER/DROP TABLE, pipelines, watchers, enrich policies - **DML** — INSERT, UPDATE, DELETE, COPY INTO -- **DQL** — SELECT with WHERE, GROUP BY, HAVING, ORDER BY, LIMIT, UNION ALL, JOIN UNNEST, window functions +- **DQL** — SELECT with WHERE, GROUP BY, HAVING, ORDER BY, LIMIT, set operators (`UNION [ALL]` / `INTERSECT` / `EXCEPT`), JOIN UNNEST, window functions - **SHOW/DESCRIBE** — Tables, pipelines, watchers, enrich policies --- @@ -324,7 +324,7 @@ Multi-cluster **federation** — joining across *separate* ES clusters — is ** ## What does NOT work yet -Subqueries (`IN (SELECT …)`, `EXISTS`, scalar, derived tables) and CTEs (`WITH`) are not supported in the current release — they arrive in a later release. Write the JOIN explicitly instead. See the Known Limitations & Roadmap (`../sql/known_limitations.md`) for the full list. +**Since engine `0.24.0`**, subqueries (`IN (SELECT …)` / `NOT IN`, `EXISTS` / `NOT EXISTS`, scalar and quantified comparisons) and derived tables (`FROM (SELECT …)`, `JOIN (SELECT …)`) are supported, correlated or not. **Non-recursive CTEs** (`WITH name AS (SELECT …)`) are supported too, at the top of a `SELECT`; `WITH RECURSIVE` and CTE column lists are refused by name. **Set operators** — `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL` — are supported as well; everything but `UNION ALL` runs on the relational engine this driver ships. See the Known Limitations & Roadmap (`../sql/known_limitations.md#subqueries-and-derived-tables`) for the forms that are still refused. --- diff --git a/documentation/sql/dql_statements.md b/documentation/sql/dql_statements.md index d5a6ccb2..b4570671 100644 --- a/documentation/sql/dql_statements.md +++ b/documentation/sql/dql_statements.md @@ -11,7 +11,7 @@ DQL supports: - `SELECT` with expressions, aliases, nested fields, STRUCT and ARRAY - `WHERE`, `GROUP BY`, `HAVING`, `ORDER BY`, `LIMIT`, `OFFSET` -- `UNION ALL` +- set operators: `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL` - cross-index JOINs (`INNER` / `LEFT` / `RIGHT` / `FULL OUTER`) across indices and clusters — see [Cross-Index JOIN](joins.md) - `JOIN UNNEST` on `ARRAY` (the single-index nested form, handled natively inside one index) - aggregations, parent-level aggregations on nested arrays @@ -29,7 +29,9 @@ DQL supports: - [WHERE](#where) - [ORDER BY](#order-by) - [LIMIT / OFFSET](#limit--offset) -- [UNION ALL](#union-all) +- [Set operators](#set-operators) +- [Subqueries and derived tables](#subqueries-and-derived-tables) +- [Common table expressions](#common-table-expressions) - [JOIN UNNEST](#join-unnest) - [Aggregations](#aggregations) - [Parent-Level Aggregations on Nested Arrays](#parent-level-aggregations-on-nested-arrays) @@ -196,9 +198,9 @@ spelling, and may carry a qualifier: ```sql SELECT category FROM `bi_events`; SELECT category FROM "bi_events"; -SELECT category FROM `elastic`.`bi_events` `bi_events`; -SELECT category FROM "elastic"."bi_events" AS e; -SELECT o.id FROM `elastic`.`orders` o JOIN `elastic`.`customers` c ON o.cid = c.id; +SELECT category FROM `prod_eu`.`bi_events` `bi_events`; +SELECT category FROM "prod_eu"."bi_events" AS e; +SELECT o.id FROM `prod_eu`.`orders` o JOIN `prod_eu`.`customers` c ON o.cid = c.id; ``` ### Quoting is what makes a dot a qualifier @@ -210,11 +212,11 @@ dot; the index name is everything after that run.** Nothing else separates the t | Written | Index read | Qualifier | | ------- | ---------- | --------- | | `FROM bi_events` | `bi_events` | — | -| `FROM elastic.bi_events` | `elastic.bi_events` | — (a bare dot is part of the name) | +| `FROM prod_eu.bi_events` | `prod_eu.bi_events` | — (a bare dot is part of the name) | | `FROM logs-2025.03` | `logs-2025.03` | — | -| `FROM "elastic".bi_events` | `bi_events` | `elastic` | -| ``FROM `elastic`.bi_events`` | `bi_events` | `elastic` | -| `FROM "elastic"."bi_events"` | `bi_events` | `elastic` | +| `FROM "prod_eu".bi_events` | `bi_events` | `prod_eu` | +| ``FROM `prod_eu`.bi_events`` | `bi_events` | `prod_eu` | +| `FROM "prod_eu"."bi_events"` | `bi_events` | `prod_eu` | | `FROM "logs-2025.03"` | `logs-2025.03` | — (the dot is *inside* the quotes) | | `FROM "elasticsearch"."prod-cluster"."bi_events"` | `bi_events` | `elasticsearch`, `prod-cluster` | @@ -255,9 +257,9 @@ Qualifiers used to be dropped from the re-rendered SQL. They are not any more statement carries the qualifier the original had, canonicalised to the ANSI double quote: ```sql -SELECT category FROM `elastic`.`bi_events` +SELECT category FROM `prod_eu`.`bi_events` -- renders as -SELECT category FROM "elastic"."bi_events" +SELECT category FROM "prod_eu"."bi_events" ``` Unlike a column name, a table name is quoted as **one lexeme** — `FROM "logs-2025.03"`, never @@ -390,7 +392,7 @@ functions of literals. Rejected with a named reason (`... requires a FROM clause - `EXCEPT(...)`, duplicate output column names, unbound `?` parameters, array literals, negative `LIMIT`/`OFFSET` -Rejected at the grammar level: `WHERE` / `GROUP BY` / `HAVING` / `ORDER BY` / `UNION ALL` +Rejected at the grammar level: `WHERE` / `GROUP BY` / `HAVING` / `ORDER BY` / a set operator after a FROM-less select-list, and `DISTINCT` literals. A constant cast works in every spelling — `CAST('125' AS BIGINT)`, `CONVERT('125', BIGINT)` and `'125'::BIGINT` all parse. Prefer `TRY_CAST('125' AS BIGINT)` when the value may not convert: `::` is always the @@ -606,22 +608,52 @@ ORDER BY age DESC LIMIT 10 OFFSET 20; ``` ---- +### SELECT TOP n — the same bound, spelled the T-SQL way + +`SELECT TOP n` is accepted as a **spelling of `LIMIT n`**, because it is what BI tools emit in their +SQL-92 dialect. It is not a second row bound: the parser folds it into the statement's `LIMIT`, so +the two statements below are the same statement, and the engine renders both as `LIMIT`. + +```sql +SELECT TOP 10 id, name FROM dql_users ORDER BY age DESC; +SELECT TOP (10) id, name FROM dql_users ORDER BY age DESC; -- parenthesised, also T-SQL +SELECT id, name FROM dql_users ORDER BY age DESC LIMIT 10; -- identical to both +``` -## UNION ALL +Four things to know: + +- **`TOP` is not reserved.** `SELECT top FROM t` still selects a column called `top`, and + `SELECT top - 1 AS x FROM t` still computes it. +- **`TOP` and `LIMIT` together are refused by name.** They are two spellings of one bound, and + guessing a precedence would be worse than saying so. +- **`TOP` carries no `OFFSET`.** For paging, use `LIMIT n OFFSET m`. +- **`TOP n PERCENT` and `TOP n WITH TIES` are refused by name** — they are real T-SQL that this + engine does not implement. See [Known limitations](known_limitations.md). + +`SELECT DISTINCT TOP n` is **not** accepted (write `SELECT DISTINCT … LIMIT n`); the order +`SELECT TOP n DISTINCT` happens to parse but is not valid T-SQL, so do not rely on it. + +--- -`UNION ALL` combines the results of multiple `SELECT` queries **without removing duplicates**. +## Set operators -> Bare `UNION` (with row de-duplication) is **not supported** and is rejected at parse time — the -> two are not synonyms. It used to be accepted silently, returning only the first leg's rows. +A set operator combines the rows of two or more `SELECT` branches. Seven spellings are accepted: -All SELECT statements in a UNION ALL must be **strictly compatible**: +| Spelling | Rows | Duplicates | Where it runs | +| --- | --- | --- | --- | +| `UNION ALL` | every row of every branch | kept | **Elasticsearch** (`_msearch`) — every venue | +| `UNION` / `UNION DISTINCT` | rows in either branch | removed | relational engine | +| `INTERSECT` | rows in **both** branches | removed | relational engine | +| `INTERSECT ALL` | rows in both branches | kept (per-branch multiplicity) | relational engine | +| `EXCEPT` | rows in the first branch and not the second | removed | relational engine | +| `EXCEPT ALL` | rows in the first branch and not the second | kept | relational engine | -- **same number of columns** -- **same column names** (after alias resolution) -- **same or implicitly compatible types** +`UNION` and `UNION DISTINCT` are the same operator — SQL's default for a bare `UNION` is +de-duplication, and `DISTINCT` just says so out loud. `UNION ALL` is the one that keeps duplicates, +and it is the only one Elasticsearch can answer by itself. -If these conditions are not met, the Gateway raises a validation error before executing the query. +> **The `EXCEPT` set operator is not the `SELECT * EXCEPT(col, …)` column-exclusion clause.** The +> first removes *rows*, the second removes *columns* from `SELECT *`. Both work; they are unrelated. **Example** @@ -631,21 +663,374 @@ UNION ALL SELECT id, name FROM dql_users WHERE age <= 30; ``` +```sql +-- Customers who ordered in both quarters, de-duplicated +SELECT customer_id FROM orders_q1 +INTERSECT +SELECT customer_id FROM orders_q2; +``` + +### Columns match by position, not by name + +Branches are matched **column by column, left to right** — SQL-92 §7.10 — and the result takes the +**first branch's** column names. The names the other branches use play no part in the matching: + +```sql +SELECT id AS x FROM left_index +UNION ALL +SELECT id AS y FROM right_index; +-- One column, named x, carrying both branches' ids. +``` + +Two consequences worth knowing: + +- Reordering a branch's `SELECT` list changes the result, even though the column names still look + right. `SELECT a, b … UNION ALL SELECT b, a …` interleaves the two. +- A branch written as a bare `SELECT *` declares no column list, so there is nothing to match + positionally. Such a branch is matched by name instead, and the engine cannot check its width. + Name the columns explicitly whenever a branch's shape matters. + +> **Changed in `0.24.0`:** before this release the engine matched branches **by name**, +> so a column the other branch did not name came back `NULL`. Positional matching is both the +> standard's rule and what every other SQL engine does. + +### Compatibility rules + +- **Same number of columns.** Checked at parse time, between the branches that declare a column list + — a mismatch is rejected before anything runs, naming both branches and their projections. +- **Compatible types, position by position,** against the first declaring branch. A bare column has + no type until the index mapping is read, so the parser lets it through; the check runs again once + the schema is attached, and a `keyword` against a `long` is refused **before any request is sent**. +- **Column names are never compared.** They are output labels, taken from the first branch. + +### Precedence and grouping + +`INTERSECT` binds tighter than `UNION` and `EXCEPT`; otherwise branches associate left to right +(SQL-92 §7.10). So this: + +```sql +SELECT a FROM t1 UNION SELECT a FROM t2 INTERSECT SELECT a FROM t3 +``` + +is `t1 UNION (t2 INTERSECT t3)`. + +**Parenthesising a set operation is rejected** — both as a branch (`(a UNION b) INTERSECT c`) and as +a whole statement. To group differently, use a derived table: + +```sql +SELECT * FROM (SELECT a FROM t1 UNION SELECT a FROM t2) AS g +INTERSECT +SELECT a FROM t3; +``` + +### ORDER BY and LIMIT + +For `UNION ALL` they apply **per branch**, exactly as they always have — each `SELECT` is its own +Elasticsearch request, and the results are concatenated in branch order. + +For every other operator a trailing `ORDER BY` / `LIMIT` after the **last** branch is **rejected**, +because a set has no branch order to inherit and an analyst writing it almost certainly meant the +whole result. Say which you meant: + +```sql +-- this branch only +SELECT a FROM t1 UNION (SELECT a FROM t2 ORDER BY a LIMIT 10); + +-- the whole result +SELECT * FROM (SELECT a FROM t1 UNION SELECT a FROM t2) AS g ORDER BY a LIMIT 10; +``` + +A first or middle branch carrying its own `ORDER BY` / `LIMIT` is unambiguous and is accepted without +parentheses. + ### Execution model -The SQL Gateway executes `UNION ALL` using **Elasticsearch Multi‑Search (`_msearch`)**: +`UNION ALL` is executed by **Elasticsearch Multi-Search (`_msearch`)**: -1. Each SELECT query is translated into an independent ES search request. -2. All requests are sent in a single `_msearch` call. -3. The Gateway concatenates the results **in order**, without deduplication. -4. ORDER BY, LIMIT, OFFSET apply **per SELECT**, not globally (unless wrapped in a subquery, which is not supported). +1. Each branch is translated into an independent ES search request. +2. All requests are sent in one `_msearch` call. +3. Results are concatenated **in branch order**, without de-duplication or sorting. -### Notes +Every other operator needs de-duplication or set arithmetic across branches, which Elasticsearch has +no operation for, so the statement runs on the **relational engine** — engine `0.24.0` with +arrow-extensions `0.3.4`, the same engine that executes cross-index JOINs and derived tables. Each +branch is executed as its own Elasticsearch query and the set operation is applied to the results. A +venue without that engine refuses the statement with a clear error rather than answering from one +branch. See [Known Limitations & Roadmap](known_limitations.md#set-operators). + +A branch may carry anything a `SELECT` can carry — `GROUP BY`, a `JOIN`, a derived table, a CTE, a +correlated subquery. A branch that needs the relational engine on its own account is planned as its +own nested query, so the whole statement routes there. + +### What is not supported + +- **A set operation as a `WHERE` subquery body.** `WHERE a IN (SELECT … UNION SELECT …)` is rejected; + write one subquery per branch. +- **A set operation across catalogs.** Combining branches with catalog-qualified names + (`` `cluster_b`.orders ``) is refused by name: catalogs are resolved by their position in the SQL + text, so a branch could silently run on the wrong cluster. Run each branch as its own statement, or + drop the catalog prefix. +- **`CORRESPONDING` / `CORRESPONDING BY`** — SQL's opt-in for name-based matching. Not implemented; + positional matching is the only mode. + +--- + +## Subqueries and derived tables + +*Since engine `0.24.0`.* A subquery is a `SELECT` nested inside another statement. Two positions are +accepted, and they behave very differently — one runs on Elasticsearch, the other on the relational +engine: + +| Position | Spelling | Where it runs | +| --- | --- | --- | +| `FROM` / `JOIN` — a **derived table** | `FROM (SELECT …) AS d`, `JOIN (SELECT …) AS d ON …` | relational engine (arrow-extensions `0.3.4`) | +| `WHERE`, **uncorrelated** | `IN` / `NOT IN` / `EXISTS` / `NOT EXISTS` / scalar / quantified | **Elasticsearch**, in two phases — every venue, the plain REPL included | +| `WHERE`, **correlated** (the body reads an outer alias) | the same spellings | relational engine | + +That middle row is the one worth remembering: *"subqueries are supported"* is true everywhere for the +uncorrelated `WHERE` forms and only at an engine venue for the other two. See +[Known Limitations & Roadmap](known_limitations.md#which-forms-need-the-relational-engine) for the +per-venue table. + +### Derived tables + +A derived table is a `SELECT` in `FROM` or `JOIN` position. **The alias is mandatory** — SQL-92 calls +it a correlation name — and leaving it off is refused by name rather than guessed at: + +```sql +SELECT d.country, d.total +FROM (SELECT country, SUM(amount) AS total FROM orders GROUP BY country) AS d +WHERE d.total > 1000; +``` + +`AS` is optional (`… ) d` is the same thing), so the SQL the BI tools generate on their own — +Tableau's connection and row-count probes, Superset's series limit — is accepted as written. + +```sql +-- in JOIN position +SELECT o.id, d.total +FROM orders o +JOIN (SELECT customer_id, SUM(amount) AS total FROM orders GROUP BY customer_id) AS d + ON o.customer_id = d.customer_id; +``` + +Bodies **nest to any depth**, and a body may itself carry a `JOIN`, a `GROUP BY`, its own `ORDER BY` +/ `LIMIT`, or a set operation: + +```sql +-- a set operation grouped by a derived table, the spelling that replaces parentheses +SELECT * FROM (SELECT a FROM t1 UNION SELECT a FROM t2) AS g; + +-- the body's LIMIT bounds the body; the outer statement is ordered and limited separately +SELECT d.total FROM (SELECT SUM(amount) AS total FROM orders GROUP BY country) AS d +ORDER BY d.total DESC LIMIT 5; +``` + +A body written as `SELECT *` is **opaque**: it declares no column list, so the engine cannot know its +width until the index mapping is read. That is legal and is exactly what the BI probes emit: + +```sql +SELECT MAX(1) AS TblMax FROM (SELECT * FROM orders) t; +``` + +#### How an unqualified name resolves — a deliberate deviation from SQL-92 + +When a name in the outer query is **not** qualified with an alias, it resolves to the sole plain +index in the `FROM` clause, **not** to the derived table — even though the derived table is also in +scope: + +```sql +SELECT country, SUM(amount) AS total +FROM orders +JOIN (SELECT country AS c2 FROM orders GROUP BY country) AS d ON country = c2 +GROUP BY country; +-- `country` and `amount` resolve against `orders`; `c2` is the only name `d` projects. +``` + +Strict SQL-92 would call `country` ambiguous. Resolving it this way is what makes Superset's +series-limit query — which projects renamed columns from its inner query precisely so the outer names +stay unambiguous — run as written. **Qualify the name** (`orders.country`, `d.c2`) whenever you want +to be explicit; a qualified name always wins. + +#### Column aliases on the derived table itself + +`FROM (SELECT id FROM orders) AS d (x)` — a column list on the correlation name — is **not +supported**. Alias the columns inside the body instead: `(SELECT id AS x FROM orders) AS d`. + +### Subqueries in `WHERE` + +Six spellings, all accepted in a `SELECT`, a `DELETE` and an `UPDATE`: + +```sql +SELECT name FROM employees WHERE department_id IN (SELECT id FROM departments WHERE region = 'EU'); +SELECT name FROM employees WHERE department_id NOT IN (SELECT id FROM departments WHERE region = 'EU'); +SELECT name FROM customers c WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id); +SELECT name FROM customers c WHERE NOT EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id); +SELECT name FROM employees WHERE salary > (SELECT AVG(salary) FROM employees); +SELECT name FROM employees WHERE department_id = ANY (SELECT id FROM departments); +SELECT name FROM employees WHERE department_id <> ALL (SELECT id FROM departments); +``` + +`SOME` is a synonym for `ANY`. The ordering quantifiers (`> ALL`, `>= ANY`, `< ALL`, `<= SOME`, …) are +accepted too and are reduced against the value list rather than sent to Elasticsearch as such. + +```sql +DELETE FROM orders WHERE customer_id IN (SELECT id FROM customers WHERE region = 'EU'); +UPDATE orders SET status = 'eu' WHERE customer_id IN (SELECT id FROM customers WHERE region = 'EU'); +``` + +#### The comparison must be written subquery-on-the-right + +A scalar or quantified subquery is accepted only on the **right-hand side** of the comparison. +`WHERE (SELECT COUNT(*) FROM orders o WHERE o.customer_id = c.id) > 5` is rejected — and the message +it produces (`Unbalanced parentheses`) does not say why. Flip the comparison: + +```sql +-- rejected +-- SELECT c.name FROM customers c WHERE (SELECT COUNT(*) FROM orders o WHERE o.customer_id = c.id) > 5; + +-- accepted, same meaning +SELECT c.name FROM customers c WHERE 5 < (SELECT COUNT(*) FROM orders o WHERE o.customer_id = c.id); +``` + +#### How an uncorrelated `WHERE` subquery executes + +In **two phases**, both on Elasticsearch, which is why these forms work at every venue including the +plain REPL: + +1. The inner query runs first, and its distinct values are collected. +2. Those values are injected into the outer query as a literal `terms` filter, and the outer query + runs as an ordinary search. + +The inner result set is therefore **bounded**: past `index.max_terms_count` (65,536 by default) the +statement fails loudly rather than truncating silently. Narrow the inner query, or raise the index +setting. + +Elasticsearch's own `terms` lookup is not what this uses and could not be: it reads **one document by +id**, and cannot express a query. + +#### `NOT IN` and `NULL` + +Standard SQL three-valued logic applies. If the inner query returns **any** `NULL`, `x NOT IN +(SELECT …)` is UNKNOWN for every row and the statement returns nothing. That is correct SQL and a +frequent surprise — add `WHERE IS NOT NULL` to the inner query when the column is nullable. + +> **One carve-out, and it fails the other way.** The engine detects the `NULL` by re-running the +> inner query with an `IS NULL` filter. When the inner query carries a `GROUP BY`, that probe is a +> grouped query too, and Elasticsearch's `terms` aggregation **drops the missing-value group** — so +> a `NULL` in a `GROUP BY` body is invisible and `NOT IN` returns rows the rule above says it +> should not. Filter the `NULL` out explicitly in that body rather than relying on the detection. + +### Correlated subqueries + +A subquery whose body reads a column from the enclosing query is **correlated**, and it runs on the +relational engine: -- `UNION ALL` does **not** sort or deduplicate results. -- Column names in the final output are taken from the **first SELECT**. -- All subsequent SELECTs must produce columns with the **same names**. -- Type mismatches should result in a validation error before execution. (⚠️ not implemented yet) +```sql +SELECT c.name FROM customers c +WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id); +``` + +One rule the engine enforces by name, and one convention nothing can enforce for you: + +- ⚠️ **The outer reference must be QUALIFIED — and this one is not refused.** A bare name inside the + body is read as the body's own column, which is a perfectly legal statement, so nothing rejects it: + `… WHERE o.customer_id = id` stops being correlated and runs as an ordinary uncorrelated subquery, + returning **different rows with HTTP 200**. It is the only mistake in this family that fails + silently. Always write the outer alias — `… WHERE o.customer_id = c.id`. +- **The outer reference must be UNQUOTED** — `"c"."id"` is refused by name. The engine rewrites an + outer reference onto the extracted leg, and a quoted identifier is not rewritten. +- **The body must be a single Elasticsearch source.** Its own `JOIN`, a comma-separated `FROM`, a + `JOIN UNNEST`, a derived table or a window function inside the body are refused; move the construct + to the outer `FROM` and correlate against it. + +A correlated subquery costs one relational operation against your plan's `maxJoins` allowance; a +derived table costs nothing on its own. See +[Known Limitations & Roadmap](known_limitations.md#licensing). + +### What is not supported + +- **A subquery in the `SELECT` list** — `SELECT (SELECT MAX(amount) FROM orders) AS m …` does not + parse. +- **A subquery in `HAVING`** — any subquery, correlated or not, refused by name. Compute the value + separately or move the condition to `WHERE`. +- **`LATERAL`** — a derived table that reads an alias from the enclosing `FROM`. Refused by name, with + the rewrite in the message. +- **A set-operation body** — `IN (SELECT … UNION ALL SELECT …)`. Write one subquery per branch. +- **A `FROM`-less body** — `IN (SELECT 1)`. Write the literal list. +- **More than one projected column** in an `IN` / quantified / scalar body. +- **A subquery in a `CASE WHEN` condition** — refused by name; filter in `WHERE`, or compute the flag + in a separate query. +- **A subquery in a `JOIN … ON` clause.** ⚠️ Its message names neither subqueries nor a rewrite — it + reads *"ON clause … must use either equality operator or AND predicate"*. Join on a plain equality + and move the subquery to `WHERE`. +- **A `WHERE` subquery inside `CREATE MATERIALIZED VIEW`** — refused by name: an Elasticsearch + transform cannot run the inner query. Resolve it into the view's source, or keep it in the queries + you run against the view. +- **A `FROM`-less `SELECT` used as a set-operation branch** — `SELECT 1 UNION ALL SELECT id FROM t` + is a syntax error. Worth knowing because `SELECT 1` is the connection-handshake idiom; it does not + compose into a set operation. + +--- + +## Common table expressions + +*Since engine `0.24.0`.* A `WITH` clause names one or more subqueries at the top of a `SELECT`: + +```sql +WITH monthly AS (SELECT category, SUM(amount) AS total FROM orders GROUP BY category) +SELECT category, total FROM monthly; +``` + +A CTE reference **is** a derived table — the same construct under a name — so a statement carrying a +`WITH` runs where derived tables run: on the relational engine, at a venue that has it. + +### Scope is left to right + +Each CTE may reference the CTEs declared **before** it, never itself and never a later one: + +```sql +WITH a AS (SELECT id FROM orders), + b AS (SELECT id FROM a) +SELECT id FROM b; +``` + +A CTE may be referenced wherever the statement reads a table — in `FROM` and in a `JOIN`: + +```sql +WITH eu AS (SELECT id FROM departments WHERE region = 'EU') +SELECT e.name FROM employees e JOIN eu ON e.department_id = eu.id; +``` + +Naming a CTE inside a **`WHERE` subquery body** parses but does **not** run: the reference inside the +body is never resolved, so the body reaches Elasticsearch asking for an index with the CTE's name and +the statement fails with a `404 index_not_found_exception` naming it. Loud, never silent. Read the +CTE in `FROM` or `JOIN` instead. + +### A CTE referenced twice is executed twice + +The semantics are **inline**, not materialised: naming a CTE does not cache it. A CTE referenced +twice runs twice, which costs twice and — on a changing index — may not return identical rows both +times. + +```sql +WITH m AS (SELECT category, SUM(amount) AS total FROM orders GROUP BY category) +SELECT l.category, l.total, r.total AS again FROM m l JOIN m r ON l.category = r.category; +``` + +### What is not supported + +- **`WITH RECURSIVE`** — refused by name. There is no rewrite that recovers arbitrary-depth + recursion; flatten the hierarchy at index time (store a path or a level on each document), or run + one statement per level. +- **CTE column lists** — `WITH a (x, y) AS (…)`. Alias the columns in the CTE's own `SELECT` list. +- **A body that names the CTE itself** — `WITH orders AS (SELECT … FROM orders)` is refused, with a + rename suggested in the message. PostgreSQL binds such a name to the base table; this engine does + not guess. Give the CTE a different name. +- **`WITH` anywhere but the top of a `SELECT`** — not in a subquery body, not in `CREATE TABLE … AS + SELECT`, not in `INSERT … SELECT`, not in a materialized view. + +A column or alias genuinely named `recursive` still works — quote it (`WITH "recursive" AS (…)`). --- @@ -1318,7 +1703,11 @@ Notes: |--------------------------------|-----|-----|-----|-----| | Basic SELECT | ✔ | ✔ | ✔ | ✔ | | Nested fields | ✔ | ✔ | ✔ | ✔ | -| UNION ALL | ✔ | ✔ | ✔ | ✔ | +| Set operators | ✔ | ✔ | ✔ | ✔ | +| Derived tables | ✔ | ✔ | ✔ | ✔ | +| `WHERE` subqueries | ✔ | ✔ | ✔ | ✔ | +| Correlated subqueries | ✔ | ✔ | ✔ | ✔ | +| CTEs (non-recursive) | ✔ | ✔ | ✔ | ✔ | | Cross-index JOINs | ✔ | ✔ | ✔ | ✔ | | JOIN UNNEST | ✔ | ✔ | ✔ | ✔ | | Aggregations | ✔ | ✔ | ✔ | ✔ | @@ -1328,17 +1717,23 @@ Notes: | Date/time functions | ✔ | ✔ | ✔ | ✔ | | String / math functions | ✔ | ✔ | ✔ | ✔ | +A ✔ here means the Elasticsearch major does not stand in the way — every row was executed against +that version. It is not a statement about the venue: derived tables, CTEs, correlated subqueries and +the set operators other than `UNION ALL` need the relational engine wherever they run, on every one +of these versions alike. See +[Known Limitations & Roadmap](known_limitations.md#which-forms-need-the-relational-engine). + --- ## Limitations -For the full picture of what works in R1, what's coming in R2a/R2b, and BI-tool workarounds, see [Known Limitations & Roadmap](known_limitations.md). +For the full picture of what works in this release, what is refused and why, and BI-tool notes, see [Known Limitations & Roadmap](known_limitations.md). Even though the DQL engine is powerful, some SQL features are not (yet) supported: - Cross-index JOINs (`INNER` / `LEFT` / `RIGHT` / `FULL OUTER`) are supported across indices and clusters — see [Cross-Index JOIN](joins.md). `JOIN UNNEST` on `ARRAY` is the single-index nested form, handled natively inside one index. -- No correlated subqueries -- No arbitrary subqueries in `SELECT` or `WHERE` (except `INSERT ... AS SELECT` in DML) +- **Since engine `0.24.0`**, subqueries in `WHERE` (`IN` / `NOT IN` / `EXISTS` / `NOT EXISTS` / scalar / quantified) and derived tables in `FROM` / `JOIN` are supported, correlated or not — see [Subqueries and derived tables](#subqueries-and-derived-tables) for the reference and [Known Limitations & Roadmap](known_limitations.md#subqueries-and-derived-tables) for the venue requirements. Set operators — `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL` — are supported too; see [Set operators](#set-operators). Still not supported: a subquery in the `SELECT` list, a subquery in `HAVING`, `LATERAL`, and a set operation used as a subquery body. +- **Since engine `0.24.0`**, non-recursive CTEs (`WITH name AS (SELECT …)`) are supported at the top of a `SELECT`, each one able to reference the CTEs declared before it — see [Common table expressions](#common-table-expressions). Still not supported: `WITH RECURSIVE`, CTE column lists (`WITH a (x, y) AS …`), a CTE body that names the CTE itself, and a `WITH` clause anywhere other than the top of a `SELECT` (not in a subquery body, CTAS, `INSERT … SELECT` or a materialized view). - No `GROUPING SETS`, `CUBE`, `ROLLUP` - No `DISTINCT ON` - No explicit window frame clauses (`ROWS BETWEEN ...`) diff --git a/documentation/sql/functions_date_time.md b/documentation/sql/functions_date_time.md index a4df2296..015ed621 100644 --- a/documentation/sql/functions_date_time.md +++ b/documentation/sql/functions_date_time.md @@ -226,19 +226,33 @@ SELECT DATE_SUB('2025-01-10'::DATE, INTERVAL 1 QUARTER) AS last_quarter; --- -#### DATETIME_ADD / DATETIMEADD +#### DATETIME_ADD / DATETIMEADD / TIMESTAMPADD Adds interval to `DATETIME` / `TIMESTAMP`. +Two argument orders are accepted, and they mean the same thing. The second — **unit first, with a +plain count instead of an `INTERVAL`** — is the ODBC/JDBC and T-SQL order, which is what a BI tool +emits. `TIMESTAMPADD` is the ODBC name for it. + **Syntax:** ```sql DATETIME_ADD(datetime_expr, INTERVAL n UNIT) DATETIMEADD(datetime_expr, INTERVAL n UNIT) + +DATETIME_ADD(UNIT, n, datetime_expr) +TIMESTAMPADD(UNIT, n, datetime_expr) ``` **Inputs:** - `datetime_expr` - `DATETIME` or `TIMESTAMP` - `INTERVAL n UNIT` - where `UNIT` is one of: `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`, `HOUR`, `MINUTE`, `SECOND` +- In the unit-first form, `n` is a plain integer and may be **negative**; the ODBC interval names + `SQL_TSI_YEAR`, `SQL_TSI_QUARTER`, `SQL_TSI_MONTH`, `SQL_TSI_WEEK`, `SQL_TSI_DAY`, `SQL_TSI_HOUR`, + `SQL_TSI_MINUTE` and `SQL_TSI_SECOND` are accepted as spellings of the units above + +> There is **no `TIMESTAMPSUB`** — it exists in neither ODBC nor MySQL. Subtract by passing a +> negative count: `TIMESTAMPADD(DAY, -89, CURRENT_DATE)`. (`DATETIME_SUB` is available for the +> `INTERVAL` form.) **Output:** - `DATETIME` @@ -249,6 +263,15 @@ DATETIMEADD(datetime_expr, INTERVAL n UNIT) SELECT DATETIME_ADD('2025-01-10T12:00:00Z'::TIMESTAMP, INTERVAL 1 DAY) AS tomorrow; -- Result: 2025-01-11T12:00:00Z +-- The same thing in the ODBC unit-first order +SELECT TIMESTAMPADD(DAY, 1, '2025-01-10T12:00:00Z'::TIMESTAMP) AS tomorrow; +-- Result: 2025-01-11T12:00:00Z + +-- ODBC interval name, and a negative count instead of a subtraction +SELECT TIMESTAMPADD(SQL_TSI_DAY, 1, '2025-01-10T12:00:00Z'::TIMESTAMP) AS tomorrow; +-- Result: 2025-01-11T12:00:00Z +SELECT TIMESTAMPADD(DAY, -89, CURRENT_DATE) AS ninety_days_ago; + -- Add 2 hours SELECT DATETIME_ADD('2025-01-10T12:00:00Z'::TIMESTAMP, INTERVAL 2 HOUR) AS later; -- Result: 2025-01-10T14:00:00Z @@ -316,62 +339,126 @@ SELECT DATETIME_SUB('2025-01-10T12:00:00Z'::TIMESTAMP, INTERVAL 1 MONTH) AS last ### Date/Time Difference Functions -#### DATEDIFF / DATE_DIFF +#### DATE_DIFF / TIMESTAMPDIFF + +Difference between 2 dates, in the specified time unit. **The result is `date2 - date1`** — the +first argument is the START and the second is the END. That is what makes +`DATE_DIFF(birthdate, CURRENT_DATE, YEAR)` an age rather than a negative one, and it is how MySQL +defines `TIMESTAMPDIFF(unit, dt1, dt2)` too. + +`TIMESTAMPDIFF` is the ODBC/JDBC spelling and takes the **unit first**. -Difference between 2 dates (date1 - date2) in the specified time unit. +> ⚠️ **`DATEDIFF` is a different function** — MySQL's — and subtracts the other way round. It has +> its own entry below. Do not assume the two names are interchangeable. **Syntax:** ```sql -DATEDIFF(date1, date2) -DATEDIFF(date1, date2, unit) DATE_DIFF(date1, date2) DATE_DIFF(date1, date2, unit) + +DATE_DIFF(unit, date1, date2) +TIMESTAMPDIFF(unit, date1, date2) ``` +> The unit-first form is the one the engine **renders back**, so `TIMESTAMPDIFF(DAY, a, b)` reads +> as `DATE_DIFF(DAY, a, b)` wherever a statement is echoed to you. + **Inputs:** -- `date1` - `DATE` or `DATETIME` -- `date2` - `DATE` or `DATETIME` -- `unit` (optional) - One of: `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`, `HOUR`, `MINUTE`, `SECOND` +- `date1` - `DATE` or `DATETIME` — the **start** +- `date2` - `DATE` or `DATETIME` — the **end** +- `unit` (optional in the date-first form) - One of: `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`, `HOUR`, `MINUTE`, `SECOND` - Default: `DAY` + - The ODBC `SQL_TSI_*` spellings are accepted here too **Output:** -- `BIGINT` +- `BIGINT` — negative when `date2` is earlier than `date1` **Examples:** ```sql -- Difference in days (default) -SELECT DATEDIFF('2025-01-10'::DATE, '2025-01-01'::DATE) AS diff; +SELECT DATE_DIFF('2025-01-01'::DATE, '2025-01-10'::DATE) AS diff; -- Result: 9 --- Difference in days (explicit) -SELECT DATEDIFF('2025-01-10'::DATE, '2025-01-01'::DATE, DAY) AS diff_days; --- Result: 9 +-- Reverse the arguments and the sign reverses +SELECT DATE_DIFF('2025-01-10'::DATE, '2025-01-01'::DATE, DAY) AS diff; +-- Result: -9 --- Difference in weeks -SELECT DATE_DIFF('2025-01-31'::DATE, '2025-01-01'::DATE, WEEK) AS diff_weeks; +-- Weeks, months, years +SELECT DATE_DIFF('2025-01-01'::DATE, '2025-01-31'::DATE, WEEK) AS diff_weeks; -- Result: 4 - --- Difference in months -SELECT DATEDIFF('2025-06-01'::DATE, '2025-01-01'::DATE, MONTH) AS diff_months; +SELECT DATE_DIFF('2025-01-01'::DATE, '2025-06-01'::DATE, MONTH) AS diff_months; -- Result: 5 - --- Difference in years -SELECT DATEDIFF('2027-01-01'::DATE, '2025-01-01'::DATE, YEAR) AS diff_years; +SELECT DATE_DIFF('2025-01-01'::DATE, '2027-01-01'::DATE, YEAR) AS diff_years; -- Result: 2 --- Difference in hours (with timestamps) -SELECT DATEDIFF('2025-01-10T14:00:00Z'::TIMESTAMP, '2025-01-10T12:00:00Z'::TIMESTAMP, HOUR) AS diff_hours; +-- Hours, minutes, seconds +SELECT DATE_DIFF('2025-01-10T12:00:00Z'::TIMESTAMP, '2025-01-10T14:00:00Z'::TIMESTAMP, HOUR) AS diff_hours; -- Result: 2 - --- Difference in minutes -SELECT DATEDIFF('2025-01-10T12:30:00Z'::TIMESTAMP, '2025-01-10T12:00:00Z'::TIMESTAMP, MINUTE) AS diff_minutes; +SELECT DATE_DIFF('2025-01-10T12:00:00Z'::TIMESTAMP, '2025-01-10T12:30:00Z'::TIMESTAMP, MINUTE) AS diff_minutes; -- Result: 30 - --- Difference in seconds -SELECT DATEDIFF('2025-01-10T12:00:45Z'::TIMESTAMP, '2025-01-10T12:00:00Z'::TIMESTAMP, SECOND) AS diff_seconds; +SELECT DATE_DIFF('2025-01-10T12:00:00Z'::TIMESTAMP, '2025-01-10T12:00:45Z'::TIMESTAMP, SECOND) AS diff_seconds; -- Result: 45 + +-- The ODBC spelling, unit first - same answer, same sign +SELECT TIMESTAMPDIFF(DAY, '2025-01-01'::DATE, '2025-01-10'::DATE) AS diff_days; +-- Result: 9 + +-- Age: start = birthdate, end = today +SELECT DATE_DIFF(birthdate, CURRENT_DATE, YEAR) AS age FROM users; +``` + +--- + +#### DATEDIFF + +MySQL's day-difference function. **`DATEDIFF(expr1, expr2)` returns `expr1 - expr2`** — the opposite +subtraction from `DATE_DIFF` above, because that is what MySQL defines and what a statement written +for MySQL expects. + +> 🔴 **The two-argument and three-argument forms subtract in opposite directions, and this is the +> one place in the dialect where that is true.** +> +> | call | result | why | +> | --- | --- | --- | +> | `DATEDIFF(a, b)` | `a - b` | MySQL's function, days only — this is the whole of what MySQL defines | +> | `DATEDIFF(a, b, unit)` | `b - a` | **not MySQL** — this engine's own extension, which follows `DATE_DIFF` | +> +> So adding a unit to a two-argument `DATEDIFF` — which looks purely clarifying — **reverses the +> sign**. If you want a unit, prefer `DATE_DIFF(start, end, unit)` or +> `TIMESTAMPDIFF(unit, start, end)`, where one rule holds at every arity. + +**Syntax:** +```sql +DATEDIFF(expr1, expr2) +DATEDIFF(date1, date2, unit) ``` +**Inputs:** +- `expr1`, `expr2` - `DATE` or `DATETIME` +- `unit` (three-argument form only) - as for `DATE_DIFF` above + +**Output:** +- `BIGINT` + +**Examples:** +```sql +-- MySQL's own two documented examples, and this engine agrees with them +SELECT DATEDIFF('2007-12-31'::DATE, '2007-12-30'::DATE) AS diff; +-- Result: 1 +SELECT DATEDIFF('2010-11-30'::DATE, '2010-12-31'::DATE) AS diff; +-- Result: -31 + +-- The two arities disagree, on purpose +SELECT DATEDIFF('2025-01-10'::DATE, '2025-01-01'::DATE) AS diff; +-- Result: 9 +SELECT DATEDIFF('2025-01-10'::DATE, '2025-01-01'::DATE, DAY) AS diff; +-- Result: -9 +``` + +> **Before engine `0.24.0`, `DATEDIFF` returned the opposite sign** — it was an alias of `DATE_DIFF` +> and inherited its order, so MySQL's `DATEDIFF('2007-12-31','2007-12-30')` answered `-1` here +> instead of `1`. Statements written against the old behaviour need their arguments swapped. + --- ### Date/Time Formatting Functions diff --git a/documentation/sql/functions_string.md b/documentation/sql/functions_string.md index 22b85a98..468d62e7 100644 --- a/documentation/sql/functions_string.md +++ b/documentation/sql/functions_string.md @@ -239,14 +239,22 @@ FROM products; ### String Measurement Functions -#### LENGTH / LEN +#### LENGTH / LEN / CHAR_LENGTH / CHARACTER_LENGTH Character length of string. +> **It counts CHARACTERS, not bytes**, and that is why `CHAR_LENGTH` is an accepted spelling here. +> MySQL's own `LENGTH` counts *bytes*, so a statement written against MySQL may say `CHAR_LENGTH` +> precisely to avoid that — and it means exactly what our `LENGTH` already did. +> `CHAR_LENGTH('café')` is **4**, not 5. `CHAR_LENGTH` and `CHARACTER_LENGTH` are the SQL-92 +> spellings; BI tools emit them for a "string length" calculated field. + **Syntax:** ```sql LENGTH(str) LEN(str) +CHAR_LENGTH(str) +CHARACTER_LENGTH(str) ``` **Inputs:** @@ -277,6 +285,12 @@ SELECT LENGTH('hello world') AS l; SELECT LENGTH('café') AS l; -- Result: 4 +-- The SQL-92 spellings, and the reason they exist: 4 characters, 5 UTF-8 bytes +SELECT CHAR_LENGTH('café') AS l; +-- Result: 4 +SELECT CHARACTER_LENGTH('hello world') AS l; +-- Result: 11 + -- Filter by length SELECT * FROM products WHERE LENGTH(name) > 20; diff --git a/documentation/sql/joins.md b/documentation/sql/joins.md index faa617f9..2836285f 100644 --- a/documentation/sql/joins.md +++ b/documentation/sql/joins.md @@ -33,7 +33,7 @@ Cross-index JOIN ships in three shapes ("rows"). Pick by where your data lives: The rule the engine actually applies: it counts the **distinct source catalogs** in the rewritten `FROM` / `JOIN` clauses versus the target catalog. Same (or no) catalog → Row 1; exactly one source catalog different from the target → Row 2; two or more source catalogs → Row 3. -> **What does NOT work yet:** Cross-index JOINs are first-class in this release, but two things are intentionally **not** here yet: arbitrary **subqueries / CTEs** in a JOIN query land in **the next release (Quarter 4 2026)**, and **heterogeneous Row-3 sources** (joining ES with Postgres, MySQL, Snowflake, …) land in **the upcoming release (Quarter 1 2027)** — this release's Row 3 is multi-**Elasticsearch** only. See [Known limitations](known_limitations.md) for the full list. +> **What does NOT work yet:** Cross-index JOINs are first-class, and since engine **`0.24.0`** so are **subqueries, derived tables, non-recursive CTEs and set operators** — including a derived table as a JOIN leg, a JOIN inside a derived body, and a JOIN inside a set-operation branch (see [Known limitations](known_limitations.md#subqueries-and-derived-tables)). One thing is intentionally **not** here yet: **heterogeneous Row-3 sources** (joining ES with Postgres, MySQL, Snowflake, …), which land in **the upcoming release (Quarter 1 2027)** — this release's Row 3 is multi-**Elasticsearch** only. --- @@ -389,7 +389,8 @@ For the full price matrix and editions, see the licensing & pricing page on the ## What does NOT work yet -- **Arbitrary subqueries and CTEs** inside a JOIN query — coming in **the next release (Quarter 4 2026)**. +- **Set operators, subqueries, derived tables and non-recursive CTEs all work since engine `0.24.0`**: a derived table can be a JOIN leg, a derived body can itself carry a JOIN, a CTE reference is a derived table and so can be a JOIN leg too, and a set-operation branch may carry a JOIN of its own. A correlated `WHERE` subquery counts as **one** unit against `maxJoins`; a derived table, a CTE reference and a set operation count none. See [Subqueries and derived tables](known_limitations.md#subqueries-and-derived-tables) and [Set operators](known_limitations.md#set-operators). +- **A correlated subquery body that is not a single Elasticsearch source** — it may not carry its own `JOIN`, comma-separated `FROM`, `JOIN UNNEST`, derived table or window function. Move the construct to the outer `FROM` and correlate against it. - **Heterogeneous Row-3 sources** (joining Elasticsearch with Postgres, MySQL, Snowflake, …) — coming in **the upcoming release (Quarter 1 2027)**; this release's Row 3 is multi-Elasticsearch only. - **JOIN inside a watcher input** (`CREATE WATCHER … FROM a JOIN b ON …`) — a watcher input is a single Elasticsearch `search` request over a list of indices, so the join is rejected at parse time. Pre-join the sources with a [materialized view](materialized_views.md) and have the watcher search the view. diff --git a/documentation/sql/keywords.md b/documentation/sql/keywords.md index d3b6a6b2..de76cf9d 100644 --- a/documentation/sql/keywords.md +++ b/documentation/sql/keywords.md @@ -2,7 +2,25 @@ # Keywords -A list of reserved words recognized by the parser for this engine. +The words the parser recognises. Two different sets live on this page, and the difference matters when +you name a column: + +- **Recognised** — the word has a meaning in the grammar. Everything listed below is recognised. +- **Reserved** — the word additionally **cannot be used as a bare identifier**. Most, but *not all*, of + the words below are reserved. + +`EXISTS` is reserved, so `SELECT exists FROM t` is a parse error. `ANY` and `SOME` are **deliberately not +reserved**, so `SELECT any, some FROM t WHERE any = 1` parses as columns — even though `x = ANY (SELECT …)` +is real grammar. The same split runs through the set-operator and CTE words: `UNION`, `INTERSECT`, `EXCEPT`, +`ALL` and `DISTINCT` are reserved, while `WITH` and `RECURSIVE` are not — `SELECT a AS recursive FROM t` +parses, `SELECT a AS intersect FROM t` does not. `TOP` is recognised (`SELECT TOP 10 id FROM t` bounds the rows) and **deliberately not reserved**, so +`SELECT top FROM t` still selects a column called `top`. `PERCENT` is recognised only so that +`SELECT TOP n PERCENT` can be refused *by name* — see +[Known limitations](known_limitations.md). + +If you have a column whose name collides with a reserved word, **quote it** rather than +renaming it: `SELECT "exists" FROM t` works, and so does the backtick spelling — see +[Quoted identifiers](dql_statements.md#quoted-identifiers). ## Main clauses COPY @@ -25,9 +43,20 @@ NULLS FIRST NULLS LAST OFFSET LIMIT +TOP +PERCENT ON CONFLICT DO +UNION ALL +UNION +UNION DISTINCT +INTERSECT +INTERSECT ALL +EXCEPT +EXCEPT ALL +WITH +RECURSIVE SHOW DESCRIBE EVERY @@ -79,6 +108,8 @@ TRIM LTRIM RTRIM LENGTH +CHAR_LENGTH +CHARACTER_LENGTH SUBSTRING SUBSTR CONCAT @@ -156,10 +187,12 @@ DATE_SUB DATESUB DATETIME_ADD DATETIMEADD +TIMESTAMPADD DATETIME_SUB DATETIMESUB DATE_DIFF DATEDIFF +TIMESTAMPDIFF DATE_FORMAT DATE_PARSE DATETIME_FORMAT @@ -181,6 +214,22 @@ NOT IN NOT BETWEEN IS NULL IS NOT NULL +EXISTS +NOT EXISTS +ALL +ANY +SOME + +`EXISTS` and `ALL` are **reserved** — a column of either name must be quoted. `ANY` and `SOME` are +recognised but **not reserved**, on purpose: a column called `any` keeps parsing, and the grammar tells the +two readings apart by what follows. + +These five words introduce the subquery predicates. `= ANY` and `= SOME` mean `IN`, and `<> ALL` means +`NOT IN` — the engine normalises them, so `WHERE customer_id = ANY (SELECT id FROM customers)` is stored +and re-rendered as `WHERE customer_id IN (SELECT id FROM customers)`. The ordering quantifiers +(`> ALL`, `>= ANY`, `< ALL`, …) keep their operator, but **`SOME` always normalises to `ANY`** — a +statement written `<= SOME (…)` is stored and re-rendered as `<= ANY (…)`. See +[Subqueries and derived tables](known_limitations.md#subqueries-and-derived-tables). ## Logical operators AND diff --git a/documentation/sql/known_limitations.md b/documentation/sql/known_limitations.md index b40f44fb..e4e2d09a 100644 --- a/documentation/sql/known_limitations.md +++ b/documentation/sql/known_limitations.md @@ -2,9 +2,11 @@ # Known Limitations & Roadmap -SoftClient4ES runs a large, practical subset of ANSI SQL on Elasticsearch — including cross-index JOINs that Elasticsearch itself cannot do. A few advanced constructs (subqueries, CTEs, set operators beyond `UNION ALL`) are not in the current release yet. This page tells you exactly what works **as of this release**, what's coming, and how to get unblocked today. +SoftClient4ES runs a large, practical subset of ANSI SQL on Elasticsearch — including cross-index JOINs, and, since engine `0.24.0`, subqueries, derived tables, non-recursive CTEs and the `UNION` / `INTERSECT` / `EXCEPT` set operators, none of which Elasticsearch can do itself. This page tells you exactly what works **as of this release**, what's coming, and how to get unblocked today. -> Great for explicit JOIN SQL — full BI-tool subquery / CTE support is coming in the next release. +> **Since engine `0.24.0`:** subqueries, derived tables, non-recursive CTEs and set operators +> (`UNION`, `UNION ALL`, `INTERSECT`, `EXCEPT`, with or without `ALL`) all work, including the nested +> SQL BI tools generate for you. ## Using a BI tool? Read this first @@ -24,31 +26,38 @@ Two different things can stop a BI tool here, and it is worth separating them. - **Looker** — Looker connects only through drivers it maintains itself, and it allowlists JDBC parameters per dialect, so a customer-supplied driver cannot be introduced. This gap is **structural, not commercial** — a licence would not close it. +- **dbt** — dbt requires a dedicated adapter plugin per platform. There is no generic JDBC or ODBC adapter, + and no SoftClient4ES adapter. -Neither is a gap we can close from our side: each needs either a change by the vendor or a driver plugin -that nobody has written. +None of these is a gap we can close from our side: each one needs either a change by the vendor or a driver +or adapter plugin that nobody has written. -*(Each blocker checked against the vendor's own connection documentation — Metabase, Microsoft Power Query -and Looker — on 2026-08-31 and 2026-09-01.)* +*(Each blocker checked against the vendor's own connection documentation — Metabase, Microsoft Power Query, +Looker and dbt — on 2026-08-31 and 2026-09-01.)* -### Tools that connect, but generate SQL we do not accept yet +### Tools that generate nested SQL for you Some BI tools auto-generate nested SQL (subqueries / derived tables) even when your logical query has none. -Until the next release lands full subquery support, send **explicit JOIN SQL** instead of letting the tool -compose nested queries — where the tool lets you: - -- **Apache Superset / DBeaver / Grafana** — you control the SQL. Write explicit JOINs for anything that - would otherwise nest, and everything in **Works in this release** below is available to you. -- **Tableau** — connecting and browsing work; queries are the constrained part. Drag-and-drop worksheets - quote and fully qualify every identifier, a form we do not accept yet, and **Custom SQL is not a way - around it**: Tableau documents that it *"must wrap the custom SQL statement within a select statement"* (Tableau's Custom SQL - documentation, checked 2026-09-01), - which turns your query into a derived table. **Extract** mode narrows the exposure but does not remove - it — the extract is still built by querying the source. - See [Tableau](../client/bi_tools.md). - -> **General rule:** prefer **explicit JOIN SQL** over tool-generated nested SQL. If you control the query, a -> cross-index JOIN is fully supported in the current release. +**Since engine `0.24.0` that form is accepted** — you no longer have to rewrite it as an explicit JOIN. + +- **Apache Superset / DBeaver / Grafana** — you control the SQL. Subqueries, derived tables and explicit + JOINs are all available; everything in **Works in this release** below applies. +- **Tableau** — connecting, browsing, previewing, aggregating, filtering and sorting work. Drag-and-drop + worksheets quote and fully qualify every identifier (backticks under the MySQL dialect, + `"schema"."table"` under Generic SQL-92) and wrap the query in a derived table; **both quoted forms + parse since engine `0.23.0`, and since `0.24.0` the derived-table wrapper also EXECUTES**, on the + relational engine. Tableau's **Custom SQL** wraps your statement too — it documents that it *"must wrap the custom + SQL statement within a select statement"* (Tableau's Custom SQL documentation, checked 2026-09-01) — and + that wrapper is a derived table, which now runs. **Extract** mode remains **untested** against + SoftClient4ES. See [Tableau](../client/bi_tools.md). + +> **One thing to check before you rely on it:** a derived table runs on the relational engine — **since +> engine `0.24.0` with arrow-extensions `0.3.4`** — so the venue executing your SQL must carry the +> `softclient4es-arrow-extensions` jar as well as the engine. See +> [Which forms need the relational engine](#which-forms-need-the-relational-engine) below. **You almost +> certainly have it already**: the JDBC driver, the ADBC driver and the Arrow Flight SQL sidecar all ship +> with it, and `install.sh` installs it with the REPL by default. Only a REPL installed explicitly with +> `--no-extensions` lacks it. **Apache Superset** (dedicated dialect), **DBeaver**, and **Grafana** (via Arrow Flight SQL) are **Tested**. **Tableau** is **Compatible** — the connection path works, but it is not yet in our formal regression suite. @@ -56,40 +65,234 @@ compose nested queries — where the tool lets you: ## Works in this release - **Cross-index JOINs**: `INNER` / `LEFT` / `RIGHT` / `FULL` / `CROSS`, plus `JOIN UNNEST` on nested arrays — something Elasticsearch cannot do natively. (See the [JOIN matrix walkthrough](joins.md) for the per-tier rows and worked examples.) +- **Subqueries in `WHERE`** — *since engine `0.24.0`*: `IN (SELECT …)` / `NOT IN`, `EXISTS` / + `NOT EXISTS`, a scalar comparison against `(SELECT …)`, and the quantified forms `= ANY | SOME`, + `<> ALL`, `> ALL`, `>= ANY`, `< ALL`, … — **correlated or not**. See + [Subqueries and derived tables](#subqueries-and-derived-tables) below. +- **Derived tables** — *since engine `0.24.0`*: `FROM (SELECT …) d` and `JOIN (SELECT …) d ON …`, nested + to any depth — including bodies that themselves carry a JOIN or another derived table. - **Aggregations** + `GROUP BY` / `HAVING`. - **Analytical SQL**: `ROW_NUMBER` / `RANK` / `DENSE_RANK`; the `STDDEV` / `VARIANCE` family (`STDDEV_POP`, `STDDEV_SAMP`, `VAR_POP`, `VAR_SAMP`); `PERCENTILE_CONT` / `PERCENTILE_DISC`; window aggregates and `FIRST_VALUE` / `LAST_VALUE` / `ARRAY_AGG` over `OVER (PARTITION BY …)`. - **Conditionals & null handling**: `CASE` / `COALESCE` / `NULLIF` / `GREATEST` / `LEAST` / `ISNULL` / `ISNOTNULL`. - `ORDER BY … NULLS FIRST | NULLS LAST`. -- `UNION ALL` (concatenate result sets — no de-duplication). -- `SELECT * EXCEPT(col, …)` — drop named columns from `SELECT *`. This is the BigQuery-style **column-exclusion** clause. It is **not** the `EXCEPT` set operator (see below). +- **Non-recursive CTEs** — *since engine `0.24.0`*: `WITH name AS (SELECT …)` at the top of a `SELECT`, + chained left to right. A CTE reference *is* a derived table, so it runs where derived tables run. +- **Set operators** — *since engine `0.24.0`*: `UNION ALL`, `UNION` / `UNION DISTINCT`, `INTERSECT` / + `INTERSECT ALL`, `EXCEPT` / `EXCEPT ALL`. See [Set operators](#set-operators) below. +- **BI dialect spellings** — *since engine `0.24.0`*: `CHAR_LENGTH` / `CHARACTER_LENGTH` (for + `LENGTH`), `TIMESTAMPADD` (for `DATETIME_ADD`) and `TIMESTAMPDIFF` (for `DATE_DIFF`) including the + ODBC unit-first argument order and the `SQL_TSI_*` interval names, and `SELECT TOP n` (for + `LIMIT n`). These are spellings, not new behaviour — each renders as the canonical form. +- `SELECT * EXCEPT(col, …)` — drop named columns from `SELECT *`. This is the BigQuery-style **column-exclusion** clause. It removes *columns*; the `EXCEPT` **set operator** removes *rows*. Both work, and they are unrelated. + +## Subqueries and derived tables + +**Since engine `0.24.0`.** Earlier releases refuse every form below, so check your engine version +before planning around them. (Where they refuse it varies by release and by form — `0.23.0`, for +instance, parses a derived table and refuses it in the engine — so do not rely on the error you get, +only on the version.) + +An **uncorrelated** `WHERE` subquery needs nothing but the engine: it executes on Elasticsearch itself, at +every venue. **Correlated subqueries and derived tables additionally need the relational engine — since +engine `0.24.0` with arrow-extensions `0.3.4`.** See +[Which forms need the relational engine](#which-forms-need-the-relational-engine). + +Every form below **parses and executes**. The examples are literal — they are the shapes the engine +accepts. -## Not in this release (coming in the next release, Quarter 4 2026) +```sql +-- IN / NOT IN over a subquery +SELECT name FROM employees +WHERE department_id IN (SELECT id FROM departments WHERE region = 'EU'); -- **Subqueries**: scalar, `IN (SELECT …)`, `EXISTS (SELECT …)`, derived tables `FROM (SELECT …)`. -- **CTEs**: `WITH name AS (SELECT …)` — recursive and non-recursive. -- **Set operators**: `UNION` (with row de-duplication), `INTERSECT`, and the `EXCEPT` **set operator**. The `EXCEPT` set operator is **distinct from** the `SELECT * EXCEPT(cols)` column-exclusion clause above — that one works; the set operator does not. -- **Positional / tiling window functions**: `NTILE`, `LAG`, `LEAD` — not yet implemented; coming with the next release's analytical-SQL work. (Note: `PERCENTILE_CONT` / `PERCENTILE_DISC` — percentile *aggregates* — already work in the current release; the positional/tiling window functions are a different family.) +-- scalar comparison +SELECT name FROM employees +WHERE salary > (SELECT AVG(salary) FROM employees); -These arrive in the next release as a driver-side enhancement — single-cluster customers get them by upgrading the driver (JDBC / ADBC / sidecar), with no infrastructure change and no federation server required. +-- quantified comparison (= ANY | SOME, <> ALL, > ALL, >= ANY, < ALL, …) +SELECT name FROM employees +WHERE salary >= ALL (SELECT salary FROM employees WHERE department = 'IT'); -### What a not-yet-supported query looks like +-- EXISTS / NOT EXISTS, correlated against the outer row +SELECT c.name FROM customers c +WHERE NOT EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id); + +-- derived table in FROM … +SELECT d.category, d.n +FROM (SELECT category, COUNT(*) AS n FROM bi_events GROUP BY category) d +WHERE d.n > 10; + +-- … and in JOIN +SELECT o.id, c.name +FROM orders o +JOIN (SELECT id, name FROM customers WHERE tier = 'gold') c ON o.customer_id = c.id; +``` + +### Which forms need the relational engine + +This is the distinction worth knowing before you plan around it. -A subquery in a `WHERE` clause is rejected by the parser today: +| Form | Runs where | Needs `softclient4es-arrow-extensions`? | +| --- | --- | --- | +| **Uncorrelated** `WHERE` subquery — `IN` / `NOT IN` / `EXISTS` / `NOT EXISTS` / scalar / quantified | Elasticsearch, in two phases: the inner statement is executed first, then the outer one is rewritten against its values | **No** — works at every venue, including a plain REPL with no extensions | +| **Correlated** `WHERE` subquery (the body reads an outer alias) | The relational engine | **Yes** — arrow-extensions `0.3.4` | +| **Derived table** in `FROM` or `JOIN` | The relational engine | **Yes** — arrow-extensions `0.3.4` | +| **Non-recursive CTE** (`WITH name AS (SELECT …)`) | The relational engine — a CTE reference *is* a derived table | **Yes** — arrow-extensions `0.3.4` | +| **`UNION ALL`** | Elasticsearch, one `_msearch`, branches concatenated in order | **No** — works at every venue | +| **`UNION` / `INTERSECT` / `EXCEPT`** (with or without `ALL`) | The relational engine | **Yes** — arrow-extensions `0.3.4` | + +A venue without that jar does not guess: it refuses the statement with an HTTP 400 naming the construct and +the jar, rather than executing it against the first index the statement mentions. **The default install has +the engine** — the JDBC driver, the ADBC driver and the Arrow Flight SQL sidecar all bundle it, and +`install.sh` installs it alongside the REPL unless you pass `--no-extensions`. + +### The bound on an uncorrelated subquery + +The two-phase path resolves the inner statement into a set of values, so it is bounded by what an +Elasticsearch `terms` query accepts — **65,536 distinct values** (`index.max_terms_count`). Past that the +statement fails loudly, naming the limit and suggesting the JOIN rewrite; it is never silently truncated. A +plain `SELECT FROM …` body is resolved with a single bounded `terms` aggregation, so the values are +already distinct and `DISTINCT` buys nothing. + +`NULL` follows ANSI: `IN` ignores NULLs in the inner values, `NOT IN` over a set containing a NULL matches no +rows, and an `EXISTS` over an empty body is false while `NOT EXISTS` over one is true. + +> ⚠️ **`NOT IN` has one carve-out, and it fails the other way.** The engine detects the `NULL` by +> re-running the inner query with an `IS NULL` filter. When that inner query carries a `GROUP BY`, +> the probe is a grouped query too, and Elasticsearch's `terms` aggregation **drops the +> missing-value group** — so a `NULL` in a `GROUP BY` body is invisible and `NOT IN` returns rows +> the rule above says it should not. Filter the `NULL` out explicitly in that body +> (`… WHERE IS NOT NULL GROUP BY …`) rather than relying on the detection. + +`DELETE` and `UPDATE` take the same `WHERE` subqueries as `SELECT`, with the same venue rules: ```sql --- Not supported in the current release: subqueries are not yet implemented. -SELECT name -FROM employees -WHERE department_id IN (SELECT id FROM departments WHERE region = 'EU'); +DELETE FROM orders WHERE customer_id IN (SELECT id FROM customers WHERE region = 'EU'); +UPDATE orders SET status = 'eu' WHERE customer_id IN (SELECT id FROM customers WHERE region = 'EU'); ``` -The parser rejects this — `IN` accepts only literal value lists today, not a nested `SELECT`. Rewrite it as an explicit JOIN (fully supported), or wait for the next release where the subquery form lands as-is. +### Subquery forms that are still refused + +Each of these is refused, never silently mis-executed. Most are refused **by name**, with the rewrite +in the message; where the refusal is a bare `end of input expected` instead, it is said so, because a +message that names nothing is the one you will need this page for: + +- **`LATERAL`** — a derived table that reads an alias from the enclosing `FROM` + (`FROM orders o, (SELECT id FROM customers WHERE id = o.customer_id) d`). Move the condition to the outer + `WHERE`, or write it as a correlated `WHERE` subquery. *Named only in that comma-`FROM` spelling: the + `LATERAL` keyword itself (`FROM o, LATERAL (…)`, `JOIN LATERAL …`) is a bare syntax error.* +- **A subquery in `HAVING`** — any subquery, correlated or not. Compute the value separately, or move the + condition to `WHERE`. +- **A subquery in the `SELECT` list** — `SELECT (SELECT MAX(amount) FROM orders) AS m …` does not parse. +- **A `UNION ALL` body** — `IN (SELECT a FROM t1 UNION ALL SELECT a FROM t2)`. Write one subquery per branch. +- **A `FROM`-less body** — `IN (SELECT 1)`. Write the literal list instead. +- **More than one projected column** — an `IN` / quantified / scalar body must project exactly one column, so + `IN (SELECT * FROM customers)` is refused. +- **A QUOTED outer reference** — `WHERE o.customer_id = "c"."id"`. Write it unquoted; the engine rewrites + an outer reference onto the extracted leg, and a quoted identifier is not rewritten. +- **A correlated body that is not a single Elasticsearch source** — its own `JOIN`, comma-separated `FROM`, + `JOIN UNNEST`, derived table or window function. Move the construct to the outer `FROM` and correlate + against it. +- **A scalar or quantified subquery on the LEFT of the comparison.** `WHERE (SELECT COUNT(*) FROM orders o + WHERE o.customer_id = c.id) > 5` is rejected, and the message it produces (`Unbalanced parentheses`) does + not say why. Flip the comparison — `WHERE 5 < (SELECT COUNT(*) …)` means the same thing and is accepted. + The subquery must be the right-hand operand. +- **A column list on the derived table's correlation name** — `FROM (SELECT id FROM orders) AS d (x)`. + Alias the columns inside the body instead: `(SELECT id AS x FROM orders) AS d`. *Syntax error, not + a named refusal.* +- **A subquery in a `CASE WHEN` condition** — refused by name, with the rewrite: filter in `WHERE`, or + compute the flag in a separate query. +- **A subquery in a `JOIN … ON` clause.** ⚠️ Its message names neither subqueries nor a rewrite — it + reads *"ON clause … must use either equality operator or AND predicate"*. Join on a plain equality + and move the subquery to `WHERE`. +- **A `WHERE` subquery in a `CREATE MATERIALIZED VIEW`** — refused by name: an Elasticsearch transform + cannot run the inner query. Resolve the subquery into the view's own source, or keep it in the + queries you run against the view. +- **A `FROM`-less `SELECT` as a set-operation branch** — `SELECT 1 UNION ALL SELECT id FROM orders`. + *Syntax error, not a named refusal.* Note this is the one place the connection-handshake idiom + `SELECT 1` does not compose. + +> ⚠️ **One mistake in this family is NOT refused, and it is the easiest one to make.** An +> **unqualified** outer reference — `WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = id)` +> instead of `… = c.id` — is a perfectly legal statement, so nothing can reject it. The bare `id` +> binds to the subquery's OWN table, the statement stops being correlated, and it runs as an +> ordinary uncorrelated subquery: **HTTP 200, and different rows from the ones you meant.** It is +> the only item here that fails silently rather than loudly. +> +> Always qualify the outer reference with the outer query's alias, and leave it unquoted. + +### Licensing + +A **correlated** subquery counts as one relational operation against your plan's `maxJoins` allowance, the +same as a JOIN clause — it is a semi-, anti- or aggregate-join the engine executes over two extracted +sources. A **derived table** costs nothing on its own; the JOINs *inside* it count, at any nesting depth. + +## Set operators + +**Since engine `0.24.0`.** Earlier releases accept `UNION ALL` only and reject every other spelling at the +parser. + +| Spelling | Duplicates | Runs where | Needs `softclient4es-arrow-extensions`? | +| --- | --- | --- | --- | +| `UNION ALL` | kept | Elasticsearch, one `_msearch`, results concatenated in branch order | **No** — every venue, a plain REPL included | +| `UNION` / `UNION DISTINCT` | removed | The relational engine | **Yes** — arrow-extensions `0.3.4` | +| `INTERSECT` / `INTERSECT ALL` | removed / kept | The relational engine | **Yes** — arrow-extensions `0.3.4` | +| `EXCEPT` / `EXCEPT ALL` | removed / kept | The relational engine | **Yes** — arrow-extensions `0.3.4` | + +Elasticsearch has no operation that de-duplicates or intersects across independent searches, so everything +but `UNION ALL` is executed by the same relational engine that runs cross-index JOINs and derived tables. A +venue without that jar refuses the statement rather than answering from one branch. + +A branch may carry anything a `SELECT` can carry — `GROUP BY`, a `JOIN`, a derived table, a CTE, a +correlated subquery. A branch that needs the relational engine on its own account routes the whole +statement there. + +Full syntax, precedence and the matching rules: [Set operators](dql_statements.md#set-operators). + +### Columns match by position + +Branches are matched **column by column**, and the result takes the **first branch's** column names — the +standard's rule (SQL-92 §7.10), and what every other SQL engine does. Column names are never compared, so +`SELECT id AS x … UNION ALL SELECT id AS y …` returns **one** column named `x` carrying both branches' ids. + +> **Changed in `0.24.0`:** before this release branches were matched **by name**, so a column +> the other branch did not name came back `NULL` — including for the first branch's own rows. If you have a +> `UNION ALL` written against the old behaviour, check that its branches project their columns in the same +> order. + +A branch written as a bare `SELECT *` declares no column list, so there is nothing to match positionally; +such a branch is matched by name instead and its width cannot be checked. Name the columns explicitly +whenever a branch's shape matters. + +### Set-operator forms that are still refused + +Each is rejected by name, never silently mis-executed: + +- **A set operation as a subquery body** — `WHERE a IN (SELECT … UNION SELECT …)`. Write one subquery per + branch. +- **A parenthesised set operation** — both `(a UNION b) INTERSECT c` and a whole statement wrapped in + parentheses. To group against the default precedence (`INTERSECT` binds tighter than `UNION` / `EXCEPT`), + use a derived table: `SELECT * FROM (a UNION b) AS g INTERSECT c`. +- **A trailing `ORDER BY` / `LIMIT` after the last branch** of a `UNION`, `INTERSECT` or `EXCEPT` — it would + silently bind to that branch alone. Parenthesise the branch to keep it there, or wrap the whole operation + in a derived table to order or limit the result. `UNION ALL` is unchanged: its `ORDER BY` / `LIMIT` have + always applied per branch. +- **A set operation across catalogs** — mixing branches with catalog-qualified names (`` `cluster_b`.orders ``). + Catalogs are resolved by their position in the SQL text, so a branch could run on the wrong cluster; the + planner refuses rather than risk it. Run each branch as its own statement, or drop the catalog prefix. +- **`CORRESPONDING` / `CORRESPONDING BY`** — SQL's opt-in for name-based matching. Not implemented; + positional matching is the only mode. + +### Licensing + +A set operation costs **nothing** against your plan's `maxJoins` allowance — like a derived table, it is the +JOINs and correlated subqueries *inside* the branches that count, at any nesting depth. ## Quoted identifiers — residual limits Quoted column names, aliases and **table names** work in both spellings — see [Quoted identifiers](dql_statements.md#quoted-identifiers) and -[Qualified and quoted table names](dql_statements.md#qualified-and-quoted-table-names). Five things +[Qualified and quoted table names](dql_statements.md#qualified-and-quoted-table-names). Six things they do **not** cover yet: - **`INSERT`, `UPDATE`, `CREATE`, `DROP` and `ALTER` names are not quotable.** @@ -107,10 +310,10 @@ they do **not** cover yet: `FROM "logs-2025.03"` — or leave it bare (`FROM logs-2025.03`). All three read the index `logs-2025.03`. -- **A qualifier must be quoted from the FIRST part.** `FROM elastic."bi_events"` mixes the +- **A qualifier must be quoted from the FIRST part.** `FROM prod_eu."bi_events"` mixes the spellings, so the leading run of quoted parts is empty and the whole operand is read as ONE index - name, `elastic.bi_events`. Quote the first part too (`FROM "elastic"."bi_events"`) if you meant - `elastic` as a qualifier, or leave both bare if you meant the dotted index name. + name, `prod_eu.bi_events`. Quote the first part too (`FROM "prod_eu"."bi_events"`) if you meant + `prod_eu` as a qualifier, or leave both bare if you meant the dotted index name. - **A dot inside a quoted COLUMN name is still a qualifier.** `` SELECT `a.b` FROM t `` is read as the column `b` qualified by `a`, exactly as `SELECT a.b` is — there is no way to address an @@ -121,7 +324,7 @@ they do **not** cover yet: qualified name; `SELECT a . b` is rejected, and so is a name left with a trailing dot (`ORDER BY b. DESC`). This is deliberate: when the dot was allowed to float, `ORDER BY b. DESC` silently parsed as a column named `b.DESC` sorted *ascending*. A **table**-name qualifier is - deliberately more tolerant (`FROM "elastic" . bi_events` is accepted), because that spelling has + deliberately more tolerant (`FROM "prod_eu" . bi_events` is accepted), because that spelling has always been accepted there and tightening it would have moved which index the statement reads. - **A qualifier shares a namespace with a real dotted index name.** When one `FROM` names the same @@ -139,6 +342,61 @@ they do **not** cover yet: > name itself unquoted** (`` `prod_us`.orders ``) until this is fixed — see > [joins.md](joins.md#row-2--cross-cluster-conveyor). +## Not yet supported + +- **Recursive CTEs** (`WITH RECURSIVE …`) and **CTE column lists** (`WITH a (x, y) AS …`), both refused by name. Plain non-recursive CTEs work since engine `0.24.0`, with two further limits: a `WITH` clause is accepted only at the top of a `SELECT` (not inside a subquery body, CTAS, `INSERT … SELECT` or a materialized view), and a CTE body may not name the CTE itself — unlike PostgreSQL, which binds such a name to the base table, this engine rejects it. +- **Positional / tiling window functions**: `NTILE`, `LAG`, `LEAD` — not yet implemented. (Note: `PERCENTILE_CONT` / `PERCENTILE_DISC` — percentile *aggregates* — already work; the positional/tiling window functions are a different family.) +- **`SELECT TOP n PERCENT` and `SELECT TOP n WITH TIES`**, both refused *by name*. `TOP n` itself + works and is a spelling of `LIMIT n`. Take the plain row count, or compute the percentage + yourself. ⚠️ Because `PERCENT` is recognised in that position, a column of that name cannot be the + sole select item directly after `TOP n` — write `SELECT TOP 5 t.percent FROM t AS t`, or quote it. +- **`SELECT DISTINCT TOP n`** — rejected. Write `SELECT DISTINCT … LIMIT n`. (The reverse order, + `SELECT TOP n DISTINCT`, happens to parse, but it is not valid T-SQL and is not a supported + spelling.) +- **The ODBC/JDBC escape sequences** — `{fn …}`, `{d '…'}`, `{ts '…'}`, `{oj …}`, `{escape '…'}`. + A tool that emits `{fn TIMESTAMPADD(SQL_TSI_DAY, -89, CURRENT_DATE)}` is refused, even though the + `TIMESTAMPADD(…)` inside it is now accepted on its own. Turn escape processing off in the client, + or write the call without the braces. +- **MySQL's null-safe equality operator `<=>`** (`a <=> b`, i.e. `a = b OR (a IS NULL AND b IS NULL)`). A BI tool set to a MySQL dialect can emit it in a `JOIN … ON`. Write the expansion, or `=` when neither side is nullable. + +When they arrive they will be a driver-side enhancement — single-cluster customers get them by upgrading the driver (JDBC / ADBC / sidecar), with no infrastructure change and no federation server required. + +### What a not-yet-supported query looks like + +A **recursive** CTE is rejected by the parser today, by name: + +```sql +-- Not supported: WITH RECURSIVE is refused — only non-recursive CTEs are accepted. +WITH RECURSIVE subordinates AS ( + SELECT id, manager_id FROM employees WHERE id = 1 + UNION ALL + SELECT e.id, e.manager_id FROM employees e JOIN subordinates s ON e.manager_id = s.id +) +SELECT id FROM subordinates; +``` + +There is no rewrite that recovers arbitrary-depth recursion. Flatten the hierarchy at index time (store a +path or a level on each document), or run one statement per level. + +The **non-recursive** CTE and the set operator below, on the other hand, both run since engine `0.24.0` — +a CTE reference is a derived table, so each executes on the relational engine and carries the same venue +requirement: + +```sql +WITH eu_departments AS (SELECT id FROM departments WHERE region = 'EU') +SELECT e.name FROM employees e JOIN eu_departments d ON e.department_id = d.id; + +SELECT customer_id FROM orders_q1 +INTERSECT +SELECT customer_id FROM orders_q2; +``` + +> ⚠️ **A CTE cannot be named inside a `WHERE` subquery body.** +> `WITH eu AS (…) SELECT name FROM employees WHERE department_id IN (SELECT id FROM eu)` parses, but +> the reference inside the body is not resolved: the body reaches Elasticsearch asking for an index +> called `eu`, and the statement fails with a `404 index_not_found_exception` naming it. Loud, never +> silent. Read the CTE in `FROM` or `JOIN`, as above. + ## Temporary tables are not supported `CREATE TEMPORARY TABLE` and `CREATE [LOCAL | GLOBAL] TEMPORARY TABLE`, with or without @@ -189,7 +447,7 @@ permanent. See [STDDEV / VARIANCE family](functions_aggregate.md#function-stddev ## Coming in the upcoming release (Quarter 1 2027) - **Heterogeneous federation**: JOIN or correlate Elasticsearch with PostgreSQL, MySQL, ClickHouse, Snowflake, and more — plus cross-cluster subqueries (e.g. correlate one cluster's data against another's). - **Not this**: correlating one Elasticsearch index against **another Elasticsearch index** — `EXISTS` / `NOT EXISTS` / `IN` / `NOT IN` / a scalar comparison against a subquery that reads the outer row — is **single-cluster** and runs through the relational engine shipped in `softclient4es-arrow-extensions`. Its one rule: the outer reference must be **qualified** with the outer table's alias (`… WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = c.id)`), because a bare column name inside a subquery is read as the subquery's own column. A venue without that jar refuses the statement with HTTP 400 rather than executing it as if it were self-contained. + **Not this**: correlating one Elasticsearch index against **another Elasticsearch index** already works in this release and is single-cluster — see [Subqueries and derived tables](#subqueries-and-derived-tables) above. What lands here is correlating across *heterogeneous* sources and across *clusters*. ## Deferred (a future release, demand-driven — tell us what you need) @@ -197,7 +455,7 @@ permanent. See [STDDEV / VARIANCE family](functions_aggregate.md#function-stddev ## Roadmap timing -We do not commit firm external dates. The next release is targeted for **Quarter 4 2026**; the upcoming release (heterogeneous federation) for **Quarter 1 2027**; the deferred items are demand-driven with no committed date. Treat the next release's feature list as *planned*, not guaranteed — its scope is gated on a function-library audit. +We do not commit firm external dates. The **Not yet supported** list above carries no target release: those items are planned, not scheduled. The upcoming heterogeneous-federation release is targeted for **Quarter 1 2027**; the deferred items are demand-driven with no committed date. ## See also @@ -206,4 +464,4 @@ We do not commit firm external dates. The next release is targeted for **Quarter --- -*This page describes SoftClient4ES **as of the current release**. Once the next release ships, the "Not in this release" list above shrinks — verify against your installed release.* +*This page describes SoftClient4ES **as of engine `0.24.0`**. Availability lines name the release a feature landed in; the **Not yet supported** list shrinks as items ship — verify against your installed release.* diff --git a/documentation/sql/materialized_views.md b/documentation/sql/materialized_views.md index 37a73d14..dbdb24df 100644 --- a/documentation/sql/materialized_views.md +++ b/documentation/sql/materialized_views.md @@ -239,6 +239,8 @@ LIMIT 100; When a `GROUP BY` clause is present, the engine generates an additional **pivot transform** for the aggregation step. +A `HAVING` in a view is more constrained than a `HAVING` in a `SELECT` — most importantly, every aggregate it reads must be published in the `SELECT` list **with an alias**, as `SUM(o.amount) AS total_amount` is above. See [HAVING in a materialized view](#having-in-a-materialized-view). + --- ## DROP MATERIALIZED VIEW @@ -531,6 +533,7 @@ DROP MATERIALIZED VIEW IF EXISTS orders_with_customers_mv; | Limitation | Details | |---------------------------------------------|----------------------------------------------------------------------| +| **`HAVING`** | More constrained than in a `SELECT`: every aggregate the clause reads must be published in the `SELECT` list with an alias, and it cannot filter on a grouping key (see below) | | **UNNEST JOIN** | Not supported in materialized views | | **`RIGHT JOIN` / `FULL OUTER JOIN`** | Not supported (see below). Use `LEFT JOIN` with swapped table order. | | **Quota limits** | Community: 1 view · Pro: 50 · Enterprise: unlimited | @@ -538,6 +541,72 @@ DROP MATERIALIZED VIEW IF EXISTS orders_with_customers_mv; | **Eventual consistency** | Data is eventually consistent based on refresh frequency and delay | | **Join cardinality** | JOINs use enrich policies which match on a single field | +### HAVING in a materialized view + +Since engine `0.24.0`, `CREATE MATERIALIZED VIEW` **refuses** a `HAVING` that the Elasticsearch transform behind the view cannot express, naming the clause and the remedy. The refusal is a parse-time rejection (HTTP 400), so it reaches every venue — REPL, JDBC, Flight SQL — and no artifact is deployed. + +The same `HAVING` usually stays valid in a plain `SELECT`. This is a limit of the **transform** a view deploys, not of the query engine, and not of Elasticsearch. + +**The case you are most likely to meet: an aggregate with no `SELECT` alias.** + +```sql +-- REFUSED: SUM(amount) carries no alias +CREATE MATERIALIZED VIEW sales_by_city_mv +AS +SELECT city, SUM(amount) +FROM orders +GROUP BY city +HAVING SUM(amount) > 5; + +-- Accepted: the same view, with the aggregate published under an alias +CREATE MATERIALIZED VIEW sales_by_city_mv +AS +SELECT city, SUM(amount) AS total +FROM orders +GROUP BY city +HAVING SUM(amount) > 5; +``` + +A view's transform names each aggregation after its `SELECT` alias, so an unaliased aggregate leaves the group filter pointing at a metric the view never creates. Adding `AS total` is the whole fix — the `HAVING` itself does not change. + +#### Why a view differs from a search + +The engine has five mechanisms for a `HAVING`; a transform's pivot offers one of them. + +| The `HAVING` names | A search applies it through | A view's pivot | +|--------------------------------------------------|------------------------------------------------------|---------------------------------------------------------------------------------| +| a grouping key — `HAVING city = 'Paris'` | the `terms` aggregation's `include` / `exclude` list | **no such channel** — `group_by` emits `{"terms":{"field":…}}` and nothing else | +| arithmetic over aggregates — `HAVING spread > 3` | a `bucket_script` | **no `bucket_script` channel** | +| a metric — `HAVING SUM(amount) > 5` | a `bucket_selector` | a `bucket_selector` — **the one that works** | + +A pivot also computes a narrower set of metrics than a search: only `MIN`, `MAX`, `SUM`, `AVG` and `COUNT` (including `COUNT(DISTINCT …)`). + +#### What is refused + +| A view's `HAVING` that… | Why | Remedy | +|-----------------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------| +| reads an aggregate with **no `SELECT` alias** — `SELECT city, SUM(amount) … HAVING SUM(amount) > 5` | the pivot names each aggregation after its `SELECT` alias, so the condition matches none of them | publish the aggregate with an alias: `SUM(amount) AS total` | +| filters on a **grouping key** — `HAVING city = 'Paris'` | a transform's `group_by` has no `include` / `exclude` channel | move the condition into the view's `WHERE` | +| has **no `GROUP BY`** | no `GROUP BY` means no pivot, and a `bucket_selector` hangs on the pivot | add a `GROUP BY`, or materialize the aggregate and filter when querying the view | +| names an aggregate **a pivot cannot compute** — `HAVING STDDEV(amount) > 1` | a pivot computes only `MIN`, `MAX`, `SUM`, `AVG` and `COUNT`, so this aggregate exists in no view | filter when querying the view | +| names an **expression over aggregates** — `SELECT MAX(amount) - MIN(amount) AS spread … HAVING spread > 3` | a search evaluates it with a `bucket_script`, which a pivot has no channel for | filter when querying the view | +| compares an aggregate with an aggregate **no other condition names on its left** — `HAVING MAX(amount) > MIN(amount)` | only the left-hand metric of each comparison is declared in `buckets_path`; the right-hand one is read as an undeclared parameter, the null guard rejects every group and the view comes out **empty** | compare with a constant, or filter when querying the view | +| names an aggregate the **`SELECT` list does not publish** — `SELECT city, COUNT(*) AS n … HAVING MAX(amount) > 100` | the filter would read a metric the view does not compute | add `MAX(amount) AS max_amount` to the `SELECT` list | +| contains a **nested, child or parent predicate** | the pivot would emit the metric condition alone — a partial filter returning the wrong groups at HTTP 200 | filter when querying the view | + +Comparing two aggregates is **not** refused as such: `HAVING SUM(amount) > MAX(amount) AND MAX(amount) > 1` is accepted, because the sibling condition is what makes the pivot declare `MAX(amount)`. + +#### What still works + +- A metric compared with a constant, when the aggregate is published under a `SELECT` alias — `SELECT city, SUM(amount) AS total … HAVING SUM(amount) > 5`. The alias form, `HAVING total > 5`, is equally accepted. +- A JOIN view with a metric-only `HAVING` — the [aggregation example](#materialized-view-with-aggregations) above. +- `COUNT(DISTINCT …)` — `SELECT city, COUNT(DISTINCT customer_id) AS customers … HAVING COUNT(DISTINCT customer_id) > 1`. +- A function grouping key — `SELECT DATE_TRUNC(order_date, MONTH) AS month, SUM(amount) AS total … GROUP BY month HAVING SUM(amount) > 100`. + +#### What changed + +Before engine `0.24.0` every refused statement above was **accepted**, and the view then materialized the wrong groups at HTTP 200: the condition was silently dropped and **every** group was materialized, or — in the two-aggregate case — the deployed `bucket_selector` read an undeclared parameter, rejected every group, and the view came out **empty**. A view bakes that missing filter into a stored transform which then feeds dashboards, which is why the refusal is preferable. A view already deployed keeps running as it is — the rules apply when a view is created, and `CREATE OR REPLACE` is validated exactly like `CREATE`. + ### Supported JOIN types Only `INNER JOIN` and `LEFT JOIN` (`LEFT OUTER JOIN`) are supported for materialized views. diff --git a/documentation/sql/operator_precedence.md b/documentation/sql/operator_precedence.md index 6132b34b..fe5f64d6 100644 --- a/documentation/sql/operator_precedence.md +++ b/documentation/sql/operator_precedence.md @@ -17,6 +17,7 @@ This page lists operator precedence used by the parser and evaluator. Operators | **5** | `<`, `<=`, `>`, `>=` | Comparison | Less than, less or equal, greater than, greater or equal | | **6** | `=`, `!=`, `<>` | Equality | Equal, not equal | | **7** | `BETWEEN`, `IN`, `LIKE`, `RLIKE` | Membership & Pattern | Range, set membership, pattern matching | +| **7** | `EXISTS` | Existential | True when the subquery returns at least one row | | **8** | `AND` | Logical AND | Logical conjunction | | **9** (Lowest) | `OR` | Logical OR | Logical disjunction | @@ -376,6 +377,15 @@ WHERE price > 50; **Seventh precedence** - Range, set membership, and pattern matching. +> **A subquery does not change the level of its operator.** `IN (SELECT …)` binds exactly where +> `IN (1, 2, 3)` binds; `x > (SELECT AVG(…) …)` and the quantified forms `x > ALL (SELECT …)` / +> `x >= ANY (SELECT …)` bind at the comparison level (5), and `= ANY | SOME` / `<> ALL` at the equality +> level (6) — where they are normalised to `IN` / `NOT IN`. `EXISTS (SELECT …)` is a unary predicate +> taking no left operand, so nothing binds to its left. All of them are leaf predicates: they group +> ahead of `AND` and `OR`, which is why `a = 1 AND b IN (SELECT id FROM u) OR c = 2` reads as +> `((a = 1) AND (b IN (SELECT id FROM u))) OR (c = 2)` — the same shape it would have with a literal +> value list. See [Subqueries and derived tables](known_limitations.md#subqueries-and-derived-tables). + **Examples:** **BETWEEN:** @@ -793,6 +803,7 @@ WHERE price BETWEEN 10 AND 100; | 5 | `<`, `<=`, `>`, `>=` | Left | `a < b` | | 6 | `=`, `!=`, `<>` | Left | `a = b` | | 7 | `BETWEEN`, `IN`, `LIKE`, `RLIKE` | N/A | `a BETWEEN 1 AND 10` | +| 7 | `EXISTS` | N/A | `EXISTS (SELECT 1 FROM t)` | | 8 | `AND` | Left | `a AND b AND c` | | 9 | `OR` | Left | `a OR b OR c` | diff --git a/documentation/sql/operators.md b/documentation/sql/operators.md index d5c6bebb..805f039a 100644 --- a/documentation/sql/operators.md +++ b/documentation/sql/operators.md @@ -577,7 +577,7 @@ WHERE event_time >= '2025-01-10 09:00:00' ### Operator: `IN` **Description:** -Membership in a set of literal or numeric values, or results of subquery (subquery support depends on implementation). +Membership in a set of literal or numeric values, or in the values a subquery returns. **Syntax:** ```sql @@ -636,24 +636,24 @@ WHERE category_id IN ( ); ``` -**Empty List:** -```sql --- Empty IN list returns false -SELECT * FROM products WHERE id IN (); --- Returns no rows -``` +Subqueries are accepted **since engine `0.24.0`**. The subquery must project **exactly one column** (`IN (SELECT * FROM …)` is refused), must read a table +(`IN (SELECT 1)` is refused) and may not be a set operation. An uncorrelated body is executed first and its +values are collected as a distinct set — bounded at **65,536** values (`index.max_terms_count`), past which +the statement fails loudly rather than truncating. A body over a plain column is resolved with a single +`terms` aggregation, so `DISTINCT` inside it buys nothing. A **correlated** body (one that reads the outer +row) runs on the relational engine — see +[Subqueries and derived tables](known_limitations.md#subqueries-and-derived-tables). **NULL Handling:** ```sql --- NULL in list -SELECT * FROM users WHERE status IN ('active', NULL); --- NULL is ignored in the list - -- Column with NULL SELECT * FROM users WHERE email IN ('test@example.com'); -- Rows with NULL email are not matched ``` +A `NULL` among the values a subquery returns is ignored by `IN` (ANSI). A literal `NULL` in a written value +list, and an empty list `IN ()`, are **parse errors** — write the condition with `IS NULL` instead. + --- ### Operator: `NOT IN` @@ -692,9 +692,10 @@ SELECT * FROM products WHERE product_id NOT IN (1, 2, 3); **With Subquery:** ```sql -- Exclude customers who have orders +-- (no DISTINCT needed — the engine collects the subquery's values as a set) SELECT * FROM customers WHERE id NOT IN ( - SELECT DISTINCT customer_id FROM orders + SELECT customer_id FROM orders ); -- Exclude inactive categories @@ -706,11 +707,8 @@ WHERE category_id NOT IN ( **NULL Handling (Important!):** ```sql --- NOT IN with NULL in list returns NULL (not true!) -SELECT * FROM users WHERE id NOT IN (1, 2, NULL); --- Returns no rows because comparison with NULL is NULL - --- Safe alternative: filter NULLs in subquery +-- NOT IN over a set that contains a NULL matches NO rows (ANSI: the comparison is UNKNOWN) +-- Filter the NULLs out in the subquery: SELECT * FROM customers WHERE id NOT IN ( SELECT customer_id FROM orders WHERE customer_id IS NOT NULL @@ -723,6 +721,12 @@ WHERE NOT EXISTS ( ); ``` +That last form is a **correlated** subquery — the body reads `c.id` from the outer row — so it runs on the +relational engine rather than on Elasticsearch alone (**since engine `0.24.0` with arrow-extensions +`0.3.4`**; the uncorrelated forms above need only the engine). The outer reference must be qualified with the outer +alias and left unquoted. See +[Which forms need the relational engine](known_limitations.md#which-forms-need-the-relational-engine). + --- ### Operator: `BETWEEN ... AND ...` @@ -1454,8 +1458,9 @@ WHERE first_name = 'John' OR last_name = 'Doe'; -- May require full table scan -- Alternative: UNION ALL (if indexes exist) --- Note: bare UNION (with de-duplication) is not supported and is rejected at --- parse time — a row matching BOTH predicates appears twice with UNION ALL. +-- UNION ALL keeps duplicates, so a row matching BOTH predicates appears twice — +-- the second branch excludes them here. A bare UNION would de-duplicate for you, +-- but it runs on the relational engine rather than on Elasticsearch directly. SELECT * FROM users WHERE first_name = 'John' UNION ALL SELECT * FROM users WHERE last_name = 'Doe' AND first_name <> 'John';