Skip to content

docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it - #346

Draft
fupelaqu wants to merge 13 commits into
mainfrom
docs/22-subqueries
Draft

fupelaqu wants to merge 13 commits into
mainfrom
docs/22-subqueries

Conversation

@fupelaqu

@fupelaqu fupelaqu commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Closes #345

The published documentation claimed subqueries and derived tables were not supported. They ship in 0.24.0, so every page saying otherwise was telling users a feature they have does not exist — most expensively the BI-tool section, which instructed them to rewrite tool-generated nested SQL as an explicit JOIN.

Sibling PR: SOFTNETWORK-APP/softclient4es-web#62 — the two carry the same capabilities, the same caveats and the same venue distinction, and are intended to merge together, during the 0.24.0 train (not before: these features are not in 0.23.0).

Version framing

The features are stamped **Since engine \0.24.0`**, and where the relational engine is required, the full house pairing **Since engine `0.24.0` with arrow-extensions `0.3.4`**, following the convention already on these pages (Since engine `0.23.0` with arrow-extensions `0.3.3`, Since `0.20.4`) — bold or code-spanned version, no date, no -SNAPSHOT. A reader on 0.23.0` can now tell from the page that their build rejects these forms.

The stamp is attached only to what is new: subqueries in WHERE, derived tables, correlated subqueries. Cross-index JOINs, UNION ALL and every other pre-existing capability carry no version line — one there would be a new false claim.

The unsupported list is deliberately date-free and version-free: the heading is Not yet supported, the roadmap-timing paragraph says that list is planned rather than scheduled, and the per-page one-liners carry no target. CTEs and set operators have not been started — both landed in this train after this PR was opened; see the two update sections at the end.

The softclient4es-arrow-extensions jar is named with its version — 0.3.4 — where it is genuinely required. That version is confirmed as the number this train will cut (arrow main is on 0.3.4-SNAPSHOT), but it has not been released yet.

🔴 Merge condition: these pages name 0.3.4, so merging them depends on arrow actually cutting 0.3.4 in this train. If the number changes, both pages need a correcting pass before they go live.

The pairing is applied only where the jar is required — derived tables and correlated subqueries. Uncorrelated WHERE subqueries execute ES-natively at every venue including a plain REPL, so they carry the engine version alone; naming an extensions version there would tell a plain-REPL user they need a jar they do not need.

Now documented as working

Form Runs where Needs the relational engine
Uncorrelated WHERE subquery — IN / NOT IN, EXISTS / NOT EXISTS, scalar comparison, quantified (= ANY|SOME, <> ALL, > ALL, >= ANY, < ALL, …) Elasticsearch, two-phase No — every venue, plain REPL included
Correlated WHERE subquery Relational engine Yes
Derived table in FROM / JOIN, nested to any depth Relational engine Yes

Plus the bounds: the two-phase path caps at 65,536 distinct values (index.max_terms_count) and fails loudly past it, never truncating; IN ignores inner NULLs while NOT IN over a NULL-bearing set matches no rows; a correlated subquery costs one maxJoins unit and a derived table costs none.

Still documented as unsupported

CTEs (WITH …); set operators beyond UNION ALL (UNION, INTERSECT, EXCEPT); LATERAL; any subquery in HAVING; a subquery in the SELECT list; a UNION ALL or FROM-less subquery body; a body projecting more than one column; an unqualified or quoted outer reference; and a correlated body carrying its own JOIN, comma-separated FROM, JOIN UNNEST, derived table or window function.

How the claims were verified

Not from a spec. A throwaway sql/src/test/…/DocSweepProbeSpec (deleted before committing) ran every statement through Parser.apply and printed relationalClosureRequired. All 12 subquery examples the docs publish parse; the WITH example is the deliberate rejection. The venue table is read from SingleSearch.relationalClosureRequired / whereSubqueries; the refusal wording from RelationalClosureGuard; the 65,536 bound from SubqueryResolver.MaxTerms; the correlated-body restrictions and the maxJoins accounting from the merged relational-engine planner.

These docs are plain Markdown with no build step, so headerCheck / scalafmtCheck prove nothing about them — the parse probe is the gate. No Scala source changed.

Two corrections found while certifying the pages, fixed here rather than filed

  • Tableau's quoted, fully-qualified identifiers were documented as "not accepted yet". Both `bi_events`.`category` and "bi_events"."category" parse, and so does Tableau's derived-table wrapper. Corrected in bi_tools.md. Tableau's tier is unchanged — still Compatible.
  • operators.md published three IN examples that do not parse or are pointless: IN () and IN ('active', NULL) are parse errors, and DISTINCT inside an IN body buys nothing (the values are collected as a set, and a plain-column body is resolved with one terms aggregation).

Keyword pages (keywords.md, operator_precedence.md)

keywords.md described itself as "a list of reserved words recognized by the parser" and omitted every word story 22.2 added. A comparison against Parser.reservedKeywords shows it was never a reserved-word list and diverges in both directions: 26 reserved words absent (EXISTS, ALL, UNION, EXCEPT, MATCH, TRUE/FALSE, PI …) and 26 listed words recognised but not reserved (CEILING, UCASE, TRY_CAST, RLIKE, OVER, POINT …). It is a grammar-keyword list wearing a reserved-word title.

🔴 Story 22.2 made that distinction load-bearing, which is why this is not "append two words". EXISTS and ALL are reserved; ANY and SOME were deliberately left unreserved so a column named any keeps parsing — while x = ANY (SELECT …) is still real grammar. Listing all four together would publish a restriction that does not exist and send users renaming columns needlessly.

Resolution: the page now states both sets and the rule rather than reclassifying 52 entries (a separate audit). The four new entries carry the distinction inline, and the remedy for a real collision is quoting, not renaming. UNION ALL, absent from the page entirely, is added.

Measured, not assumed — every claim probed through Parser.apply:

statement verdict
SELECT any, some FROM t WHERE any = 1 parses (columns)
SELECT exists FROM t / SELECT all FROM t rejected
SELECT "exists" FROM t / SELECT `exists` FROM t parse; both render "exists"
WHERE customer_id = ANY (SELECT …) renders as IN (SELECT …)
WHERE amount > ALL (SELECT …) keeps its spelling

operator_precedence.md gains EXISTS as a level-7 unary predicate and a note that a subquery does not change the level of its operator. That note's grouping claim was verified on the AST, not by eye: a = 1 AND b IN (SELECT id FROM u) OR c = 2 parses as ((GenericExpression AND InSubquery) OR GenericExpression) — the same shape as the literal-list form. The page already carried two IN (SELECT …) examples that only became valid with 22.2.

Cross-surface consistency — no code defect. The core SQLKeywords registry lists EXISTS, ANY, SOME, ALL in clauseTokens, and the REPL completer derives its words from SQLKeywords.highlightedWords via ReplKeywords.all, so all four already tab-complete. The registry and the completer agreed with the parser; only the docs page was stale. known_limitations makes no reserved-word claim, so nothing there is contradicted.

Files

README.md (DQL section + roadmap — an omission, not a false claim: it never mentioned subqueries), documentation/sql/{known_limitations,dql_statements,joins,operators,keywords,operator_precedence}.md, documentation/client/{bi_tools,jdbc,adbc_driver,arrow_flight_sql}.md.

Deliberate twin divergence

known_limitations.md keeps its elasticsql-only "Quoted identifiers — residual limits" section, and operators.md, keywords.md and operator_precedence.md have no web twin — the site has no keywords or precedence page at all. All of that pre-dates this PR, so the keyword work here is deliberately elasticsql-only and PR #62 carries none of it; creating a site keywords page is a content decision, not a correction. The subquery sections themselves are identical across the two repos apart from link syntax.


Update 2 — set operators (story 22.6) and positional column matching

Story 22.6 shipped UNION / UNION DISTINCT / INTERSECT [ALL] / EXCEPT [ALL], and the follow-up fix (#357, closing #354 and #355) changed how branch columns are matched. Ten places in this PR still said "set operators beyond UNION ALL are not yet supported", so they would have repeated — one story later — exactly the defect this PR exists to correct, as the CTE commit before it already had to.

What changed

  • The UNION ALL section becomes Set operators: the seven spellings with where each runs, the compatibility rules, precedence and the grouping refusal, the ORDER BY / LIMIT split, and the execution model. UNION ALL keeps its own row — it is the one Elasticsearch answers by itself with a single _msearch, at every venue; everything else needs arrow-extensions 0.3.4, exactly like a derived table.
  • The venue table gains the two rows a reader actually looks for: a CTE reference (relational engine — it is a derived table) and the UNION ALL vs UNION/INTERSECT/EXCEPT split.
  • Columns match by position, published as a change rather than as a rule, because it is one. Before this release branches were matched by name, so SELECT id AS x … UNION ALL SELECT id AS y … returned a null column for every row — branch 1's own included. A reader with a UNION ALL written against the old behaviour has to check their branch order, and the note says so.
  • keywords.md gains the six new words plus the reserved-vs-recognised line for them: UNION / INTERSECT / EXCEPT / ALL / DISTINCT are reserved, WITH and RECURSIVE are not.

Three claims on these pages that are no longer true, found by probing rather than by reading:

Claim Status
"a UNION ALL cannot be a derived-table body" (so it cannot be wrapped to sort globally) False now — it parses; the line is gone
a subquery body "may not be a UNION ALL" Too narrow — every set-operator spelling is refused there
the "what a not-yet-supported query looks like" example was the de-duplicating UNION Replaced by WITH RECURSIVE, which is genuinely refused and has no rewrite

Verification. Every example that ships — including the ones asserted to fail — was parsed against merged main (058f27e2 + 7187c7d9) by a throwaway probe reporting Parser(stmt) and relationalClosureRequired, deleted afterwards. The venue column in every table is that probe's closure value, not a reading of the spec. The refusals asserted: WITH RECURSIVE, CORRESPONDING, a parenthesised group, a trailing ORDER BY after the last branch, a set operator after a FROM-less select list, and a set operation as a subquery body. The reserved-word claims were checked in the alias position, which is where the distinction bites — SELECT a AS recursive FROM t parses, SELECT a AS intersect FROM t does not.

The far-side refusals were read in arrow, not inferred from core: a catalog-qualified set operation is refused by name in JoinPlanner.plan, because catalogs resolve by text position and a branch would otherwise run on the wrong cluster. That refusal is documented.


Update — the BI dialect spellings from #360 (same 0.24.0 train)

This PR now also documents CHAR_LENGTH / CHARACTER_LENGTH, TIMESTAMPADD, TIMESTAMPDIFF, the ODBC SQL_TSI_* interval names and SELECT TOP n, plus the new refusals (TOP n PERCENT, TOP n WITH TIES, SELECT DISTINCT TOP n, and the ODBC {fn …} escape family, which is still refused even though the call inside it now parses).

🔴 And one correction that has nothing to do with the new spellings. functions_date_time.md said DATEDIFF returns "(date1 - date2)" and published eight examples built on that reading. The engine returns date2 - date1. Measured on real Elasticsearch 8.18:

DATE_DIFF('2025-01-10', '2025-01-01', DAY)  =>  -9     (the page claimed 9)
DATE_DIFF('2025-01-01', '2025-01-10', DAY)  =>   9

It surfaced because TIMESTAMPDIFF had to go in that same block and MySQL defines TIMESTAMPDIFF(unit, dt1, dt2) as dt2 - dt1 — the page and the alias could not both be right. The merged pin for DATE_DIFF(birthdate, CURRENT_DATE, YEAR) settles which: it emits ChronoUnit.YEARS.between(birthdate, now), which is what makes it an age. Every example in the block has been re-executed and shows its real result, with the reversed-sign case published beside them.

Gate. These docs have no build, so the parse probe is the gate: all 24 SQL examples added or changed were parsed against a parser built from origin/main after #360 merged — 0 rejected, 0 non-round-tripping — and the date/length results were executed against real Elasticsearch 8.18 rather than computed by hand. The probe also showed that DATE_DIFF(unit, a, b) is what the engine renders back, so that spelling is documented too.

Version stamped 0.24.0, checked against build.sbt on main rather than assumed.


⚠️ The DATEDIFF section is accurate today, and may change again — see #363

A correction to how the DATEDIFF change above was first described in this PR.

The page was not sloppy. MySQL 8.4 defines DATEDIFF(expr1, expr2) as expr1 − expr2, and the page — description and all eight examples — described MySQL faithfully. What it did not describe is this engine, which returns expr2 − expr1 because DATEDIFF is an alias word on the DATE_DIFF token and inherits BigQuery's DATE_DIFF(start, end, unit) order.

Run against MySQL's own two documented examples, on real Elasticsearch 8.18:

statement MySQL 8.4 this engine
DATEDIFF('2007-12-31','2007-12-30') 1 -1
DATEDIFF('2010-11-30','2010-12-31') -31 31

So the real defect is arguably in the engine, not the page: a MySQL-named function that contradicts MySQL, silently. That is now filed as #363. Note we already match MySQL on TIMESTAMPDIFF (dt2 − dt1), so DATEDIFF is the single odd one out rather than a house convention.

What this PR does: documents the engine as it behaves today, which is the accurate thing to publish right now. If #363 is fixed, this section changes again — that is expected, and is why the note is here rather than in a commit message.

…y it

The published documentation said subqueries, derived tables and correlated
subqueries were "not in this release". They ship in 0.24.0, so every page
carrying that claim was steering users away from a feature they have — most
expensively the BI-tool section, which told them to rewrite tool-generated
nested SQL as an explicit JOIN.

Every claim here was measured against the merged engine, not a spec: each
published example was run through Parser.apply, and the venue split was read
from relationalClosureRequired.

- known_limitations.md gains a "Subqueries and derived tables" reference
  section: the accepted forms, the venue table (an UNCORRELATED WHERE
  subquery runs ES-natively at every venue including a plain REPL; a
  CORRELATED one and a derived table need the relational engine), the
  65,536-value bound on the two-phase path, the ANSI NULL rules, the forms
  still refused by name, and the maxJoins accounting.
- The false worked example is replaced by a CTE — which really is rejected —
  and its accepted rewrite.
- README gains the capability in the DQL section and the roadmap, leading
  with the form that works everywhere.
- CTEs, set operators beyond UNION ALL, LATERAL, a subquery in HAVING, a
  subquery in the SELECT list, a UNION ALL or FROM-less body, a body
  projecting more than one column, and a correlated body carrying its own
  JOIN / derived table / window function stay documented as unsupported.

Also corrected while certifying the pages: Tableau's quoted, fully-qualified
identifiers and its derived-table wrapper do parse (both spellings measured),
and operators.md published three IN examples that do not parse or are
pointless (IN (), IN ('active', NULL), a DISTINCT inside an IN body).

Closes #345
…ne 0.24.0

A reader on 0.23.0 could not tell from the pages whether their build has
these forms; they would try one and get a parse error, which is the failure
the pages exist to prevent.

Follows the house convention already on these pages ("Since engine `0.23.0`",
"Since `0.20.4`") — bold or code-spanned version, no date, no -SNAPSHOT.

The stamp is attached ONLY to what is new in 0.24.0: subqueries in WHERE,
derived tables, correlated subqueries. Cross-index JOINs, UNION ALL and every
other pre-existing capability keep no version line, since one there would be a
new false claim.

The unsupported list is deliberately date-free and version-free. CTEs and set
operators have not been started, so "coming in the next release (Quarter 4
2026)" is removed from the prose: the heading becomes "Not yet supported", the
roadmap-timing paragraph says the list is planned rather than scheduled, and
the per-page one-liners say "not yet supported" with no target. A published
target becomes a broken promise the moment priorities move.

Closes #345
Applies the full house pairing ("Since engine `0.23.0` with arrow-extensions
`0.3.3`" is the existing form on joins.md) at the six places that state the
relational-engine requirement.

The pairing is attached ONLY where the jar is genuinely required — derived
tables and correlated subqueries. Uncorrelated WHERE subqueries execute
ES-natively at every venue including a plain REPL, so they carry the engine
version alone. Naming an extensions version there would tell a plain-REPL user
they need a jar they do not need, which is the inverse of the over-claim this
sweep exists to remove.

The venue table makes the split explicit per row: the uncorrelated row reads
"No — works at every venue, including a plain REPL with no extensions"; only
the correlated and derived-table rows read "Yes — arrow-extensions `0.3.4`".

Closes #345
keywords.md called itself "a list of reserved words recognized by the parser".
It is not one, and a comparison against Parser.reservedKeywords shows the page
diverges in BOTH directions: 26 reserved words are absent (EXISTS, ALL, UNION
among them) and 26 listed words are recognised but not reserved (CEILING,
UCASE, TRY_CAST, RLIKE, OVER …). So the page has always been a list of GRAMMAR
keywords wearing a reserved-word title.

Story 22.2 made that difference load-bearing rather than pedantic: EXISTS and
ALL are reserved, while ANY and SOME were deliberately left unreserved so a
column named `any` keeps parsing — even though `x = ANY (SELECT …)` is real
grammar. A reader taking this page as "words I cannot use as a column name" is
right about `exists` and wrong about `any`.

Rather than reclassify 52 entries, the page now states the two sets and the
rule, and the four new entries carry the distinction where it bites. The remedy
for a genuine collision is quoting, not renaming.

Measured, not assumed: `SELECT any, some FROM t WHERE any = 1` parses;
`SELECT exists FROM t` and `SELECT all FROM t` are rejected; `SELECT "exists"`
and the backtick spelling both parse; `= ANY (SELECT …)` is normalised to
`IN (SELECT …)` in the render, while `> ALL` keeps its spelling.

Also adds UNION ALL, which the page omitted entirely.

operator_precedence.md gains EXISTS as a level-7 unary predicate and a note
that a subquery does not change the level of its operator. The grouping claim
in that note was verified on the AST, not by eye:
`a = 1 AND b IN (SELECT id FROM u) OR c = 2` parses as
`((GenericExpression AND InSubquery) OR GenericExpression)` — the same shape as
the literal-list form.

Closes #345
fupelaqu and others added 2 commits September 16, 2026 10:26
… them

Story 22.5 puts non-recursive CTEs in this train, so the fourteen places
this PR was about to publish "CTEs are not yet supported" would have
repeated, one story later, exactly the defect the PR exists to correct.

A CTE reference IS a derived table, so every line reads like the
derived-table lines beside it, venue requirement included: it runs on the
relational engine, which means arrow-extensions 0.3.4.

What the engine actually refuses is stated rather than generalised, because
"CTEs work" would be an over-claim: WITH RECURSIVE and CTE column lists
(WITH a (x, y) AS …) are refused by name; a WITH clause is accepted only at
the top of a SELECT, never inside a subquery body, CTAS, INSERT … SELECT or
a materialized view; and a CTE body may not name the CTE itself — unlike
PostgreSQL, which binds such a name to the base table.

known_limitations.md's "what a not-yet-supported query looks like" example
was the CTE that now works, so it is replaced by a de-duplicating UNION —
still refused — and the CTE is shown as the thing that runs.

Set operators beyond UNION ALL stay listed as unsupported: story 22.6 is in
this train but not yet built, so those lines are corrected when it lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… by position

Story 22.6 puts UNION / UNION DISTINCT / INTERSECT [ALL] / EXCEPT [ALL] in
this train, so the ten places this PR was about to publish "set operators
beyond UNION ALL are not supported" would have repeated, one story later,
exactly the defect the PR exists to correct — as the CTE commit before it did.

dql_statements.md's "UNION ALL" section becomes "Set operators": the seven
spellings with where each one runs, the compatibility rules, precedence and
the grouping refusal, the ORDER BY / LIMIT split, and the execution model.
UNION ALL keeps its own row and its own paragraph because it is the one the
engine answers with a single _msearch, at every venue, with no relational
engine in the picture — everything else needs arrow-extensions 0.3.4, exactly
like a derived table.

Columns match BY POSITION, and that is documented as a CHANGE rather than as
a rule, because it is one: before this release branches were matched by name,
so SELECT id AS x ... UNION ALL SELECT id AS y ... returned {x -> null,
y -> ...} for every row, branch 1's own included. A reader with a UNION ALL
written against the old behaviour needs to check their branch order, so the
note says so.

Three statements the pages made that are no longer true, found by probing
rather than by reading: a UNION ALL CAN now be a derived-table body (so the
"wrapping it to sort globally is not supported either" line went), a subquery
body "may not be a UNION ALL" is now "may not be a set operation" (every
spelling is refused there, not just the one), and known_limitations.md's
"what a not-yet-supported query looks like" example WAS the de-duplicating
UNION — replaced by WITH RECURSIVE, which is genuinely refused and has no
rewrite.

keywords.md gains the six new words and the reserved-vs-recognised line for
them: UNION / INTERSECT / EXCEPT / ALL / DISTINCT are reserved, WITH and
RECURSIVE are not. Verified in the ALIAS position, which is where the
distinction bites: SELECT a AS recursive FROM t parses, SELECT a AS intersect
FROM t does not, and quoting rescues either.

Every example that ships was parsed against merged main (058f27e + 7187c7d)
through a throwaway probe reporting Parser(stmt) and relationalClosureRequired,
deleted afterwards — including the ones asserted to FAIL: WITH RECURSIVE,
CORRESPONDING, a parenthesised group, a trailing ORDER BY, a set operator
after a FROM-less select list, and a set operation as a subquery body. The
venue column is that probe's closure value, not a reading of the spec.

Refs #354

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@fupelaqu fupelaqu changed the title docs: subqueries and derived tables ship in 0.24.0 — correct the pages that deny it docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it Sep 17, 2026
fupelaqu and others added 6 commits September 17, 2026 13:07
The "Which forms need the relational engine" table listed the uncorrelated
subquery, the correlated subquery and the derived table — the three forms
story 22.2/22.4 added — and stopped there. A reader arriving after 22.5 and
22.6 looks up a CTE and a set operator in that table and finds neither, even
though the answer for both is decided by the same rule.

A CTE reference IS a derived table, so it gets the derived table's row verbatim.
The set operators split, and the split is the point: UNION ALL is the one
Elasticsearch answers itself, in one _msearch, at every venue — everything else
needs the jar. Publishing them as one row would hide exactly the distinction a
plain-REPL user needs.

"Works in this release" gains the non-recursive CTE bullet it never had; the
web twin already carried it, and these two pages are meant to agree sentence
for sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sweep on this branch corrected every page that DENIED these constructs; what it did not do is
document them. `dql_statements.md` had `## Set operators` and nothing else — no reference for the
two constructs the release leads with.

Two new sections, both mirrored on the web:

- **Subqueries and derived tables** — the venue table first, because "subqueries are supported" is
  wrong for a plain-REPL user about everything but the uncorrelated `WHERE` forms; the mandatory
  correlation name; opaque `SELECT *` bodies; the unqualified-name rule as the documented SQL-92
  deviation it is; the two-phase execution and its 65,536 bound; `NOT IN` and `NULL`; the
  correlated rules and what they cost against `maxJoins`.
- **Common table expressions** — left-to-right scope, and the one that surprises people: a CTE
  referenced twice is EXECUTED twice.

Every SQL statement in both sections was parse-probed AS PUBLISHED — extracted back out of the
rendered files, not out of a draft — and the probe was falsified by breaking a fence.

Four claims changed because the probe refused them:

- A CTE may not shadow the index it reads. The draft said it could.
- A scalar or quantified subquery is accepted only as the RIGHT operand; `WHERE (SELECT …) > 5` is
  rejected as `Unbalanced parentheses`, a message naming nothing. Published with the rewrite.
- A derived table's correlation name takes no column list.
- `NOT IN` + `NULL` has a carve-out that fails the OTHER way: the engine detects the NULL by
  re-running the body with an `IS NULL` filter, and under a `GROUP BY` body Elasticsearch's `terms`
  aggregation drops the missing-value group — so the NULL is invisible and rows come back that the
  rule says should not.

Also: `<=>` is named in the not-yet list (it was in neither refusal list and it blocks a captured
Tableau statement); `dbt` is added to match the web page; and `elastic` is gone as an example
catalog qualifier in twelve places — since story 20.3 the qualifier a BI tool emits is the CLUSTER
NAME, and the two pages had drifted apart on it.

`joins.md:114` is a since/before line and was SKIPPED, not edited: BIDC-8 owns that sentence and
has landed.

Refs #358
Story 22.7 pass 1 deliberately left these out: they are claims about what runs on which
Elasticsearch major, and no run had produced them yet. Pass 2 ran the matrix — the corpus statements
execute on 6.8, 7.17, 8.18 and 9.0 through both the REPL extension and the Flight SQL sidecar — so
the rows can be published from a measurement instead of an intention.

Derived tables, `WHERE` subqueries, correlated subqueries and non-recursive CTEs, all ✔ on all four.
Correlated is ✔ and not "pending": story 22.3b shipped it on both transports.

The note under the table is the part that stops the rows being misread. A ✔ means the Elasticsearch
major does not stand in the way; it is NOT a statement about the venue. Derived tables, CTEs,
correlated subqueries and every set operator but `UNION ALL` need the relational engine wherever they
run, on all four versions alike — and a reader who takes a version matrix as a capability matrix is
exactly who ends up filing "it works on 8.18 but not in the REPL".

Refs #358
…stamps are corrected

Seventeen findings from a full-diff review that ran 112 statements through the parser. The substance
held; these are defects of PRECISION in what we claim is refused, plus two stamps that cost a
customer a capability they already have. Every engine claim below was re-measured.

B1 — the one "rejected by name" item that is a SILENT WRONG ANSWER. An UNQUALIFIED outer reference
is not refused and cannot be: `WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = id)` is a
legal statement whose bare `id` binds to the body's own table. MEASURED: it parses with
`relationalClosureRequired = false`, i.e. it stops being correlated and runs as an ordinary
uncorrelated subquery — HTTP 200, different rows. The bullet refuted itself: its own second sentence
explained why nothing could reject it. Lifted out of the refusal list into its own callout on all
four pages, and the DQL lead-in becomes "one rule the engine enforces by name, and one convention
nothing can enforce for you". The QUOTED half of that bullet IS refused, so it stays, alone.

B2 — the `NOT IN` NULL rule was unconditional on the page that exists to hold exceptions. The
carve-out now appears there too: a `GROUP BY` body makes the NULL invisible, because the engine
detects it by re-running the body with an `IS NULL` filter and ES's `terms` aggregation drops the
missing-value group.

B3 — quoted, fully-qualified identifiers were stamped `0.24.0`; they shipped in `0.23.0`. Verified
by ancestry rather than by reading: `762023f7`, `8c426253` and `7c41c313` are all ancestors of tag
`v0.23.0`. Now "both quoted forms parse since 0.23.0; since 0.24.0 the derived-table wrapper also
EXECUTES". Also softened "earlier releases reject every form below AT THE PARSER" — false for
0.23.0, which parses a derived table and refuses it in the engine — to say only that they refuse it.

S6 — the two pages contradicted each other on a CTE named inside a WHERE subquery body, and the one
that committed published it as a WORKED EXAMPLE. RESOLVED BY EXECUTING IT at an engine venue, not by
choosing a side: it does NOT run. The reference is never resolved, the body asks Elasticsearch for
an index with the CTE's name, and it fails 404 `index_not_found_exception`. Loud, never silent. The
example is replaced with the JOIN form and both pages now say so definitely; the hedge is gone.

S9 — "rejected by name" over-claimed. Measured: the `LATERAL` KEYWORD (`FROM o, LATERAL (…)`,
`JOIN LATERAL`), `CORRESPONDING` and a derived-table column list all fail with a bare
`end of input expected`. Only the comma-`FROM` LATERAL spelling is named. The promise is narrowed to
the spellings that earn it, and the bare ones are labelled as syntax errors — a message that names
nothing is exactly what a reader needs this page for.

S8 — four refusals a "subqueries are supported" reader will hit, documented nowhere: a subquery in
`CASE WHEN` (named), in `JOIN … ON` (whose message names neither subqueries nor a rewrite — the
worst in the family), a WHERE subquery inside `CREATE MATERIALIZED VIEW` (named; a flagship feature,
so the first thing such a reader tries), and a FROM-less `SELECT 1` as a set-operation branch. 🔴 Two
of the review's candidates are NOT published: `CREATE WATCHER` (my probe was malformed — I will not
publish a refusal I did not measure) and a CTE over a catalog-qualified table, which MEASURED
EXECUTES, 5 rows, at an engine venue.

S4 broken anchor (core only; the web twin was already right). S7 the engine is installed BY DEFAULT —
said positively, on the page people read before buying, instead of only "a REPL with
`--no-extensions` does not". S10 `DECIMAL` was listed as Deferred on the web while the MD had it
accepted-as-approximate; the MD is right (`parser/type/package.scala`). S11 `DELETE` and `UPDATE`
take the same WHERE subqueries, with examples. N13 "Five things" followed by six bullets. N14 a
pointer to "what's coming in R2a/R2b" left behind on a page that no longer uses that vocabulary.
N15 `keywords.md` claimed the ordering quantifiers keep their own spelling; measured, `SOME`
normalises to `ANY` everywhere (`<= SOME` re-renders `<= ANY`).

MD↔MDX parity re-checked by diffing heading lists: identical but for the STDDEV section, which the
web deliberately carries as a bullet (pass 1's measured decision, unchanged).

Refs #358
…kwards

Adds the spellings that shipped with SoftClient4ES#360 in the same 0.24.0 train
this PR documents: CHAR_LENGTH / CHARACTER_LENGTH, TIMESTAMPADD, TIMESTAMPDIFF,
the ODBC SQL_TSI_ interval names, and SELECT TOP n.

🔴 And one correction that is not about the new spellings at all.
`functions_date_time.md` said DATEDIFF returns *"(date1 - date2)"* and published
EIGHT examples built on that reading, every one of them claiming a positive
result for arguments in the order later-date-first. The engine returns
`date2 - date1`. Measured on real Elasticsearch 8.18:

  DATE_DIFF('2025-01-10', '2025-01-01', DAY)  =>  -9     (the page claimed 9)
  DATE_DIFF('2025-01-01', '2025-01-10', DAY)  =>   9

It was found because TIMESTAMPDIFF had to be added to that same block, and MySQL
defines TIMESTAMPDIFF(unit, dt1, dt2) as dt2 - dt1 — so the page and the alias
could not both be right. The merged pin for
`DATE_DIFF(birthdate, CURRENT_DATE, YEAR)` settles which: it emits
`ChronoUnit.YEARS.between(birthdate, now)`, which is what makes it an AGE rather
than a negative one. Every example in the block has now been EXECUTED and shows
its real result, and a worked example of the reversed sign is published beside
them so the rule is visible rather than implied.

Documented for TOP, in `dql_statements.md`: it is a SPELLING of LIMIT, not a
second bound; `TOP (n)` works; `TOP` is NOT reserved, so `SELECT top FROM t`
still selects a column and `SELECT top - 1 AS x FROM t` still computes one;
TOP+LIMIT is refused by name; TOP carries no OFFSET.

New entries under "Not yet supported": `TOP n PERCENT` and `TOP n WITH TIES`
(refused by name), `SELECT DISTINCT TOP n`, and the ODBC/JDBC escape sequences
`{fn …}` / `{d …}` / `{ts …}` — a tool emitting
`{fn TIMESTAMPADD(SQL_TSI_DAY, -89, CURRENT_DATE)}` is still refused even though
the call inside it now parses. Also recorded: because PERCENT is recognised in
that position, a column of that name cannot be the sole select item straight
after `TOP n` — qualify or quote it.

🔴 There is NO TIMESTAMPSUB, in ODBC or in MySQL, and the docs say so: subtract
with a negative count.

Gate: these docs have no build, so the parse probe IS the gate. All 24 SQL
examples added or changed here were parsed against a parser built from
`origin/main` AFTER #360 merged — 0 rejected, 0 non-round-tripping — and the
date/length results were executed against real Elasticsearch 8.18 rather than
computed by hand. The unit-first `DATE_DIFF(unit, a, b)` spelling is documented
because the probe showed it is what the engine renders BACK.

Version stamped `0.24.0`, checked against `build.sbt` on main rather than
assumed — #360 is an ancestor of main and main is 0.24.0-SNAPSHOT.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… two

Rewrites the block for the behaviour issue #363 gives them, so the section is
written ONCE rather than published as-is today and revised in a few days.

⚠️ THIS PR MUST NOT MERGE BEFORE SoftClient4ES#363. Until that lands, `DATEDIFF`
still returns the old sign and this page is ahead of the engine.

- `DATE_DIFF` / `TIMESTAMPDIFF` keep their own entry: `date2 - date1`, first
  argument is the START. Unchanged behaviour, and it is what makes
  `DATE_DIFF(birthdate, CURRENT_DATE, YEAR)` an age.
- `DATEDIFF` gets its own entry, because after #363 it is its own function:
  MySQL's, `expr1 - expr2`, days only.
- 🔴 The one genuinely surprising thing is stated in a table rather than left to
  be discovered: on `DATEDIFF` alone, the TWO-argument and THREE-argument forms
  subtract in opposite directions, because the 3-arg form is this engine's own
  extension and is not MySQL. Adding a unit — which looks purely clarifying —
  reverses the sign. The page says so, and says to prefer DATE_DIFF or
  TIMESTAMPDIFF if you want a unit, where one rule holds at every arity.
- An upgrade note: before 0.24.0 `DATEDIFF` returned the opposite sign, so
  statements written against the old behaviour need their arguments swapped.

The examples are MySQL's OWN two documented calls plus the arity pair, all of
them executed rather than reasoned about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@fupelaqu fupelaqu changed the title docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it ⛔ HOLD until SoftClient4ES#363 merges — docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it Sep 18, 2026
@fupelaqu fupelaqu changed the title ⛔ HOLD until SoftClient4ES#363 merges — docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it docs: subqueries, derived tables, CTEs and set operators ship in 0.24.0 — correct the pages that deny it Sep 18, 2026
…ss (#398)

`materialized_views.md` mentioned `HAVING` exactly once — inside the worked
aggregation example — while a released build refuses seven shapes of it, so
a user meeting the refusal had nowhere to read about it.

The new `### HAVING in a materialized view` says, in this order: the case
almost everyone meets first (an aggregate with no `SELECT` alias — the view's
transform names each aggregation after that alias, so an unaliased one leaves
the group filter naming a metric the view never creates); why a view differs
from a search at all (five mechanisms carry a `HAVING`, a transform's pivot
offers one — no `terms` `include`/`exclude` channel, no `bucket_script`, and
only MIN/MAX/SUM/AVG/COUNT computed); the eight refusals with a reason and a
remedy each; what still works; and what changed.

Every SQL statement published here was MEASURED against the merged guard, in
the exact spelling the page prints (semicolon and `CREATE OR REPLACE` forms
included): the refused ones are refused, the accepted ones accepted, and the
page's pre-existing JOIN example is still accepted unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Documentation states subqueries and derived tables are unsupported; they ship in 0.24.0

1 participant