Conversation
…y it
The published documentation said subqueries, derived tables and correlated
subqueries were "not in this release". They ship in 0.24.0, so every page
carrying that claim was steering users away from a feature they have — most
expensively the BI-tool section, which told them to rewrite tool-generated
nested SQL as an explicit JOIN.
Every claim here was measured against the merged engine, not a spec: each
published example was run through Parser.apply, and the venue split was read
from relationalClosureRequired.
- known_limitations.md gains a "Subqueries and derived tables" reference
section: the accepted forms, the venue table (an UNCORRELATED WHERE
subquery runs ES-natively at every venue including a plain REPL; a
CORRELATED one and a derived table need the relational engine), the
65,536-value bound on the two-phase path, the ANSI NULL rules, the forms
still refused by name, and the maxJoins accounting.
- The false worked example is replaced by a CTE — which really is rejected —
and its accepted rewrite.
- README gains the capability in the DQL section and the roadmap, leading
with the form that works everywhere.
- CTEs, set operators beyond UNION ALL, LATERAL, a subquery in HAVING, a
subquery in the SELECT list, a UNION ALL or FROM-less body, a body
projecting more than one column, and a correlated body carrying its own
JOIN / derived table / window function stay documented as unsupported.
Also corrected while certifying the pages: Tableau's quoted, fully-qualified
identifiers and its derived-table wrapper do parse (both spellings measured),
and operators.md published three IN examples that do not parse or are
pointless (IN (), IN ('active', NULL), a DISTINCT inside an IN body).
Closes #345
…ne 0.24.0
A reader on 0.23.0 could not tell from the pages whether their build has
these forms; they would try one and get a parse error, which is the failure
the pages exist to prevent.
Follows the house convention already on these pages ("Since engine `0.23.0`",
"Since `0.20.4`") — bold or code-spanned version, no date, no -SNAPSHOT.
The stamp is attached ONLY to what is new in 0.24.0: subqueries in WHERE,
derived tables, correlated subqueries. Cross-index JOINs, UNION ALL and every
other pre-existing capability keep no version line, since one there would be a
new false claim.
The unsupported list is deliberately date-free and version-free. CTEs and set
operators have not been started, so "coming in the next release (Quarter 4
2026)" is removed from the prose: the heading becomes "Not yet supported", the
roadmap-timing paragraph says the list is planned rather than scheduled, and
the per-page one-liners say "not yet supported" with no target. A published
target becomes a broken promise the moment priorities move.
Closes #345
Applies the full house pairing ("Since engine `0.23.0` with arrow-extensions
`0.3.3`" is the existing form on joins.md) at the six places that state the
relational-engine requirement.
The pairing is attached ONLY where the jar is genuinely required — derived
tables and correlated subqueries. Uncorrelated WHERE subqueries execute
ES-natively at every venue including a plain REPL, so they carry the engine
version alone. Naming an extensions version there would tell a plain-REPL user
they need a jar they do not need, which is the inverse of the over-claim this
sweep exists to remove.
The venue table makes the split explicit per row: the uncorrelated row reads
"No — works at every venue, including a plain REPL with no extensions"; only
the correlated and derived-table rows read "Yes — arrow-extensions `0.3.4`".
Closes #345
keywords.md called itself "a list of reserved words recognized by the parser". It is not one, and a comparison against Parser.reservedKeywords shows the page diverges in BOTH directions: 26 reserved words are absent (EXISTS, ALL, UNION among them) and 26 listed words are recognised but not reserved (CEILING, UCASE, TRY_CAST, RLIKE, OVER …). So the page has always been a list of GRAMMAR keywords wearing a reserved-word title. Story 22.2 made that difference load-bearing rather than pedantic: EXISTS and ALL are reserved, while ANY and SOME were deliberately left unreserved so a column named `any` keeps parsing — even though `x = ANY (SELECT …)` is real grammar. A reader taking this page as "words I cannot use as a column name" is right about `exists` and wrong about `any`. Rather than reclassify 52 entries, the page now states the two sets and the rule, and the four new entries carry the distinction where it bites. The remedy for a genuine collision is quoting, not renaming. Measured, not assumed: `SELECT any, some FROM t WHERE any = 1` parses; `SELECT exists FROM t` and `SELECT all FROM t` are rejected; `SELECT "exists"` and the backtick spelling both parse; `= ANY (SELECT …)` is normalised to `IN (SELECT …)` in the render, while `> ALL` keeps its spelling. Also adds UNION ALL, which the page omitted entirely. operator_precedence.md gains EXISTS as a level-7 unary predicate and a note that a subquery does not change the level of its operator. The grouping claim in that note was verified on the AST, not by eye: `a = 1 AND b IN (SELECT id FROM u) OR c = 2` parses as `((GenericExpression AND InSubquery) OR GenericExpression)` — the same shape as the literal-list form. Closes #345
… them Story 22.5 puts non-recursive CTEs in this train, so the fourteen places this PR was about to publish "CTEs are not yet supported" would have repeated, one story later, exactly the defect the PR exists to correct. A CTE reference IS a derived table, so every line reads like the derived-table lines beside it, venue requirement included: it runs on the relational engine, which means arrow-extensions 0.3.4. What the engine actually refuses is stated rather than generalised, because "CTEs work" would be an over-claim: WITH RECURSIVE and CTE column lists (WITH a (x, y) AS …) are refused by name; a WITH clause is accepted only at the top of a SELECT, never inside a subquery body, CTAS, INSERT … SELECT or a materialized view; and a CTE body may not name the CTE itself — unlike PostgreSQL, which binds such a name to the base table. known_limitations.md's "what a not-yet-supported query looks like" example was the CTE that now works, so it is replaced by a de-duplicating UNION — still refused — and the CTE is shown as the thing that runs. Set operators beyond UNION ALL stay listed as unsupported: story 22.6 is in this train but not yet built, so those lines are corrected when it lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… by position
Story 22.6 puts UNION / UNION DISTINCT / INTERSECT [ALL] / EXCEPT [ALL] in
this train, so the ten places this PR was about to publish "set operators
beyond UNION ALL are not supported" would have repeated, one story later,
exactly the defect the PR exists to correct — as the CTE commit before it did.
dql_statements.md's "UNION ALL" section becomes "Set operators": the seven
spellings with where each one runs, the compatibility rules, precedence and
the grouping refusal, the ORDER BY / LIMIT split, and the execution model.
UNION ALL keeps its own row and its own paragraph because it is the one the
engine answers with a single _msearch, at every venue, with no relational
engine in the picture — everything else needs arrow-extensions 0.3.4, exactly
like a derived table.
Columns match BY POSITION, and that is documented as a CHANGE rather than as
a rule, because it is one: before this release branches were matched by name,
so SELECT id AS x ... UNION ALL SELECT id AS y ... returned {x -> null,
y -> ...} for every row, branch 1's own included. A reader with a UNION ALL
written against the old behaviour needs to check their branch order, so the
note says so.
Three statements the pages made that are no longer true, found by probing
rather than by reading: a UNION ALL CAN now be a derived-table body (so the
"wrapping it to sort globally is not supported either" line went), a subquery
body "may not be a UNION ALL" is now "may not be a set operation" (every
spelling is refused there, not just the one), and known_limitations.md's
"what a not-yet-supported query looks like" example WAS the de-duplicating
UNION — replaced by WITH RECURSIVE, which is genuinely refused and has no
rewrite.
keywords.md gains the six new words and the reserved-vs-recognised line for
them: UNION / INTERSECT / EXCEPT / ALL / DISTINCT are reserved, WITH and
RECURSIVE are not. Verified in the ALIAS position, which is where the
distinction bites: SELECT a AS recursive FROM t parses, SELECT a AS intersect
FROM t does not, and quoting rescues either.
Every example that ships was parsed against merged main (058f27e + 7187c7d)
through a throwaway probe reporting Parser(stmt) and relationalClosureRequired,
deleted afterwards — including the ones asserted to FAIL: WITH RECURSIVE,
CORRESPONDING, a parenthesised group, a trailing ORDER BY, a set operator
after a FROM-less select list, and a set operation as a subquery body. The
venue column is that probe's closure value, not a reading of the spec.
Refs #354
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "Which forms need the relational engine" table listed the uncorrelated subquery, the correlated subquery and the derived table — the three forms story 22.2/22.4 added — and stopped there. A reader arriving after 22.5 and 22.6 looks up a CTE and a set operator in that table and finds neither, even though the answer for both is decided by the same rule. A CTE reference IS a derived table, so it gets the derived table's row verbatim. The set operators split, and the split is the point: UNION ALL is the one Elasticsearch answers itself, in one _msearch, at every venue — everything else needs the jar. Publishing them as one row would hide exactly the distinction a plain-REPL user needs. "Works in this release" gains the non-recursive CTE bullet it never had; the web twin already carried it, and these two pages are meant to agree sentence for sentence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sweep on this branch corrected every page that DENIED these constructs; what it did not do is document them. `dql_statements.md` had `## Set operators` and nothing else — no reference for the two constructs the release leads with. Two new sections, both mirrored on the web: - **Subqueries and derived tables** — the venue table first, because "subqueries are supported" is wrong for a plain-REPL user about everything but the uncorrelated `WHERE` forms; the mandatory correlation name; opaque `SELECT *` bodies; the unqualified-name rule as the documented SQL-92 deviation it is; the two-phase execution and its 65,536 bound; `NOT IN` and `NULL`; the correlated rules and what they cost against `maxJoins`. - **Common table expressions** — left-to-right scope, and the one that surprises people: a CTE referenced twice is EXECUTED twice. Every SQL statement in both sections was parse-probed AS PUBLISHED — extracted back out of the rendered files, not out of a draft — and the probe was falsified by breaking a fence. Four claims changed because the probe refused them: - A CTE may not shadow the index it reads. The draft said it could. - A scalar or quantified subquery is accepted only as the RIGHT operand; `WHERE (SELECT …) > 5` is rejected as `Unbalanced parentheses`, a message naming nothing. Published with the rewrite. - A derived table's correlation name takes no column list. - `NOT IN` + `NULL` has a carve-out that fails the OTHER way: the engine detects the NULL by re-running the body with an `IS NULL` filter, and under a `GROUP BY` body Elasticsearch's `terms` aggregation drops the missing-value group — so the NULL is invisible and rows come back that the rule says should not. Also: `<=>` is named in the not-yet list (it was in neither refusal list and it blocks a captured Tableau statement); `dbt` is added to match the web page; and `elastic` is gone as an example catalog qualifier in twelve places — since story 20.3 the qualifier a BI tool emits is the CLUSTER NAME, and the two pages had drifted apart on it. `joins.md:114` is a since/before line and was SKIPPED, not edited: BIDC-8 owns that sentence and has landed. Refs #358
Story 22.7 pass 1 deliberately left these out: they are claims about what runs on which Elasticsearch major, and no run had produced them yet. Pass 2 ran the matrix — the corpus statements execute on 6.8, 7.17, 8.18 and 9.0 through both the REPL extension and the Flight SQL sidecar — so the rows can be published from a measurement instead of an intention. Derived tables, `WHERE` subqueries, correlated subqueries and non-recursive CTEs, all ✔ on all four. Correlated is ✔ and not "pending": story 22.3b shipped it on both transports. The note under the table is the part that stops the rows being misread. A ✔ means the Elasticsearch major does not stand in the way; it is NOT a statement about the venue. Derived tables, CTEs, correlated subqueries and every set operator but `UNION ALL` need the relational engine wherever they run, on all four versions alike — and a reader who takes a version matrix as a capability matrix is exactly who ends up filing "it works on 8.18 but not in the REPL". Refs #358
…stamps are corrected Seventeen findings from a full-diff review that ran 112 statements through the parser. The substance held; these are defects of PRECISION in what we claim is refused, plus two stamps that cost a customer a capability they already have. Every engine claim below was re-measured. B1 — the one "rejected by name" item that is a SILENT WRONG ANSWER. An UNQUALIFIED outer reference is not refused and cannot be: `WHERE EXISTS (SELECT 1 FROM orders o WHERE o.customer_id = id)` is a legal statement whose bare `id` binds to the body's own table. MEASURED: it parses with `relationalClosureRequired = false`, i.e. it stops being correlated and runs as an ordinary uncorrelated subquery — HTTP 200, different rows. The bullet refuted itself: its own second sentence explained why nothing could reject it. Lifted out of the refusal list into its own callout on all four pages, and the DQL lead-in becomes "one rule the engine enforces by name, and one convention nothing can enforce for you". The QUOTED half of that bullet IS refused, so it stays, alone. B2 — the `NOT IN` NULL rule was unconditional on the page that exists to hold exceptions. The carve-out now appears there too: a `GROUP BY` body makes the NULL invisible, because the engine detects it by re-running the body with an `IS NULL` filter and ES's `terms` aggregation drops the missing-value group. B3 — quoted, fully-qualified identifiers were stamped `0.24.0`; they shipped in `0.23.0`. Verified by ancestry rather than by reading: `762023f7`, `8c426253` and `7c41c313` are all ancestors of tag `v0.23.0`. Now "both quoted forms parse since 0.23.0; since 0.24.0 the derived-table wrapper also EXECUTES". Also softened "earlier releases reject every form below AT THE PARSER" — false for 0.23.0, which parses a derived table and refuses it in the engine — to say only that they refuse it. S6 — the two pages contradicted each other on a CTE named inside a WHERE subquery body, and the one that committed published it as a WORKED EXAMPLE. RESOLVED BY EXECUTING IT at an engine venue, not by choosing a side: it does NOT run. The reference is never resolved, the body asks Elasticsearch for an index with the CTE's name, and it fails 404 `index_not_found_exception`. Loud, never silent. The example is replaced with the JOIN form and both pages now say so definitely; the hedge is gone. S9 — "rejected by name" over-claimed. Measured: the `LATERAL` KEYWORD (`FROM o, LATERAL (…)`, `JOIN LATERAL`), `CORRESPONDING` and a derived-table column list all fail with a bare `end of input expected`. Only the comma-`FROM` LATERAL spelling is named. The promise is narrowed to the spellings that earn it, and the bare ones are labelled as syntax errors — a message that names nothing is exactly what a reader needs this page for. S8 — four refusals a "subqueries are supported" reader will hit, documented nowhere: a subquery in `CASE WHEN` (named), in `JOIN … ON` (whose message names neither subqueries nor a rewrite — the worst in the family), a WHERE subquery inside `CREATE MATERIALIZED VIEW` (named; a flagship feature, so the first thing such a reader tries), and a FROM-less `SELECT 1` as a set-operation branch. 🔴 Two of the review's candidates are NOT published: `CREATE WATCHER` (my probe was malformed — I will not publish a refusal I did not measure) and a CTE over a catalog-qualified table, which MEASURED EXECUTES, 5 rows, at an engine venue. S4 broken anchor (core only; the web twin was already right). S7 the engine is installed BY DEFAULT — said positively, on the page people read before buying, instead of only "a REPL with `--no-extensions` does not". S10 `DECIMAL` was listed as Deferred on the web while the MD had it accepted-as-approximate; the MD is right (`parser/type/package.scala`). S11 `DELETE` and `UPDATE` take the same WHERE subqueries, with examples. N13 "Five things" followed by six bullets. N14 a pointer to "what's coming in R2a/R2b" left behind on a page that no longer uses that vocabulary. N15 `keywords.md` claimed the ordering quantifiers keep their own spelling; measured, `SOME` normalises to `ANY` everywhere (`<= SOME` re-renders `<= ANY`). MD↔MDX parity re-checked by diffing heading lists: identical but for the STDDEV section, which the web deliberately carries as a bullet (pass 1's measured decision, unchanged). Refs #358
…kwards
Adds the spellings that shipped with SoftClient4ES#360 in the same 0.24.0 train
this PR documents: CHAR_LENGTH / CHARACTER_LENGTH, TIMESTAMPADD, TIMESTAMPDIFF,
the ODBC SQL_TSI_ interval names, and SELECT TOP n.
🔴 And one correction that is not about the new spellings at all.
`functions_date_time.md` said DATEDIFF returns *"(date1 - date2)"* and published
EIGHT examples built on that reading, every one of them claiming a positive
result for arguments in the order later-date-first. The engine returns
`date2 - date1`. Measured on real Elasticsearch 8.18:
DATE_DIFF('2025-01-10', '2025-01-01', DAY) => -9 (the page claimed 9)
DATE_DIFF('2025-01-01', '2025-01-10', DAY) => 9
It was found because TIMESTAMPDIFF had to be added to that same block, and MySQL
defines TIMESTAMPDIFF(unit, dt1, dt2) as dt2 - dt1 — so the page and the alias
could not both be right. The merged pin for
`DATE_DIFF(birthdate, CURRENT_DATE, YEAR)` settles which: it emits
`ChronoUnit.YEARS.between(birthdate, now)`, which is what makes it an AGE rather
than a negative one. Every example in the block has now been EXECUTED and shows
its real result, and a worked example of the reversed sign is published beside
them so the rule is visible rather than implied.
Documented for TOP, in `dql_statements.md`: it is a SPELLING of LIMIT, not a
second bound; `TOP (n)` works; `TOP` is NOT reserved, so `SELECT top FROM t`
still selects a column and `SELECT top - 1 AS x FROM t` still computes one;
TOP+LIMIT is refused by name; TOP carries no OFFSET.
New entries under "Not yet supported": `TOP n PERCENT` and `TOP n WITH TIES`
(refused by name), `SELECT DISTINCT TOP n`, and the ODBC/JDBC escape sequences
`{fn …}` / `{d …}` / `{ts …}` — a tool emitting
`{fn TIMESTAMPADD(SQL_TSI_DAY, -89, CURRENT_DATE)}` is still refused even though
the call inside it now parses. Also recorded: because PERCENT is recognised in
that position, a column of that name cannot be the sole select item straight
after `TOP n` — qualify or quote it.
🔴 There is NO TIMESTAMPSUB, in ODBC or in MySQL, and the docs say so: subtract
with a negative count.
Gate: these docs have no build, so the parse probe IS the gate. All 24 SQL
examples added or changed here were parsed against a parser built from
`origin/main` AFTER #360 merged — 0 rejected, 0 non-round-tripping — and the
date/length results were executed against real Elasticsearch 8.18 rather than
computed by hand. The unit-first `DATE_DIFF(unit, a, b)` spelling is documented
because the probe showed it is what the engine renders BACK.
Version stamped `0.24.0`, checked against `build.sbt` on main rather than
assumed — #360 is an ancestor of main and main is 0.24.0-SNAPSHOT.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… two Rewrites the block for the behaviour issue #363 gives them, so the section is written ONCE rather than published as-is today and revised in a few days.⚠️ THIS PR MUST NOT MERGE BEFORE SoftClient4ES#363. Until that lands, `DATEDIFF` still returns the old sign and this page is ahead of the engine. - `DATE_DIFF` / `TIMESTAMPDIFF` keep their own entry: `date2 - date1`, first argument is the START. Unchanged behaviour, and it is what makes `DATE_DIFF(birthdate, CURRENT_DATE, YEAR)` an age. - `DATEDIFF` gets its own entry, because after #363 it is its own function: MySQL's, `expr1 - expr2`, days only. - 🔴 The one genuinely surprising thing is stated in a table rather than left to be discovered: on `DATEDIFF` alone, the TWO-argument and THREE-argument forms subtract in opposite directions, because the 3-arg form is this engine's own extension and is not MySQL. Adding a unit — which looks purely clarifying — reverses the sign. The page says so, and says to prefer DATE_DIFF or TIMESTAMPDIFF if you want a unit, where one rule holds at every arity. - An upgrade note: before 0.24.0 `DATEDIFF` returned the opposite sign, so statements written against the old behaviour need their arguments swapped. The examples are MySQL's OWN two documented calls plus the arity pair, all of them executed rather than reasoned about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ss (#398) `materialized_views.md` mentioned `HAVING` exactly once — inside the worked aggregation example — while a released build refuses seven shapes of it, so a user meeting the refusal had nowhere to read about it. The new `### HAVING in a materialized view` says, in this order: the case almost everyone meets first (an aggregate with no `SELECT` alias — the view's transform names each aggregation after that alias, so an unaliased one leaves the group filter naming a metric the view never creates); why a view differs from a search at all (five mechanisms carry a `HAVING`, a transform's pivot offers one — no `terms` `include`/`exclude` channel, no `bucket_script`, and only MIN/MAX/SUM/AVG/COUNT computed); the eight refusals with a reason and a remedy each; what still works; and what changed. Every SQL statement published here was MEASURED against the merged guard, in the exact spelling the page prints (semicolon and `CREATE OR REPLACE` forms included): the refused ones are refused, the accepted ones accepted, and the page's pre-existing JOIN example is still accepted unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #345
The published documentation claimed subqueries and derived tables were not supported. They ship in
0.24.0, so every page saying otherwise was telling users a feature they have does not exist — most expensively the BI-tool section, which instructed them to rewrite tool-generated nested SQL as an explicit JOIN.Sibling PR: SOFTNETWORK-APP/softclient4es-web#62 — the two carry the same capabilities, the same caveats and the same venue distinction, and are intended to merge together, during the 0.24.0 train (not before: these features are not in
0.23.0).Version framing
The features are stamped **
Since engine \0.24.0`**, and where the relational engine is required, the full house pairing **Since engine `0.24.0` with arrow-extensions `0.3.4`**, following the convention already on these pages (Since engine `0.23.0` with arrow-extensions `0.3.3`,Since `0.20.4`) — bold or code-spanned version, no date, no-SNAPSHOT. A reader on0.23.0` can now tell from the page that their build rejects these forms.The stamp is attached only to what is new: subqueries in
WHERE, derived tables, correlated subqueries. Cross-index JOINs,UNION ALLand every other pre-existing capability carry no version line — one there would be a new false claim.The unsupported list is deliberately date-free and version-free: the heading is Not yet supported, the roadmap-timing paragraph says that list is planned rather than scheduled, and the per-page one-liners carry no target.
CTEs and set operators have not been started— both landed in this train after this PR was opened; see the two update sections at the end.The
softclient4es-arrow-extensionsjar is named with its version —0.3.4— where it is genuinely required. That version is confirmed as the number this train will cut (arrow main is on0.3.4-SNAPSHOT), but it has not been released yet.🔴 Merge condition: these pages name
0.3.4, so merging them depends on arrow actually cutting0.3.4in this train. If the number changes, both pages need a correcting pass before they go live.The pairing is applied only where the jar is required — derived tables and correlated subqueries. Uncorrelated
WHEREsubqueries execute ES-natively at every venue including a plain REPL, so they carry the engine version alone; naming an extensions version there would tell a plain-REPL user they need a jar they do not need.Now documented as working
WHEREsubquery —IN/NOT IN,EXISTS/NOT EXISTS, scalar comparison, quantified (= ANY|SOME,<> ALL,> ALL,>= ANY,< ALL, …)WHEREsubqueryFROM/JOIN, nested to any depthPlus the bounds: the two-phase path caps at 65,536 distinct values (
index.max_terms_count) and fails loudly past it, never truncating;INignores inner NULLs whileNOT INover a NULL-bearing set matches no rows; a correlated subquery costs onemaxJoinsunit and a derived table costs none.Still documented as unsupported
CTEs (
WITH …); set operators beyondUNION ALL(UNION,INTERSECT,EXCEPT);LATERAL; any subquery inHAVING; a subquery in theSELECTlist; aUNION ALLorFROM-less subquery body; a body projecting more than one column; an unqualified or quoted outer reference; and a correlated body carrying its ownJOIN, comma-separatedFROM,JOIN UNNEST, derived table or window function.How the claims were verified
Not from a spec. A throwaway
sql/src/test/…/DocSweepProbeSpec(deleted before committing) ran every statement throughParser.applyand printedrelationalClosureRequired. All 12 subquery examples the docs publish parse; theWITHexample is the deliberate rejection. The venue table is read fromSingleSearch.relationalClosureRequired/whereSubqueries; the refusal wording fromRelationalClosureGuard; the 65,536 bound fromSubqueryResolver.MaxTerms; the correlated-body restrictions and themaxJoinsaccounting from the merged relational-engine planner.These docs are plain Markdown with no build step, so
headerCheck/scalafmtCheckprove nothing about them — the parse probe is the gate. No Scala source changed.Two corrections found while certifying the pages, fixed here rather than filed
`bi_events`.`category`and"bi_events"."category"parse, and so does Tableau's derived-table wrapper. Corrected inbi_tools.md. Tableau's tier is unchanged — still Compatible.operators.mdpublished threeINexamples that do not parse or are pointless:IN ()andIN ('active', NULL)are parse errors, andDISTINCTinside anINbody buys nothing (the values are collected as a set, and a plain-column body is resolved with onetermsaggregation).Keyword pages (
keywords.md,operator_precedence.md)keywords.mddescribed itself as "a list of reserved words recognized by the parser" and omitted every word story 22.2 added. A comparison againstParser.reservedKeywordsshows it was never a reserved-word list and diverges in both directions: 26 reserved words absent (EXISTS,ALL,UNION,EXCEPT,MATCH,TRUE/FALSE,PI…) and 26 listed words recognised but not reserved (CEILING,UCASE,TRY_CAST,RLIKE,OVER,POINT…). It is a grammar-keyword list wearing a reserved-word title.🔴 Story 22.2 made that distinction load-bearing, which is why this is not "append two words".
EXISTSandALLare reserved;ANYandSOMEwere deliberately left unreserved so a column namedanykeeps parsing — whilex = ANY (SELECT …)is still real grammar. Listing all four together would publish a restriction that does not exist and send users renaming columns needlessly.Resolution: the page now states both sets and the rule rather than reclassifying 52 entries (a separate audit). The four new entries carry the distinction inline, and the remedy for a real collision is quoting, not renaming.
UNION ALL, absent from the page entirely, is added.Measured, not assumed — every claim probed through
Parser.apply:SELECT any, some FROM t WHERE any = 1SELECT exists FROM t/SELECT all FROM tSELECT "exists" FROM t/SELECT `exists` FROM t"exists"WHERE customer_id = ANY (SELECT …)IN (SELECT …)WHERE amount > ALL (SELECT …)operator_precedence.mdgainsEXISTSas a level-7 unary predicate and a note that a subquery does not change the level of its operator. That note's grouping claim was verified on the AST, not by eye:a = 1 AND b IN (SELECT id FROM u) OR c = 2parses as((GenericExpression AND InSubquery) OR GenericExpression)— the same shape as the literal-list form. The page already carried twoIN (SELECT …)examples that only became valid with 22.2.Cross-surface consistency — no code defect. The core
SQLKeywordsregistry listsEXISTS,ANY,SOME,ALLinclauseTokens, and the REPL completer derives its words fromSQLKeywords.highlightedWordsviaReplKeywords.all, so all four already tab-complete. The registry and the completer agreed with the parser; only the docs page was stale.known_limitationsmakes no reserved-word claim, so nothing there is contradicted.Files
README.md(DQL section + roadmap — an omission, not a false claim: it never mentioned subqueries),documentation/sql/{known_limitations,dql_statements,joins,operators,keywords,operator_precedence}.md,documentation/client/{bi_tools,jdbc,adbc_driver,arrow_flight_sql}.md.Deliberate twin divergence
known_limitations.mdkeeps its elasticsql-only "Quoted identifiers — residual limits" section, andoperators.md,keywords.mdandoperator_precedence.mdhave no web twin — the site has no keywords or precedence page at all. All of that pre-dates this PR, so the keyword work here is deliberately elasticsql-only and PR #62 carries none of it; creating a site keywords page is a content decision, not a correction. The subquery sections themselves are identical across the two repos apart from link syntax.Update 2 — set operators (story 22.6) and positional column matching
Story 22.6 shipped
UNION/UNION DISTINCT/INTERSECT [ALL]/EXCEPT [ALL], and the follow-up fix (#357, closing #354 and #355) changed how branch columns are matched. Ten places in this PR still said "set operators beyondUNION ALLare not yet supported", so they would have repeated — one story later — exactly the defect this PR exists to correct, as the CTE commit before it already had to.What changed
UNION ALLsection becomes Set operators: the seven spellings with where each runs, the compatibility rules, precedence and the grouping refusal, theORDER BY/LIMITsplit, and the execution model.UNION ALLkeeps its own row — it is the one Elasticsearch answers by itself with a single_msearch, at every venue; everything else needs arrow-extensions0.3.4, exactly like a derived table.UNION ALLvsUNION/INTERSECT/EXCEPTsplit.SELECT id AS x … UNION ALL SELECT id AS y …returned a null column for every row — branch 1's own included. A reader with aUNION ALLwritten against the old behaviour has to check their branch order, and the note says so.keywords.mdgains the six new words plus the reserved-vs-recognised line for them:UNION/INTERSECT/EXCEPT/ALL/DISTINCTare reserved,WITHandRECURSIVEare not.Three claims on these pages that are no longer true, found by probing rather than by reading:
UNION ALLcannot be a derived-table body" (so it cannot be wrapped to sort globally)UNION ALL"UNIONWITH RECURSIVE, which is genuinely refused and has no rewriteVerification. Every example that ships — including the ones asserted to fail — was parsed against merged
main(058f27e2+7187c7d9) by a throwaway probe reportingParser(stmt)andrelationalClosureRequired, deleted afterwards. The venue column in every table is that probe's closure value, not a reading of the spec. The refusals asserted:WITH RECURSIVE,CORRESPONDING, a parenthesised group, a trailingORDER BYafter the last branch, a set operator after a FROM-less select list, and a set operation as a subquery body. The reserved-word claims were checked in the alias position, which is where the distinction bites —SELECT a AS recursive FROM tparses,SELECT a AS intersect FROM tdoes not.The far-side refusals were read in arrow, not inferred from core: a catalog-qualified set operation is refused by name in
JoinPlanner.plan, because catalogs resolve by text position and a branch would otherwise run on the wrong cluster. That refusal is documented.Update — the BI dialect spellings from #360 (same 0.24.0 train)
This PR now also documents
CHAR_LENGTH/CHARACTER_LENGTH,TIMESTAMPADD,TIMESTAMPDIFF, the ODBCSQL_TSI_*interval names andSELECT TOP n, plus the new refusals (TOP n PERCENT,TOP n WITH TIES,SELECT DISTINCT TOP n, and the ODBC{fn …}escape family, which is still refused even though the call inside it now parses).🔴 And one correction that has nothing to do with the new spellings.
functions_date_time.mdsaidDATEDIFFreturns "(date1 - date2)" and published eight examples built on that reading. The engine returnsdate2 - date1. Measured on real Elasticsearch 8.18:It surfaced because
TIMESTAMPDIFFhad to go in that same block and MySQL definesTIMESTAMPDIFF(unit, dt1, dt2)asdt2 - dt1— the page and the alias could not both be right. The merged pin forDATE_DIFF(birthdate, CURRENT_DATE, YEAR)settles which: it emitsChronoUnit.YEARS.between(birthdate, now), which is what makes it an age. Every example in the block has been re-executed and shows its real result, with the reversed-sign case published beside them.Gate. These docs have no build, so the parse probe is the gate: all 24 SQL examples added or changed were parsed against a parser built from
origin/mainafter #360 merged — 0 rejected, 0 non-round-tripping — and the date/length results were executed against real Elasticsearch 8.18 rather than computed by hand. The probe also showed thatDATE_DIFF(unit, a, b)is what the engine renders back, so that spelling is documented too.Version stamped
0.24.0, checked againstbuild.sbton main rather than assumed.DATEDIFFsection is accurate today, and may change again — see #363A correction to how the
DATEDIFFchange above was first described in this PR.The page was not sloppy. MySQL 8.4 defines
DATEDIFF(expr1, expr2)asexpr1 − expr2, and the page — description and all eight examples — described MySQL faithfully. What it did not describe is this engine, which returnsexpr2 − expr1becauseDATEDIFFis an alias word on theDATE_DIFFtoken and inherits BigQuery'sDATE_DIFF(start, end, unit)order.Run against MySQL's own two documented examples, on real Elasticsearch 8.18:
DATEDIFF('2007-12-31','2007-12-30')1-1DATEDIFF('2010-11-30','2010-12-31')-3131So the real defect is arguably in the engine, not the page: a MySQL-named function that contradicts MySQL, silently. That is now filed as #363. Note we already match MySQL on
TIMESTAMPDIFF(dt2 − dt1), soDATEDIFFis the single odd one out rather than a house convention.What this PR does: documents the engine as it behaves today, which is the accurate thing to publish right now. If #363 is fixed, this section changes again — that is expected, and is why the note is here rather than in a commit message.