Skip to content

Story 22.7 — the corpus replay becomes a tracked series, and Epic 22's split is measured at 70/99 - #359

Merged
fupelaqu merged 4 commits into
mainfrom
feature/22.7
Sep 18, 2026
Merged

fupelaqu merged 4 commits into
mainfrom
feature/22.7

Conversation

@fupelaqu

Copy link
Copy Markdown
Contributor

Epic 22's acceptance story, elasticsql half. Extends story 21.6's replay harness rather than forking it, publishes the Epic-22 split as a measured number with both denominators, and turns the scoreboard into a series Epic 23 appends to.

Closes #358

The headline the suite prints

[22.7] corpus replay: SCORES 70/99 (was 56 after Epic 21, 12 before it); 70/75 of the statements we intend to answer; 24 capability probes excluded from scoring; Epic 22 rows: 10 of 12 scored fixed; stories 22.2 / 22.3 / 22.6 corpus credit: 0 / 0 / 0. 91 PARSE — the 21-row difference is never counted (21 capability probes + 0 that parse but answer wrongly, incompletely or UNMEASURED).

That trailing + 0 is the sentence worth keeping: nothing in the corpus parses and answers wrongly any more. The gap is entirely SQL we deliberately decline to score.

The five statements that remain refused are all dialect or by-design, not structural gaps: T-SQL TOP n, the ODBC {fn …} escape, the CHAR_LENGTH spelling, MySQL's null-safe <=>, and an alias-less derived table carrying Oracle ROWNUM.

What is in the diff

  • CorpusReplaySpec extended with the Epic-22 owners, and four gates on top of 21.6's: G4b pins the by-design rejections in compiled code and asserts they STAY rejected; G7 gains the rejected_by_design biconditional (without it a row could be scored under the wrong owner where G4b, which filters on owner, could never see it); G8 asserts the series head equals the live run; G9 proves the corpus resource and story 22.1's DerivedTableCorpusSpec carry byte-identical statements for the eleven derived ids — two files, two authors, one text.
  • corpus/series.csv — three measured rows (pre-epic21 · pre-epic22 · post-epic22), append-only by convention, head gated by G8.
  • corpus/baseline-pre-epic22.csv — measured in a detached worktree at the first parent of the first Epic-22 merge.
  • select.json — two honest limitation lines added, and one corrected: the trailing-ORDER BY rule it published was the opposite of what the engine does. Command syntax[] is never probed by the build, so it was hand-probed on every arm.

G8 is bookkeeping, not safety — and the README says so

Asked the standard question of it: what is the smallest edit to a data file that makes this pass while the thing it guards is broken? Answer: edit the head row, because the expectation lives in the file the author maintains. That is acceptable only because the verdicts are pinned independently — by G2 against the attribution table and by G4b against compiled ids — so a silent scoreboard move is impossible without also moving a row those gates police. Anyone reading G8 as the guard has mis-read it, which is why it is written down.

Evidence

12 of 12 gates falsified — each mutated until it reddened, then restored by bytes. Including the three added after an independent review: G9 reading story 22.1's own literal rather than a local copy of it, the compiled headline pin, and the pre-Epic-22 cross-file invariant.

A finding from that review worth recording: a two-cell CSV edit could move the published total past everything except G8 — the one gate the README says can never be the guard. Hence the compiled pin.

Parse cost

No engine source is touched by this branch — the diff is test resources, one test suite, one help JSON. git diff --name-only origin/main...HEAD -- sql/src/main core/src/main/scala is empty, so the grammar cannot have moved.

Measured anyway, before and after: ParserSpec 842 ms → 827 ms median of 10 (ranges 807–878 → 792–878, fully overlapping ⇒ no signal). Apple-silicon macOS, Zulu 11. No timing assertion — a wall-clock assertion on a shared runner is a CI-flake liability.

Verification

sql/test 1330/1330 · the four census suites 57/57 · HelpCorpusSpec 19/19 after core/clean · ++ 2.12.20 sql/Test/compile · the CI lint line byte-for-byte.

Not in this PR

The arrow and jdbc halves of story 22.7 — the corpus statements executing on real Elasticsearch 6.8 / 7.17 / 8.18 / 9.0, the capability declarations, and the epic close-out — are on feature/22.7 in their own repositories, held until core 0.24.0-SNAPSHOT is published. They are done and green locally; what they cannot do yet is resolve core from JFrog.

🤖 Generated with Claude Code

…split is published at 64/99

Story 21.6's scoreboard is EXTENDED, never forked: a second baseline measured at `40c8c63e` (the
first parent of the first Epic-22 merge, verified both ways), a tracked `series.csv` with one row
per measured tree, the `epic22` and `rejected_by_design` owners, and a code-pinned scoring
partition of the twelve corpus statements Epic 22 owns.

SCORES 64/99 (was 56 after Epic 21, 12 before it); 64/75 of the statements we intend to answer;
91 PARSE. Stories 22.2 / 22.3 / 22.6 contribute a measured ZERO and the headline says so.

Two rows are deliberately NOT scored. `tableau.mysql.w1.019` and `tableau.sql92.wx.009` wrap a
fully-quoted qualifier inside the derived body and every merged witness executes a BARE index
name, so nothing has measured that spelling. They are `epic22` / `residual` behind a code pin
whose failure message says what to do when the measurement lands — not `fixed` (a declaration
whose only evidence is an unrun test), not a `local:` defect slug (that publishes a failure
nobody has seen).

Four defects fixed on the way:

- The historical number was derived from a baseline file, and a verdict file cannot answer a
  scoring question. It had already diverged — 81 parsed at the pre-Epic-22 tree, 60 of them
  non-probe, against a published Epic-21 headline of 56. The summary now reads its history from
  `series.csv` and names it as the source.
- `tableau.sql92.wx.012` is refused for its missing correlation name, not for `ROWNUM`; `ROWNUM`
  round-trips as an ordinary identifier, which is the HAZARD rather than the refusal.
- `select.json` read as licensing a parenthesised whole set operation (rejected), and never said
  a CORRELATED subquery needs the relational engine. All 31 `syntax[]` lines hand-probed —
  nothing in the build probes them (HELP-1a).
- `TableauLiveReplaySpec` kept `epic22a_derived_table` alive with zero rows.

Twelve gate mutations RED, control green either side. Three of the gates exist because an
independent review asked the standing question of the others: G9 compared against a local copy of
story 22.1's literals and now reads 22.1's own; the published total is pinned in compiled code,
because a two-cell CSV edit previously passed every gate except the one the README says can never
be the guard; and the pre-Epic-22 no-regression gate is defended by a cross-file invariant rather
than a row count.

`sql/test` 1330/1330, `HelpCorpusSpec` green after `core/clean`, CI lint line byte-for-byte green.
`ParserSpec` x10: 842 ms -> 827 ms median (807-878 -> 792-878, overlapping; no signal), Apple
silicon, Zulu 11. No timing assertion.

Closed Issue #358
… headline is 68/99

The four `issue:328` rows carried a note handing the decision to this story. I declined and routed
it up; the lead ruled to take it.

The reasoning, recorded because the rule it bends is load-bearing: the published sentence claims
what the engine ANSWERS, not which epic earned it. Correctness here is asserted by a merged,
five-client suite against real Elasticsearch — GroupByCompletenessSpec "corpus shape: HAVING with
no GROUP BY" (PR #327), one row for a true predicate and ZERO rows for a false one — which is
exactly AD-10's bar. Understating Tableau's own per-data-source probe by four rows was safe in
direction but inaccurate in a headline that is about BI SQL.

SCORES 68/99, 68/75. `91 PARSE` is unchanged and the gap is re-derived: 91 - 68 = 23 = 21 parsing
capability probes + the 2 UNMEASURED E5 rows.

🔴 The admission is a COMPILED SINGLETON (`Issue328FixedIds`), never a loosened predicate and never
`owner.startsWith("issue:")`. G7 exists so that `fixed` cannot quietly come to mean "it parsed"; a
widened predicate would hand that meaning to every future `issue:` owner in silence. An exception a
reader can ENUMERATE keeps the rule a rule, and it is checked BOTH ways so a dead exception cannot
survive either.

🔴 The series head row was EDITED, not appended, and the reasoning now sits in the README beside the
append-only sentence. Append-only exists so a MERGE that flips a verdict cannot be hidden by moving
the expectation; all three of its preconditions failed here — nothing had merged, the number had
never left the branch, and what changed was the SCORING POLICY rather than the tree. `series.csv`
has one row per measured TREE and no column that could tell two policies on one commit apart, so a
second row for 7187c7d would read as a contradiction and manufacture a history nobody measured. G8
going red until the head moved is the mechanism working. A change to the TREE never qualifies.

The four notes were REWRITTEN in place. They previously argued the rows were "DELIBERATELY NOT
re-scored" and that the headline "UNDERSTATES the corpus by these four rows"; left standing, the
artefact would contradict itself and the next reader would quote whichever half they hit first.

Falsified, six mutations, all RED with the control green either side: the exception in BOTH
directions; both bumped count pins reverted; the series head reverted; and — the one that proves
the widening did not leak — a different, non-excepted `local:` owner scored `fixed` is still
refused by G7.

sql/test 1330/1330; the four census suites 57/57; CI lint line green; ++ 2.12.20 sql/Test/compile
green. Neither docs branch quotes a corpus number, so the two docs commits do not move.

Closed Issue #358
…; the headline is 70/99

Pass 1 could not run E5 and refused to guess it: `tableau.mysql.w1.019` and `tableau.sql92.wx.009`
wrap a fully-quoted qualifier inside the derived body, every merged witness executed a BARE index
name, so the two rows were scored `residual` under a code pin rather than declared `fixed` on an
unrun test.

Pass 2 ran them. Rows E5a / E5b / E5c in softclient4es-arrow's JoinExtensionIntegrationSpec and
`corpus E5-jdbc` in the jdbc testkit return exactly one row with TblMax = 1 on real Elasticsearch
6.8, 7.17, 8.18 and 9.0, and DuckDB answers the same text as ordinary schema-qualified SQL.

That SETTLES the 22.4 review's prediction that such a qualifier folds into the index name and fails
with `index_not_found`. It does not: the grammar keeps the qualifier in `Table.parts` and the bare
last part as `Table.name`, so the read reaches the index. Measured, not argued — which is why the
row existed.

SCORES 70/99, 70/75. `91 PARSE` is unchanged and the gap is now 21 — every one of them a temp-table
capability probe, and NOTHING else. For the first time no statement parses and answers wrongly.

`Epic22UnmeasuredIds` is kept as an EMPTY set with its history rather than deleted: that shape —
`epic22` / `residual` behind a code pin, with a failure message naming the next action — is the one
to reuse the next time a row parses before anyone has executed it.

The series head is edited rather than appended, for the second and last time on this unpublished
branch, and the README now states the rule with both instances: the TREE did not move (the grammar
at 7187c7d is the same in all three readings) — what moved was first a scoring policy and then the
EVIDENCE about rows that already parsed. A change to the tree is always an append.

sql/test 1330/1330; CorpusReplaySpec 17/17; CI lint line green.

Closed Issue #358
… ORDER BY

`select.json`'s set-operator block carried "-- ORDER BY / LIMIT after the last branch belong to that
branch". MEASURED: that is true for `UNION ALL` and FALSE for the other three —
`SELECT a FROM t1 UNION SELECT a FROM t2 ORDER BY a LIMIT 10` is REJECTED, with the very message the
documentation pages describe. The line sat under a block covering all six operators, so the REPL was
telling customers the opposite of the docs for five of them.

Command `syntax[]` is NEVER probed by the build (HELP-1a), which is why a contradiction between the
shipped help and the shipped engine survived. Hand-probed, and the corrected text now matches the
engine on every arm: `UNION ALL` + trailing ORDER BY ACCEPTS; `UNION`, `INTERSECT` and `EXCEPT` +
trailing REJECT; both remedies the comment names — parenthesising the branch, and the derived-table
wrapper — ACCEPT.

`limitations[]` was already correct and is untouched. `HelpCorpusSpec` green after `core/clean`.

Closed Issue #358
@fupelaqu
fupelaqu marked this pull request as ready for review September 18, 2026 09:33
@fupelaqu
fupelaqu merged commit 767ec07 into main Sep 18, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Story 22.7 (pass 1): corpus replay becomes a tracked series, and the Epic-22 split is published at 64/99

1 participant