Skip to content

fix(sql): a HAVING over a function of an aggregate must filter, or fail — never vanish (#389) - #397

Merged
fupelaqu merged 6 commits into
mainfrom
feature/389
Sep 25, 2026
Merged

fupelaqu merged 6 commits into
mainfrom
feature/389

Conversation

@fupelaqu

@fupelaqu fupelaqu commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Closes #389

A HAVING predicate whose aggregate sits inside a function argument emitted no bucket_selector: the clause vanished and every group came back, HTTP 200. The #205 / #253 silent-wrong-answer family.

The cause was not the one the issue named

#389 blamed ElasticAggregation.metricSelectorForBucket's if (fullScript.isEmpty) return "". Measured: that is the second line of defence. One rung up, in sql, MetricSelectorScript.metricSelector falls to its case _ => "1 == 1" catch-all because Expression.referencesBucketMetric is false — a function never looks inside its own arguments, so COUNT(*) within ABS(...) / COALESCE(...) / a CASE branch is invisible to isAggregation and hasAggregation. That single fact explained the dropped clause, the missing auxiliary aggregation, and why the whole-table guard whose comment promises "nothing may fall through silently" did not fire.

Three things the issue does not record, all measured on main: an AND conjunct was silently dropped (a partial filter — the answer looks filtered and is wrong), an OR disjunct likewise, and the whole-table HAVING took the same path.

What ships

A function-wrapped aggregate in HAVING now emits a correct bucket_selector, or the statement is refused by name. Never dropped, never partially applied.

  • HAVING COALESCE(COUNT(*), 0) > 1 and GREATEST(...) forms emit and run.
  • NULLIF / ROUND / CASE / boxing forms are refused, naming the clause and the reason. The gate is structural, on the rendered Painless — not a function-name list, which would be a fourth derivation of the renderer. ⚠️ The bucket-pipeline Painless context has a different whitelist from the query context: Double.valueOf(x) compiles in one and not the other. Only execution settled that.
  • Conditions on the GROUP BY key are honoured at bucket level by the terms include/exclude filter, or refused. The channel unions, it never intersects — an OR across two grouping levels is a conjunction, because Elasticsearch nests the terms aggregations.
  • Four shapes that reached Elasticsearch as HTTP 400 are now named rejections — one of them leaked raw SQL into the Painless source, one surfaced Internal parser error to the user.
  • Having.script is restored (@deprecated): it has a production caller in softclient4es-extensions (graph/Stage.scala:291), contrary to what an earlier round of this branch asserted.
  • Error plumbing: a refusal now reaches the caller on search, searchAsync, searchWithInnerHits and ScrollApi, not just search.

Verification

sql 1635 · core 1141 · bridge 238 · es6bridge 234 · + compile (2.12.20 + 2.13.16) · the exact CI lint line — all green. Integration re-run after the final change: 47/47 on real ES 8.18.3, 47/47 on real ES 6.8.23.

Four independent fresh-context reviews, each over the commits the previous one had not seen, executing against hand-computed oracles on real clusters. The last measured 0 surviving wrong answers over its own independently generated 309-statement population.

Behaviour changes — release-note material

  1. Statements that returned every group now return the right groups, or fail. Dashboards built on the old answer will show fewer rows.
  2. Previously-accepted statements are now refused by name — and, contrary to what this PR body originally said, they were not working. ⚠️ CORRECTION, measured on 455433ae after the merge (verdict AND emitted query): GROUP BY t.status HAVING t.status = 'a', GROUP BY t.status HAVING status = 'a', GROUP BY age HAVING age > 0, GROUP BY status HAVING status IS NOT NULL and the JOIN form GROUP BY d.category HAVING d.category = 'x' were all ACCEPTED on main and emitted no include at all — the condition was silently dropped and every group was returned. That is HAVING <function>(<aggregate>) silently drops the entire HAVING clause — every group is returned #389 itself, in the group-key population; the refusal makes the same defect loud and nothing regressed. Only the UNQUALIFIED direct comparison (HAVING x = 'v', <>, IN, LIKE) was ever correct, and it is unchanged. The original wording ("correct only coincidentally") relayed a reviewer's classification that had not been measured. Aggregation over JOINs is unaffected — GROUP BY d.category HAVING COUNT(*) > 1, GROUP BY d.category, o.kind HAVING COUNT(*) > 1 and WHERE age > 0 GROUP BY age HAVING COUNT(*) > 1 are all accepted (measured on the merged tree). Residual, now filed as a follow-up: a qualified key should resolve to its bucket and does not — same family as SQL parser: support backtick-qualified catalog names natively #85 / Two tables differing only by qualifier collapse into one alias, silently dropping an alias #292.
  3. HAVING <key> LIKE now uses the shared translation. _ becomes . and regex metacharacters are escaped, so LIKE 'a_' — which matched nothing — now matches, and LIKE 'cat_3%' emits cat.3.* rather than cat_3.*. No test pin anywhere used such a literal on this channel, which is how it survived.
  4. Binary compatibility: additive. No downstream rebuild required by this PR.

Recorded, not fixed — pre-existing and byte-identical to main

  • ES 6.8 only: a single-valued include/exclude renders as a bare string that ES 6.8 reads as a regex, so HAVING status = 'a.b' returns bucket axb. Wider than first recorded: a pattern in one channel beside a multi-value list in the other is HTTP 400 there (three further shapes).
  • child(...) / nested(...) puts an include on the parent bucket key.
  • HAVING k = 'a' OR child(LENGTH(k) = 1) executes as its first branch.
  • The vanished GROUP BY in a JOIN UNNEST + root-metric shape — a separate defect on a different mechanism. That statement is no longer answered 200-with-a-wrong-answer; it is refused.

A capability that was built and then dropped

A case-fold inversion (HAVING UPPER(key) = 'A' → a terms include regex) was built across two rounds and removed by decision. It produced three silent wrong answers and a remotely-triggerable OutOfMemoryError from a ~30-character string literal. The structural reason it was dropped: it inverts a case fold performed by a JVM the client has never seen. Reading the client JVM made answers depend on the client's JDK (96 patterns differ between zulu-11 and zulu-21); freezing a table relocated the divergence to the node (ES 7.17 → Unicode 15, ES 8.18/9.0 → Unicode 16, versus a pinned Unicode 13 — 134 extra cased code points, with a missing bucket proved end-to-end). The work is preserved locally and the reasoning is recorded so it is not rebuilt the same way.

One lesson worth carrying

Every defect found on this branch — across six dev rounds and four reviews, including two rounds where the fix shipped something worse than what it fixed — was a hole in the population under test, never in the assertions. The mutation matrices were sound each time. The failures were all absences — a bucket that is not there — which only an independently computed expected set can see.

🤖 Generated with Claude Code

Cost — measured on this tree (1a6a23a1), reported, not asserted

9 interleaved rounds per side against a 455433ae control in one session, with six must-be-zero baseline shapes (bare projection, WHERE-only, wide WHERE, plain GROUP BY, 2-key GROUP BY, ORDER BY) carrying the noise floor. No timing assertion is added anywhere (#269/#270).

Search — nothing to degrade. Of the 99 well-formed HAVING statements present in the control tree's own sources, 99 emit byte-identical queries. Zero changed, zero newly refused, zero refused→emitted. The only queries that move are ones the control emitted wrong (a bucket_selector reading params.max_x with no max_x aggregation — that query fails at runtime) or did not emit at all (#389). Their extra ES work is one metric sub-aggregation plus a coordinating-node pipeline. RowCostProbe was deliberately not run: every core/src/main hunk here is in a failure branch (invalidSearchRequest and its 5 call sites, ScrollApi's Source.failed), so nothing on the row-extraction, paging or response-parsing path is touched.

Parse + validate (µs, median of 9 rounds × 500 runs; floor over the 6 baseline shapes = −4.6 … +1.7):

shape control branch Δ verdict
HAVING COALESCE(COUNT(*),0) > 1 318.8 329.3 +10.5 signal, +3.3% (paired +12.5; ranges non-overlapping)
HAVING COUNT(*) > 1 270.0 272.9 +2.9 marginal (~1%)
HAVING COUNT(*) > 1 AND MAX(a) < 9 323.6 325.7 +2.1 marginal (~0.6%)
HAVING n > 1 (alias) 258.7 260.0 +1.3 no signal
HAVING code = 'x' (GROUP BY key) 262.5 256.9 −5.6 no signal
every non-HAVING shape — — inside floor no signal

Resolution columns, same run (floors: resolve ±0.4, validate ±0.2): a metric-bearing GROUP BY … HAVING resolves in 2.3 → 6.6 µs, and COALESCE(COUNT(*),0) in 2.4 → 8.9 µs. A ratio of ~3×, ~4–6 µs in absolute terms against a ~270 µs parse. A GROUP-BY-key HAVING pays nothing (1.5 → 1.8 µs).

Rendering / emission (fresh parse + ElasticSearchRequest.query; floor ±1.3 µs on min): metric-bearing HAVING costs +3.7 to +15.2 µs on a 329–348 µs parse+emit (≈1–4.5%), scaling with metric-leaf count — 1 leaf +6.0, 2 leaves +15.2. Key-only HAVING (=, IN, LIKE, <>) and every non-HAVING shape sit inside the floor.

Mechanism (reasoning, corroborated by the leaf-count scaling): MetricSelectorScript.selector now calls representable(e) per metric leaf — a painless(None) render plus a StringBuilder scan. That is the gate that converts four HTTP 400s and the silent drop into named refusals. It runs once per metric leaf at parse/emit time, never per document. Key-only leaves return NoFilter before reaching it, which is why they measure flat.

Two caveats stated rather than hidden: the branch ran first in every round (a fixed position, absorbed by the baseline floor — which is why the floor is reported), and the newly-refused statements are not counted as "slower" — a refusal short-circuits emission, and two of them measure faster.

fupelaqu and others added 6 commits September 24, 2026 10:19
…il — never vanish

`HAVING ABS(COUNT(*)) > 1` parsed, ran, and returned EVERY group with HTTP 200 — and so did
`COALESCE`, `NULLIF`, `ROUND`, `GREATEST`, a `CASE`, an aggregate on the RIGHT of the comparison,
and the whole-table form. The #205/#209/#253 silent-wrong-answer family.

Cause, in one sentence: a function never looks inside its own ARGUMENTS. `FunctionChain
.hasAggregation` reads the chain's OWN links, so `COUNT(*)` inside `ABS(...)` made the tree answer
"no aggregate here" — no aggregation was created, `buckets_path` published nothing, and the
selector fell to its `1 == 1` catch-all, which the caller reads as "nothing to filter".

ONE derivation, `Identifier.bucketMetrics`, now feeds all three consumers that used to disagree:
what `buckets_path` publishes, which aggregations get created, and the `bucket_selector` null
guard. It needs TWO walks — `funIdentifiers` reaches a function argument and a `CASE`'s THEN
result, while a `CASE`'s WHEN conditions are reachable only through `Case.conditionsOf`.

Everything the engine can express is emitted; everything it cannot is REFUSED BY NAME at
`SingleSearch.validate()`, inside `Parser.apply`, so every venue sees it. A partial emission counts
as a wrong answer: one un-expressible conjunct refuses the whole statement.

The representability gate decides on the RENDERING, never on a list of function names — a name list
is a fourth derivation of what the renderer does, and it could not have expressed the rule that
only EXECUTION produced: the bucket-pipeline Painless context does NOT compile `Double.valueOf(…)`,
so `ABS`/`FLOOR`/`CEIL`/`SQRT`/`EXP`/`LOG`/`POWER` over an aggregate are refused, even though the
same function works in WHERE, in the SELECT list and in ORDER BY. The spec asserted the opposite;
eighteen shapes were rendered and executed against a real cluster to settle it.

Fixed on the way: `HAVING COUNT(*) > ABS(MAX(x))` emitted a script reading `params.max_x` with no
`max_x` aggregation anywhere; the nested form emitted NO aggregation at all and returned raw
documents; an OR-disjunct was dropped as well as an AND-conjunct; and `WHERE ABS(COUNT(*)) > 1` now
hits the existing aggregates-in-WHERE rejection.

No bridge file changed — the whole fix is in `sql`, proved by holding the hand-maintained es6 twin
to the same emission matrix.

Verified: 19/19 mutations RED · sql 1607 / core 1141 / bridge 225 / es6bridge 221 · `+ compile`
both Scala legs · GrammarDiffProbe over 18,165 inputs = 4 differing rows, every one intended ·
integration 32/32 on real Elasticsearch 8.18.3 and 32/32 on 6.8.23 · cost, 5 interleaved rounds vs
a 455433a control: parse+emit −0.39 % against a must-be-zero baseline reading −0.08 % ⇒ no signal.

Closed Issue #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… both operands (#389 round 2)

Round 2 on an independent review that executed the branch against a 455433a control and a real
Elasticsearch 8.18.3.

LEAD RULING — `HAVING <function>(<group key>)` must WORK, not be refused. It does now. The
equivalence it rests on was verified before building on it: a function of the GROUP BY key is
CONSTANT within a bucket, so filtering the documents by it keeps exactly the satisfying buckets and
changes nothing in a surviving one. Executed on real data, a surviving bucket's doc_count and every
metric are byte-identical to the unfiltered run, including for a multi-level grouping. The predicate
is carried by `SingleSearch.searchCriteria` and read by ONE line in each bridge; the AST and the
render are untouched, because `MaterializedViewExtension` persists that render.

What the equivalence does not license is refused by name: a predicate reading neither an aggregate
nor a key, a NESTED key (a document filter keeps the whole parent), and an OR mixing a key with an
aggregate — which closes a PRE-EXISTING silent wrong answer, measured on `main`:
`HAVING COUNT(*) > 1 OR status = 'b'` returned NO buckets where the disjunction is two of them,
because the key became a terms `include` and the metric a `bucket_selector`, so the OR executed as
an AND.

🔴 The review also found that the representability gate was bypassed whenever the LEFT operand was a
bare aggregate — so everything refused on the left was EMITTED on the right. Measured and executed:
`> ROUND(MAX(a),2)`, `> NULLIF(MAX(a),0)` and `BETWEEN 1 AND ABS(MAX(a))` (which leaked RAW SQL into
the Painless source) each returned HTTP 400, and `> CASE WHEN …` surfaced `Internal parser error` to
the user. The gate now runs on both arms. That also CORRECTS this branch's own record: the
`COUNT(*) > ABS(MAX(x))` row was documented as refused and was in fact running — the test for it
asserted detection and never a verdict.

Also closed: a metric compared as TEXT is refused rather than reaching Elasticsearch as a raw
class-cast; a metric computed outside a nested grouping is refused rather than answering HTTP 200
with no grouping at all; the gate is re-asked at `validateResolved()` so a post-resolution
divergence stays a 400 naming the clause; and the refusal reason now reaches the caller instead of
the client's generic sentence.

Two defects in this branch's own new code, both found by the mutation matrix reporting a refusal as
unguarded: HAVING has FOUR mechanisms and scoping past the nested filter refused a shipped fixture;
and `Identifier.bucket` is a stale copy reporting `nestedElement = None`, which made the nested
refusal silent and pushed the one predicate the equivalence forbids.

The emission matrix is now DERIVED (9 wrappers x 3 aggregates x 8 operand positions), each cell
asserted for a verdict — the hand-written lists it replaces excluded the whole family above.

Verified, 247-cell derived matrix EXECUTED on ES 8.18.3, control vs branch:
  runs with the filter applied        74 -> 111
  silently dropped (200, wrong)      149 ->   0
  HTTP 400 at the user                16 ->   0
  `Internal parser error` label        3 ->   0
  refused by name                      5 -> 136

13/13 mutations RED · GrammarDiffProbe 18,062 control inputs, 0 differing · sql 1618 / core 1141 /
bridge 231 / es6bridge 227 · `+ compile` both legs · the exact CI lint line · es8java
*GroupByCompleteness* 37/37 on real ES 8.18.3. es6/es7/es9 NOT re-run after round 2.
Cost: parse+emit +1.69% inside a +1.28% must-be-zero floor; `validateResolved` costs +2.6 us on the
one `group by + having` shape and ±0.1 us on the other sixteen.

Closed Issue #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… or refused (#389 round 3)

Round 2 honoured `HAVING <function>(<group key>)` by pushing the predicate into the
QUERY, on the argument that a function of the key is constant within a bucket. That
argument holds only when the key is single-valued AND the bucket key is the column
itself, and both premises fail in the shipped grammar. Two HIGH silent wrong answers,
both reproduced on real ES 8.18.3:

  * a MULTI-VALUED key filtered on `doc['tags'].value` reads the FIRST value only, so a
    document is kept by one of its values and every OTHER value still forms its own
    bucket -- bucket `b` survived a filter naming only `A`;
  * `GROUP BY DAY(d) HAVING YEAR(d) = 2025` removes DOCUMENTS from buckets it does not
    remove, so a SURVIVING bucket's `doc_count` and every metric in it change.

A document filter is not a bucket filter. The push-down (`pushedHavingKeys`,
`searchCriteria`) is DELETED and both bridges' query seam is byte-identical to
`455433ae` again.

What replaces it is ONE classifier, `SingleSearch.keyPredicateOutcome`, with three
outcomes -- `TermsFilter`, `Refused(reason)` and a `ScriptedTerms` that is specified but
NOT BUILT, so its population is refused by name rather than silently dropped. Both
`validate()` and `validateResolved()` read that one value, and `keyExpressibleByTerms`
CALLS the real `Expression.includes` / `excludes` instead of re-deriving what the terms
filter can spell, so the classifier and the emission cannot disagree by construction.

Also in this scope:

  * `Having.script` is NOT dead code -- it has a production caller in
    softclient4es-extensions (`graph/Stage.scala:291`), which builds a materialized
    view's `bucket_selector` from it. Restored `@deprecated`, now gated on
    `unrepresentable`, which is made public so that caller can read the REASON. The
    extensions-side repair needs a published core and is NOT made here.
  * the Unscoped rule is UNGATED: a HAVING over a column that is neither an aggregate
    nor a GROUP BY key is refused whether or not the statement groups.
  * the OR-homogeneity rule is re-derived over leaves that INCLUDE relation leaves, and
    its dead `unsupported-key` answer -- unreachable because the key rule refuses that
    population first -- is deleted rather than kept beside a live copy.
  * the refusal reason reaches the caller at every venue: `search`, `searchAsync`,
    `searchWithInnerHits` and `ScrollApi`.
  * `havingLeaves` is a `lazy val`; four validation rules walk it.

Measured: 8/8 mutations RED; sql 1620 / core 1141 / bridge 231 / es6bridge 227; cross
compile 2.12.20 + 2.13.16; the exact CI lint line; GrammarDiffProbe over 18,062 corpus
inputs against the `455433ae` control with 0 differing rows; `GroupByCompletenessSpec`
38/38 on real ES 8.18.3 and 38/38 on ES 6.8.23. Cost isolates to one shape --
`group by + having` validate 0.40 -> 3.10 us, resolve 2.00 -> 6.10 us -- with parse and
emission showing overlapping ranges, i.e. no signal.

Closed Issue #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
)

Elasticsearch applies a `terms` filter through ONE list of kept values and ONE list
of removed values, and each is a UNION. So the kept list can express a DISJUNCTION and
the removed list a CONJUNCTION -- and a combination needing the other one is not
expressible at all. Three shapes were emitting a plausible-looking WRONG answer,
measured on `455433ae` and on real Elasticsearch:

  HAVING k <> 'a' OR  k <> 'b'   emitted exclude:["a","b"] -- the SAME filter as the
                                 AND, though the disjunction is true for EVERY group,
                                 so two groups were dropped. HTTP 200.
  HAVING k =  'a' AND k =  'b'   emitted include:["a","b"] -- two groups returned where
                                 the SQL means none.
  HAVING k =  'a' OR  k <> 'b'   emitted include:["a"] AND exclude:["b"] -- a
                                 conjunction where the SQL says disjunction.

All three are now refused by name. What already worked is untouched, byte for byte:
`= a OR = b`, `<> a AND <> b` (including `NOT IN`, and chains of three), and one kept
list beside one removed list.

🔴 `Criteria.excludes` IS `Criteria.includes(bucket, !not, …)` -- one method read in two
senses -- so `SingleSearch.keyChannelConflict` is ONE derivation that asks BOTH methods
about BOTH sides of every predicate, per bucket, with the polarity the emission itself
threads. A rule about these two channels that consults only one of them is wrong by
construction. The polarity is now named, `Predicate.includePolarityOfRight`, and the
emission and the rule call the same member -- the repo's own `PainlessNullSurvivalSpec`
source gate reddened on the first draft, which is exactly what it exists for.

🔴 No De Morgan arm, and that is MEASURED rather than assumed: `Predicate.maybeNot`
negates the RIGHT OPERAND only and `NOT ( … )` around a group is rejected by the grammar
(here and on `455433ae`), so no `Predicate` is ever reached with a flipped polarity. An
arm dualising the operator would be unreachable for every possible input -- the trap
this issue already sprang once.

Falsification, because the reason the defect class was invisible is the POPULATION and
not the assertions: no test in this branch or in the repository combined two
EXCLUDE-side key predicates under AND or OR, though `excludes` is literally the same
method. The fix ships a DERIVED matrix -- `=` / `IN` / `LIKE` against `<>` / `NOT IN` /
`NOT LIKE`, times AND / OR, same-channel and mixed, each verdict computed from the
principle rather than listed -- at three levels: `sql` validation, both bridges'
emission, and the testkit EXECUTED on a real index against bucket sets computed by hand
from the fixture.

Measured: 5/5 mutations RED, including "the rule reads only the include sense"; sql 1628
/ core 1141 / bridge 235 / es6bridge 231; cross compile 2.12.20 + 2.13.16; the exact CI
lint line; an emission diff over 20 shapes chosen to EXERCISE the two channels against
the `455433ae` control -- 6 moved, every one from a silent wrong answer to a refusal, 14
byte-identical; `GroupByCompletenessSpec` 44/44 on real ES 8.18.3 and 44/44 on ES
6.8.23.

⚠️ The 18,062-input corpus replay is deliberately NOT cited as the guard for this
surface: the corpus holds 4 HAVING rows and none with `<>`, `NOT IN`, `LIKE` or `OR`.

Closed Issue #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sted (#389)

The final review of `74b99e0f..e674014` found the commit's CODE sound and its own
DOCUMENTATION false in three places. Both defects it measured are fixed here, and each
false claim is corrected where it was written.

HIGH-2 -- a pattern silently replaces a value list on the same key. The emission matches
the regex arm FIRST and discards the values, and a second pattern is lost to `orElse`.
MEASURED on ES 8.18.3 over the buckets `a`, `b1`, `c`:

  HAVING status = 'a' OR status LIKE 'b%'  ->  {"terms":{"include":"b.*"}}  ->  ['b1']
                                               and the SQL means ['a','b1']

🔴 That shape is an OR of two KEPT-value contributions -- exactly the row this branch's
own documentation table blessed as supported. **Being a union is necessary, not
sufficient**: each list holds either a set of values or ONE pattern. A new arm refuses
any same-channel collision, with its own remedy: a single RLIKE, or splitting the query.

HIGH-1 -- an OR across TWO GROUP BY keys executed as an AND. `keyChannelConflict` is per
bucket, so each level only ever sees one contributing side, and `mechanismOf` collapsed
every grouping level to one label, so the OR rule saw a single mechanism and passed.
Elasticsearch NESTS the two `terms` aggregations, which IS the conjunction that rule
exists to catch. MEASURED over (a,a) (a,b) (x,b) (x,y):

  GROUP BY status, city HAVING status = 'a' OR city = 'b'  ->  ONE group (a,b),
                                                               and the SQL means THREE

A leaf's STAGE is now its mechanism plus, for a key predicate, the bucket it addresses,
and the two-level case has its own message -- "a group filter, a key filter and a nested
filter are separate stages" is not what two grouping levels are.

🔴 The suite PINNED HIGH-2 as supported, and the axis is why: the derived matrix paired
each operator with a FIXED partner and varied only channel and combiner, so `=` was
never crossed with `LIKE`. The operators are CROSSED now -- 72 cells, each verdict
derived from BOTH properties of the channel. Every defect found on this branch was a
hole in the population, never in the assertions.

Also: `PainlessNullSurvivalSpec`'s probe no longer saw the site it caught two rounds ago
(`p.includePolarityOfRight(not)` does not match `<name>.not`), so that file passed
vacuously; the named fold now counts as a use and as a classification, and a raw `p.not`
reddens it again.

Measured, re-derived rather than applied on trust: a generated cross-product of 366
statements -- every operator against every operator, AND and OR, one key and two, two
leaves and three -- each cell carrying its own oracle evaluated independently and
executed on ES 8.18.3 through the real emission: WRONG 64 -> 0. Cross-tabulated against
`455433ae`: 229 were wrong on main and are now refused, 91 are correct and byte-identical,
and 46 were correct on main by COINCIDENCE (23 repeat a leaf, 21 are contradictions, 2
genuinely distinct) and are now refused -- accepted knowingly, because their emission
cannot be told apart from the wrong one. 9/9 mutations RED; sql 1632 / core 1141 /
bridge 237 / es6bridge 233; cross compile 2.12.20 + 2.13.16; the exact CI lint line;
`GroupByCompletenessSpec` 46/46 on real ES 8.18.3 and 46/46 on ES 6.8.23.

⚠️ Corrects the previous commit's release note "everything that worked keeps working,
byte-identical": false, as the 46 above show. The 20-shape sample it rested on was not
representative of the moved population.

Closes #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The delta review of `893e2ce2` returned ready-to-push and disclosed five findings. Two
are fixed here.

F1 -- the terms include/exclude channel was a THIRD derivation of the LIKE -> regex
translation: `v.value.replaceAll("%", ".*")`. It neither translated `_` nor escaped a
regex metacharacter, while the query-DSL path uses the shared `toRegex` -- and
`metricSelector`'s own scaladoc asserts the shared one is used, so the code contradicted
its own documentation. MEASURED on ES 8.18.3 over the buckets `a.bZ`, `axbZ`, `ab`, `a1`:

  status LIKE 'a_'    WHERE -> [a1, ab]   HAVING -> NO BUCKETS
  status LIKE 'a.b%'  WHERE -> [a.bZ]     HAVING -> [a.bZ, axbZ]

Pre-existing and byte-identical to `455433ae`, so not a regression -- but a silent wrong
answer in the very channel this branch reasons about, which the standing rule says is
fixed in the branch. `RLIKE` is raw regex by definition and is untouched.

⚠️ This CHANGES EMISSIONS. Measured against the control over a LIKE population: five of
thirteen move, all of them in the `_`/metacharacter class (`a_` -> `a.`, `a.b%` ->
`a\.b.*`, `a+b%` -> `a\+b.*`, `a_b` -> `a.b`, `cat_3%` -> `cat.3.*`); `LIKE 'a%'`,
`LIKE '%a%'`, every `RLIKE` form, `=` and `<>` are byte-identical. 🔴 PINS THAT MOVED:
NONE -- no pin anywhere used a LIKE literal with `_` or a metacharacter on this channel,
which is exactly why the defect survived. New pins in both bridges and in `sql`, plus an
EXECUTED row: `HAVING category LIKE 'cat_0_'` selected nothing before (no category is
named `cat_0_`) and now selects cat_01 .. cat_09.

F4 -- two unpinned decisions in the previous commit. The equal-pattern exemption in
`collides` (`LIKE 'a%' OR LIKE 'a%'` is one pattern twice, so nothing is dropped) is live
and correct: PINNED, with its complement. The `e.nested ||` guard in `levelOf` is
DELETED: it made two leaves on different NESTED grouping levels collapse to one stage
while two flat levels are refused -- the very asymmetry that rule exists to remove -- and
it could not be pinned, because with no schema attached `GROUP BY e.name` under
`JOIN UNNEST` emits no `terms` at all on either tree. An unobservable guard in a symmetry
rule is worse than none.

F5/F6 -- a no-op `.replace` deleted from the matrix, and the matrix's `collides` model
documented as COARSER than production (it cannot express the equal-pattern exemption and
is sound over the 72 cells only because every crossed partner differs).

🔴 Also corrects a FALSE claim this branch shipped twice. The release note and §13.C
characterised the newly-refused population from a generated sample whose alphabet --
six operators over one keyword column -- contains no `IS NULL`, no range comparison, no
qualified column name and no empty-string literal. Both halves were wrong: `= 'a' AND
<> 'a'` is ACCEPTED (one include, one exclude), and the real set contains
`HAVING <key> IS NOT NULL`, `HAVING age > 0` on a grouping key, and `HAVING t.status =
'a'` with `GROUP BY t.status`. The release notes now NAME those families, and the spec
states the alphabet beside the population so the next reader can judge its reach.

Recorded, not fixed (ES 6.8.23 only, byte-identical to `455433ae`, correct on 8.18.3): a
pattern in one channel beside a MULTI-value list in the other is also inexpressible
there -- `LIKE 'a%' AND NOT IN ('b','x')` and two siblings answer HTTP 400. Wider than
the single-element residual already recorded.

Measured: the 366-statement cross-product re-run against its independent oracle after
the change -- WRONG 0; sql 1635 / core 1141 / bridge 238 / es6bridge 234; cross compile
2.12.20 + 2.13.16; the exact CI lint line; `GroupByCompletenessSpec` 47/47 on real ES
8.18.3 and 47/47 on ES 6.8.23.

Closes #389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@fupelaqu
fupelaqu marked this pull request as ready for review September 25, 2026 03:15
@fupelaqu
fupelaqu merged commit 04b1c69 into main Sep 25, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

HAVING <function>(<aggregate>) silently drops the entire HAVING clause — every group is returned

1 participant