Repository navigation
fix(sql): a HAVING over a function of an aggregate must filter, or fail — never vanish (#389) - #397
Merged
Merged
Conversation
…il — never vanish `HAVING ABS(COUNT(*)) > 1` parsed, ran, and returned EVERY group with HTTP 200 — and so did `COALESCE`, `NULLIF`, `ROUND`, `GREATEST`, a `CASE`, an aggregate on the RIGHT of the comparison, and the whole-table form. The #205/#209/#253 silent-wrong-answer family. Cause, in one sentence: a function never looks inside its own ARGUMENTS. `FunctionChain .hasAggregation` reads the chain's OWN links, so `COUNT(*)` inside `ABS(...)` made the tree answer "no aggregate here" — no aggregation was created, `buckets_path` published nothing, and the selector fell to its `1 == 1` catch-all, which the caller reads as "nothing to filter". ONE derivation, `Identifier.bucketMetrics`, now feeds all three consumers that used to disagree: what `buckets_path` publishes, which aggregations get created, and the `bucket_selector` null guard. It needs TWO walks — `funIdentifiers` reaches a function argument and a `CASE`'s THEN result, while a `CASE`'s WHEN conditions are reachable only through `Case.conditionsOf`. Everything the engine can express is emitted; everything it cannot is REFUSED BY NAME at `SingleSearch.validate()`, inside `Parser.apply`, so every venue sees it. A partial emission counts as a wrong answer: one un-expressible conjunct refuses the whole statement. The representability gate decides on the RENDERING, never on a list of function names — a name list is a fourth derivation of what the renderer does, and it could not have expressed the rule that only EXECUTION produced: the bucket-pipeline Painless context does NOT compile `Double.valueOf(…)`, so `ABS`/`FLOOR`/`CEIL`/`SQRT`/`EXP`/`LOG`/`POWER` over an aggregate are refused, even though the same function works in WHERE, in the SELECT list and in ORDER BY. The spec asserted the opposite; eighteen shapes were rendered and executed against a real cluster to settle it. Fixed on the way: `HAVING COUNT(*) > ABS(MAX(x))` emitted a script reading `params.max_x` with no `max_x` aggregation anywhere; the nested form emitted NO aggregation at all and returned raw documents; an OR-disjunct was dropped as well as an AND-conjunct; and `WHERE ABS(COUNT(*)) > 1` now hits the existing aggregates-in-WHERE rejection. No bridge file changed — the whole fix is in `sql`, proved by holding the hand-maintained es6 twin to the same emission matrix. Verified: 19/19 mutations RED · sql 1607 / core 1141 / bridge 225 / es6bridge 221 · `+ compile` both Scala legs · GrammarDiffProbe over 18,165 inputs = 4 differing rows, every one intended · integration 32/32 on real Elasticsearch 8.18.3 and 32/32 on 6.8.23 · cost, 5 interleaved rounds vs a 455433a control: parse+emit −0.39 % against a must-be-zero baseline reading −0.08 % ⇒ no signal. Closed Issue #389 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… both operands (#389 round 2) Round 2 on an independent review that executed the branch against a 455433a control and a real Elasticsearch 8.18.3. LEAD RULING — `HAVING <function>(<group key>)` must WORK, not be refused. It does now. The equivalence it rests on was verified before building on it: a function of the GROUP BY key is CONSTANT within a bucket, so filtering the documents by it keeps exactly the satisfying buckets and changes nothing in a surviving one. Executed on real data, a surviving bucket's doc_count and every metric are byte-identical to the unfiltered run, including for a multi-level grouping. The predicate is carried by `SingleSearch.searchCriteria` and read by ONE line in each bridge; the AST and the render are untouched, because `MaterializedViewExtension` persists that render. What the equivalence does not license is refused by name: a predicate reading neither an aggregate nor a key, a NESTED key (a document filter keeps the whole parent), and an OR mixing a key with an aggregate — which closes a PRE-EXISTING silent wrong answer, measured on `main`: `HAVING COUNT(*) > 1 OR status = 'b'` returned NO buckets where the disjunction is two of them, because the key became a terms `include` and the metric a `bucket_selector`, so the OR executed as an AND. 🔴 The review also found that the representability gate was bypassed whenever the LEFT operand was a bare aggregate — so everything refused on the left was EMITTED on the right. Measured and executed: `> ROUND(MAX(a),2)`, `> NULLIF(MAX(a),0)` and `BETWEEN 1 AND ABS(MAX(a))` (which leaked RAW SQL into the Painless source) each returned HTTP 400, and `> CASE WHEN …` surfaced `Internal parser error` to the user. The gate now runs on both arms. That also CORRECTS this branch's own record: the `COUNT(*) > ABS(MAX(x))` row was documented as refused and was in fact running — the test for it asserted detection and never a verdict. Also closed: a metric compared as TEXT is refused rather than reaching Elasticsearch as a raw class-cast; a metric computed outside a nested grouping is refused rather than answering HTTP 200 with no grouping at all; the gate is re-asked at `validateResolved()` so a post-resolution divergence stays a 400 naming the clause; and the refusal reason now reaches the caller instead of the client's generic sentence. Two defects in this branch's own new code, both found by the mutation matrix reporting a refusal as unguarded: HAVING has FOUR mechanisms and scoping past the nested filter refused a shipped fixture; and `Identifier.bucket` is a stale copy reporting `nestedElement = None`, which made the nested refusal silent and pushed the one predicate the equivalence forbids. The emission matrix is now DERIVED (9 wrappers x 3 aggregates x 8 operand positions), each cell asserted for a verdict — the hand-written lists it replaces excluded the whole family above. Verified, 247-cell derived matrix EXECUTED on ES 8.18.3, control vs branch: runs with the filter applied 74 -> 111 silently dropped (200, wrong) 149 -> 0 HTTP 400 at the user 16 -> 0 `Internal parser error` label 3 -> 0 refused by name 5 -> 136 13/13 mutations RED · GrammarDiffProbe 18,062 control inputs, 0 differing · sql 1618 / core 1141 / bridge 231 / es6bridge 227 · `+ compile` both legs · the exact CI lint line · es8java *GroupByCompleteness* 37/37 on real ES 8.18.3. es6/es7/es9 NOT re-run after round 2. Cost: parse+emit +1.69% inside a +1.28% must-be-zero floor; `validateResolved` costs +2.6 us on the one `group by + having` shape and ±0.1 us on the other sixteen. Closed Issue #389 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… or refused (#389 round 3) Round 2 honoured `HAVING <function>(<group key>)` by pushing the predicate into the QUERY, on the argument that a function of the key is constant within a bucket. That argument holds only when the key is single-valued AND the bucket key is the column itself, and both premises fail in the shipped grammar. Two HIGH silent wrong answers, both reproduced on real ES 8.18.3: * a MULTI-VALUED key filtered on `doc['tags'].value` reads the FIRST value only, so a document is kept by one of its values and every OTHER value still forms its own bucket -- bucket `b` survived a filter naming only `A`; * `GROUP BY DAY(d) HAVING YEAR(d) = 2025` removes DOCUMENTS from buckets it does not remove, so a SURVIVING bucket's `doc_count` and every metric in it change. A document filter is not a bucket filter. The push-down (`pushedHavingKeys`, `searchCriteria`) is DELETED and both bridges' query seam is byte-identical to `455433ae` again. What replaces it is ONE classifier, `SingleSearch.keyPredicateOutcome`, with three outcomes -- `TermsFilter`, `Refused(reason)` and a `ScriptedTerms` that is specified but NOT BUILT, so its population is refused by name rather than silently dropped. Both `validate()` and `validateResolved()` read that one value, and `keyExpressibleByTerms` CALLS the real `Expression.includes` / `excludes` instead of re-deriving what the terms filter can spell, so the classifier and the emission cannot disagree by construction. Also in this scope: * `Having.script` is NOT dead code -- it has a production caller in softclient4es-extensions (`graph/Stage.scala:291`), which builds a materialized view's `bucket_selector` from it. Restored `@deprecated`, now gated on `unrepresentable`, which is made public so that caller can read the REASON. The extensions-side repair needs a published core and is NOT made here. * the Unscoped rule is UNGATED: a HAVING over a column that is neither an aggregate nor a GROUP BY key is refused whether or not the statement groups. * the OR-homogeneity rule is re-derived over leaves that INCLUDE relation leaves, and its dead `unsupported-key` answer -- unreachable because the key rule refuses that population first -- is deleted rather than kept beside a live copy. * the refusal reason reaches the caller at every venue: `search`, `searchAsync`, `searchWithInnerHits` and `ScrollApi`. * `havingLeaves` is a `lazy val`; four validation rules walk it. Measured: 8/8 mutations RED; sql 1620 / core 1141 / bridge 231 / es6bridge 227; cross compile 2.12.20 + 2.13.16; the exact CI lint line; GrammarDiffProbe over 18,062 corpus inputs against the `455433ae` control with 0 differing rows; `GroupByCompletenessSpec` 38/38 on real ES 8.18.3 and 38/38 on ES 6.8.23. Cost isolates to one shape -- `group by + having` validate 0.40 -> 3.10 us, resolve 2.00 -> 6.10 us -- with parse and emission showing overlapping ranges, i.e. no signal. Closed Issue #389 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
) Elasticsearch applies a `terms` filter through ONE list of kept values and ONE list of removed values, and each is a UNION. So the kept list can express a DISJUNCTION and the removed list a CONJUNCTION -- and a combination needing the other one is not expressible at all. Three shapes were emitting a plausible-looking WRONG answer, measured on `455433ae` and on real Elasticsearch: HAVING k <> 'a' OR k <> 'b' emitted exclude:["a","b"] -- the SAME filter as the AND, though the disjunction is true for EVERY group, so two groups were dropped. HTTP 200. HAVING k = 'a' AND k = 'b' emitted include:["a","b"] -- two groups returned where the SQL means none. HAVING k = 'a' OR k <> 'b' emitted include:["a"] AND exclude:["b"] -- a conjunction where the SQL says disjunction. All three are now refused by name. What already worked is untouched, byte for byte: `= a OR = b`, `<> a AND <> b` (including `NOT IN`, and chains of three), and one kept list beside one removed list. 🔴 `Criteria.excludes` IS `Criteria.includes(bucket, !not, …)` -- one method read in two senses -- so `SingleSearch.keyChannelConflict` is ONE derivation that asks BOTH methods about BOTH sides of every predicate, per bucket, with the polarity the emission itself threads. A rule about these two channels that consults only one of them is wrong by construction. The polarity is now named, `Predicate.includePolarityOfRight`, and the emission and the rule call the same member -- the repo's own `PainlessNullSurvivalSpec` source gate reddened on the first draft, which is exactly what it exists for. 🔴 No De Morgan arm, and that is MEASURED rather than assumed: `Predicate.maybeNot` negates the RIGHT OPERAND only and `NOT ( … )` around a group is rejected by the grammar (here and on `455433ae`), so no `Predicate` is ever reached with a flipped polarity. An arm dualising the operator would be unreachable for every possible input -- the trap this issue already sprang once. Falsification, because the reason the defect class was invisible is the POPULATION and not the assertions: no test in this branch or in the repository combined two EXCLUDE-side key predicates under AND or OR, though `excludes` is literally the same method. The fix ships a DERIVED matrix -- `=` / `IN` / `LIKE` against `<>` / `NOT IN` / `NOT LIKE`, times AND / OR, same-channel and mixed, each verdict computed from the principle rather than listed -- at three levels: `sql` validation, both bridges' emission, and the testkit EXECUTED on a real index against bucket sets computed by hand from the fixture. Measured: 5/5 mutations RED, including "the rule reads only the include sense"; sql 1628 / core 1141 / bridge 235 / es6bridge 231; cross compile 2.12.20 + 2.13.16; the exact CI lint line; an emission diff over 20 shapes chosen to EXERCISE the two channels against the `455433ae` control -- 6 moved, every one from a silent wrong answer to a refusal, 14 byte-identical; `GroupByCompletenessSpec` 44/44 on real ES 8.18.3 and 44/44 on ES 6.8.23.⚠️ The 18,062-input corpus replay is deliberately NOT cited as the guard for this surface: the corpus holds 4 HAVING rows and none with `<>`, `NOT IN`, `LIKE` or `OR`. Closed Issue #389 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sted (#389) The final review of `74b99e0f..e674014` found the commit's CODE sound and its own DOCUMENTATION false in three places. Both defects it measured are fixed here, and each false claim is corrected where it was written. HIGH-2 -- a pattern silently replaces a value list on the same key. The emission matches the regex arm FIRST and discards the values, and a second pattern is lost to `orElse`. MEASURED on ES 8.18.3 over the buckets `a`, `b1`, `c`: HAVING status = 'a' OR status LIKE 'b%' -> {"terms":{"include":"b.*"}} -> ['b1'] and the SQL means ['a','b1'] 🔴 That shape is an OR of two KEPT-value contributions -- exactly the row this branch's own documentation table blessed as supported. **Being a union is necessary, not sufficient**: each list holds either a set of values or ONE pattern. A new arm refuses any same-channel collision, with its own remedy: a single RLIKE, or splitting the query. HIGH-1 -- an OR across TWO GROUP BY keys executed as an AND. `keyChannelConflict` is per bucket, so each level only ever sees one contributing side, and `mechanismOf` collapsed every grouping level to one label, so the OR rule saw a single mechanism and passed. Elasticsearch NESTS the two `terms` aggregations, which IS the conjunction that rule exists to catch. MEASURED over (a,a) (a,b) (x,b) (x,y): GROUP BY status, city HAVING status = 'a' OR city = 'b' -> ONE group (a,b), and the SQL means THREE A leaf's STAGE is now its mechanism plus, for a key predicate, the bucket it addresses, and the two-level case has its own message -- "a group filter, a key filter and a nested filter are separate stages" is not what two grouping levels are. 🔴 The suite PINNED HIGH-2 as supported, and the axis is why: the derived matrix paired each operator with a FIXED partner and varied only channel and combiner, so `=` was never crossed with `LIKE`. The operators are CROSSED now -- 72 cells, each verdict derived from BOTH properties of the channel. Every defect found on this branch was a hole in the population, never in the assertions. Also: `PainlessNullSurvivalSpec`'s probe no longer saw the site it caught two rounds ago (`p.includePolarityOfRight(not)` does not match `<name>.not`), so that file passed vacuously; the named fold now counts as a use and as a classification, and a raw `p.not` reddens it again. Measured, re-derived rather than applied on trust: a generated cross-product of 366 statements -- every operator against every operator, AND and OR, one key and two, two leaves and three -- each cell carrying its own oracle evaluated independently and executed on ES 8.18.3 through the real emission: WRONG 64 -> 0. Cross-tabulated against `455433ae`: 229 were wrong on main and are now refused, 91 are correct and byte-identical, and 46 were correct on main by COINCIDENCE (23 repeat a leaf, 21 are contradictions, 2 genuinely distinct) and are now refused -- accepted knowingly, because their emission cannot be told apart from the wrong one. 9/9 mutations RED; sql 1632 / core 1141 / bridge 237 / es6bridge 233; cross compile 2.12.20 + 2.13.16; the exact CI lint line; `GroupByCompletenessSpec` 46/46 on real ES 8.18.3 and 46/46 on ES 6.8.23.⚠️ Corrects the previous commit's release note "everything that worked keeps working, byte-identical": false, as the 46 above show. The 20-shape sample it rested on was not representative of the moved population. Closes #389 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The delta review of `893e2ce2` returned ready-to-push and disclosed five findings. Two
are fixed here.
F1 -- the terms include/exclude channel was a THIRD derivation of the LIKE -> regex
translation: `v.value.replaceAll("%", ".*")`. It neither translated `_` nor escaped a
regex metacharacter, while the query-DSL path uses the shared `toRegex` -- and
`metricSelector`'s own scaladoc asserts the shared one is used, so the code contradicted
its own documentation. MEASURED on ES 8.18.3 over the buckets `a.bZ`, `axbZ`, `ab`, `a1`:
status LIKE 'a_' WHERE -> [a1, ab] HAVING -> NO BUCKETS
status LIKE 'a.b%' WHERE -> [a.bZ] HAVING -> [a.bZ, axbZ]
Pre-existing and byte-identical to `455433ae`, so not a regression -- but a silent wrong
answer in the very channel this branch reasons about, which the standing rule says is
fixed in the branch. `RLIKE` is raw regex by definition and is untouched.
⚠️ This CHANGES EMISSIONS. Measured against the control over a LIKE population: five of
thirteen move, all of them in the `_`/metacharacter class (`a_` -> `a.`, `a.b%` ->
`a\.b.*`, `a+b%` -> `a\+b.*`, `a_b` -> `a.b`, `cat_3%` -> `cat.3.*`); `LIKE 'a%'`,
`LIKE '%a%'`, every `RLIKE` form, `=` and `<>` are byte-identical. 🔴 PINS THAT MOVED:
NONE -- no pin anywhere used a LIKE literal with `_` or a metacharacter on this channel,
which is exactly why the defect survived. New pins in both bridges and in `sql`, plus an
EXECUTED row: `HAVING category LIKE 'cat_0_'` selected nothing before (no category is
named `cat_0_`) and now selects cat_01 .. cat_09.
F4 -- two unpinned decisions in the previous commit. The equal-pattern exemption in
`collides` (`LIKE 'a%' OR LIKE 'a%'` is one pattern twice, so nothing is dropped) is live
and correct: PINNED, with its complement. The `e.nested ||` guard in `levelOf` is
DELETED: it made two leaves on different NESTED grouping levels collapse to one stage
while two flat levels are refused -- the very asymmetry that rule exists to remove -- and
it could not be pinned, because with no schema attached `GROUP BY e.name` under
`JOIN UNNEST` emits no `terms` at all on either tree. An unobservable guard in a symmetry
rule is worse than none.
F5/F6 -- a no-op `.replace` deleted from the matrix, and the matrix's `collides` model
documented as COARSER than production (it cannot express the equal-pattern exemption and
is sound over the 72 cells only because every crossed partner differs).
🔴 Also corrects a FALSE claim this branch shipped twice. The release note and §13.C
characterised the newly-refused population from a generated sample whose alphabet --
six operators over one keyword column -- contains no `IS NULL`, no range comparison, no
qualified column name and no empty-string literal. Both halves were wrong: `= 'a' AND
<> 'a'` is ACCEPTED (one include, one exclude), and the real set contains
`HAVING <key> IS NOT NULL`, `HAVING age > 0` on a grouping key, and `HAVING t.status =
'a'` with `GROUP BY t.status`. The release notes now NAME those families, and the spec
states the alphabet beside the population so the next reader can judge its reach.
Recorded, not fixed (ES 6.8.23 only, byte-identical to `455433ae`, correct on 8.18.3): a
pattern in one channel beside a MULTI-value list in the other is also inexpressible
there -- `LIKE 'a%' AND NOT IN ('b','x')` and two siblings answer HTTP 400. Wider than
the single-element residual already recorded.
Measured: the 366-statement cross-product re-run against its independent oracle after
the change -- WRONG 0; sql 1635 / core 1141 / bridge 238 / es6bridge 234; cross compile
2.12.20 + 2.13.16; the exact CI lint line; `GroupByCompletenessSpec` 47/47 on real ES
8.18.3 and 47/47 on ES 6.8.23.
Closes #389
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
fupelaqu
marked this pull request as ready for review
September 25, 2026 03:15
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #389
A
HAVINGpredicate whose aggregate sits inside a function argument emitted nobucket_selector: the clause vanished and every group came back, HTTP 200. The #205 / #253 silent-wrong-answer family.The cause was not the one the issue named
#389 blamed
ElasticAggregation.metricSelectorForBucket'sif (fullScript.isEmpty) return "". Measured: that is the second line of defence. One rung up, insql,MetricSelectorScript.metricSelectorfalls to itscase _ => "1 == 1"catch-all becauseExpression.referencesBucketMetricisfalse— a function never looks inside its own arguments, soCOUNT(*)withinABS(...)/COALESCE(...)/ aCASEbranch is invisible toisAggregationandhasAggregation. That single fact explained the dropped clause, the missing auxiliary aggregation, and why the whole-table guard whose comment promises "nothing may fall through silently" did not fire.Three things the issue does not record, all measured on
main: an AND conjunct was silently dropped (a partial filter — the answer looks filtered and is wrong), an OR disjunct likewise, and the whole-tableHAVINGtook the same path.What ships
A function-wrapped aggregate in
HAVINGnow emits a correctbucket_selector, or the statement is refused by name. Never dropped, never partially applied.HAVING COALESCE(COUNT(*), 0) > 1andGREATEST(...)forms emit and run.NULLIF/ROUND/CASE/ boxing forms are refused, naming the clause and the reason. The gate is structural, on the rendered Painless — not a function-name list, which would be a fourth derivation of the renderer.Double.valueOf(x)compiles in one and not the other. Only execution settled that.termsinclude/exclude filter, or refused. The channel unions, it never intersects — anORacross two grouping levels is a conjunction, because Elasticsearch nests thetermsaggregations.Internal parser errorto the user.Having.scriptis restored (@deprecated): it has a production caller insoftclient4es-extensions(graph/Stage.scala:291), contrary to what an earlier round of this branch asserted.search,searchAsync,searchWithInnerHitsandScrollApi, not justsearch.Verification
sql1635 ·core1141 · bridge 238 · es6bridge 234 ·+ compile(2.12.20 + 2.13.16) · the exact CI lint line — all green. Integration re-run after the final change: 47/47 on real ES 8.18.3, 47/47 on real ES 6.8.23.Four independent fresh-context reviews, each over the commits the previous one had not seen, executing against hand-computed oracles on real clusters. The last measured 0 surviving wrong answers over its own independently generated 309-statement population.
Behaviour changes — release-note material
455433aeafter the merge (verdict AND emitted query):GROUP BY t.status HAVING t.status = 'a',GROUP BY t.status HAVING status = 'a',GROUP BY age HAVING age > 0,GROUP BY status HAVING status IS NOT NULLand the JOIN formGROUP BY d.category HAVING d.category = 'x'were all ACCEPTED onmainand emitted noincludeat all — the condition was silently dropped and every group was returned. That is HAVING <function>(<aggregate>) silently drops the entire HAVING clause — every group is returned #389 itself, in the group-key population; the refusal makes the same defect loud and nothing regressed. Only the UNQUALIFIED direct comparison (HAVING x = 'v',<>,IN,LIKE) was ever correct, and it is unchanged. The original wording ("correct only coincidentally") relayed a reviewer's classification that had not been measured. Aggregation over JOINs is unaffected —GROUP BY d.category HAVING COUNT(*) > 1,GROUP BY d.category, o.kind HAVING COUNT(*) > 1andWHERE age > 0 GROUP BY age HAVING COUNT(*) > 1are all accepted (measured on the merged tree). Residual, now filed as a follow-up: a qualified key should resolve to its bucket and does not — same family as SQL parser: support backtick-qualified catalog names natively #85 / Two tables differing only by qualifier collapse into one alias, silently dropping an alias #292.HAVING <key> LIKEnow uses the shared translation._becomes.and regex metacharacters are escaped, soLIKE 'a_'— which matched nothing — now matches, andLIKE 'cat_3%'emitscat.3.*rather thancat_3.*. No test pin anywhere used such a literal on this channel, which is how it survived.Recorded, not fixed — pre-existing and byte-identical to
mainHAVING status = 'a.b'returns bucketaxb. Wider than first recorded: a pattern in one channel beside a multi-value list in the other is HTTP 400 there (three further shapes).child(...)/nested(...)puts anincludeon the parent bucket key.HAVING k = 'a' OR child(LENGTH(k) = 1)executes as its first branch.GROUP BYin aJOIN UNNEST+ root-metric shape — a separate defect on a different mechanism. That statement is no longer answered 200-with-a-wrong-answer; it is refused.A capability that was built and then dropped
A case-fold inversion (
HAVING UPPER(key) = 'A'→ atermsinclude regex) was built across two rounds and removed by decision. It produced three silent wrong answers and a remotely-triggerableOutOfMemoryErrorfrom a ~30-character string literal. The structural reason it was dropped: it inverts a case fold performed by a JVM the client has never seen. Reading the client JVM made answers depend on the client's JDK (96 patterns differ between zulu-11 and zulu-21); freezing a table relocated the divergence to the node (ES 7.17 → Unicode 15, ES 8.18/9.0 → Unicode 16, versus a pinned Unicode 13 — 134 extra cased code points, with a missing bucket proved end-to-end). The work is preserved locally and the reasoning is recorded so it is not rebuilt the same way.One lesson worth carrying
Every defect found on this branch — across six dev rounds and four reviews, including two rounds where the fix shipped something worse than what it fixed — was a hole in the population under test, never in the assertions. The mutation matrices were sound each time. The failures were all absences — a bucket that is not there — which only an independently computed expected set can see.
🤖 Generated with Claude Code
Cost — measured on this tree (
1a6a23a1), reported, not asserted9 interleaved rounds per side against a
455433aecontrol in one session, with six must-be-zero baseline shapes (bare projection,WHERE-only, wideWHERE, plainGROUP BY, 2-keyGROUP BY,ORDER BY) carrying the noise floor. No timing assertion is added anywhere (#269/#270).Search — nothing to degrade. Of the 99 well-formed
HAVINGstatements present in the control tree's own sources, 99 emit byte-identical queries. Zero changed, zero newly refused, zero refused→emitted. The only queries that move are ones the control emitted wrong (abucket_selectorreadingparams.max_xwith nomax_xaggregation — that query fails at runtime) or did not emit at all (#389). Their extra ES work is one metric sub-aggregation plus a coordinating-node pipeline.RowCostProbewas deliberately not run: everycore/src/mainhunk here is in a failure branch (invalidSearchRequestand its 5 call sites,ScrollApi'sSource.failed), so nothing on the row-extraction, paging or response-parsing path is touched.Parse + validate (µs, median of 9 rounds × 500 runs; floor over the 6 baseline shapes = −4.6 … +1.7):
HAVING COALESCE(COUNT(*),0) > 1HAVING COUNT(*) > 1HAVING COUNT(*) > 1 AND MAX(a) < 9HAVING n > 1(alias)HAVING code = 'x'(GROUP BY key)Resolution columns, same run (floors:
resolve±0.4,validate±0.2): a metric-bearingGROUP BY … HAVINGresolves in 2.3 → 6.6 µs, andCOALESCE(COUNT(*),0)in 2.4 → 8.9 µs. A ratio of ~3×, ~4–6 µs in absolute terms against a ~270 µs parse. A GROUP-BY-key HAVING pays nothing (1.5 → 1.8 µs).Rendering / emission (fresh parse +
ElasticSearchRequest.query; floor ±1.3 µs onmin): metric-bearing HAVING costs +3.7 to +15.2 µs on a 329–348 µs parse+emit (≈1–4.5%), scaling with metric-leaf count — 1 leaf +6.0, 2 leaves +15.2. Key-only HAVING (=,IN,LIKE,<>) and every non-HAVING shape sit inside the floor.Mechanism (reasoning, corroborated by the leaf-count scaling):
MetricSelectorScript.selectornow callsrepresentable(e)per metric leaf — apainless(None)render plus aStringBuilderscan. That is the gate that converts four HTTP 400s and the silent drop into named refusals. It runs once per metric leaf at parse/emit time, never per document. Key-only leaves returnNoFilterbefore reaching it, which is why they measure flat.Two caveats stated rather than hidden: the branch ran first in every round (a fixed position, absorbed by the baseline floor — which is why the floor is reported), and the newly-refused statements are not counted as "slower" — a refusal short-circuits emission, and two of them measure faster.