Two findings of the same family on the UNION ALL execution paths, both measured, neither a
blocker. Introduced or left behind by story 22.6 (PR for #353).
1. The licensed cap reports truncation when the budget lands on a leg boundary and every remaining leg is EMPTY
Introduced by 65e2ed55 (deliberately — see "Why it is this way" below).
CoreDqlExtension.cappedUnionAllRows detects truncation in two places: inside a leg (each leg is
asked for one row more than it may contribute) and AT a leg boundary (the budget is spent and a leg
is dropped unexecuted). The boundary bit is a GUESS: dropping a leg is treated as truncation without
asking whether that leg had any rows.
When every remaining leg is empty, nothing was actually cut and the statement still reports a cap.
Measured (real Elasticsearch 8.18)
| statement |
quota |
fixture |
rows |
truncated |
warning |
cap-hits |
SELECT a FROM t1 LIMIT 10 UNION ALL SELECT a FROM t2 |
10 |
t1 = 100 docs, t2 empty |
10 |
true ❌ |
non-empty ❌ |
1 ❌ |
| 3 legs of (5, 5, 0) rows |
10 |
last leg empty |
10 |
true ❌ |
non-empty ❌ |
1 ❌ |
| 2 legs of (5, 0) rows |
5 |
last leg empty |
5 |
true ❌ |
non-empty ❌ |
1 ❌ |
The false positive is byte-identical in outcome to a genuine cut, so a consumer cannot tell them
apart, and the warning advises adding a LIMIT the analyst has often already added.
The trigger is more reachable than "an empty index": a first branch whose LIMIT equals the quota,
or legs summing to exactly the quota, followed by a branch that matches nothing.
The over-report is NARROW, and that was verified rather than assumed
| shape |
quota |
result |
| legs (0, 5) |
5 |
truncated = false, 0 cap-hits ✅ |
| legs (5, 0, 5) |
5 |
truncated = true, 1 cap-hit — and genuinely truncated ✅ |
| every leg empty |
5 |
truncated = false, 0 cap-hits ✅ |
So the over-report happens only when the budget is spent exactly at a boundary AND every remaining
leg is empty.
Why it is this way
The previous shape of this code had a per-leg probe against a GLOBAL budget, which UNDER-reported:
when a leg returned exactly the remaining budget there was no probe row, later legs were skipped,
and a genuinely truncated statement reported truncated = false, an empty warning and zero
cap-hits — the flag and the meter agreeing with each other and both wrong, which nothing detects.
For a Community licence (maxQueryResults = 10000) that was any UNION ALL whose first leg held
10,000 documents. Trading that for a narrow over-report was a deliberate decision.
Suggested fix
Turn the boundary bit from a guess into a fact: before reporting truncation, run a size: 0 /
count-only search on the remaining legs. No scroll is opened and no rows are materialised, so it
does not violate the property that a leg the budget cannot pay for is never scrolled — and it makes
the flag and the meter exact in both directions.
2. unionAllByLeg hands already-resolved legs to search, which resolves them again
Pre-existing on the un-capped per-leg route (SearchApi.scala, unionAllByLeg).
search(multiple) resolves the whole MultiSearch at the seam, then unionAllByLeg iterates the
RESOLVED legs and calls search(leg) on each — and search(single) resolves again.
resolveWithSchema(single) is not a pure check: when a leg carries a WHERE subquery it runs
SubqueryResolver.resolve, which EXECUTES the inner statement against Elasticsearch, and there is
no cache. The same double execution was measured and fixed on the licensed cap path (2 inner
executions vs 1), but this route still has it.
The re-resolution is idempotent in outcome — phase one has already rewritten the subquery into
literals, so the second pass finds nothing to execute for THAT shape — but the schema lookup and the
walk are repeated per leg, and any future resolution step that is not idempotent inherits the
double call.
Suggested fix
Same as the cap path: execute the legs the seam already resolved, rather than resolving each one a
second time inside its own search.
Two findings of the same family on the
UNION ALLexecution paths, both measured, neither ablocker. Introduced or left behind by story 22.6 (PR for #353).
1. The licensed cap reports truncation when the budget lands on a leg boundary and every remaining leg is EMPTY
Introduced by
65e2ed55(deliberately — see "Why it is this way" below).CoreDqlExtension.cappedUnionAllRowsdetects truncation in two places: inside a leg (each leg isasked for one row more than it may contribute) and AT a leg boundary (the budget is spent and a leg
is dropped unexecuted). The boundary bit is a GUESS: dropping a leg is treated as truncation without
asking whether that leg had any rows.
When every remaining leg is empty, nothing was actually cut and the statement still reports a cap.
Measured (real Elasticsearch 8.18)
truncatedSELECT a FROM t1 LIMIT 10 UNION ALL SELECT a FROM t2The false positive is byte-identical in outcome to a genuine cut, so a consumer cannot tell them
apart, and the warning advises adding a
LIMITthe analyst has often already added.The trigger is more reachable than "an empty index": a first branch whose
LIMITequals the quota,or legs summing to exactly the quota, followed by a branch that matches nothing.
The over-report is NARROW, and that was verified rather than assumed
truncated = false, 0 cap-hits ✅truncated = true, 1 cap-hit — and genuinely truncated ✅truncated = false, 0 cap-hits ✅So the over-report happens only when the budget is spent exactly at a boundary AND every remaining
leg is empty.
Why it is this way
The previous shape of this code had a per-leg probe against a GLOBAL budget, which UNDER-reported:
when a leg returned exactly the remaining budget there was no probe row, later legs were skipped,
and a genuinely truncated statement reported
truncated = false, an empty warning and zerocap-hits — the flag and the meter agreeing with each other and both wrong, which nothing detects.
For a Community licence (
maxQueryResults = 10000) that was anyUNION ALLwhose first leg held10,000 documents. Trading that for a narrow over-report was a deliberate decision.
Suggested fix
Turn the boundary bit from a guess into a fact: before reporting truncation, run a
size: 0/count-only search on the remaining legs. No scroll is opened and no rows are materialised, so it
does not violate the property that a leg the budget cannot pay for is never scrolled — and it makes
the flag and the meter exact in both directions.
2.
unionAllByLeghands already-resolved legs tosearch, which resolves them againPre-existing on the un-capped per-leg route (
SearchApi.scala,unionAllByLeg).search(multiple)resolves the wholeMultiSearchat the seam, thenunionAllByLegiterates theRESOLVED legs and calls
search(leg)on each — andsearch(single)resolves again.resolveWithSchema(single)is not a pure check: when a leg carries a WHERE subquery it runsSubqueryResolver.resolve, which EXECUTES the inner statement against Elasticsearch, and there isno cache. The same double execution was measured and fixed on the licensed cap path (2 inner
executions vs 1), but this route still has it.
The re-resolution is idempotent in outcome — phase one has already rewritten the subquery into
literals, so the second pass finds nothing to execute for THAT shape — but the schema lookup and the
walk are repeated per leg, and any future resolution step that is not idempotent inherits the
double call.
Suggested fix
Same as the cap path: execute the legs the seam already resolved, rather than resolving each one a
second time inside its own
search.