Found while validating story 21.5 Parts C1/C2 on a real index. Those parts fix SQLTypeUtils.coerce's arms correctly (tracked in #304), and the arms are unreachable for a column operand.
The defect
GenericIdentifier.baseType is col.map(_.dataType).getOrElse(SQLTypes.Any), and col is populated only inside update from request.schema. Grepped across sql / core / bridge: SingleSearch.update(Some(...)) has exactly one production call site — Table.mergeWithSearch (schema/package.scala), which uses it to infer a CTAS target table's columns and returns a Table, not a search to execute. SearchApi.resolveTemporalLiterals (#276) loads a schema at the SingleSearch -> ElasticQuery seam but rewrites WHERE literals only; it never attaches the schema to the statement.
So every identifier reaches coerce as SQLTypes.Any, and no cast over a column has ever emitted a conversion, for any source type.
Measured on real ES 8.18, after #304's fix
CAST(code AS BIGINT) over a keyword column still returns the string "125"
CAST(amount AS BIGINT) over a double column still returns 1.9
CAST(age AS DOUBLE) over an int column has always returned the raw number (visible in pre-existing testkit output as def param1 = ...; param1)
Pinned as a known limitation in GatewayApiIntegrationSpec ("should NOT yet convert a cast over a COLUMN") with a delete-me-when-fixed note. It pins a defect, not a contract — the test goes red the instant a schema reaches the execution path.
Why this is its own story rather than a fold-in
Attaching the schema at the execution seam is one line at the resolveTemporalLiterals site, and its blast radius is the whole scripting surface:
coerce's case SQLTypes.Any if !ctx.isProcessor branch treats every untyped identifier as a ZonedDateTime in Painless context and injects .toLocalDate() / .toLocalTime(). With a real schema that branch stops firing, so date scripting changes everywhere.
- Every widening arm (
Int -> BigInt, Int -> Double, ...) starts emitting casts it does not emit today, changing the generated script_fields for a large fraction of queries.
- The bridge's emitted JSON changes with it, so every downstream repo pinning generated JSON moves.
That needs its own measurement campaign — re-baseline every emitted-script pin and the bridge JSON, re-verify across all five clients — and its own release note.
Reproduction note worth keeping
The schema-less unit fixture cannot see this: with no schema every column emits no conversion, so "an identifier operand's Painless is just the raw doc value" looks like the whole story. Three independent measurements all built the schema by hand because the reproduction recipe said to. A recipe the production path does not follow reproduces something else.
Found while validating story 21.5 Parts C1/C2 on a real index. Those parts fix
SQLTypeUtils.coerce's arms correctly (tracked in #304), and the arms are unreachable for a column operand.The defect
GenericIdentifier.baseTypeiscol.map(_.dataType).getOrElse(SQLTypes.Any), andcolis populated only insideupdatefromrequest.schema. Grepped acrosssql/core/bridge:SingleSearch.update(Some(...))has exactly one production call site —Table.mergeWithSearch(schema/package.scala), which uses it to infer a CTAS target table's columns and returns aTable, not a search to execute.SearchApi.resolveTemporalLiterals(#276) loads a schema at theSingleSearch -> ElasticQueryseam but rewrites WHERE literals only; it never attaches the schema to the statement.So every identifier reaches
coerceasSQLTypes.Any, and no cast over a column has ever emitted a conversion, for any source type.Measured on real ES 8.18, after #304's fix
CAST(code AS BIGINT)over akeywordcolumn still returns the string"125"CAST(amount AS BIGINT)over adoublecolumn still returns1.9CAST(age AS DOUBLE)over anintcolumn has always returned the raw number (visible in pre-existing testkit output asdef param1 = ...; param1)Pinned as a known limitation in
GatewayApiIntegrationSpec("should NOT yet convert a cast over a COLUMN") with a delete-me-when-fixed note. It pins a defect, not a contract — the test goes red the instant a schema reaches the execution path.Why this is its own story rather than a fold-in
Attaching the schema at the execution seam is one line at the
resolveTemporalLiteralssite, and its blast radius is the whole scripting surface:coerce'scase SQLTypes.Any if !ctx.isProcessorbranch treats every untyped identifier as aZonedDateTimein Painless context and injects.toLocalDate()/.toLocalTime(). With a real schema that branch stops firing, so date scripting changes everywhere.Int -> BigInt,Int -> Double, ...) starts emitting casts it does not emit today, changing the generatedscript_fieldsfor a large fraction of queries.That needs its own measurement campaign — re-baseline every emitted-script pin and the bridge JSON, re-verify across all five clients — and its own release note.
Reproduction note worth keeping
The schema-less unit fixture cannot see this: with no schema every column emits no conversion, so "an identifier operand's Painless is just the raw doc value" looks like the whole story. Three independent measurements all built the schema by hand because the reproduction recipe said to. A recipe the production path does not follow reproduces something else.