Skip to content

Keywords documentation page is hand-maintained and cannot express which words are reserved #347

Description

@fupelaqu

documentation/sql/keywords.md is hand-maintained, has drifted in both directions, and cannot say which words are reserved.

The page claims one thing and lists another

Its first line says "A list of reserved words recognized by the parser". It is not that list. It is a
grammar-keyword list wearing a reserved-word title, and it has always been one — so it is wrong in
both directions at once.

Measured on main, treating each bare keyword line of the page as one entry (145 entries), against
Parser.reservedKeywords:

26 reserved words the page does not list — a user reading the page will name a column after one of
these and get a parse error the page told them was safe:

AGAINST, ALL, BY, COLUMN, CROSS, EXCEPT, EXISTS, FALSE, FIRST, FORMAT_DATE, FORMAT_DATETIME, FULL, GROUP, INNER, LAST, LEFT, MATCH, ORDER, OUTER, PARSE_DATE, PARSE_DATETIME, PI, RIGHT, TO, TRUE, UNION

26 words the page lists that are recognised but NOT reserved — the page tells the user to rename a
column that never needed renaming:

CEILING, CONVERT, CURDATE, CURTIME, DATEADD, DATEDIFF, DATESUB, DATETIMEADD, DATETIMESUB, DATETIME_FORMAT, DATETIME_PARSE, DATE_FORMAT, DATE_PARSE, LCASE, OVER, POINT, POSITION, POWER, REGEXP, REGEXP_LIKE, REVERSE, RLIKE, SAFE_CAST, ST_DISTANCE, TRY_CAST, UCASE

It also lists only 145 entries where the engine recognises 331 distinct words — 186 recognised words
appear nowhere on it — and it has a block of 17 date/time functions commented out with [//]: #.

And "reserved / not reserved" turns out not to be the whole answer either. Four words are neither
reserved nor usable: CURDATE, CURTIME, NULL and RANDOM are shadowed by a zero-argument function
or a literal, so SELECT curdate FROM t silently returns the current date instead of your column. The
page tells the reader they are safe.

Why the distinction is now load-bearing

Story 22.2 made "recognised" and "reserved" two different answers a user has to be able to look up.
EXISTS is reserved, so SELECT exists FROM t is a parse error; ANY and SOME were deliberately
left unreserved
so a column named any keeps parsing. A one-dimensional list cannot express that, so
the page cannot be fixed by adding entries to it — the missing column is the defect.

Why hand-editing it again is not the fix

The page has drifted once per keyword-touching story for as long as it has existed, and nothing anywhere
reads it. Copying today's correct list into it restarts the same clock, and copying it a second time into
the web docs would start a second one.

Proposed fix

Generate the tables from the engine and guard the result:

  • one table, every word, with a per-word Reserved flag derived from Parser.reservedKeywords and the
    words derived from SQLKeywords;
  • the hand-written explanation stays hand-written — only the region between two markers is generated;
  • a test that fails when the committed page and the engine disagree, and whose failure message names the
    command that fixes it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions