Skip to content

fix(splitter): accept unreserved and column name keywords as CTE names - #837

Merged
psteinroe merged 2 commits into
mainfrom
fix/parser-unreserved-cte-names
Oct 6, 2026
Merged

psteinroe merged 2 commits into
mainfrom
fix/parser-unreserved-cte-names

Conversation

@psteinroe

@psteinroe psteinroe commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

The statement splitter required the CTE name to be a plain IDENT token. The lexer gives every keyword its own kind, so CTEs named constraints, indexes, comments, functions (or any other keyword) failed with Expected IDENT / Expected AS_KW / Expected L_PAREN, although Postgres accepts them. libpg_query was never involved.

Keyword categories in the lexer

The lexer codegen already reads kwlist.h but dropped each keyword's category. It now keeps it, and SyntaxKind::keyword_category() returns a KeywordCategory (Unreserved, ColName, TypeFuncName, Reserved). The splitter uses it to accept a ColId for the CTE name: an identifier (quoted or not), an unreserved keyword or a column name keyword, as common_table_expr does. Reserved and type/function name keywords (select, left) are still reported, as Postgres rejects them.

WITH recursive AS (...)

RECURSIVE is itself unreserved, and Postgres accepts WITH recursive AS (...) and WITH recursive(n) AS (...) as a non-recursive CTE named recursive. The splitter now only reads RECURSIVE as the modifier when it isn't followed by AS or (.

Decisions

  • The lexer keeps its vendored libpg_query 17-latest kwlist.h, matching the parser in use. Compared with REL_18_6, every shared keyword is in the same category. The only differences are four unreserved keywords that PG 18 adds (enforced, objects, period, virtual) and one it removes (recheck), so ColId gives the same answers either way.
  • The CTE name was the only place the splitter required an identifier. I found two related splitter issues and left them for follow-ups, since fixing them means changing the statement-boundary heuristics that error recovery relies on:
    • Unreserved statement keywords used as identifiers still split a statement, e.g. SELECT t.update FROM t, ... WHERE insert = 1, SELECT alter FROM t.
    • The CTE SEARCH / CYCLE clauses split the statement before the main SELECT.

Regression fixtures

The Postgres regression fixtures are re-recorded for 15 to 18. cluster.sql (WITH rows AS), with.sql (with ordinality as) and, in 18, numeric.sql (WITH rows AS) name a CTE after an unreserved keyword. The old splitter swallowed the rest of each of those files into one statement that Postgres rejected. Those statements now get their own verdicts. The typecheck snapshots only change their statement counts; the finding counts are unchanged.

All statements in the new tests run on PostgreSQL 18.6 (and 15). postgres-language-server check on the SQL from the issue now reports no errors.

Fixes #836

@psteinroe
psteinroe merged commit f8e4935 into main Oct 6, 2026
9 checks passed
@psteinroe
psteinroe deleted the fix/parser-unreserved-cte-names branch October 6, 2026 05:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

parser: unreserved keywords (constraints, indexes, comments, functions) rejected as CTE names

1 participant