Skip to content

Postgres: Support U&"..." quoted identifiers and UESCAPE clause - #2509

Open
BenSatori wants to merge 3 commits into
apache:mainfrom
BenSatori:postgres-unicode-parsing
Open

BenSatori wants to merge 3 commits into
apache:mainfrom
BenSatori:postgres-unicode-parsing

Conversation

@BenSatori

Copy link
Copy Markdown
Contributor

Example

Before this PR, none of these valid Postgres queries parsed:

SELECT 1 AS U&"d\0061ta", * from customers;
SELECT * from U&"c\0075stomers" limit 10;
SELECT * from U&"c!0075stomers" UESCAPE '!' limit 10;

U&"..." Unicode-escaped quoted identifiers were not tokenized at all (only the U&'...' string literal form was supported), and the UESCAPE '<char>' clause that can follow either a U&'...' string or a U&"..." identifier to override the default \ escape character was not implemented, even though the UESCAPE keyword already existed in keywords.rs.

Fix

U&"..." identifiers now tokenize into a regular quoted identifier (decoded, e.g. U&"d\0061ta" -> "data"), and both U&'...' and U&"..." now support an optional trailing UESCAPE '<char>' clause.

See the Postgres docs (identifiers) and UESCAPE docs (escape clause). Relates to #1354, which added the U&'...' string literal support this PR extends.

This PR was authored with the assistance of GitHub Copilot.

PostgreSQL supports Unicode-escaped quoted identifiers (U&"...") in
addition to Unicode string literals (U&'...'), and both accept an
optional trailing UESCAPE '<char>' clause to override the default
backslash escape character. Neither was previously supported by the
parser.

See https://www.postgresql.org/docs/current/sql-syntax-lexical.html#SQL-SYNTAX-IDENTIFIERS

Co-authored-by: GitHub Copilot <copilot@github.com>
@codecov-commenter

codecov-commenter commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 75.28090% with 22 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.97%. Comparing base (9ae00e7) to head (42546fa).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/tokenizer.rs 75.28% 13 Missing and 9 partials ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2509      +/-   ##
==========================================
+ Coverage   80.96%   80.97%   +0.01%     
==========================================
  Files          42       42              
  Lines       33359    33417      +58     
  Branches    33359    33417      +58     
==========================================
+ Hits        27009    27060      +51     
- Misses       2790     2791       +1     
- Partials     3560     3566       +6     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread src/tokenizer.rs
chars: &mut State,
) -> Result<String, TokenizerError> {
let error_loc = chars.location();
let raw = self.tokenize_single_quoted_string(chars, '\'', false)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You should collapse the doubled quotes when unescape is off

Suggested change
let raw = self.tokenize_single_quoted_string(chars, '\'', false)?;
let mut raw = self.tokenize_single_quoted_string(chars, '\'', false)?;
if !self.unescape {
raw = raw.replace("''", "'");
}

fn test_unicode_string_literal_uescape() {
// Custom escape character via UESCAPE, see the postgres docs example
pg_and_generic().expr_parses_to(r#"U&'d!0061t!+000061' UESCAPE '!'"#, "U&'data'");
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a red test relative to the other note.

Suggested change
}
}
#[test]
fn test_unicode_string_literal_no_unescape() {
TestedDialects::new_with_options(
vec![Box::new(PostgreSqlDialect {})],
sqlparser::parser::ParserOptions::new().with_unescape(false),
)
.verified_expr("U&'a''b'");
}

@LucaCappelletti94 LucaCappelletti94 added the waiting on contributor The review needs further refinements by its author label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

PostgreSQL waiting on contributor The review needs further refinements by its author

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants