Fix tokenizer handling of escaped identifiers - #2123
Open
deepakganesh78 wants to merge 1 commit into
Open
deepakganesh78 wants to merge 1 commit into
deepakganesh78 wants to merge 1 commit into
Conversation
Keep valid CSS escapes inside the surrounding word token so escaped punctuation in identifiers is not split into separate tokens. Fixes postcss#1349 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Member
|
De-ASI-INTERFACE
approved these changes
Aug 16, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1349
Reproduction
Before this change, tokenizing the issue examples split one escaped identifier into many
wordtokens:Current
mainproducedC,\(,\#,0280ae,\)as separate words instead of one identifier token.Root cause
The word scanner stopped at every backslash because
RE_WORD_ENDtreats\as a word boundary. The dedicated backslash branch consumed only the escape sequence itself, so escaped punctuation inside identifiers could never stay attached to the surrounding identifier text.Fix
Reuse the existing escape-consumption rules while scanning word endings. When the scanner reaches a valid CSS escape, it skips over that escape and continues the same word token; invalid escapes keep the previous token boundaries. This also keeps consecutive hexadecimal escapes in the same identifier token.
Compatibility
No new dependencies. Parsing/stringification behavior is unchanged for normal CSS; the tokenizer now matches CSS identifier semantics for valid escaped punctuation.
Validation
node -r ts-node/register/transpile-only test/tokenize.test.js: 34/34 passed.lib/tokenize.jsreverted: failed as expected (keeps escaped punctuation in identifiers).pnpm run test:coverage: 682/682 unit tests passed.pnpm run test:lint: passed with 0 errors and one existing warning intest/visitor.test.ts.pnpm run test:types: passed.pnpm run test:version: passed.pnpm run test:integration: 33 real-world CSS fixtures passed.pnpm run test:size: passed, 16.37 kB / 16.5 kB.Note: on Windows, raw
pnpm testfails before running because the package script uses POSIXFORCE_COLOR=1; I ran the equivalenttest:*scripts individually with$env:FORCE_COLOR=1.