Repository navigation
Tokenizer module does not handle backslash characters correctly #90432
Description
Activity
A source of one or more backslash-escaped newlines, and one final newline, is not tokenized the same as a source where those lines are "manually joined".
The source
\ \ \produces the tokens NEWLINE, ENDMARKER when piped to the tokenize module.
Whereas the source
produces the tokens NL, ENDMARKER.
What I expect is to receive only one NL token from both sources. As per the documentation "Two or more physical lines may be joined into logical lines using backslash characters" ... "A logical line that contains only spaces, tabs, formfeeds and possibly a comment, is ignored (i.e., no NEWLINE token is generated)"
And, because these logical lines are not being ignored, if there are spaces/tabs, INDENT and DEDENT tokens are also being unexpectedly produced.
The source
\produces the tokens INDENT, NEWLINE, DEDENT, ENDMARKER.
Whereas the source (with spaces)
produces the tokens NL, ENDMARKER.
- added3.7 (EOL)end of lifeend of life3.8 (EOL)end of lifeend of life3.9 (EOL)end of lifeend of life3.10 (EOL)end of lifeend of life3.11only security fixesonly security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jan 5, 2022 - changed the title
[-]backslash creating statement out of nothing[/-][+]Tokenizer module does not handle backslash characters correctly[/+]on Jan 6, 2022 - changed the title
[-]backslash creating statement out of nothing[/-][+]Tokenizer module does not handle backslash characters correctly[/+]on Jan 6, 2022 7 remaining items
@lysnikolaou it seems to fix that -- but also includes a regression(?) (this is 3.12.0b1):
$ diff -u <(python3.11 -mtokenize t.py) <(python3.12 -mtokenize t.py) --- /dev/fd/63 2023-05-24 08:44:07.429455783 -0400 +++ /dev/fd/62 2023-05-24 08:44:07.429455783 -0400 @@ -2,7 +2,7 @@ 1,0-1,6: NAME 'import' 1,7-1,10: NAME 'sys' 1,10-1,11: NEWLINE '\n' -3,0-3,1: NEWLINE '\n' +3,0-3,1: NL '\n' 4,0-4,3: NAME 'def' 4,4-4,5: NAME 'f' 4,5-4,6: OP '(' @@ -12,5 +12,5 @@ 5,0-5,4: INDENT ' ' 5,4-5,8: NAME 'pass' 5,8-5,9: NEWLINE '\n' -6,0-6,0: DEDENT '' +5,9-5,9: DEDENT '' 6,0-6,0: ENDMARKER ''
I think that the 3.12 version is the correct one regarding the
NEWLINE/NLtokens. Regarding the location of theDEDENTtoken, I don't have a strong opinion, but feel that it could stay this way. @pablogsal @mgmacias95 What do you think?This is now covered in the 3.12 what's new:
Aditionally, there may be some minor behavioral changes as a consecuence of the changes required to support PEP 701. Some of these changes include:
Some final DEDENT tokens are now emitted within the bounds of the input. This means that for a file containing 3 lines, the old version of the tokenizer returned a DEDENT token in line 4 whilst the new version returns the token in line 3.
The type attribute of the tokens emitted when tokenizing some invalid Python characters such as ! has changed from ERRORTOKEN to OP.Reacted by Lysandros NikolaouI think we can close this issue
Thanks @pablogsal!
Singe the latest fix, the only different now is the DEDENT:
❯ diff -u <( ../3.11/python.exe -mtokenize lol.py) <(./python.exe -mtokenize lol.py ) --- /dev/fd/11 2023-05-25 17:08:47 +++ /dev/fd/12 2023-05-25 17:08:47 @@ -12,5 +12,5 @@ 4,0-4,4: INDENT ' ' 4,4-4,8: NAME 'pass' 4,8-4,9: NEWLINE '\n' -5,0-5,0: DEDENT '' +4,9-4,9: DEDENT '' 5,0-5,0: ENDMARKER ''@pablogsal then this should be reopened
@pablogsal then this should be reopened
Why? That is an expected difference as I mentioned in my previous comment.
the NEWLINE token should be a NL if this issue is resolved
I'm confused -- I think we're talking about different things -- are you using the source here?
are you using the source here?
I am. The difference I was mentioning was between Python 3.11 and 3.12. Just to be clear, given
$cat source1.py import sys def f(): pass $ cat source2.py import sys \ def f(): passwith 3.12:
./python -m tokenize < source1.py 1,0-1,6: NAME 'import' 1,7-1,10: NAME 'sys' 1,10-1,11: NEWLINE '\n' 2,0-2,1: NL '\n' 3,0-3,3: NAME 'def' 3,4-3,5: NAME 'f' 3,5-3,6: OP '(' 3,6-3,7: OP ')' 3,7-3,8: OP ':' 3,8-3,9: NEWLINE '\n' 4,0-4,4: INDENT ' ' 4,4-4,8: NAME 'pass' 4,8-4,9: NEWLINE '\n' 4,9-4,9: DEDENT '' 5,0-5,0: ENDMARKER '' ./python -m tokenize < source2.py 1,0-1,6: NAME 'import' 1,7-1,10: NAME 'sys' 1,10-1,11: NEWLINE '\n' 3,0-3,1: NL '\n' 4,0-4,3: NAME 'def' 4,4-4,5: NAME 'f' 4,5-4,6: OP '(' 4,6-4,7: OP ')' 4,7-4,8: OP ':' 4,8-4,9: NEWLINE '\n' 5,0-5,4: INDENT ' ' 5,4-5,8: NAME 'pass' 5,8-5,9: NEWLINE '\n' 5,9-5,9: DEDENT '' 6,0-6,0: ENDMARKER ''With 3.11:
python -m tokenize < source1.py 1,0-1,6: NAME 'import' 1,7-1,10: NAME 'sys' 1,10-1,11: NEWLINE '\n' 2,0-2,1: NL '\n' 3,0-3,3: NAME 'def' 3,4-3,5: NAME 'f' 3,5-3,6: OP '(' 3,6-3,7: OP ')' 3,7-3,8: OP ':' 3,8-3,9: NEWLINE '\n' 4,0-4,4: INDENT ' ' 4,4-4,8: NAME 'pass' 4,8-4,9: NEWLINE '\n' 5,0-5,0: DEDENT '' 5,0-5,0: ENDMARKER '' python -m tokenize < source2.py 1,0-1,6: NAME 'import' 1,7-1,10: NAME 'sys' 1,10-1,11: NEWLINE '\n' 3,0-3,1: NEWLINE '\n' 4,0-4,3: NAME 'def' 4,4-4,5: NAME 'f' 4,5-4,6: OP '(' 4,6-4,7: OP ')' 4,7-4,8: OP ':' 4,8-4,9: NEWLINE '\n' 5,0-5,4: INDENT ' ' 5,4-5,8: NAME 'pass' 5,8-5,9: NEWLINE '\n' 6,0-6,0: DEDENT '' 6,0-6,0: ENDMARKER ''The difference between 3.11 and 3.12 for the "correct" version is still the location of the last DEDENT token, which is expected.
hmmm ok but this diff doesn't show the NEWLINE -> NL: #90432 (comment)
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: