Bug description
This is a fairly minor bug in tomllib: when a document ends with an unfinished backslash escape inside a double-quoted string or key, the exception reports that the error is past the end of the document.
import tomllib
document = '"\\' # A quote followed by a backslash.
try:
tomllib.loads(document)
except tomllib.TOMLDecodeError as exc:
print(len(exc.doc), exc.pos, exc.lineno, exc.colno)
# 2 3 1 4
# Doc length = 2
# Error position = 3 (1 past the EOF)
# Line number = 1 (correct)
# Column number = 4 (1 past EOF)
The document has two characters, but pos is 3. I would expect the position to point either to the backslash or to EOF, which would be index 2 and column 3. As a practical matter the formatted message already says "at end of document", so this only affects callers using the location attributes.
The same thing happens for a trailing backslash in a basic string value or a multiline basic string. An unterminated string without the backslash, such as '"abc', reports pos == len(doc) as expected.
The cause is parse_basic_str_escape advancing the position by two before checking the escape, even when only the backslash remains. That position is then used to construct TOMLDecodeError.
I found this while adding Hypothesis tests in gh-158565. The PR includes a focused expected-failure property and the minimized example above. This is related to but distinct from Tomli #307, which concerns the location reported for invalid Unicode escape digits.
CPython versions tested on
- Python 3.14.7.
- Main at
eb30e3d9d4b523e3f8de2cd1485c9ae9a996ab03, using a local Python 3.16.0a0 build to run that checkout's tomllib.
Operating systems tested on
Linux.
Bug description
This is a fairly minor bug in
tomllib: when a document ends with an unfinished backslash escape inside a double-quoted string or key, the exception reports that the error is past the end of the document.The document has two characters, but
posis 3. I would expect the position to point either to the backslash or to EOF, which would be index 2 and column 3. As a practical matter the formatted message already says "at end of document", so this only affects callers using the location attributes.The same thing happens for a trailing backslash in a basic string value or a multiline basic string. An unterminated string without the backslash, such as
'"abc', reportspos == len(doc)as expected.The cause is
parse_basic_str_escapeadvancing the position by two before checking the escape, even when only the backslash remains. That position is then used to constructTOMLDecodeError.I found this while adding Hypothesis tests in gh-158565. The PR includes a focused expected-failure property and the minimized example above. This is related to but distinct from Tomli #307, which concerns the location reported for invalid Unicode escape digits.
CPython versions tested on
eb30e3d9d4b523e3f8de2cd1485c9ae9a996ab03, using a local Python 3.16.0a0 build to run that checkout'stomllib.Operating systems tested on
Linux.