Repository navigation
os.path.normcase() is inconsistent with Windows file system #86824
Description
Activity
On Windows file system, U+03A9 (Greek capital letter Omega) and U+2126 (Ohm sign) are distinguished. In fact, two distinct files "\u03A9.txt" and "\u2126.txt" can exist side by side in the same folder. But os.path.normcase() transforms both U+03A9 and U+2126 to U+03C9 (Greek small letter omega).
MSDN reads they use CompareStringOrdinal() to compare NTFS file names: https://docs.microsoft.com/en-us/windows/win32/intl/handling-sorting-in-your-applications#sort-strings-ordinally . This document also says "the function maps case using the operating system *uppercasing* table." But I made an experiment and found that at least in the Basic Multilingual Plane, "lowercase two strings by means of LCMapStringEx() and then wcscmp the two" always gives the same result as "compare the two strings with CompareStringOrdinal()". Though this fact is not explicitly mentioned in MSDN https://docs.microsoft.com/en-us/windows/win32/api/winnls/nf-winnls-lcmapstringex , the description of LCMAP_LINGUISTIC_CASING in this page implies that casing rules conform to file system's unless LCMAP_LINGUISTIC_CASING is used.
Therefore, I believe that os.path.normcase() should probably call LCMapStringEx(), with the first argument LOCALE_NAME_INVARIANT and the second argument LCMAP_LOWERCASE.
- added3.9 (EOL)end of lifeend of lifetype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Dec 16, 2020 "lowercase two strings by means of LCMapStringEx() and then wcscmp
the two" always gives the same result as "compare the two strings
with CompareStringOrdinal()"For checking case-insensitive equality, it shouldn't matter whether names are converted to uppercase or lowercase when using invariant non-linguistic casing. It's based on symmetric mappings between pairs of uppercase and lowercase codes, which avoids problems such as 'ϴ' (U+03F4) and 'Θ' (U+0398) both lowercasing as 'θ' (U+03B8), or 'ß' uppercasing as 'SS'.
That said, when sorting filenames, you need to use LCMAP_UPPERCASE in order to match the case-insensitive sort order of Windows. For example, 'Ÿ' (U+0178) is greater than 'Ŷ' (U+0176), but -- respectively lowercase -- 'ÿ' (U+00FF) is less than 'ŷ' (U+0177). In particular, if you have an NTFS directory with two files named 'ÿ' and 'ŷ', the listing will be ['ŷ', 'ÿ'] -- in uppercase order. (An NTFS directory is stored on disk as a b-tree sorted by uppercase filenames.)
For the implementation, _winapi.LCMapStringEx and related constants could be added.
- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory3.8 (EOL)end of lifeend of life3.10 (EOL)end of lifeend of life
on Mar 9, 2021 - added3.11only security fixesonly security fixes3.12only security fixesonly security fixesand removed3.9 (EOL)end of lifeend of life3.8 (EOL)end of lifeend of life
on Jun 6, 2022 Fix is committed for 3.12, but needs some help backporting. I can't get to it today, so if anyone wants to do it, go ahead. I'll try and get to it later this week.
I started backporting to 3.11, but the memory leak is back in the
test_embedtests. I don't see where it could be, and it doesn't occur in 3.12, so we need to fix it before backporting.I started backporting to 3.11, but the memory leak is back in the
test_embedtests. I don't see where it could be, and it doesn't occur in 3.12, so we need to fix it before backporting.The leak is caused by calling
PyUnicode_AsUnicodeAndSizewith empty string. SincePyUnicode_AsUnicodeAndSizeis removed in 3.12, should we fix it in former versions?Who's calling AsUnicodeAndSize? Is it inside one of the parameter converters? If so, we should definitely fix it in earlier versions.
I updated #93591 to use
PyUnicode_AsWideCharStringinstead, which avoids the leak. Presumably we're (at least potentially) leaking elsewhere, but it doesn't seem to have come up, so less urgent to track it down. I did look through, and it wasn't anything real obvious.Reacted by An LongThe leak is caused by calling
PyUnicode_AsUnicodeAndSizewith empty string.It is very strange. I see nothing suspicious in the code. Does it only leak for empty strings? Functions in the
osmodule use this converter to convert path arguments, so it should be easy to reproduce. Also it is used inwinreg.SetValue()and indirectly in other functions which use "u" or "Z" format units inPyArg_Parse*()functions.I didn't see anything suspicious either, so maybe it's just a false positive in the test. Either way, the code is gone and AC just hasn't been updated yet for null-embedded strings, so it doesn't really matter.
Closing?
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: