Repository navigation
C API: Add PyObject_GenericHash() function #113024
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Dec 12, 2023 Just added Py_HashPointer() function can be used to implement the default Python object hashing function (object.hash()), but only in CPython. In other Python implementations the default object hash can depend not on the object address, but on its identity.
What is an object identity? The id() function? How is it computed? Is it different than the object address? Which Python implementations are you talking about?
If Py_HashPointer() should not be used to write "portable" C extensions, maybe only PyObject_GenericHash() should be used?
Reacted by Steve DowerIt is an entity that remains the same during the lifetime of the object and is unique for co-existing objects. The
isoperator is used to compare identities,a is bis true iffaandbis the same object.id()returns an integer representation of identity. In CPython it is the address of the object, but in other implementations (especially if the GC can move objects) it can be something other. For example in Jython it could not be an address. I don't know about PyPy.Py_HashPointer()can still be used for hashing pointers to C functions or other external pointers, which are not managed by the Python GC.We need also more portable C API for
id()(and the reverse function?), but I am not sure about name and interface yet.In PR #112449, it was proposed to replace calls to the private
_Py_HashDouble(obj, value)function with:Py_hash_t hash_double(PyObject *obj, double value) { if (Py_IS_NAN(value)) { return Py_HashPointer(obj); } else { return Py_HashDouble(value); } }
According to what you wrote, using
Py_HashPointer()here is not portable and a bad idea.PyObject_GenericHash()should be used instead.PyObject_GenericHash(obj)seems to be more complicated thathash(id(obj)): depending on howid()is implemented, a different hash function should be used:- CPython implements
id()by returning the object memory address.PyObject_GenericHash(PyObject *obj)is justPy_HashPointer(obj). - PyPy has a more complicated
id()implementation, and usingPy_HashPointer()onid()may not be the best hash funciton to reduce the risk of hash collision.
- CPython implements
IMO, in current API,
PyObject*is the identity a Python object. You can dereference it. You can compare it with C==.
Implementations where object identity is decoupled from the address can't really usePyObject*at all.So, it seems to me that this should be something like
PyObject_GenericHash(PyRef), with a HPy-like “reference” type. It's not something we have in current CPython.Or is there an implementation where
PyObject_GenericHash(PyObject)would be useful?So, it seems to me that this should be something like PyObject_GenericHash(PyRef), with a HPy-like “reference” type. It's not something we have in current CPython.
In HPy, a reference is just a
PyObject*in the "CPython build mode".Or is there an implementation where PyObject_GenericHash(PyObject) would be useful?
To replace the private
_Py_HashDouble()function, it looks useful if you care about other Python implementations than CPython, if I understood correctly.In HPy, a reference is just a PyObject* in the "CPython build mode".
No, it still has reference semantics. You cannot compare HPy references with
==, period. In non-debug “CPython build mode” it'll work, but that's an unfortunate implementation detail.Proposition for the C API WG: capi-workgroup/decisions#5.
Answering to @vstinner's suggestion on the PR (to not lose it in review comments). Note that it is not directly depending on
id(). I just do not know how to say it better than "only depends on the object's identity". Having the same identity is not the same as having the same id. Even if two objects in different time can have the same id, their hashes can be different. The implementation that generates a random hash and preserves it during the lifetime of the object will be good too.IMO you should be more specific and say that the hash value only depends on the memory address.
That's the thing, it's not. The address of the object can be changed during its lifetime in other Python implementations. This is why
PyObject_GenericHash()is introduce at first place.id()does not return the object's identity. It returns an integer that only depends on the object's identity and is unique for all co-existing objects.PyObject_GenericHash()also only depends on the object's identity, but without the uniqueness condition. There is no direct relation between them.- added 2 commits that reference this issue
on Apr 2, 2024
Feature or enhancement
Just added
Py_HashPointer()function can be used to implement the default Python object hashing function (object.__hash__()), but only in CPython. In other Python implementations the default object hash can depend not on the object address, but on its identity.I think that we need a new function,
PyObject_GenericHash()(similar toPyObject_GenericGetAttr()etc), that does not depend on CPython implementation details.Linked PRs