Skip to content

PySequence_Fast needs new macros to be safe in a nogil world #119247

Description

@MojoVampire

Feature or enhancement

Proposal:

Right now, most uses of PySequence_Fast are invalid in a nogil context when it is passed an existing list; PySequence_FAST_ITEMS returns a reference to the internal array of PyObject*s that can be resized at any time if other threads add or delete items, PySequence_FAST_GET_SIZE similarly reports a size that is invalid an instant after it's reported. Similarly, if individual items are replaced without changing size, you'd have similar issues.

But when the argument passed is a tuple (incref-ed and returned unchanged, but safe due to immutability) or any non-list type (converted to new list) no lock is needed. Per conversation with Dino, going to create macros, to be called after a call to PySequence_Fast, to conditionally lock and unlock the original list when applicable, while avoiding locks in all other cases, before any other PySequence* APIs are used.

Preliminary (subject to bike-shedding) macro names are:

Py_BEGIN_CRITICAL_SECTION_SEQUENCE_FAST
Py_END_CRITICAL_SECTION_SEQUENCE_FAST

both defined in pycore_critical_section.h.

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

Discussion occurred with Dino during CPython core sprints.

Linked PRs

Activity

  1. self-assigned this
    on May 20, 2024
  2. added a commit that references this issue on May 22, 2024
  3. added a commit that references this issue on May 22, 2024
  4. MojoVampire commented on May 22, 2024

    @MojoVampire
    ContributorAuthor

    Hey @DinoV and @colesbury: Now that the base patch is merged, should I:

    1. Close this as complete and open new issue(s) to apply similar fixes to other uses of PySequence_Fast* APIs?
    2. Reuse this issue for similar changes?
    3. Skip working on this task specifically, and instead make more core Python objects work free-threaded (e.g. I started working on bytearray to make its use of PySequence_Fast* stuff safe, but it's clearly not free-threaded safe in general yet, so maybe I just work on making bytearray as a whole free-threaded safe?).
  5. colesbury commented on May 22, 2024

    @colesbury
    Contributor

    Hi Josh - thanks for fixing str.join. We typically reuse the same issue for related changes. As to whether to continue with PySequence_Fast* or work other changes like bytearray, that's up to you -- both sound really useful.

  6. added a commit that references this issue on May 22, 2024
  7. eendebakpt commented on May 24, 2024

    @eendebakpt
    Contributor

    @MojoVampire @colesbury In the _json module the PySequence_Fast API is used in a way that is not thread-safe. I am working on a PR, but could use some advice. The new macro Py_BEGIN_CRITICAL_SECTION_SEQUENCE_FAST starts with a { (I suspect to avoid name clashes), but that makes it hard to use when the critical section used return statements or goto statements.

    In #119438 I added an extra macro Py_RETURN_CRITICAL_SECTION_SEQUENCE_FAST to address this, but I am not convinced this is a nice solution. Another option would be to refactor the code in question to eliminate the goto and return statements inside the critical part, but that will result in quite some churn of the code (and perhaps less readable code)

  8. added a commit that references this issue on Jul 17, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions