Feature or enhancement
bytes.join() and bytearray.join() fill a Py_buffer and take a reference for every item, even when all items are exact bytes. With more than 10 items, that also means allocating the array of Py_buffers. On the free-threaded build the increfs are atomic when the items are shared between threads, and since gh-158803 they run while the list's lock is held.
When every item is exact bytes and the result is below the 1 MiB threshold at which join() releases the GIL, nothing can run Python code or suspend the critical section. So the items can be copied straight from the sequence, without a Py_buffer or a reference for each one. Other inputs keep the current path.
Release builds, b",".join(seq) with distinct 8-byte items:
|
FT main |
FT patched |
default main |
default patched |
| list of 10 |
170 ns |
92 ns |
138 ns |
79 ns |
| list of 1000 |
11.0 µs |
4.3 µs |
10.0 µs |
4.0 µs |
| 8 threads joining one shared list of 100 |
0.60 s |
0.18 s |
|
|
Lists with non-bytes items take the general path, which can be a bit slower depending on the compiler's inlining (see the discussion in #159096).
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
#158910, #159096
Linked PRs
Feature or enhancement
bytes.join()andbytearray.join()fill aPy_bufferand take a reference for every item, even when all items are exactbytes. With more than 10 items, that also means allocating the array ofPy_buffers. On the free-threaded build the increfs are atomic when the items are shared between threads, and since gh-158803 they run while the list's lock is held.When every item is exact
bytesand the result is below the 1 MiB threshold at whichjoin()releases the GIL, nothing can run Python code or suspend the critical section. So the items can be copied straight from the sequence, without aPy_bufferor a reference for each one. Other inputs keep the current path.Release builds,
b",".join(seq)with distinct 8-byte items:Lists with non-bytes items take the general path, which can be a bit slower depending on the compiler's inlining (see the discussion in #159096).
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
#158910, #159096
Linked PRs