Repository navigation
test_concurrent_futures.test_interpreter_pool failing #125716
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error3.14bugs and security fixesbugs and security fixes
on Oct 18, 2024 I've been looking into this for the past hour or so, and I wasn't able to reproduce a segfault, nor do any debuggers detect any foul play.
However, I was able to get this assertion to fail after sending a CTRL+C to unittest. That, and running
test_interpreter_poolwith a TSan build causes it to absolutely explode with errors, but I wouldn't be surprised if they were all false positives--TSan is far from perfect (do we even support it on GIL-ful builds?)Oh, it turns out this is a BUG. I triggered this while implementing the PR.
Reacted by Peter Biermaericsnowcurrently commented
on Oct 21, 2024 MemberAuthorMore actionsThe biggest clue is what the USAN buildbot tells us:
- a NULL pointer is being passed to
sem_wait() - the pointer points to uninitialized/deallocated memory
- the failure happens when the executor's worker context is initialized and calls
_interpqueues.create()
Python/thread_pthread.h:555:42: runtime error: null pointer passed as argument 1, which is declared to never be null /usr/include/semaphore.h:55:36: note: nonnull attribute specified here SUMMARY: UndefinedBehaviorSanitizer: undefined-behavior Python/thread_pthread.h:555:42 in Current thread 0x00007f89af7fe6c0 (most recent call first): File "Lib/concurrent/futures/interpreter.py", line 137 in initialize ==2260212==ERROR: UndefinedBehaviorSanitizer: SEGV on unknown address 0x03e800227cf4 ==2260212==The signal is caused by a READ memory access. #4 0x7f89b64e7504 in sem_wait (/usr/lib/libc.so.6+0x96504) (BuildId: 915eeec6439cfded1125deefc44a8d73e57873d9) #5 0x555a15c3e86b in PyThread_acquire_lock_timed Python/thread_pthread.h:555:33 #6 0x7f89b47308c6 in _queues_add Modules/_interpqueuesmodule.c:909:5 #7 0x7f89b47308c6 in queue_create Modules/_interpqueuesmodule.c:1103:19 #8 0x7f89b47308c6 in queuesmod_create Modules/_interpqueuesmodule.c:1487:19The pointer in question is the mutex created in
_globals_init(), which is called by the module exec function the first time the module is loaded. The mutex is cleared (in_globals_fini()) when the last copy of the module is cleared.There is a unlikely-but-possible race there in
_globals_init()with the module count, which may play a part here. It would make sense to do the following:- make the lock a
PyMutex - use atomic operations for the module count
That may improve the situation. However, it would be best if we could definitively determine why the mutex is NULL when it shouldn't be.
- a NULL pointer is being passed to
Something that could possibly be related: I've noticed an issue with exceptions inside subinterpreters created by
_interpretersthat causes them to finalize earlier than they should--depending on what's going on, that could explain theNULLmutex. I'm investigating to see if that's an issue on my end, or something deeper.Reacted by Eric Snow- added a commit that references this issue
on Oct 21, 2024 - added a commit that references this issue
on Oct 21, 2024 ericsnowcurrently commented
on Oct 22, 2024 MemberAuthorMore actionsIt looks like the latest fix has helped. I'll reopen this if there are any intermittent failures.
Reacted by Peter Bierma
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Bug report
Bug description:
I've seen 4 kinds of failure which I'm failure sure have the same cause:
WorkerContext.initialize()(line 137)The failures have happened in different test methods. Different failures have happened during the retry. Sometimes the retry passes. In all cases the architecture is AMD64, but across a variety of builders and non-Windows operating systems. The failures have all been on either refleaks buildbots or the USAN buildbot.
FWIW, it looks like
InterpreterPoolExecutorhas only exposed an underlying problem in the _interpqueues module, which means any fix would need to target 3.13 also (and maybe 3.12).Here are the buildbots where I've seen failures:
Here's the failure text:
segfault
hang 1
hang 2
test failed
USAN
CC @encukou
CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs