Possible race condition resolving virtual display device ID after creation #1554
ObviouslyDanCodes
started this conversation in
General
Replies: 1 comment
|
libdisplaydevice is problematic and unreliable, so I don't recommend using it generally. That's why I collapsed the configuration in the webUI by default. Windows is unreliable for getting the newly created display name, so I have to poll it with a timeout. Usually it's fine but on some machines that timeout still isn't enough. If you're mainly using virtual display, disable advanced display devices settings. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
First off, thanks for Apollo. I was actually just trying to stream games from my PC to a Galaxy Tab while watching Netflix, and somehow ended up spending the evening digging through Apollo's source with Codex.
I think I may have isolated what appears to be a race condition in proc_t::execute() affecting virtual display selection.
From what I observed:
The configured output_name GUID is parsed correctly from sunshine.conf.
After the virtual display is created (\.\DISPLAY6 on my system), Apollo immediately performs a reverse lookup to resolve that GDI display name to its stable device ID.
That lookup can occur before Windows has finished enumerating the new display, causing map_display_name() to return an empty string.
That empty value then overwrites the valid output_name.
From there, libdisplaydevice intentionally interprets an empty device ID as "configure the primary device" for the current topology snapshot, so the display configuration is applied to a physical monitor instead of the newly created virtual display.
My logs consistently showed something similar to:
Virtual display created (DISPLAY6)
[]
Display configuration built with device_id = ""
[]
~69 ms later Windows enumeration can finally resolve DISPLAY6 to the correct GUID
By then the display configuration had already been constructed using the empty device ID.
One thing I'm less certain about
I don't think I've completely explained why the fallback ultimately targeted the physical monitor it did.
Before launching Apollo, my MSI ultrawide (DISPLAY2) was my primary display.
After the virtual display was created, the topology snapshot used during display configuration instead identified my OMEN (DISPLAY1) as primary, so that became the fallback target.
I'm not sure whether that's:
Windows temporarily reporting a different primary while rebuilding display topology,
an expected side effect of the display configuration sequence,
or another piece of the initialization process that I haven't fully traced.
So I don't want to overstate that part of the diagnosis.
I implemented and tested a local fix with Codex that:
retries the DISPLAYx -> device_id lookup using a short bounded exponential backoff (20/40/80/160/320 ms),
only replaces output_name after a successful non-empty lookup,
skips display-device configuration if no stable device ID can be resolved,
leaves proc.display_name unchanged so capture continues to work normally.
That resolved the issue on my machine.
I also added regression tests covering:
delayed device visibility, and
permanent mapping failure.
Before I clean everything up into a PR, I wanted to ask:
Does this diagnosis line up with your understanding of this initialization path, or is there another reason the immediate reverse lookup is intentional?
If this aligns with how you'd like to address it, I'd be happy to clean up the patch and submit a pull request.
All reactions