Skip to content

vmgr leaks forwarded TCP sockets (state CLOSED) → "too many open files" after ~40 h; container DNS fails with "no memory" #2709

Description

@peatoe

Describe the bug

OrbStack Helper vmgr never closes some of the host TCP sockets it opens for container connections. They stay open in state CLOSED, 2.2–2.9 per minute for each Twingate connector container, until vmgr has used all 10,240 of its file descriptors. After that it cannot create sockets. DNS inside containers fails, published ports stop answering, and docker ps hangs. Restarting OrbStack recovers.

With two connectors running, the last two OrbStack runs failed 39 h 32 m and 39 h 46 m after vmgr started. Those times run from the first line of vmgr.log to its first no memory DNS failure.

vmgr.log once it happens:

level=warning msg="DNS query failed" error="no memory" name=<network>.twingate.com. type=A
level=error msg="UDP dial failed: dial udp 8.8.8.8:53: socket: too many open files"

The unified log at the same time:

OrbStack Helper[55394:3438e7] (libsystem_dnssd.dylib) dnssd_clientstub ConnectToServer: socket failed 24 Too many open files

dnssd_clientstub.c returns kDNSServiceErr_NoMemory when socket() fails. That is where the misleading "no memory" comes from.

At failure, lsof -p <vmgr> showed every fd from 0 to 10239 in use (10,240). Of those, 10,113 were TCP sockets (9,976 CLOSED, 129 ESTABLISHED, 8 LISTEN), 67 were pipes and 23 were regular files. The rest were a handful of kqueue, unix and other descriptors.

To Reproduce

  1. Run twingate/connector (1.92.0 or 1.93.0) in OrbStack 2.2.3 with a valid network token.
  2. Watch lsof -n -P -p $(pgrep -f "Helper[ ]vmgr" | head -1) -a -i TCP | grep -c CLOSED. The count rises steadily and never falls.
  3. With two connectors it hit the fd limit about 40 h after OrbStack started.

Leak rates measured on 2026-09-23 over 7–9 minute windows:

Setup Leaked sockets per minute
2 connectors, 1.92.0 4.4
1 connector, 1.92.0 2.3
1 connector, 1.93.0 2.9

At failure, 9,897 of the 9,976 CLOSED sockets (99.2%) connected to the connector's two HTTPS endpoints, both on :443: <network>.twingate.com (the controller) and relays-do.twingate.com (Cloudflare). The rest were 39 to a service on the Mac itself and 40 to three AWS addresses.

Lifecycle of a leaking connection. Host sockets were sampled every 3 s. Guest packets were captured with tcpdump in the connector's network namespace.

  1. The remote end sends FIN first, 54–60 s after the connection opened. The host socket goes ESTABLISHED → CLOSE_WAIT.
  2. The connector sends its own FIN later. Host sockets stayed in CLOSE_WAIT for 12 s to 3 min.
  3. That FIN is never acknowledged. The guest retransmits it 8 times over 57 s, with gaps of 0.23, 0.45, 0.91, 1.79, 3.58, 7.17, 14.34 and 28.67 s, and then gives up. No RST comes back on these connections.
  4. The host socket moves from CLOSE_WAIT to CLOSED, and vmgr never closes its fd.

Each connector produced two such connections per minute, one per endpoint. That matches the leak rate.

Synthetic clients did not reproduce it. Each of these leaked 0 sockets, using python in a container:

  • connect, then close at once: 1.1.1.1:443, 20×
  • read until the server closes, then close: 1.0.0.1:80 after an HTTP/1.0 request (40×), and a local server that sends 2 bytes and closes (20×)
  • after the remote's FIN, write 31 bytes, then close, either straight away or after 0.5 s. The socket was left half-closed for 0.5–75 s first. Targets: 1.0.0.1:80, Twingate's own servers, and a local server that had fully closed its socket, 10–20× each.
  • a silent client that the server drops after 5 s: Twingate's server, 10×

So the trigger involves something else about the connector's connections. Since connector 1.91, its HTTP backend has been libcurl.

Expected behavior

vmgr closes the host socket once both directions are finished or on error, and ACKs the guest's FIN.

Diagnostic report (REQUIRED)

Not attached yet. The two restarts rotated the log that covers the failure out of ~/.orbstack/log. I kept copies of that log, the full lsof of vmgr at failure, and the tcpdump output. I can send them, for example to logs@orbstack.dev, or upload orb report once the count is high again. Tell me which you prefer.

Screenshots and additional context (optional)

  • This is not the file-sharing fd leak from Too many open file descriptors #2347 / Too many open file descriptors #1253. Here, 10,113 of the 10,240 fds are TCP sockets.
  • OrbStack 2.2.3 (20963), macOS 27.0 (26A428), Mac mini M1 (Macmini9,1), 8 GB RAM, memory_mib 5120.
  • The host's default route is a full-tunnel WireGuard VPN (Private Internet Access). The leaked sockets' local address is the tunnel address.
  • The guest kernel logs Huh VM_FAULT_OOM leaked out to the #PF handler. Retrying PF from the first minute after boot: 106 times in the failing run. It is probably unrelated, because it starts long before the fds run out.
  • Workaround in use: a launchd job restarts OrbStack at 4 AM when vmgr holds more than 3,000 CLOSED sockets, and at any hour above 8,000.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions