Skip to content

feat(sandbox): restore peer addresses on kernels without WAIT_KILLABLE_RECV - #4087

Open
russellb wants to merge 4 commits into
NVIDIA:mainfrom
russellb:fix/4058-legacy-accept-peer-address/russellb
Open

russellb wants to merge 4 commits into
NVIDIA:mainfrom
russellb:fix/4058-legacy-accept-peer-address/russellb

Conversation

@russellb

@russellb russellb commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

On kernels before Linux 5.19 the seccomp broker cannot write into workload memory, so it fails closed with EOPNOTSUPP on any mediated syscall whose result is an output buffer. accept(fd, &addr, &len) has exactly that shape, which makes server workloads unusable on RHEL 9.x and RHCOS. This restores them by rewriting the address-bearing call into a form the broker can satisfy, and fixes getpeername on directly connected sockets for every runtime.

Related Issue

Closes #4058

Changes

  • Add openshell-accept-shim, a freestanding preloadable object that rewrites an address-bearing accept into accept4(fd, NULL, NULL, flags) followed by getpeername on the accepted descriptor. The first call has no output buffer, so the broker injects the descriptor without touching workload memory; the second is answered with SECCOMP_USER_NOTIF_FLAG_CONTINUE, so the kernel writes the address into the caller's own buffer. The cross-process TOCTOU that WAIT_KILLABLE_RECV exists to close does not apply, because the workload writes its own memory.
  • Install the object at /run/openshell-compat only when the listener reports writes disabled, composing LD_PRELOAD so a workload's own value is preserved. Installation failure is non-fatal and logged as an OCSF config state change; the existing fail-closed behavior is retained.
  • Admit the object and its directory to Landlock inside prepare_child_sandbox. Landlock gates open independently of the file mode, so a world-readable object is not by itself reachable — the loader reports cannot open shared object file and silently drops the shim. Workload children are launched from two places (the sandbox entrypoint and sandbox exec through the boundary), each preparing Landlock with its own runtime paths; admitting at the single shared entry point means a launch path cannot set LD_PRELOAD without also letting the loader open what it points at.
  • Answer getpeername from the kernel for directly connected sockets, rather than substituting the socket the broker knows about.
  • Resolve the shim's C compiler through the cc crate so CC_<target>, TARGET_CC, and CC apply, along with the per-target wrappers cargo-zigbuild installs when release binaries are cross-compiled from a non-Linux host. Verify the emitted object's ELF machine against CARGO_CFG_TARGET_ARCH and fail the build on a mismatch — a misresolved compiler otherwise produces a valid object for the wrong architecture, which the loader reports identically to a policy denial.
  • Update the support matrix and OpenShift docs.

Scope and limits

This is a compatibility aid, not a security control. Nothing in OpenShell trusts its output: loopback and authorization decisions continue to use the kernel's peer address from the broker's own accept4. A workload that unsets LD_PRELOAD, links statically, or issues raw syscalls gains nothing it did not already have.

Coverage follows from LD_PRELOAD interposing symbols rather than syscalls. Bun and CPython are covered. Node.js does not need the shim — libuv passes a null address and resolves the peer lazily through getpeername. Go and statically linked binaries remain unsupported for address-bearing accept. getpeername is fixed for every runtime because the kernel answers it.

Testing

  • Checks appropriate to the affected code and behavior pass
  • Unit tests added/updated (if applicable)
  • E2E tests added/updated (if applicable)

mise run pre-commit passes.

Unit and integration tests run on Linux as a non-root user: 191 passed across openshell-sandbox and openshell-accept-shim. One pre-existing failure, canonical_tty_environment_replaces_supervisor_identity_defaults, is unrelated to this change — it fails in a bare container whenever the running uid has no /etc/passwd entry, on this branch and without it alike.

New coverage:

  • shim_behavior.rs asserts the object's ELF invariants (no DT_NEEDED, no TEXTREL, __errno_location as the sole undefined symbol, exactly two exported functions) and the rewrite's behavior.
  • preloaded_shim_remains_readable_after_landlock_for_non_root_workload forks, drops privileges, enforces Landlock, and confirms the loader can still open the object while an unadmitted neighbor stays denied.
  • runtime_read_only_composition_admits_caller_paths_and_the_shim pins the composition so a launch path cannot drop the admission.

End-to-end verification on affected hardware: e2e/rust/tests/peer_address.rs was run against an OpenShift cluster on RHCOS 9.6 (kernel 5.14.0-570.141.1.el9_6.x86_64, no WAIT_KILLABLE_RECV), via the Kubernetes compute driver with an image built from this branch. It passes. The same test fails on that cluster without the change, with errno 95 and ld.so: object '/run/openshell-compat/accept_shim.so' from LD_PRELOAD cannot be preloaded. The test asserts what the workload observes rather than how the sandbox arranges it, so it passes unchanged on both kernel generations.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture docs updated (if applicable)

…ed sockets

getpeername(2) was always answered by writing the sockaddr into the
workload's address space from the broker. On kernels without
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV the broker runs in LegacyReadOnly
mode, where that cross-process write fails closed with EOPNOTSUPP, so
getpeername failed for every socket.

For SocketState::Local and SocketState::AcceptedLocal the descriptor is
genuinely connected to the recorded peer, so the kernel's own answer is
identical to the broker's. Respond with SECCOMP_USER_NOTIF_FLAG_CONTINUE
for those states and let the kernel write the sockaddr in the workload's
own address space. That needs no cross-process task-memory write and
therefore works on kernels before 5.19.

SocketState::Connected keeps the existing substitution. Those descriptors
are connected to a loopback relay rather than to the destination the
workload requested, so the kernel's answer would leak the relay address.
A regression test pins that distinction.

This is the path libuv-based runtimes depend on: libuv passes a null peer
address to accept4(2) and resolves the peer lazily through getpeername(2)
when the application reads remoteAddress.

Refs NVIDIA#4058

Signed-off-by: Russell Bryant <rbryant@redhat.com>
…E_RECV

Kernels before Linux 5.19 lack SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV, so
the broker cannot hold a notified workload thread in a kill-only wait and
cannot safely write into workload memory. It fails closed with EOPNOTSUPP
on every mediated syscall whose result is an output buffer. accept(fd,
&addr, &len) has exactly that shape, so server workloads on RHEL 9.x and
RHCOS see EOPNOTSUPP where they expect a connection.

Add openshell-accept-shim, a freestanding preloadable library that rewrites
an address-bearing accept into accept4(fd, NULL, NULL, flags) followed by
getpeername on the accepted descriptor. The first call has no output buffer,
so the broker injects the descriptor without touching workload memory; the
second is answered with SECCOMP_USER_NOTIF_FLAG_CONTINUE, so the kernel
writes the address into the caller's own buffer. Because the workload writes
its own memory, the cross-process TOCTOU that WAIT_KILLABLE_RECV exists to
close does not apply.

The sandbox installs the object at /run/openshell-compat only when the
listener reports writes disabled, and composes LD_PRELOAD so a workload's
own value is preserved. Installation failure is non-fatal and logged as an
OCSF config state change; the sandbox keeps the existing fail-closed
behavior.

Landlock gates open independently of the file mode, so a world-readable
object is not by itself reachable: without an explicit admission the loader
reports "cannot open shared object file" and silently drops the shim.
Workload children are launched from more than one place -- the sandbox
entrypoint and sandbox exec through the boundary -- and each prepares
Landlock with its own runtime paths. Admit the object and its directory
inside prepare_child_sandbox, the single production entry point, so a
launch path cannot set LD_PRELOAD without also letting the loader open what
it points at.

Resolve the shim's C compiler through the cc crate so the standard
CC_<target>, TARGET_CC, and CC overrides apply, along with the per-target
wrappers cargo-zigbuild installs when release binaries are cross-compiled
from a non-Linux host. Verify the emitted object's ELF machine against
CARGO_CFG_TARGET_ARCH and fail the build on a mismatch: a misresolved
compiler otherwise produces a valid object for the wrong architecture,
which the loader reports as "cannot open shared object file" and is
indistinguishable from a policy denial.

This is a compatibility aid, not a security control. Nothing in OpenShell
trusts its output: loopback and authorization decisions continue to use the
kernel's peer address from the broker's own accept4. A workload that unsets
LD_PRELOAD, links statically, or issues raw syscalls gains nothing it did
not already have.

Coverage follows from LD_PRELOAD interposing symbols rather than syscalls:
Bun and CPython are covered, Node.js does not need it because libuv passes
a null address and resolves the peer lazily through getpeername, and Go and
statically linked binaries remain unsupported for address-bearing accept.
getpeername is fixed for every runtime because the kernel answers it.

The object is built freestanding with -nostdlib so one build per
architecture loads under both glibc and musl, carries no DT_NEEDED entry
and no text relocations, and leaves __errno_location as its sole undefined
symbol.

Add an e2e test that asserts what the workload observes rather than how the
sandbox arranges it: a listener accepts a connection from a client in the
same sandbox and must report the client's address, not EOPNOTSUPP and not
its own. It passes unchanged on both kernel generations.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Russell Bryant <rbryant@redhat.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(sandbox): support peer-address accept() for server workloads on pre-5.19 kernels (legacy read-only mode)

1 participant