Skip to content

Add SOCK_SEQPACKET support to the vsock device. - #8938

Open
beshleman wants to merge 12 commits into
cloud-hypervisor:mainfrom
beshleman:vsock-seqpacket
Open

beshleman wants to merge 12 commits into
cloud-hypervisor:mainfrom
beshleman:vsock-seqpacket

Conversation

@beshleman

@beshleman beshleman commented Sep 25, 2026 •

Copy link
Copy Markdown

Add SOCK_SEQPACKET support to the vsock device.

Today the vsock implementation and muxer do not support SOCK_SEQPACKET. We teach the vsock backend to accept VSOCK_TYPE_SEQPACKET and back it with an AF_UNIX/SOCK_SEQPACKET socket on the host. Message boundaries carry through end to end.

Some caveats and quirks to be aware of:

  1. AF_UNIX does not support EOR, so this implementation silently drops the flag. EOM is respected as usual.
  2. Its possible for users to try sending a message that exceeds its peer's credit limit. This is a problem because sockets do not release credits until the entire message has been delivered to the user, so if the message partially queues and credits exhaust, there are not enough credits to continue queuing, and so the the EOM flag is never received, creating a deadlock. In the Linux implementation, the socket layer avoids this by simply rejecting the sendmsg() with an error. On our host-to-guest path, the muxer is unable to deliver errors over the unix socket and we cannot know that the message is undeliverable until it has already been delivered to us. We respond to this situation by sending an RST to the guest and closing the connection with the host peer. Its imperfect, so definitely open to suggestions.

Best,
Bobby

Fixes: #8894

copy_from_slice() and copy_buf_from_slice() were gated on cfg(test)
since only the unit tests needed to fill in a packet buffer.

The seqpacket receive path needs to do the same outside of tests: a
host-side datagram is read into a local buffer and then copied into the
guest RX buffer, possibly across several RW packets. Un-gate both
helpers so they can be used by the device code.

No functional change.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
The seqpacket support added over the following patches exchanges whole
datagrams with the host socket, which the byte-stream Read/Write
interface cannot express. Those reads and writes go through recvmsg()
and sendmsg(), neither of which the vsock thread's seccomp filter
currently allows, so a seqpacket connection would take a SIGSYS on its
first transfer.

Allow both. Widening the filter ahead of its first caller keeps the
series bisectable at runtime, not just at build time.

The vsock thread already holds the host socket fd and already speaks to
it with recvfrom()/sendto(), so this grants no reach the thread did not
already have -- only a second way to exercise it.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
VsockConnection hardcodes VSOCK_TYPE_STREAM into every packet header it
builds, which is correct only while the stream type is the only one the
device supports.

Give the connection a sock_type variable, set at construction and
written into outgoing packets in place of the constant. This prepares
for future callers to pass in other types (e.g., for seqpacket).

No functional change: every caller passes VSOCK_TYPE_STREAM, which is
the value that used to be hardcoded. The seqpacket type that makes this
field useful arrives in a later patch.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
SEQPACKET connections must preserve message boundaries end to end, which
the byte-stream Read/Write interface the connection state machine is
built on cannot express: a write that happens to cover one message is
indistinguishable from part of a larger one.

Add a SeqPacketStream trait exposing datagram-based send and receive,
and implement it for UnixStream. std only offers UnixStream
(SOCK_STREAM) and UnixDatagram (SOCK_DGRAM), with no UnixSeqpacket and
no way to choose the socket type, so there is no ready-made type to
import... consequently, we have no choice but directly call
recvmsg()/sendmsg().

Record boundaries (MSG_EOR and vsock's VSOCK_SEQ_EOR flag) are not
implemented. The only host transport this backend has is
AF_UNIX/SOCK_SEQPACKET, and Linux's AF_UNIX/SOCK_SEQPACKET
implementation silently ignores MSG_EOR, the flag is never transmitted
to the unix socket peer. Message boundaries are unaffected.

Calls into the trait are introduced in subsequent patches in this
series.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Add the guest-to-host half of SOCK_SEQPACKET: take the RW packets the
guest sends and turn them back into the messages they came from.

A message may span several RW packets, so payloads are accumulated
until the packet carrying VSOCK_SEQ_EOM completes one, at which point
the whole thing is enqueued and written to the host socket as a single
SOCK_SEQPACKET datagram. Datagrams the socket will not take right away
stay queued and drain on EPOLLOUT.

Seqpacket handles OP_RW differently than streams, but credit accounting
and control ops are the same. For this reason, e replace tx_buf length
checks inside the control and credit paths with calls to the new helper
has_pending_tx() which works for both seqpacket and streams.

The number of vq bytes we can buffer is limited to CONN_TX_BUF_SIZE. If
an unterminated message (lacking VSOCK_SEQ_EOM) exceeds this, we reset
the connection.

When the guest passes us a message that exceeds our advertised
CONN_TX_BUF_SIZE, we reset the connection via sending an RST. This is to
avoid the connection from hanging infinitely. We can't release credits
until receiving an EOM, but the guest now has zero credits and is unable
to send an EOM. A well-behaving guest kernel should never send such a
message (Linux's implementation returns -EMSGSIZE to the sender), but
unfortunately the spec doesn't speak on this scenario. We try to give
the buggy guest a hint via the RST instead of hang.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Add the host-to-guest half of SOCK_SEQPACKET.

Seqpacket connections read whole datagrams into a per-connection buffer
and then doles it out in chunks in one or more RW packets. The final
chunk carries the VSOCK_SEQ_EOM flag. In the case that we have a
partially delivered datagram (one that has entirely been received by the
unix sender already, but not completely transferred to the guest), we
hook into the guest -> host path to detect that the peer has enough
credits and flip on PendingRx::Rw so that the remainder of the message
is later sent out without the need for another EPOLLIN from the unix
side.

A zero-length recvmsg() means either an empty datagram or EOF, and the
return value cannot tell them apart. To detect EOF, we watch for
EPOLLRDHUP.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
The muxer rejects any packet that isn't VSOCK_TYPE_STREAM, so a guest
can't open a SOCK_SEQPACKET connection. Back these with a host socket of
the same type, which preserves message boundaries end to end.

std only creates SOCK_STREAM sockets, so we build the seqpacket ones
ourselves with socket(2), bind(2) and connect(2).

Guest-initiated connections take the type from the request.
Host-initiated ones have no request to read it from, so we bind a second
listener at "<uds>_seqpacket" and use the type of whichever listener the
connection arrived on. Both listeners share the same accept path, since
nothing after the accept differs. The device unlinks the new socket file
on shutdown alongside the stream one.

Reading the "connect <port>" command differs by type. On a stream we
size the read so we don't consume bytes past the command. On seqpacket
the command is its own datagram, so we just read it whole.

The RSTs the muxer generates need a type too. The guest drops any packet
whose type doesn't match the target socket, so a seqpacket connection
reset with a stream RST would stay established on the guest side. We
carry the type through MuxerRx::RstPkt, taking it from the connection,
the offending packet, or the request. A packet of some other type still
gets an RST, with its own type echoed back.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Now that the device and the unix backend can carry message-oriented
connections, offer VIRTIO_VSOCK_F_SEQPACKET during feature negotiation.
The guest driver only allows SOCK_SEQPACKET sockets once this feature
has been negotiated.

Update test_virtio_device() to test for the new bit.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Cover the seqpacket data path in the state machine, against a mock host
stream so the cases stay independent of the host transport:

- message boundaries in both directions: payloads are only flushed to
  the host once VSOCK_SEQ_EOM arrives, and each host datagram comes back
  as RW packet(s) whose last one carries EOM;
- a datagram larger than the guest RX buffer is chunked across several
  RW packets, with only the final chunk delimiting the message;
- an orderly host close yields a full bidirectional shutdown, while a
  zero-length datagram without EPOLLRDHUP stays an empty message;
- with no peer credit, the receive path asks for a credit update once
  and resumes once the guest grants space, including mid-datagram;
- a datagram larger than the peer's window, or a peer that shrinks its
  window mid-message, resets rather than stalling;
- an unterminated message may not exceed the advertised window, but one
  that exactly fills it still completes;
- a datagram that cannot be written right away is queued, EPOLLOUT is
  requested, and it is flushed when the host becomes writable.

Also add the stream counterpart of the window test, since both paths now
enforce the same bound.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Assisted-by: Claude:Opus-5
Exercise the unix backend over real AF_UNIX/SOCK_SEQPACKET host sockets,
in both connection directions:

- guest-initiated: the guest connects to a host seqpacket listener at
  "<uds>_<port>"; two messages sent by the guest arrive as two distinct
  datagrams, and a host datagram comes back as an RW packet flagged EOM;
- host-initiated: the host connects to the "<uds>_seqpacket" listener,
  the "CONNECT <port>" handshake is answered with "OK <port>", and
  messages then flow in both directions with their boundaries intact;
- every RST the muxer emits for a seqpacket connection is itself typed
  seqpacket, whether it comes from a refused connect, an orphan packet,
  or a bulk reset; a packet of an unsupported type has that type echoed
  back instead.

Export SeqPacketStream from the csm module for the tests, so they can
send and receive datagrams on the host end.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Assisted-by: Claude:Opus-5
The integration tests only cover stream vsock, so nothing here exercises
the seqpacket paths against a real guest.

Add test_virtio_vsock_seqpacket, which boots a guest and drives both
directions over real AF_UNIX/SOCK_SEQPACKET sockets:

- host -> guest: we connect to the "_seqpacket" listener, do the
  "connect <port>" handshake, and send a message that the guest reads
  off a SOCK_SEQPACKET vsock socket.
- guest -> host: we listen at "<uds>_<port>" and the muxer connects out
  to us. Two guest sends arrive as two datagrams, so boundaries survive.

We also re-check the host -> guest direction after a reboot, and check
that stream connections still work on the same device.

The handshake has to be its own datagram. The muxer reads a whole
datagram for the command and drops the rest of it, so a payload tacked
onto the same send would be lost.

std has no SOCK_SEQPACKET type, so test_infra grows SeqpacketSocket and
SeqpacketListener on top of libc, the same way the unix backend does.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
Assisted-by: Claude:Opus-5
The page still said Cloud Hypervisor only supports stream vsock sockets.
Describe the seqpacket interface instead: the "_seqpacket" listener for
host-initiated connections, the shared per-port path for guest-initiated
ones, and the message size limit.

Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
@beshleman
beshleman requested a review from a team as a code owner September 25, 2026 16:44
@rbradford

Copy link
Copy Markdown
Member

@beshleman Thanks! Please take a moment to review the CONTRIBUTING.md and maybe revisit if you need quite so many commits?

@beshleman

Copy link
Copy Markdown
Author

@beshleman Thanks! Please take a moment to review the CONTRIBUTING.md and maybe revisit if you need quite so many commits?

@rbradford np! I'm thinking maybe squash each test with the corresponding patch under test, and same for the few preparation-style commits (e.g., the seccomp rules for sendmsg/recvmsg can squash with the commit that adds the sendmsg/recvmsg calls)?

Thanks, I'll review CONTRIBUTING.md and see what I overlooked.

@rbradford

Copy link
Copy Markdown
Member

@beshleman Thanks! Please take a moment to review the CONTRIBUTING.md and maybe revisit if you need quite so many commits?

@rbradford np! I'm thinking maybe squash each test with the corresponding patch under test, and same for the few preparation-style commits (e.g., the seccomp rules for sendmsg/recvmsg can squash with the commit that adds the sendmsg/recvmsg calls)?

Whatever you think best.

Thanks, I'll review CONTRIBUTING.md and see what I overlooked.

You will need to rustup override set nightly to get cargo fmt to do the right thing and you'll also need nightly for the fuzz.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SOCK_SEQPACKET support in AF_VSOCK

2 participants