Skip to content

Report a failed non-blocking connect through select and SO_ERROR (#8786) - #9661

Merged
headius merged 4 commits into
jruby:jruby-10.0from
aminmansuri:fix-8786-nonblocking-connect-select
Sep 11, 2026
Merged

headius merged 4 commits into
jruby:jruby-10.0from
aminmansuri:fix-8786-nonblocking-connect-select

Conversation

@aminmansuri

Copy link
Copy Markdown
Contributor

Fixes #8786

It looks like it was incorrect error handling.

It happens because:

  1. JRuby 10, because it ships Ruby 3.4's socket library with the Happy Eyeballs connect loop. JRuby 9.4 has the old sequential loop and is not affected.
  2. A call to Socket.tcp with a hostname. An IP literal takes a different, older path. Libraries that use TCPSocket.new, such as redis 4.4.0 and Net::HTTP, never enter this loop.
  3. A hostname that resolves to more than one address, in practice both an IPv6 and an IPv4 one. localhost is the everyday case: most systems map it to ::1 and 127.0.0.1. A dual-stack DNS name does the same.
  4. One of those connects is refused. The everyday case again: the server listens on IPv4 only, which is Redis's default bind 127.0.0.1, so the attempt on ::1 is refused. The loop races both addresses; when select sees the refused one, JRuby raises instead of reporting it, and the whole call fails even though the IPv4 connect would have worked.

It does not happen when the server listens on both families, when the connect fails by timeout rather than refusal, when the JVM has IPv6 disabled with the IPv4 flag, or on an IP literal.

Right now: select tries to complete the refused connect, the JDK throws and closes that channel, JRuby turns the exception into a plain IOError, and the loop dies right there: it never learns which socket failed, never sees the IPv4 socket finish, and its cleanup then trips over the channel the JDK already closed. The caller gets an error and no connection.

The Fix: select catches the JDK's exception, keeps it on the refused socket's descriptor, and returns both sockets as writable, which is what the operating system would do. The loop asks each one SO_ERROR. The refused socket now answers "connection refused" instead of zero, so the loop drops it, closes it, which no longer fails, and asks the IPv4 socket, which answers zero. That socket is handed back and the first write succeeds. If the refused address had been the only one, the loop would have reported "connection refused" with the right errno instead of a generic IO error.

aminmansuri and others added 2 commits September 10, 2026 17:19
select(2) reports a socket whose non-blocking connect failed as writable,
leaves the error for getsockopt(SOL_SOCKET, SO_ERROR), and leaves the
descriptor valid until the application closes it. SocketChannel.finishConnect()
throws and closes the channel instead. SelectExecutor calls it while building
IO.select's writable set, so the failure escaped select as a plain IOError,
getsockopt(SO_ERROR) answered a hardcoded 0, and closing the socket raised
Errno::EBADF. Socket.tcp's Happy Eyeballs v2 path depends on all three, so
Socket.tcp("localhost", port) failed whenever "localhost" resolved to ::1 as
well as 127.0.0.1 and the server listened on 127.0.0.1 only.

* SelectExecutor: keep a failed finishConnect() on the ChannelFD and still
  report the socket writable.
* BasicSocket#getsockopt(SOL_SOCKET, SO_ERROR): report that error.
* ChannelFD: allow closing a channel the JDK closed on connect failure.
* Helpers.errnoFromException: map java.net.ConnectException to ECONNREFUSED.

Fixes jruby#8786

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q6FreJjoCrTqixEqUQqcUB
…re again

It was excluded as "needs investigation": Socket.tcp raised Errno::EBADF
instead of Errno::ECONNREFUSED when the first connect was refused. That is
jruby#8786, fixed by the previous commit; the test passes now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q6FreJjoCrTqixEqUQqcUB
@aminmansuri

Copy link
Copy Markdown
Contributor Author

was able to remove test_tcp_socket_hostname_resolution_failed_after_connection_failure exclusion.

@headius

headius commented Sep 10, 2026

Copy link
Copy Markdown
Member

The one failing job looked unrelated so I ran it again.

@headius headius left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a good find and closes the loop (pun intended) with the new Happy Eyeballs code and JRuby's emulation of BSD socket behaviors. A few questions and a few changes and we can move forward with this.

Comment thread test/jruby/test_socket.rb Outdated
Comment thread core/src/main/java/org/jruby/util/io/SelectExecutor.java Outdated
Comment thread core/src/main/java/org/jruby/util/io/ChannelFD.java Outdated
aminmansuri and others added 2 commits September 10, 2026 20:05
Review follow-up. The stash for a failed non-blocking connect does not
belong on ChannelFD, which every IO object carries; only a socket can have
a connect fail. The error now lives on RubyBasicSocket, and ChannelFD is
left untouched by this branch.

SelectExecutor already has the Ruby IO object in hand at the connectable
branch (write_io), so it hands the IOException to the socket instead of to
the descriptor. An IO that is not a BasicSocket has nowhere to keep the
error, so for those the exception is rethrown and the select fails exactly
as it did before.

The close tolerance moves with it. finishConnect() closes the channel when
it fails, so the normal close path reports EBADF for a descriptor POSIX
still considers open. Rather than teach ChannelFD about connect failures,
RubyBasicSocket overrides rbIoClose. That is the single funnel every
socket close goes through - IO#close, Socket#close, and close_read /
close_write once the other half is already shut - and when a connect error
is stashed and the channel is already closed it releases the descriptor
with OpenFile#cleanup(runtime, true), without raising. That is what
RubySocket.tryConnect's finally already does for a connect that fails
synchronously, so a refused connect is now cleaned up the same way whether
the failure arrives out of connect() or out of select().

Helpers.errnoFromException keeps the ConnectException -> ECONNREFUSED
mapping. It is runtime-wide, but it is the general "which errno is this
IOException" helper, the JDK's message here ("Connection refused (connect
failed)") is not one the existing message table matches, and that table's
own FIXME already flags message scraping as fragile. A socket-local copy
would only duplicate it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q6FreJjoCrTqixEqUQqcUB
Review follow-up. Nothing these two tests assert is JRuby-specific - they
are the POSIX contract behind Happy Eyeballs v2 - so they belong in
spec/ruby/library/socket, where every implementation runs them, rather
than in test/jruby/test_socket.rb.

spec/ruby/library/socket/basicsocket/getsockopt_spec.rb gains a
"using Socket::SO_ERROR" context: SO_ERROR is 0 on a fresh socket; after a
refused non-blocking connect IO.select reports the socket writable and
SO_ERROR is Errno::ECONNREFUSED::Errno; and the socket stays valid until
it is closed, with close not raising. The select assertion stays in the
socket specs rather than moving to spec/ruby/core/io/select_spec.rb -
that file is about pipes and does not require 'socket', and what is being
specified here is the socket's state, not IO.select's contract.
Both connect examples tolerate a platform that refuses a loopback connect
synchronously, in which case there is no pending connect to report; that
is the same shape as MRI's own test_socket_connect_nonblock in
test/socket/test_addrinfo.rb, which wraps its SO_ERROR assertion in
rescue IO::WaitWritable.

spec/ruby/library/socket/socket/tcp_spec.rb gains the end-to-end case, in
its own describe with its own listener: the existing block builds its
server with Socket#listen, which JRuby does not support on a plain Socket,
so all six of its examples are tagged out here. The new one uses
TCPServer, and is guarded on a non-Windows platform where "localhost"
resolves to both an IPv4 and an IPv6 address - the situation the algorithm
needs before it can race two families. It asserts that Socket.tcp reaches
the IPv4 server (remote_address.ip_address == "127.0.0.1") and that a
write gets through. No resolv_timeout: or fast_fallback: keyword is used.

No spec tags are added: all four examples pass on this branch. On its base
three of the four fail - the two connect examples with the IOError out of
IO.select and an Errno::EBADF out of close, and the Socket.tcp example
with the Errno::EBADF that jruby#8786 reports.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q6FreJjoCrTqixEqUQqcUB
@aminmansuri
aminmansuri requested a review from headius September 11, 2026 00:25

@headius headius left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@headius
headius merged commit 3faf62f into jruby:jruby-10.0 Sep 11, 2026
224 of 226 checks passed
@headius headius added this to the JRuby 10.0.7.0 milestone Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants