Skip to content

Optimize create-and-resume usage of Fiber in Enumerator - #9552

Draft
headius wants to merge 1 commit into
jruby:masterfrom
headius:thread_fiber_optz
Draft

headius wants to merge 1 commit into
jruby:masterfrom
headius:thread_fiber_optz

Conversation

@headius

@headius headius commented Jul 27, 2026

Copy link
Copy Markdown
Member

When constructing a Fiber in support of Enumerator#next, we know it will be immediately resumed, so we can feed that initial resume value directly into the first execution of the virtual thread, avoiding a potentially costly and unnecessary vthread suspension at create time. This avoids such create-and-resume cases triggering the initial allocation of vthread state (StackChunk etc) based on a much shallower stack than they will actually need, and a subsequent throwing out of that initial state.

This patch was created with assistance from an agent and has a few unresolved issues:

  • Repeatedly checking for null initialBlock and initialRequest even for cases that are not using lazy initialization (such as typical Fiber.new usage). This could be eliminated by abstracting the first exchange into a function object, replaced with the simple version on first resume, but my first attempt to clean it up got messy.
  • The method is private and must be called with #send, but it remains potentially visible to user code; a cleaner implementation would move this internal method to an internal location that will only be called from internal code.
  • This only improves the performance of an initial Enumerator#next, which would be dwarfed by subsequent #next calls whenever there's more than one. This is a severe edge case which has many other strikes against it, so the extra complexity here may not be worth speeding up this rare and discouraged case.

The positive improvements, for reference:

  • Because the new vthread does not suspend until it actually yields its first result, the StackChunk allocated at that point may be exactly as large as it needs to be (assuming that yield lives at the same level as later yields, which is the case for the Fiber created by Enumerator#next).
  • It proves a fused create-and-resume case can perform better than creating and resuming a Fiber separately to get a single object. This may be an interesting case combined with fiber schedulers, since that scenario may have many single-shot fibers expecting to be called only once (and presumably created to encapsulate a single blocking call).

Further cleanup and experimentation are needed to move forward with this change.

This relates to ruby/csv#361 and improves the performance of the Enumerator#next benchmark there by roughly 30-40%.

When constructing a Fiber in support of Enumerator#next, we know it
will be immediately resumed, so we can feed that initial resume
value directly into the first execution of the virtual thread,
avoiding a potentially costly and unnecessary vthread suspension
at create time. This avoids such create-and-resume cases triggering
the initial allocation of vthread state (StackChunk etc) based on
a much shallower stack than they will actually need, and a
subsequent throwing out of that initial state.

This patch was created with assistance from an agent and has a few
unresolved issues:

* Repeatedly checking for null initialBlock and initialRequest even
  for cases that are not using lazy initialization (such as typical
  Fiber.new usage). This could be eliminated by abstracting the
  first exchange into a function object, replaced with the simple
  version on first resume, but my first attempt to clean it up got
  messy.
* The method is private and must be called with #send, but it
  remains potentially visible to user code; a cleaner
  implementation would move this internal method to an internal
  location that will only be called from internal code.
* This only improves the performance of an initial Enumerator#next,
  which would be dwarfed by subsequent #next calls whenever there's
  more than one. This is a severe edge case which has many other
  strikes against it, so the extra complexity here may not be worth
  speeding up this rare and discouraged case.

The positive improvements, for reference:

* Because the new vthread does not suspend until it actually yields
  its first result, the StackChunk allocated at that point may be
  exactly as large as it needs to be (assuming that yield lives at
  the same level as later yields, which is the case for the Fiber
  created by Enumerator#next).
* It proves a fused create-and-resume case can perform better than
  creating and resuming a Fiber separately to get a single object.
  This may be an interesting case combined with fiber schedulers,
  since that scenario may have many single-shot fibers expecting to
  be called only once (and presumably created to encapsulate a
  single blocking call).

Further cleanup and experimentation are needed to move forward with
this change.

This relates to ruby/csv#361 and improves the performance of the
Enumerator#next benchmark there by roughly 30-40%.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant