Skip to content

Enable old space GC when recording/replaying - #161

Draft
Andarist wants to merge 9 commits into
claude/gc-finalizers-replayfrom
claude/node-gc-chromium-fork-c994c3
Draft

Andarist wants to merge 9 commits into
claude/gc-finalizers-replayfrom
claude/node-gc-chromium-fork-c994c3

Conversation

@Andarist

@Andarist Andarist commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Enables old-space (major) GC in replay-node. Until now mark-compact collections were skipped entirely when recording or replaying, so the old generation only ever grew. With this PR major GCs run while recording, triggered at the normal allocation limit, and the resulting recordings replay. In a churn test that pushes ~1.7GB through the old generation with ~120MB live, peak RSS under recording drops from ~1.9GB to ~0.6GB.

When replaying, incremental marking is off and the heap is always allowed to expand, so replay doesn't depend on reproducing the recording's GC timing.

Ported from the Chromium fork

Enabling major GC:

  • Stop skipping mark-compact collections in Heap::PerformGarbageCollection (7b72420d07e).
  • Enable incremental marking while recording and keep it disabled while replaying (fb7aa5108d1) or when the v8-flags-gc feature is off (f4bfbecd7cf). Without it Heap::ShouldExpandOldGenerationOnSlowAllocation keeps expanding the old generation and a major GC only runs at --max-old-space-size.
  • Stop forcing --never-compact (3104d7cad00, a build fix after the flag was removed upstream; in V8 9.4 the flag still exists).

AutoDisallowEvents scopes. The labels are the ones the Chromium fork uses (1917249bb49). Where V8 9.4 names a function differently, the scope is on the 9.4 counterpart and keeps the Chromium fork's label:

  • LocalHeap::UnparkSlowPath, LocalHeap::SafepointSlowPath (da7a5cd8b0b), covering the whole function as in the Chromium fork's current code.
  • LocalHeap::TryPerformCollection (f417cac55f3). The function no longer exists in the Chromium fork's V8.
  • LocalHeap::PerformCollectionAndAllocateAgain (bc3ca2ba1a1).
  • Heap::MemoryPressureNotification (8de2fbfb228).
  • PagedSpace::RawRefillLabMain, labelled PagedSpaceBase::RawRefillLabMain (7138afe0be7).
  • PagedSpace::RawRefillLabBackground, the 9.4 counterpart of ConcurrentAllocator::AllocateFromSpaceFreeList (e39ca9b782a).
  • Heap::EnsureSweepingCompleted(HeapObject), the 9.4 counterpart of Heap::EnsureSweepingCompletedForObject (0d24bd38639).
  • MarkCompactCollector::EnsureSweepingCompleted, the 9.4 counterpart of Heap::EnsureSweepingCompleted (993ce4f2437).
  • cppgc SweeperImpl::Start (1ca67bc0762).
  • AllocationCounter::InvokeAllocationObservers, ConcurrentMarking::ScheduleJob, StatsCollector::AllocatedObjectSizeSafepointImpl, Heap::ReportExternalMemoryPressure, Heap::StartIncrementalMarking, Heap::StartIncrementalMarkingIfAllocationLimitIsReached, IncrementalMarkingJob::ScheduleTask, IncrementalMarkingJob::Task::RunInternal, IncrementalMarking::MarkBlackAndVisitObjectDueToLayoutChange (fb7aa5108d1).
  • Heap::ActivateMemoryReducerIfNeeded (6e683f09032).
  • Heap::CollectAllAvailableGarbage (7e0f242a8eb).
  • Heap::StartTearDown, cppgc SweeperImpl::FinishIfRunning (2810b09546e).
  • Heap::CollectGarbageOnMemoryPressure (6ab2e8798a8).
  • cppgc SweeperImpl::SweepForAllocationIfRunning, as in the Chromium fork's current code.

Behavior that would otherwise depend on GC timing:

  • While replaying, let background threads expand the old generation at the allocation limit instead of waiting for a main thread GC, in Heap::ShouldExpandOldGenerationOnSlowAllocation (58b6be11dd5). Waiting can deadlock when the main thread is blocked on an ordered lock that the background thread is next to acquire. The hard limit in Heap::CanExpandOldGenerationBackground still applies.
  • Don't use the eval compilation cache when recording/replaying (no-eval-cache), in CompilationCache::LookupEval and CompilationCache::PutEval. Major GCs age entries out of the cache, so whether an eval creates a new script would depend on GC timing.
  • Ignore interrupts in StackGuard::HandleInterrupts while events are disallowed (076d23fbc40).
  • Defer Managed<T> destructors to Isolate::ReleaseSharedPtrs when the finalizer runs with events disallowed, plus the list-length asserts (ffe66fe960a). Gated on the leak-references feature.
  • Keep WasmCode objects alive forever (b8142479974), gated on the leak-references feature as in the Chromium fork's current code.

Asserts, diagnostics and feature tags:

  • Assert the use count in Managed<T>::Destructor (9563c03f088).
  • Diagnostics for heap allocation failures on background threads, in OldLargeObjectSpace::AllocateRawBackground and PagedSpace::RawRefillLabBackground (4ea00745613).
  • The ~LocalHeap diagnostic (c920c27399c).
  • Tag the existing GC behavior changes with the gc-changes feature (Heap::ShouldExpandOldGenerationOnSlowAllocation, MemoryReducer::ScheduleTimer, GCTracer::Print, HeapSnapshotJSONSerializer::SerializeImpl) and label the existing AutoDisallowEvents scopes in heap.cc.

Added on top of the Chromium fork

  • AutoDisallowEvents in Heap::HandleGCRequest, Heap::FinalizeIncrementalMarkingIfComplete and Heap::FinalizeIncrementalMarkingIncrementally. V8 9.4 finalizes incremental marking in a separate step run from the stack guard, outside every scope listed above. It made recorded calls that don't happen when replaying (Mismatched call expected clock_gettime got gettimeofday). The Chromium fork's V8 has no Heap::FinalizeIncrementalMarkingIncrementally; its Heap::HandleGCRequest and Heap::FinalizeIncrementalMarkingIfComplete only call functions that already carry a scope, so they have none of their own.
  • A null check on local_heap in Heap::ShouldExpandOldGenerationOnSlowAllocation, which can be null in 9.4.

Intentional divergences from the Chromium fork

  • Concurrent and parallel GC stay disabled when recording. The Chromium fork leaves them enabled while recording and only disables them when replaying or when v8-flags-gc is off. V8 runs this work through Platform::PostJob, almost always from inside a GC, where events are disallowed.
    • Chromium implements PostJob on its own thread pool (base::CreateJob), which carries Replay patches so that worker threads can pick up a job posted while events are disallowed (base/task/thread_pool/job_task_source.cc).
    • Node uses V8's default job implementation, where DefaultJobState::NotifyConcurrencyIncrease returns early while events are disallowed. A job posted during a GC is created but no worker thread is ever told to run it.
    • Observed with the Chromium fork's flags: the page unmapper job, gated by --concurrent-sweeping, ran once and never again. Freed pages were only returned to the OS at teardown, and RSS stayed at ~1.6GB while the JS heap shrank. With --no-concurrent-sweeping RSS follows the heap.
    • The remaining flags were not shown to fail individually. They stay off, as they were before this PR, because they depend on the same job mechanism: --concurrent-marking, --parallel-compaction, --parallel-marking, --parallel-pointer-update and --parallel-scavenge.
    • --concurrent-array-buffer-sweeping also stays disabled in Node. The Chromium fork no longer sets it in either mode.
  • GC tasks stay disabled when recording (--incremental-marking-task, --scavenge-task). Node's platform drops tasks marked IsRecordReplayNonDeterministic, so posting them is pointless. Incremental marking advances on allocation and is finalized from the stack guard instead.
  • Second-pass weak callbacks still run synchronously inside the GC. The Chromium fork posts a task for them. Node's PostTask makes recorded calls, and the post would happen at a non-deterministic point. As a consequence the Managed<T> deferral applies to every finalizer, so with leak-references enabled those objects live until isolate teardown.
  • Heap::CanExpandOldGeneration always returns true when replaying. An existing Node patch that the Chromium fork no longer carries. Left as is, since it keeps a replay from running out of memory when it uses more heap than the recording did.

Testing

  • Heap churn script under recording: heap used at exit 1762MB → 294MB, peak RSS 1945MB → 594MB.
  • Backend test/node-recording suite: 25 of 25 pass, including a new gc/old-space-churn case that records, checks the heap stays under 800MB, replays and evaluates.
  • Node core test test-domain-error-types (--stress-compaction) aborted under recording before this PR and passes now, so it can leave the backend's NodeTestIgnoreList.
  • Node core test test/report/test-report-fatal-error.js timed out once on CI while recording (120s). It passes locally with the exact CI binary in every run. Its children run into --max-old-space-size=20 on purpose and now take ~0.9s each instead of ~0.3s, because they do real GC work before aborting. Not root-caused.

Follow-ups

  • WeakRef and FinalizationRegistry. Both become observable now that old-space objects die, and Node's own internals use them (event_target.js, abort_controller.js, domain.js). To be ported from the Chromium fork once the FinalizationRegistry work there lands.
  • Node embedder finalizers. Weak callbacks run inside the GC with events disallowed, so their calls are passed through unrecorded. They run only while recording and can change state the replay depends on: FileHandle/DirHandle closing their fd and scheduling a warning, N-API finalizers, ObjectWrap destructors. Needs either the Chromium fork's leak-references pattern per site, or recording which handles were finalized and replaying that.
  • AsyncWrap::EmitDestroy is a no-op under record/replay, so GC-driven async_hooks destroy hooks never fire. The test-gc-http-client* core tests hang on this and stay in the backend's NodeTestIgnoreList.
  • JS-visible weak handles in Node (node_util WeakReference) have the same problem as WeakRef.
  • Managed<T> memory. Native objects behind Managed<T> (ICU objects from Intl, WebAssembly modules) are kept until isolate teardown. The list-length asserts in Isolate::ReleaseSharedPtrs and Isolate::RegisterManagedPtrDestructor walk that list on every registration, also when not recording.
  • CollectionBarrier. A background thread requesting a GC while recording posts a foreground task through Node's PostTask. The Chromium fork has the same code and no handling; not exercised here.
  • Concurrent GC while recording. Would need Node's worker thread job path to work while events are disallowed.
  • The test-report-fatal-error CI timeout above.
  • eval combined with heap churn crashes on replay (MissingReport), also on a build without this PR. Found while writing a test for the eval cache change, which therefore has no test.
  • Coverage. Not yet exercised under GC pressure: worker threads, native addons, WebAssembly.
  • Backend. Commit the gc/old-space-churn recording test and the NodeTestIgnoreList change.

Andarist and others added 5 commits October 2, 2026 11:58
Port the Chromium fork's handling of major GCs to node's V8 (9.4):

- Stop skipping mark-compact collections in
  Heap::PerformGarbageCollection (chromium v8 7b72420d07e).
- Stop forcing --never-compact, as in the Chromium fork.
- Disallow events in the GC, sweeping, memory pressure, teardown and
  background allocation paths that the Chromium fork covers, using its
  labels on the V8 9.4 equivalents of those functions.
- Never trigger a GC from a background thread allocation while
  replaying (58b6be11dd5).
- Defer Managed<T> destructors to Isolate::ReleaseSharedPtrs when the
  finalizer runs with events disallowed (ffe66fe960a).
- Keep WasmCode objects alive forever (b8142479974).
- Tag the existing GC behavior changes with the "gc-changes" feature.

Concurrent, parallel and incremental GC stay disabled when both
recording and replaying; the Chromium fork only disables them when
replaying. WeakRef and FinalizationRegistry handling and node's own
embedder finalizers are not covered here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With incremental marking disabled the old generation always expands
when its allocation limit is reached, so major GCs only ran once the
heap hit --max-old-space-size. Enable incremental marking while
recording, as the Chromium fork does (chromium v8 fb7aa5108d1), and
keep it disabled while replaying.

The concurrent and parallel GC flags stay disabled in both modes.
DefaultJobState::NotifyConcurrencyIncrease doesn't post to worker
threads while events are disallowed, so e.g. the page unmapper job
posted during a GC never ran and freed pages were not returned to the
OS until teardown.

Disallow events in Heap::HandleGCRequest and in the incremental
finalization steps it runs. V8 9.4 still finalizes incremental marking
from the stack guard, which made recorded calls that don't happen when
replaying.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Port of chromium v8 4ea00745613. V8 9.4 has no
ConcurrentAllocator::AllocateFromSpaceFreeList, so those diagnostics
go into its counterpart PagedSpace::RawRefillLabBackground.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Found by comparing the GC related record/replay code of both forks:

- Don't use the eval compilation cache when recording/replaying. Major
  GCs age entries out of the cache, so whether an eval creates a new
  script depended on GC timing.
- Ignore interrupts in StackGuard::HandleInterrupts while events are
  disallowed.
- Disallow events for the whole of LocalHeap::UnparkSlowPath and
  LocalHeap::SafepointSlowPath, not only for background threads.
- Assert the use count in Managed<T>::Destructor.
- Add the LocalHeap destruction diagnostic.
- Tag the heap snapshot serializer change with the "gc-changes"
  feature.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Incremental marking is no longer disabled while recording.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@bhackett1024
bhackett1024 force-pushed the claude/node-gc-chromium-fork-c994c3 branch from 3a8c1bf to 5303c56 Compare October 2, 2026 11:58
@Andarist
Andarist changed the base branch from master to claude/gc-finalizers-replay October 2, 2026 12:04
@Andarist
Andarist added this pull request to stack #165 October 2, 2026 12:04
Andarist and others added 4 commits October 2, 2026 14:06
…fork

- Don't use the compilation cache when replaying, or when the
  "v8-flags-compilation-cache" feature is off. Major GCs age entries
  out of the script cache, and only a cache hit while recording is
  replayed, so a hit that only happens when replaying would skip
  creating a script.
- Gate recordreplay::AreEventsDisallowed and AreEventsPassedThrough on
  the "disallow-events" and "pass-through-events" features and pass
  the label on, instead of ignoring both.
- Disallow events in cppgc's HeapBase::Terminate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Incremental marking and the compilation cache are enabled when
recording and disabled when replaying, and the flag hash covers every
non-default flag, so it differed between the two. It is exposed by
v8.cachedDataVersionTag() and decides whether a code cache is
accepted, so both behaved differently when replaying.

Record the hash when it is computed and use the recorded value when
replaying.

Also update the comment in
Heap::ShouldExpandOldGenerationOnSlowAllocation, which described the
record/replay case as incremental marking being disabled.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The measurement is reported after a GC, at a point and with sizes
which differ between recording and replaying.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…registration"

Port of replayio/chromium-v8#338.

RegisterManagedPtrDestructor and ReleaseSharedPtrs passed the length
of the destructor list to recordreplay::Assert, walking the whole list
on every call, also when not recording. Registering N Managed objects
cost O(N^2): 40,000 Intl.NumberFormat instances took 17.5 s instead of
1.5 s.

Keep a count next to the list head instead.

Only assert the count when leak-references is enabled. Without it
ManagedObjectFinalizerSecondPass unregisters destructors at GC time,
so the count is not the same when replaying.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant