Conversation
Port the Chromium fork's handling of major GCs to node's V8 (9.4): - Stop skipping mark-compact collections in Heap::PerformGarbageCollection (chromium v8 7b72420d07e). - Stop forcing --never-compact, as in the Chromium fork. - Disallow events in the GC, sweeping, memory pressure, teardown and background allocation paths that the Chromium fork covers, using its labels on the V8 9.4 equivalents of those functions. - Never trigger a GC from a background thread allocation while replaying (58b6be11dd5). - Defer Managed<T> destructors to Isolate::ReleaseSharedPtrs when the finalizer runs with events disallowed (ffe66fe960a). - Keep WasmCode objects alive forever (b8142479974). - Tag the existing GC behavior changes with the "gc-changes" feature. Concurrent, parallel and incremental GC stay disabled when both recording and replaying; the Chromium fork only disables them when replaying. WeakRef and FinalizationRegistry handling and node's own embedder finalizers are not covered here. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With incremental marking disabled the old generation always expands when its allocation limit is reached, so major GCs only ran once the heap hit --max-old-space-size. Enable incremental marking while recording, as the Chromium fork does (chromium v8 fb7aa5108d1), and keep it disabled while replaying. The concurrent and parallel GC flags stay disabled in both modes. DefaultJobState::NotifyConcurrencyIncrease doesn't post to worker threads while events are disallowed, so e.g. the page unmapper job posted during a GC never ran and freed pages were not returned to the OS until teardown. Disallow events in Heap::HandleGCRequest and in the incremental finalization steps it runs. V8 9.4 still finalizes incremental marking from the stack guard, which made recorded calls that don't happen when replaying. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Port of chromium v8 4ea00745613. V8 9.4 has no ConcurrentAllocator::AllocateFromSpaceFreeList, so those diagnostics go into its counterpart PagedSpace::RawRefillLabBackground. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Found by comparing the GC related record/replay code of both forks: - Don't use the eval compilation cache when recording/replaying. Major GCs age entries out of the cache, so whether an eval creates a new script depended on GC timing. - Ignore interrupts in StackGuard::HandleInterrupts while events are disallowed. - Disallow events for the whole of LocalHeap::UnparkSlowPath and LocalHeap::SafepointSlowPath, not only for background threads. - Assert the use count in Managed<T>::Destructor. - Add the LocalHeap destruction diagnostic. - Tag the heap snapshot serializer change with the "gc-changes" feature. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Incremental marking is no longer disabled while recording. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
bhackett1024
force-pushed
the
claude/node-gc-chromium-fork-c994c3
branch
from
October 2, 2026 11:58
3a8c1bf to
5303c56
Compare
Andarist
added this pull request to stack #165
October 2, 2026 12:04
…fork - Don't use the compilation cache when replaying, or when the "v8-flags-compilation-cache" feature is off. Major GCs age entries out of the script cache, and only a cache hit while recording is replayed, so a hit that only happens when replaying would skip creating a script. - Gate recordreplay::AreEventsDisallowed and AreEventsPassedThrough on the "disallow-events" and "pass-through-events" features and pass the label on, instead of ignoring both. - Disallow events in cppgc's HeapBase::Terminate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Incremental marking and the compilation cache are enabled when recording and disabled when replaying, and the flag hash covers every non-default flag, so it differed between the two. It is exposed by v8.cachedDataVersionTag() and decides whether a code cache is accepted, so both behaved differently when replaying. Record the hash when it is computed and use the recorded value when replaying. Also update the comment in Heap::ShouldExpandOldGenerationOnSlowAllocation, which described the record/replay case as incremental marking being disabled. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The measurement is reported after a GC, at a point and with sizes which differ between recording and replaying. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…registration" Port of replayio/chromium-v8#338. RegisterManagedPtrDestructor and ReleaseSharedPtrs passed the length of the destructor list to recordreplay::Assert, walking the whole list on every call, also when not recording. Registering N Managed objects cost O(N^2): 40,000 Intl.NumberFormat instances took 17.5 s instead of 1.5 s. Keep a count next to the list head instead. Only assert the count when leak-references is enabled. Without it ManagedObjectFinalizerSecondPass unregisters destructors at GC time, so the count is not the same when replaying. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Enables old-space (major) GC in replay-node. Until now mark-compact collections were skipped entirely when recording or replaying, so the old generation only ever grew. With this PR major GCs run while recording, triggered at the normal allocation limit, and the resulting recordings replay. In a churn test that pushes ~1.7GB through the old generation with ~120MB live, peak RSS under recording drops from ~1.9GB to ~0.6GB.
When replaying, incremental marking is off and the heap is always allowed to expand, so replay doesn't depend on reproducing the recording's GC timing.
Ported from the Chromium fork
Enabling major GC:
Heap::PerformGarbageCollection(7b72420d07e).fb7aa5108d1) or when thev8-flags-gcfeature is off (f4bfbecd7cf). Without itHeap::ShouldExpandOldGenerationOnSlowAllocationkeeps expanding the old generation and a major GC only runs at--max-old-space-size.--never-compact(3104d7cad00, a build fix after the flag was removed upstream; in V8 9.4 the flag still exists).AutoDisallowEventsscopes. The labels are the ones the Chromium fork uses (1917249bb49). Where V8 9.4 names a function differently, the scope is on the 9.4 counterpart and keeps the Chromium fork's label:LocalHeap::UnparkSlowPath,LocalHeap::SafepointSlowPath(da7a5cd8b0b), covering the whole function as in the Chromium fork's current code.LocalHeap::TryPerformCollection(f417cac55f3). The function no longer exists in the Chromium fork's V8.LocalHeap::PerformCollectionAndAllocateAgain(bc3ca2ba1a1).Heap::MemoryPressureNotification(8de2fbfb228).PagedSpace::RawRefillLabMain, labelledPagedSpaceBase::RawRefillLabMain(7138afe0be7).PagedSpace::RawRefillLabBackground, the 9.4 counterpart ofConcurrentAllocator::AllocateFromSpaceFreeList(e39ca9b782a).Heap::EnsureSweepingCompleted(HeapObject), the 9.4 counterpart ofHeap::EnsureSweepingCompletedForObject(0d24bd38639).MarkCompactCollector::EnsureSweepingCompleted, the 9.4 counterpart ofHeap::EnsureSweepingCompleted(993ce4f2437).SweeperImpl::Start(1ca67bc0762).AllocationCounter::InvokeAllocationObservers,ConcurrentMarking::ScheduleJob,StatsCollector::AllocatedObjectSizeSafepointImpl,Heap::ReportExternalMemoryPressure,Heap::StartIncrementalMarking,Heap::StartIncrementalMarkingIfAllocationLimitIsReached,IncrementalMarkingJob::ScheduleTask,IncrementalMarkingJob::Task::RunInternal,IncrementalMarking::MarkBlackAndVisitObjectDueToLayoutChange(fb7aa5108d1).Heap::ActivateMemoryReducerIfNeeded(6e683f09032).Heap::CollectAllAvailableGarbage(7e0f242a8eb).Heap::StartTearDown, cppgcSweeperImpl::FinishIfRunning(2810b09546e).Heap::CollectGarbageOnMemoryPressure(6ab2e8798a8).SweeperImpl::SweepForAllocationIfRunning, as in the Chromium fork's current code.Behavior that would otherwise depend on GC timing:
Heap::ShouldExpandOldGenerationOnSlowAllocation(58b6be11dd5). Waiting can deadlock when the main thread is blocked on an ordered lock that the background thread is next to acquire. The hard limit inHeap::CanExpandOldGenerationBackgroundstill applies.no-eval-cache), inCompilationCache::LookupEvalandCompilationCache::PutEval. Major GCs age entries out of the cache, so whether anevalcreates a new script would depend on GC timing.StackGuard::HandleInterruptswhile events are disallowed (076d23fbc40).Managed<T>destructors toIsolate::ReleaseSharedPtrswhen the finalizer runs with events disallowed, plus the list-length asserts (ffe66fe960a). Gated on theleak-referencesfeature.WasmCodeobjects alive forever (b8142479974), gated on theleak-referencesfeature as in the Chromium fork's current code.Asserts, diagnostics and feature tags:
Managed<T>::Destructor(9563c03f088).OldLargeObjectSpace::AllocateRawBackgroundandPagedSpace::RawRefillLabBackground(4ea00745613).~LocalHeapdiagnostic (c920c27399c).gc-changesfeature (Heap::ShouldExpandOldGenerationOnSlowAllocation,MemoryReducer::ScheduleTimer,GCTracer::Print,HeapSnapshotJSONSerializer::SerializeImpl) and label the existingAutoDisallowEventsscopes inheap.cc.Added on top of the Chromium fork
AutoDisallowEventsinHeap::HandleGCRequest,Heap::FinalizeIncrementalMarkingIfCompleteandHeap::FinalizeIncrementalMarkingIncrementally. V8 9.4 finalizes incremental marking in a separate step run from the stack guard, outside every scope listed above. It made recorded calls that don't happen when replaying (Mismatched call expected clock_gettime got gettimeofday). The Chromium fork's V8 has noHeap::FinalizeIncrementalMarkingIncrementally; itsHeap::HandleGCRequestandHeap::FinalizeIncrementalMarkingIfCompleteonly call functions that already carry a scope, so they have none of their own.local_heapinHeap::ShouldExpandOldGenerationOnSlowAllocation, which can be null in 9.4.Intentional divergences from the Chromium fork
v8-flags-gcis off. V8 runs this work throughPlatform::PostJob, almost always from inside a GC, where events are disallowed.PostJobon its own thread pool (base::CreateJob), which carries Replay patches so that worker threads can pick up a job posted while events are disallowed (base/task/thread_pool/job_task_source.cc).DefaultJobState::NotifyConcurrencyIncreasereturns early while events are disallowed. A job posted during a GC is created but no worker thread is ever told to run it.--concurrent-sweeping, ran once and never again. Freed pages were only returned to the OS at teardown, and RSS stayed at ~1.6GB while the JS heap shrank. With--no-concurrent-sweepingRSS follows the heap.--concurrent-marking,--parallel-compaction,--parallel-marking,--parallel-pointer-updateand--parallel-scavenge.--concurrent-array-buffer-sweepingalso stays disabled in Node. The Chromium fork no longer sets it in either mode.--incremental-marking-task,--scavenge-task). Node's platform drops tasks markedIsRecordReplayNonDeterministic, so posting them is pointless. Incremental marking advances on allocation and is finalized from the stack guard instead.PostTaskmakes recorded calls, and the post would happen at a non-deterministic point. As a consequence theManaged<T>deferral applies to every finalizer, so withleak-referencesenabled those objects live until isolate teardown.Heap::CanExpandOldGenerationalways returns true when replaying. An existing Node patch that the Chromium fork no longer carries. Left as is, since it keeps a replay from running out of memory when it uses more heap than the recording did.Testing
test/node-recordingsuite: 25 of 25 pass, including a newgc/old-space-churncase that records, checks the heap stays under 800MB, replays and evaluates.test-domain-error-types(--stress-compaction) aborted under recording before this PR and passes now, so it can leave the backend'sNodeTestIgnoreList.test/report/test-report-fatal-error.jstimed out once on CI while recording (120s). It passes locally with the exact CI binary in every run. Its children run into--max-old-space-size=20on purpose and now take ~0.9s each instead of ~0.3s, because they do real GC work before aborting. Not root-caused.Follow-ups
WeakRefandFinalizationRegistry. Both become observable now that old-space objects die, and Node's own internals use them (event_target.js,abort_controller.js,domain.js). To be ported from the Chromium fork once theFinalizationRegistrywork there lands.FileHandle/DirHandleclosing their fd and scheduling a warning, N-API finalizers,ObjectWrapdestructors. Needs either the Chromium fork'sleak-referencespattern per site, or recording which handles were finalized and replaying that.AsyncWrap::EmitDestroyis a no-op under record/replay, so GC-drivenasync_hooksdestroy hooks never fire. Thetest-gc-http-client*core tests hang on this and stay in the backend'sNodeTestIgnoreList.node_utilWeakReference) have the same problem asWeakRef.Managed<T>memory. Native objects behindManaged<T>(ICU objects fromIntl, WebAssembly modules) are kept until isolate teardown. The list-length asserts inIsolate::ReleaseSharedPtrsandIsolate::RegisterManagedPtrDestructorwalk that list on every registration, also when not recording.CollectionBarrier. A background thread requesting a GC while recording posts a foreground task through Node'sPostTask. The Chromium fork has the same code and no handling; not exercised here.test-report-fatal-errorCI timeout above.evalcombined with heap churn crashes on replay (MissingReport), also on a build without this PR. Found while writing a test for the eval cache change, which therefore has no test.gc/old-space-churnrecording test and theNodeTestIgnoreListchange.