perf(beacon): keep inactivity scores and the ring buffers on the tree - #650
MegaRedHand wants to merge 3 commits into
Conversation
Moving more BeaconState fields onto the tree would otherwise mean editing apply_pending_mutations and has_pending_mutations by hand for each one, and the two could drift apart (a field flushed but not counted as pending, or the reverse). An object-safe `Buffered` trait over `List` and `Vector`, and a `tree_fields!` list per fork that generates both views, keep them in step. H256 now forwards `is_basic_type`, as H160 already does: the transparent derive drops it, so a tree-backed collection of roots was treated as composite and kept a second cached copy of every root. The root is unchanged (libssz packs 32-byte basic elements one per chunk); tests compare `SszList`/`SszVector` of H256 against `[u8; 32]` and pin a lean container. New ssz-tree properties cover multi-leaf u8 lists, u64 and 32-byte-newtype vectors, their decode parity, and rebase of mostly-zero lists and vectors.
Inactivity scores are the largest flat field of a state (8 bytes per validator) and every cached state held a private deep copy, rehashed in full on each state root. On the tree, a derived state shares the list with its parent and rehashes only the leaves it touched. That only pays if epoch processing stops rewriting every eligible score: a write rebuilds its leaf even for an equal value. Outside a leak a missed epoch adds the bias and recovery takes it back, so nearly every score ends where it started. process_inactivity_updates now computes the new score with the same arithmetic as before (same checked add, same errors) and writes it only if it differs. The field joins the per-fork flush list, and BeaconState::rebase_on pairs the two states' scores whenever both forks carry them. Writes are buffered in a BTreeMap: sparse outside a leak. A leak writes about every score of the registry, which costs one B-tree insert each until a bulk write path exists.
block_roots, state_roots, randao_mixes and slashings are fixed-size ring buffers that change by a handful of entries per block or epoch, yet every cached state held a private deep copy of each and rehashed it in full on every state root. As tree vectors a derived state shares the unchanged leaves with its parent and rehashes only the touched paths. The small lists (eth1_data_votes, historical_roots, historical_summaries) move by the same alias swap, so no flat field is left that a derived state copies in bulk apart from the participation lists and the queues. All join the per-fork flush list and BeaconState::rebase_on, so a decoded state shares them with its resident parent. process_slot writes its two roots after its flush, so they stay buffered until the next flush (the next slot's, or process_slots' last step); the historical summaries update hashes them with that write pending, which is correct and cheap. Block production hashed the post-state without flushing it first, so every node it computed was thrown away; it flushes now, as the state transition does. Test-only sites that sliced or iterated these fields in place (the RPC test states, the storage delta fixtures) write through indexing instead.
Import-replay benchmarkMethod.
This matches the estimates: peak RSS was expected at 4.0 → 3.3 GiB and the ordinary-block saving at ~30 ms (47 ms measured). It is the only one of these PRs that lowers memory. RSS at the end of the import was not recorded, only the peak. The peak was consistent across both rounds (3.31 / 3.31 GiB). All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (
In that build #633's fused epoch pass reads and writes the tree-backed inactivity scores in order ( |
…3-64-633-636-638-gloas-live Conflicted files: - crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs - crates/blockchain/state_transition/src/beacon/stf/mod.rs - crates/common/ssz-tree/tests/model.rs - crates/common/types/src/beacon/containers/mod.rs - crates/common/types/src/beacon/containers/shared.rs Gloas adaptations: - tree_fields! gets a gloas line: validators and balances (progressive trees) plus block_roots, state_roots, historical_roots, eth1_data_votes, randao_mixes, slashings and historical_summaries, which gloas takes from the shared aliases. inactivity_scores is left out on purpose: gloas declares its own flat libssz ProgressiveList for it, which buffers nothing, so there is nothing to flush or rebase. - ssz-tree: ProgressiveList implements Buffered, so the field list can hold the progressive registry beside the bounded fields. - apply_pending_mutations / has_pending_mutations dispatch over all eight beacon forks through buffered()/buffered_mut(); the lean guard on has_pending_mutations is kept (the merge had dropped it with the old registry match). - rebase_on: the validators/balances list-kind match stays (so a gloas state still rebases onto a gloas base only); the other fields are shared types in every fork and rebase unconditionally; inactivity scores rebase only for a Tree/Tree pair. - Inactivity scores are a tree List before gloas and a flat slice-like list in gloas, so no &[u64] can serve both. altair_validator_lists now returns an InactivityScoresRef view (len, get, in-order iter, ptr_eq, has_pending_updates, PartialEq) as its third element, and inactivity_scores_mut (a &mut [u64]) becomes inactivity_score_mut(index), an element write that dispatches over both kinds. - stf/epoch/altair.rs update_inactivity_scores: decides over an in-order walk of the scores (in step with the summary) and writes only changed scores afterwards, which subsumes both #633's "write if different" and #650's no-op skip; the error order (first eligible validator without a score, then checked-add overflow) is unchanged. - stf/mod.rs tests: BEACON_FORKS now includes Gloas, the historical summaries write covers gloas, and the inactivity score case expects no pending write on gloas. process_slot's registry check uses the new BeaconState::registry_has_pending_updates, since the per-slot roots writes stay buffered by design. - storage store.rs test and the epoch-processing fixture runner comments follow the renamed accessors. Other: shared.rs imports both List/Vector and the progressive alias; model.rs keeps both sides' tests (the base side was empty).
Motivation
After lambdaclass/ethlambda_private#37 (persistent trees for
validatorsandbalances) every cached state still holds a private deep copy of the other large fields and rehashes them in full on each state root.inactivity_scoresrandao_mixesblock_roots,state_roots,slashingseth1_data_votes,historical_roots,historical_summariesEstimates, not measurements: ~26 MiB of flat fields per state and ~0.85M Merkle nodes rehashed per root, of which this PR moves the large majority of the nodes (participation lists and queues stay flat). The import-replay benchmark was not run for this PR, and no memory or node-count number was measured; the estimates come from the field sizes above.
Changes
perf(ssz-tree): flush tree fields through one Buffered list per forkBufferedtrait (impls forListandVector);tree_fields!lists each fork's tree fields once and generates bothapply_pending_mutationsandhas_pending_mutations.H256writesHashTreeRootby hand and forwardsis_basic_type(asH160does). New ssz-tree properties: multi-leafu8lists,u64and 32-byte-newtype vectors, decode parity, rebase of mostly-zero lists and vectors.perf(beacon): keep inactivity scores on the tree and skip no-op writesInactivityScoresis a tree list (BTreeMapupdate map).process_inactivity_updatescomputes the same value with the same checked arithmetic and writes onlyif new != score. Joins the flush list andBeaconState::rebase_on.perf(beacon): keep the roots and randao ring buffers on the treeBlockRoots,StateRoots,RandaoMixes,Slashingsbecome tree vectors;Eth1DataVotes,HistoricalRoots,HistoricalSummariesbecome tree lists (alias swap). All join the flush list andrebase_on. Block production now flushes before it hashes the post-state. Test-only sites that sliced or iterated in place use indexing.SSZ bytes, roots and the storage delta format are unchanged.
Design decisions
H256change done locally, not in libssz'stransparentderive (no dependency bump). The lean chain sharesH256, soethlambda-typeshas unit tests that anSszList/SszVectorofH256hashes like one of[u8; 32], and that a container holdingH256collections keeps its root. The lean state's pinned genesis root test passes unchanged.BTreeMapupdate map for inactivity scores. Outside a leak almost no score changes, so writes are sparse. A leak writes nearly every score, which costs about one B-tree insert per validator per epoch (~1M at the benchmark's registry size) until a bulk write path exists. A dense map (VecMap) was the alternative; it allocates 16 B per validator on the first push or write.process_inactivity_updates: read the score, compute the new one exactly as before (overflow still raisesArithmeticOverflowbefore the recovery step), write only if different. A write rebuilds its leaf even for an equal value, which would unshare the list from the parent state's.rebase_onpairs scores through the altair accessors, so it works across forks and skips phase0 (no scores).process_slotleaves its two roots writes buffered until the next flush (the next slot's, orprocess_slots' last step).process_historical_summaries_updatehashes the roots vectors with that write pending: correct, and the vectors are small. Tests pin this.produce_blockhashed the post-state withoutapply_pending_mutations, so the computed hashes were discarded; one added call.Bufferedandtree_fields!are additive in ssz-tree andcontainers/mod.rs, to keep a merge with other work in the same files easy.Not in this PR
previous_/current_epoch_participation)get_unslashed_participating_indices(#632): its per-index reads would cost ~2M tree descents per block in the pulled-up tip. Also needsList::filledfor an O(1) rotation.pending_depositson the treeArc+ cached-root newtype)previousequals the base'scurrent)ifinrebase_on, left for the participation PR.Testing
cargo test -p ethlambda-ssz-tree --profile release-fastcargo test -p ethlambda-types --profile release-fast --libcargo test -p ethlambda-storage --profile release-fast --libcargo test -p ethlambda-state-transition --profile release-fast --libcargo test -p ethlambda-rpc --profile release-fast --libcargo clippy(workspace, and state-transition withbeacon-spec-tests) with-D warnings,cargo fmt --check,cargo check -p ethlambdaThe 4 failures in each spec run are the
gossip/*trials, which report "matched no fixture cases": the gossip fixture directory is intentionally empty in the local setup. Everyssz_static,sanity,finality,random,epoch_processingandtransitiontrial passes. The leanssz_spectestsand the import-replay benchmark were not run.