Skip to content

perf(evmonly): parallelize OCC validation and merge behind the serial acceptance barrier - #4457

Merged
bdchatham merged 5 commits into
mainfrom
brandon2/port-parallel-occ-to-main
Oct 5, 2026
Merged

bdchatham merged 5 commits into
mainfrom
brandon2/port-parallel-occ-to-main

Conversation

@bdchatham

@bdchatham bdchatham commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Backport: this ports #4273, the reviewable version of #4261, to main from the giga-1 PR stack #4270–#4273. The stack targeted stacked giga-1 branches and never reached main; only its first PR landed on main, as #4366. @shemnon approved #4273 at 1143167. Since that commit, the changes are main's renames (baseAccount, the phaseOCC* constants, takeAccessSets) and two review fixes: merge fragments are cleared before they return to the pool, and each parallel pass's look-ahead is bounded.

On giga-testnet-2, occ_validate and occ_merge took about a third of the executor main loop, all of it on the calling goroutine. Most of that time went to transactions that the frontier accepts as they stand. A block of about 1,800 transactions holds less than one conflict. On giga-1, #4261 moved this work onto the OCC pool. It cut validate and merge from about 8.5 ms to 4.1 ms per block. That gave about 20% more blocks per second during the 95k tx/s window. #4269 reverted it for the reviewable stack, and that stack (#4273) never reached main. This PR ports #4261 and the #4273 follow-up to main (PLT-1377).

Acceptance stays a serial barrier in block order. stateAccessIndex and blockSTMState split into 64 address shards. indexResults records the writes of every speculative result up front, and conflictsWithin(key, sourcePrefix, txIndex) replaces conflictsWithAfter. acceptValidatedPrefix finds the first result that the serial frontier would not accept, and applyRange folds the run before it into the prefix shard by shard. validateBlockSTMFrontier then handles only that one result, and serialBackoff keeps dependency chains on the calling goroutine. Each pass looks ahead at most twice as far as the previous pass accepted, and at least 2,048 results, so the work of a block's passes stays linear in its size. changeSetIntoParallel builds the changeset of each shard on the pool and joins the shards in address order. Its per-shard base-row cache replaces prefetchBaseAccounts. The port needs nothing from #4258 or #4260; the only conflicts came from the baseAccount and phase-constant renames on main.

Non-app-hash-breaking: changesets, receipts, and tx results stay byte-identical to main. Shards are contiguous ranges of the first address byte, so shard order is canonical address order. Each address lives in one shard, so the apply order in a shard is block order. The new index can only add conflicts, from stale writes of an earlier incarnation. An extra rerun executes against the exact accepted prefix, so it gives the sequential result, and the extra reruns show only in OCC metrics. Review stmFrontierAccepts most closely, because it must mirror needsSTMRerun, and touchedShards, because it must cover every address that applyOwned and indexResults touch. TestOCCRandomizedConflictingBlocksMatchSequential checks seeded dense and sparse blocks at 3 and 8 workers against the sequential executor, down to the encoded FlatKV pairs. A seeded digest run gave identical output on main and this branch for 864 blocks, including 3,000- and 6,000-transaction blocks that cross a pass's look-ahead.

bdchatham and others added 2 commits October 5, 2026 11:47
… acceptance barrier

Port of #4261 and the #4273 follow-up from giga-1 onto main. Acceptance
stays in block order on the calling goroutine. The executor splits the
write index and the accepted prefix into 64 address shards. The OCC pool
indexes every speculative result up front. A parallel pass accepts the
longest run that the serial frontier would accept. The frontier then
handles only the first result that needs a rerun or fails. The merge
builds each shard's changeset on the pool and joins the shards in
address order.

A per-shard base-row cache in the merge replaces prefetchBaseAccounts.
The port keeps main's renames (baseAccount, baseAccountReader, phase
constants). It needs no change from #4258 or #4260.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tial executor

Seeded dense and sparse blocks mix hot recipients, read-modify-write
slots, hot-balance reads, nonce chains, mid-block funding, contract
creation, a create-and-destroy child, reverts, logs, and rejected
nonces. OCC results at 3 and 8 workers must equal the sequential
executor's results, including the encoded FlatKV pairs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bdchatham

Copy link
Copy Markdown
Contributor Author

@seidroid review

@cursor

cursor Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

PR Summary

High Risk
Changes core optimistic block execution, conflict detection, and final changeset construction; correctness depends on stmFrontierAccepts matching the serial frontier, though broad equivalence tests mitigate risk.

Overview
Parallelizes Block-STM validation and changeset merge while keeping block-order acceptance on the calling goroutine. The accepted prefix and write index are split into 64 address shards; speculative writes are indexed up front, and conflict checks use conflictsWithin(key, sourcePrefix, txIndex) (writes in [sourcePrefix, txIndex)) instead of comparing only against post-prefix writes.

Validation alternates parallel prefix passes (acceptValidatedPrefix / firstUnacceptedResult, bounded look-ahead, serialBackoff on dependency chains) with the existing serial frontier for the first rejected transaction. Accepted runs are folded into the prefix via applyRange / applyOwned, using per-incarnation shard sets so workers skip unrelated results.

Merge drops prefetchBaseAccounts in favor of changeSetIntoParallel, which builds per-shard changeset fragments on the OCC pool and concatenates them in canonical address order (small blocks still merge serially). A closed worker pool during merge triggers sequential fallback.

Docs add expected speedups/regressions by workload shape. Tests/benchmarks add shard and backoff coverage, a large mixed OCC-vs-sequential test, seeded OCC equivalence against sequential execution (including FlatKV encoding), validate/merge benchmarks, and an ERC20 load-test scenario.

Reviewed by Cursor Bugbot for commit 043d815. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedOct 5, 2026, 9:13 PM

@codecov

codecov Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.42379% with 15 lines in your changes missing coverage. Please review.
✅ Project coverage is 56.80%. Comparing base (880848f) to head (043d815).

Files with missing lines Patch % Lines
giga/evmonly/occ_shards.go 93.61% 9 Missing ⚠️
giga/evmonly/occ.go 95.31% 6 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #4457      +/-   ##
==========================================
+ Coverage   56.77%   56.80%   +0.03%     
==========================================
  Files        2126     2127       +1     
  Lines      166876   167018     +142     
==========================================
+ Hits        94737    94871     +134     
- Misses      72134    72142       +8     
  Partials        5        5              
Flag Coverage Δ
sei-chain 54.98% <94.42%> (+0.03%) ⬆️
sei-db 75.17% <ø> (ø)
sei-db-state-db 78.85% <ø> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
giga/evmonly/occ.go 92.67% <95.31%> (+2.71%) ⬆️
giga/evmonly/occ_shards.go 93.61% <93.61%> (ø)

... and 27 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@seidroid seidroid Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ports the parallel OCC validation and merge to main: a sharded write index filled up front, a parallel pass that accepts runs ahead of the serial frontier, sharded prefix apply, and a per-shard parallel changeset merge. Approving: the full index only ever adds conflicts (the accepted incarnation's writes are always indexed), stmFrontierAccepts matches needsSTMRerun, touchedShards covers every address applyOwned and indexResults touch, and shard concatenation keeps canonical address order. Reviewed against the PR merge tree but not built (no Go toolchain in the sandbox); the codex scout returned no reading.

1 nit, not posted on the code
  • giga/evmonly/occ_shards.go:272 — The fragments go back to occFragmentPool still holding the last block's *big.Int balances and cloned code slices. resetForReuse only runs on the next Get, so that data stays reachable until the next parallel merge. Resetting each fragment before Put would release it right away.

seidroid review · decision approve · session 40d2fa1050864bf9b6ac911728d0533d · turn resp_claude_64ad0acff67424e7ec6a8ae3f513d8c7 · item bdcfe70d0c6f5028a0d050546a331766

Findings: 0 blocking | 0 non-blocking | 0 posted inline

…pool

The pool now holds only cleared fragments, so the last block's balances and
code slices stop being reachable as soon as the merge ends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bdchatham

Copy link
Copy Markdown
Contributor Author

@seidroid review

@seidroid seidroid Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the last review, releaseOCCFragments clears each fragment before it goes back to occFragmentPool, which fixes my earlier nit about the pool keeping balances and code reachable; nothing else changed, so the earlier correctness analysis still holds and this is approved. I kept codex's finding about the gas prefix being rebuilt on every validation pass, downgraded to a suggestion because the cost is small next to EVM execution at current block sizes; this review is based on the current merge tree, not a build.

Non-blocking

1 finding on the changed lines, as inline comments.

seidroid review · decision approve · session 40d2fa1050864bf9b6ac911728d0533d · turn resp_claude_3543caff4c3abc1f1e350032a5b069cf · item b846d0d9a9915c06bd7749cf795454df

Findings: 0 blocking | 1 non-blocking | 1 posted inline

Comment thread giga/evmonly/occ_shards.go Outdated
A parallel pass computed the cumulative gas of every result to the end
of the block and allocated a new buffer for it. When conflicts are just over
64 results apart, each pass still accepts enough to reset the serial
backoff. This serial work then grows with the square of the block size.

A pass now looks ahead at most twice as far as the previous pass
accepted, and at least occMinPassLookahead (2048) results. The
cumulative-gas buffer is one per block. The accepted runs and every
frontier decision stay the same.

BenchmarkOCCValidateAndMerge adds a 20,000-transaction block with a
conflict every 65 results. The equivalence test adds two long profiles
that cross the look-ahead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bdchatham

Copy link
Copy Markdown
Contributor Author

@seidroid review

@wen-coding wen-coding left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving since this is a backport, some questions inside.

// occStateShards is the number of address shards the accepted prefix and the write index are split
// into. Shards are contiguous address ranges, so concatenating per-shard output in shard order keeps
// it in canonical address order.
const occStateShards = 64

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this constant decided? Same for constants below, how were they decided, maybe put those in comments?

@bdchatham bdchatham Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

256 and 64 copy the package's existing pool thresholds (minPrefetchedAccounts, occParallelReceiptThreshold), 64 shards fits the uint64 shard bitmask, and 2048 sizes one pass to a testnet block. Agreed the comments should say so.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 043d815: each constant now says where its value comes from (design choice, shared pool threshold, or sized for testnet blocks) and that none of them is measured for this path.


// occShardOf returns the shard that holds addr.
func occShardOf(addr common.Address) int {
return int(addr[0]) * occStateShards / 256

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How do we guarantee the loads are evenly distributed among shards?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't guarantee it: addr[0] is uniform for hash-derived addresses, but a hot contract or leading-zero vanity addresses land in one shard, which costs only speed (it degrades toward the old serial merge); sub-sharding storage by slot range, flagged back on #4261, is the fix.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 043d815: occShardOf now names the skew (hot contract, leading-zero addresses) and that it costs parallelism, not correctness. Balancing it, by slot-range sub-sharding or per-block ranges, is tracked in PLT-1379.

Comment thread giga/evmonly/README.md
between that incarnation's source prefix and its own index.

Acceptance is the serial barrier, but the work behind it is spread across the
pool: every incarnation's writes are indexed by transaction index up front, a

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I guess this will work very well on the transfer benchmark, not sure it works well for all workloads. Should we add pros and cons here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed: the gain is largest on transfer-like blocks (validate+merge −20% conflict-free, −11% sparse) and flat to slightly worse under contention (+3.8% dense, +1.7% hot recipient), so a pros/cons note with those numbers belongs here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 043d815: the Block-STM section now says when parallel acceptance helps and when it costs a few percent, with the local benchmark range.

… sharding helps

Each OCC shard constant now says whether it is a design choice, a threshold
shared with the package's other pool fan-outs, or sized for testnet blocks.
occShardOf names the skew a hot contract or leading-zero addresses cause, and
the README gives the workloads where parallel acceptance helps or costs a few
percent, with the local benchmark range.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bdchatham

Copy link
Copy Markdown
Contributor Author

@seidroid review

@seidroid seidroid Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the last review, each parallel validation pass now looks ahead at most max(2048, 2 × the previous pass's accepted count) and reuses one cumulative-gas buffer per block, so a block's passes do linear work; that addresses the suggestion I raised from codex, already marked resolved, and accept/reject decisions are unchanged. The fragment-pool fix still holds and nothing in the update breaks correctness, so this is approved; codex found nothing this round and only a nit remains, in the new benchmark.

1 nit, not posted on the code
  • giga/evmonly/occ_bench_test.go:71 — With GOMAXPROCS=1, OCCWorkers is 1, so NewExecutor never creates occPool. runRanges then calls Run on a nil pool and the benchmark panics. Clamping workers to at least 2, or skipping when the pool is nil, avoids that.

seidroid review · decision approve · session 40d2fa1050864bf9b6ac911728d0533d · turn resp_claude_5bd475cc8fd4b55b6d28452cf0ff83f3 · item e504d5e785a25288be4a182c5ee7bfd0

Findings: 0 blocking | 0 non-blocking | 0 posted inline

@bdchatham

Copy link
Copy Markdown
Contributor Author

@seidroid review

@seidroid seidroid Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing changed since the last review: same merge commit (145a203), diff, description and review standards, and the look-ahead cap and reused gas buffer behind the resolved occ_shards.go:159 thread are still in place. Approved as before, with only the GOMAXPROCS=1 benchmark nit open; codex found nothing this round, so there was nothing from it to keep or drop.

1 nit, not posted on the code
  • giga/evmonly/occ_bench_test.go:71 — With GOMAXPROCS=1, OCCWorkers is 1, so NewExecutor never creates occPool, and runRanges calls Run on a nil pool, so the benchmark panics. Clamping workers to at least 2, or skipping when the pool is nil, avoids that.

seidroid review · decision approve · session 40d2fa1050864bf9b6ac911728d0533d · turn resp_claude_9e5bffc3f586defc744ac94a9c272279 · item 87705234d5a853ef9c2cfc24e6ab9ded

Findings: 0 blocking | 0 non-blocking | 0 posted inline

@bdchatham
bdchatham added this pull request to the merge queue Oct 5, 2026
Merged via the queue into main with commit 94967eb Oct 5, 2026
86 checks passed
@bdchatham
bdchatham deleted the brandon2/port-parallel-occ-to-main branch October 5, 2026 22:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants