Skip to content

evmonly: parallelize OCC validation and merge behind the serial acceptance barrier - #4273

Closed
bdchatham wants to merge 2 commits into
devin/1789943071-stack-3-prepare-pipelinefrom
devin/1789943071-stack-4-parallel-occ
Closed

bdchatham wants to merge 2 commits into
devin/1789943071-stack-3-prepare-pipelinefrom
devin/1789943071-stack-4-parallel-occ

Conversation

@bdchatham

Copy link
Copy Markdown
Contributor

Describe your changes and provide context

Re-opens #4261 on top of #4272 (stack 4/4). With this PR applied the tree is byte-identical to today's giga-1.

OCC acceptance stays a serial barrier in block order, but the work behind it is spread across the existing worker pool (occ_shards.go):

  • every incarnation's writes are indexed by tx index up front (stateAccessIndex);
  • a parallel pass checks the pending run of transactions against that index (conflictsWithin) and reports the first one the frontier would not accept (firstUnacceptedResult); the frontier handles only that tx on the calling goroutine, then reruns from there (with a serialBackoff to avoid thrashing on hot contracts);
  • the accepted run is folded into the prefix shard by shard (occShardOf: contiguous address ranges), each incarnation recording which shards it touched so workers skip results holding nothing of theirs;
  • mergeOCCResults emits the changeset one shard at a time on the pool and concatenates shards in canonical address order, so the output is identical to the serial merge; a prefix with few keys is merged on the calling goroutine.

Determinism: acceptance decisions and the emitted changeset/receipts are the same as the serial path by construction (parity tests compare both).

Testing performed to validate your change

  • scripts/ramtest.sh -race ./giga/evmonly/... — parity vs sequential validation/merge, conflictsWithin bounds, occShardOf, cumulativeGasFrom, firstUnacceptedResult, boundary rejection at results[0].
  • make fmtcheck, make lint.

Link to Devin session: https://app.devin.ai/sessions/ff612badcded4aa5914ea408dbb41888
Open in Devin Desktop: https://app.devin.ai/desktop/session/ff612badcded4aa5914ea408dbb41888?variant=devin
Requested by: @bdchatham

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #4274 September 20, 2026 22:34
@cursor

cursor Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

PR Summary

High Risk
Changes core optimistic block execution, conflict detection, and state changeset output; correctness relies on parity with the serial path, with span-based conflicts allowing extra reruns only.

Overview
Block-STM acceptance stays serial in tx order, but indexing, bulk validation, prefix application, and changeset emission now use the existing OCC worker pool and 64-way address sharding (occ_shards.go).

The write index is built up front (per-tx, sharded) and conflicts are checked for writes in [sourcePrefix, txIndex) via conflictsWithin and first/last txIndexSpan (safe over-approximation). Parallel passes (acceptValidatedPrefix, firstUnacceptedResult) accept contiguous runs until the first tx the frontier would reject; accepted runs are folded with applyRange. serialBackoff lengthens serial stretches when parallel passes accept too little. Merge drops account prefetch in favor of changeSetIntoParallel (per-shard base reads, concatenated in canonical order; small blocks stay on the calling goroutine). Closed worker pool during merge falls back to sequential execution.

Docs/README and tests add large mixed-block OCC vs sequential parity, shard/conflict coverage, and an ERC20 loadtest benchmark scenario.

Reviewed by Cursor Bugbot for commit 1143167. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

github-actions Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedSep 21, 2026, 5:24 PM

@codecov

codecov Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.07692% with 18 lines in your changes missing coverage. Please review.
✅ Project coverage is 84.71%. Comparing base (453eb65) to head (1143167).

Files with missing lines Patch % Lines
giga/evmonly/occ_shards.go 92.59% 10 Missing ⚠️
giga/evmonly/occ.go 93.60% 8 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@                              Coverage Diff                              @@
##           devin/1789943071-stack-3-prepare-pipeline    #4273      +/-   ##
=============================================================================
+ Coverage                                      83.02%   84.71%   +1.69%     
=============================================================================
  Files                                             30       27       -3     
  Lines                                           2992     4357    +1365     
=============================================================================
+ Hits                                            2484     3691    +1207     
- Misses                                           507      666     +159     
+ Partials                                           1        0       -1     
Flag Coverage Δ
sei-chain-pr 84.71% <93.07%> (+0.34%) ⬆️
sei-db ?

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
giga/evmonly/occ.go 91.31% <93.60%> (ø)
giga/evmonly/occ_shards.go 92.59% <92.59%> (ø)

... and 55 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The parallel OCC validation and sharded merge look correct: the pre-built span index is equivalent to the old accept-as-you-go index for the only two sourcePrefix values that occur, stmFrontierAccepts reproduces the serial accept condition, shard ownership is disjoint and totally covering, and shard order preserves canonical address order. Only two non-blocking efficiency observations.

Findings: 0 blocking | 2 non-blocking | 1 posted inline

Blockers

  • None at the file/PR level.

Non-blocking

  • [suggestion] occShardOf keys on addr[0] only, which is required for order-preserving concatenation but means a single-hot-contract workload puts all of that contract's storage into one shard. The erc20_single_contract benchmark this PR adds is exactly that shape, so the sharded merge will parallelize little there — worth calling out in the README alongside the existing shard description so the next reader doesn't expect merge speedup on single-contract blocks.
  • 1 suggestion(s)/nit(s) flagged inline on specific lines.

if len(results)-from < occMinParallelValidation {
return 0, nil
}
cumulative, to := cumulativeGasFrom(results, from, state.cumulativeGasUsed)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] cumulativeGasFrom walks and allocates for the entire remaining block (len(results)-from+1 entries) on every parallel pass, on the calling goroutine — i.e. inside the serial barrier this PR is trying to shrink.

The scan in firstUnacceptedResult is self-limiting (workers bail once stop drops below their index), so a pass typically advances only a short distance, but the prefix sum is paid in full regardless. With one conflict roughly every occMinParallelValidation+ transactions the backoff never engages, so you get ~C passes each doing O(N) serial work: O(C·N) total. At N=100k with C≈1k that is ~100M writes and ~800MB of allocation churn per block, all serial.

Capping the pass window (to = min(to, from+window)) would bound both the prefix sum and the atomic stop.Load() contention, at the cost of an extra pass on wide conflict-free runs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changed here by design: this stack re-opens the already-merged code unmodified so the team can review what actually runs on giga-1 (the stack top equals current giga-1). Noted as a follow-up.

@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789943071-stack-4-parallel-occ branch from a004fe1 to a96c76d Compare September 20, 2026 22:50
@bdchatham bdchatham changed the title [stack 4/4] evmonly: parallelize OCC validation and merge behind the serial acceptance barrier evmonly: parallelize OCC validation and merge behind the serial acceptance barrier Sep 20, 2026
@devin-ai-integration
devin-ai-integration Bot removed this pull request from stack #4274 September 20, 2026 23:18
…the serial acceptance barrier (#4261)""

This reverts commit 1a6a64f.
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789943071-stack-4-parallel-occ branch from a96c76d to 971643c Compare September 20, 2026 23:18
@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #4275 September 20, 2026 23:18
Comment thread giga/evmonly/occ.go Outdated
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@suryuhh

suryuhh commented Sep 30, 2026

Copy link
Copy Markdown

I tested this PR at commit 11431679 with recorded Sei and Ethereum blocks. On blocks with many dependent transactions, its parallel path was slower than sequential execution. #4392 samples up to 256 transactions and uses sequential execution when dependencies are high. Its tests compare receipts and state hashes with sequential execution; the PR links the benchmark inputs and commands. #4393 is a separate, larger Block-STM follow-up for blocks with independent chains. It uses more CPU per transaction. Could you review the smaller routing change against this branch?

@bdchatham

Copy link
Copy Markdown
Contributor Author

Superseded by #4457, a backport of this change (#4261) to main. This stack targeted stacked giga-1 branches and could not merge to main; #4457 is based on main, resolves the conflicts with main's renames, and adds two review fixes (merge fragments are cleared before returning to the pool, and each parallel pass's look-ahead is bounded).

@bdchatham bdchatham closed this Oct 5, 2026
revofusion pushed a commit to revofusion/sei-chain that referenced this pull request Oct 6, 2026
… acceptance barrier (sei-protocol#4457)

Backport: this ports sei-protocol#4273, the reviewable version of sei-protocol#4261, to main
from the giga-1 PR stack sei-protocol#4270–sei-protocol#4273. The stack targeted stacked giga-1
branches and never reached main; only its first PR landed on main, as
sei-protocol#4366. @shemnon approved sei-protocol#4273 at 1143167. Since that commit, the
changes are main's renames (`baseAccount`, the `phaseOCC*` constants,
`takeAccessSets`) and two review fixes: merge fragments are cleared
before they return to the pool, and each parallel pass's look-ahead is
bounded.

On giga-testnet-2, `occ_validate` and `occ_merge` took about a third of
the executor main loop, all of it on the calling goroutine. Most of that
time went to transactions that the frontier accepts as they stand. A
block of about 1,800 transactions holds less than one conflict. On
`giga-1`, sei-protocol#4261 moved this work onto the OCC pool. It cut validate and
merge from about 8.5 ms to 4.1 ms per block. That gave about 20% more
blocks per second during the 95k tx/s window. sei-protocol#4269 reverted it for the
reviewable stack, and that stack (sei-protocol#4273) never reached main. This PR
ports sei-protocol#4261 and the sei-protocol#4273 follow-up to main (PLT-1377).

Acceptance stays a serial barrier in block order. `stateAccessIndex` and
`blockSTMState` split into 64 address shards. `indexResults` records the
writes of every speculative result up front, and `conflictsWithin(key,
sourcePrefix, txIndex)` replaces `conflictsWithAfter`.
`acceptValidatedPrefix` finds the first result that the serial frontier
would not accept, and `applyRange` folds the run before it into the
prefix shard by shard. `validateBlockSTMFrontier` then handles only that
one result, and `serialBackoff` keeps dependency chains on the calling
goroutine. Each pass looks ahead at most twice as far as the previous
pass accepted, and at least 2,048 results, so the work of a block's
passes stays linear in its size. `changeSetIntoParallel` builds the
changeset of each shard on the pool and joins the shards in address
order. Its per-shard base-row cache replaces `prefetchBaseAccounts`. The
port needs nothing from sei-protocol#4258 or sei-protocol#4260; the only conflicts came from the
`baseAccount` and phase-constant renames on main.

Non-app-hash-breaking: changesets, receipts, and tx results stay
byte-identical to main. Shards are contiguous ranges of the first
address byte, so shard order is canonical address order. Each address
lives in one shard, so the apply order in a shard is block order. The
new index can only add conflicts, from stale writes of an earlier
incarnation. An extra rerun executes against the exact accepted prefix,
so it gives the sequential result, and the extra reruns show only in OCC
metrics. Review `stmFrontierAccepts` most closely, because it must
mirror `needsSTMRerun`, and `touchedShards`, because it must cover every
address that `applyOwned` and `indexResults` touch.
`TestOCCRandomizedConflictingBlocksMatchSequential` checks seeded dense
and sparse blocks at 3 and 8 workers against the sequential executor,
down to the encoded FlatKV pairs. A seeded digest run gave identical
output on main and this branch for 864 blocks, including 3,000- and
6,000-transaction blocks that cross a pass's look-ahead.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants