Skip to content

evmonlyapp: land a block's state commit behind FinalizeBlock - #4366

Merged
bdchatham merged 3 commits into
mainfrom
devin/1790610643-plt-1312-commit-behind-finalize
Sep 28, 2026
Merged

bdchatham merged 3 commits into
mainfrom
devin/1790610643-plt-1312-commit-behind-finalize

Conversation

@bdchatham

@bdchatham bdchatham commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The EVM-only executor on main can already pipeline commits (PrepareBlock, ExecutePreparedBlock, AwaitCommits, from #4317), but evmOnlyApplication.FinalizeBlock still calls the synchronous ExecuteBlock. That call waits for the FlatKV commit of block N before it returns, so the node can't start N+1 until the write lands. This PR brings the application-layer half of #4270 over from giga-1 (PLT-1312), adapted to main's evmOnlyState / sender-cache layout instead of the older cursor/executor split that #4270 assumes.

FinalizeBlock now runs PrepareBlock and then ExecutePreparedBlock (in executeBlockPipelined), and returns once the state commit has started. The next block still waits on the previous commit inside the executor before it starts its own. Committed-state readers settle the in-flight commit first. EvmNonce, EvmBalance and EvmCode go through openSettledView, which calls AwaitCommits on the executor published through a lock-free settler, so they never block on FinalizeBlock. currentExecutionContext, used by EvmCall and EvmEstimateGas, awaits while it holds state, so no new block can start between the settle and the snapshot, and it returns a failed commit as an error. openSettledView logs a failed commit once and serves the last version that landed. Its return values have no error channel, and the failure halts the node through the next FinalizeBlock. Proxy.AwaitCommits exposes the settle, and nodeImpl.closeGigaStorage calls it before manager.Close() so shutdown never closes the store under an in-flight write.

Consensus and app hash are unchanged: the app hash is still computed from the execution result, not the store write. The behaviour change is that a failed store commit now surfaces one block later, from the next FinalizeBlock, rather than from its own. Between FinalizeBlock and Commit, EvmNonce and friends already see the finalized block's state; before this PR they could not, because the write had already landed by then anyway. New tests cover reads settling across four pipelined blocks and a commit failure surfacing from the next block, from EvmCall and from AwaitCommits, and a -race test that drives EvmNonce/EvmCall from a goroutine while blocks are finalized and committed, asserting the nonce never goes backwards. The test helpers now settle before closing storage. I validated with scripts/ramtest.sh -race over evmonlyapp, proxy, node and giga/evmonly, plus make fmtcheck and make lint.

Link to Devin session: https://app.devin.ai/sessions/d808db90106341278344623de10b92c0
Open in Devin Desktop: https://app.devin.ai/desktop/session/d808db90106341278344623de10b92c0?variant=devin
Requested by: @bdchatham

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedSep 28, 2026, 4:24 PM

@cursor

cursor Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

PR Summary

High Risk
Changes core block finalization, async state persistence, and RPC read semantics; commit failures are deferred one block, which can affect halt timing and operator debugging.

Overview
FinalizeBlock no longer waits for the FlatKV write to finish. It runs PrepareBlock then ExecutePreparedBlock and returns once the state commit has started, so store I/O can overlap with the next block while the executor still serializes commits between heights.

Committed-state RPC paths (EvmNonce, EvmBalance, EvmCode) go through openSettledView, which AwaitCommits via a lock-free settler before opening a view—so callers see finalized state without blocking FinalizeBlock, and failed commits are logged once while serving the last landed version. EvmCall / gas estimation AwaitCommits under the app state lock so the next block cannot race the settle.

Proxy.AwaitCommits exposes settlement for shutdown; closeGigaStorage awaits in-flight commits before closing the manager. App hash / consensus behavior is unchanged; a failed store commit now surfaces on the next FinalizeBlock (or read/settle paths) rather than on the block that started the write.

Tests cover pipelined reads, deferred commit failure, concurrent reads vs finalize/commit, and test teardown that settles before closing storage. OTel adds a prepare finalize phase.

Reviewed by Cursor Bugbot for commit fa79805. Bugbot is set up for automated code reviews on this repo. Configure here.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR changes FinalizeBlock to run PrepareBlock and then ExecutePreparedBlock, so the block's state commit can land in the background. Readers of committed state wait for that commit before reading: through the atomically published executor in openSettledView, and under the state lock in currentExecutionContext. Shutdown also waits for the commit before closing Giga storage. I found no issues: the executor's AwaitCommits is safe to call from several goroutines, closeGigaStorage runs only after giga.Run has returned, and the new tests cover pipelined reads, a failed commit surfacing from the next block, and reads racing FinalizeBlock.

Findings: 0 blocking | 0 non-blocking | 0 posted inline

Blockers

  • None at the file/PR level.

Non-blocking

  • None at the file/PR level.

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 70.96774% with 9 lines in your changes missing coverage. Please review.
✅ Project coverage is 66.54%. Comparing base (31f473b) to head (fa79805).

Files with missing lines Patch % Lines
sei-tendermint/internal/proxy/proxy.go 0.00% 5 Missing ⚠️
sei-tendermint/internal/evmonlyapp/app.go 91.66% 2 Missing ⚠️
sei-tendermint/node/node.go 0.00% 2 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #4366      +/-   ##
==========================================
- Coverage   67.70%   66.54%   -1.16%     
==========================================
  Files        2168     2045     -123     
  Lines      168283   156003   -12280     
==========================================
- Hits       113928   103818   -10110     
+ Misses      54346    52176    -2170     
  Partials        9        9              
Flag Coverage Δ
sei-chain-pr 76.22% <70.96%> (?)
sei-db 74.75% <ø> (+0.24%) ⬆️
sei-db-state-db ?

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
sei-tendermint/internal/evmonlyapp/app.go 86.86% <91.66%> (+0.79%) ⬆️
sei-tendermint/node/node.go 74.73% <0.00%> (-0.33%) ⬇️
sei-tendermint/internal/proxy/proxy.go 89.65% <0.00%> (-5.47%) ⬇️

... and 163 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit fa79805. Configure here.

}
// The executor's own timer breaks execution down further.
a.finalizePhases.SetPhase(finalizePhaseExecute)
return executor.ExecutePreparedBlock(ctx, prepared)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale execution after settled commit failure

Medium Severity

FinalizeBlock no longer waits for its own store write, so a failed commit is first observed on the next block. openSettledView can settle that failure first and clear the executor overlay. The next executeBlockPipelined then runs against the last version that landed, and the executor can persist receipts for that execution before returning the latched error.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit fa79805. Configure here.

@shemnon shemnon left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@bdchatham
bdchatham added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit b2baf9f Sep 28, 2026
66 checks passed
@bdchatham
bdchatham deleted the devin/1790610643-plt-1312-commit-behind-finalize branch September 28, 2026 20:20
revofusion pushed a commit to revofusion/sei-chain that referenced this pull request Oct 6, 2026
… acceptance barrier (sei-protocol#4457)

Backport: this ports sei-protocol#4273, the reviewable version of sei-protocol#4261, to main
from the giga-1 PR stack sei-protocol#4270–sei-protocol#4273. The stack targeted stacked giga-1
branches and never reached main; only its first PR landed on main, as
sei-protocol#4366. @shemnon approved sei-protocol#4273 at 1143167. Since that commit, the
changes are main's renames (`baseAccount`, the `phaseOCC*` constants,
`takeAccessSets`) and two review fixes: merge fragments are cleared
before they return to the pool, and each parallel pass's look-ahead is
bounded.

On giga-testnet-2, `occ_validate` and `occ_merge` took about a third of
the executor main loop, all of it on the calling goroutine. Most of that
time went to transactions that the frontier accepts as they stand. A
block of about 1,800 transactions holds less than one conflict. On
`giga-1`, sei-protocol#4261 moved this work onto the OCC pool. It cut validate and
merge from about 8.5 ms to 4.1 ms per block. That gave about 20% more
blocks per second during the 95k tx/s window. sei-protocol#4269 reverted it for the
reviewable stack, and that stack (sei-protocol#4273) never reached main. This PR
ports sei-protocol#4261 and the sei-protocol#4273 follow-up to main (PLT-1377).

Acceptance stays a serial barrier in block order. `stateAccessIndex` and
`blockSTMState` split into 64 address shards. `indexResults` records the
writes of every speculative result up front, and `conflictsWithin(key,
sourcePrefix, txIndex)` replaces `conflictsWithAfter`.
`acceptValidatedPrefix` finds the first result that the serial frontier
would not accept, and `applyRange` folds the run before it into the
prefix shard by shard. `validateBlockSTMFrontier` then handles only that
one result, and `serialBackoff` keeps dependency chains on the calling
goroutine. Each pass looks ahead at most twice as far as the previous
pass accepted, and at least 2,048 results, so the work of a block's
passes stays linear in its size. `changeSetIntoParallel` builds the
changeset of each shard on the pool and joins the shards in address
order. Its per-shard base-row cache replaces `prefetchBaseAccounts`. The
port needs nothing from sei-protocol#4258 or sei-protocol#4260; the only conflicts came from the
`baseAccount` and phase-constant renames on main.

Non-app-hash-breaking: changesets, receipts, and tx results stay
byte-identical to main. Shards are contiguous ranges of the first
address byte, so shard order is canonical address order. Each address
lives in one shard, so the apply order in a shard is block order. The
new index can only add conflicts, from stale writes of an earlier
incarnation. An extra rerun executes against the exact accepted prefix,
so it gives the sequential result, and the extra reruns show only in OCC
metrics. Review `stmFrontierAccepts` most closely, because it must
mirror `needsSTMRerun`, and `touchedShards`, because it must cover every
address that `applyOwned` and `indexResults` touch.
`TestOCCRandomizedConflictingBlocksMatchSequential` checks seeded dense
and sparse blocks at 3 and 8 workers against the sequential executor,
down to the encoded FlatKV pairs. A seeded digest run gave identical
output on main and this branch for 864 blocks, including 3,000- and
6,000-transaction blocks that cross a pass's look-ahead.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants