Skip to content

fix(seidb): keep writes in the old DB until the migration boundary first moves - #4369

Merged
blindchaser merged 3 commits into
mainfrom
fix/migrate-evm-paused-routing
Sep 28, 2026
Merged

blindchaser merged 3 commits into
mainfrom
fix/migrate-evm-paused-routing

Conversation

@blindchaser

@blindchaser blindchaser commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

In MigrateEVM, flatkv's lattice hash joins the AppHash only after the migration boundary first moves (shouldAppendLatticeHash). But MigrationManager.shouldForwardWriteToNewDB sent new keys to flatkv before that. With batch size 0 (the NumKeysToMigratePerBlock default), the boundary never moves, so these writes were committed but not hashed. Only a node pinned to sc-write-mode = migrate_evm can reach this state. Auto switches modes only with a positive batch size, and that block's first batch moves the boundary.

Now, while the boundary is NotStarted, every caller write goes to the old DB. Thus flatkv stays empty until the boundary moves, which is what the gate assumes. The iterator migrates keys written during the pause, and after the migration starts, new keys go to flatkv again.

Upgrade note: the AppHash changes only for nodes pinned to migrate_evm, migrate_all_but_bank, or migrate_bank with a paused migration that has not started. Auto nodes and mainnet do not change. Pinned networks in that state must upgrade together, and a node that ran the old binary in that state must resync. New tests in migration_manager_test.go and store_migration_test.go cover the paused routing, the first-batch ordering, a pause mid-migration (also across a restart), and AppHash parity with MemiavlOnly.

…rst moves

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedSep 28, 2026, 7:43 PM

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 66.81%. Comparing base (0a017df) to head (fc94de0).

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #4369      +/-   ##
==========================================
- Coverage   67.70%   66.81%   -0.90%     
==========================================
  Files        2168     2064     -104     
  Lines      168283   157912   -10371     
==========================================
- Hits       113938   105509    -8429     
+ Misses      54336    52394    -1942     
  Partials        9        9              
Flag Coverage Δ
sei-db 74.50% <ø> (ø)
sei-db-state-db ?
sei-db-state-db-pr 86.31% <100.00%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
sei-db/state_db/sc/migration/migration_manager.go 97.59% <100.00%> (+0.02%) ⬆️

... and 104 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

blindchaser and others added 2 commits September 28, 2026 14:50
Co-authored-by: Cursor <cursoragent@cursor.com>
@blindchaser
blindchaser marked this pull request as ready for review September 28, 2026 18:58
@cursor

cursor Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

PR Summary

High Risk
Changes write routing and AppHash for pinned migrate modes with a paused, not-started migration—consensus-critical; affected validator sets must coordinate upgrades and resync if they already committed unhashed flatkv writes.

Overview
Fixes a consensus bug where MigrateEVM nodes with migration batch size 0 could commit EVM writes to flatkv while evm_lattice was still excluded from the AppHash, because shouldForwardWriteToNewDB routed brand-new keys to the new DB before the boundary ever moved.

MigrationManager.shouldForwardWriteToNewDB now keeps all caller writes on the old DB while the boundary is MigrationNotStarted; after the first batch advances the boundary, routing is unchanged (migrated / absent-in-old → new DB). Unit tests cover pre-start writes, first-batch ordering, and pause-after-start; composite tests assert AppHash parity with MemiavlOnly when paused before the boundary and that lattice hashes update when paused mid-migration (including across restart).

Upgrade note: AppHash changes only for networks pinned to a migrate write mode with a not-yet-started paused migration; affected nodes must upgrade together and may need resync if they ran the old binary in that state.

Reviewed by Cursor Bugbot for commit fc94de0. Bugbot is set up for automated code reviews on this repo. Configure here.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T19:04:49.369478Z fc94de0 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fix is placed at the one function every caller write passes through, shouldForwardWriteToNewDB. While the boundary is NotStarted, all writes now go to the old DB, which matches the rule that the lattice hash joins the AppHash only once the boundary first moves. ApplyChangeSets advances the boundary before it routes caller writes, so the first batch correctly switches routing to the started behaviour. Read still falls back to the new DB. The tests cover a paused migration that never started, the block that starts the migration, a pause mid-migration (including across a restart), and AppHash parity with MemiavlOnly. I found no issues.

Findings: 0 blocking | 0 non-blocking | 0 posted inline

Blockers

  • None at the file/PR level.

Non-blocking

  • None at the file/PR level.

@blindchaser blindchaser added the backport release/v6.7 Backport to release v6.7 label Sep 28, 2026
@blindchaser
blindchaser added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit b32f042 Sep 28, 2026
70 of 72 checks passed
@blindchaser
blindchaser deleted the fix/migrate-evm-paused-routing branch September 28, 2026 19:57
@seidroid

seidroid Bot commented Sep 28, 2026

Copy link
Copy Markdown

Created backport PR for release/v6.7:

Please cherry-pick the changes locally and resolve any conflicts.

git fetch origin backport-4369-to-release/v6.7
git worktree add --checkout .worktree/backport-4369-to-release/v6.7 backport-4369-to-release/v6.7
cd .worktree/backport-4369-to-release/v6.7
git reset --hard HEAD^
git cherry-pick -x b32f0425b242152d16d3eea9cb7df73989e4be72
git push --force-with-lease

blindchaser added a commit that referenced this pull request Sep 28, 2026
…rst moves (#4369)

In `MigrateEVM`, flatkv's lattice hash joins the AppHash only after the
migration boundary first moves (`shouldAppendLatticeHash`). But
`MigrationManager.shouldForwardWriteToNewDB` sent new keys to flatkv
before that. With batch size 0 (the `NumKeysToMigratePerBlock` default),
the boundary never moves, so these writes were committed but not hashed.
Only a node pinned to `sc-write-mode = migrate_evm` can reach this
state. Auto switches modes only with a positive batch size, and that
block's first batch moves the boundary.

Now, while the boundary is `NotStarted`, every caller write goes to the
old DB. Thus flatkv stays empty until the boundary moves, which is what
the gate assumes. The iterator migrates keys written during the pause,
and after the migration starts, new keys go to flatkv again.

Upgrade note: the AppHash changes only for nodes pinned to
`migrate_evm`, `migrate_all_but_bank`, or `migrate_bank` with a paused
migration that has not started. Auto nodes and mainnet do not change.
Pinned networks in that state must upgrade together, and a node that ran
the old binary in that state must resync. New tests in
`migration_manager_test.go` and `store_migration_test.go` cover the
paused routing, the first-batch ordering, a pause mid-migration (also
across a restart), and AppHash parity with `MemiavlOnly`.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit b32f042)
blindchaser added a commit that referenced this pull request Sep 29, 2026
…the migration boundary first moves (#4371)

Backport of #4369 to `release/v6.7`.

Co-authored-by: yirenz <blindchaser@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
masih pushed a commit that referenced this pull request Oct 1, 2026
## Summary
- Bump `version.json` from `v6.7.0-rc3` to `v6.7.0-rc4` to cut the
fourth `v6.7` release candidate.

Contents since rc3: #4411, #4410, #4409, #4397, #4378, #4377, #4371,
plus the rc4 changelog update (#4415, with the conflict-marker fix
#4416). The changelog has already landed, so the `v6.7.0-rc4` tag will
include it.

- All seven are labeled `non-app-hash-breaking`. #4371 changes the
AppHash only for nodes pinned to a `migrate_*` write mode while the
migration hasn't started (see #4369), and the hard-fork handlers added
by #4409 and #4410 are not registered for any chain in this release.
Unlike rc3, moving from rc3 to rc4 should not need a coordinated
validator switch.
- Push `v6.7.0-rc4` by hand on this PR's merge commit once it has
merged; the tagging ruleset stops `uci-release-publish` from creating
it. The rc3 tag was pushed before #4339 merged and sits on `0d56aaeef`,
where `version.json` still reads `v6.7.0-rc2`.
- #4378 raises GoReleaser's timeout to 2h. The rc3 tag push hit the 1h
limit during the emulated arm64 build and attached nothing, so rc4 is
the first `v6.7` release candidate that should get binaries.

## Test plan
- [x] `git diff --check`

Made with [Cursor](https://cursor.com)

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants