Skip to content

Backport release/v6.7: Flush MemIAVL changelog before exiting on an upgrade panic - #4292

Merged
alexander-sei merged 1 commit into
release/v6.7from
backport-4287-to-release/v6.7
Sep 22, 2026
Merged

alexander-sei merged 1 commit into
release/v6.7from
backport-4287-to-release/v6.7

Conversation

@seidroid

@seidroid seidroid Bot commented Sep 22, 2026

Copy link
Copy Markdown

Backport of #4287 to release/v6.7.

@seidroid seidroid Bot assigned masih Sep 22, 2026
@seidroid seidroid Bot added the backport label Sep 22, 2026
@seidroid

seidroid Bot commented Sep 22, 2026

Copy link
Copy Markdown
Author

Please cherry-pick the changes locally and resolve any conflicts.

git fetch origin backport-4287-to-release/v6.7
git worktree add --checkout .worktree/backport-4287-to-release/v6.7 backport-4287-to-release/v6.7
cd .worktree/backport-4287-to-release/v6.7
git reset --hard HEAD^
git cherry-pick -x 53a681b964a4b2a22f202040dc72929281c8c1be
git push --force-with-lease

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedSep 22, 2026, 8:14 AM

`TestUpgradeMajor` fails when the last validator to upgrade comes back
with `RUNNING_UPGRADED_NODE_3=FAIL`. MemIAVL writes its changelog WAL
asynchronously (`sc-async-commit-buffer` defaults to 100), so `Commit`
returns before the entry for that block is on disk. A node that is
catching up executes the block before the upgrade height and the upgrade
height back to back, and the `UPGRADE "v2.0.0" NEEDED` panic kills the
process before the writer goroutine has flushed the previous block.
FlatKV, whose block WAL is flushed synchronously in `Commit`, comes back
at height 149 while MemIAVL comes back at 148;
`CompositeCommitStore.reconcileVersions` resolves the disagreement by
rolling FlatKV back to 148. The upgraded binary then re-executes block
149, one below the plan height, and `x/upgrade` correctly panics with
`BINARY UPDATED BEFORE TRIGGER`, so every restart attempt in
`seid_upgrade.sh` dies within its one-second grace period and
`verify_running.sh` never sees a live process. In the failed run node
3's log shows exactly this: MemIAVL and FlatKV both at 148 with the
early-upgrade panic on every attempt, while nodes 1 and 2 restarted at
149.

The fix makes the intentional upgrade exit durable. `sctypes.Committer`
gains `Flush()`, which blocks until every committed version is
persisted; `memiavl.DB.Flush` reuses the wait that
`checkBackgroundSnapshotRewrite` already performed inline (now
`waitForPendingWALWrites`), `CompositeCommitStore.Flush` delegates to
MemIAVL since FlatKV already persists synchronously, and
`rootmulti.Store.Flush` exposes it to the app. `App.ProcessBlock` calls
`flushCommittedStateForUpgradeExit` right before re-panicking on an
upgrade panic, so the process only dies once the last committed block is
in every backend's WAL. Closing the multistore instead was rejected
because the in-process `upgrade_v67` tests recover the panic and keep
reading state from the same app.

Flaked in:
https://github.com/sei-protocol/sei-chain/actions/runs/35609798417/job/106367298402

(cherry picked from commit 53a681b)
@codecov

codecov Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 47.91667% with 25 lines in your changes missing coverage. Please review.
✅ Project coverage is 60.90%. Comparing base (2cf6111) to head (d69d1d0).

Files with missing lines Patch % Lines
sei-db/state_db/sc/memiavl/db.go 67.85% 5 Missing and 4 partials ⚠️
sei-db/state_db/sc/composite/store.go 0.00% 8 Missing ⚠️
sei-db/state_db/sc/memiavl/store.go 0.00% 4 Missing ⚠️
app/app.go 50.00% 1 Missing and 1 partial ⚠️
sei-db/state_db/sc/flatkv/store.go 0.00% 2 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@               Coverage Diff                @@
##           release/v6.7    #4292      +/-   ##
================================================
- Coverage         61.37%   60.90%   -0.48%     
================================================
  Files              2163     2105      -58     
  Lines            189024   183786    -5238     
================================================
- Hits             116017   111933    -4084     
+ Misses            62280    61523     -757     
+ Partials          10727    10330     -397     
Flag Coverage Δ
sei-chain-pr 60.80% <66.66%> (?)
sei-db 70.02% <ø> (+0.21%) ⬆️
sei-db-state-db ?
sei-db-state-db-pr 76.24% <45.23%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
sei-cosmos/storev2/rootmulti/store.go 69.52% <100.00%> (+0.08%) ⬆️
app/app.go 71.71% <50.00%> (+0.25%) ⬆️
sei-db/state_db/sc/flatkv/store.go 80.04% <0.00%> (-0.39%) ⬇️
sei-db/state_db/sc/memiavl/store.go 88.81% <0.00%> (-2.56%) ⬇️
sei-db/state_db/sc/composite/store.go 70.53% <0.00%> (-0.90%) ⬇️
sei-db/state_db/sc/memiavl/db.go 70.85% <67.85%> (+0.54%) ⬆️

... and 61 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@masih
masih marked this pull request as ready for review September 22, 2026 08:28
@cursor

cursor Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Touches commit-store durability and the upgrade exit path; incorrect flush behavior could affect restart correctness, though scope is limited to upgrade panics and WAL catch-up with tests.

Overview
Ensures committed state is persisted to backend WALs before the node exits when an upgrade is triggered via panic, so a restart after Cosmovisor does not lose blocks that were committed in memory but not yet written to the MemIAVL changelog.

Adds a Flush() path on the state-commit stack (Committer → composite → memIAVL / flatKV; flatKV is a no-op because commit is synchronous). MemIAVL Flush() blocks until the changelog WAL reflects every committed version (10s timeout), and snapshot rewrite now reuses shared waitForPendingWALWrites logic.

On upgrade-related panics in ProcessBlock, the app calls flushCommittedStateForUpgradeExit() (via rootStore.Flush()) before re-panicking. An integration test asserts that after an UPGRADE NEEDED panic, the on-disk MemIAVL version matches the last successfully committed height.

Reviewed by Cursor Bugbot for commit d69d1d0. Bugbot is set up for automated code reviews on this repo. Configure here.

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adds a Flush() path through the SC commit-store stack and invokes it from ProcessBlock's upgrade-panic re-panic so the async memiavl changelog WAL catches up before the process exits. The mechanism is sound (no lock inversion with the WAL writer goroutine, bounded by a 10s deadline, flatkv's no-op Flush is justified by its synchronous commit-time WAL flush); the notes below are non-blocking.

Findings: 0 blocking | 3 non-blocking | 1 posted inline

Blockers

  • None at the file/PR level.

Non-blocking

  • [suggestion] Neither new test is likely to fail without the production change. In app/upgrade_exit_flush_test.go, the whole of FinalizeBlock(3) (BeginBlock + upgrade panic) runs between the last Commit() and the GetLatestVersion assertion, so the async WAL writer has almost certainly drained regardless of the flush. TestFlushWaitsForAsyncWALWrites has the same shape — it asserts a post-condition the writer reaches on its own. Both are useful smoke tests, but they don't guard the regression; making the WAL writer demonstrably lag (or asserting the flush was reached through a seam) would.
  • 1 suggestion(s)/nit(s) flagged inline on specific lines.
  • 1 non-blocking pre-existing issue(s) listed below under pre-existing issues.

Pre-existing issues

  • [suggestion] checkBackgroundSnapshotRewrite's wait for the WAL to catch up (now extracted as waitForPendingWALWrites in sei-db/state_db/sc/memiavl/db.go) is unbounded and busy-spins with a 1ns sleep, all while Commit holds db.mtx. A wedged or errored-then-stalled WAL writer hangs the consensus goroutine indefinitely. This PR gives the new Flush a flushTimeout; the older path has no equivalent.

Inline comments (could not post inline; listed here)

  • sei-db/state_db/sc/composite/store.go:1409 (RIGHT) -- [suggestion] Flush returns on the first backend error, so a memIAVL failure (e.g. the new flushTimeout firing) skips the flatKV flush entirely. Close just below deliberately joins errors from both backends instead. Today flatKV's Flush is a no-op so there's no behavioural impact, but the interface contract allows a real implementation, and this is an exit path where best-effort on both backends is what you want. Consider collecting into errs and joining, matching Close.

@alexander-sei
alexander-sei merged commit daea498 into release/v6.7 Sep 22, 2026
79 of 81 checks passed
@alexander-sei
alexander-sei deleted the backport-4287-to-release/v6.7 branch September 22, 2026 08:36
masih pushed a commit that referenced this pull request Sep 22, 2026
Adds the `release/v6.7` entries merged since the rc1 changelog (#4110),
in prep to cut **v6.7.0-rc2**:

- [#4292](#4292) — Flush
MemIAVL changelog before exiting on an upgrade panic
- [#4285](#4285) —
fix(seidb): refuse a corrupted changelog in digest replay instead of
repairing it
- [#4255](#4255) —
feat(seidb): Add JSON output to evm-logical-digest and inspect a FlatKV
migration in flight
- [#4116](#4116) — rc1
version bump
- [#4113](#4113) — rc1
changelog backport

Regenerated with `./scripts/generate-changelog.sh release/v6.6
release/v6.7`; only the `## v6.7` PR list changes, so the `backport
release/v6.7` cherry-pick applies cleanly (verified with `git apply
--check` against `origin/release/v6.7`). Docs-only; no code change.

Made with [Cursor](https://cursor.com)

Co-authored-by: Cursor <cursoragent@cursor.com>
alexander-sei added a commit that referenced this pull request Sep 22, 2026
## Summary
- Bump `version.json` from `v6.7.0-rc1` to `v6.7.0-rc2` to cut the
second `v6.7` release candidate.

Contents since rc1: #4285, #4255, #4292 (all `sei-db` fixes/tooling),
plus the rc2 changelog update (#4294). Merge after #4294 so the tag cut
by `uci-release-publish` on this `version.json` change includes the
updated changelog.

## Test plan
- [x] `git diff --check`

Made with [Cursor](https://cursor.com)

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants