Skip to content

Backport release/v6.7: fix(seidb): refuse a corrupted changelog in digest replay instead of repairing it - #4282

Closed
blindchaser wants to merge 1 commit into
release/v6.7from
backport-3983-to-release/v6.7
Closed

blindchaser wants to merge 1 commit into
release/v6.7from
backport-3983-to-release/v6.7

Conversation

@blindchaser

Copy link
Copy Markdown
Contributor

Backport of #3983 to release/v6.7.

Why this goes to v6.7

release/v6.7 already ships evm-logical-digest --memiavl-open-mode replay, and on this branch that open path still repairs a changelog it cannot read. The repair is a truncation, and a record a running node is midway through writing is indistinguishable from a corrupt one, so pointing the tool at a live v6.7 node can discard a block that node has committed.

#4255 (backport of #4166) widens who reaches that path: inspect mode and the composite backend both start accepting --memiavl-open-mode replay. Land this PR before #4255.

Conflicts

The cherry-pick applied without conflicts. One change was needed on top:

  • sei-db/tools/cmd/seidb/operations/memiavl_open_test.go — CommitStore.Commit takes a version argument on main and no argument on release/v6.7, so store.Commit(store.Version() + 1) becomes store.Commit(). This matches how the rest of the package already calls it on this branch.

No CHANGELOG entry, following the convention on this branch since the v6.7 changelog cut (#4113).

Tests

scripts/ramtest.sh ./sei-db/wal/... ./sei-db/tools/cmd/seidb/operations/... ./sei-db/state_db/sc/memiavl/...

All three packages pass. Both new tests run and pass on this branch:

  • TestOpenMemiAVLReplayReadOnlyRefusesATornChangelogWithoutTruncatingIt
  • TestOpenMemiAVLReplayReadOnlyReplaysAnIntactChangelog

go build ./sei-db/..., gofmt -s -l, and goimports -l are clean.

…repairing it (#3983)

## Summary

`seidb evm-logical-digest --memiavl-open-mode replay` opens the
changelog through the same opener a writer uses, and that opener repairs
a tail ending mid-record by truncating it. `Options.ReadOnly` does not
prevent this: it gates the DB API, not the changelog open.

On a live node the repair fires on a healthy log. A tail ending
mid-record is usually the writer mid-append — `write` is not atomic
against a concurrent reader, so a reader sees a page-aligned prefix of a
multi-page changeset. Truncating it discards a record `seid` has
committed, and the writer's descriptor keeps its old offset, leaving a
zero-filled hole the binary decoder reads as valid zero-length records.
Nothing fails at the time; it surfaces when the node next reopens and
replays.

Nothing about the directory distinguishes that from real damage, so this
stops trying to. Replay refuses instead, leaving the tail where it found
it, and reports that a rerun is the next step. The window is a single
write, so a rerun clears it; a tail that survives repeated runs is real
damage, and the message says so.

- `sei-db/wal/wal.go`, `utils.go`: add `Config.NoRepairOnOpen`, which
returns `ErrCorrupt` from `open` instead of truncating, and export
`ErrCorrupt` so a caller can name the outcome. Default is unset, so
every existing caller — `seid` included — keeps repairing as before.
- `sei-db/state_db/sc/memiavl/opts.go`, `db.go`: add
`Options.NoChangelogRepair` and pass it through. `ReadOnly` alone still
repairs, because a reader that cannot rerun is better served by
proceeding.
- `sei-db/tools/cmd/seidb/operations/evm_logical_digest.go`: set it for
replay mode, and translate `ErrCorrupt` into what the operator should
do. The translation also covers a torn record found during the replay
rather than at the open: `computeWALIndexDelta` and `Catchup` read the
changelog after the open returns, and they surface the same sentinel,
where a rerun is the same right answer.

### What this does not cover

`wal.Open` also finishes an interrupted truncation before it checks for
a torn record, by removing and renaming segments — a `.START` file left
by `TruncateFront`, or a `.END` file left by `TruncateBack`. It reports
no error for either, so there is nothing to gate on and no rerun to
prompt. Finishing one underneath the writer makes the writer's own
remove fail, and tidwall sets `l.corrupt` past that point, so every
later append fails until `seid` restarts. The operator message therefore
promises only the changelog tail, which does survive, rather than the
whole directory.

Both windows are narrower than the mid-record one, which recurs every
block. `TruncateFront` runs after a snapshot rewrite, roughly hourly at
the default interval. `TruncateBack` runs only inside a
`LoadForOverwriting` open, so reaching that window needs a crash during
a rollback and a replay before the node restarts. Both are left open
deliberately — closing them needs either a check before the open, which
the writer can invalidate in the gap, or the directory's `LOCK`, which
would require the node stopped and take away the case replay exists for.

Replay stays usable on a live node, which is the point: it is the
fallback for a height with no snapshot, and on a live migrating node the
two backends rarely retain a common snapshot height.

## Test plan

- `sei-db/tools/cmd/seidb/operations/memiavl_open_test.go`: tear the
last changelog segment, then require the open to report `ErrCorrupt`
with the rerun guidance *and* the segment's bytes to be unchanged, since
the refusal is worth nothing if the tail does not survive it. Replaying
an intact changelog to the requested height is the positive control.
Reverting `NoChangelogRepair` makes the first test fail on a successful
open, so it pins the behavior rather than restating it.
- `sei-db/wal/wal_test.go`: existing `TestOpenAndCorruptedTail` still
passes with repair requested, pinning that the default path is
untouched.
- `go test ./sei-db/wal/... ./sei-db/tools/cmd/seidb/operations/...`,
`scripts/ramtest.sh ./sei-db/state_db/sc/memiavl/...`
- `make dblint`

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit 7273a43)
Co-authored-by: Cursor <cursoragent@cursor.com>
@cursor

cursor Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes WAL/memiavl open behavior for callers that opt in; default node opens still auto-repair, but misconfigured tools could fail opens that previously succeeded after silent truncation.

Overview
Adds an opt-out for automatic changelog tail repair on WAL open, so read-only diagnostic tools do not truncate a live node’s changelog when the tail looks corrupt mid-write.

The WAL layer gains Config.NoRepairOnOpen: if the log ends mid-record, open returns wal.ErrCorrupt and leaves files unchanged; default opens still repair by truncating the tail. MemIAVL exposes this as Options.NoChangelogRepair, passed through when opening the changelog WAL.

evm-logical-digest replay mode sets NoChangelogRepair on read-only MemIAVL open and maps ErrCorrupt to a message that suggests rerun (in-flight write) vs real corruption. Docs note replay on a live node can hit this case.

New tests assert a torn changelog is refused without mutating the segment, and intact changelogs still replay.

Reviewed by Cursor Bugbot for commit b829937. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed✅ passed✅ passed✅ passedSep 21, 2026, 2:17 PM

@codecov

codecov Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 60.60%. Comparing base (e59189e) to head (b829937).

Additional details and impacted files

Impacted file tree graph

@@               Coverage Diff                @@
##           release/v6.7    #4282      +/-   ##
================================================
- Coverage         61.34%   60.60%   -0.74%     
================================================
  Files              2163     2083      -80     
  Lines            188785   180185    -8600     
================================================
- Hits             115813   109206    -6607     
+ Misses            62254    60970    -1284     
+ Partials          10718    10009     -709     
Flag Coverage Δ
sei-chain-pr 34.90% <100.00%> (?)
sei-db 69.80% <ø> (ø)
sei-db-state-db ?
sei-db-state-db-pr 77.23% <100.00%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
sei-db/state_db/sc/memiavl/db.go 69.90% <100.00%> (-0.36%) ⬇️
sei-db/state_db/sc/memiavl/opts.go 100.00% <ø> (ø)
...b/tools/cmd/seidb/operations/evm_logical_digest.go 24.30% <100.00%> (+1.39%) ⬆️
sei-db/wal/utils.go 56.71% <100.00%> (ø)
sei-db/wal/wal.go 71.47% <100.00%> (+1.50%) ⬆️

... and 82 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@blindchaser

Copy link
Copy Markdown
Contributor Author

Closing: this was created by hand. The right path is the seidroid backport bot, so #3983 gets the backport + release/v6.7 labels and a /backport comment re-triggers the action. Reopening via the bot so the branch and PR match the convention.

Keeping the note about the one fix the cherry-pick needs: CommitStore.Commit takes a version argument on main and none on release/v6.7, so memiavl_open_test.go needs store.Commit() instead of store.Commit(store.Version() + 1).

@blindchaser
blindchaser deleted the backport-3983-to-release/v6.7 branch September 21, 2026 14:18

@seidroid seidroid Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Faithful backport of #3983 — the diff matches upstream exactly apart from the documented store.Commit() signature adaptation, and the NoRepairOnOpen → wal.ErrCorrupt plumbing through memiavl to the digest tool is correct and covered by a positive/negative test pair. Only a doc-accuracy point on the newly exported ErrCorrupt and a test-coverage gap at the wal layer.

Findings: 0 blocking | 3 non-blocking | 1 posted inline

Blockers

  • None at the file/PR level.

Non-blocking

  • [suggestion] The new Config.NoRepairOnOpen field has no direct test in sei-db/wal; wal_test.go was only updated mechanically to pass false. TestOpenAndCorruptedTail is already a table test over corrupt-tail cases, so adding a noRepair: true row that asserts open returns wal.ErrCorrupt and leaves the segment bytes untouched would pin the new behaviour at the layer that implements it, rather than only end-to-end through the seidb tool.
  • 1 suggestion(s)/nit(s) flagged inline on specific lines.
  • 1 non-blocking pre-existing issue(s) listed below under pre-existing issues.

Pre-existing issues

  • [suggestion] truncateCorruptedTail (sei-db/wal/utils.go:58) ends with if pos != len(data), but data is resliced inside the loop (data = data[n:]), so the check compares bytes consumed against bytes remaining instead of against the original file length. When those two happen to be equal the truncation is silently skipped and the retried wal.Open fails with ErrCorrupt even though repair was requested. Comparing against the original length (captured before the loop) is the intended condition.

Comment thread sei-db/wal/wal.go
const defaultWriteBatchSize = 64

// ErrCorrupt reports that the log ends mid-record. An open returns it only under
// Config.NoRepairOnOpen; otherwise the tail is truncated and the open succeeds.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] The word "only" overstates the guarantee. open also returns wal.ErrCorrupt on the repair path: when no segment file is found (len(lastSeg) == 0 returns the original err), and when the retried wal.Open after truncateCorruptedTail still reports corruption. Since this is a newly exported sentinel that callers will branch on with errors.Is, the doc as written invites the inference "ErrCorrupt ⇒ NoRepairOnOpen was set ⇒ the log was not modified," which does not hold. Something like "NoRepairOnOpen makes an open return it instead of truncating the tail; it can also surface when the repair itself fails" keeps the contract honest.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-21T14:21:31.725544Z b829937 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant