test(state): cross-store fault-injection E2E for the unified state control plane (issue #5314) - #5318
Merged
qqeasonchen merged 3 commits intoAug 31, 2026
Conversation
Four in-process primitives shared by the six apache#5314 scenarios: MetaPartitionSwitch - MetaStore wrapper; open()/close() simulates a network partition (mutating ops throw MetaPartitionException, reads continue against the pre-partition snapshot). CrossStoreRaceProbe - ordered log of cross-store operations (DELIVERY_PUT/ REMOVE, OFFSET_WRITE/READ, TASK_UPDATE) with a monotonic seq so a test can assert happens-before relationships. JvmCrashHarness - child JVM + sentinel-file-driven destroyForcibly() (SIGKILL / TerminateProcess), then relaunch against the same on-disk stores. Gated on ENABLE_JVM_CRASH_HARNESS. InMemorySubscriptionStore- ConcurrentHashMap-backed SubscriptionStore for the split-brain scenario (two views of the world). All four are test-only and run fully in-process: no Nacos, no Docker, no Testcontainers. See §13.2.12 of the architecture doc (added in a follow-up commit).
…pache#5314) CrossStoreFaultInjectionTest covers the fault modes that span two or more stores (or the runtime + cluster-shared Meta). The individual store contract tests (apache#5310/apache#5311/apache#5312/apache#5313) cannot observe these, because the invariant lives at the seam: 1. CrashMidAckReAck - crash after offset-write, before MQ-ACK callback; recovery retires without re-invoking the channel (issue apache#5291 idempotency). 2. MetaPartitionDuringDlq - Meta unreachable while dead-letter recording; the store throws MetaPartitionException rather than silently no-op'ing, so the dispatcher keeps the delivery in flight and retries on heal (apache#5292). 3. A2aCancelMidStream - cancel lands between PENDING and RUNNING; the taskEpoch guard rejects stale-epoch late transitions and the task converges on one terminal state (apache#5302). 4. SubscriptionReRegisterAfterSplit - update during a Meta partition; after heal the latest write wins, nothing dropped or duplicated (apache#5288, apache#5301 SubscriptionStore). 5. OffsetStoreRaceVsDeliveryStore - cross-thread offset-advance vs retire race; the probe log proves every DELIVERY_REMOVE is preceded by an OFFSET_WRITE at the same offset (apache#5289 at-least-once). 6. A2aDispatchRaceVsTaskStore - two dispatchers race on one task record; stale-epoch writes are rejected and createTask yields exactly one winner (apache#5291). Every scenario runs in-process and deterministically. The JvmCrashHarness from the previous commit remains the optional cross-JVM verification path. Note on the taskEpoch contract exercised by scenarios 3 and 6: the epoch is set at createTask and never reset, so updateStatus rejects any epoch that differs from the record's. Same-epoch writes are last-writer-wins by design - the Runtime dispatcher is the sole writer and the epoch guards against a restarted instance's stale handle, not against intra-JVM ordering.
…pache#5314) Documents the 6-scenario CrossStoreFaultInjectionTest: why single-store contract tests cannot cover cross-store invariants, the four harness primitives, the per-scenario assertion table, why Testcontainers is not used, and the two extension points (RocksDB-backed crash scenario, MetaBackedOffsetStore going active). Numbered 13.2.12 because §13.2.11 was taken by the dual-topology matrix in PR apache#5317.
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #5314 (one-shot, supersedes the earlier plan to split into multiple PRs).
This PR delivers the cross-store fault-injection test suite for the unified state control
plane introduced by #5301. The four individual Sub-PRs (#5310/#5311/#5312/#5313) each ship a
contract test for their own store, but no test exists for invariants that span two or more
stores — and those are exactly the invariants the runtime depends on. The 6 scenarios here
exercise the failure modes at the seams between stores, the runtime, and the cluster-shared Meta.
What ships
4 test-harness primitives (
eventmesh-runtime/src/test/java/.../state/fault/)MetaPartitionSwitch— wraps aMetaStore, lets a test open/close a simulated networkpartition. Mutating ops throw
MetaPartitionException; reads continue against thepre-partition snapshot.
CrossStoreRaceProbe— ordered log of cross-store operations (DELIVERY_PUT/REMOVE,OFFSET_WRITE/READ,TASK_UPDATE) with a monotonicseq, used to asserthappens-before relationships.
JvmCrashHarness— spawns a child JVM, sentinel-file-drivendestroyForcibly()(SIGKILLon POSIX, TerminateProcess on Windows), then relaunches against the same on-disk stores.
Gated on
ENABLE_JVM_CRASH_HARNESS=true; the in-process simulation in scenario 1 coversthe same property without spawning a child JVM and runs everywhere.
InMemorySubscriptionStore—ConcurrentHashMap-backedSubscriptionStorefor the split-brain scenario, so a test can hold two views of the world.
6 fault-injection scenarios (
CrossStoreFaultInjectionTest, 596 LoC)CrashMidAckReAckMetaPartitionDuringDlqMetaBackedDeadLetterStorethrowsMetaPartitionException— a loud failure, not a silent no-op — so the dispatcher keeps the delivery in flight and retries on heal (#5292)A2aCancelMidStreamtaskEpochguard rejects the late transition; the task converges on a single terminal state (#5302)SubscriptionReRegisterAfterSplitSubscriptionStore)OffsetStoreRaceVsDeliveryStoreDELIVERY_REMOVEis preceded by anOFFSET_WRITEat the same offset (#5289 at-least-once)A2aDispatchRaceVsTaskStorecreateTaskyields exactly one winner (#5291)All 6 scenarios run fully in-process and deterministically. No Testcontainers / no Nacos /
no Docker required.
Documentation
docs/eventmesh-uni-architecture-redesign.md: new §13.2.12 "Cross-store fault-injectionverification ([Architecture Review][P1] #5301 Sub-PR D: cross-store fault-injection E2E #5314 Sub-PR D2)" — scope, the 4 harness primitives, the per-scenario
assertion table, why Testcontainers is not used, and the two extension points.
Numbered 13.2.12 because §13.2.11 was taken by the dual-topology matrix in PR feat(cluster): PARTITION_OWNED_PULL delivery topology (issue #5309, closes) #5317.
Why one-shot, not split
The 6 scenarios share the same 4 harness primitives. Splitting them would either (a)
duplicate the harness across PRs, or (b) require a pre-PR that lands the harness first and
has no assertions of its own. Keeping the harness slim (414 LoC across 4 files) and bundling
the scenarios keeps the review scope manageable while avoiding the
pre-PR-without-assertions anti-pattern.
Why no Testcontainers
#5314originally listed 6 Testcontainers-based scenarios, but the in-process approach isstrictly better for these failure modes:
a Nacos/Kafka container adds a network-latency surface that obscures the assertion rather
than strengthening it.
the Testcontainers variant would be skipped in CI; the in-process variant runs everywhere.
([Bug] Make RocksDB and Meta offset writes atomically monotonic #5289, [Bug] Prevent stale ACKs from matching reused delivery IDs #5291, [Bug] Make DLQ transition durable before retiring source delivery #5292); running them against real Nacos with network delays does not change
the assertion outcome.
How to verify locally
Relationship to existing suites (no overlap)
SubscriptionStoreTest/SessionStoreTest/DeadLetterStoreTest/TaskStoreTest/DeliveryStateStoreTest/MetaBackedDeadLetterStoreTest/MetaBackedTaskStoreTestcover single-store contracts.
DeliveryRecoveryTestcovers single-store recovery (dispatcher.recover()).ClusterDeliveryFaultTestcovers single-store faults (partition ownership failure).CrossStoreFaultInjectionTest(this PR) covers cross-store collaborative faults.Notes for reviewers
in this branch is zero — the diff is 6 files, +1063 lines, 0 deletions.
taskEpochguard's documented contract (set atcreateTask, neverreset): it rejects cross-restart staleness, not intra-JVM double-writes. Scenario 6's
concurrentDispatchersConvergeOnFreshestWriteasserts thecreateTaskputIfAbsentguarantee rather than inventing a second epoch-bump semantics.
JvmCrashHarnessis shipped but inert unlessENABLE_JVM_CRASH_HARNESS=true, so it cannotmake CI flaky.
Checklist
apache/develop(includes PR feat(cluster): PARTITION_OWNED_PULL delivery topology (issue #5309, closes) #5317)