Skip to content

test(state): cross-store fault-injection E2E for the unified state control plane (issue #5314) - #5318

Merged
qqeasonchen merged 3 commits into
apache:developfrom
qqeasonchen:feat/5314-cross-store-fault
Aug 31, 2026
Merged

qqeasonchen merged 3 commits into
apache:developfrom
qqeasonchen:feat/5314-cross-store-fault

Conversation

@qqeasonchen

Copy link
Copy Markdown
Contributor

Summary

Closes #5314 (one-shot, supersedes the earlier plan to split into multiple PRs).

This PR delivers the cross-store fault-injection test suite for the unified state control
plane introduced by #5301. The four individual Sub-PRs (#5310/#5311/#5312/#5313) each ship a
contract test for their own store, but no test exists for invariants that span two or more
stores
— and those are exactly the invariants the runtime depends on. The 6 scenarios here
exercise the failure modes at the seams between stores, the runtime, and the cluster-shared Meta.

What ships

4 test-harness primitives (eventmesh-runtime/src/test/java/.../state/fault/)

  • MetaPartitionSwitch — wraps a MetaStore, lets a test open/close a simulated network
    partition. Mutating ops throw MetaPartitionException; reads continue against the
    pre-partition snapshot.
  • CrossStoreRaceProbe — ordered log of cross-store operations (DELIVERY_PUT/REMOVE,
    OFFSET_WRITE/READ, TASK_UPDATE) with a monotonic seq, used to assert
    happens-before relationships.
  • JvmCrashHarness — spawns a child JVM, sentinel-file-driven destroyForcibly() (SIGKILL
    on POSIX, TerminateProcess on Windows), then relaunches against the same on-disk stores.
    Gated on ENABLE_JVM_CRASH_HARNESS=true; the in-process simulation in scenario 1 covers
    the same property without spawning a child JVM and runs everywhere.
  • InMemorySubscriptionStore — ConcurrentHashMap-backed SubscriptionStore for the split-
    brain scenario, so a test can hold two views of the world.

6 fault-injection scenarios (CrossStoreFaultInjectionTest, 596 LoC)

# Scenario Failure mode Core invariant asserted
1 CrashMidAckReAck crash after offset-write, before MQ-ACK callback recovery on a fresh dispatcher retires the delivery without re-invoking the channel (#5291 idempotency)
2 MetaPartitionDuringDlq Meta unreachable while dead-letter recording MetaBackedDeadLetterStore throws MetaPartitionException — a loud failure, not a silent no-op — so the dispatcher keeps the delivery in flight and retries on heal (#5292)
3 A2aCancelMidStream cancel lands between PENDING and RUNNING the taskEpoch guard rejects the late transition; the task converges on a single terminal state (#5302)
4 SubscriptionReRegisterAfterSplit subscription update during a Meta partition after heal the latest write wins — nothing dropped, nothing duplicated (#5288, #5301 SubscriptionStore)
5 OffsetStoreRaceVsDeliveryStore cross-thread offset-advance vs retire race the probe log proves every DELIVERY_REMOVE is preceded by an OFFSET_WRITE at the same offset (#5289 at-least-once)
6 A2aDispatchRaceVsTaskStore two dispatchers race on one task record stale-epoch writes are rejected; concurrent createTask yields exactly one winner (#5291)

All 6 scenarios run fully in-process and deterministically. No Testcontainers / no Nacos /
no Docker required.

Documentation

Why one-shot, not split

The 6 scenarios share the same 4 harness primitives. Splitting them would either (a)
duplicate the harness across PRs, or (b) require a pre-PR that lands the harness first and
has no assertions of its own. Keeping the harness slim (414 LoC across 4 files) and bundling
the scenarios keeps the review scope manageable while avoiding the
pre-PR-without-assertions anti-pattern.

Why no Testcontainers

#5314 originally listed 6 Testcontainers-based scenarios, but the in-process approach is
strictly better for these failure modes:

How to verify locally

# In-process: always runs
./gradlew :eventmesh-runtime:test --tests \
  org.apache.eventmesh.runtime.state.CrossStoreFaultInjectionTest

# Cross-JVM (optional; requires process control / SIGKILL)
ENABLE_JVM_CRASH_HARNESS=true ./gradlew :eventmesh-runtime:test --tests \
  org.apache.eventmesh.runtime.state.fault.JvmCrashHarness

Relationship to existing suites (no overlap)

  • SubscriptionStoreTest / SessionStoreTest / DeadLetterStoreTest / TaskStoreTest /
    DeliveryStateStoreTest / MetaBackedDeadLetterStoreTest / MetaBackedTaskStoreTest
    cover single-store contracts.
  • DeliveryRecoveryTest covers single-store recovery (dispatcher.recover()).
  • ClusterDeliveryFaultTest covers single-store faults (partition ownership failure).
  • CrossStoreFaultInjectionTest (this PR) covers cross-store collaborative faults.

Notes for reviewers

  • The 6 scenarios are additive: no production code is touched. The only main-source change
    in this branch is zero — the diff is 6 files, +1063 lines, 0 deletions.
  • Scenario 3/6 model the taskEpoch guard's documented contract (set at createTask, never
    reset): it rejects cross-restart staleness, not intra-JVM double-writes. Scenario 6's
    concurrentDispatchersConvergeOnFreshestWrite asserts the createTask putIfAbsent
    guarantee rather than inventing a second epoch-bump semantics.
  • JvmCrashHarness is shipped but inert unless ENABLE_JVM_CRASH_HARNESS=true, so it cannot
    make CI flaky.

Checklist

Four in-process primitives shared by the six apache#5314 scenarios:

  MetaPartitionSwitch      - MetaStore wrapper; open()/close() simulates a network
                             partition (mutating ops throw MetaPartitionException,
                             reads continue against the pre-partition snapshot).
  CrossStoreRaceProbe      - ordered log of cross-store operations (DELIVERY_PUT/
                             REMOVE, OFFSET_WRITE/READ, TASK_UPDATE) with a monotonic
                             seq so a test can assert happens-before relationships.
  JvmCrashHarness          - child JVM + sentinel-file-driven destroyForcibly()
                             (SIGKILL / TerminateProcess), then relaunch against the
                             same on-disk stores. Gated on ENABLE_JVM_CRASH_HARNESS.
  InMemorySubscriptionStore- ConcurrentHashMap-backed SubscriptionStore for the
                             split-brain scenario (two views of the world).

All four are test-only and run fully in-process: no Nacos, no Docker, no
Testcontainers. See §13.2.12 of the architecture doc (added in a follow-up commit).
…pache#5314)

CrossStoreFaultInjectionTest covers the fault modes that span two or more
stores (or the runtime + cluster-shared Meta). The individual store contract
tests (apache#5310/apache#5311/apache#5312/apache#5313) cannot observe these, because the invariant
lives at the seam:

  1. CrashMidAckReAck  - crash after offset-write, before MQ-ACK callback;
                        recovery retires without re-invoking the channel
                        (issue apache#5291 idempotency).
  2. MetaPartitionDuringDlq - Meta unreachable while dead-letter recording;
                        the store throws MetaPartitionException rather than
                        silently no-op'ing, so the dispatcher keeps the
                        delivery in flight and retries on heal (apache#5292).
  3. A2aCancelMidStream - cancel lands between PENDING and RUNNING; the
                        taskEpoch guard rejects stale-epoch late transitions
                        and the task converges on one terminal state (apache#5302).
  4. SubscriptionReRegisterAfterSplit - update during a Meta partition; after
                        heal the latest write wins, nothing dropped or
                        duplicated (apache#5288, apache#5301 SubscriptionStore).
  5. OffsetStoreRaceVsDeliveryStore - cross-thread offset-advance vs retire
                        race; the probe log proves every DELIVERY_REMOVE is
                        preceded by an OFFSET_WRITE at the same offset
                        (apache#5289 at-least-once).
  6. A2aDispatchRaceVsTaskStore - two dispatchers race on one task record;
                        stale-epoch writes are rejected and createTask yields
                        exactly one winner (apache#5291).

Every scenario runs in-process and deterministically. The JvmCrashHarness
from the previous commit remains the optional cross-JVM verification path.

Note on the taskEpoch contract exercised by scenarios 3 and 6: the epoch is set
at createTask and never reset, so updateStatus rejects any epoch that differs
from the record's. Same-epoch writes are last-writer-wins by design - the
Runtime dispatcher is the sole writer and the epoch guards against a restarted
instance's stale handle, not against intra-JVM ordering.
…pache#5314)

Documents the 6-scenario CrossStoreFaultInjectionTest: why single-store contract
tests cannot cover cross-store invariants, the four harness primitives, the
per-scenario assertion table, why Testcontainers is not used, and the two
extension points (RocksDB-backed crash scenario, MetaBackedOffsetStore going
active).

Numbered 13.2.12 because §13.2.11 was taken by the dual-topology matrix in
PR apache#5317.
@qqeasonchen
qqeasonchen merged commit 7abca88 into apache:develop Aug 31, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Architecture Review][P1] #5301 Sub-PR D: cross-store fault-injection E2E

1 participant