feat(state): control-plane store SPI + deprecate zombie code (#5301 Sub-PR A) - #5310
Merged
Merged
Conversation
qqeasonchen
force-pushed
the
feat/state-stores-spi
branch
2 times, most recently
from
August 26, 2026 07:49
1625149 to
0b00f83
Compare
Sub-PR A) This is the first of four sub-PRs that deliver issue apache#5301 (unified state control-plane stores). Sub-PR A adds the SPI surface and deprecates zombie code, without changing runtime behavior. New public interfaces in org.apache.eventmesh.runtime.state: - SubscriptionStore : cluster-shared subscription registry (Meta prefix-watch with local cache) - SessionStore : cluster-shared agent/binding/session registry (Meta prefix-watch with local cache) - DeadLetterStore : durable ledger of dead-lettered deliveries (Meta CAS, idempotent recordDeadLetter) - TaskStore : A2A task state with epoch-protected status transitions (Meta CAS, TTL via expireStale) Existing classes adapted to the new SPI: - ClusterSubscriptionStore now implements SubscriptionStore; remove() returns the boolean result from Meta.delete so callers can detect when a subscription was already gone. - SessionRegistry now implements SessionStore; register/markReady/ unregister are renamed to registerAgent/markAgentReady/ unregisterAgent to match the interface contract. Zombie code marked @deprecated(forRemoval = true) pending issue apache#5309: - cluster/PartitionOwnership (only callable from a closed delivery mode re-introduced by PARTITION_OWNED_PULL) - cluster/ClusterCoordinator (superseded by the new Meta CAS path) - cluster/MetaBackedOffsetStore (superseded by the OffsetStore in Sub-PR B) Unit tests cover the new interfaces and the adapted implementations with in-process test doubles (no external dependencies): - SubscriptionStoreTest (5 tests) - SessionStoreTest (6 tests) - DeadLetterStoreTest (2 tests) - TaskStoreTest (3 tests) Refs: apache#5301, apache#5309
qqeasonchen
force-pushed
the
feat/state-stores-spi
branch
from
August 26, 2026 08:02
0b00f83 to
6418c55
Compare
This was referenced Aug 27, 2026
qqeasonchen
added a commit
to qqeasonchen/eventmesh
that referenced
this pull request
Aug 31, 2026
…pache#5314) CrossStoreFaultInjectionTest covers the fault modes that span two or more stores (or the runtime + cluster-shared Meta). The individual store contract tests (apache#5310/apache#5311/apache#5312/apache#5313) cannot observe these, because the invariant lives at the seam: 1. CrashMidAckReAck - crash after offset-write, before MQ-ACK callback; recovery retires without re-invoking the channel (issue apache#5291 idempotency). 2. MetaPartitionDuringDlq - Meta unreachable while dead-letter recording; the store throws MetaPartitionException rather than silently no-op'ing, so the dispatcher keeps the delivery in flight and retries on heal (apache#5292). 3. A2aCancelMidStream - cancel lands between PENDING and RUNNING; the taskEpoch guard rejects the late transition and the task converges on one terminal state (apache#5302). 4. SubscriptionReRegisterAfterSplit - update during a Meta partition; after heal the latest write wins, nothing dropped or duplicated (apache#5288, apache#5301 SubscriptionStore). 5. OffsetStoreRaceVsDeliveryStore - cross-thread offset-advance vs retire race; the probe log proves every DELIVERY_REMOVE is preceded by an OFFSET_WRITE at the same offset (apache#5289 at-least-once). 6. A2aDispatchRaceVsTaskStore - two dispatchers race on one task record; stale-epoch writes are rejected and createTask yields exactly one winner (apache#5291). Every scenario runs in-process and deterministically. The JvmCrashHarness from the previous commit remains the optional cross-JVM verification path.
qqeasonchen
added a commit
to qqeasonchen/eventmesh
that referenced
this pull request
Aug 31, 2026
…pache#5314) CrossStoreFaultInjectionTest covers the fault modes that span two or more stores (or the runtime + cluster-shared Meta). The individual store contract tests (apache#5310/apache#5311/apache#5312/apache#5313) cannot observe these, because the invariant lives at the seam: 1. CrashMidAckReAck - crash after offset-write, before MQ-ACK callback; recovery retires without re-invoking the channel (issue apache#5291 idempotency). 2. MetaPartitionDuringDlq - Meta unreachable while dead-letter recording; the store throws MetaPartitionException rather than silently no-op'ing, so the dispatcher keeps the delivery in flight and retries on heal (apache#5292). 3. A2aCancelMidStream - cancel lands between PENDING and RUNNING; the taskEpoch guard rejects stale-epoch late transitions and the task converges on one terminal state (apache#5302). 4. SubscriptionReRegisterAfterSplit - update during a Meta partition; after heal the latest write wins, nothing dropped or duplicated (apache#5288, apache#5301 SubscriptionStore). 5. OffsetStoreRaceVsDeliveryStore - cross-thread offset-advance vs retire race; the probe log proves every DELIVERY_REMOVE is preceded by an OFFSET_WRITE at the same offset (apache#5289 at-least-once). 6. A2aDispatchRaceVsTaskStore - two dispatchers race on one task record; stale-epoch writes are rejected and createTask yields exactly one winner (apache#5291). Every scenario runs in-process and deterministically. The JvmCrashHarness from the previous commit remains the optional cross-JVM verification path. Note on the taskEpoch contract exercised by scenarios 3 and 6: the epoch is set at createTask and never reset, so updateStatus rejects any epoch that differs from the record's. Same-epoch writes are last-writer-wins by design - the Runtime dispatcher is the sole writer and the epoch guards against a restarted instance's stale handle, not against intra-JVM ordering.
qqeasonchen
added a commit
that referenced
this pull request
Aug 31, 2026
…ntrol plane (issue #5314) (#5318) * test(state): add cross-store fault-injection harness (issue #5314) Four in-process primitives shared by the six #5314 scenarios: MetaPartitionSwitch - MetaStore wrapper; open()/close() simulates a network partition (mutating ops throw MetaPartitionException, reads continue against the pre-partition snapshot). CrossStoreRaceProbe - ordered log of cross-store operations (DELIVERY_PUT/ REMOVE, OFFSET_WRITE/READ, TASK_UPDATE) with a monotonic seq so a test can assert happens-before relationships. JvmCrashHarness - child JVM + sentinel-file-driven destroyForcibly() (SIGKILL / TerminateProcess), then relaunch against the same on-disk stores. Gated on ENABLE_JVM_CRASH_HARNESS. InMemorySubscriptionStore- ConcurrentHashMap-backed SubscriptionStore for the split-brain scenario (two views of the world). All four are test-only and run fully in-process: no Nacos, no Docker, no Testcontainers. See §13.2.12 of the architecture doc (added in a follow-up commit). * test(state): add cross-store fault-injection 6-scenario test (issue #5314) CrossStoreFaultInjectionTest covers the fault modes that span two or more stores (or the runtime + cluster-shared Meta). The individual store contract tests (#5310/#5311/#5312/#5313) cannot observe these, because the invariant lives at the seam: 1. CrashMidAckReAck - crash after offset-write, before MQ-ACK callback; recovery retires without re-invoking the channel (issue #5291 idempotency). 2. MetaPartitionDuringDlq - Meta unreachable while dead-letter recording; the store throws MetaPartitionException rather than silently no-op'ing, so the dispatcher keeps the delivery in flight and retries on heal (#5292). 3. A2aCancelMidStream - cancel lands between PENDING and RUNNING; the taskEpoch guard rejects stale-epoch late transitions and the task converges on one terminal state (#5302). 4. SubscriptionReRegisterAfterSplit - update during a Meta partition; after heal the latest write wins, nothing dropped or duplicated (#5288, #5301 SubscriptionStore). 5. OffsetStoreRaceVsDeliveryStore - cross-thread offset-advance vs retire race; the probe log proves every DELIVERY_REMOVE is preceded by an OFFSET_WRITE at the same offset (#5289 at-least-once). 6. A2aDispatchRaceVsTaskStore - two dispatchers race on one task record; stale-epoch writes are rejected and createTask yields exactly one winner (#5291). Every scenario runs in-process and deterministically. The JvmCrashHarness from the previous commit remains the optional cross-JVM verification path. Note on the taskEpoch contract exercised by scenarios 3 and 6: the epoch is set at createTask and never reset, so updateStatus rejects any epoch that differs from the record's. Same-epoch writes are last-writer-wins by design - the Runtime dispatcher is the sole writer and the epoch guards against a restarted instance's stale handle, not against intra-JVM ordering. * docs(architecture): add §13.2.12 cross-store fault-injection (issue #5314) Documents the 6-scenario CrossStoreFaultInjectionTest: why single-store contract tests cannot cover cross-store invariants, the four harness primitives, the per-scenario assertion table, why Testcontainers is not used, and the two extension points (RocksDB-backed crash scenario, MetaBackedOffsetStore going active). Numbered 13.2.12 because §13.2.11 was taken by the dual-topology matrix in PR #5317.
4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sub-PR A: Control-plane store SPI + deprecate zombie code (issue #5301)
This is the first of four sub-PRs that deliver issue #5301
("Unified state control-plane stores"). Sub-PR A adds the SPI surface, adapts
two existing implementations to it, and deprecates three zombie-code classes
that the new architecture will eventually replace.
It does not change runtime behavior: nothing currently calls the new
interfaces yet (call sites move over in Sub-PRs B and C).
What's in this PR
New public interfaces (in
org.apache.eventmesh.runtime.state)SubscriptionStoreSessionStoreDeadLetterStoreTaskStoreexpireStale)OffsetStoreandDeliveryStateStoreare intentionally not added in thisPR — they are introduced in Sub-PRs B and C with the actual storage backends.
Adapted implementations
ClusterSubscriptionStorenowimplements SubscriptionStore.remove(...)returns the boolean result from
Meta.delete(...)so callers can detectwhen a subscription was already gone (idempotent teardown).
SessionRegistrynowimplements SessionStore.register/markReady/ unregisterare renamed toregisterAgent/markAgentReady/unregisterAgentto match the interface contract (no callers existed in the current tree
for the old names).
Zombie code marked
@Deprecated(forRemoval = true)Pending issue #5309
("PARTITION_OWNED_PULL delivery topology"):
cluster.PartitionOwnership— only callable from the second delivery modethe closed [Architecture Review][P0] Configurable DeliveryTopology for cluster delivery #5300 originally enabled; will return as part of [Architecture Review][P0] PARTITION_OWNED_PULL delivery topology (reopens #5300 second mode) #5309.
cluster.ClusterCoordinator— superseded by the new Meta CAS path.cluster.MetaBackedOffsetStore— superseded by theOffsetStoreinSub-PR B.
Unit tests
SubscriptionStoreTest(5),SessionStoreTest(6),DeadLetterStoreTest(2),
TaskStoreTest(3) — total 16 tests, all in-process with hand-rolled test doubles (no Nacos/RocksDB/Testcontainers needed). They pin the
SPI contract so subsequent sub-PRs can build on it.
Out of scope (handled in follow-up sub-PRs)
DeliveryStateStore(RocksDB) +OffsetStore(Meta asyncflush); wire
DeliveryStateStoreinto the new consumer path.DeadLetterStoreMeta-backed implementation +TaskStoreMeta-backed implementation; wire A2A
TaskRegistryoverTaskStore.transfer, Meta outage, DLQ replay, and TaskStore epoch conflict.
Open questions for the maintainer
(Inherited from the #5301 issue body — please review there.)
DeadLetterStore.recordDeadLettercarry adlqOffset(currentdesign), or just the
dlqTopicso the offset is derived?TaskStore.TaskRecordcarry fullinput/outputpayloads, oronly a
payloadRefthat points to a separate blob store?SubscriptionStore.put, do we want a separateputIfAbsentvariantto make subscription reconciliation idempotent without overwriting?
SessionStoreinterface include aWatch/Subscriptioncallback hook (similar to Nacos
Listener) so callers can react tosession-state changes without polling?
Testing
This PR was developed on the uni-runtime branch (
developat7260581). The fulleventmesh-runtime:compileJavaplus the four newtest classes pass locally with JDK 21 + Gradle 8.7. Upstream CI will be
the source of truth — if anything in the SPI needs to change, the four
sub-PRs are independent so a fix here will not block B/C/D.