Skip to content

[P0] propagate FencingToken epoch through dispatch / ACK / offset / DLQ #5360

Description

@qqeasonchen

Tracking sub-issue of #5354. Part of the production-HA acceptance plan (Phase 1, P0).

What this PR changes

Every offset write, broker ACK, DLQ transition, and dispatch call carries the owner's FencingToken. A new StaleOwnerException is raised when the in-memory token is below the value in Meta; the dispatcher catches it and stops processing the partition. PartitionOwnership re-tries ownership acquisition on the next tick.

Concretely

  • Add FencingToken currentToken() to ReliableDispatcher, set at construction from the PartitionOwnership scheduler.
  • Pass currentToken to every OffsetStore.commit, StoragePlugin.ack, DeadLetterSink.dispatch, and AckCallback.onAck call.
  • Before the write, the storage / offset / DLQ layer calls metaStore.get("em/assignments/" + topic + "#" + partition) and compares the persisted token to the local one. If the persisted token is higher, throw StaleOwnerException (new class in cluster package).
  • ReliableDispatcher.dispatch catches StaleOwnerException, removes the partition from the local owned set, increments a new fencedPartitions metric, and re-subscribes via the PartitionOwnership reaper.
  • Add boundedDelayedAckQueue cleanup on instance shutdown: pending POP ACKs are flushed to the OffsetStore before JVM exit; a unit test asserts the flush runs on graceful shutdown but NOT on kill -9 (the next instance's startup must rebuild the queue from the broker, then from the local OffsetStore).

Acceptance criteria

  • ReliableDispatcher has a FencingToken constructor argument and forwards it on every storage/offset/DLQ write.
  • A new StaleOwnerIntegrationTest runs two Runtime instances against an in-memory MetaStore, has instance A claim partitions, then bumps the token in Meta to simulate takeover. Instance A's next dispatch throws StaleOwnerException and is excluded from ownedPartitions on the next tick.
  • OffsetStore.commit is unreachable from ReliableDispatcher without a FencingToken (ArchUnit rule mirrors the runtime contract).
  • boundedDelayedAckQueue flush test: start instance, enqueue 10 POP ACKs, send SIGTERM, restart, assert all 10 ACKs are still pending (not lost) AND not double-applied (the broker is queried, not local state).

Verification

./gradlew :eventmesh-runtime:test --tests "*StaleOwnerIntegrationTest*"
./gradlew :eventmesh-runtime:test --tests "*ReliableDispatcherTest*"
./gradlew :eventmesh-runtime:test --tests "*BoundedDelayedAckQueueTest*"

Depends on / blocks

References

  • eventmesh-runtime/.../cluster/FencingToken.java
  • eventmesh-runtime/.../cluster/PartitionOwnership.java
  • eventmesh-runtime/.../cluster/MetaStore.java
  • eventmesh-runtime/.../delivery/ReliableDispatcher.java
  • eventmesh-runtime/.../offset/OffsetStore.java

Part of the production-HA topology in #5354. See also #5352 and #5353.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions