Skip to content

fix(client): decouple runtime cluster add from readiness - #660

Merged
SunSi12138 merged 1 commit into
devfrom
feat/652-add-cluster-lifecycle
Sep 11, 2026
Merged

SunSi12138 merged 1 commit into
devfrom
feat/652-add-cluster-lifecycle

Conversation

@SunSi12138

@SunSi12138 SunSi12138 commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Closes #652

Summary

  • Make runtime AddClusterAsync commit on local lifecycle/publication facts instead of remote readiness.
  • Split Add and Replace candidate activation policy: Add starts the child runtime and publishes; Replace keeps connect/ready-before-swap behavior.
  • Allow Add while the parent lifecycle is Running even when the legacy aggregate connectivity state is Connecting.
  • Capture the compatibility ConnectAsync slot batch at operation start so a later Add is not grafted into, extended by, or stopped by that older batch.
  • Preserve Created-state Add behavior: prepare/publish locally without proactively connecting until the parent runtime starts.
  • Preserve route/connection-budget preflight before child Start and atomic local revalidation/rollback before publication.
  • Migrate integration/package consumers that require immediate RPC availability to the canonical StartAsync + scoped WaitForReadyAsync flow instead of relying on Add completion as a readiness signal.
  • Keep [api][multi-cluster] 为 Add/Replace/Remove 提供各自的 structured result contract #656 structured Add/Replace/Remove result contracts out of this PR.
  • Keep development versions unchanged; no package/project/assembly/file/diagnostics schema version bump.

Lifecycle / publication semantics

Running Add now follows:

local validation + route/budget preflight
-> prepare candidate
-> child StartAsync
-> atomic local revalidation
-> publish cluster/routes
-> transaction committed
-> child connectivity supervisor continues independently

A first dial/handshake failure may happen before or after parent publication and does not make a locally valid Add fail. The published child can remain Connecting / Reconnecting / NotReady until its own supervisor converges.

Caller cancellation or parent lifecycle closure before publication rolls back and stops the started candidate. Cancellation after publication does not revoke coordinator ownership.

Created-state Add remains local-only: it prepares/publishes the slot without proactively starting connectivity; later parent StartAsync starts child runtimes normally.

Compatibility ConnectAsync

A legacy ConnectAsync captures its slot snapshot when the operation starts. Concurrent Add publication:

  • does not extend or rewrite that captured batch;
  • starts the new child under its own runtime supervisor when the parent lifecycle is Running;
  • does not let an old compatibility failure stop a newly published child;
  • projects aggregate State/Readiness from the latest snapshot after the old operation completes.

Pure legacy Created + Connecting mutation rejection remains unchanged. The Running + aggregate Connecting exception is scoped to Add only; Replace/Remove are not broadened by this PR.

Add vs Replace

  • Add: Start -> publish -> background readiness convergence.
  • Replace: retains Start/Connect -> ready -> atomic swap -> retire old availability-first behavior.

The implementation uses separate activation helpers so future changes cannot accidentally couple these policies again.

Regression coverage

New/updated tests cover:

  • Running Add publishes a started candidate without waiting for Ready.
  • Immediate first dial failure still commits Add and leaves the child NotReady / parent Degraded.
  • Running + legacy aggregate Connecting permits Add.
  • An in-flight compatibility ConnectAsync batch does not grow when a later Add publishes.
  • Failure of the old compatibility batch does not stop a slot published after its snapshot.
  • Pre-publication caller cancellation after child Start rolls back and disposes the candidate.
  • Parent Stop winning after candidate Start but before publication rolls back the candidate and leaves the snapshot unchanged.
  • Post-publication caller cancellation does not undo ownership.
  • Concurrent same-key Adds remain serialized through the publication boundary.
  • Ready + added NotReady projects aggregate Degraded; an added NotReady cluster as the only published slot projects aggregate NotReady and later converges without another ConnectAsync call.
  • Replace retains connect-before-swap semantics and unavailable replacement candidates preserve the old authoritative cluster through existing coverage.
  • Integration/TLS/package smoke callers that issue RPC immediately after runtime Add now explicitly wait for scoped readiness.

Validation

Implementation base: exact dev d5a74be2bc20080ccc95ff688be90f19b7625cce.

Development/focused validation:

  • exact-base / intended-file scope guards: passed
  • git diff --check: passed
  • maintainability debt gate: passed without new baseline allowance
  • version/schema mutation guard: passed
  • Release Unit coverage after final test additions: 1693 tests on the formal PR path
  • serialized Integration validation after lifecycle/readiness callsite migration: 451/451 passed

Final exact-head candidate: 5fbfb0959ae96fbae4aea8b497c3125f067e7208.

Formal PR CI on that exact SHA:

  • PR Fast #1507: passed — project-reference boundary, formatting, maintainability, Release build, generated-assembly dependency check, Unit, Generator, LoadTest, allocation gate.
  • PR Quick #2844: passed — formatting/maintainability, Debug + Release builds, Unit, Generator, LoadTest, 451/451 Integration, NativeAOT transport/topology smoke, pack/package contract verification, NuGet package smoke, demo/load/chaos smoke.
  • CodeQL #3885: passed.

Final self-review

  • Branch is one commit ahead of exact dev and zero behind.
  • Final diff contains 13 intended production/test/docs files and no temporary workflow/script files.
  • Running Add publication never inspects candidate readiness/connectivity as a commit gate.
  • Add and Replace activation policies are explicitly separate.
  • Compatibility ConnectAsync uses a captured slot batch and does not own later Add publications.
  • Created-state Add still does not proactively connect.
  • Pre-publication rollback and post-publication ownership boundaries are covered, including parent Stop and caller cancellation races.
  • Aggregate readiness is computed from the latest published snapshot.
  • [api][multi-cluster] 为 Add/Replace/Remove 提供各自的 structured result contract #656 structured mutation result work is not included.
  • No development version or diagnostics schema version was changed.

@SunSi12138
SunSi12138 force-pushed the feat/652-add-cluster-lifecycle branch 3 times, most recently from 5d99819 to 0164f3f Compare September 11, 2026 12:57
@SunSi12138
SunSi12138 force-pushed the feat/652-add-cluster-lifecycle branch from 2aa5cf1 to 5fbfb09 Compare September 11, 2026 13:08
@SunSi12138
SunSi12138 marked this pull request as ready for review September 11, 2026 13:18
@SunSi12138
SunSi12138 merged commit f7b51d0 into dev Sep 11, 2026
27 checks passed
@SunSi12138
SunSi12138 deleted the feat/652-add-cluster-lifecycle branch September 11, 2026 13:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant