fix(client): decouple runtime cluster add from readiness - #660
Merged
Merged
Conversation
SunSi12138
force-pushed
the
feat/652-add-cluster-lifecycle
branch
3 times, most recently
from
September 11, 2026 12:57
5d99819 to
0164f3f
Compare
SunSi12138
force-pushed
the
feat/652-add-cluster-lifecycle
branch
from
September 11, 2026 13:08
2aa5cf1 to
5fbfb09
Compare
SunSi12138
marked this pull request as ready for review
September 11, 2026 13:18
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #652
Summary
AddClusterAsynccommit on local lifecycle/publication facts instead of remote readiness.Runningeven when the legacy aggregate connectivity state isConnecting.ConnectAsyncslot batch at operation start so a later Add is not grafted into, extended by, or stopped by that older batch.StartAsync+ scopedWaitForReadyAsyncflow instead of relying on Add completion as a readiness signal.Lifecycle / publication semantics
Running Add now follows:
A first dial/handshake failure may happen before or after parent publication and does not make a locally valid Add fail. The published child can remain
Connecting/Reconnecting/NotReadyuntil its own supervisor converges.Caller cancellation or parent lifecycle closure before publication rolls back and stops the started candidate. Cancellation after publication does not revoke coordinator ownership.
Created-state Add remains local-only: it prepares/publishes the slot without proactively starting connectivity; later parent
StartAsyncstarts child runtimes normally.Compatibility ConnectAsync
A legacy
ConnectAsynccaptures its slot snapshot when the operation starts. Concurrent Add publication:State/Readinessfrom the latest snapshot after the old operation completes.Pure legacy
Created + Connectingmutation rejection remains unchanged. TheRunning + aggregate Connectingexception is scoped to Add only; Replace/Remove are not broadened by this PR.Add vs Replace
Start -> publish -> background readiness convergence.Start/Connect -> ready -> atomic swap -> retire oldavailability-first behavior.The implementation uses separate activation helpers so future changes cannot accidentally couple these policies again.
Regression coverage
New/updated tests cover:
Running + legacy aggregate Connectingpermits Add.ConnectAsyncbatch does not grow when a later Add publishes.ConnectAsynccall.Validation
Implementation base: exact
devd5a74be2bc20080ccc95ff688be90f19b7625cce.Development/focused validation:
git diff --check: passedFinal exact-head candidate:
5fbfb0959ae96fbae4aea8b497c3125f067e7208.Formal PR CI on that exact SHA:
Final self-review
devand zero behind.ConnectAsyncuses a captured slot batch and does not own later Add publications.