Skip to content

test(client): synchronize dynamic reconnect replacement timer - #699

Merged
SunSi12138 merged 2 commits into
devfrom
fix/ci-flake-387-dynamic-reconnect-20260917
Sep 18, 2026
Merged

SunSi12138 merged 2 commits into
devfrom
fix/ci-flake-387-dynamic-reconnect-20260917

Conversation

@SunSi12138

Copy link
Copy Markdown
Owner

Refs #387.

New record analyzed

Nightly Regression #285 reported SharpLinkReconnectPolicyLifecycleRaceTests.DynamicGenerationReplacementShouldRetireTheOldReconnectOwnerAcrossPolicyWake on macOS. After generation 2 was materialized / snapshot-processed, the test advanced manual time by 999ms + 1ms but observed old=1, new=0.

Static root cause

This is a test synchronization race around reconnect-policy wake, not a demonstrated production ownership failure.

DynamicClusterReconnectCoordinator.EnsureReconnect starts ReconnectAsync synchronously up to its first reconnect-delay await. If generation replacement is processed just before UpdateReconnectPolicy, the replacement owner can first arm the old policy's 10-second timer. The policy update then cancels that generation's delay; re-entering the reconnect loop and arming the new 1-second timer happens on the async cancellation continuation.

PublishedSnapshotProcessed proves topology application plus the initial reconnect reconciliation returned, but it does not prove that this policy-wake continuation has run. Likewise, CreatedTimerCount > beforeReplacement can be satisfied by the now-cancelled 10-second timer. The test can therefore advance fake time to t=1 before the live one-second timer exists; if the continuation then arms the timer at t=1, it is due at t=2 and the test never advances again.

Fix

Use the test-local ManualTimeProvider.WaitForPendingTimerAsync(TimeSpan.FromSeconds(1)) seam, already present in this file for the analogous reconnect timer-registration race, before advancing fake time. This requires a live timer at exactly the current manual timestamp + one second and cannot be satisfied by the cancelled old-policy 10-second wait.

The existing assertions remain intact:

  • no replacement dial before 1 second;
  • replacement dial at the 1-second boundary;
  • retired old endpoint never dials again;
  • current generation owns the subsequent reconnect attempt.

No production code, reconnect semantics, timeout widening, retries, sleeps, or workflow changes.

Validation

Static analysis / diff review locally; executable validation is delegated to PR CI.

@SunSi12138
SunSi12138 merged commit e1edc26 into dev Sep 18, 2026
4 checks passed
@SunSi12138
SunSi12138 deleted the fix/ci-flake-387-dynamic-reconnect-20260917 branch September 18, 2026 10:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant