perf(client): sample the resolve-stage deadline once and drop the redundant checkpoint - #689
Merged
Merged
Conversation
…undant checkpoint A plain timed Unary read the monotonic clock ten times per invocation. Three of those samples carried no information the call had not already established: - ResolveCallControl created the deadline from one sample and then validated it with a second read of the same clock a few lines later, - the invoker checked deadline progress before connection selection, which cannot have moved past the stage-local validation it just repeated, - the pre-registration checkpoint after selection remains the first sample that can observe blocking work. The resolve stage now keeps one timestamp and validates the selected deadline against it, including the inherited-deadline branch, which already had the comparison sample in hand. The pre-selection checkpoint is gone; the post-selection checkpoint and the publication gate still take their own authoritative samples. Reads per timed Unary invocation drop from 10 to 8, and the samples that can fail a call are unchanged.
This was referenced Sep 13, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
A plain timed Unary read the monotonic clock ten times per invocation on
dev(measured, see Evidence). Two of those samples cannot change the outcome:ResolveCallControlcreates the deadline fromtimeProvider.GetTimestamp(), then validates it a few lines later withdeadline.IsExpired(timeProvider)- a second read of the same clock.InvokeUnaryCoreAsynccallsEnsureLogicalCallProgress(control)before connection selection. That step cannot have blocked yet, so the sample can only repeat the stage-local validation that just ran. The checkpoint after selection is the first sample that can observe blocking work, and it stays.The removed work is not only the two samples: each checkpoint also pulls the
ResolvedCallControl→RpcDeadline→TimeProviderchain through the hottest path in the client. Per counter evidence below that chain costs 116 instructions and 23 branches per call.Changes
ResolveCallControlkeeps one timestamp for the whole resolve stage and validates the selected deadline against it. The inherited-deadline branch reuses the comparison sample it already took, so it does not add a read either.EnsureLogicalCallProgress(control)inInvokeUnaryCoreAsyncis removed, with the reason recorded next to the remaining checkpoint.No behavior boundary moves. The samples that are allowed to fail a call are unchanged:
Compatibility
TimeBudgetfield and its sampling point are untouched.Validation
dotnet build Sharplink.slnx -c Release, 0 warnings / 0 errors).SharpLink.UnitTests1846/1846, including the newSharpLinkClientDeadlineClockReadTests. The new test fails ondevwithtook 10 samplesand is stable at 8 samples over five consecutive runs.CHANGELOG.md: not required, no user-visible or compatibility-relevant change.Additional gates run locally:
dotnet format whitespace --verify-no-changes,python3 eng/check-project-reference-boundaries.py,bash eng/check-maintainability.sh(1157 files checked, within baseline).Evidence
Clock samples per invocation (deterministic)
ManualTimeProvider.TimestampReadCountaround oneInvokeUnaryAsyncon a timed method, same test file compiled against both revisions:dev(c061169)Stack13ClockTestsadditionally pins the resolve stage itself at exactly one sample.The remaining eight split as: resolve (1), pre-registration checkpoint (1), publication gate (1), response terminal (1), send pump (2), session activity stamp (2). The last four are per-batch infrastructure shared with untimed traffic; an untimed call reaches only those.
Per-op work (
perf stat, 11 s window, same client workload)Both orders were measured (dev→PR and PR→dev); the direction is the same in both.
perf recordover the same workload puts the vDSO clock frame at 10.72 % → 9.82 % of client samples.For scale: one
TimeProvider.GetTimestamp()costs 12.2–12.6 ns in a tight loop on an Apple M4 host, and the vDSO share above corresponds to roughly 20 ns per call on the Linux host - consistent with the 65-cycle reduction.End-to-end throughput (bare metal, client pinned to cores 0-3, server to cores 4-7, UDS, concurrency 64,
add-64, 25 s + 5 s warmup, alternating variants)Two independent alternating series were run, because the host's own CPU time per op drifts by a few percent over minutes.
Series A (first session, 6 pairs):
dev+5.0 %, with the two ranges not overlapping (dev max 1,456,264 < PR min 1,463,435). Allocation is unchanged (383 → 382 B/op) and p99 is unchanged within noise.
Series B (later session, 6 pairs, host state drifted):
Series B starts with two pairs the change loses; those are also the two runs where
devis far above its own range (1.59 M and 1.53 M against 1.40–1.46 M in every earlier and later run), which is a warm-up artefact of a freshly started series rather than a property of the change. The remaining four pairs all favour this PR. Taking the two series together, the honest claim is +2 to +5 % throughput, −0.09 to −0.10 µs/op of client CPU, against a counter-verified reduction of 2 clock reads, 116 instructions and 23 branches per call. The percentages are host-state dependent and should not be read as a fixed number.echo-4096shows no change in either series (498,467 → 502,712 and 506,971 → 497,922 ops/s, both inside the run-to-run spread): a 4 KB payload is dominated by frame copy and socket write.Artifacts on the bare-metal host:
/tmp/ab-*.jsonand/tmp/ac-*.json(series A),/tmp/ax-*.json(series B),/tmp/pg2-{X,W}-add.data(perf record),/tmp/psa-{X,W}.txt(perf stat).