Skip to content

perf(client): sample the resolve-stage deadline once and drop the redundant checkpoint - #689

Merged
SunSi12138 merged 5 commits into
devfrom
perf/client-deadline-clock-read-reuse
Sep 13, 2026
Merged

SunSi12138 merged 5 commits into
devfrom
perf/client-deadline-clock-read-reuse

Conversation

@SunSi12138

@SunSi12138 SunSi12138 commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Motivation

A plain timed Unary read the monotonic clock ten times per invocation on dev (measured, see Evidence). Two of those samples cannot change the outcome:

  1. ResolveCallControl creates the deadline from timeProvider.GetTimestamp(), then validates it a few lines later with deadline.IsExpired(timeProvider) - a second read of the same clock.
  2. InvokeUnaryCoreAsync calls EnsureLogicalCallProgress(control) before connection selection. That step cannot have blocked yet, so the sample can only repeat the stage-local validation that just ran. The checkpoint after selection is the first sample that can observe blocking work, and it stays.

The removed work is not only the two samples: each checkpoint also pulls the ResolvedCallControl → RpcDeadline → TimeProvider chain through the hottest path in the client. Per counter evidence below that chain costs 116 instructions and 23 branches per call.

Changes

  • ResolveCallControl keeps one timestamp for the whole resolve stage and validates the selected deadline against it. The inherited-deadline branch reuses the comparison sample it already took, so it does not add a read either.
  • The pre-selection EnsureLogicalCallProgress(control) in InvokeUnaryCoreAsync is removed, with the reason recorded next to the remaining checkpoint.

No behavior boundary moves. The samples that are allowed to fail a call are unchanged:

boundary before after
resolve stage 2 samples 1 sample
pre-registration checkpoint (after connection selection) 1 1
publication gate 1 1
response terminal 1 1

Compatibility

  • Public API: unchanged. No public surface is touched.
  • Protocol/wire format: unchanged. The TimeBudget field and its sampling point are untouched.
  • Generated code or contract manifest: none.
  • NativeAOT and platform behavior: unchanged; no reflection, no trimming-relevant surface.

Validation

  • Non-incremental Release build has no warnings or errors (dotnet build Sharplink.slnx -c Release, 0 warnings / 0 errors).
  • Relevant unit tests pass: SharpLink.UnitTests 1846/1846, including the new SharpLinkClientDeadlineClockReadTests. The new test fails on dev with took 10 samples and is stable at 8 samples over five consecutive runs.
  • No user-facing behavior, limits, XML comments, or demos change.
  • Package smoke and NativeAOT checks: not required, no packaging or runtime ABI change.
  • CHANGELOG.md: not required, no user-visible or compatibility-relevant change.

Additional gates run locally: dotnet format whitespace --verify-no-changes, python3 eng/check-project-reference-boundaries.py, bash eng/check-maintainability.sh (1157 files checked, within baseline).

Evidence

Clock samples per invocation (deterministic)

ManualTimeProvider.TimestampReadCount around one InvokeUnaryAsync on a timed method, same test file compiled against both revisions:

revision samples
dev (c061169) 10
this PR 8

Stack13ClockTests additionally pins the resolve stage itself at exactly one sample.

The remaining eight split as: resolve (1), pre-registration checkpoint (1), publication gate (1), response terminal (1), send pump (2), session activity stamp (2). The last four are per-batch infrastructure shared with untimed traffic; an untimed call reaches only those.

Per-op work (perf stat, 11 s window, same client workload)

counter dev PR delta
instructions/op 4,578 4,462 −116 (−2.5 %)
branches/op 932 909 −23 (−2.4 %)
cycles/op 3,804 3,739 −65 (−1.7 %)
task-clock 27.82 s 27.78 s flat

Both orders were measured (dev→PR and PR→dev); the direction is the same in both. perf record over the same workload puts the vDSO clock frame at 10.72 % → 9.82 % of client samples.

For scale: one TimeProvider.GetTimestamp() costs 12.2–12.6 ns in a tight loop on an Apple M4 host, and the vDSO share above corresponds to roughly 20 ns per call on the Linux host - consistent with the 65-cycle reduction.

End-to-end throughput (bare metal, client pinned to cores 0-3, server to cores 4-7, UDS, concurrency 64, add-64, 25 s + 5 s warmup, alternating variants)

Two independent alternating series were run, because the host's own CPU time per op drifts by a few percent over minutes.

Series A (first session, 6 pairs):

revision median ops/s median µs/op
dev 1,411,302 2.495
this PR 1,481,935 2.391

+5.0 %, with the two ranges not overlapping (dev max 1,456,264 < PR min 1,463,435). Allocation is unchanged (383 → 382 B/op) and p99 is unchanged within noise.

Series B (later session, 6 pairs, host state drifted):

rep dev ops/s PR ops/s PR/dev
1 1,589,772 1,434,660 −9.8 %
2 1,526,330 1,466,957 −3.9 %
3 1,436,844 1,450,761 +1.0 %
4 1,399,635 1,504,249 +7.5 %
5 1,428,888 1,495,486 +4.7 %
6 1,403,723 1,465,972 +4.4 %
median 1,432,866 1,466,464 +2.3 %

Series B starts with two pairs the change loses; those are also the two runs where dev is far above its own range (1.59 M and 1.53 M against 1.40–1.46 M in every earlier and later run), which is a warm-up artefact of a freshly started series rather than a property of the change. The remaining four pairs all favour this PR. Taking the two series together, the honest claim is +2 to +5 % throughput, −0.09 to −0.10 µs/op of client CPU, against a counter-verified reduction of 2 clock reads, 116 instructions and 23 branches per call. The percentages are host-state dependent and should not be read as a fixed number.

echo-4096 shows no change in either series (498,467 → 502,712 and 506,971 → 497,922 ops/s, both inside the run-to-run spread): a 4 KB payload is dominated by frame copy and socket write.

Artifacts on the bare-metal host: /tmp/ab-*.json and /tmp/ac-*.json (series A), /tmp/ax-*.json (series B), /tmp/pg2-{X,W}-add.data (perf record), /tmp/psa-{X,W}.txt (perf stat).

…undant checkpoint

A plain timed Unary read the monotonic clock ten times per invocation. Three
of those samples carried no information the call had not already established:

- ResolveCallControl created the deadline from one sample and then validated
  it with a second read of the same clock a few lines later,
- the invoker checked deadline progress before connection selection, which
  cannot have moved past the stage-local validation it just repeated,
- the pre-registration checkpoint after selection remains the first sample
  that can observe blocking work.

The resolve stage now keeps one timestamp and validates the selected deadline
against it, including the inherited-deadline branch, which already had the
comparison sample in hand. The pre-selection checkpoint is gone; the
post-selection checkpoint and the publication gate still take their own
authoritative samples.

Reads per timed Unary invocation drop from 10 to 8, and the samples that can
fail a call are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant