Skip to content

refactor(runtime): sample the request time budget once at frame creation - #684

Merged
SunSi12138 merged 10 commits into
devfrom
refactor/client-timeout-single-sample
Sep 13, 2026
Merged

SunSi12138 merged 10 commits into
devfrom
refactor/client-timeout-single-sample

Conversation

@SunSi12138

@SunSi12138 SunSi12138 commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Problem

Every Request frame carried a process-local RpcDeadline from the caller all the way into the send pump. The pump used it to

  • stamp the wire TimeBudget at emission time, in a bounded pass that re-sampled GetRemaining per frame inside an already-acquired output span,
  • defer writes for any batch holding a timed frame, which bypassed the batching window and produced a second write path with its own writtenCount bookkeeping,
  • compact frames whose deadline elapsed while they waited in the local queue and complete them with a synthetic DeadlineExceeded,
  • retain the deadline on the pooled writer lease (PooledByteBufferWriter.EmissionDeadline) or, for custom writers, in a boxed completion state inside OwnedFrame.

None of that changes the caller-visible outcome: the caller's failure is driven by the client's pending-deadline timer and by cancellation, not by the identity of the frame that carried the deadline.

Change

The remaining budget is sampled once, when the request frame is built, and written into the TimeBudget field of the frame. The pump never samples, rewrites, or compacts a deadline.

Removed: deferred timed writes and the deferWrites bookkeeping, WriteRetainedBatchAtEmission, the per-frame GetRemaining re-stamp, expired-frame compaction with CompleteExpiredBatch, OwnedFrame.Deadline and its DeadlineState/CompletionDeadlineState/FailureDeadlineState wrappers, and PooledByteBufferWriter.EmissionDeadline with all its lease-reset sites.

Kept, because it is what actually enforces the deadline end to end:

  • the client pending-deadline timer, so a call whose Request never reaches the transport (a blocked writer, for example) still fails at its deadline, and
  • transport cancellation, so a remote call that already started is terminated.

Wire consequence, and the reason this needs a changelog entry: the server-visible TimeBudget is sampled earlier than before, so it is a strictly more generous value than the caller's own remaining lifetime. The sampling point is the frame header, which is written before the request payload and metadata are serialized, so three things are no longer deducted from the value the server sees:

not deducted why
the rest of the request serialization the budget is written before the codec and metadata payloads
the client send queue the frame can wait for the batch window or for a stalled transport
the network and the server receive path unchanged in kind, now additive to the two above

The server budget is therefore advisory: it bounds how long the server is willing to keep working, and it is not the caller's deadline. The caller's own end-to-end deadline stays authoritative, which is the gRPC model - a deadline is fixed at invocation time and is not refreshed as the call moves down the stack. Two client-side guarantees hold that up, and both have deterministic tests in this PR:

  • A call with a pending entry (unary, streaming, and OneWay with client streams) fails through its pending deadline even when its Request never leaves the send queue, because the deadline is armed when the call is registered rather than when a frame is emitted. Publication is arbitrated against that same terminal transition, and admission to the send queue is the linearization point: frame preparation - including a user-provided compression provider - runs before the gate, the frame is marked published only for a frame the queue accepted, and the transport cancel is only emitted for a Request that did reach the queue. The gate is also the deadline authority: it re-checks the call's deadline while it holds the same completion gate that arbitrates publication, so an elapsed deadline refuses the Request even when the deadline scheduler callback has not run yet, and a frame can never reach the peer carrying a budget that already elapsed on the client. Without that arbitration the deadline produced Cancel and then Request - a pair the peer answers by discarding the cancel for an unknown request and dispatching the Request anyway - and a frame the queue refused was still reported as published, so its terminal transition emitted a cancel for a Request the peer never received.
  • A plain OneWay registers no pending entry, so nothing else enforces its lifetime. A Request whose budget already elapsed while its payload was still being serialized is not published at all, and one that was published races its emission wait against its own sampled deadline: losing that race fails the caller with DeadlineExceeded instead of blocking past its lifetime, and publishes the deadline cancel behind the Request it follows in the same normal queue, so the peer stops work whose caller has already given up.
  • The client-producer shapes - OneWay with client streams, client streaming and duplex streaming - now also wait for the emission of their head Request on the pending call's own producer token, so a deadline, a caller cancellation or a closed connection ends the invocation from a stalled transport write instead of leaving it parked in the send path. A pending call only creates that token for those three shapes; unary returns the pending operation without awaiting an emission at all, and server streaming has no producer token, so its start task can still sit in the emission waiter until the transport recovers. That last one is pre-existing behaviour and not a caller-visible deadline gap - the pending terminal still completes its dispatcher - so it is left as a follow-up. A plain OneWay re-checks its budget after preparation, so a Request that expired while the compression provider ran is refused at the linearization point instead of being published and then cancelled.

A timed Request also now waits for the configured batch window like any other frame when the caller sets an explicit RpcSessionFlushOptions.MaxLatency; under the default profiles the batch window is zero, so timed frames are still emitted as soon as the queue drains.

The batch is still copied into one output span. The deadline machinery was also doing the transport a favour: the frames it deferred were serialized in a single GetSpan at flush time, while untimed frames were copied one at a time as they arrived. Removing the deferral without replacing that would have cost real throughput - the first measurement below shows echo-4096 losing 15% against 2.0.0 with arrival-time copies. The pump now coalesces the whole batch into one span at flush time (WritePendingBatch) and inspects no deadline while doing it.

Measurement (baremetal UDS, client cores 0-3 / server cores 4-7, c64, 20 s + 5 s warmup, 3 reps)

Variants: A = 1.1.1, B = 2.0.0 main, T = main + budget change only (no batch coalescing, no #681/#682), D = dev, U = dev + budget change, V = dev + budget change + batch coalescing (this branch).

Attribution on the 2.0.0 base

case A (1.1.1) B (2.0.0) T (budget only) T vs B
add 1,844,193 1,454,059 1,493,676 +2.7%
oneway 5,172,417 4,247,111 4,238,958 -0.2%
echo-256 1,462,573 1,250,264 1,217,409 -2.6%
echo-4096 316,773 363,371 308,768 -15.0%
upload-4096 45,942 45,317 46,081 +1.7%
download-4096 44,828 45,328 45,275 -0.1%
duplex-4096 38,837 40,022 39,946 -0.2%

Confirmation on the dev base

U against D isolates the same change on top of #681/#682 and reproduces the regression, so neither the allocation work nor #682 is the cause:

case D (dev) U (dev + budget) U vs D
add 1,461,073 1,474,875 +0.9%
oneway 4,228,935 4,229,974 0.0%
echo-256 1,212,917 1,202,703 -0.8%
echo-4096 377,955 309,083 -18.2%
upload-4096 45,578 45,481 -0.2%
download-4096 45,204 45,537 +0.7%
duplex-4096 39,723 39,779 +0.1%

Final result

case A (1.1.1) B (2.0.0) V (this branch) V vs B V vs A
add 1,844,193 1,454,059 1,478,748 +1.7% -19.8%
oneway 5,172,417 4,247,111 4,196,992 -1.2% -18.9%
echo-256 1,462,573 1,250,264 1,305,866 +4.4% -10.7%
echo-4096 316,773 363,371 504,663 +38.9% +59.3%
upload-4096 45,942 45,317 87,554 +93.2% +90.6%
download-4096 44,828 45,328 94,653 +108.8% +111.2%
duplex-4096 38,837 40,022 67,058 +67.6% +72.7%

Per-operation allocation is unchanged from D/U (add 382.7 B/op against 382.1 for 1.1.1, echo-256 677.7 against 679.6), so none of this is bought with pooling or with a smaller working set. Latency follows throughput: echo-4096 p50 155 -> 101 us, upload-4096 p50 1463 -> 674 us, download-4096 p50 1446 -> 606 us, duplex-4096 p99 3474 -> 2716 us (B -> V).

The size of the win is explained by what the old path did at 4 KiB: Output is a StreamPipeWriter over the socket stream with the default 4 KiB buffer, so a frame plus its header never fits in what is left of a buffer. Writing each frame as it arrived therefore committed the buffered bytes to the socket once per frame - the pump's 16 KiB flush threshold only decided how often it awaited FlushAsync, not how often the socket was actually written. Counting syscalls with the same client pinned to the same cores (10 s, c64, upload-4096, so strace overhead dominates the absolute numbers):

variant ops completed sendto calls socket writes per op
B 46,694 476,430 10.2
V 187,518 441,374 2.4

Under strace the four-times-smaller syscall count per operation shows up as four times the throughput (4.0x, against 1.9x without strace). The total number of bytes copied into the pipe is the same either way; only the destination layout changed, so this is fewer syscalls rather than a trade against the working set.

Timeout-enabled benchmark

The probe contract puts no [Timeout] on any of its methods, and only the unary entry points inherit the client's default request timeout (InvokeUnaryAsync resolves its control with includeClientDefault: true; OneWay, client-streaming, server-streaming, and duplex resolve with includeClientDefault: false). So the deadline machinery is live for add and echo-4096, and the OneWay / streaming rows are deadline-free workloads that no configured client default reaches - the -1.3% on oneway is not a measurement of this change. That split is asserted on the wire rather than inferred: ClientDefaultTimeoutShouldReachOnlyTheShapesThatOptIn runs a client with a 30 s default over all five shapes and checks that only the unary Request carries HasTimeBudget, with the 30 s value, while the one-way, client-streaming, server-streaming and duplex Requests carry none; ClientDefaultTimeoutShouldNotEnforceAPlainOneWayDeadline stalls the transport, lets that default elapse, and shows the untimed OneWay neither fails nor publishes a cancel. To show the live rows are not an artifact of a 30 s budget, the probe was rebuilt with an env-controlled timeout and rerun with a 1 s client default budget, same cores, same server, 20 s + 5 s warmup:

case deadline live A (1.1.1) B (2.0.0) V (this branch) V vs B
add yes 1,832,853 1,487,064 1,503,115 +1.1%
echo-4096 yes 315,707 348,506 495,185 +42.1%
oneway no 5,157,796 4,353,866 4,298,222 -1.3%
upload-4096 no 47,405 46,205 87,645 +89.7%
download-4096 no - - 93,084 -

The same probe also re-measures this branch after the publication and plain-OneWay deadline fixes at the original 30 s budget: add 1,541,890, oneway 4,397,626, echo-4096 514,681, upload-4096 87,604, download-4096 95,220. Those are within the run-to-run spread of the pre-fix numbers, so closing the client-side gaps does not cost measurable throughput.

What the timeout change itself earns

Removing the deadline machinery does not show up while the call rate is limited by the server. On add (64 concurrent calls, client CPU measured through wait4 rusage) B is 2.779 us/op against 2.823 for T, which is inside run-to-run noise. On a payload path the same measurement separates the two effects:

variant ops/s client CPU us/op socket writes per call
B (2.0.0) 349,301 10.521 0.31
T (budget change only) 280,351 13.737 1.18
V (this branch) 462,344 7.955 0.30

(echo-4096, 15 s, c64, one run each; the write counts are from the strace run above at the same op.)

The middle row is why this PR needed the second commit. The old pump batched as a side effect: the moment a batch held a timed frame, the frames after it were copied together in one span at emission, and every request in this benchmark carries the probe's 30 s timeout, so in practice every frame was batched. Dropping the deferral sent each frame back to an arrival-time GetSpan, and a 4 KiB frame does not fit in what is left of the 4 KiB writer buffer, so each frame forced its own socket write - 1.18 per call against 0.31, which costs 3.2 us/op more.

V restores the batched copy explicitly (0.30 writes per call) and lands 2.6 us/op below B at an equal write count. That residual gap is the per-frame emission pass this PR deletes: the HasTimeBudget header inspection, the clock sample, the in-span budget patch and the expired-frame compaction with its RemoveAt bookkeeping. So the deletion is not a throughput change on its own; it removes a pass that only becomes measurable once the copy layout is fixed, and it removes a second terminal writer that could complete a call with a synthetic DeadlineExceeded while the pending-request table, the authoritative terminal, had already chosen another reason.

The lightweight control paths (add, oneway, echo-256) are untouched by this change and stay below 1.1.1 by 10-20%; that gap predates this PR and is not affected by it.

Validation

  • Solution builds with 0 warnings / 0 errors.
  • Unit 1845/1845, Integration 451/451, maintainability debt gate and whitespace format gates pass.
  • The send-pump budget tests were rewritten to the new contract: the wire budget is the creation-time value even when the clock advances before flush, an already-expired request is still published in order instead of being compacted, a timed call blocked in the transport write queue still fails through its pending deadline, and timed unary / client-stream / oneway-client-stream calls publish the creation-time budget.
  • SendPumpTimedWaitStopTests was adapted: the pump is still blocked inside the transport write when the stop is latched (its own flush triggers that write), so the "stop latched before the timed wait arms" race is still covered.
  • The syscall comparison above was taken with strace -f -c -e trace=write,writev,sendto,sendmsg on the client process only.
  • SharpLinkClientDeadlinePublicationTests covers the publication and deadline guarantees with eleven tests, five of which fail on the commit before these fixes:
    • a timed OneWay with client streams whose frame the send queue refuses fails through its elapsed deadline and emits neither a Request nor a Cancel (the refused frame used to be marked published and cancelled);
    • a deadline that elapses inside a user-provided compression provider leaves the Request unpublished (preparation used to run inside the completion gate);
    • a timed unary whose codec moves the clock past its deadline without running the deadline timer is refused by the publication gate itself, and emits neither a Request nor a Cancel (the gate used to read the elapsed deadline off the scheduler callback);
    • a timed OneWay with client streams fails at its deadline while the transport write is stalled, without starting its producer, and still publishes Request then Cancel once the transport recovers;
    • a timed client-streaming call behaves the same way, and a timed duplex call additionally leaves no DuplexStreamingInvoker framework task behind;
    • a timed plain OneWay whose budget elapsed during serialization is not published at all, and one whose budget elapses inside the compression provider is refused at the linearization point;
    • a timed plain OneWay that loses its emission race fails at its deadline and publishes Request then Cancel with the negotiated DeadlineExceeded reason, in that order;
    • plus the tests already in the PR: a timed plain OneWay failing at its own deadline while the transport write is stalled (its already-published Request still carries the creation-time budget), and a deadline that elapses inside a custom codec's Serialize failing without publishing either half of the pair.
  • SharpLinkClientTimeBudgetTests now also pins the benchmark shape to the wire: only the unary entry point resolves the client default timeout, so add/echo Requests carry a budget and the one-way / client-streaming / server-streaming / duplex Requests carry none, and an untimed OneWay does not fail or cancel when that default elapses.

@SunSi12138
SunSi12138 force-pushed the refactor/client-timeout-single-sample branch 2 times, most recently from 7c86db0 to b7feb2d Compare September 13, 2026 06:22
The send pump owned a per-frame process-local RpcDeadline so it could stamp the
wire TimeBudget at emission, defer timed writes past the batching window, and
compact requests whose deadline elapsed while they waited in the local queue.

Sample the remaining budget once, when the frame is built, and let the pump copy
frames verbatim. The end-to-end deadline stays authoritative on the caller: the
client pending-deadline timer still fails a call whose Request never reaches the
transport, and cancellation still terminates a remote call that already started.
Writing each frame into the transport pipe as it arrives costs one GetSpan per
frame, which fragments the outgoing segments at larger payload sizes. Copy the
batch into a single span at flush time instead; the frames were already
serialized with their wire TimeBudget, so this copy still inspects no deadline.
@SunSi12138
SunSi12138 force-pushed the refactor/client-timeout-single-sample branch from b7feb2d to acf1ec0 Compare September 13, 2026 06:27
The flush policy gate no longer relates to per-frame deadlines after the
time budget is sampled once at frame creation, so rename the flag and the
affected test fixture to describe the explicit batch window.
Two deadline gaps were left open once the send pump stopped compacting
expired frames:

- A plain OneWay registers no pending entry, so nothing enforced the
  caller's deadline while the pump held its Request. Race the emission
  wait against the sampled deadline; the frame is already published, so
  losing the race only stops the caller from blocking past its own
  lifetime.
- A cancellable call armed its pending deadline before the Request was
  serialized and queued. A deadline that elapsed during serialization
  emitted a Cancel first and then published the Request after it, which
  the peer dispatches - the cancel is discarded for an unknown request.
  Publish the frame under the owning call's completion gate instead, and
  emit the transport cancel only when the Request actually reached the
  queue. Local dispatcher cleanup still runs for every terminal reason.
Move TryPublishRequest into its own partial file so
PendingRequestTable.cs stays within the maintainability baseline.
Frame preparation now runs before the publication decision instead of inside
the completion gate, so a user-provided compression provider never runs while a
call's terminal arbitration is blocked and its faults surface synchronously,
before the caller decides whether the frame was published. The frame is marked
published only for a frame the send queue accepted, so a refused Request can no
longer produce a cancel for a Request the peer never received.

A plain OneWay that loses its emission race against its own deadline now also
publishes the deadline cancel behind its Request, so the peer stops work whose
caller has already given up, and one whose budget elapsed while its payload was
still being serialized is not published at all.
…he transport

A call that owns a pending entry waited for the emission of its head Request
on the bare caller token. The deadline, a caller cancellation and a closed
connection all end such a call by completing its pending entry and cancelling
its producer token, but none of them cancelled the token the emission wait was
observing, so the invocation stayed parked in the send path until the transport
recovered even though the call already had a terminal reason.

The OneWay-with-client-stream, client-streaming and duplex starts now wait on
the pending call's own producer token, which is the token that ownership model
cancels. A shape without a producer entry keeps the caller token, and a shape
that waits on nothing (server streaming, whose caller is not blocked on this
path) is unchanged.

A plain OneWay also only checked its budget before frame preparation, so a
Request whose budget elapsed inside the compression provider was published and
then cancelled. It now re-checks after preparation, immediately before the
linearization point, and returns the prepared frame instead of admitting it.
TryPublishRequest only checked that the slot was still alive before
admitting the frame, so whether an elapsed deadline had already won was
read off the PendingDeadlineScheduler callback rather than decided under the
same completion gate. A scheduler callback that runs late let a Request whose
budget had already elapsed reach the peer, carrying the creation-time budget
it captured before a slow codec or compression provider.

The gate now decides both boundaries together, exactly as an inbound response
already does in TryTakeMatchingCall: while it holds the completion gate it
checks IsExpired, atomically removes the slot when the deadline won, and
completes the call with DeadlineExceeded outside the gate. Only a live
deadline can admit a frame and mark it published.

The refused-publication path now reaches the client-stream producer with an
already cancelled producer token, so the test writer honours that token
instead of parking the invocation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant