refactor(runtime): sample the request time budget once at frame creation - #684
Merged
Merged
Conversation
SunSi12138
force-pushed
the
refactor/client-timeout-single-sample
branch
2 times, most recently
from
September 13, 2026 06:22
7c86db0 to
b7feb2d
Compare
The send pump owned a per-frame process-local RpcDeadline so it could stamp the wire TimeBudget at emission, defer timed writes past the batching window, and compact requests whose deadline elapsed while they waited in the local queue. Sample the remaining budget once, when the frame is built, and let the pump copy frames verbatim. The end-to-end deadline stays authoritative on the caller: the client pending-deadline timer still fails a call whose Request never reaches the transport, and cancellation still terminates a remote call that already started.
Writing each frame into the transport pipe as it arrives costs one GetSpan per frame, which fragments the outgoing segments at larger payload sizes. Copy the batch into a single span at flush time instead; the frames were already serialized with their wire TimeBudget, so this copy still inspects no deadline.
SunSi12138
force-pushed
the
refactor/client-timeout-single-sample
branch
from
September 13, 2026 06:27
b7feb2d to
acf1ec0
Compare
The flush policy gate no longer relates to per-frame deadlines after the time budget is sampled once at frame creation, so rename the flag and the affected test fixture to describe the explicit batch window.
Two deadline gaps were left open once the send pump stopped compacting expired frames: - A plain OneWay registers no pending entry, so nothing enforced the caller's deadline while the pump held its Request. Race the emission wait against the sampled deadline; the frame is already published, so losing the race only stops the caller from blocking past its own lifetime. - A cancellable call armed its pending deadline before the Request was serialized and queued. A deadline that elapsed during serialization emitted a Cancel first and then published the Request after it, which the peer dispatches - the cancel is discarded for an unknown request. Publish the frame under the owning call's completion gate instead, and emit the transport cancel only when the Request actually reached the queue. Local dispatcher cleanup still runs for every terminal reason.
Move TryPublishRequest into its own partial file so PendingRequestTable.cs stays within the maintainability baseline.
Frame preparation now runs before the publication decision instead of inside the completion gate, so a user-provided compression provider never runs while a call's terminal arbitration is blocked and its faults surface synchronously, before the caller decides whether the frame was published. The frame is marked published only for a frame the send queue accepted, so a refused Request can no longer produce a cancel for a Request the peer never received. A plain OneWay that loses its emission race against its own deadline now also publishes the deadline cancel behind its Request, so the peer stops work whose caller has already given up, and one whose budget elapsed while its payload was still being serialized is not published at all.
…he transport A call that owns a pending entry waited for the emission of its head Request on the bare caller token. The deadline, a caller cancellation and a closed connection all end such a call by completing its pending entry and cancelling its producer token, but none of them cancelled the token the emission wait was observing, so the invocation stayed parked in the send path until the transport recovered even though the call already had a terminal reason. The OneWay-with-client-stream, client-streaming and duplex starts now wait on the pending call's own producer token, which is the token that ownership model cancels. A shape without a producer entry keeps the caller token, and a shape that waits on nothing (server streaming, whose caller is not blocked on this path) is unchanged. A plain OneWay also only checked its budget before frame preparation, so a Request whose budget elapsed inside the compression provider was published and then cancelled. It now re-checks after preparation, immediately before the linearization point, and returns the prepared frame instead of admitting it.
TryPublishRequest only checked that the slot was still alive before admitting the frame, so whether an elapsed deadline had already won was read off the PendingDeadlineScheduler callback rather than decided under the same completion gate. A scheduler callback that runs late let a Request whose budget had already elapsed reach the peer, carrying the creation-time budget it captured before a slow codec or compression provider. The gate now decides both boundaries together, exactly as an inbound response already does in TryTakeMatchingCall: while it holds the completion gate it checks IsExpired, atomically removes the slot when the deadline won, and completes the call with DeadlineExceeded outside the gate. Only a live deadline can admit a frame and mark it published. The refused-publication path now reaches the client-stream producer with an already cancelled producer token, so the test writer honours that token instead of parking the invocation.
This was referenced Sep 13, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Every Request frame carried a process-local
RpcDeadlinefrom the caller all the way into the send pump. The pump used it toTimeBudgetat emission time, in a bounded pass that re-sampledGetRemainingper frame inside an already-acquired output span,writtenCountbookkeeping,DeadlineExceeded,PooledByteBufferWriter.EmissionDeadline) or, for custom writers, in a boxed completion state insideOwnedFrame.None of that changes the caller-visible outcome: the caller's failure is driven by the client's pending-deadline timer and by cancellation, not by the identity of the frame that carried the deadline.
Change
The remaining budget is sampled once, when the request frame is built, and written into the
TimeBudgetfield of the frame. The pump never samples, rewrites, or compacts a deadline.Removed: deferred timed writes and the
deferWritesbookkeeping,WriteRetainedBatchAtEmission, the per-frameGetRemainingre-stamp, expired-frame compaction withCompleteExpiredBatch,OwnedFrame.Deadlineand itsDeadlineState/CompletionDeadlineState/FailureDeadlineStatewrappers, andPooledByteBufferWriter.EmissionDeadlinewith all its lease-reset sites.Kept, because it is what actually enforces the deadline end to end:
Wire consequence, and the reason this needs a changelog entry: the server-visible
TimeBudgetis sampled earlier than before, so it is a strictly more generous value than the caller's own remaining lifetime. The sampling point is the frame header, which is written before the request payload and metadata are serialized, so three things are no longer deducted from the value the server sees:The server budget is therefore advisory: it bounds how long the server is willing to keep working, and it is not the caller's deadline. The caller's own end-to-end deadline stays authoritative, which is the gRPC model - a deadline is fixed at invocation time and is not refreshed as the call moves down the stack. Two client-side guarantees hold that up, and both have deterministic tests in this PR:
Canceland thenRequest- a pair the peer answers by discarding the cancel for an unknown request and dispatching the Request anyway - and a frame the queue refused was still reported as published, so its terminal transition emitted a cancel for a Request the peer never received.DeadlineExceededinstead of blocking past its lifetime, and publishes the deadline cancel behind the Request it follows in the same normal queue, so the peer stops work whose caller has already given up.A timed Request also now waits for the configured batch window like any other frame when the caller sets an explicit
RpcSessionFlushOptions.MaxLatency; under the default profiles the batch window is zero, so timed frames are still emitted as soon as the queue drains.The batch is still copied into one output span. The deadline machinery was also doing the transport a favour: the frames it deferred were serialized in a single
GetSpanat flush time, while untimed frames were copied one at a time as they arrived. Removing the deferral without replacing that would have cost real throughput - the first measurement below showsecho-4096losing 15% against 2.0.0 with arrival-time copies. The pump now coalesces the whole batch into one span at flush time (WritePendingBatch) and inspects no deadline while doing it.Measurement (baremetal UDS, client cores 0-3 / server cores 4-7, c64, 20 s + 5 s warmup, 3 reps)
Variants:
A= 1.1.1,B= 2.0.0main,T=main+ budget change only (no batch coalescing, no #681/#682),D=dev,U=dev+ budget change,V=dev+ budget change + batch coalescing (this branch).Attribution on the 2.0.0 base
Confirmation on the dev base
UagainstDisolates the same change on top of #681/#682 and reproduces the regression, so neither the allocation work nor #682 is the cause:Final result
Per-operation allocation is unchanged from
D/U(add 382.7 B/op against 382.1 for 1.1.1, echo-256 677.7 against 679.6), so none of this is bought with pooling or with a smaller working set. Latency follows throughput:echo-4096p50 155 -> 101 us,upload-4096p50 1463 -> 674 us,download-4096p50 1446 -> 606 us,duplex-4096p99 3474 -> 2716 us (B -> V).The size of the win is explained by what the old path did at 4 KiB:
Outputis aStreamPipeWriterover the socket stream with the default 4 KiB buffer, so a frame plus its header never fits in what is left of a buffer. Writing each frame as it arrived therefore committed the buffered bytes to the socket once per frame - the pump's 16 KiB flush threshold only decided how often it awaitedFlushAsync, not how often the socket was actually written. Counting syscalls with the same client pinned to the same cores (10 s, c64,upload-4096, sostraceoverhead dominates the absolute numbers):sendtocallsUnder
stracethe four-times-smaller syscall count per operation shows up as four times the throughput (4.0x, against 1.9x withoutstrace). The total number of bytes copied into the pipe is the same either way; only the destination layout changed, so this is fewer syscalls rather than a trade against the working set.Timeout-enabled benchmark
The probe contract puts no
[Timeout]on any of its methods, and only the unary entry points inherit the client's default request timeout (InvokeUnaryAsyncresolves its control withincludeClientDefault: true; OneWay, client-streaming, server-streaming, and duplex resolve withincludeClientDefault: false). So the deadline machinery is live foraddandecho-4096, and the OneWay / streaming rows are deadline-free workloads that no configured client default reaches - the-1.3%ononewayis not a measurement of this change. That split is asserted on the wire rather than inferred:ClientDefaultTimeoutShouldReachOnlyTheShapesThatOptInruns a client with a 30 s default over all five shapes and checks that only the unary Request carriesHasTimeBudget, with the 30 s value, while the one-way, client-streaming, server-streaming and duplex Requests carry none;ClientDefaultTimeoutShouldNotEnforceAPlainOneWayDeadlinestalls the transport, lets that default elapse, and shows the untimed OneWay neither fails nor publishes a cancel. To show the live rows are not an artifact of a 30 s budget, the probe was rebuilt with an env-controlled timeout and rerun with a 1 s client default budget, same cores, same server, 20 s + 5 s warmup:The same probe also re-measures this branch after the publication and plain-OneWay deadline fixes at the original 30 s budget: add 1,541,890, oneway 4,397,626, echo-4096 514,681, upload-4096 87,604, download-4096 95,220. Those are within the run-to-run spread of the pre-fix numbers, so closing the client-side gaps does not cost measurable throughput.
What the timeout change itself earns
Removing the deadline machinery does not show up while the call rate is limited by the server. On
add(64 concurrent calls, client CPU measured throughwait4rusage)Bis 2.779 us/op against 2.823 forT, which is inside run-to-run noise. On a payload path the same measurement separates the two effects:(
echo-4096, 15 s, c64, one run each; the write counts are from thestracerun above at the same op.)The middle row is why this PR needed the second commit. The old pump batched as a side effect: the moment a batch held a timed frame, the frames after it were copied together in one span at emission, and every request in this benchmark carries the probe's 30 s timeout, so in practice every frame was batched. Dropping the deferral sent each frame back to an arrival-time
GetSpan, and a 4 KiB frame does not fit in what is left of the 4 KiB writer buffer, so each frame forced its own socket write - 1.18 per call against 0.31, which costs 3.2 us/op more.Vrestores the batched copy explicitly (0.30 writes per call) and lands 2.6 us/op belowBat an equal write count. That residual gap is the per-frame emission pass this PR deletes: theHasTimeBudgetheader inspection, the clock sample, the in-span budget patch and the expired-frame compaction with itsRemoveAtbookkeeping. So the deletion is not a throughput change on its own; it removes a pass that only becomes measurable once the copy layout is fixed, and it removes a second terminal writer that could complete a call with a syntheticDeadlineExceededwhile the pending-request table, the authoritative terminal, had already chosen another reason.The lightweight control paths (
add,oneway,echo-256) are untouched by this change and stay below 1.1.1 by 10-20%; that gap predates this PR and is not affected by it.Validation
SendPumpTimedWaitStopTestswas adapted: the pump is still blocked inside the transport write when the stop is latched (its own flush triggers that write), so the "stop latched before the timed wait arms" race is still covered.strace -f -c -e trace=write,writev,sendto,sendmsgon the client process only.SharpLinkClientDeadlinePublicationTestscovers the publication and deadline guarantees with eleven tests, five of which fail on the commit before these fixes:Requestnor aCancel(the refused frame used to be marked published and cancelled);Requestnor aCancel(the gate used to read the elapsed deadline off the scheduler callback);RequestthenCancelonce the transport recovers;DuplexStreamingInvokerframework task behind;RequestthenCancelwith the negotiatedDeadlineExceededreason, in that order;Serializefailing without publishing either half of the pair.SharpLinkClientTimeBudgetTestsnow also pins the benchmark shape to the wire: only the unary entry point resolves the client default timeout, soadd/echoRequests carry a budget and the one-way / client-streaming / server-streaming / duplex Requests carry none, and an untimed OneWay does not fail or cancel when that default elapses.