Skip to content

perf(frame): borrow the payload from the buffer pool and send the frame as one slice - #222

Open
darakanoit wants to merge 12 commits into
roadrunner-server:masterfrom
darakanoit:perf/frame-borrowed-payload
Open

darakanoit wants to merge 12 commits into
roadrunner-server:masterfrom
darakanoit:perf/frame-borrowed-payload

Conversation

@darakanoit

@darakanoit darakanoit commented Sep 22, 2026 •

Copy link
Copy Markdown

Reason for This PR

Design discussion: roadrunner-server/roadrunner#2402. Absorbs #221 (roadrunner-server/roadrunner#2401) at the maintainer's request, since both touch the same files.

Two problems in the tiered buffer pool, which since #220 serves both the send and the receive path.

The smallest tier is 1 MB. Every payload up to 1 MB, and every options read of 4 to 40 bytes, takes a 1 MB buffer while in flight, so resident memory grows with the number of workers, not with the size of the data: 64 concurrent 200-byte receives hold 67 MB. put routes by the requested size, not by the buffer, so a buffer above 10 MB allocated with make ends up in the 10 MB tier and is handed out and retained from then on.

The pool is a staging area. ReceiveFrame reads the body into a pooled buffer and copies it into the frame, SendFrame copies header and payload into a pooled buffer to write once. Both copies go away when the pooled buffer is the payload, the way mem.BufferSlice works in grpc-go.

Description of Changes

  • Tiers of 4 KB, 16 KB, 64 KB and 256 KB below the existing 1 MB, 5 MB and 10 MB, one pool per tier selected with a switch, and Put routed by the buffer's capacity: a buffer goes back only to the tier whose size it has, anything else is dropped. Same rule as putDataBufferChunk in x/net/http2 and BinaryTieredBufferPool.Put in grpc-go.
  • internal/bpool is a package now, with Get and Put exported and the tiers initialized statically. Preallocate is gone.
  • Frame borrows its payload from the pool. AllocPayload(n) resizes the payload in place for the caller to fill; WritePayload copies through it. The pooled pointer is the ownership token: memory from From, ReadFrame or ReadHeader is written in place when it fits, as before, and never enters the pool.
  • Reset returns a buffer above 4 KB to its tier and keeps a 4 KB one on the frame. A frame no longer retains the largest payload it ever carried.
  • A pooled payload starts Headroom (52 bytes) into its buffer. Wire() writes the header into that room and returns the contiguous frame; SendFrame writes it with one call at any size, no copy, no scratch.
  • ReceiveFrame reads the options into the header's spare capacity and the body straight into AllocPayload.
  • Frames over caller memory keep the assembled path up to 64 KB and go through a pooled net.Buffers above it: one writev on a socket, two writes on a pipe. The threshold is measured, BenchmarkPipeSendStrategy and BenchmarkSocketSendStrategy stay in the tree.
  • Wire format unchanged. Payload() is valid until the next write or Reset, as documented since perf: reuse frame buffers on the send and receive paths #220.

Contract change: Payload() is empty after Reset, nil when the buffer went back to the pool. Every consumer in the org copies before Reset (pool worker.receiveFrame, both RPC codecs).

Against master after #220 for the pool change and against the pool change for the rest, interleaved runs, benchstat over 10 samples each, M3. First the pool change alone:

Path master with the tiers
ReceivePath 1 KB 95.3 ns 71.5 ns
ReceivePath 64 KB 2.02 µs 1.73 µs
SendPath 1 KB 115 ns 91.6 ns
64 concurrent 200-byte receives 67.2 MB 0.3 MB

Then the borrowed payload on top of it:

Path with the tiers this PR
ReceivePath 1 KB 68.9 ns 43.1 ns
ReceivePath 64 KB 1.69 µs 0.92 µs
ReceivePath 1 MB 37.3 µs 18.9 µs
ReceivePath + clone 200 B 94.5 ns 76.7 ns
ReceivePath + clone 4 KB 425 ns 377 ns
ReceivePath + clone 64 KB 6.66 µs 4.70 µs
ReceivePath + clone 1 MB 72.2 µs 52.5 µs
SendPath 1 KB 89.9 ns 75.1 ns
SendPath 64 KB 4.13 µs 1.76 µs
SendPath 1 MB 55.9 µs 37.5 µs

allocs/op identical in every row. BenchmarkReceivePathWithClone mirrors worker.receiveFrame in pool including its two bytes.Clone; the clones are most of what is left.

Resident memory with 64 concurrent 200-byte receives: 0.39 MB.

Tests: pointer identity across Put and Get is checked in !race files under GOMAXPROCS(1), so Makefile and the Linux workflow run pkg/frame without the race detector too, as #220 did for internal. Round trips of a 1 MB frame over a real pipe and a TCP loopback, a header that announces a 4 GB payload, Wire refusing a header longer than the headroom, and Reset never pooling caller memory are covered.

Second step, separate PR in pool: sendFrame writes into AllocPayload instead of a bytes.Buffer, receiveFrame hands the borrowed buffer on through payload.Payload, the http plugin releases it after the response is written.

closes: roadrunner-server/roadrunner#2401
closes: roadrunner-server/roadrunner#2402

License Acceptance

By submitting this pull request, I confirm that my contribution is made under
the terms of the MIT license.

PR Checklist

  • All commits in this PR are signed (git commit -s) or (git commit -S).
  • The reason for this PR is clearly provided (issue no. or explanation).
  • The description of changes is clear and encompassing.
  • Any required documentation changes (code and docs) are included in this PR.
  • Any user-facing changes are mentioned in CHANGELOG.md.
  • All added/changed functionality is tested.

The buffer pool served every request up to 1 MB from the 1 MB tier, so a
200-byte payload held a 1 MB buffer while in flight. With 64 workers
receiving at once that is 64 MB of resident memory for a few kilobytes
of data, and every pool miss zeroes a full megabyte.

put routed a buffer by the requested size, not by its capacity. A buffer
above 10 MB is allocated with make and was stored in the 10 MB tier,
where it was handed out for every later request in that tier and never
released while the pool stayed warm.

Add 4 KB, 16 KB, 64 KB and 256 KB tiers below the existing 1 MB, 5 MB
and 10 MB ones, hold them in an array instead of a sync.Map, and route
put by the buffer's capacity: a buffer goes back only to the tier whose
size it has, anything else is dropped.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
The frame package will borrow payload buffers from the pool, and it
cannot import internal without a cycle. bpool now lives in
internal/bpool with an exported Get and Put. The tiers are initialized
in the package variable, so Preallocate and the sync.Once are gone
together with the calls in the pipe and socket relays.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
A frame now takes its payload buffer from bpool on the first write and
returns it in Reset. The pooled pointer is the ownership token: a
payload that aliases caller memory, as built by From, ReadFrame or
ReadHeader, has none and is only truncated by Reset, so memory that
was not taken from the pool never enters it.

AllocPayload resizes the payload for the caller to fill in place, which
lets the relay read the body straight into the frame. WritePayload
copies through it and keeps its semantics.

A pooled frame no longer keeps the largest payload it ever carried:
the buffer goes back to its tier and is shared by every frame.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
ReceiveFrame no longer stages the options and the payload in pool
buffers and copies them into the frame. The options are read into
the header's spare capacity and the body into the buffer that
AllocPayload borrows for the frame, one copy less per frame. A
payload-less frame leaves the payload empty regardless of what the
frame held before.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
…them

SendFrame copies a frame into a pooled buffer only up to 64 KB, where
one write beats a memcpy of the body. Above that the header and the
payload go to the writer as net.Buffers: a single writev on a socket,
two writes on a pipe, and no copy of the body either way. The vector
itself comes from a pool, so the path stays allocation free.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
…and socket

The threshold in SendFrame is a measured value, not a guess. The
benchmark keeps the pre-vector strategy next to writeVector so the
comparison can be rerun on any machine. On an M3 the vector wins on
a pipe from 64 KB up and is level on a TCP socket.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
Returning every payload buffer in Reset cost a pool round trip per
frame, which the 1 KB send path paid for without saving a copy: +22 %
on BenchmarkSendPath/1KB. A buffer of the smallest tier now stays on
the frame, bounded at 4 KB per pooled frame, and only larger buffers
go back to the pool. AllocPayload keeps the reuse check first and
moves the remaining cases to a helper.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
…s one slice

A pooled payload now starts Headroom bytes into its buffer, enough
for a header with the maximum of 10 options. Wire writes the header
into that room and returns the contiguous frame, so SendFrame hands
it to a single Write at any size: no scratch buffer, no pool round
trip, no copy of the body. The assembled and vectored paths remain
for frames over caller memory, which have no headroom.

BenchmarkSendPath/1KB, which paid a pool round trip since the payload
became pooled, is now 18 % faster than before the change.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
…cannot wrap

AllocPayload asked the pool for uint32(n + Headroom). A header that
announces a payload of 0xFFFFFFFF bytes, which any peer can send, made
that wrap to 51, and slicing the 4 KB buffer to the announced length
panicked in the receive path. The base returned a read error for the
same input. bpool.Get now takes an int, and the request is computed in
int all the way. Regression test added, skipped on Windows because the
announced length is reserved before the read fails.

Also covered: Wire refuses a header longer than Headroom, and Reset
returns the buffer that a From frame grew into.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
Comment thread internal/bpool/bpool.go Outdated
…er the tiers

Review note on roadrunner-server#222. The loop cost grew with the tier: Get plus Put
took 7.7 ns on the 4 KB tier, 10.0 ns on 64 KB and 11.7 ns on 5 MB.
With one pool per tier and a switch in both directions the cost is
flat at about 7.7 ns, 23 % less on 64 KB and 29 % less on 5 MB.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
darakanoit added a commit to darakanoit/goridge that referenced this pull request Sep 22, 2026
… over the tiers

Review note on roadrunner-server#222, which carries this file. The loop cost grew with
the tier: get plus put took 7.7 ns on the 4 KB tier, 10.0 ns on 64 KB
and 11.7 ns on 5 MB. With one pool per tier and a switch in both
directions the cost is flat at about 7.7 ns.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
@codecov

codecov Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.30137% with 10 lines in your changes missing coverage. Please review.
✅ Project coverage is 77.49%. Comparing base (d1d60f5) to head (20438cb).

Files with missing lines Patch % Lines
pkg/frame/pool.go 63.63% 8 Missing ⚠️
pkg/frame/frame.go 92.30% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master     #222      +/-   ##
==========================================
- Coverage   78.56%   77.49%   -1.07%     
==========================================
  Files           9        9              
  Lines         695      702       +7     
==========================================
- Hits          546      544       -2     
- Misses        149      158       +9     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@darakanoit
darakanoit marked this pull request as ready for review September 23, 2026 10:53
Comment thread internal/bpool/bpool_test.go Outdated
Comment thread internal/alloc_test.go Outdated
Review notes on roadrunner-server#222. The pool cannot go back to internal, which
imports frame: the frame is the only thing that borrows buffers now,
so the pool lives next to it, unexported, and internal no longer
touches it. With the headroom in place the assembled send path only
served frames over caller memory, which nothing in the org builds;
those go through net.Buffers at any size, and the 64 KB threshold
and its strategy benchmark are gone with it.

The pool tests and the extra allocation tests in internal are
removed as requested.

Signed-off-by: darakanoit <dara.kamaliev@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants