Skip to content

Refactor: hbg records and ships Definitions in one retained block - #1988

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:refactor/hbg-definitions-in-one-retained-block
Aug 24, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:refactor/hbg-definitions-in-one-retained-block

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Summary

Each distinct Graph Definition was uploaded as its own device object: for every one, bind acquired a retained device block, allocated a host staging vector, assembled [GraphDefinitionHeader][Definition image] into it, issued its own rtMemcpy, and freed the vector — and the image it copied from was itself a vector the recording had allocated and would return. A dsv4 bind has eight Definitions, so that is eight allocate-copy-free rounds for 1.0 MB, on a path whose cost is per-call rather than per-byte.

The objects now share one device block and one host staging block, both retained per pipeline slot and grown to a high-water mark, and the recordings build their images directly into that staging:

  • bind reads the retained staging before orchestration and hands it to the Graph host state as a GraphDefinitionArena. The capacity on offer is whatever the previous bind left, because a run's total is not known until every recording has ended.
  • graph_build_definition splits at the point the size becomes known. graph_layout_definition settles the counts, section offsets and total_bytes without writing anything; the recording thread claims an object of that size; graph_fill_definition writes the image at the claimed offset.
  • the upload writes the object headers into the room the arena left in front of each image and ships the used prefix with one copy_to_device. A steady-state bind copies no image at all, and acquires no host memory for its Definitions.

Three properties make the arena safe to write into from several recording threads while the submitting thread is still submitting shells:

  • the claim is a compare-exchange that fails rather than advancing the cursor past the capacity, so a Definition that does not fit costs the run its own slot and no one else's. It is built in a buffer of its own instead and copied in by the upload, which then grows the block — so outgrowing the retained capacity costs a copy, not the run, and the next bind is sized for it. graph_upload's new spilled= counts those, and is 0 on every bind after the first.
  • the arena never moves during orchestration, since a recorder holds addresses inside it. Only the upload may grow it, which preserves the content and is why a published record names an offset rather than an address; by then nothing is recording.
  • an object is aligned and padded to GRAPH_DEFINITION_OBJECT_ALIGN, a named constant beside the header whose size is a multiple of it, so every image base carries the alignment its section offsets assume. The fill asserts that rather than trusting the arena's provider.

The image is still zero-filled before it is written, and now for a second reason: the alignment slack between sections is inside the hashed range, so a slot a previous bind wrote would give two structurally identical Definitions unequal content hashes.

HostApi grows get_graph_definition_staging and replaces acquire_graph_definition_buffer with acquire_graph_definition_block, which takes no Definition key and returns both blocks, since the layout inside them belongs to the caller. The device block is still zeroed when it is (re)allocated, so a region no upload has covered reads as zero and fails the magic gate instead of being read as a stale object.

Measurement

dsv4, interleaved base, measure, base, measure at six rounds each per .claude/skills/hbg-bind-phases; both arms ran the same command string (the logs' [stamp] lines differ only in the commit) and did the same work — 8 Definitions, 86 Graph submissions, 129 host tasks per bind. Two figures per cell, one per repetition:

per warm bind base this branch
graph_upload minflt, min / median / max 2 / 38 / 47, 34 / 43 / 133 0 / 1 / 3, 0 / 1 / 5
graph_upload dur min (ms) 0.225, 0.282 0.141, 0.106
host_orch minflt, min / median 1100 / 1248, 1047 / 1138 1062 / 1107, 1082 / 1124
control plane, min of sums (ms) 1.194, 1.365 1.549, 1.199

Only graph_upload moved, and it moved consistently in both repetitions: the ~40 minor faults per bind it took allocating one staging vector per Definition are gone, and so is the copy.

host_orch did not move, and the investigation entry now records why: the two repetitions disagree in sign on both its fault count and its duration, which is that file's own criterion for "not resolvable". A freed 126 KB block is reused from the heap without re-faulting, so an image vector was never part of that segment's fault tail — size against glibc's mmap and trim thresholds, not byte count, decides what shows up there. The control-plane total is not resolvable either (+0.355 ms then −0.166 ms), which is what a ~0.1 ms change looks like on a box whose load average sat between 40 and 66.

Testing

  • Simulation tests pass — examples tests/st on a2a3sim (21/21 cases) and a5sim (17/17), both after the rebase onto 66ba5c4a1
  • Hardware tests pass — examples tests/st -m 'not sdma' --platform a2a3 --exclude-level 4 under task-submit, 150 selected / 57 resource-phase cases, twice
  • ctest -LE requires_hardware — 119/119, including the new test_hbg_graph_definition_arena for a2a3 and a5
  • pytest tests/ut -m "not requires_hardware" — 1886 passed, 7 skipped
  • spilled= verified across rounds: 2 of 12 binds (one per rank, cold) spill; the other 10 report spilled=0

test_hbg_graph_definition_arena covers what the scene tests cannot see, since they pass either way: objects land at aligned, disjoint offsets that the claimed prefix exactly covers, and a run with no arena or with one too small still publishes a valid image.

One caveat from the hardware runs, recorded rather than swept under: the first corpus run on this branch hit a device-side AIV ub address out of bounds in worker_async_endpoint, which cascaded to 507018 / SCHEDULER_TIMEOUT. It did not reproduce — that case passed 3/3 alone on 66ba5c4a1, the same corpus passed 57/57 on 66ba5c4a1, and the second corpus run on this branch passed 57/57 — and the case records no Graph at all, so this diff does not execute in it. Worth an eye on CI: a UB fault is not a flaky timeout.

Each distinct Graph Definition was uploaded as its own device object: for
every one, bind acquired a retained device block, allocated a host staging
vector, assembled [GraphDefinitionHeader][Definition image] into it, issued
its own rtMemcpy, and freed the vector -- and the image it copied from was
itself a vector the recording had allocated and would return. A dsv4 bind
has eight Definitions, so that is eight allocate-copy-free rounds for
1.0 MB, on a path whose cost is per-call rather than per-byte.

The objects now share one device block and one host staging block, both
retained per pipeline slot and grown to a high-water mark, and the
recordings build their images directly into that staging:

  - bind reads the retained staging before orchestration and hands it to the
    Graph host state as a GraphDefinitionArena. The capacity on offer is
    whatever the previous bind left, because a run's total is not known until
    every recording has ended.
  - graph_build_definition splits at the point the size becomes known.
    graph_layout_definition settles the counts, section offsets and
    total_bytes without writing anything; the recording thread claims an
    object of that size; graph_fill_definition writes the image at the
    claimed offset.
  - the upload writes the object headers into the room the arena left in
    front of each image and ships the used prefix with one copy_to_device.
    A steady-state bind therefore copies no image at all, and acquires no
    host memory for its Definitions.

Three properties make the arena safe to write into from several recording
threads while the submitting thread is still submitting shells:

  - the claim is a compare-exchange that fails rather than advancing the
    cursor past the capacity, so a Definition that does not fit costs the run
    its own slot and no one else's. It is built in a buffer of its own
    instead and copied in by the upload, which then grows the block -- so
    outgrowing the retained capacity costs a copy, not the run, and the next
    bind is sized for it. graph_upload's new `spilled=` counts those, and is
    0 on every bind after the first.
  - the arena never moves during orchestration, since a recorder holds
    addresses inside it. Only the upload may grow it, which preserves the
    content and is why a published record names an offset rather than an
    address; by then nothing is recording.
  - an object is aligned and padded to GRAPH_DEFINITION_OBJECT_ALIGN, a named
    constant beside the header whose size is a multiple of it, so every image
    base carries the alignment its section offsets assume. The fill asserts
    that rather than trusting the arena's provider.

The image is still zero-filled before it is written, and now for a second
reason: the alignment slack between sections is inside the hashed range, so a
slot a previous bind wrote would give two structurally identical Definitions
unequal content hashes.

HostApi grows get_graph_definition_staging and replaces
acquire_graph_definition_buffer with acquire_graph_definition_block, which
takes no Definition key and returns both blocks, since the layout inside them
belongs to the caller. The device block is still zeroed when it is
(re)allocated, so a region no upload has covered reads as zero and fails the
magic gate instead of being read as a stale object.

Measured on dsv4, interleaved base/measure/base/measure at six rounds each
(8 Definitions, 86 Graph submissions, 129 host tasks per bind on both arms):
graph_upload loses the ~40 minor faults per bind it took allocating those
staging vectors (median 38 and 43 -> 1 and 1) and its minimum drops 0.225 ->
0.141 ms and 0.282 -> 0.106 ms. host_orch does not move -- the two
repetitions disagree in sign on both its fault count and its duration -- and
the investigation entry records why: a freed 126 KB block is reused from the
heap without re-faulting, so an image vector was never part of that segment's
fault tail. The control-plane total is not resolvable from this A/B either
(+0.355 ms then -0.166 ms), which is what a ~0.1 ms change looks like on a
box whose load average sat between 40 and 66.

test_hbg_graph_definition_arena covers what the scene tests cannot see, since
they pass either way: objects land at aligned, disjoint offsets that the
claimed prefix exactly covers, and a run with no arena or with one too small
still publishes a valid image.
@coderabbitai

coderabbitai Bot commented Aug 24, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 5 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 49339433-3121-4578-a3b9-35ff44caa192

📥 Commits

Reviewing files that changed from the base of the PR and between 66ba5c4 and 311fcd7.

📒 Files selected for processing (23)
  • docs/dfx/hbg-bind-phases.md
  • docs/investigations/2026-08-hbg-graph-definition-single-upload.md
  • docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md
  • docs/investigations/README.md
  • src/a2a3/platform/sim/host/device_runner.cpp
  • src/a2a3/runtime/host_build_graph/host/runtime_maker.cpp
  • src/a2a3/runtime/host_build_graph/runtime/orchestrator_core/orchestrator.cpp
  • src/a5/platform/sim/host/device_runner.cpp
  • src/a5/runtime/host_build_graph/host/runtime_maker.cpp
  • src/a5/runtime/host_build_graph/runtime/orchestrator_core/orchestrator.cpp
  • src/common/host_build_graph/graph_execution.h
  • src/common/host_build_graph/graph_host_state.h
  • src/common/platform/include/common/host_api.h
  • src/common/platform/onboard/host/c_api_shared.cpp
  • src/common/platform/onboard/host/device_runner_base.cpp
  • src/common/platform/onboard/host/device_runner_base.h
  • src/common/platform/sim/host/c_api_shared.cpp
  • src/common/platform/sim/host/device_runner_base.cpp
  • src/common/platform/sim/host/device_runner_base.h
  • tests/ut/cpp/CMakeLists.txt
  • tests/ut/cpp/common/test_hbg_graph_definition_arena.cpp
  • tests/ut/cpp/common/test_hbg_graph_submit_failure.cpp
  • tests/ut/cpp/common/test_hbg_slot_claim.cpp

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ChaoWao
ChaoWao merged commit a7510ca into hw-native-sys:main Aug 24, 2026
20 checks passed
@ChaoWao
ChaoWao deleted the refactor/hbg-definitions-in-one-retained-block branch August 24, 2026 11:39
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the
heap and freed it on the way out. The mirror is dimensioned for the run's
configured task count, which is 82 MB at the dsv4 case's 16384 — far above
every glibc reuse threshold, so each bind was one mmap and one guaranteed
munmap of that size. A bind's own faults then cost 14-33 us each instead of
~1.7 us, because the unmap holds mmap_lock for write and excludes every
faulting thread in the address space
(docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md).

The mirror is now the platform runner's, one buffer per pipeline slot,
grown to the high-water mark and mapped for the life of the process — the
same shape hw-native-sys#1988 gave the Graph Definition staging block, reached through a
new HostApi::acquire_sm_mirror.

The buffer is an uninitialized owning block rather than a container: the
layout is init-on-write, so a value-initializing resize would fault in
every page of a capacity whose live prefix is a fraction of it, and would
copy bytes that mean nothing between binds. First touch still commits the
mirror, so a run pays only for the bytes it writes.

Reuse needs no new invariant. A bind already cleared only the fixed-size
header and relied on init-on-write for the per-slot segments, which is why
compact_live_image can ship each segment's used prefix: prepare_task writes
the slot state and clears the completion flag as it claims a slot, and
TaskPayload::init is the single payload-init point. The one shipped range no
device-side reader reaches is the alignment padding the fanin cursor rounds
past, which lies outside every payload's fanin_count.

The buffer is released at Worker finalization on both the ordinary and the
force-reset path, since it is host memory that no device failure
invalidates.

Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to
its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and
the cold-bind fault count matches the per-bind allocation it replaces.
host_orch's warm-bind fault count and the control-plane duration resolve in
neither direction — the mirror is ~6 THP faults of a ~1200-fault bind.

This closes the investigation entry's "Where a fix would go" list, so the
entry is amended to say so: items 1-3 shipped as hw-native-sys#1981 and item 4 as hw-native-sys#1988
plus this change, and none of them reached the ~1100 faults the submitting
thread takes per bind. hw-native-sys#1981 already reported those unmoved when the
recording's allocations went away; this is the second measurement of the
same thing. Attributing them needs mincore() on a buffer's pages before the
write, not another guess at which allocation it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the
heap and freed it on the way out. The mirror is dimensioned for the run's
configured task count, which is 82 MB at the dsv4 case's 16384 — far above
every glibc reuse threshold, so each bind was one mmap and one guaranteed
munmap of that size. A bind's own faults then cost 14-33 us each instead of
~1.7 us, because the unmap holds mmap_lock for write and excludes every
faulting thread in the address space
(docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md).

The mirror is now the platform runner's, one buffer per pipeline slot,
grown to the high-water mark and held across binds until Worker
finalization — the same shape hw-native-sys#1988 gave the Graph Definition staging
block, reached through a new HostApi::acquire_sm_mirror.

The buffer is an uninitialized owning block rather than a container: the
layout is init-on-write, so a value-initializing resize would fault in
every page of a capacity whose live prefix is a fraction of it, and would
copy bytes that mean nothing between binds. First touch still commits the
mirror, so a run pays only for the bytes it writes.

Reuse needs no new invariant. A bind already cleared only the fixed-size
header and relied on init-on-write for the per-slot segments, which is why
compact_live_image can ship each segment's used prefix: prepare_task writes
the slot state and clears the completion flag as it claims a slot, and
TaskPayload::init is the single payload-init point. The one shipped range no
device-side reader reaches is the alignment padding the fanin cursor rounds
past, which lies outside every payload's fanin_count.

The buffer is released at Worker finalization on both the ordinary and the
force-reset path, since it is host memory that no device failure
invalidates.

Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to
its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and
the cold-bind fault count matches the per-bind allocation it replaces.
host_orch's warm-bind fault count and the control-plane duration resolve in
neither direction — the mirror is ~6 THP faults of a ~1200-fault bind.

This closes the investigation entry's "Where a fix would go" list, so the
entry is amended to say so: items 1-3 shipped as hw-native-sys#1981 and item 4 as hw-native-sys#1988
plus this change, and none of them reached the ~1100 faults the submitting
thread takes per bind. hw-native-sys#1981 already reported those unmoved when the
recording's allocations went away; this is the second measurement of the
same thing. Attributing them needs mincore() on a buffer's pages before the
write, not another guess at which allocation it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit that referenced this pull request Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the
heap and freed it on the way out. The mirror is dimensioned for the run's
configured task count, which is 82 MB at the dsv4 case's 16384 — far above
every glibc reuse threshold, so each bind was one mmap and one guaranteed
munmap of that size. A bind's own faults then cost 14-33 us each instead of
~1.7 us, because the unmap holds mmap_lock for write and excludes every
faulting thread in the address space
(docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md).

The mirror is now the platform runner's, one buffer per pipeline slot,
grown to the high-water mark and held across binds until Worker
finalization — the same shape #1988 gave the Graph Definition staging
block, reached through a new HostApi::acquire_sm_mirror.

The buffer is an uninitialized owning block rather than a container: the
layout is init-on-write, so a value-initializing resize would fault in
every page of a capacity whose live prefix is a fraction of it, and would
copy bytes that mean nothing between binds. First touch still commits the
mirror, so a run pays only for the bytes it writes.

Reuse needs no new invariant. A bind already cleared only the fixed-size
header and relied on init-on-write for the per-slot segments, which is why
compact_live_image can ship each segment's used prefix: prepare_task writes
the slot state and clears the completion flag as it claims a slot, and
TaskPayload::init is the single payload-init point. The one shipped range no
device-side reader reaches is the alignment padding the fanin cursor rounds
past, which lies outside every payload's fanin_count.

The buffer is released at Worker finalization on both the ordinary and the
force-reset path, since it is host memory that no device failure
invalidates.

Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to
its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and
the cold-bind fault count matches the per-bind allocation it replaces.
host_orch's warm-bind fault count and the control-plane duration resolve in
neither direction — the mirror is ~6 THP faults of a ~1200-fault bind.

This closes the investigation entry's "Where a fix would go" list, so the
entry is amended to say so: items 1-3 shipped as #1981 and item 4 as #1988
plus this change, and none of them reached the ~1100 faults the submitting
thread takes per bind. #1981 already reported those unmoved when the
recording's allocations went away; this is the second measurement of the
same thing. Attributing them needs mincore() on a buffer's pages before the
write, not another guess at which allocation it is.
ChaoWao added a commit that referenced this pull request Aug 27, 2026
…ate (#2036)

Every per-bind fault count in the host-orch investigation is a warm-up
figure under a steady-state label, and the correction is not a smaller
number: the quantity being divided was never per-bind.

Measured at f40cacf, with #1988, #2013, #2015, #2019 and #2022 all in,
host_orch's minflt per bind in arrival order over three runs at different
round counts:

  --rounds 2 (4 binds)    992, 949 (cold), then 165, 130
  --rounds 5 (10 binds)   989, 983 (cold), then 114, 173, 54, 3, 13, 13, 11, 8
  --rounds 8 (16 binds)   931, 837 (cold), then 112, 216, 164, 10, 2, 1, 1,
                          8, 57, 10, 8, 11, 0, 0

`args` follows the same curve (699310 cold, then 0-2) and so does
graph_upload (246 cold, 0 after). The tail decays over roughly six binds
and then reaches zero, on three independent runs. So nothing "survived"
the four changes the entry says it did, and the mechanism left
undetermined for three rounds turned out not to need determining.

The entry's counts came from three-to-six-round runs divided by the bind
count, so each one averaged two cold binds and three or four
still-decaying ones. The reusable half of that mistake goes to the
measurement guide, because the guide's own framing invited it: dropping
one cold bind per rank is what the parser does, and the decay behind that
bind is what it does not. Two traps added -- reading a warm bind as a
steady-state one, and dividing a total by the bind count -- and the
Reference-numbers preamble no longer calls its post-cold binds
steady-state, since a four-round session spends most of them inside the
decay.

Also records where the control plane stands at that commit, with the
caveat the same trap implies: 0.529 ms min / 0.570 median over 8 warm
binds, of which host_orch 0.341/0.364, graph_upload 0.129/0.133,
arena_h2d 0.056/0.058 -- and two of those 8 are still decaying, so it
needs --rounds 12 to be clean.

What survives from the amendment above it: the tunable arm and #2015 do
point in opposite directions, so inside the warm-up window the count is
not a lever and the price is.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 10, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty
tensors exactly as TRB's does, so the packing predicate and the slicing
predicate stay word for word identical: a tensor counted by one but not
the other shifts every later slice off the offsets the buffer was sized
from.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, and a failed grow leaving the
slot cleared rather than naming the buffer it just freed.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 10, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty
tensors exactly as TRB's does, so the packing predicate and the slicing
predicate stay word for word identical: a tensor counted by one but not
the other shifts every later slice off the offsets the buffer was sized
from.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 11, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty
tensors exactly as TRB's does, so the packing predicate and the slicing
predicate stay word for word identical: a tensor counted by one but not
the other shifts every later slice off the offsets the buffer was sized
from.

RetainedTempBump aligns the base it hands out rather than assuming the
backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard
device_malloc happens to return them, but the sim backend is std::malloc,
which guarantees only max_align_t, so aligning slice offsets alone would
have left every sim slice as misaligned as its base. begin() over-allocates
by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer
in the slot because that is what device_free must receive.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed. Both that fake and the TRB
one allocate with plain std::malloc, so a base the bump failed to align
would show up rather than being hidden by an over-aligning test double.

test_hbg_bind_ledger.cpp is the barrier for the ledger clear: it binds one
Runtime twice with no validate between, as a finalize whose
attach_current_thread failed leaves it, and asserts the second bind's
ledger holds only its own lease and that validate leaves the first run's
host buffer untouched. It drives the real bind because bind takes the
resolved host-orch entry points as a parameter, so the test supplies its
own two function pointers instead of an orchestration .so.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 14, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer.

The packing that sizes the buffer is utils/staged_tensor_plan.h, one copy
for both runtimes and both architectures. It is the half of a pair that
must not drift — a tensor it counts but the staging loop does not slice,
or the reverse, shifts every later slice off the offsets the size was
computed from, and nothing detects that until a kernel reads the wrong
bytes. A predicate whose failure is silent is the wrong thing to keep four
copies of. It is a header of its own rather than part of
retained_temp_bump.h so that one keeps needing only <cstddef>, and its
unit test keeps compiling without the task_interface include path.

hbg's staging loop now skips empty tensors as TRB's does, and that is a
change in what hbg accepts, not only a predicate alignment. A zero-byte
non-child tensor that was not a pure OUT used to fail the bind: the loop
handed it to HostTensorAccessor::add, which rejects an empty region, and
the bind reported "no host view for tensor N". It is now passed through
with a null address, which is what TRB has always done. Both hbg
RUNTIME_LOGIC.md files say so, and a unit test pins it.

RetainedTempBump aligns the base it hands out rather than assuming the
backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard
device_malloc happens to return them, but the sim backend is std::malloc,
which guarantees only max_align_t, so aligning slice offsets alone would
have left every sim slice as misaligned as its base. begin() over-allocates
by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer
in the slot because that is what device_free must receive.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed. Both that fake and the TRB
one allocate with plain std::malloc, so a base the bump failed to align
would show up rather than being hidden by an over-aligning test double.

hbg's validate no longer logs "Freed %d device allocations" at INFO: the
release is now the shared helper's LOG_DEBUG tally, so that line is absent
at the default verbosity.

test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the
empty-tensor pass-through: it binds one
Runtime twice with no validate between, as a finalize whose
attach_current_thread failed leaves it, and asserts the second bind's
ledger holds only its own lease and that validate leaves the first run's
host buffer untouched. It drives the real bind because bind takes the
resolved host-orch entry points as a parameter, so the test supplies its
own two function pointers instead of an orchestration .so.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 14, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer.

The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for
both runtimes and both architectures. It is the half of a pair that
must not drift — a tensor it counts but the staging loop does not slice,
or the reverse, shifts every later slice off the offsets the size was
computed from, and nothing detects that until a kernel reads the wrong
bytes. A predicate whose failure is silent is the wrong thing to keep four
copies of. It is a header of its own rather than part of
retained_temp_bump.h so that one keeps needing only <cstddef>, and its
unit test keeps compiling without the task_interface include path. Its
predicate is not the set hw-native-sys#2213 named h2d=: a pure OUT tensor takes a slice
without a copy-in.

hbg's argument loop now skips empty tensors as TRB's does, and that is a
change in what hbg accepts, not only a predicate alignment. A zero-byte
non-child tensor that was not a pure OUT used to fail the bind: the loop
handed it to HostTensorAccessor::add, which rejects an empty region, and
the bind reported "no host view for tensor N". It is now passed through
with a null address, which is what TRB has always done. Both hbg
RUNTIME_LOGIC.md files say so, and a unit test pins it.

RetainedTempBump aligns the base it hands out rather than assuming the
backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard
device_malloc happens to return them, but the sim backend is std::malloc,
which guarantees only max_align_t, so aligning slice offsets alone would
have left every sim slice as misaligned as its base. begin() over-allocates
by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer
in the slot because that is what device_free must receive.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed. Both that fake and the TRB
one allocate with plain std::malloc, so a base the bump failed to align
would show up rather than being hidden by an over-aligning test double.

hbg's validate no longer logs "Freed %d device allocations" at INFO: the
release is now the shared helper's LOG_DEBUG tally, so that line is absent
at the default verbosity.

test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the
empty-tensor pass-through: it binds one
Runtime twice with no validate between, as a finalize whose
attach_current_thread failed leaves it, and asserts the second bind's
ledger holds only its own lease and that validate leaves the first run's
host buffer untouched. It drives the real bind because bind takes the
resolved host-orch entry points as a parameter, so the test supplies its
own two function pointers instead of an orchestration .so.
ChaoZheng109 added a commit to ChaoZheng109/simpler that referenced this pull request Sep 14, 2026
Fixes hw-native-sys#2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer.

The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for
both runtimes and both architectures. It is the half of a pair that
must not drift — a tensor it counts but the staging loop does not slice,
or the reverse, shifts every later slice off the offsets the size was
computed from, and nothing detects that until a kernel reads the wrong
bytes. A predicate whose failure is silent is the wrong thing to keep four
copies of. It is a header of its own rather than part of
retained_temp_bump.h so that one keeps needing only <cstddef>, and its
unit test keeps compiling without the task_interface include path. Its
predicate is not the set hw-native-sys#2213 named h2d=: a pure OUT tensor takes a slice
without a copy-in.

hbg's argument loop now skips empty tensors as TRB's does, and that is a
change in what hbg accepts, not only a predicate alignment. A zero-byte
non-child tensor that was not a pure OUT used to fail the bind: the loop
handed it to HostTensorAccessor::add, which rejects an empty region, and
the bind reported "no host view for tensor N". It is now passed through
with a null address, which is what TRB has always done. Both hbg
RUNTIME_LOGIC.md files say so, and a unit test pins it.

RetainedTempBump aligns the base it hands out rather than assuming the
backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard
device_malloc happens to return them, but the sim backend is std::malloc,
which guarantees only max_align_t, so aligning slice offsets alone left
every sim slice as misaligned as its base. That was already true of TRB
before this change and no test could see it, because a sim "device" pointer
is host memory and never faults. begin() now over-allocates by
kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in
the slot because that is what device_free must receive. TRB therefore does
change: it allocates 1023 bytes more and re-aligns, which is why three
assertions in test_trb_runtime_temp_buffer.cpp move with it.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed. Both that fake and the TRB
one allocate with plain std::malloc, so a base the bump failed to align
would show up rather than being hidden by an over-aligning test double.

hbg's validate no longer logs "Freed %d device allocations" at INFO: the
release is now the shared helper's LOG_DEBUG tally, so that line is absent
at the default verbosity.

test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the
empty-tensor pass-through: it binds one
Runtime twice with no validate between, as a finalize whose
attach_current_thread failed leaves it, and asserts the second bind's
ledger holds only its own lease and that validate leaves the first run's
host buffer untouched. It drives the real bind because bind takes the
resolved host-orch entry points as a parameter, so the test supplies its
own two function pointers instead of an orchestration .so.
poursoul pushed a commit that referenced this pull request Sep 15, 2026
#2168)

Fixes #2151

host_build_graph's bind path did a device_malloc + H2D + device_free for
every host-side tensor argument on every run. tensormap_and_ringbuffer
stopped doing that in #1198, which added a runner-scoped retained staging
buffer that grows to the high-water packed size and is bump-sliced each
run, so steady state converges to zero temporary device allocations. The
platform side of that mechanism has never been TRB-specific — HostApi has
exposed get/set_retained_temp_buffer generically since #1198, and hbg
already consumes three of DeviceRunnerBase's four per-pipeline-slot
retained storages. This makes it consume the fourth.

The mechanism now has one definition. RetainedTempBump moves out of the
two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h,
and the TensorLease ledger out of the three runtime.h files into
src/common/utils/tensor_lease.h, with its release loop in
utils/tensor_lease_release.h. The lease type is split from its release
because both runtime.h files are included by AICore and AICPU translation
units, which cannot see the host API. RetainedTempBump::begin() takes a
byte count rather than ChipStorageTaskArgs, which keeps that header free
of task_interface types and leaves each caller's packing loop next to the
staging loop it has to mirror; the packing rule, the 1024-byte slice
alignment and the grow rule are unchanged, so TRB behaves exactly as
before. Diagnostics the two helpers used to log are now emitted by the
callers, which have the tensor index and the logging backend.

hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a
TensorReleaseKind, and validate releases through release_tensor_leases
instead of device_free'ing every recorded pointer — a bump slice is not a
separate allocation and must not be freed. Slices are recorded as
BufferNoop; the buffer itself is freed once per Worker in
DeviceRunnerBase::clear_temporary_buffer.

The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for
both runtimes and both architectures. It is the half of a pair that
must not drift — a tensor it counts but the staging loop does not slice,
or the reverse, shifts every later slice off the offsets the size was
computed from, and nothing detects that until a kernel reads the wrong
bytes. A predicate whose failure is silent is the wrong thing to keep four
copies of. It is a header of its own rather than part of
retained_temp_bump.h so that one keeps needing only <cstddef>, and its
unit test keeps compiling without the task_interface include path. Its
predicate is not the set #2213 named h2d=: a pure OUT tensor takes a slice
without a copy-in.

hbg's argument loop now skips empty tensors as TRB's does, and that is a
change in what hbg accepts, not only a predicate alignment. A zero-byte
non-child tensor that was not a pure OUT used to fail the bind: the loop
handed it to HostTensorAccessor::add, which rejects an empty region, and
the bind reported "no host view for tensor N". It is now passed through
with a null address, which is what TRB has always done. Both hbg
RUNTIME_LOGIC.md files say so, and a unit test pins it.

RetainedTempBump aligns the base it hands out rather than assuming the
backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard
device_malloc happens to return them, but the sim backend is std::malloc,
which guarantees only max_align_t, so aligning slice offsets alone left
every sim slice as misaligned as its base. That was already true of TRB
before this change and no test could see it, because a sim "device" pointer
is host memory and never faults. begin() now over-allocates by
kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in
the slot because that is what device_free must receive. TRB therefore does
change: it allocates 1023 bytes more and re-aligns, which is why three
assertions in test_trb_runtime_temp_buffer.cpp move with it.

A slice is bounded as `bytes > capacity_ - aligned` rather than as
`aligned + bytes > capacity_`. A caller's byte count comes from
ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap
and compare small — the slice would then be handed out as a pointer past
the buffer for copy_to_device to write through. Every overflow upstream,
in a size, in a packed total, or in align_up itself, now lands as an
ordinary slice miss instead. align_up stays unchecked: it is the same
failure one step earlier, and checking it would thread an error path
through both packing loops.

bind now clears the ledger on entry. A run whose validate never runs
otherwise leaves its entries for the next bind — on the onboard native
path a failed prepare still reaches validate_runtime_impl through
cleanup_failed_prepare, but a finalize whose attach_current_thread fails
skips validation entirely. Stale entries were a leak before; with a reused
buffer they name offsets the next bind re-slices, so validate would copy
that run's bytes back to the earlier run's host pointer.

The buffer is grown inside the args bind-phase span rather than before it,
so the one device allocation the phase can make is attributed to it and
"steady state allocates nothing here" is a measurement rather than a
definition.

Correctness rests on the retained slot being per pipeline_slot: a native
run holds its slot from bind through validate, and a concurrent
reservation is admitted only on a distinct slot, so no other run can
re-slice a buffer whose slices are live. That is the same property hbg's
Graph Definition blocks (#1988) and SM mirror (#2013) already rely on.
The H2D of a staged tensor still precedes its registration with the run's
host accessor, so a reused slice cannot expose the previous run's bytes to
orchestration. A pure OUT tensor is staged but never copied in and never
zero-filled, so on a run that reuses the buffer unchanged and packs to the
same offsets it reads the previous round's own bytes; a first or grown
allocation carries uninitialized allocator residue as before, since
begin() neither preserves nor initializes a buffer it replaces. Both
runtimes' RUNTIME_LOGIC.md now say so.

test_depth_two_slots_own_separate_resources asserted that neither hbg
pipeline slot held a retained buffer. It now asserts both hold one and
that the two are distinct, which is the regression barrier for the change.
test_retained_temp_bump.cpp covers the mechanism itself against a fake
HostApi: first allocation, reuse without allocation, grow, slice
alignment and disjointness, slice miss, an oversized request that would
have wrapped past the buffer, and a failed grow leaving the slot cleared
rather than naming the buffer it just freed. Both that fake and the TRB
one allocate with plain std::malloc, so a base the bump failed to align
would show up rather than being hidden by an over-aligning test double.

hbg's validate no longer logs "Freed %d device allocations" at INFO: the
release is now the shared helper's LOG_DEBUG tally, so that line is absent
at the default verbosity.

test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the
empty-tensor pass-through: it binds one
Runtime twice with no validate between, as a finalize whose
attach_current_thread failed leaves it, and asserts the second bind's
ledger holds only its own lease and that validate leaves the first run's
host buffer untouched. It drives the real bind because bind takes the
resolved host-orch entry points as a parameter, so the test supplies its
own two function pointers instead of an orchestration .so.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant