Conversation
Each distinct Graph Definition was uploaded as its own device object: for
every one, bind acquired a retained device block, allocated a host staging
vector, assembled [GraphDefinitionHeader][Definition image] into it, issued
its own rtMemcpy, and freed the vector -- and the image it copied from was
itself a vector the recording had allocated and would return. A dsv4 bind
has eight Definitions, so that is eight allocate-copy-free rounds for
1.0 MB, on a path whose cost is per-call rather than per-byte.
The objects now share one device block and one host staging block, both
retained per pipeline slot and grown to a high-water mark, and the
recordings build their images directly into that staging:
- bind reads the retained staging before orchestration and hands it to the
Graph host state as a GraphDefinitionArena. The capacity on offer is
whatever the previous bind left, because a run's total is not known until
every recording has ended.
- graph_build_definition splits at the point the size becomes known.
graph_layout_definition settles the counts, section offsets and
total_bytes without writing anything; the recording thread claims an
object of that size; graph_fill_definition writes the image at the
claimed offset.
- the upload writes the object headers into the room the arena left in
front of each image and ships the used prefix with one copy_to_device.
A steady-state bind therefore copies no image at all, and acquires no
host memory for its Definitions.
Three properties make the arena safe to write into from several recording
threads while the submitting thread is still submitting shells:
- the claim is a compare-exchange that fails rather than advancing the
cursor past the capacity, so a Definition that does not fit costs the run
its own slot and no one else's. It is built in a buffer of its own
instead and copied in by the upload, which then grows the block -- so
outgrowing the retained capacity costs a copy, not the run, and the next
bind is sized for it. graph_upload's new `spilled=` counts those, and is
0 on every bind after the first.
- the arena never moves during orchestration, since a recorder holds
addresses inside it. Only the upload may grow it, which preserves the
content and is why a published record names an offset rather than an
address; by then nothing is recording.
- an object is aligned and padded to GRAPH_DEFINITION_OBJECT_ALIGN, a named
constant beside the header whose size is a multiple of it, so every image
base carries the alignment its section offsets assume. The fill asserts
that rather than trusting the arena's provider.
The image is still zero-filled before it is written, and now for a second
reason: the alignment slack between sections is inside the hashed range, so a
slot a previous bind wrote would give two structurally identical Definitions
unequal content hashes.
HostApi grows get_graph_definition_staging and replaces
acquire_graph_definition_buffer with acquire_graph_definition_block, which
takes no Definition key and returns both blocks, since the layout inside them
belongs to the caller. The device block is still zeroed when it is
(re)allocated, so a region no upload has covered reads as zero and fails the
magic gate instead of being read as a stale object.
Measured on dsv4, interleaved base/measure/base/measure at six rounds each
(8 Definitions, 86 Graph submissions, 129 host tasks per bind on both arms):
graph_upload loses the ~40 minor faults per bind it took allocating those
staging vectors (median 38 and 43 -> 1 and 1) and its minimum drops 0.225 ->
0.141 ms and 0.282 -> 0.106 ms. host_orch does not move -- the two
repetitions disagree in sign on both its fault count and its duration -- and
the investigation entry records why: a freed 126 KB block is reused from the
heap without re-faulting, so an image vector was never part of that segment's
fault tail. The control-plane total is not resolvable from this A/B either
(+0.355 ms then -0.166 ms), which is what a ~0.1 ms change looks like on a
box whose load average sat between 40 and 66.
test_hbg_graph_definition_arena covers what the scene tests cannot see, since
they pass either way: objects land at aligned, disjoint offsets that the
claimed prefix exactly covers, and a run with no arena or with one too small
still publishes a valid image.
|
Warning Review limit reachedNext included review available in 5 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (23)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
4 of 5 tasks
6 tasks done
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the heap and freed it on the way out. The mirror is dimensioned for the run's configured task count, which is 82 MB at the dsv4 case's 16384 — far above every glibc reuse threshold, so each bind was one mmap and one guaranteed munmap of that size. A bind's own faults then cost 14-33 us each instead of ~1.7 us, because the unmap holds mmap_lock for write and excludes every faulting thread in the address space (docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md). The mirror is now the platform runner's, one buffer per pipeline slot, grown to the high-water mark and mapped for the life of the process — the same shape hw-native-sys#1988 gave the Graph Definition staging block, reached through a new HostApi::acquire_sm_mirror. The buffer is an uninitialized owning block rather than a container: the layout is init-on-write, so a value-initializing resize would fault in every page of a capacity whose live prefix is a fraction of it, and would copy bytes that mean nothing between binds. First touch still commits the mirror, so a run pays only for the bytes it writes. Reuse needs no new invariant. A bind already cleared only the fixed-size header and relied on init-on-write for the per-slot segments, which is why compact_live_image can ship each segment's used prefix: prepare_task writes the slot state and clears the completion flag as it claims a slot, and TaskPayload::init is the single payload-init point. The one shipped range no device-side reader reaches is the alignment padding the fanin cursor rounds past, which lies outside every payload's fanin_count. The buffer is released at Worker finalization on both the ordinary and the force-reset path, since it is host memory that no device failure invalidates. Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and the cold-bind fault count matches the per-bind allocation it replaces. host_orch's warm-bind fault count and the control-plane duration resolve in neither direction — the mirror is ~6 THP faults of a ~1200-fault bind. This closes the investigation entry's "Where a fix would go" list, so the entry is amended to say so: items 1-3 shipped as hw-native-sys#1981 and item 4 as hw-native-sys#1988 plus this change, and none of them reached the ~1100 faults the submitting thread takes per bind. hw-native-sys#1981 already reported those unmoved when the recording's allocations went away; this is the second measurement of the same thing. Attributing them needs mincore() on a buffer's pages before the write, not another guess at which allocation it is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the heap and freed it on the way out. The mirror is dimensioned for the run's configured task count, which is 82 MB at the dsv4 case's 16384 — far above every glibc reuse threshold, so each bind was one mmap and one guaranteed munmap of that size. A bind's own faults then cost 14-33 us each instead of ~1.7 us, because the unmap holds mmap_lock for write and excludes every faulting thread in the address space (docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md). The mirror is now the platform runner's, one buffer per pipeline slot, grown to the high-water mark and held across binds until Worker finalization — the same shape hw-native-sys#1988 gave the Graph Definition staging block, reached through a new HostApi::acquire_sm_mirror. The buffer is an uninitialized owning block rather than a container: the layout is init-on-write, so a value-initializing resize would fault in every page of a capacity whose live prefix is a fraction of it, and would copy bytes that mean nothing between binds. First touch still commits the mirror, so a run pays only for the bytes it writes. Reuse needs no new invariant. A bind already cleared only the fixed-size header and relied on init-on-write for the per-slot segments, which is why compact_live_image can ship each segment's used prefix: prepare_task writes the slot state and clears the completion flag as it claims a slot, and TaskPayload::init is the single payload-init point. The one shipped range no device-side reader reaches is the alignment padding the fanin cursor rounds past, which lies outside every payload's fanin_count. The buffer is released at Worker finalization on both the ordinary and the force-reset path, since it is host memory that no device failure invalidates. Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and the cold-bind fault count matches the per-bind allocation it replaces. host_orch's warm-bind fault count and the control-plane duration resolve in neither direction — the mirror is ~6 THP faults of a ~1200-fault bind. This closes the investigation entry's "Where a fix would go" list, so the entry is amended to say so: items 1-3 shipped as hw-native-sys#1981 and item 4 as hw-native-sys#1988 plus this change, and none of them reached the ~1100 faults the submitting thread takes per bind. hw-native-sys#1981 already reported those unmoved when the recording's allocations went away; this is the second measurement of the same thing. Attributing them needs mincore() on a buffer's pages before the write, not another guess at which allocation it is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao
added a commit
that referenced
this pull request
Aug 25, 2026
Every bind allocated the host mirror of the runtime shared memory from the heap and freed it on the way out. The mirror is dimensioned for the run's configured task count, which is 82 MB at the dsv4 case's 16384 — far above every glibc reuse threshold, so each bind was one mmap and one guaranteed munmap of that size. A bind's own faults then cost 14-33 us each instead of ~1.7 us, because the unmap holds mmap_lock for write and excludes every faulting thread in the address space (docs/investigations/2026-08-host-orch-phase-tail-is-page-faults.md). The mirror is now the platform runner's, one buffer per pipeline slot, grown to the high-water mark and held across binds until Worker finalization — the same shape #1988 gave the Graph Definition staging block, reached through a new HostApi::acquire_sm_mirror. The buffer is an uninitialized owning block rather than a container: the layout is init-on-write, so a value-initializing resize would fault in every page of a capacity whose live prefix is a fraction of it, and would copy bytes that mean nothing between binds. First touch still commits the mirror, so a run pays only for the bytes it writes. Reuse needs no new invariant. A bind already cleared only the fixed-size header and relied on init-on-write for the per-slot segments, which is why compact_live_image can ship each segment's used prefix: prepare_task writes the slot state and clears the completion flag as it claims a slot, and TaskPayload::init is the single payload-init point. The one shipped range no device-side reader reaches is the alignment padding the fanin cursor rounds past, which lies outside every payload's fanin_count. The buffer is released at Worker finalization on both the ordinary and the force-reset path, since it is host memory that no device failure invalidates. Measured on dsv4, interleaved twice: mallinfo2's hblkhd stops returning to its pre-bind value on 6 of 6 binds, which is the map/unmap going away, and the cold-bind fault count matches the per-bind allocation it replaces. host_orch's warm-bind fault count and the control-plane duration resolve in neither direction — the mirror is ~6 THP faults of a ~1200-fault bind. This closes the investigation entry's "Where a fix would go" list, so the entry is amended to say so: items 1-3 shipped as #1981 and item 4 as #1988 plus this change, and none of them reached the ~1100 faults the submitting thread takes per bind. #1981 already reported those unmoved when the recording's allocations went away; this is the second measurement of the same thing. Attributing them needs mincore() on a buffer's pages before the write, not another guess at which allocation it is.
This was referenced Aug 26, 2026
ChaoWao
added a commit
that referenced
this pull request
Aug 27, 2026
…ate (#2036) Every per-bind fault count in the host-orch investigation is a warm-up figure under a steady-state label, and the correction is not a smaller number: the quantity being divided was never per-bind. Measured at f40cacf, with #1988, #2013, #2015, #2019 and #2022 all in, host_orch's minflt per bind in arrival order over three runs at different round counts: --rounds 2 (4 binds) 992, 949 (cold), then 165, 130 --rounds 5 (10 binds) 989, 983 (cold), then 114, 173, 54, 3, 13, 13, 11, 8 --rounds 8 (16 binds) 931, 837 (cold), then 112, 216, 164, 10, 2, 1, 1, 8, 57, 10, 8, 11, 0, 0 `args` follows the same curve (699310 cold, then 0-2) and so does graph_upload (246 cold, 0 after). The tail decays over roughly six binds and then reaches zero, on three independent runs. So nothing "survived" the four changes the entry says it did, and the mechanism left undetermined for three rounds turned out not to need determining. The entry's counts came from three-to-six-round runs divided by the bind count, so each one averaged two cold binds and three or four still-decaying ones. The reusable half of that mistake goes to the measurement guide, because the guide's own framing invited it: dropping one cold bind per rank is what the parser does, and the decay behind that bind is what it does not. Two traps added -- reading a warm bind as a steady-state one, and dividing a total by the bind count -- and the Reference-numbers preamble no longer calls its post-cold binds steady-state, since a four-round session spends most of them inside the decay. Also records where the control plane stands at that commit, with the caveat the same trap implies: 0.529 ms min / 0.570 median over 8 warm binds, of which host_orch 0.341/0.364, graph_upload 0.129/0.133, arena_h2d 0.056/0.058 -- and two of those 8 are still decaying, so it needs --rounds 12 to be clean. What survives from the amendment above it: the tunable arm and #2015 do point in opposite directions, so inside the warm-up window the count is not a lever and the price is.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 10, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty tensors exactly as TRB's does, so the packing predicate and the slicing predicate stay word for word identical: a tensor counted by one but not the other shifts every later slice off the offsets the buffer was sized from. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, and a failed grow leaving the slot cleared rather than naming the buffer it just freed.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 10, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty tensors exactly as TRB's does, so the packing predicate and the slicing predicate stay word for word identical: a tensor counted by one but not the other shifts every later slice off the offsets the buffer was sized from. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 11, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. hbg's staging loop skips empty tensors exactly as TRB's does, so the packing predicate and the slicing predicate stay word for word identical: a tensor counted by one but not the other shifts every later slice off the offsets the buffer was sized from. RetainedTempBump aligns the base it hands out rather than assuming the backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard device_malloc happens to return them, but the sim backend is std::malloc, which guarantees only max_align_t, so aligning slice offsets alone would have left every sim slice as misaligned as its base. begin() over-allocates by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in the slot because that is what device_free must receive. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed. Both that fake and the TRB one allocate with plain std::malloc, so a base the bump failed to align would show up rather than being hidden by an over-aligning test double. test_hbg_bind_ledger.cpp is the barrier for the ledger clear: it binds one Runtime twice with no validate between, as a finalize whose attach_current_thread failed leaves it, and asserts the second bind's ledger holds only its own lease and that validate leaves the first run's host buffer untouched. It drives the real bind because bind takes the resolved host-orch entry points as a parameter, so the test supplies its own two function pointers instead of an orchestration .so.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 14, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. The packing that sizes the buffer is utils/staged_tensor_plan.h, one copy for both runtimes and both architectures. It is the half of a pair that must not drift — a tensor it counts but the staging loop does not slice, or the reverse, shifts every later slice off the offsets the size was computed from, and nothing detects that until a kernel reads the wrong bytes. A predicate whose failure is silent is the wrong thing to keep four copies of. It is a header of its own rather than part of retained_temp_bump.h so that one keeps needing only <cstddef>, and its unit test keeps compiling without the task_interface include path. hbg's staging loop now skips empty tensors as TRB's does, and that is a change in what hbg accepts, not only a predicate alignment. A zero-byte non-child tensor that was not a pure OUT used to fail the bind: the loop handed it to HostTensorAccessor::add, which rejects an empty region, and the bind reported "no host view for tensor N". It is now passed through with a null address, which is what TRB has always done. Both hbg RUNTIME_LOGIC.md files say so, and a unit test pins it. RetainedTempBump aligns the base it hands out rather than assuming the backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard device_malloc happens to return them, but the sim backend is std::malloc, which guarantees only max_align_t, so aligning slice offsets alone would have left every sim slice as misaligned as its base. begin() over-allocates by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in the slot because that is what device_free must receive. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed. Both that fake and the TRB one allocate with plain std::malloc, so a base the bump failed to align would show up rather than being hidden by an over-aligning test double. hbg's validate no longer logs "Freed %d device allocations" at INFO: the release is now the shared helper's LOG_DEBUG tally, so that line is absent at the default verbosity. test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the empty-tensor pass-through: it binds one Runtime twice with no validate between, as a finalize whose attach_current_thread failed leaves it, and asserts the second bind's ledger holds only its own lease and that validate leaves the first run's host buffer untouched. It drives the real bind because bind takes the resolved host-orch entry points as a parameter, so the test supplies its own two function pointers instead of an orchestration .so.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 14, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for both runtimes and both architectures. It is the half of a pair that must not drift — a tensor it counts but the staging loop does not slice, or the reverse, shifts every later slice off the offsets the size was computed from, and nothing detects that until a kernel reads the wrong bytes. A predicate whose failure is silent is the wrong thing to keep four copies of. It is a header of its own rather than part of retained_temp_bump.h so that one keeps needing only <cstddef>, and its unit test keeps compiling without the task_interface include path. Its predicate is not the set hw-native-sys#2213 named h2d=: a pure OUT tensor takes a slice without a copy-in. hbg's argument loop now skips empty tensors as TRB's does, and that is a change in what hbg accepts, not only a predicate alignment. A zero-byte non-child tensor that was not a pure OUT used to fail the bind: the loop handed it to HostTensorAccessor::add, which rejects an empty region, and the bind reported "no host view for tensor N". It is now passed through with a null address, which is what TRB has always done. Both hbg RUNTIME_LOGIC.md files say so, and a unit test pins it. RetainedTempBump aligns the base it hands out rather than assuming the backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard device_malloc happens to return them, but the sim backend is std::malloc, which guarantees only max_align_t, so aligning slice offsets alone would have left every sim slice as misaligned as its base. begin() over-allocates by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in the slot because that is what device_free must receive. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed. Both that fake and the TRB one allocate with plain std::malloc, so a base the bump failed to align would show up rather than being hidden by an over-aligning test double. hbg's validate no longer logs "Freed %d device allocations" at INFO: the release is now the shared helper's LOG_DEBUG tally, so that line is absent at the default verbosity. test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the empty-tensor pass-through: it binds one Runtime twice with no validate between, as a finalize whose attach_current_thread failed leaves it, and asserts the second bind's ledger holds only its own lease and that validate leaves the first run's host buffer untouched. It drives the real bind because bind takes the resolved host-orch entry points as a parameter, so the test supplies its own two function pointers instead of an orchestration .so.
ChaoZheng109
added a commit
to ChaoZheng109/simpler
that referenced
this pull request
Sep 14, 2026
Fixes hw-native-sys#2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in hw-native-sys#1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since hw-native-sys#1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for both runtimes and both architectures. It is the half of a pair that must not drift — a tensor it counts but the staging loop does not slice, or the reverse, shifts every later slice off the offsets the size was computed from, and nothing detects that until a kernel reads the wrong bytes. A predicate whose failure is silent is the wrong thing to keep four copies of. It is a header of its own rather than part of retained_temp_bump.h so that one keeps needing only <cstddef>, and its unit test keeps compiling without the task_interface include path. Its predicate is not the set hw-native-sys#2213 named h2d=: a pure OUT tensor takes a slice without a copy-in. hbg's argument loop now skips empty tensors as TRB's does, and that is a change in what hbg accepts, not only a predicate alignment. A zero-byte non-child tensor that was not a pure OUT used to fail the bind: the loop handed it to HostTensorAccessor::add, which rejects an empty region, and the bind reported "no host view for tensor N". It is now passed through with a null address, which is what TRB has always done. Both hbg RUNTIME_LOGIC.md files say so, and a unit test pins it. RetainedTempBump aligns the base it hands out rather than assuming the backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard device_malloc happens to return them, but the sim backend is std::malloc, which guarantees only max_align_t, so aligning slice offsets alone left every sim slice as misaligned as its base. That was already true of TRB before this change and no test could see it, because a sim "device" pointer is host memory and never faults. begin() now over-allocates by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in the slot because that is what device_free must receive. TRB therefore does change: it allocates 1023 bytes more and re-aligns, which is why three assertions in test_trb_runtime_temp_buffer.cpp move with it. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (hw-native-sys#1988) and SM mirror (hw-native-sys#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed. Both that fake and the TRB one allocate with plain std::malloc, so a base the bump failed to align would show up rather than being hidden by an over-aligning test double. hbg's validate no longer logs "Freed %d device allocations" at INFO: the release is now the shared helper's LOG_DEBUG tally, so that line is absent at the default verbosity. test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the empty-tensor pass-through: it binds one Runtime twice with no validate between, as a finalize whose attach_current_thread failed leaves it, and asserts the second bind's ledger holds only its own lease and that validate leaves the first run's host buffer untouched. It drives the real bind because bind takes the resolved host-orch entry points as a parameter, so the test supplies its own two function pointers instead of an orchestration .so.
poursoul
pushed a commit
that referenced
this pull request
Sep 15, 2026
#2168) Fixes #2151 host_build_graph's bind path did a device_malloc + H2D + device_free for every host-side tensor argument on every run. tensormap_and_ringbuffer stopped doing that in #1198, which added a runner-scoped retained staging buffer that grows to the high-water packed size and is bump-sliced each run, so steady state converges to zero temporary device allocations. The platform side of that mechanism has never been TRB-specific — HostApi has exposed get/set_retained_temp_buffer generically since #1198, and hbg already consumes three of DeviceRunnerBase's four per-pipeline-slot retained storages. This makes it consume the fourth. The mechanism now has one definition. RetainedTempBump moves out of the two TRB runtime_maker.cpp files into src/common/utils/retained_temp_bump.h, and the TensorLease ledger out of the three runtime.h files into src/common/utils/tensor_lease.h, with its release loop in utils/tensor_lease_release.h. The lease type is split from its release because both runtime.h files are included by AICore and AICPU translation units, which cannot see the host API. RetainedTempBump::begin() takes a byte count rather than ChipStorageTaskArgs, which keeps that header free of task_interface types and leaves each caller's packing loop next to the staging loop it has to mirror; the packing rule, the 1024-byte slice alignment and the grow rule are unchanged, so TRB behaves exactly as before. Diagnostics the two helpers used to log are now emitted by the callers, which have the tensor index and the logging backend. hbg's ledger takes TRB's shape: TensorPair becomes TensorLease and gains a TensorReleaseKind, and validate releases through release_tensor_leases instead of device_free'ing every recorded pointer — a bump slice is not a separate allocation and must not be freed. Slices are recorded as BufferNoop; the buffer itself is freed once per Worker in DeviceRunnerBase::clear_temporary_buffer. The packing that sizes the buffer is utils/temp_buffer_plan.h, one copy for both runtimes and both architectures. It is the half of a pair that must not drift — a tensor it counts but the staging loop does not slice, or the reverse, shifts every later slice off the offsets the size was computed from, and nothing detects that until a kernel reads the wrong bytes. A predicate whose failure is silent is the wrong thing to keep four copies of. It is a header of its own rather than part of retained_temp_bump.h so that one keeps needing only <cstddef>, and its unit test keeps compiling without the task_interface include path. Its predicate is not the set #2213 named h2d=: a pure OUT tensor takes a slice without a copy-in. hbg's argument loop now skips empty tensors as TRB's does, and that is a change in what hbg accepts, not only a predicate alignment. A zero-byte non-child tensor that was not a pure OUT used to fail the bind: the loop handed it to HostTensorAccessor::add, which rejects an empty region, and the bind reported "no host view for tensor N". It is now passed through with a null address, which is what TRB has always done. Both hbg RUNTIME_LOGIC.md files say so, and a unit test pins it. RetainedTempBump aligns the base it hands out rather than assuming the backend aligned it. Kernels need 1024-byte-aligned device pointers; onboard device_malloc happens to return them, but the sim backend is std::malloc, which guarantees only max_align_t, so aligning slice offsets alone left every sim slice as misaligned as its base. That was already true of TRB before this change and no test could see it, because a sim "device" pointer is host memory and never faults. begin() now over-allocates by kAlignment - 1 and aligns inside the allocation, keeping the raw pointer in the slot because that is what device_free must receive. TRB therefore does change: it allocates 1023 bytes more and re-aligns, which is why three assertions in test_trb_runtime_temp_buffer.cpp move with it. A slice is bounded as `bytes > capacity_ - aligned` rather than as `aligned + bytes > capacity_`. A caller's byte count comes from ChipTensor::nbytes(), an unchecked uint64_t product, so that sum can wrap and compare small — the slice would then be handed out as a pointer past the buffer for copy_to_device to write through. Every overflow upstream, in a size, in a packed total, or in align_up itself, now lands as an ordinary slice miss instead. align_up stays unchecked: it is the same failure one step earlier, and checking it would thread an error path through both packing loops. bind now clears the ledger on entry. A run whose validate never runs otherwise leaves its entries for the next bind — on the onboard native path a failed prepare still reaches validate_runtime_impl through cleanup_failed_prepare, but a finalize whose attach_current_thread fails skips validation entirely. Stale entries were a leak before; with a reused buffer they name offsets the next bind re-slices, so validate would copy that run's bytes back to the earlier run's host pointer. The buffer is grown inside the args bind-phase span rather than before it, so the one device allocation the phase can make is attributed to it and "steady state allocates nothing here" is a measurement rather than a definition. Correctness rests on the retained slot being per pipeline_slot: a native run holds its slot from bind through validate, and a concurrent reservation is admitted only on a distinct slot, so no other run can re-slice a buffer whose slices are live. That is the same property hbg's Graph Definition blocks (#1988) and SM mirror (#2013) already rely on. The H2D of a staged tensor still precedes its registration with the run's host accessor, so a reused slice cannot expose the previous run's bytes to orchestration. A pure OUT tensor is staged but never copied in and never zero-filled, so on a run that reuses the buffer unchanged and packs to the same offsets it reads the previous round's own bytes; a first or grown allocation carries uninitialized allocator residue as before, since begin() neither preserves nor initializes a buffer it replaces. Both runtimes' RUNTIME_LOGIC.md now say so. test_depth_two_slots_own_separate_resources asserted that neither hbg pipeline slot held a retained buffer. It now asserts both hold one and that the two are distinct, which is the regression barrier for the change. test_retained_temp_bump.cpp covers the mechanism itself against a fake HostApi: first allocation, reuse without allocation, grow, slice alignment and disjointness, slice miss, an oversized request that would have wrapped past the buffer, and a failed grow leaving the slot cleared rather than naming the buffer it just freed. Both that fake and the TRB one allocate with plain std::malloc, so a base the bump failed to align would show up rather than being hidden by an over-aligning test double. hbg's validate no longer logs "Freed %d device allocations" at INFO: the release is now the shared helper's LOG_DEBUG tally, so that line is absent at the default verbosity. test_hbg_bind_ledger.cpp is the barrier for the ledger clear and for the empty-tensor pass-through: it binds one Runtime twice with no validate between, as a finalize whose attach_current_thread failed leaves it, and asserts the second bind's ledger holds only its own lease and that validate leaves the first run's host buffer untouched. It drives the real bind because bind takes the resolved host-orch entry points as a parameter, so the test supplies its own two function pointers instead of an orchestration .so.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Each distinct Graph Definition was uploaded as its own device object: for every one, bind acquired a retained device block, allocated a host staging vector, assembled
[GraphDefinitionHeader][Definition image]into it, issued its ownrtMemcpy, and freed the vector — and the image it copied from was itself a vector the recording had allocated and would return. A dsv4 bind has eight Definitions, so that is eight allocate-copy-free rounds for 1.0 MB, on a path whose cost is per-call rather than per-byte.The objects now share one device block and one host staging block, both retained per pipeline slot and grown to a high-water mark, and the recordings build their images directly into that staging:
GraphDefinitionArena. The capacity on offer is whatever the previous bind left, because a run's total is not known until every recording has ended.graph_build_definitionsplits at the point the size becomes known.graph_layout_definitionsettles the counts, section offsets andtotal_byteswithout writing anything; the recording thread claims an object of that size;graph_fill_definitionwrites the image at the claimed offset.copy_to_device. A steady-state bind copies no image at all, and acquires no host memory for its Definitions.Three properties make the arena safe to write into from several recording threads while the submitting thread is still submitting shells:
graph_upload's newspilled=counts those, and is 0 on every bind after the first.GRAPH_DEFINITION_OBJECT_ALIGN, a named constant beside the header whose size is a multiple of it, so every image base carries the alignment its section offsets assume. The fill asserts that rather than trusting the arena's provider.The image is still zero-filled before it is written, and now for a second reason: the alignment slack between sections is inside the hashed range, so a slot a previous bind wrote would give two structurally identical Definitions unequal content hashes.
HostApigrowsget_graph_definition_stagingand replacesacquire_graph_definition_bufferwithacquire_graph_definition_block, which takes no Definition key and returns both blocks, since the layout inside them belongs to the caller. The device block is still zeroed when it is (re)allocated, so a region no upload has covered reads as zero and fails the magic gate instead of being read as a stale object.Measurement
dsv4, interleaved
base, measure, base, measureat six rounds each per.claude/skills/hbg-bind-phases; both arms ran the same command string (the logs'[stamp]lines differ only in the commit) and did the same work — 8 Definitions, 86 Graph submissions, 129 host tasks per bind. Two figures per cell, one per repetition:graph_uploadminflt, min / median / maxgraph_uploaddur min (ms)host_orchminflt, min / medianOnly
graph_uploadmoved, and it moved consistently in both repetitions: the ~40 minor faults per bind it took allocating one staging vector per Definition are gone, and so is the copy.host_orchdid not move, and the investigation entry now records why: the two repetitions disagree in sign on both its fault count and its duration, which is that file's own criterion for "not resolvable". A freed 126 KB block is reused from the heap without re-faulting, so an image vector was never part of that segment's fault tail — size against glibc's mmap and trim thresholds, not byte count, decides what shows up there. The control-plane total is not resolvable either (+0.355 ms then −0.166 ms), which is what a ~0.1 ms change looks like on a box whose load average sat between 40 and 66.Testing
examples tests/stona2a3sim(21/21 cases) anda5sim(17/17), both after the rebase onto66ba5c4a1examples tests/st -m 'not sdma' --platform a2a3 --exclude-level 4undertask-submit, 150 selected / 57 resource-phase cases, twicectest -LE requires_hardware— 119/119, including the newtest_hbg_graph_definition_arenafor a2a3 and a5pytest tests/ut -m "not requires_hardware"— 1886 passed, 7 skippedspilled=verified across rounds: 2 of 12 binds (one per rank, cold) spill; the other 10 reportspilled=0test_hbg_graph_definition_arenacovers what the scene tests cannot see, since they pass either way: objects land at aligned, disjoint offsets that the claimed prefix exactly covers, and a run with no arena or with one too small still publishes a valid image.One caveat from the hardware runs, recorded rather than swept under: the first corpus run on this branch hit a device-side AIV
ub address out of boundsinworker_async_endpoint, which cascaded to507018/SCHEDULER_TIMEOUT. It did not reproduce — that case passed 3/3 alone on66ba5c4a1, the same corpus passed 57/57 on66ba5c4a1, and the second corpus run on this branch passed 57/57 — and the case records no Graph at all, so this diff does not execute in it. Worth an eye on CI: a UB fault is not a flaky timeout.