Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
nalinaly
left a comment
There was a problem hiding this comment.
请调整 callable 驻留容量及 device 内存分配方式:8192 项规格,参考 CANN 的 2 MiB 分块按需增长。具体意见见行内评论。
| class KernelCallableCache { | ||
| public: | ||
| static constexpr size_t kDescriptorBytes = MAX_REGISTERED_CALLABLE_IDS * sizeof(KernelCallableDeviceResidency); | ||
| static constexpr size_t kByteLimit = 512ULL * 1024 * 1024; |
There was a problem hiding this comment.
建议将 kernel 模式的驻留 callable 容量规格调整为 8192,并将条目规格与 device 内存预分配解耦。
当前实现存在两个问题:第 65 个不同 callable 会在仍有可用 device 内存时被拒绝;第一个很小的 callable 也会触发约 512 MiB 的 arena 分配。仅把 64 调大、继续保留这个 arena 模型,不能满足大量 kernel 按需注册的场景。
建议本 PR 按以下规格修改:
- 最多支持 8192 个不同的驻留 callable。 需要同步处理 Host cache/resolve、runner、A2A3/A5 AICPU 注册表和 Device dispatch 的相关边界,不能只修改 cache 一处;8192 是条目容量规格,不代表预分配 8192 套 device 执行资源。
- Device 内存按实际驻留需求申请。 可参考 CANN runtime 的 2 MiB 分块池:小镜像在已有块内子分配,空间不足才追加新块;超过单块大小的镜像按实际大小加必要对齐独立分配。2 MiB 是共享分配块的粒度,不是每个 callable 至少占 2 MiB。取消固定 512 MiB 首次预留,按需增长。
- 驻留代码缓存的 device 总容量上限暂定为 2 GiB。 这是容量上限,不是首次预分配量;仍按上述方式按需申请。达到 8192 项或 2 GiB 容量上限时返回明确的容量错误,不影响已经成功驻留的条目。
- 扩容只追加,不迁移已发布的代码和 descriptor。 保持已有地址及 handle 有效;新增分配、上传和注册仍放在 capture 外的 prepare,launch/replay 不承担扩容工作。
Host 内存目前不是本条 review 的关注点,现有 Host 镜像保留方式可以不改。
参考 CANN 官方源码:BinaryLoad 按实际加载大小分配、2 MiB 单池定义、池不足时追加新池。
There was a problem hiding this comment.
已按建议调整:
- 条目上限改为 8192,统一使用
MAX_REGISTERED_CALLABLE_IDS,覆盖 Host cache/resolve、runner 和 A2A3/A5 AICPU 注册表及 dispatch 边界;不预分配 8192 套执行资源。 - 取消首次固定预留 512 MiB。完整 ChipCallable 镜像不超过 2 MiB 时,按 64 字节对齐在已有块内子分配;已有块都放不下时才追加一个 2 MiB 块。各块独立分配,不要求地址连续,单个镜像不跨块。
- 大于 2 MiB 的镜像独立申请
align_up(size, 64)空间;已有小块的剩余空间仍可供后续小镜像使用。 - 2 GiB 上限按所有已分配代码块的容量统计,包含对齐及空闲尾部,与 8192 项上限独立检查。超限返回容量错误,已有驻留条目不受影响。
- 扩容只在 capture 外的 prepare 追加分配,不搬迁已发布镜像,已有地址和 handle 保持有效。每次成功 prepare 仍注册新 callable,上层负责 handle 复用。
验证:5 个相关 C++ 测试目标、14 项缓存 ASan/UBSan 测试、296 项 Python/ABI 回归及 18 项 A3 设备生命周期测试通过;两种 runtime 的 program 模式模拟执行通过。设备测试覆盖 65 次连续注册、首次小块分配、按块增长及 close 重试;8192 项边界和容量超限不变性由 C++ 单测覆盖。中文调用链文档已同步更新。
Provision persistent kernel streams, events, runtime arguments, and callable uploads on top of the K1 ABI. Keep preparation independent of caller streams and retain partial allocations when rollback fails so explicit close can retry ownership-preserving teardown. Validate complete callable images and surface finalize failures through the public worker API.
2c87647 to
05a2471
Compare
05a2471 to
daa1fdf
Compare
e13d019 to
5f9895d
Compare
9c085aa to
08f7f92
Compare
…gnment, 8192 ids (#2241) * CI: bump pinned pto-isa to 03e45c4b (#2219) * Refactor: unify scalar reads across host_build_graph and TaskArgsTpl (#2197) host_build_graph's Arg::scalar(i) returned InheritableScalar unconditionally and exposed the raw uint64_t only through a deprecated implicit conversion, kept for callers that had not yet moved to args.scalar(i).to<T>(). TaskArgsTpl::scalar(i) (TMR and every other Arg built directly on it) returned S itself with no template parameter at all. The two spellings disagreed on what a caller writes for a static read even though both slots are the same uint64_t. scalar(i) is now a template on both sides, its parameter spelled ScalarT in both -- TaskArgsTpl's T is its tensor type. host_build_graph defaults it to InheritableScalar, so a bare scalar(i) still forwards with its origin; TaskArgsTpl defaults it to S (uint64_t), so a bare scalar(i) there is unchanged. Either side accepts an explicit type for a static value read (args.scalar<T>(i)), and TaskArgsTpl bounds that read by sizeof(S) -- the slot the value has to fit in -- rather than by a hardcoded 8. task_args.h includes data_type.h for the from_u64 that read applies, rather than reaching it through tensor.h. The deprecated operator uint64_t() is removed, which is stronger than the deprecation it replaces: with no conversion left to suppress, a value read is a compile error rather than a warning. That closes the blind spot the deprecation had, where a read instantiated inside a system header -- EXPECT_EQ(args.scalar(i), v) -- was silently exempt. InheritableScalar::to<T>() stays: it is the only way to read a handle that was passed on as a function argument, where Arg::scalar<T>(i) is unavailable because the Arg is not in hand. Every call site that triggered the deprecation warning (98 across 20 orchestration files) moves to the explicit-T spelling. All of them were already value reads, so no call site changes meaning; which of them ought to forward instead is a question about each example's semantics and is tracked separately. The tensormap_and_ringbuffer orchestrations that read a slot as some type other than the slot's own move with them, so both runtimes spell that read the same way and the non-S branch of TaskArgsTpl::scalar has in-tree instantiations -- a template body is only checked when it is instantiated, and until now every scalar<T> call site was host_build_graph, whose Arg hides the base's scalar with its own. A bare uint64_t read is left alone: that is what scalar(i) already answers, so naming the type would add nothing. Python's add_scalar took a pre-encoded uint64_t, pushing scalar_to_uint64(value) onto every caller. It now takes the value directly -- int, float, bool, or a ctypes scalar -- and encodes it natively (encode_scalar in the bindings, exposed to Python as scalar_to_uint64 for callers that still want the raw bits). scene_test.py's three add_scalar call sites drop their scalar_to_uint64 wrapping accordingly. That encoder matches C++ to_u64() bit for bit, which the previous Python implementation did not: it read a ctypes scalar through its `.value`, which ctypes has already sign-extended for a signed type, so c_int8(-1) produced 0xFFFF'FFFF'FFFF'FFFF where to_u64(int8_t{-1}) is 0xFF. A ctypes scalar is now read through the buffer protocol at its own width and zero-extended, which is what to_u64's union does. Reading raw bytes makes the scalar's byte order load-bearing, and its buffer format is where that order is stated. ctypes admits byte-order-qualified variants -- c_uint32.__ctype_be__ carries the same _type_ as c_uint32 and differs only in the format prefix -- whose bytes for the value 1 are 00 00 00 01, which a raw copy would store as 0x01000000. A reversed-order scalar has no native C++ counterpart to agree with, so the format decides admission: host byte order plus one of the integer widths, f, d or ?. The pointer and character types (P, z, Z, c, u) and long double are refused with it; c_void_p and c_char_p would otherwise encode a host pointer into a device-bound slot. A subclass inherits its base's format and so is admitted with the base, which dispatching on the type's __name__ would not do. A native Python float still narrows to IEEE-754 single precision, and a finite value out of that range now raises where a narrowing conversion would produce an infinity -- struct.pack("<f", 1e100) raised too, and storing that infinity would silently be a different number. inf and NaN pass through as themselves. This is the one encoding that cannot align with its C++ counterpart, because a Python float carries no width where to_u64(1.5) is a double; ctypes.c_double is the spelling for full precision. Two Python C API returns that signal failure are checked rather than used: PyObject_IsInstance answers -1, which is truthy, and PyNumber_Index answers nullptr with the caller's own exception pending. The common Python-int case is tested first, so the hot path costs one PyLong_CheckExact and no attribute lookup. test_task_interface.py pins every encoding across scalar_to_uint64, TaskArgs.add_scalar and ChipStorageTaskArgs.add_scalar, including the zero-extension widths, a c_double subclass, both byte-order qualifiers, the single-precision range boundary against struct.pack's own verdict, out-of-range integers, and an __index__ that raises. * Fix: check args-dump payload truth at every level, on both arches (#2214) The args-dump scene test validated structure only — entry counts, arg indices, offsets, sizes, the manifest schema — and a dump whose payload is correctly sized and entirely zero satisfies every one of those. #1560 is exactly that failure: a5sim `--dump-args 2` wrote correctly sized, all-zero tensor payloads and the test reported PASSED. The a2a3 test did carry one payload-truth assertion, but only under `if level == 3`, so levels 1 and 2 had none. The a5 test had that block deleted, its docstring recording that "the A5 payload values remain untrusted until #1560 is fixed". The assertion now runs at every level that writes a payload, in both tests. It anchors on task 0's `a + b` over the 2.0 / 3.0 inputs, whose bytes are known before the run, rather than on "not all zero": the case also passes `torch.zeros` as an argument, so an all-zero payload is a legitimate value somewhere in this dump and a blanket non-zero rule would be wrong. a5 also regains the level-3 payload-selection check. The two files now differ only in `KERNELS_BASE` and their platform lists. The underlying bug is gone. Running #1560's own per-record check — payload sliced by each entry's `bin_offset` / `bin_size` — over a fresh a5sim level-2 run gives 13 tensor records, 0 of them all-zero, matching the a2a3sim control exactly. Verified at levels 1, 2 and 3 on a2a3sim, a5sim and a2a3 onboard: 9 runs, all green. The negative control expects 6.0 instead of 5.0 and fails all six sim combinations, including levels 1 and 2 — which is the point, since those two had no payload assertion to fail before. Note that CI exercises level 1 only: `_st-sim-a2a3.yml`, `_st-sim-a5.yml` and `run-onboard-dfx-smokes` all pass a bare `--dump-args`, and `conftest.py` declares that flag `nargs="?", const=1`. So levels 2 and 3 still have no coverage there, including the level-3 mode `core_swimlane.py` consumes. * Refactor: name the per-run H2D copy-in instead of overloading "staged" (#2213) "stage" carries four unrelated meanings in this repo and is defined nowhere. Two of them meet inside one file: host_build_graph's runtime_maker.cpp reports `staged=%d` for caller tensors it copied to the device, and a hundred lines away holds a `staging` block that is the host scratch the Graph Definitions are assembled in. The scheduler adds a third (`staged_core_mask`, cores held ready before release) and `drain_stage` a fourth (a step in a sequence). A reader cannot tell which is meant without following the code, and error text shown to users inherits the ambiguity. This renames one sense: the per-run copy of a caller tensor into device memory. The other senses keep the word -- a staging buffer, a pipeline stage, a staged IPC frame and the args-dump capture stage are all ordinary uses of it, and the sense renamed here had the weakest claim, since the device buffer it produces is the live one the kernel reads rather than somewhere bytes pass through. The replacement reuses names the repo already has rather than coining any. A tensor on this path is `child_memory=False`, i.e. `AddressSpace::HOST`, so it is a host-memory tensor; the operation is an H2D copy-in, and `h2d` is already how the sibling bind phase spells it (`BindArenaH2d`, `arena_h2d`). So `bind.args` now reports `h2d=%d bytes=%llu`, and tensormap_and_ringbuffer's `stage_device_args` becomes `copy_in_device_args`. Both runtimes carry this sense, so both move, and each arch sibling moves with its pair. Nothing parses the attribute string programmatically -- `h2d=` replaces `staged=` in logs and in docs/dfx/hbg-bind-phases.md only. The remaining occurrences were classified by reading them rather than by pattern: a regex over the obvious spellings missed tmr's "Failed to stage tensor" diagnostic, SCALAR_DATA_ACCESS.md's "used to stage", and the prose in task-flow.md and buffer-abi.md. `docs/investigations/` is left as written, since those are dated records of past measurements. docs/testing.md also said the default copies every tensor in on every round, which overstates it: pure OUT buffers skip the H2D. No behaviour change. * Update: take the PR-head refresh that does not re-open an adjudication The heads this integration was audited against were frozen on 2026-09-11. Since then #2064 merged to main and six of the thirteen remaining kernel-mode PRs moved. Four of those moves are adoptable as they stand; the rest reverse decisions this line already made. D15 in INTEGRATION-LOG.md records the survey and says what each deferred item needs before it can be taken. Adopted: - From #2177, build_kernel_pipeline_contract_impl rejects an invalid CallConfig with INVALID_ARGUMENT and keeps INTERNAL for a null output or a contract it generated but cannot validate. This extends the K1 numbering D14 adopted for the entry points to the hook behind them. - From #2176, KernelContextOps drops event_flag; create_event takes (context, event) and the onboard implementation owns ACL_EVENT_SYNC, so the platform constant stays with the platform. - From #2176, a PersistentKernelArgs rollback whose own release fails now has coverage: a host-only test for the block the owner retains, and an rtFree fault hook behind the new persistent_free_close lifecycle scenario. - From #2185, the Python entry test terminates a hung child instead of leaving it for the CI job timeout, and destroying the borrowed stream after a refused kernel_init is an assertion rather than a swallowed exception. persistent_free_close asserts that the retry re-attempts and completes rather than that it makes exactly one more rtFree call. A kernel context here releases several device blocks and the injected failure stops the first pass partway, so the retry covers the failed block plus everything the first pass never reached: eight calls across the two passes where the first made three. The count-exact form the source PR uses holds only for its own allocation set. #2176, #2177 and #2185 also drop kernel_execution_state.cpp from a platform source list, which does not apply here: kernel_resource_requirements.h calls bind_resources_for_launch from the H chain, so the simulation host runtimes need the translation unit too. Validation is recorded in docs/kernel-integration-validation.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update: rebase the kernel-mode line onto the target call-flow design The target call-flow design settles three of the four adjudications D15 left open, and names this line as still carrying the old shape. D16 in INTEGRATION-LOG.md records the decisions; they are one package rather than three independent choices. Registration mints the id. simpler_kernel_mode_prepare_callable takes (ctx, callable, size, int32_t *out) and writes a context-local id on success, -1 on every failure. It takes no caller_stream: registration enqueues on the context's own AICPU stream, which every later launch also enqueues on, so stream FIFO carries the ordering. Only launch takes a stream, per call. Registration is pure. There is no deduplication and no lookup, so the same image registered twice takes two ids, two uploads and two charges, and arena capacity is spent per registration rather than per distinct image. No callable generation survives. SimplerCallableHandle, PTO_RUNTIME_ERR_CALLABLE_STALE, KernelCallableDeviceResidency::generation and SimplerKernelInvocationHeader::generation are gone; the wire header is 32 bytes and the device descriptor 24. Three properties replace the guard: an id is minted once and never reused within a context, close invalidates every id the context minted, and a closed worker accepts no launch. PTO_RUNTIME_ERR_CAPACITY_EXCEEDED takes the vacated BASE - 8. The launch sequence is chained rather than sibling. Events are Start, AicoreStart, AicoreDone, AicpuDone, SerialTail: the caller records Start, AICPU waits it, clears the handshake and records AicoreStart, AICore waits that and records AicoreDone, AICPU launches with HostArgs, joins AicoreDone and records AicpuDone, and the caller joins that and records SerialTail. Caller and AICore share no event, so ACLGraph capture propagates in two hops and no wait crosses a capture boundary. AicoreStart must precede the AICPU launch, since the AICPU orchestrator spins on AICore's handshake report. Compensation now cancels only between the AICore and AICPU launches and drives the chain back through AICPU, which is the caller's only path to a tail. #2190's block allocator is not taken: its cache is host-only, while this line's AICPU entry resolves a descriptor at arena_ + callable_id * sizeof(descriptor). The single arena and its descriptor prefix stay. Validation is recorded in docs/kernel-integration-validation.md. The chained sequence is what the a2a3 capture probe drives, and it completes 100 captured replays with verified buffers; the public prepare/launch entries are still not exercised inside a captured graph by any test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update: carry the callable image span in the launch packet #2190 restored its device entry without a residency descriptor: the packet carries chip_callable_address and chip_callable_bytes, and the consumer receives the ChipCallable reference. This line follows that shape. D17 in INTEGRATION-LOG.md records why. The descriptor bought two things and only one of them was real. Its stated purpose was revocation — the entry re-read the slot on every invocation including graph replay, so the host had one small location it could invalidate. Nothing ever wrote a descriptor after commit: ids are never evicted, kernel-mode unregister_callable returns INVALID_STATE, and clear() only drops host metadata. It would not have covered the case that matters either, since after close the freed arena leaves the descriptor address dangling exactly as the image address does. What it did buy was keeping the extent out of the per-call packet, so a packet could name which descriptor but not widen the window the device parses — an argument that assumes an untrusted packet, when the packet is built by the host binder from cache.resolve(). SimplerKernelDispatchArgs replaces residency_address with chip_callable_address and chip_callable_bytes. This line keeps the four binding fields #2190 has no source for — binding_address, context_generation, sm_bytes, arena_bytes — because its TMR consumer reads them, so the prefix is 88 bytes against #2190's 56. consume_kernel_invocation takes (args, const ChipCallable &, callable_bytes, payload, payload_bytes), passing the whole prefix for the same reason. The entry validates the span — non-null, alignof(ChipCallable), at least sizeof(ChipCallable), no wraparound — and performs no device read; cache visibility for the image moves to the consumer, where the image is parsed. KernelCallableDeviceResidency and its header are deleted, the code arena loses its descriptor prefix, and KernelDispatchStatus retires NotResident and Stale rather than reusing 2 and 3. docs/zh-cn/kernel-mode-integration-test.md also had a stale launch diagram: it still showed the sibling topology the previous commit replaced with the chained one. Its device-side section and interface-adjudication table follow the packet change. Validation is recorded in docs/kernel-integration-validation.md. Every scene and unit count matches the pre-change baseline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update: raise MAX_REGISTERED_CALLABLE_IDS to 8192 This was the one part of #2190's contract this line had not taken. Both lines already register purely — no deduplication, no lookup, no reuse before close — so every prepare spends an id permanently and 64 was the binding limit for a caller covering all of a model's specializations. The capacity unit test stops naming 64: it fills MAX_REGISTERED_CALLABLE_IDS residents and static_asserts that one arena still holds them at that image size, so the case keeps measuring the count limit rather than silently becoming a byte-limit test. The AICPU cost is real and measured, not incidental. The TMR AICPU executor holds orch_so_table_[MAX_REGISTERED_CALLABLE_IDS] of ~296-byte entries, so libaicpu_kernel.so's .bss grows from 450 KiB to 2.73 MiB. The device loads and runs it: the kernel C API and capture suites pass unchanged on a2a3. The entry is 86% char path[256], which the kernel path never uses since it registers by dev_orch_so_addr, so moving path out of the resident table would return about 2 MiB. That is left alone here because it is program-mode registration code. docs/zh-cn/kernel-callable-residency.md records which of the two admission limits binds first, since that now depends on image size: sizeof(ChipCallable) is 9376 bytes, so 8192 empty images occupy 73.5 MiB and the count binds, while 1 MiB images exhaust the 512 MiB budget at 512 registrations and the bytes bind. Validation is in docs/kernel-integration-validation.md, including the triage of four cases that failed once on a device that dropped its host channel (507901, hdc disconnect) and pass 9/9 on re-run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update: align the callable cache and image validation with #2190 Two things come over together. Device memory for callable images is taken in 2 MiB blocks against a 2 GiB budget instead of one 512 MiB arena committed on the first registration, and the structural image validation moves into `validate_kernel_callable_image` so the entry and the cache stop carrying two copies of it. The block allocator was blocked until now by the residency descriptor, which indexed a fixed prefix of the single arena by callable_id. That descriptor is gone, so nothing requires the images to be contiguous from one base. The first registration now takes max(charged, 2 MiB); a registration larger than a block gets one sized exactly to it while earlier blocks keep their usable tails; block slack is charged against the budget through allocated_bytes() while resident_bytes() still counts only charged image bytes; and a published device_address never moves as blocks are added. With the id cap at 8192 the crossover between the two admission limits lands at exactly 2 GiB / 8192 = 256 KiB per image. The weak stub for consume_kernel_invocation is now linked into both dispatch test targets, so test_kernel_dispatch's own strong consumer proves the override resolves rather than colliding. That is the link error a runtime-specific consumer would hit if the fallback were not weak. Two deliberate divergences from #2190 remain. validate_kernel_callable_image keeps rejecting a child func_id that is out of range or repeated within one image. #2190's extraction lost those two checks, and without them such an image reaches the device consumer, which rejects it at launch and poisons the context rather than failing the registration cleanly. The simulation prepare_callable calls the image validator too, so both entries agree that structural checks precede the lifecycle refusal. Without it an image whose size clears the header floor but whose variable tail does not add up reports INVALID_STATE instead of INVALID_ARGUMENT, which is what test_kernel_entries_reject_a_context_with_no_kernel_claim caught. kernel_arena_change_is_forbidden is a third difference but not one against #2190. It is mainline code from the merged K1 (#2064), so every branch off main carries it. This line deleted it in D14 item 4 and routes the same rule through K3's commit_static_arena_bank, which applies it to onboard and simulation from one place. That consolidation removes a guard main uses today in favour of a mechanism that exists only in an unmerged PR — static_arena_bank.h is on this line and on #2193 and nowhere else — so it is recorded in the validation doc as a risk rather than a settled question. Validation is in docs/kernel-integration-validation.md. Every scene and unit count matches the pre-change baseline with no re-runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Fix: bound the packet length against the mapped prefix, not SIZE_MAX `simpler_aicpu_kernel_exec` treats `packet_bytes` as the length of the region starting at `arg`, and derives the payload span from it. The upper bound compared it against `std::numeric_limits<size_t>::max()`, which a `uint64_t` field can never exceed on a 64-bit AICPU, so the check was dead and `arg + packet_bytes` could wrap. Bound it against the space actually left above `arg` instead, so the payload span the consumer receives is guaranteed not to wrap the address space. This is the shape #2190 carries at `9c085aa0`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update: bound a child func_id by the runtime function table `validate_kernel_callable_image` rejected a child `func_id` outside the extent of `ChipCallable::child_func_ids_`. That array's capacity is the number of children an image may carry, not the size of the table the id indexes; the two agree at 1024 only by coincidence, so raising the child capacity would have silently widened the accepted id range past what the device function table holds. Name the real bound, `KERNEL_MAX_FUNC_ID`, next to the registration-id cap it is independent of, and tie it to `RUNTIME_MAX_FUNC_ID` with a `static_assert` in both onboard and simulation `device_runner_base.cpp`, so the two cannot drift apart without a compile error. This is the shape #2190 carries at `08f7f921`, adopted along with the test it added; the accepted set is unchanged today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Docs: record the newest-head review and its regression The frozen-heads table now shows #2190 at `08f7f9218728`, and the validation doc carries the sweep run against the binaries this line actually ships: 166 and 168 C++ unit tests, 2428 Python unit tests, 2 C++ hardware tests, 21 scheduled hardware unit cases plus the ten that marker run drops, 13 kernel C API cases, the capture replay, and all four scene platforms. Every count matches the pre-change baseline. D18 records why the child `func_id` bound moved from the extent of `ChipCallable::child_func_ids_` to `KERNEL_MAX_FUNC_ID`, and the divergence note in the earlier section is corrected: #2190 restored those checks, so only the simulation `prepare_callable` ordering still differs. Two harness details are written down because each cost a false red: the C++ unit build takes no `CMAKE_BUILD_TYPE`, since `Release` defines `NDEBUG` and silences the `debug_assert` three tests assert on; and the C++ hardware tests need a `--resource-spec-file` built from `TASK_DEVICE`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Fix: drop the duplicated kernel_device_resources entry in the a2a3 sim list `HOST_RUNTIME_SOURCES` in the a2a3 simulation host list named `common/platform/shared/host/kernel_device_resources.cpp` on two consecutive lines. The a5 simulation list and both onboard lists name it once, and CMake dedups a target's source list, so this changes no build output — it is a copy/paste slip from the integration commit, removed so the four lists read the same. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: chenshengxin2026 <hw_chenshengxin@163.com> Co-authored-by: poursoul <49787929+poursoul@users.noreply.github.com> Co-authored-by: Chao Wang <26245345+ChaoWao@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
08f7f92 to
a8b19e8
Compare
Upload and register every kernel callable independently, including identical content, and return a fresh context-local integer ID. Keep ID reuse in the caller and retain each registration until close. Support 8192 registrations across host and AICPU protocol bounds. Allocate shared 2 MiB code blocks on demand and give larger images independent aligned allocations. Cap total allocated code blocks, including unused tails, at 2 GiB without moving published addresses. Preserve existing residents on admission failure and check ID residency before kernel launch. Keep K2 context resources and preparation intact. Expose a separate AICPU kernel entry for immutable invocation packets, carry the resident ChipCallable device address and length to the AICPU consumer for per-invocation function-table reconstruction. Validate envelope and callable span bounds before consumption, and return UnsupportedPayload from a weak fallback that runtime-specific strong consumers can override in the same link target. Keep the program KernelArgs entry separate. Use an ID-only callable contract without generation tracking. Cover block growth, capacity boundaries, rollback and registration lifetime, and document the call chain in the Chinese documentation navigation. Reject out-of-range and duplicate child function IDs during canonical image validation, before allocation or upload. Assert that the kernel protocol bound matches each runtime function table and cover boundaries. Allow preparation inside capture without stream synchronization. Upload context-owned images with aclrtMemcpy under a temporary relaxed thread capture mode, restoring the previous mode on success and failure. Enqueue kernel AICPU registration without waiting; keep program registration synchronous. Cover mode restoration and first/repeated preparation in GLOBAL and THREAD_LOCAL capture on hardware.
…2290) #2190 moved to a8b19e8, which makes prepare_callable callable inside an ACLGraph capture: context-owned H2D goes through capture_memcpy_h2d, which exchanges the thread capture mode to RELAXED around a synchronous aclrtMemcpy, and preparation's stream wait is gone. CANN's capture guide is categorical that synchronizing during capture is illegal in every mode, so the wait cannot be kept in any form. Removing it exposed what it had been paying for. K7's registration payload dlopens the orchestration SO into the AICPU process, and dispatch refuses a callable whose table slot is empty; the two are ordered only by the hidden AICPU stream's FIFO, which an eager launch is submitted to and a replayed graph is not. cold_unsynced, the one scenario draining nothing between prepare and the first replay, reported InvalidBinding on all six AICPU threads 192 us after the registration ran. Moving preparation inside the capture does not help: a registration enqueued before the first launch's Start event still executes eagerly. The load moves into prepare_kernel_round, the leader-only phase of the launch round whose status the gate publishes to every other thread. It runs exactly once, inside the task that needs the orchestration, so no cross-stream edge is required. - register_kernel_callable records the image span and child function table and no longer loads. Dispatch already receives the span in the packet, so it establishes residency itself when the slot is empty; a launch no longer depends on the registration payload having executed. - OrchSoEntry gains needs_load. A revoke clears residency but not the dlopen handle, so a non-null handle does not mean the current image is loaded. - load_orch_so resets only its dlopen state. Resetting the whole entry dangled the KernelCallableView a dispatch already held, observed as a device-side Scheduler fatal error (code=5). - capture_memcpy_h2d covers the persistent args copy, the callable cache copy and init_aicore_register_addresses on both arches, plus two uploads #2190's prepare does not reach: the prebuilt runtime arena, through HostApi::copy_to_device in kernel mode only, and K7's coordination image. - The revoke event joins the five KernelEventKind events created at init, instead of being created and destroyed on its own in preparation. - The capture observer's two sync scopes collapse into one, since nothing in preparation may wait. Its fault injection and gate placement are unchanged. A device-side dlopen refusal is now reported by neither preparation nor a drain, only by the first launch round; register_failure asserts that. Program-mode registration is untouched and still loads and waits. Decisions are D22 in the integration log. On a2a3: 30/30 kernel_mode_capture scenarios including cold_unsynced with both waits gone, 30/30 hardware Python unit cases, 177/177 C++ unit tests, 2529 non-hardware Python tests, and pre-commit on every changed file.
关闭:交付路线已替代(superseded)根据维护者确认,关闭这份旧 kernel-mode PR。 **原因:**有界、稳定地址、不自动淘汰的注册方向仍符合目标。但当前整包同时引入 capture 内 prepare、RELAXED 同步上传及独立 dispatch 接线;集成线 #2290 / INTEGRATION-LOG D22 已吸收并继续修改。这不是当前 main 路线可直接合入的单元。 **仍需保留/迁移:**后续拆出面向当前 main 的有界代码池、稳定 callable handle、容量失败保护和保活测试;注册等冷路径在 capture 前完成,错误与退休接共同执行层。关闭不取消常驻注册能力,也不否认现有硬件证据。 当前确定路线是 Pipeline A → kernel eager → capture/replay:两种资源模式接共同执行与退休层,每次 eager 重新准备并提交一次。集成分支中的采用不等于已合入 main,也不代表本 PR 的全部目标已交付。 本次只关闭旧 PR,保留分支、提交、作者贡献、测试及实验记录,供后续按能力迁移。 本次核查 head: |
Each successful kernel
prepare_callablecall uploads and registers a new callable and returns a fresh context-localint32_tID, even for identical content. The caller owns ID reuse and content deduplication. Registrations retain independent device addresses until context close.Kernel residency supports 8192 registrations. Small images share 2 MiB device blocks, while images larger than a block receive an independent allocation rounded to 64 bytes. Blocks grow on demand only during prepare; published device addresses never move. The 2 GiB cap applies to all allocated code blocks, including alignment and unused block tails. Entry count and device capacity are independent limits; neither preallocates execution resources.
The shared callable-ID protocol bound covers Host runner checks and the A2A3/A5 AICPU registration tables and dispatch guards. Host kernel launch resolves IDs directly without allocation. Admission failures preserve existing residents; failed unpublished uploads retain reusable block storage within the same budget.
Kernel preparation is supported inside GLOBAL and THREAD_LOCAL capture after initialization outside capture. Callable images, persistent Runtime/KernelArgs, and register-address tables use synchronous
aclrtMemcpyunder a temporaryACL_MODEL_RI_CAPTURE_MODE_RELAXEDthread mode, restored after the copy even on failure. These uploads execute immediately and their device allocations remain valid through graph replay. Prepare performs no stream/device synchronization: success guarantees upload and registration enqueue, while AICPU execution errors surface through subsequent warmup and synchronization. The kernel Host launch binder is still unconnected, so this PR does not validate full kernel graph execution.Interface and lifetime
Failed preparation writes
-1. Successful IDs are never reused within a context and become invalid at close. Callers must pair IDs with their issuing context and drain executions and destroy referencing graphs before close. There is no callable generation field or device residency descriptor. K9 carries the callable ID and payload metadata without a callable generation.The separate device entry
simpler_aicpu_kernel_execaccepts aSimplerKernelDispatchArgsprefix followed by an immutable runtime payload. It validates packet size, address arithmetic, kernel mode, ID range, argument counts and reserved fields before callingconsume_kernel_invocation. The packet also carrieschip_callable_addressandchip_callable_bytes, filled from the issuing context's committed residency. After null/alignment/size/overflow checks, the entry passes aconst ChipCallable &and image length to the consumer. This gives the AICPU access to child function IDs and patched device code addresses for per-invocation function-table reconstruction. The consumer must establish cache visibility and validate child offsets/signatures before consuming the image. The producer must supply an actual allocation matchingpacket_bytes; the entry ABI has no independent allocation length. This protocol does not share the program-modesimpler_aicpu_exec/KernelArgsentry. The default consumer is weak and returnsUnsupportedPayload; a runtime-specific strong consumer overrides it in the same link target without duplicate symbols. In this branch, runtime-specific function-table reconstruction and payload execution remain unconnected, as does the Host launch binder.Dependencies and scope
abbe07f5; dependency commits remain separate from K10a.docs/zh-cn/kernel-callable-residency.md.Validation
Capture regressions reproduce a failure before this fix and pass after it: first prepare plus 64 additional registrations inside GLOBAL and THREAD_LOCAL capture, both runtimes, A3 hardware. The full kernel C ABI suite passes 34 tests (22 platform cases deselected). A preload guard rejects and counts synchronization calls during preparation; all prepare paths record zero calls.
Capture-copy unit tests cover all thread modes, exchange failure, copy failure with restoration, and restore failure propagation; all five relevant C++ targets pass. Program vector execution passes on both a2a3sim and A3 hardware for both runtimes. CPU-only C ABI checks pass (12 passed, 44 skipped).
Canonical image validation rejects child function IDs outside
[0, 1024)and duplicates within an image before allocation/upload. The bound is checked against every runtime at compile time. All 33 relevant C++ tests pass, including the new regression that failed before this fix.Integration link regression: the test links both the default consumer and a strong runtime substitute. It reproduces duplicate symbols before the weak attribute and passes after the fix; the default-only consumer test also passes. All eight AICPU builds expose the fallback as a weak symbol.
Removed the always-false uint64 device-address comparison with uintptr maximum; retained the address-plus-length overflow check.
Editable runtime rebuild succeeds; all pre-commit checks pass.
All eight AICPU runtime libraries export
simpler_aicpu_kernel_exec. The two device-entry test targets pass; six entry-validation tests also pass ASan/UBSan, including A/B/A callable switching and invalid device-image spans. The updated kernel C API/Host ABI regression passes (16 passed, 40 skipped).Five relevant C++ targets pass. Four new capacity/pool tests failed against the previous implementation and pass with this change.
Fourteen registration tests pass ASan/UBSan: 8192 IDs, small-image sharing, large-image allocation, stable addresses, block-slack accounting, allocation/upload failures and rollback.
Python/C ABI/callable-identity/task-interface/worker ABI regression: 296 passed, 40 skipped.
A3 onboard kernel lifecycle: 18 passed. Both runtimes register 65 identical images, check bounded first allocation and incremental block accounting, and verify persistent-free retry and zero memory after close.
Program-mode vector execution passes on a2a3sim for both runtimes. Additional program-mode hardware smoke was cancelled while all devices were occupied.
The Chinese call-chain document is included in site navigation; the new strict docs CI build passes (run 34817495346).
Full PR CI reruns after push. Prior-head macOS remote transport and A3 scene failures were observed; this change does not claim those checks green. Actual kernel launch/binder and graph replay remain outside the validation scope.