Skip to content

refactor: extract AICPU device profiler engine - #1276

Merged
ChaoZheng109 merged 2 commits into
hw-native-sys:mainfrom
vegetabledoww:fix/issue-1247-aicpu-device-engine
Jul 8, 2026
Merged

ChaoZheng109 merged 2 commits into
hw-native-sys:mainfrom
vegetabledoww:fix/issue-1247-aicpu-device-engine

Conversation

@vegetabledoww

@vegetabledoww vegetabledoww commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #1247.

This PR introduces a shared AICPU-side profiling device engine, DeviceProfilerEngine<Module>, as the device-side counterpart to the host ProfilerAlgorithms<Module> pattern. The goal is not just to extract low-level wait helpers; it is to centralize the common producer-side enqueue/pop/switch buffer operation layer that had been duplicated across AICPU collectors.

The shared engine now owns the common ready-queue handoff, free-queue pop/install, and current-buffer switch flow. Each collector keeps only its subsystem-specific record fill, flush/finalize, state shape, and local hooks.

What Changed

  • Added src/common/platform/include/aicpu/profiler_device_engine.h.
  • Migrated the common AICPU producer buffer flow for:
    • ScopeStats
    • DepGen on a2a3 and a5
    • TensorDump / ArgsDump
    • PMU on a2a3 and a5
    • the standard L2Swimlane AICPU task buffer path on a2a3 and a5
  • Updated docs/profiling-framework.md to document the new device-side profiling layer.
  • Updated the issue plan under the external working notes with the implementation and validation status.

Design Notes

DeviceProfilerEngine<Module> is intentionally header-only and trait-driven, matching the existing host-side design style. The module trait supplies the collector-specific types, queue sizes, wait limits, pointer/sequence accessors, ready-entry writer, and event hooks.

The engine covers the shared flow:

  1. wait for ready-queue space
  2. enqueue the current buffer to the ready queue
  3. wait for a free-queue entry
  4. pop and install the next writable buffer
  5. preserve the existing drop-accounting and memory-barrier behavior through module hooks

L2Swimlane is only partially normalized by design. Its AICPU task buffer path now uses the shared engine, while phase-pool switching/recovery and AICore rotation remain local because they have different sequence, retry, drop, and AICore-visible head semantics. Forcing those paths into the generic engine would make the common abstraction less clear and risk changing L2-specific behavior.

Validation

Build and static checks:

  • python -m pip install --no-build-isolation -e .
  • cmake --build /data/jinzongquan/simpler/build/ut_cpp_issue1253 -j2
  • git diff --check
  • commit-time pre-commit hooks: headers, English-only, EOF, trailing whitespace, clang-format, clang-tidy, cpplint, markdownlint

Unit tests:

  • test_a2a3_orchestrator_fanin
  • test_a5_orchestrator_fanin
  • test_scope_stats_collector
  • test_a2a3_scope_stats_collector

a2a3 onboard ST:

  • ScopeStats
  • DepGen
  • DepGen chain
  • ArgsDump / TensorDump with --dump-args 1, --dump-args 2, and --dump-args 3
  • PMU
  • L2Swimlane
  • L2Swimlane mixed

a5sim ST:

  • DepGen
  • ArgsDump
  • PMU

Known validation note:

  • a5sim L2Swimlane currently segfaults with --enable-l2-swimlane 1/2/3/4 on this branch.
  • The same --enable-l2-swimlane 1 scenario was reproduced on a clean main worktree, so this appears to be a pre-existing main-branch issue rather than a regression from this PR.

Risk / Follow-up

  • The refactor touches device-side queue handoff and memory-barrier-sensitive paths. The implementation keeps the barriers explicit in the shared engine and was onboard-validated on a2a3, but a5 onboard validation is still recommended before merge if available.
  • L2Swimlane is intentionally not fully collapsed into the generic engine. The remaining local code is the part where L2 behavior genuinely differs from ScopeStats, DepGen, TensorDump, and PMU.

@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfdffee2-a5fa-4ed0-8068-205f8147322a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Introduces a shared profiling_device::DeviceProfilerEngine<Module> template header encapsulating AICPU-side ready/free queue backpressure, enqueue, pop, and buffer-switch logic. Six AICPU collectors (dep_gen, l2_swimlane, pmu across a2a3/a5, plus common scope_stats and tensor_dump) are refactored to delegate to this engine via per-subsystem module adapters. Documentation updated accordingly.

Changes

AICPU Device Profiler Engine Extraction

Layer / File(s) Summary
DeviceProfilerEngine template header
src/common/platform/include/aicpu/profiler_device_engine.h
New template struct providing wait_for_ready_queue_space, wait_for_free_queue_entry, enqueue_ready, pop_free, and switch_buffer static methods driven by module-provided type aliases and hook callbacks.
dep_gen collector migration
src/a2a3/.../dep_gen_collector_aicpu.cpp, src/a5/.../dep_gen_collector_aicpu.cpp
Replaces local queue wait/enqueue/pop/switch implementations with DepGenDeviceModule adapter and DepGenEngine delegation.
l2_swimlane collector migration
src/a2a3/.../l2_swimlane_collector_aicpu.cpp, src/a5/.../l2_swimlane_collector_aicpu.cpp
Adds L2SwimlaneTaskDeviceModule/L2SwimlaneTaskEngine adapter; enqueue_ready_buffer, try_pop_records_buffer, and switch_records_buffer now delegate to the engine.
pmu collector migration
src/a2a3/.../pmu_collector_aicpu.cpp, src/a5/.../pmu_collector_aicpu.cpp
Adds PmuDeviceModule/PmuEngine adapter; ready/free queue helpers and pmu_switch_buffer delegate to PmuEngine.
scope_stats and tensor_dump collector migration
src/common/.../scope_stats_collector_aicpu.cpp, src/common/.../tensor_dump_aicpu.cpp
Adds ScopeStatsDeviceModule/ScopeStatsEngine and DumpDeviceModule/DumpEngine adapters replacing manual queue/switch logic.
Profiling framework documentation update
docs/profiling-framework.md
Documents the device-side DeviceProfilerEngine<Module>, adds a layered-view diagram, a new §3.5 section, and a collector-authoring step.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Collector as AICPU Collector
  participant Engine as DeviceProfilerEngine
  participant Ready as Ready Queue
  participant Free as Free Queue

  Collector->>Engine: switch_buffer(ctx, state)
  Engine->>Ready: enqueue_ready(full buffer)
  alt enqueue succeeds
    Engine->>Engine: on_current_cleared()
  else ready queue full
    Engine->>Engine: on_enqueue_dropped()
  end
  Engine->>Free: pop_free(next_seq)
  alt free buffer available
    Engine->>Collector: on_switch_complete(new buffer)
  else no replacement
    Engine->>Collector: on_no_replacement()
  end
Loading

Possibly related issues

Possibly related PRs

  • hw-native-sys/simpler#916: Builds on the same L2 swimlane buffer lifecycle hot-path this PR routes through DeviceProfilerEngine.
  • hw-native-sys/simpler#1058: Modifies the a5 L2 swimlane rotation state that this PR's engine migration wraps.
  • hw-native-sys/simpler#1162: Reworks the same L2 swimlane ready/free enqueue/pop/switch and drop-accounting logic this PR migrates to the engine.

Poem

A rabbit hopped through queues so tight,
Ready, free, and switch — all set right.
One engine now to rule the hop,
No more copy-paste, the drift will stop.
🐇 Buffers dance in tidy rows,
Where the DeviceProfilerEngine goes!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed It clearly summarizes the main change: extracting a shared AICPU device profiler engine.
Description check ✅ Passed It accurately describes the refactor, affected components, and validation context.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a unified AICPU-side producer algorithm layer, DeviceProfilerEngine, in profiler_device_engine.h and refactors various profiling collectors (DepGen, L2Swimlane, PMU, ScopeStats, and TensorDump) across the a2a3 and a5 platforms to use it, reducing code duplication. Feedback on the new engine highlights a performance concern in the tight spin loops of wait_for_ready_queue_space and wait_for_free_queue_entry, where calling get_sys_cnt_aicpu() on every iteration can cause high CPU overhead due to expensive MMIO reads; gating these checks to run periodically is recommended.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +40 to +53
const uint64_t start = get_sys_cnt_aicpu();
do {
uint32_t current_tail = header->queue_tails[thread_idx];
uint32_t current_head = header->queue_heads[thread_idx];
uint32_t next_tail = (current_tail + 1) % Module::kReadyQueueSize;
if (next_tail != current_head) {
*tail_out = current_tail;
*head_out = current_head;
return true;
}
if (get_sys_cnt_aicpu() - start >= Module::kBackpressureWaitCycles) {
break;
}
} while (true);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Reading system counters via get_sys_cnt_aicpu() on every iteration of a tight spin loop can introduce significant CPU overhead and latency due to expensive MMIO reads. It is highly recommended to gate these checks to run periodically (e.g., every 1024 spins) to minimize overhead.

        const uint64_t start = get_sys_cnt_aicpu();
        uint32_t spin_count = 0;
        do {
            uint32_t current_tail = header->queue_tails[thread_idx];
            uint32_t current_head = header->queue_heads[thread_idx];
            uint32_t next_tail = (current_tail + 1) % Module::kReadyQueueSize;
            if (next_tail != current_head) {
                *tail_out = current_tail;
                *head_out = current_head;
                return true;
            }
            if (++spin_count % 1024 == 0) {
                if (get_sys_cnt_aicpu() - start >= Module::kBackpressureWaitCycles) {
                    break;
                }
            }
        } while (true);
References
  1. Avoid reading system counters (which can be expensive MMIO reads) or performing complex structural checks on every iteration of a tight spin loop. Instead, gate these checks to run periodically (e.g., every 1024 spins) to minimize CPU overhead and latency impact.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in eac1111 by polling the backpressure timeout every 1024 spins instead of reading get_sys_cnt_aicpu() on every wait-loop iteration. The zero-wait-cycle case still exits immediately.

Comment on lines +63 to +76
const uint64_t start = get_sys_cnt_aicpu();
do {
uint32_t head = free_queue->head;
uint32_t tail = free_queue->tail;
if (head != tail) {
*head_out = head;
*tail_out = tail;
rmb(); // acquire: order the tail read above before the caller's buffer_ptrs read
return true;
}
if (get_sys_cnt_aicpu() - start >= Module::kBackpressureWaitCycles) {
break;
}
} while (true);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Reading system counters via get_sys_cnt_aicpu() on every iteration of a tight spin loop can introduce significant CPU overhead and latency due to expensive MMIO reads. It is highly recommended to gate these checks to run periodically (e.g., every 1024 spins) to minimize overhead.

        const uint64_t start = get_sys_cnt_aicpu();
        uint32_t spin_count = 0;
        do {
            uint32_t head = free_queue->head;
            uint32_t tail = free_queue->tail;
            if (head != tail) {
                *head_out = head;
                *tail_out = tail;
                rmb();  // acquire: order the tail read above before the caller's buffer_ptrs read
                return true;
            }
            if (++spin_count % 1024 == 0) {
                if (get_sys_cnt_aicpu() - start >= Module::kBackpressureWaitCycles) {
                    break;
                }
            }
        } while (true);
References
  1. Avoid reading system counters (which can be expensive MMIO reads) or performing complex structural checks on every iteration of a tight spin loop. Instead, gate these checks to run periodically (e.g., every 1024 spins) to minimize CPU overhead and latency impact.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in eac1111 with the same 1024-spin gated timeout polling in wait_for_free_queue_entry(), avoiding a get_sys_cnt_aicpu() read on every tight-loop iteration while keeping the memory barrier behavior unchanged.

@vegetabledoww
vegetabledoww force-pushed the fix/issue-1247-aicpu-device-engine branch 2 times, most recently from ab7e64e to 13b39eb Compare July 7, 2026 11:02
@vegetabledoww

Copy link
Copy Markdown
Contributor Author
image

@ChaoZheng109

Copy link
Copy Markdown
Collaborator
image

🐮

if (Module::kBackpressureWaitCycles == 0) {
break;
}
if ((++spin_count & kCounterPollSpinMask) != 0) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Undocumented behavior change: the new kCounterPollSpinMask alters the bounded-wait, contrary to the PR's "preserve existing behavior" claim.

Both wait_for_ready_queue_space and wait_for_free_queue_entry now gate the timeout check behind (++spin_count & kCounterPollSpinMask) != 0 continue — so get_sys_cnt_aicpu() is only sampled once per 1024 spins. The four common collectors (ScopeStats / DepGen / TensorDump / PMU) and the old L2 path all checked get_sys_cnt_aicpu() - start >= kBackpressureWaitCycles on every iteration and had no poll-mask before this PR (verified against merge-base). So this is a genuine, new behavior for all of them.

I don't think it's a correctness issue: the timeout overshoot is bounded to ~1024 cheap head/tail reads, and cutting the sys-cnt sampling rate is almost certainly a perf win on the AICPU. But two things:

  1. It's not "preserving the existing barrier/timeout behavior" as the description states — worth calling out explicitly so a future reader doesn't treat it as accidental drift.
  2. Same note applies to the rmb() that pop_free now always issues after reading buffer_ptrs[head]: ScopeStats/DepGen/TensorDump/PMU previously had no such barrier there (only old L2 did). It's a strengthening (safe), but it's another silent semantic unification.

Could you add a one-line comment on kCounterPollSpinMask (and on that rmb()) noting these are deliberate, so the intent is captured? Alternatively, if bit-for-bit behavior preservation was the goal, drop the poll-mask and keep per-iteration timeout checks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for calling this out. I chose the bit-for-bit timeout-preservation path here: removed kCounterPollSpinMask/spin_count and restored per-iteration get_sys_cnt_aicpu() timeout checks in both wait loops, since the backpressure budget is only ~20us. I kept the extra rmb() after buffer_ptrs[head] and added a comment marking it as an intentional acquire-strengthening before advancing free_queue.head.

ChaoZheng109
ChaoZheng109 previously approved these changes Jul 8, 2026
Introduce a shared DeviceProfilerEngine for AICPU-side profiling buffer handoff and switch logic. Migrate ScopeStats, DepGen, TensorDump, PMU, and the standard L2Swimlane AICPU task path while keeping subsystem-specific record fill, flush, and L2 AICore rotation semantics local. Update profiling framework docs for the new device-side layer. Throttle wait-loop timeout checks to avoid reading get_sys_cnt_aicpu() on every spin.
@ChaoZheng109
ChaoZheng109 merged commit 87252d4 into hw-native-sys:main Jul 8, 2026
16 checks passed
ccyywwen pushed a commit to ccyywwen/ptoruntime-simpler that referenced this pull request Jul 9, 2026
refactor: extract AICPU device profiler engine

Introduce a shared DeviceProfilerEngine for AICPU-side profiling buffer handoff and switch logic. Migrate ScopeStats, DepGen, TensorDump, PMU, and the standard L2Swimlane AICPU task path while keeping subsystem-specific record fill, flush, and L2 AICore rotation semantics local. Update profiling framework docs for the new device-side layer. Throttle wait-loop timeout checks to avoid reading get_sys_cnt_aicpu() on every spin.
@vegetabledoww
vegetabledoww deleted the fix/issue-1247-aicpu-device-engine branch July 14, 2026 01:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants