Skip to content

[Performance] A5 TMR: batch progress publication every 16 advances - #1575

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
yanghaoran29:exp/pr906-F3
Aug 18, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
yanghaoran29:exp/pr906-F3

Conversation

@yanghaoran29

@yanghaoran29 yanghaoran29 commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Batch A5 TMR ring->fc.last_task_alive publication every 16 Scheduler-local advances on the non-blocking path.
  • Keep Scheduler advance_lock contention on the original advance_pending_mask, drained only from no-progress iterations.
  • Use a cache-line-separated publication_request_mask only when an Orchestrator reclaim consumer has made no progress for 10 ms; Scheduler thread 0 then force-publishes and acknowledges from productive or no-progress iterations.
  • Enable batching only after the allocator, fanin-pool, and dependency-pool request/ack wiring is validated against the current Scheduler; otherwise fall back to per-advance publication.
  • Classify structural open-scope deadlock by exact ring-local task identity, and only against a freshly acknowledged watermark; retain the 500 ms wall-clock backstop.

Scope

This PR changes only:

  • A5 tensormap_and_ringbuffer Scheduler progress publication and Orchestrator reclaim back-pressure;
  • A5 runtime initialization/reuse and prebuilt-arena pointer wiring;
  • A5 C++ unit coverage, plus the shared test_scope_deadlock_detection fixture that the exact-identity head check requires; and
  • the A5 runtime, A2/A3-versus-A5 comparison, and reclaim-triage documentation.

It does not change host-build-graph, A2/A3 runtime behavior, benchmark tooling, PTO-ISA checkout, or CI workflows.

Why A5 only

Property A2/A3 A5
Physical AICPU cores per cluster 4 2
Default active roles 1 Orchestrator + 3 Schedulers = 4 1 Orchestrator + 4 Schedulers = 5
Normal placement All roles fit in one cluster Roles must span clusters and may span dies
Progress-publication cost Normally cluster-local Can transfer/invalidate the Orchestrator's cache line across clusters or dies
Port result Mean Effective change -0.24%; all results within approximately +/-2% Batch Case1 change -3.83%
Decision Keep per-advance publication Enable K=16 batching with a pressure-only liveness escape

Any Scheduler may win advance_lock and publish the shared watermark that the Orchestrator repeatedly reads. K=16 does not remove Scheduler-to-Scheduler contention; it reduces Scheduler-to-Orchestrator publication frequency by up to 16x on A5's distributed placement.

The implementation is not unsafe or unsupported on A2/A3. A local A2/A3 port passed its tests, but batching added reclamation lag without a measurable payoff, so A2/A3 remains unchanged. See the A2/A3 benchmark follow-up.

Correctness

last_task_alive is not a progress statistic: it is the allocation credit for task slots, heap bytes, dependency entries, fanin spill entries, and TensorMap entries. A blocked reclaim consumer can itself prevent an open scope from ending, so a withheld batch must never depend on full-ring drain to be published.

  • Scheduler-local last_task_alive remains authoritative; the published watermark is a lower bound.
  • K=16 applies on the non-blocking path. A resource-short consumer first waits for ordinary batched publication; only 10 ms of continuous no-progress sets its exact-publication request bit.
  • All five reclaim consumers participate: task slots and heap (PTO2TaskAllocator::alloc), PTO2DepListPool::ensure_space, PTO2FaninPool::ensure_space, and ensure_tensormap_capacity.
  • Scheduler try-lock contention and Orchestrator pressure use separate cache lines and separate drain semantics. A deferred Scheduler retry cannot acknowledge an Orchestrator request.
  • Scheduler thread 0 services pressure requests during productive and no-progress loops, so a genuinely blocked ring does not wait for global Scheduler idleness. One servicer suffices because request bits stay latched until acknowledged.
  • Batching is self-disabling. publication_batching_enabled defaults false at initialization, arena relocation, reuse, and teardown. runtime_wire_arena_pointers and runtime_reset_for_reuse enable it only after verifying every ring's allocator, fanin-pool, and dependency-pool request/ack pointers against the live Scheduler. Incomplete wiring degrades to pre-change per-advance publication rather than batching without a liveness escape.
  • Structural open-scope classification runs only after a fresh exact-publication acknowledgement, and compares exact ring-local task identity rather than a slot address, so a wrapped slot cannot alias the reclaim head. Natural watermark progress invalidates an older acknowledgement.
  • Terminal drain still force-publishes the final partial batch, and initialization/reuse resets local, published, deferred, request, and acknowledgement state.
  • Arena relocation rewires dependency-pool request/ack pointers from the host prebuilt image to the AICPU runtime address.

Validation

Rebased on main at a59ffde75c2907819a45690eee065056857330dc.

Measured on the current head:

  • Complete independent C++ suite: 100/100 passed.
  • Local commit-stage pre-commit: passed (clang-format, clang-tidy, cpplint, markdownlint).

Measured on the immediately preceding head, which differs from the current one only by removing three files unrelated to this change:

  • Full CI matrix green, including st-onboard-a5, st-sim-a5sim, ut-a5, ut-a2a3, and packaging-matrix.
  • Timing-coupled pressure suites under parallel load (--repeat until-fail:12 -j 4 across wiring, fanin_pool, orchestrator_fanin, task_allocator, shared_memory, scheduler_state, scope_deadlock_detection): 96/96 passed, no flakes.

Measured on earlier heads of this branch:

  • A5 simulator paged_attention_ringbuffer: 258/258 checks passed.
  • Locked A5 paged_attention_ringbuffer with ring_task_window=64, ring_heap=4 MiB, ring_dep_pool=256: passed.
  • Locked A5 paged_attention_manual_scope with PTO2_RING_TASK_WINDOW=64: 4/4 cases passed — the small-window plus manual-scope combination that exercises the pressure path.
  • Locked A5 Batch Paged Attention Case1 with golden comparison: passed.

A5 performance

Batch Paged Attention Case1 on Ascend950PR device 0. Four groups, each running baseline immediately followed by this branch, 100 measured rounds per side. Negative deltas mean this branch is faster.

Metric Main mean (us) PR mean (us) Change
Device 7342.7 7062.5 -3.82%
Effective 7310.0 7029.9 -3.83%
Orchestrator 6309.7 5991.3 -5.05%
Scheduler 7301.9 7022.0 -3.83%

The four paired Effective changes were -4.92%, -2.29%, -4.20%, and -3.91%. Profiling showed that an earlier implementation promoted 36 short-lived allocation-pressure episodes per Case1 run into exact publication handshakes; the 10 ms no-progress criterion promoted zero in this workload while preserving the liveness path for a genuinely blocked Orchestrator.

Directional context from a seven-workload sweep:

Workload Effective change
Alternating matmul add -5.30%
BGEMM -10.85%
Manual attention Case1 -3.21%
Manual attention Case2 -0.21%
Unroll attention Case1 -8.53%
Unroll attention Case2 -3.52%
Qwen3 14B decode (five-run median) -0.07%

Measurement provenance. Every number above was collected before the wiring-validation gate and the exact-identity head check landed; the seven-workload rows predate the Case1 table. They are not a single-SHA aggregate. Both later additions are off the steady-state path — a bool read on a cache line the hot path already writes, and one extra indirection inside the 1024-spin cold block — so no steady-state change is expected, but that expectation is unmeasured. A re-run of the Case1 table on the current head would close it. See the full earlier evidence.

@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The PR adds an OCCUPY-based AICPU topology fallback, throttles ring scheduler updates to shared memory, and introduces A5-specific benchmark mappings with a single-example benchmark wrapper.

Changes

AICPU topology fallback

Layer / File(s) Summary
OCCUPY topology synthesis and probing
src/a5/platform/onboard/host/aicpu_topology_probe.cpp
CPU topology probing now intersects CPU_TOPO results with OCCUPY data and synthesizes entries from occupied CPU IDs when no usable CPU_TOPO entries exist.

Ring scheduler synchronization

Layer / File(s) Summary
Throttled SM publication
src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/pto_scheduler.h
RingSchedState tracks the last published task counter and updates shared memory only after a 16-task advance.

A5 benchmark selection

Layer / File(s) Summary
A5 example configuration
tools/benchmark_rounds.sh
Adds A5 example cases and ordering, selects them for the A5 architecture, and supports validating a single requested example.
Case1 benchmark wrapper
tools/benchmark_a5_case1.sh
Adds strict option parsing and runs the A5 tensormap-and-ringbuffer benchmark for paged_attention_unroll.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Probe as probe_aicpu_topology_uncached
  participant CPU_TOPO
  participant OCCUPY as OCCUPY bitmap
  Probe->>CPU_TOPO: query_cpu_topo
  CPU_TOPO-->>Probe: entries or failure
  Probe->>OCCUPY: intersect reported CPU IDs
  OCCUPY-->>Probe: usable IDs or empty result
  Probe->>OCCUPY: synthesize pool when needed
Loading

Possibly related PRs

Poem

A rabbit found CPUs in a bitmap bright,
And taught the scheduler to hop just right.
Sixteen steps, then memory sings,
A5 benchmarks flutter on speedy wings.
paged_attention_unroll leads the way—
“Binky!” says Bunny, “let’s test today!”

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the A5 TMR performance change and the 16-advance publication batching interval.
Description check ✅ Passed The description directly explains the batching optimization, liveness safeguards, validation, scope, and measured performance results.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/pto_scheduler.h`:
- Line 475: Update advance_ring_pointers() and its sync_to_sm() ordering so
ring->fc.last_task_alive is stored with last_task_alive once PUBLISH_INTERVAL_K
slots have been consumed, avoiding the extra-call throttle; when reuse resets
last_published_to_sm to 0, publish immediately after the local lifetime advance.
- Around line 467-478: Update the scheduler’s terminal, idle, and drained-work
paths around advance_ring_pointers() to force-publish the current
last_task_alive watermark, including sub-interval advances of 1–15 steps.
Preserve batched sync_to_sm() behavior for ongoing work, and add coverage for 1,
15, 16, and 17 advances to verify final publication and threshold batching.

In `@tools/benchmark_a5_case1.sh`:
- Around line 17-21: Update the option handling case for -d|--device and
-n|--rounds in the argument parser to validate that a following value exists
before reading $2 or executing shift 2. When either option is missing its value,
emit the script’s usage error and terminate consistently; preserve the existing
assignments and shifts for valid values.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: f3f19174-72e0-4e19-bda6-6ddebea50fd0

📥 Commits

Reviewing files that changed from the base of the PR and between b079a23 and 511a165.

📒 Files selected for processing (4)
  • src/a5/platform/onboard/host/aicpu_topology_probe.cpp
  • src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/pto_scheduler.h
  • tools/benchmark_a5_case1.sh
  • tools/benchmark_rounds.sh

Comment thread src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/pto_scheduler.h Outdated
Comment thread src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/pto_scheduler.h Outdated
Comment thread tools/benchmark_a5_case1.sh Outdated
@yanghaoran29 yanghaoran29 changed the title a5: K=16 batched sync_to_sm (PR906 F3) + OCCUPY topo fallback a5: batch sync_to_sm every K=16 local advances Jul 29, 2026
@yanghaoran29

Copy link
Copy Markdown
Contributor Author

A5 performance re-test

Re-tested paged_attention_unroll Case1 with 100 rounds on the same device (device 0), running the main baseline first and PR #1575 second under one task-submit lock.

Metric Main baseline PR #1575 Delta Change
Effective 2061.2 us 1965.6 us -95.6 us -4.64%
Orch 1906.7 us 1799.1 us -107.6 us -5.64%
Sched 2056.8 us 1960.3 us -96.5 us -4.69%
Device 2092.9 us 1996.9 us -96.0 us -4.59%
Host 148042.1 us 148592.8 us +550.7 us +0.37%

Test details:

  • Baseline: cb32999c (the main parent of the PR commit)
  • PR: 526eb8bb
  • PTO-ISA pin: 83d01313d9bfc247c4b7c8bcf969d1019f0d106f
  • Device/Effective coefficient of variation: approximately 1.5%-1.9%, so the ~4.6% device-side improvement is above the observed run-to-run noise.

On this host, both CPU topology query paths return 65534, so unmodified main and the current PR cannot pass the AICPU affinity gate. For measurement only, the same OCCUPY-only topology fallback was applied to temporary baseline and PR worktrees. It was therefore a shared control variable; the only A/B code difference was the K=16 scheduler publication change. The fallback was not added back to this PR.

@yanghaoran29

Copy link
Copy Markdown
Contributor Author

A2/A3 port benchmark follow-up

I ported PR #1575's A5 scheduler progress-publication batching to A2/A3 locally: publish ring->fc.last_task_alive every 16 local advances, force-publish at current_task_index and on idle pending-advance drain, and reset last_published_to_sm on initialization/reuse.

Benchmark setup:

Workload Baseline Effective (us) Port Effective (us) Change
alternating_matmul_add (Case1) 772.0 765.7 -0.82%
benchmark_bgemm (Case0) 720.3 724.2 +0.54%
paged_attention_unroll (Case1) 1098.7 1094.2 -0.41%
paged_attention_unroll (Case2) 572.1 568.3 -0.66%
paged_attention_unroll_manual_scope (Case1) 1090.1 1084.1 -0.55%
paged_attention_unroll_manual_scope (Case2) 560.3 568.1 +1.39%
batch_paged_attention (Case1) 3565.1 3496.3 -1.93%
qwen3_14b_decode (StressBatch16Seq3500) 35385.3 35557.1 +0.49%

Effective time improved in 5 workloads and regressed in 3. Every change stayed within the +/-2% noise margin; the unweighted mean Effective change was -0.24%. On A2/A3 this establishes no clear performance benefit and no significant regression, unlike the A5 result recorded in this PR.

Validation of the port: targeted test_wiring passed, the complete no-hardware C++ suite passed 99/99, and pre-commit passed.

@yanghaoran29 yanghaoran29 changed the title a5: batch sync_to_sm every K=16 local advances [Performance] A5 TMR: batch progress publication every 16 advances Aug 17, 2026
@yanghaoran29

Copy link
Copy Markdown
Contributor Author

Complete A5 benchmark evidence

Measured on the same Ascend950PR device 0 under a single-device task-submit lock. Each complete pair contains 100 measured rounds for both main and this PR after warm-up. Values are arithmetic means across complete paired groups; negative percentages mean the PR is faster.

Workload Paired groups Host main -> PR (us) Device main -> PR (us) Effective main -> PR (us) Orch main -> PR (us) Sched main -> PR (us)
alternating matmul add Case1 3 162431.3 -> 155732.3 (-4.12%) 1514.2 -> 1453.1 (-4.03%) 1484.6 -> 1423.1 (-4.14%) 1480.1 -> 1418.4 (-4.17%) 1476.4 -> 1414.7 (-4.18%)
batch paged attention Case1 3 178352.4 -> 185378.2 (+3.94%) 7174.6 -> 6569.1 (-8.44%) 7142.5 -> 6536.4 (-8.48%) 6125.8 -> 5357.6 (-12.54%) 7134.4 -> 6528.2 (-8.50%)
benchmark BGEMM Case0 3 54468.7 -> 58915.1 (+8.16%) 1707.9 -> 1617.1 (-5.31%) 1677.6 -> 1586.7 (-5.41%) 1587.2 -> 1493.5 (-5.90%) 1669.0 -> 1578.3 (-5.43%)
paged attention manual Case1 3 179923.5 -> 171354.9 (-4.76%) 1918.9 -> 1898.9 (-1.04%) 1888.5 -> 1868.6 (-1.05%) 1524.1 -> 1408.9 (-7.56%) 1880.2 -> 1859.8 (-1.09%)
paged attention manual Case2 3 39997.3 -> 45469.4 (+13.68%) 1078.1 -> 1076.0 (-0.19%) 1048.8 -> 1047.3 (-0.14%) 622.6 -> 609.1 (-2.16%) 1040.5 -> 1038.6 (-0.19%)
paged attention unroll Case1 3 167619.5 -> 172768.8 (+3.07%) 1959.7 -> 1942.7 (-0.87%) 1929.6 -> 1912.7 (-0.87%) 1643.2 -> 1620.8 (-1.36%) 1921.4 -> 1904.3 (-0.89%)
paged attention unroll Case2 3 40778.4 -> 42082.7 (+3.20%) 1086.3 -> 1068.2 (-1.67%) 1056.5 -> 1039.6 (-1.60%) 750.8 -> 753.9 (+0.41%) 1047.8 -> 1030.8 (-1.62%)
Qwen3 14B decode Stress B16/S3500 2 8282247.9 -> 7818213.1 (-5.60%) 36681.9 -> 36402.8 (-0.76%) 36647.0 -> 36368.1 (-0.76%) 11635.5 -> 11294.6 (-2.93%) 36639.3 -> 36359.7 (-0.76%)

Unweighted mean delta across the eight workloads:

  • Host: +2.20%
  • Device: -2.79%
  • Effective: -2.81%
  • Orch: -4.53%
  • Sched: -2.83%

For Qwen, two complete paired groups are included. A third group completed only the main side before the remaining run was stopped, so that unmatched result is excluded.

For comparison, the portable A2/A3 implementation produced an unweighted mean Effective change of -0.24%, with all eight workload changes within approximately +/-2%. See the A2/A3 benchmark follow-up.

@yanghaoran29 yanghaoran29 reopened this Aug 17, 2026
@yanghaoran29
yanghaoran29 marked this pull request as ready for review August 17, 2026 03:31
@ChaoWao

ChaoWao commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

Review

The optimization is well-motivated and the A5-vs-A2/A3 asymmetry is properly measured — but I think the correctness argument has a hole that turns some previously-working runs into a fatal. Details below.

Stated vs. real goal

They match. The diff does exactly what the body says: adds a last_published_to_sm shadow, gates the release store on Δ ≥ 16 || force, forces on ring-drained and on the deferred-advance drain path, resets the shadow in both init paths, updates the divergence doc, adds three UTs. No scope creep; A2/A3 is genuinely untouched.

Why the blast radius is bigger than 19 lines of core logic

ring->fc.last_task_alive is not a progress statistic — it is the flow-control credit the orchestrator blocks on. Three orchestrator-side waits spin on it, and pto_orchestrator.cpp:519 says so explicitly ("all three share last_task_alive"):

Waiter Blocks on
PTO2TaskAllocator::alloc (pto_ring_buffer.h:141) task slots and heap bytes (via update_heap_tail(last_alive))
PTO2DepListPool::ensure_space (pto_ring_buffer.cpp:188) dep-list pool entries
PTO2FaninPool reclaim wait (pto_ring_buffer.cpp:60) fanin spill pool

Each has the same two detectors: a structural test (is the reclaim head the oldest task pinned by an open scope?) and a ~500 ms wall-clock backstop that latches a fatal.

The load-bearing fact is in PTO2TaskSlotState::reset_for_reuse (pto_runtime2_types.h:592): every task is stamped fanout_count = PTO2_FANOUT_SCOPE_BIT, cleared only by release_producer_scope from on_scope_end — which runs on the orchestrator (pto_orchestrator.cpp:753). So no task of an open scope can become CONSUMED while the orchestrator is blocked in one of those spins. That is exactly why the structural detector exists, and its comment states the invariant: "no open scope can end while this orchestrator is blocked here, so that head cannot become CONSUMED."

Pre-PR, unconditional publication gave the waiters an invariant they silently depend on: the orchestrator sees every slot the scheduler has actually reclaimed. This PR replaces that with two compensating force-publish paths. I believe one of them doesn't work.


Must fix

1. The "idle" force-publish never fires for a withheld batch

pto_scheduler.h:570-572:

bool drain_pending_ring_advances() {
    uint32_t pending = advance_pending_mask.load(std::memory_order_acquire);
    if (pending == 0) return false;      // <-- early-out

advance_pending_mask is set only by mark_ring_advance_pending(), called only from the else branch where advance_lock.compare_exchange_strong failed (pto_scheduler.h:642, :681). A partial batch left behind by a successful advance sets no bit, so the idle path returns immediately and never flushes it. The mask tracks deferred advances, not withheld publications — different things.

So this claim in the body and the doc is not accurate for the idle half:

Force-publish terminal and idle/drained progress so a final partial batch cannot stall slot, heap, or TensorMap reclamation.

The surviving backstop is last_task_alive == current_task_index — the ring must fully drain.

2. A blocked orchestrator can prevent the full-drain backstop it depends on

Combining #1 with the scope-pinning fact above:

  1. Orchestrator submits scope A, ends it; submits scope B, leaves it open.
  2. Orchestrator calls alloc() for the next task; the ring is full as measured against the published watermark → it blocks.
  3. Scope A's tasks complete and consume → last_task_alive advances locally by k.
  4. If 1 ≤ k ≤ 15 and the ring is not fully drained (scope B's tasks are live) → publication withheld.
  5. Scope B cannot end — the orchestrator that would end it is blocked at step 2. No further advance is possible. Neither backstop fires.
  6. ~500 ms later the wall-clock backstop latches and reports FATAL: Task Allocator Deadlock.

Pre-PR, step 3 published immediately, local_task_id_ - last_alive + 1 < window_size_ became true, and the orchestrator proceeded. This converts a working run into a fatal.

Note what does not bound the exposure: alloc() only blocks when the ring is full, and when it blocks the withheld slots are exactly the ones it needs. Window size only sets how easily you reach that state — and PTO2_RING_TASK_WINDOW is validated down to 4 (runtime_maker.cpp:238), with real in-tree configs at 16 and 64 (examples/workers/l3/per_task_runtime_env/main.py:77, examples/a5/.../paged_attention_ringbuffer at 64, tests/st/a5/.../test_prewarm_config.py at 64). At window 16, Δ ≥ 16 can never be reached from within a window's worth of in-flight tasks at all — publication depends entirely on full drain, which serializes submission against execution and inverts the intent of the change.

Green st-onboard-a5 is not evidence against this. CI exercises manual scopes (paged_attention_manual_scope, default 16384 window) and undersized rings (paged_attention_ringbuffer, window 64) — but never both together, which is the conjunction required.

Suggested fix — make batching headroom-aware. The scheduler already reads current_task_index, and task_window_size is in the SM ring header (pto_shared_memory.h:90):

// Publish immediately once the orchestrator is within one batch of the
// window limit: below that headroom the watermark is its allocation credit,
// not a progress statistic.
bool near_full = current_task_index - last_published_to_sm >=
                 ring->task_window_size - PUBLISH_INTERVAL_K;
sync_to_sm(force_publish || near_full || last_task_alive == current_task_index);

This keeps the full win in the regime you measured (deep ring, plenty of headroom — where all 8 benchmarks live) and restores the pre-PR guarantee exactly where liveness depends on it. Alternatives: have the orchestrator set a relaxed waiting flag in SM before it spins and force on that; or drop the pending == 0 early-out — though that last one alone is insufficient, since the scheduler need not be idle while a blocked orchestrator waits on another ring.

3. Both structural deadlock detectors lose their exactness

head_is_oldest_open_task(last_alive, oldest_open_task) (pto_ring_buffer.h:408, plus the inline copy at pto_ring_buffer.cpp:208) compares oldest_open_task against &slot_states_[last_alive & window_mask_] using the published watermark:

  • False negative (common). When the true head is the oldest open-scope task but the published head lags, the pointers differ and the precise detector misses. The failure falls through to the wall-clock backstop and is reported with scope_gated=false — a provable head-of-line deadlock misreported as a reclaim timeout. That distinction is what the repo's triage guidance uses to classify 507018, so this silently degrades the documented signal.
  • False positive (narrower, but on a fatal path). Pre-PR the test was exact: to commit id local_task_id_-1 the admission test required local_task_id_ - T < W, so every live id sat within one window of the true head and id & mask == T & mask ⟺ id == T. With lag L = T - P, task P + W aliases the published head's slot and is live whenever P + W < local_task_id_ < P + W + L — reachable for L ≥ 2. If that aliased task is the oldest open-scope task, the detector latches a spurious "provable" deadlock fatal while reclaim is progressing normally.

Whichever fix lands for #2, these detectors need a head they can trust, and the comments asserting the old reasoning (pto_ring_buffer.h:184-190, pto_ring_buffer.cpp:205-209) should be updated with it.


Should fix

4. Heap reclamation is throttled by the same withheld watermark. alloc() calls update_heap_tail(last_alive) on the published value, so up to 15 tasks' worth of output bytes stay invisible. For large-output tasks that is a lot of bytes, and a heap-tight workload can hit FATAL: Task Allocator Deadlock - Heap Exhausted! where it previously fit. The headroom fix above covers the slot axis but not this one.

5. The Tasks in current scope >= task_window_size guard now checks the wrong threshold. pto_orchestrator.cpp:559 admits any scope with scope_task_count < window_size() - 1. With up to 15 slots invisible the real limit is window_size() - 1 - K, so the guard passes and the run then deadlocks with a less actionable fatal — losing the good diagnostic (FATAL: Scope Deadlock Detected!, which names the cause and prints three remedies).

6. The benchmark set omits the regime this PR changes. All 8 workloads run the default 16384 window, where the ring never fills and back-pressure never engages — which is why they all improved. paged_attention_ringbuffer is the repo's designated "deliberately undersized rings … stress test for rotation and reclamation" and is the example whose behavior this PR most directly alters; it's absent from the table. Please add it (window 64) plus a manual-scope case under a small window, and report them even if they regress — a documented regression in a stress config is a fine trade to state explicitly; an unmeasured one is not.

7. No test covers the interaction. The three new UTs are the right cases for the arithmetic — boundaries 1/15/16/17, drained-tail forcing, and RingReuseResetsPublicationShadow is a good catch. But none exercises a consumer of the watermark. A UT that drives PTO2TaskAllocator::alloc's admission test against a deliberately withheld watermark with an open-scope head should fail on the current diff and pass after the fix.

8. Doc and body wording. Besides the idle-path claim in #1: "The published watermark is a conservative lower bound" holds for validity (TensorMap) but not for credit — for a blocked allocator a lower bound isn't conservative, it's a false negative on available space. Also, the divergence doc's baseline blockquote was changed from a maintenance baseline to "compares main … with the changes introduced by PR #1575"; that re-frames a standing reference doc as a PR artifact, and the PR reference becomes noise once merged. Suggest keeping it a maintenance baseline pinned at the post-merge SHA.


Consider

9. pto_scheduler.h:445 hardcodes "at most 15" while PUBLISH_INTERVAL_K = 16 is declared 20 lines below inside sync_to_sm. Lifting the constant to struct scope and phrasing the comment as PUBLISH_INTERVAL_K - 1 keeps them from drifting. (The comment is otherwise good — a present-tense invariant.)

10. K=16 is unmotivated — no sweep reported. If the win is mostly cache-line traffic, K=4 or 8 may capture most of it at a quarter of the lag, which also shrinks the hazard surface.

11. advance_ring_pointers(true) at the call site is opaque; a two-value enum would read better. Matches surrounding style, so optional.

12. docs/investigations/2026-06-cross-task-batched-publish.md records a different batched publish (AICore MMIO handles) that measured well and then broke spmd_sync_start_stress 2/5 — root-caused months later to a liveness bug that batching only widened. Same shape: delaying a publication another party waits on. A pointer to it in the divergence doc would help the next reader, and this PR's own outcome probably belongs there too.


Things I checked that are fine

Noting these so they don't get re-litigated:

  • TensorMap validity (pto_tensormap.h:698, sync_validity) — a lagging watermark makes producers look live longer. Conservative in the safe direction.
  • STALL diagnostics (scheduler_cold_path.cpp:245) — scanning from the stale tail now covers up to 15 retired ids, but reset_for_reuse leaves task_state == CONSUMED and the scan does if (st >= PTO2_TASK_COMPLETED) continue. No phantom state=WAIT entries.
  • Eager slot reset safety (pto_scheduler.h:500-503) — the argument holds, and strengthens: a slot the orchestrator cannot see cannot be reused.
  • Stale current_task_index in the drained test — loaded before the while loop, so it can only under-report; errs toward publishing.
  • A5-only decision — declining the port on a −0.24% result is right, the topology rationale (A5: 2 AICPU cores/cluster × 5 roles; A2/A3: 4 × 4) is a real mechanism, and updating the divergence doc in the same commit is exactly right.

Verdict

Request changes — on #1/#2/#3, which are one fix.

The performance work is sound and the measurements support it for the regime measured. The problem is that the stated safety argument names two backstops and one of them doesn't fire, leaving full-ring drain as the only one — which a blocked orchestrator can itself prevent, since scope-end only runs on the orchestrator. All eight benchmarks and green onboard CI are consistent with the bug being latent rather than absent: every measured workload uses the 16384 default window and never reaches back-pressure.

Happy to be shown wrong on the reachability of #2 — if there's a path that force-publishes on a blocked allocator that I've missed, that resolves #1 and #2 together.

@yanghaoran29

Copy link
Copy Markdown
Contributor Author

@ChaoWao Thanks for the detailed review. I addressed the selected items as follows:

I did not add #10's K sweep in this correctness fix: K=16 is the already measured configuration, and changing it would require a new paired performance study. I also left #11's optional enum out; the existing bool force_publish is local and consistent with the surrounding API.

Validation is clean: independent C++ suite 100/100, A5 simulator ringbuffer 258/258 checks, both locked A5 small-window runs above, and local pre-commit including clang-format, clang-tidy, cpplint, and markdownlint.

@yanghaoran29

yanghaoran29 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Final implementation A5 performance re-test

This re-test supersedes the earlier performance summary for the current correctness implementation.

Setup:

  • Baseline: the PR merge-base with main; comparison: the current PR head.
  • Platform: Ascend950PR, device 0, with baseline and PR run sequentially inside the same task-submit allocation.
  • PTO-ISA pin: f51c92f610827daad0ddfb383072e03d514b4ae9.
  • Seven ordinary workloads use 100 measured rounds per side and arithmetic means.
  • batch_paged_attention is the mean of four tightly adjacent baseline-to-PR pairs, each with 100 measured rounds per side, because the first final-implementation run showed a regression greater than 5%.
  • As requested, Qwen3 uses exactly five measured runs per side and reports the median of each metric.
  • Negative percentages mean the PR is faster. Effective is the headline metric; Host is included for context and is visibly noisier.
Workload Host main -> PR Device main -> PR Effective main -> PR Orch main -> PR Sched main -> PR
alternating matmul add Case1 141699.4 -> 319254.9 (+125.30%) 1518.0 -> 1439.6 (-5.16%) 1488.4 -> 1409.5 (-5.30%) 1484.1 -> 1406.7 (-5.22%) 1480.3 -> 1401.4 (-5.33%)
benchmark BGEMM Case0 72738.5 -> 68909.6 (-5.26%) 1796.3 -> 1603.3 (-10.74%) 1766.1 -> 1574.4 (-10.85%) 1682.4 -> 1493.2 (-11.25%) 1758.0 -> 1565.2 (-10.97%)
paged attention unroll Case1 187132.3 -> 176474.0 (-5.70%) 2028.9 -> 1858.8 (-8.38%) 1999.0 -> 1828.5 (-8.53%) 1799.9 -> 1397.3 (-22.37%) 1991.0 -> 1820.5 (-8.56%)
paged attention unroll Case2 44183.3 -> 44822.1 (+1.45%) 1107.1 -> 1069.6 (-3.39%) 1078.1 -> 1040.2 (-3.52%) 805.8 -> 690.1 (-14.36%) 1069.4 -> 1029.6 (-3.72%)
paged attention manual Case1 164546.5 -> 161092.9 (-2.10%) 1930.3 -> 1869.3 (-3.16%) 1900.9 -> 1839.8 (-3.21%) 1559.8 -> 1266.2 (-18.82%) 1892.9 -> 1831.6 (-3.24%)
paged attention manual Case2 40455.7 -> 43549.0 (+7.65%) 1091.2 -> 1089.4 (-0.16%) 1062.5 -> 1060.3 (-0.21%) 649.1 -> 578.7 (-10.85%) 1053.5 -> 1051.2 (-0.22%)
batch paged attention Case1 (4-pair mean) 180144.3 -> 186756.8 (+3.67%) 7368.6 -> 7971.8 (+8.19%) 7336.1 -> 7939.3 (+8.22%) 6338.1 -> 7039.3 (+11.06%) 7327.9 -> 7931.1 (+8.23%)
Qwen3 14B decode Stress B16/S3500 (5-run median) 35091338.6 -> 22445701.0 (-36.04%) 36920.4 -> 36900.4 (-0.05%) 36886.4 -> 36860.7 (-0.07%) 11784.7 -> 10120.2 (-14.12%) 36880.8 -> 36857.4 (-0.06%)

Effective improved in 7 of 8 workloads, with an unweighted mean change of -2.93%. Five improvements exceed the +/-2% noise band; manual Case2 and Qwen are effectively flat.

The exception is batch_paged_attention: its four adjacent Effective pairs were 7337.2 -> 7928.7 (+8.06%), 7315.8 -> 7939.1 (+8.52%), 7361.9 -> 7932.5 (+7.75%), and 7329.4 -> 7957.0 (+8.56%). The +8.22% mean regression is therefore reproducible and is reported explicitly rather than retaining the pre-correctness-fix improvement number.

@yanghaoran29

Copy link
Copy Markdown
Contributor Author

Final split-mask Batch Paged Attention Case1 recheck on A5 device 0 (Ascend950PR): four adjacent baseline -> PR groups, 100 measured rounds per side, same task-submit allocation, PTO ISA f51c92f610827daad0ddfb383072e03d514b4ae9.

Group Side Host (us) Device (us) Effective (us) Orch (us) Sched (us)
1 main ad4acaf2 164495.4 7348.4 7315.8 6314.1 7307.2
1 PR 9605f5e1 154180.7 6988.1 6955.7 5898.1 6947.3
2 main ad4acaf2 185362.0 7312.4 7279.7 6276.3 7271.2
2 PR 9605f5e1 193725.5 7144.9 7113.1 6102.2 7105.4
3 main ad4acaf2 151174.2 7334.1 7301.3 6295.2 7293.4
3 PR 9605f5e1 168216.8 7027.2 6994.6 5948.4 6986.6
4 main ad4acaf2 161870.0 7375.7 7343.2 6353.3 7335.7
4 PR 9605f5e1 157826.5 7089.7 7056.3 6016.3 7048.5

Four-group means:

  • Device: 7342.7 -> 7062.5 us (-3.82%)
  • Effective: 7310.0 -> 7029.9 us (-3.83%)
  • Orchestrator: 6309.7 -> 5991.3 us (-5.05%)
  • Scheduler: 7301.9 -> 7022.0 us (-3.83%)

Per-group Effective changes: -4.92%, -2.29%, -4.20%, -3.91%.

The earlier +8.22% regression is gone. Profiling isolated the cause: the combined/eager handshake promoted 36 ordinary allocation-pressure episodes per run into exact publication. With Scheduler contention on advance_pending_mask, Orchestrator pressure on the separate publication_request_mask, and the productive-loop handshake gated by 10 ms of continuous no-progress, Case1 issues zero forced requests while the real blocked-Orchestrator liveness path remains covered by pressure UTs.

Correctness/quality evidence on the final code:

  • Batch Case1 golden comparison: passed on locked A5.
  • Complete independent non-hardware C++ suite: 100/100 passed.
  • Split-mask targeted suites after the final conservative ack invalidation: 4/4 passed.
  • Commit-stage pre-commit: passed.

@yanghaoran29
yanghaoran29 force-pushed the exp/pr906-F3 branch 5 times, most recently from 50f4e22 to 9847d12 Compare August 18, 2026 02:31
@ChaoWao

ChaoWao commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator

Re-review (force-push 50f4e22c, rebased onto ad4acaf2)

The three Must-fixes from the previous round are genuinely resolved. This is a much stronger change: a real request/ack handshake rather than a patch over the symptom, backed by tests that actually exercise a blocked allocator. Remaining items are one doc-consistency defect, two should-fixes, and some hardening.

Status of the previous findings

# Prior finding Status
M1 "Idle" force-publish never fires for a withheld batch ✅ Fixed — dedicated publication_request_mask / publication_ack_mask, serviced by thread 0 from both productive and idle iterations. The old advance_pending_mask keeps its idle-only try-lock-retry semantics, correctly kept separate.
M2 Blocked orchestrator can prevent the full-drain backstop ✅ Fixed — 10 ms no-progress escape wired into all four waiters: alloc (slots and heap), PTO2DepListPool::ensure_space, PTO2FaninPool::ensure_space, ensure_tensormap_capacity.
M3 Structural detectors lose exactness ✅ Fixed, and better than I asked for. Gating on a fresh ack is the right move: an ack can only enable the check when the forced publish was a no-op (a productive ack raises the watermark → the last_alive > prev_last_alive branch fires → watermark_synchronized is cleared → check skipped). So the check runs precisely when published == scheduler-local head, which restores the pre-PR aliasing impossibility, since admission bounds live ids to [T, T+W) around an exact head.
S4 Heap throttled by the same watermark ✅ Fixed — same waiter, update_heap_tail refreshed on ack, dedicated HeapPressurePublishesWithheldProgress test.
S5 Scope-size guard threshold now wrong ✅ Resolved as a consequence — the escape restores the full effective window, so window_size() - 1 is the correct threshold again.
S6 Benchmark set omits the back-pressure regime 🟡 Partly — see should-fix 2.
S7 No test covers the interaction ✅ Fixed — four pressure tests spawn a real blocked allocator thread, plus independent-mask and arena-relocation tests.
S8 Doc/body claims inaccurate ✅ Fixed — the false idle-path claim is gone, the baseline blockquote is a date-verified maintenance baseline again, and the investigation doc records the first design's flaw. One new defect: see should-fix 1.
C9 K constant / hardcoded "15" ✅ Fixed — static constexpr PUBLISH_INTERVAL_K at struct scope; comment derives from it.
C12 Cite the prior batched-publish investigation ✅ Done, with an honest write-up of why the first design was incomplete.

Verification I ran

  • Configured and built tests/ut/cpp; ctest -LE requires_hardware: 100/100 passed.
  • The four new pressure tests are coupled to two real wall-clock constants, so I stress-ran them: --repeat until-fail:15 -j 4 over test_a5_{wiring,fanin_pool,orchestrator_fanin,task_allocator} → 60/60 passed, no flakes.
  • I suspected a defect in the two pool waiters: their ack branch assigns the refreshed watermark straight into prev_last_alive, which hides that advance from the cur > prev branch that re-anchors block_cycle0 — where alloc() avoids this by using a separate local. I wrote a scratch A/B harness (ack-driven vs. unprompted publication, permanently starved pool, 2 s budget) to see whether a progressing waiter could reach the 500 ms backstop. It can't — any subsequent advance re-anchors, so only the single final advance before a genuine stall is affected, and declaring a deadlock there is correct. Not a defect; the asymmetry with alloc() is cosmetic.

Should fix

1. docs/MULTI_RING.md is now misclassified in the divergence doc — by this commit

The PR adds 2 lines to src/a5/.../docs/MULTI_RING.md, so the a2a3/a5 pair is no longer byte-identical, but docs/tensormap-and-ringbuffer-a2a3-vs-a5.md still lists it under "Byte-Identical Files". Measured, not derived:

$ diff -rqs src/a2a3/runtime/tensormap_and_ringbuffer \
            src/a5/runtime/tensormap_and_ringbuffer | grep -c 'are identical'
24
$ cmp -s src/a2a3/runtime/tensormap_and_ringbuffer/docs/MULTI_RING.md \
         src/a5/runtime/tensormap_and_ringbuffer/docs/MULTI_RING.md   # → DIFFERS

The doc claims 25. The whole off-by-one is this one file: 24 identical + 29 differing = 53 matching pairs, which only reconciles with the doc's 25 + 11 + 17 once MULTI_RING.md moves. docs/RUNTIME_LOGIC.md was correctly moved into the functional list for the same reason, so this looks like a simple miss.

Fix: byte-identical 25 → 24, drop MULTI_RING.md from that list, add docs/MULTI_RING.md to the functional list (17 → 18 matching, 19 → 20 total). The doc's own definition of functional covers files that "document differences in runtime behavior", and its new maintenance note — "Recompute the counts and update the affected sections whenever the files or constants described here change" — is exactly the instruction that got missed.

2. Only one workload is measured on the final SHA, and the PR added a shared cache line to thread 0's productive dispatch path

The seven-row table is explicitly from the preceding SHA (thanks for labelling that honestly), which leaves Batch Case1 as the sole final-SHA measurement. That matters more than usual here, because drain_publication_requests() now runs on every thread-0 dispatch iteration:

if (thread_idx == 0 && made_progress && sched_->drain_publication_requests()) {

and it opens with publication_request_mask.load(acquire) — a line the orchestrator writes. It should stay clean in thread 0's cache while no request is outstanding, and −3.83% on Case1 suggests it is. But a PR whose thesis is "remove a Scheduler↔Orchestrator cache line from the hot path" shouldn't leave a newly-added Scheduler↔Orchestrator line in the hot path validated on one workload. Please re-run the seven on the final SHA.

Relatedly: the 10 ms criterion is well justified for Case1 — 36 promotions → 0 is a genuinely good piece of profiling — but the configs where promotions should fire are the small-window ones, and those appear only as correctness passes. A promotion count plus an Effective number for paged_attention_ringbuffer at ring_task_window=64 would close this properly. Credit where due: the locked paged_attention_manual_scope at window 64 run is exactly the manual-scope × small-window conjunction I said CI never covers.

3. The wiring is fail-silent in the unsafe direction

If reclaim_request_mask_ / reclaim_ack_mask_ are ever left null, enabled() returns false, so watermark_synchronized is initialized true and never cleared. That simultaneously (a) disables the liveness escape and (b) lets the structural check run on a lagging watermark — i.e. it silently restores both bugs from the last round — while K-batching stays enabled. The fail-safe direction is inverted.

I checked the current wiring and it is correct: runtime_init_data_from_layout (which nulls via init) precedes runtime_wire_arena_pointers → set_scheduler on the host path; the AICPU path runs wire_arena_pointers then reset_for_reuse, and PTO2OrchestratorState::reset_for_reuse re-wires immediately after task_allocator.init, while both pools' reset_for_reuse preserve the pointers. So this is hardening, not a live bug. But the safe shape is for "escape not wired" to mean "don't batch" (or to latch a fatal at init), rather than "batch without an escape".


Consider

4. Compare task ids, not slot pointers, in the structural check. head_is_oldest_open_task compares oldest_open_task == &slot_states_[head_task_id & window_mask_]. The ack gate makes this exact at the instant of the ack, but the check runs up to 1024 spins later and the scheduler may have advanced locally by δ without publishing in that window; for δ ≥ 2, task P + W is live and aliases the published head's slot. It additionally needs the oldest open-scope task to sit exactly there, so it's a very narrow coincidence — but it's on a fatal path, and comparing the local task id closes it unconditionally.

5. ReclaimPublicationRequest computes 1u << ring_id directly, bypassing PTO2SchedulerState::ring_advance_pending_bit() and its static_assert(PTO2_MAX_RING_DEPTH <= 32). Same bit derivation in two places, only one guarded.

6. ensure_tensormap_capacity hand-rolls the handshake with a lambda over all 32 ring bits instead of reusing ReclaimPublicationRequest, and never consumes the ack. That is correct — it has no structural head check, so it only needs the watermark to move — but it reads as an oversight next to the other three call sites. A one-line comment ("no ack needed: no structural head check here") would settle it.

7. Document why a single servicer suffices. Thread 0 services from both loop branches, so the common case is covered. But thread 0 is not in the loop while inside a handle_drain_mode spin, which delays servicing; that resolves on its own (drain waits on cores freeing, independent of the orchestrator) and the 500 ms backstop leaves ~490 ms of margin. Worth stating at the call site, since "why thread 0 only" is the first question a reader will have.

8. Genuine open-scope deadlock classification is now ≥10 ms slower (was ~1024 spins). Fine for a fatal path, but docs/troubleshooting/device-error-codes.md and the running-onboard.md triage table describe the structural-vs-timeout distinction, so "A5 takes ≥10 ms to reach the structural verdict" is now part of that story.

9. The four pressure UTs depend on two real wall-clock constants. 15× clean at -j 4 here, and the 490 ms margin plus the latched-request design make them robust. Still, making PTO2_PUBLICATION_REQUEST_TIMEOUT_CYCLES injectable would decouple them from runner load and cut ~40 ms of hot spinning from the suite.


Verdict

Approve, conditional on st-onboard-a5 going green on 50f4e22c (CI was still mid-flight when I looked — pre-commit pending), with should-fix 1 requested before merge: it's a two-line doc correction and the count is verifiably wrong.

Should-fix 2 and 3 are worth doing but neither affects the correctness of what's here. The design is now the right shape: batching on the non-blocking path, an explicit receiver-visible escape when a consumer is actually starved, separate masks so scheduler lock contention cannot forge an acknowledgment, and structural classification gated on a watermark that is provably exact. The investigation-doc entry generalizing the lesson — a batched publication needs a receiver-visible escape whenever the receiver depends on it for forward progress — is the most valuable artifact in the diff.

@yanghaoran29

Copy link
Copy Markdown
Contributor Author

Implemented the requested follow-up and rebased the single PR commit onto current upstream/main.

  • Fixed the divergence inventory: 24 byte-identical files; 18 matching functional paths plus 2 A5-only paths (20 total), with docs/MULTI_RING.md moved to the functional list.
  • Made batching fail-safe: every ring starts with per-advance publication. K=16 is enabled only after the task allocator, fanin pool, and dependency pool are all verified against the current scheduler request/ack masks. Arena relocation and runtime reuse disable batching first and re-enable it only after the same validation.
  • Replaced slot-pointer structural predicates with an exact mixed task-id comparison (ring + local id), shared by task allocation and dependency-pool classification, including a slot-alias regression test.
  • Added one guarded ring_mask_bit() helper and reused it for both publication requests and scheduler masks.
  • Documented why TensorMap does not consume an ack, why thread 0 is sufficient even though drain mode can delay it, and that A5 structural classification requires at least 10 ms plus an exact-publication ack.
  • Updated the A5 runtime docs to state the incomplete-wiring per-advance fallback.
  • Kept the timeout constant non-injectable (refactor: Convert all variable/function names to snake_case #9): changing it would broaden the configuration/test interface for no production correctness gain; the existing real-time pressure tests remain stable.

Validation:

  • Targeted A5/A2A3 C++ regressions: 5/5 passed.
  • Complete non-hardware C++ suite: 100/100 passed.
  • Pressure tests (wiring, fanin_pool, orchestrator_fanin, task_allocator): 60/60 passed with until-fail:15 -j4.
  • A5sim small-window prewarm config: passed.
  • Repository-wide pre-commit: passed.

The final-SHA A5 hardware correctness/performance rerun is still blocked locally by the mandatory architecture gate: inside a single-device task-submit allocation, npu-smi returned neither Chip Name nor NPU Name, so onboard-arch-precheck correctly refused the run. I did not bypass that safety check. The new push has started CI, including the A5 hardware lane once pre-commit completes.

@yanghaoran29
yanghaoran29 force-pushed the exp/pr906-F3 branch 2 times, most recently from 4863559 to f66f976 Compare August 18, 2026 07:58
Batch last_task_alive publication every 16 local advances on the non-blocking path. Keep scheduler try-lock retries on their original idle-only deferred mask, and isolate orchestrator publication requests and acknowledgments on a separate cache line.

Promote task, heap, dependency, fanin, and TensorMap pressure to the productive-loop handshake only after 10 ms without reclaim progress. Service those requests from scheduler thread 0, require a fresh acknowledgment before structural deadlock classification, and preserve the 500 ms wall-clock backstop.

Enable batching only after allocator, fanin, and dependency-pool request/ack wiring is validated against the current scheduler. Default initialization, arena relocation, and reuse to per-advance publication until validation succeeds.

Compare exact ring-local task identities during structural classification, share the guarded ring-bit derivation, document the liveness and diagnostic contracts, and cover batching fallback, relocation, reuse, pressure, and slot-alias cases.
@ChaoWao

ChaoWao commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator

I pushed the two mechanical fixes from my review directly (f66f9765 → 357d1307). No functional code changed — every file that remains in the diff is byte-identical to f66f9765, the SHA that went green across the full matrix.

1. Dropped three files unrelated to this change

.claude/lib/github/branch-naming.md      ordered-list renumbering
.claude/skills/github-pr/SKILL.md        blank lines around fences/lists
src/common/platform_comm/comm_sim.cpp    std::lock_guard → std::scoped_lock ×3

Restored from the merge-base a59ffde7 (not from upstream/main, which has since moved to f4ed1045 — restoring from there would have pulled unrelated newer content into the branch). I also removed the commit-message paragraph that described them ("Keep the repository-wide lint baseline clean…"), since it no longer refers to anything in the diff.

CI's pre-commit runs with --from-ref <merge-base> --to-ref <head>, so these files leave the linted set entirely rather than reverting to a state that would fail lint. The scoped_lock substitution is behavior-identical for a single mutex, so this is purely about scope — if you want that cleanup it stands fine as its own one-line PR.

Correction to my review: I claimed dropping these would return the a2a3 jobs to skipping. That was wrong, and I checked it by replaying the gate locally rather than re-reading the regex:

$ git diff a59ffde7...HEAD --name-only \
    | grep -vE '^(src/a5/|examples/a5/|tests/(st|ut/cpp)/a5/)' \
    | grep -vE '<NON_CODE>'
tests/ut/cpp/common/test_scope_deadlock_detection.cpp

That shared fixture had to change for the exact-identity head check, it belongs to no arch partition, and it is compiled against both runtimes — so it correctly flips a2a3_changed and the a2a3 jobs still run. That is the fail-safe direction working as designed, and it means the CI-cost half of my objection didn't apply. The discipline.md §2 scope argument, and the conflict with this PR's own "does not change A2/A3 runtime behavior" statement, still do.

2. Rewrote the PR body

It was two rebases stale (still citing ad4acaf2) and never described the two strongest parts of the change. The new body:

  • cites the correct base a59ffde7;
  • documents the self-disabling batching gate and the exact ring-local identity comparison under Correctness;
  • states why last_task_alive is allocation credit rather than a progress statistic, since that is the premise the whole design rests on;
  • separates Validation into what was measured on the current head, on the preceding head, and on earlier heads, instead of presenting them as one aggregate;
  • adds a Measurement provenance paragraph stating plainly that every perf number predates the wiring-validation gate and the identity check, why both are expected to be steady-state-neutral, and that the expectation is unmeasured.

I did not invent numbers for the current head. Re-running the Case1 table on 357d1307 is the one thing still outstanding, and it's yours to run — I have no a5 silicon on this box.

I deliberately did not rebase onto the newer upstream/main. The PR reports MERGEABLE, and rebasing would pull unrelated changes into an A5-only diff and discard the CI run this push started.

Two notes on what I chose not to touch: I left the title alone (the commit subject didn't change), and I added no Co-Authored-By trailer — the commit is yours, the amendment is a three-file revert, and the recent history here doesn't carry that trailer.

My approval from the previous comment stands, now with both should-fixes closed.

@ChaoWao
ChaoWao merged commit be25a3b into hw-native-sys:main Aug 18, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants