Skip to content

Add: opt-in async-DMA SDMA workspace via Worker enable_sdma - #1406

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
lyfne123:feat-sdma-prefetch-workspace-addr
Jul 22, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
lyfne123:feat-sdma-prefetch-workspace-addr

Conversation

@lyfne123

@lyfne123 lyfne123 commented Jul 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Provision the PTO-ISA async-SDMA workspace only when a Worker is created with enable_sdma=True, replacing the earlier demand-driven design.

  • Worker(level=2, …, enable_sdma=True) → ChipWorker.init(dma_workspace_mask) → simpler_provision_dma_workspace → dma_workspace_provision. Provisions the 48 STARS streams once at init, latches the addresses into InitArgs, and re-latches simpler_aicpu_init so g_dma_workspace_addr feeds every core's GlobalContext (get_dma_workspace).
  • A Worker without the flag creates no SDMA streams; its kernels read a zero workspace address. a5/sim/host_build_graph have no SDMA provider (dma_workspace_supported_mask()==0), so enable_sdma there fails fast — the shared plumbing still compiles for every arch.

Why (issue #1425)

The 48 device-only STARS streams sit in the device fault/sync domain, so a fault on an SDMA-provisioned device pays a multi-minute CANN teardown. Making SDMA an explicit per-Worker opt-in confines that risk to Workers that ask for it; ordinary Workers keep main's fast (~0.3s) recovery and ordinary 507046 error code.

What changed vs the demand-driven design

  • Replace the comm run-gate (dma_workspace_run_begin/run_end/reset_begin/reset_end/abandon + DmaWorkspaceState) with dma_workspace_provision / dma_workspace_release.
  • Add simpler_provision_dma_workspace to the runtime C ABI.
  • Drop the callable required_dma_workspaces declaration and the requires_device_generation_recycle pool signal; revert arch finalize() and chip_worker to their upstream teardown.

CI isolation

Both SDMA demos run in a dedicated task-submit task, excluded from the general a2a3 sweep — so their teardown never shares a device with the aicore_op_timeout fault-injection test.

Testing

  • Clean rebuild of all 4 platforms (a2a3sim/a5sim/a2a3/a5)
  • tools/verify_packaging.sh (clean wheel build + packaging) passes
  • 60 cpp UT, 726 non-hardware Python UT
  • Hardware (a2a3): prefetch demo provisions 48 STARS streams at init, injected workspace reaches the kernel, output bit-exact

Mitigates #1425. Faults after genuine SDMA use still require a CANN runtime/driver teardown fix.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the sdma_prefetch_workspace_addr function and its corresponding bindings across the codebase, including C++ platform implementations, Python bindings, and worker interfaces. This function retrieves the device address of the process-global PTO-ISA async-SDMA scratch workspace, which is used by single-card kernels via pto.tprefetch_async. There are no review comments provided, so we have no feedback to offer.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@coderabbitai

coderabbitai Bot commented Jul 20, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 25bc9b52-8ac8-4d8d-b0d1-8bdaa4da08d5

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds an async-SDMA workspace address API across platform backends, ChipWorker runtime forwarding, nanobind bindings, and Python Worker interfaces. The API returns zero when the workspace is unavailable and supports level-2 workers.

Changes

Async SDMA workspace address

Layer / File(s) Summary
Backend workspace API
src/common/platform_comm/comm.h, src/common/platform_comm/comm_sim.cpp, src/a2a3/platform/onboard/host/comm_hccl.cpp, src/a5/platform/onboard/host/comm_hccl.cpp
Defines the backend-neutral function and implements workspace provisioning, unavailable-as-zero behavior, and null-pointer validation across supported backends.
ChipWorker runtime bridge
src/common/worker/chip_worker.h, src/common/worker/chip_worker.cpp
Adds runtime symbol resolution, cleanup, and address retrieval through the new host-runtime API.
Python API exposure
python/simpler/task_interface.py, python/bindings/task_interface.cpp, python/simpler/worker.py
Exposes the address through Python ChipWorker and level-2 Worker methods, with unsupported levels raising NotImplementedError.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Worker as Python Worker
  participant ChipWorker
  participant RuntimeAPI as sdma_prefetch_workspace_addr
  participant WorkspaceManager as SdmaWorkspaceManager
  Worker->>ChipWorker: sdma_prefetch_workspace_addr()
  ChipWorker->>RuntimeAPI: retrieve workspace address
  RuntimeAPI->>WorkspaceManager: initialize workspace
  WorkspaceManager-->>RuntimeAPI: address or zero
  RuntimeAPI-->>ChipWorker: status and address
  ChipWorker-->>Worker: return address
Loading

Possibly related PRs

Poem

A bunny hops where scratch bytes gleam,
Fetching an address from an async dream.
Zero means “not here,” no need to fear,
Level two carries the workspace near.
PTO-ISA kernels thump and play—
Safe little pointers lead the way.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 10.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title matches the main change by calling out the opt-in SDMA workspace and the Worker enable_sdma path.
Description check ✅ Passed The description is clearly about the same SDMA workspace provisioning and Python/worker plumbing changes in the patch.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@lyfne123

Copy link
Copy Markdown
Contributor Author

Added a single-card ST so the new accessor is covered inside this repo, not only via the PyPTO consumer.

examples/a2a3/tensormap_and_ringbuffer/prefetch_async_demo/ — kernel + orchestration + test:

  • The host provisions the workspace and threads its device address into the kernel as a scalar, the same way sdma_async_completion_demo threads its CommContext pointer. That avoids needing a child_memory tensor path: the workspace is a runtime-owned device buffer, not user data to stage H2D.
  • The kernel prefetches in into L2, waits on the returned event, then copies in to out through a tile. Since the prefetch changes no value, bit-exact out == in is the property under test — plus the wait actually completing rather than hanging.
  • Needs no comm domain (unlike the SDMA completion demo), because the accessor is comm-independent. One device is enough.
  • Mirrors the pto-isa single-card ST at tests/npu/a2a3/src/st/testcase/tprefetch_async.
  • When the address is 0 the kernel skips the prefetch and still runs the copy, so the demo stays green — with reduced coverage — where SDMA is unavailable.

Verified on a2a3 hardware that it exercises the real SDMA path, not the skip branch:

[SDMA] Created 48 STARS streams OK
[SDMA] STARS query AICPU kernel completed OK
[prefetch_async_demo] sdma workspace addr = 0x12c0c001a000
[INFO] prefetch_async_demo: out matches src after prefetch + copy
1 passed

a5 is intentionally not covered: its SDMA overlay is gated off by default (SIMPLER_ENABLE_PTO_SDMA_WORKSPACE, see docs/a5-sdma-overlay.md), so the demo would only ever take the skip branch there.

@ChaoWao
ChaoWao force-pushed the feat-sdma-prefetch-workspace-addr branch from 303c9d0 to 6300fef Compare July 21, 2026 04:32
@ChaoWao ChaoWao changed the title feat(sdma): expose the async-SDMA prefetch workspace address feat(sdma): runtime-inject async-DMA workspace (engine/domain split) Jul 21, 2026
@ChaoWao

ChaoWao commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

/gemini review

@gemini-code-assist

Copy link
Copy Markdown

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@ChaoWao
ChaoWao force-pushed the feat-sdma-prefetch-workspace-addr branch from 6300fef to c8004a1 Compare July 21, 2026 15:07
@ChaoWao ChaoWao changed the title feat(sdma): runtime-inject async-DMA workspace (engine/domain split) Add: demand-driven async-DMA workspace injection Jul 21, 2026
@ChaoWao
ChaoWao force-pushed the feat-sdma-prefetch-workspace-addr branch 5 times, most recently from 5ef139c to 80b277f Compare July 22, 2026 09:39
@ChaoWao ChaoWao changed the title Add: demand-driven async-DMA workspace injection Add: opt-in async-DMA SDMA workspace via Worker enable_sdma Jul 22, 2026
@ChaoWao
ChaoWao force-pushed the feat-sdma-prefetch-workspace-addr branch 2 times, most recently from d99bebd to 0504a2d Compare July 22, 2026 13:47
Provision the PTO-ISA async-SDMA workspace only when a Worker is created
with enable_sdma=True. The runtime provisions the 48 STARS streams once at
init, latches their device addresses into InitArgs and re-latches the
resident simpler_aicpu_init, so g_dma_workspace_addr feeds every core's
GlobalContext (get_dma_workspace). A Worker without the flag creates no SDMA
streams and its kernels read a zero workspace address. On an L2 Worker the
flag reaches ChipWorker.init directly; on an L3 Worker it is threaded across
the fork to each chip child so every device provisions its own workspace.

Because those streams sit in the device fault/sync domain, a fault on an
SDMA-provisioned device pays a multi-minute CANN teardown (hw-native-sys#1425). Making
SDMA an explicit per-Worker opt-in confines that risk to Workers that ask
for it; ordinary Workers keep main's fast recovery. This replaces the prior
demand-driven design (auto-provision on the first callable declaring an SDMA
mask, plus a device-generation quarantine/reset fence):

- Replace the comm run-gate (dma_workspace_run_begin/run_end/reset_begin/
  reset_end/abandon + DmaWorkspaceState) with dma_workspace_provision/
  dma_workspace_release; the provider handle is owned by the runner and
  released by finalize_common on ordinary teardown.
- Add simpler_provision_dma_workspace to the runtime C ABI; ChipWorker.init
  gains a dma_workspace_mask derived from the Python enable_sdma bool.
- Drop the callable required_dma_workspaces declaration (API, binding,
  artifact, runner) and the requires_device_generation_recycle pool signal;
  revert arch finalize() and chip_worker to their upstream teardown.

a5 has no SDMA provider (dma_workspace_supported_mask returns 0), so
enable_sdma fails fast on a5/sim/host_build_graph; the shared plumbing still
compiles for every arch.

CI runs the SDMA demos (prefetch_async_demo, sdma_async_completion_demo) in a
dedicated task-submit "SDMA pytest" task, excluded from the general sweep so
their teardown never mixes with the aicore_op_timeout fault-injection test.

Validated on a2a3 hardware: the L2 prefetch demo and the 2-device L3
sdma_async_completion demo both provision the workspace, the injected address
reaches the kernel, and output is bit-exact.
@ChaoWao
ChaoWao merged commit 717d703 into hw-native-sys:main Jul 22, 2026
16 checks passed
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 31, 2026
The SDMA quarantine was a pair of `--ignore=<path>` arguments. `pytest
--ignore` on a path that no longer exists **exits 0 and says nothing** —
measured — so moving either demo directory would have dropped the quarantine
silently and landed both tests back beside the `aicore_op_timeout`
fault-injection test. That collision is the whole hazard: provisioning the
SDMA workspace creates 48 device-only STARS streams that sit in the device
fault domain, so a later AICore fault on that device costs ~306 s instead of
~0.3 s (hw-native-sys#1425, contained in hw-native-sys#1406).

`@pytest.mark.sdma` replaces the path matching, and carries the whole
consequence rather than one arbitrary piece of it. One declaration now drives
three things:

- **The Worker is built with `enable_sdma=True`.** Both pytest construction
  sites in `conftest.py` and both standalone sites in `scene_test.py` read it,
  the latter from `cls.pytestmark` so `python test_x.py` behaves like the
  pytest path.
- **The L2 Worker pool stops mixing capabilities.** The key gains the flag and
  reuse tests it, because an enable_sdma Worker holds its STARS streams for
  life: handing it to a test that did not ask for them would spread the
  teardown hazard to every later L2 case on that device. The existing
  same-device retire loop then performs the swap, so no new teardown path.
- **SDMA sorts last.** `sort_key` gains a term, keyed off the marker rather
  than the class because the fault-injection tests are plain functions with no
  `_st_level`. Fault injection therefore always runs on a device that has never
  provisioned — which is what actually addresses the interaction, rather than
  merely quarantining it.

With the capability expressible, `prefetch_async_demo` becomes an ordinary L2
`@scene_test` class. It was a hand-rolled Worker only because `CASES` had no
channel to `Worker.__init__`, and it paid for that by forfeiting golden
comparison, case parametrization, `--rounds`, `--case` and the dispatcher's
device allocation. 160 lines become 89, and its verification —
`torch.equal(out, src)` — is now a one-line `compute_golden`. The framework's
default orchestration includes turned out to be sufficient, so no new
passthrough was needed there.

`sdma_async_completion_demo` keeps its hand-rolled L3 Worker for now and takes
the marker for ordering and CI selection only; converting it means expressing a
comm domain through `CASES`, which is a larger change and independent of this
one.

CI selects with `-m sdma` / `-m "not sdma"`, so the sweep and the dedicated
step can no longer drift apart. **The dedicated step stays until hw-native-sys#1425 is
fixed**: in-session ordering separates the cases, but a fault on a device that
has already provisioned still costs minutes, so they must not share a device.

Separately, `docs/capability-survey.md` carried four `ci.yml:<line>` citations,
three already invalidated by this session's edits — `:607` lands on a blank
line, `:884` on a `cmake --build`, `:628-643` on a dep_gen comment, after the
qwen step was deleted in hw-native-sys#1601 and `detect-changes` reworked in hw-native-sys#1589 / hw-native-sys#1590 /
hw-native-sys#1607. Line numbers into a file this active cannot be maintained, so all four
become greppable anchors. `grep -rIn 'ci\.yml:[0-9]'` is now empty repo-wide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 31, 2026
The SDMA quarantine was a pair of `--ignore=<path>` arguments. `pytest
--ignore` on a path that no longer exists **exits 0 and says nothing** —
measured — so moving either demo directory would have dropped the quarantine
silently and landed both tests back beside the `aicore_op_timeout`
fault-injection test. That collision is the whole hazard: provisioning the
SDMA workspace creates 48 device-only STARS streams that sit in the device
fault domain, so a later AICore fault on that device costs ~306 s instead of
~0.3 s (hw-native-sys#1425, contained in hw-native-sys#1406).

`@pytest.mark.sdma` replaces the path matching, and carries the whole
consequence rather than one arbitrary piece of it. One declaration now drives
three things:

- **The Worker is built with `enable_sdma=True`.** Both pytest construction
  sites in `conftest.py` and both standalone sites in `scene_test.py` read it,
  the latter from `cls.pytestmark` so `python test_x.py` behaves like the
  pytest path.
- **The L2 Worker pool stops mixing capabilities.** The key gains the flag and
  reuse tests it, because an enable_sdma Worker holds its STARS streams for
  life: handing it to a test that did not ask for them would spread the
  teardown hazard to every later L2 case on that device. The existing
  same-device retire loop then performs the swap, so no new teardown path.
- **SDMA sorts last.** `sort_key` gains a term, keyed off the marker rather
  than the class because the fault-injection tests are plain functions with no
  `_st_level`. This is what will make merging the dedicated CI step back into
  the sweep safe once hw-native-sys#1425 is fixed; today `-m` already separates them, so it
  matters for local full runs.

With the capability expressible, `prefetch_async_demo` becomes an ordinary L2
`@scene_test` class. It was a hand-rolled Worker only because `CASES` had no
channel to `Worker.__init__`, and it paid for that by forfeiting golden
comparison, case parametrization, `--rounds`, `--case` and the dispatcher's
device allocation. 160 lines become 89, and its verification —
`torch.equal(out, src)` — is now a one-line `compute_golden`. The framework's
default orchestration includes turned out to be sufficient, so no new
passthrough was needed there. The conversion also moves it from the Resource
phase to the L2 phase, which runs after it, so it provisions on a device the
fault-injection cases have already finished with.

`sdma_async_completion_demo` keeps its hand-rolled L3 Worker for now and takes
the marker for ordering and CI selection only; converting it means expressing a
comm domain through `CASES`, which is a larger change and independent of this
one.

CI selects with `-m sdma` / `-m "not sdma"`, so the sweep and the dedicated
step can no longer drift apart. **The dedicated step stays until hw-native-sys#1425 is
fixed**: the two phases run their jobs in parallel across devices, so only
separate device pools keep an SDMA provisioning away from a fault injection.

`.claude/skills/testing/SKILL.md` moves with the mechanism throughout — not
only the mirror-CI command, but the two instructions that told readers to
extract `--ignore` sets from `ci.yml` and to grep it when a test passes alone
and fails in the sweep. Both now name the marker; following the skill
reproduces CI rather than a mechanism that no longer exists.

Separately, `docs/capability-survey.md` carried four `ci.yml:<line>` citations,
three already invalidated by this session's edits — `:607` lands on a blank
line, `:884` on a `cmake --build`, `:628-643` on a dep_gen comment, after the
qwen step was deleted in hw-native-sys#1601 and `detect-changes` reworked in hw-native-sys#1589 / hw-native-sys#1590 /
hw-native-sys#1607. Line numbers into a file this active cannot be maintained, so all four
become greppable anchors. `grep -rIn 'ci\.yml:[0-9]'` is now empty repo-wide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 31, 2026
The SDMA quarantine was a pair of `--ignore=<path>` arguments. `pytest
--ignore` on a path that no longer exists **exits 0 and says nothing** —
measured — so moving either demo directory would have dropped the quarantine
silently and landed both tests back beside the `aicore_op_timeout`
fault-injection test. That collision is the whole hazard: provisioning the
SDMA workspace creates 48 device-only STARS streams that sit in the device
fault domain, so a later AICore fault on that device costs ~306 s instead of
~0.3 s (hw-native-sys#1425, contained in hw-native-sys#1406).

`@pytest.mark.sdma` replaces the path matching, and carries the whole
consequence rather than one arbitrary piece of it. One declaration now drives
three things:

- **The Worker is built with `enable_sdma=True`.** Both pytest construction
  sites in `conftest.py` and both standalone sites in `scene_test.py` read it,
  the latter from `cls.pytestmark` so `python test_x.py` behaves like the
  pytest path.
- **The L2 Worker pool stops mixing capabilities.** The key gains the flag and
  reuse tests it, because an enable_sdma Worker holds its STARS streams for
  life: handing it to a test that did not ask for them would spread the
  teardown hazard to every later L2 case on that device. The existing
  same-device retire loop then performs the swap, so no new teardown path.
- **SDMA sorts last.** `sort_key` gains a term, keyed off the marker rather
  than the class because the fault-injection tests are plain functions with no
  `_st_level`. This is what will make merging the dedicated CI step back into
  the sweep safe once hw-native-sys#1425 is fixed; today `-m` already separates them, so it
  matters for local full runs.

With the capability expressible, `prefetch_async_demo` becomes an ordinary L2
`@scene_test` class. It was a hand-rolled Worker only because `CASES` had no
channel to `Worker.__init__`, and it paid for that by forfeiting golden
comparison, case parametrization, `--rounds`, `--case` and the dispatcher's
device allocation. 160 lines become 89, and its verification —
`torch.equal(out, src)` — is now a one-line `compute_golden`. The framework's
default orchestration includes turned out to be sufficient, so no new
passthrough was needed there. The conversion also moves it from the Resource
phase to the L2 phase, which runs after it, so it provisions on a device the
fault-injection cases have already finished with.

`sdma_async_completion_demo` keeps its hand-rolled L3 Worker for now and takes
the marker for ordering and CI selection only; converting it means expressing a
comm domain through `CASES`, which is a larger change and independent of this
one.

CI selects with `-m sdma` / `-m "not sdma"`, so the sweep and the dedicated
step can no longer drift apart. **The dedicated step stays until hw-native-sys#1425 is
fixed**: the two phases run their jobs in parallel across devices, so only
separate device pools keep an SDMA provisioning away from a fault injection.

`.claude/skills/testing/SKILL.md` moves with the mechanism throughout — not
only the mirror-CI command, but the two instructions that told readers to
extract `--ignore` sets from `ci.yml` and to grep it when a test passes alone
and fails in the sweep. Both now name the marker; following the skill
reproduces CI rather than a mechanism that no longer exists.

Separately, `docs/capability-survey.md` carried four `ci.yml:<line>` citations,
three already invalidated by this session's edits — `:607` lands on a blank
line, `:884` on a `cmake --build`, `:628-643` on a dep_gen comment, after the
qwen step was deleted in hw-native-sys#1601 and `detect-changes` reworked in hw-native-sys#1589 / hw-native-sys#1590 /
hw-native-sys#1607. Line numbers into a file this active cannot be maintained, so all four
become greppable anchors. `grep -rIn 'ci\.yml:[0-9]'` is now empty repo-wide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ChaoWao added a commit that referenced this pull request Jul 31, 2026
…#1609)

The SDMA quarantine was a pair of `--ignore=<path>` arguments. `pytest
--ignore` on a path that no longer exists **exits 0 and says nothing** —
measured — so moving either demo directory would have dropped the quarantine
silently and landed both tests back beside the `aicore_op_timeout`
fault-injection test. That collision is the whole hazard: provisioning the
SDMA workspace creates 48 device-only STARS streams that sit in the device
fault domain, so a later AICore fault on that device costs ~306 s instead of
~0.3 s (#1425, contained in #1406).

`@pytest.mark.sdma` replaces the path matching, and carries the whole
consequence rather than one arbitrary piece of it. One declaration now drives
three things:

- **The Worker is built with `enable_sdma=True`.** Both pytest construction
  sites in `conftest.py` and both standalone sites in `scene_test.py` read it,
  the latter from `cls.pytestmark` so `python test_x.py` behaves like the
  pytest path.
- **The L2 Worker pool stops mixing capabilities.** The key gains the flag and
  reuse tests it, because an enable_sdma Worker holds its STARS streams for
  life: handing it to a test that did not ask for them would spread the
  teardown hazard to every later L2 case on that device. The existing
  same-device retire loop then performs the swap, so no new teardown path.
- **SDMA sorts last.** `sort_key` gains a term, keyed off the marker rather
  than the class because the fault-injection tests are plain functions with no
  `_st_level`. This is what will make merging the dedicated CI step back into
  the sweep safe once #1425 is fixed; today `-m` already separates them, so it
  matters for local full runs.

With the capability expressible, `prefetch_async_demo` becomes an ordinary L2
`@scene_test` class. It was a hand-rolled Worker only because `CASES` had no
channel to `Worker.__init__`, and it paid for that by forfeiting golden
comparison, case parametrization, `--rounds`, `--case` and the dispatcher's
device allocation. 160 lines become 89, and its verification —
`torch.equal(out, src)` — is now a one-line `compute_golden`. The framework's
default orchestration includes turned out to be sufficient, so no new
passthrough was needed there. The conversion also moves it from the Resource
phase to the L2 phase, which runs after it, so it provisions on a device the
fault-injection cases have already finished with.

`sdma_async_completion_demo` keeps its hand-rolled L3 Worker for now and takes
the marker for ordering and CI selection only; converting it means expressing a
comm domain through `CASES`, which is a larger change and independent of this
one.

CI selects with `-m sdma` / `-m "not sdma"`, so the sweep and the dedicated
step can no longer drift apart. **The dedicated step stays until #1425 is
fixed**: the two phases run their jobs in parallel across devices, so only
separate device pools keep an SDMA provisioning away from a fault injection.

`.claude/skills/testing/SKILL.md` moves with the mechanism throughout — not
only the mirror-CI command, but the two instructions that told readers to
extract `--ignore` sets from `ci.yml` and to grep it when a test passes alone
and fails in the sweep. Both now name the marker; following the skill
reproduces CI rather than a mechanism that no longer exists.

Separately, `docs/capability-survey.md` carried four `ci.yml:<line>` citations,
three already invalidated by this session's edits — `:607` lands on a blank
line, `:884` on a `cmake --build`, `:628-643` on a dep_gen comment, after the
qwen step was deleted in #1601 and `detect-changes` reworked in #1589 / #1590 /
#1607. Line numbers into a file this active cannot be maintained, so all four
become greppable anchors. `grep -rIn 'ci\.yml:[0-9]'` is now empty repo-wide.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants