Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions docs/buffer-abi.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ collapse that into a single type were considered and dropped.

### Rejected: merge `Tensor` into `ChipTensor`

Drop `buffer.addr`, add the buffer descriptor, and have the H2D staging step
Drop `buffer.addr`, add the buffer descriptor, and have the H2D copy-in step
rewrite the backend tag and body (and mint a fresh identity for the device copy).
That is self-consistent, but it charges the device for host-side fields:

Expand Down Expand Up @@ -132,8 +132,8 @@ task.
> OverlapMap by it rather than by `buffer.addr` would make two views of one
> backing bucket together by construction. That needs 32 B — which fits the
> existing `_pad_cl2[36]` at `sizeof == 128`, i.e. **without** merging anything
> else. If it is ever done, the H2D staging step must mint a *new* identity for
> each staged copy, because the device buffer is a distinct backing from the host
> else. If it is ever done, the H2D copy-in step must mint a *new* identity for
> each copy, because the device buffer is a distinct backing from the host
> one it was copied from.

### Rejected: keep the wire type transport-only, use `ChipTensor` in the L3 orch
Expand Down Expand Up @@ -337,7 +337,7 @@ values are final, at submit:
named as an output and then silently losing every write in the child.
- **No overlapping writes within one task.** Two arguments of one task that name
intersecting bytes of the same backing are rejected: they belong to one node,
so there is no order between them to express, and a device-staged copy of a
so there is no order between them to express, and a device-side copy of a
host backing does not even alias on the device for the L2 overlap map to
notice. Disjoint slices of one buffer stay legal — that is what `byte_offset`
is for, and this check runs the same two-stage comparison dependency
Expand Down
20 changes: 10 additions & 10 deletions docs/dfx/hbg-bind-phases.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# The `host_build_graph` bind phases

`host_build_graph` builds the whole task graph on the host before the device
executes anything, so the host-side **`bind` stage** — argument staging,
executes anything, so the host-side **`bind` stage** — argument copy-in,
orchestration, the Graph Definition, and every H2D copy — is a first-class cost.
`bind` is the `chip.run.bind` `[STRACE]` span both runtimes emit; only this one
subdivides it into **segments**, one `chip.run.bind.<segment>` span each. This
Expand All @@ -26,7 +26,7 @@ the `chip.run.bind` span:

| Segment | What it covers |
| ------- | -------------- |
| `args` | staging readable caller tensors H2D and exposing their existing host buffers to orchestration; pure outputs skip both |
| `args` | copying readable caller tensors in H2D and exposing their existing host buffers to orchestration; pure outputs skip both |
| `arena_build`, `static_arena`, `gm_heap`, `shared_mem`, `runtime_init` | arena layout, GM heap and shared-memory bring-up |
| `host_orch` | **all** orchestration: every task submitted, every in-graph task recorded, the Definition built |
| `graph_upload` | one H2D of the block holding every Definition object, and binding each Graph task to the one with its key. The recorders built the objects in that block's host staging during `host_orch`, so this segment writes their headers and copies in only what did not fit |
Expand Down Expand Up @@ -462,7 +462,7 @@ meant to outlive it.
| `arena_h2d` † | 0.035–0.039 ms / 632 B | 0.03–0.10 ms / 632 B |
| `heap_used` | 127,673,344 | 2,038,508,544 |
| device wall | 39.3 ms | does not complete yet (`sched_error_code=5 INVALID_ARGS`) |
| `args` (excluded) | 1.37 s / 40.9 GB, 19 of 20 staged | 1.48 s / 45.8 GB, 77 of 92 staged |
| `args` (excluded) | 1.37 s / 40.9 GB, 19 of 20 copied in | 1.48 s / 45.8 GB, 77 of 92 copied in |
| `host_view_close` (excluded, legacy mapping path) | 0.25 s / 40.9 GB | 0.28 s / 45.8 GB |

† The three upload rows are the markers as they read at that commit, before the
Expand All @@ -474,35 +474,35 @@ remaining regions in `arena_h2d` — so the same case reports different figures
the same work.

**dsv4's `args` and `host_view_close` rows no longer describe that case at this
scale.** Both are per-byte costs over what a bind stages, and dsv4's parameters
scale.** Both are per-byte costs over what a bind copies in, and dsv4's parameters
now live in child memory: allocated once before the first round, and passed
through without malloc, H2D or a host view. What still crosses is
`num_tokens_per_owner`, the one caller tensor the host orchestrator has to read —
so a bind stages **1 of its 92 tensors, 8 bytes**. On `dcf7559e8`, 12 binds
so a bind copies in **1 of its 92 tensors, 8 bytes**. On `dcf7559e8`, 12 binds
(`--rounds 6`, both ranks) measure `args` at 0.036–0.075 ms and
`host_view_close` at 0.0012–0.0030 ms with `count=0 bytes=0`, against 1.48 s and
0.28 s over 45.8 GB above. The same run peaks at 1.31 GiB of host RSS across the
whole process tree under `--skip-golden`, and at 23.4 GiB when the fixture is
streamed in, where the row above cost ~45.5 GB per rank. qwen still stages its
streamed in, where the row above cost ~45.5 GB per rank. qwen still copies in its
fixture.

The rows also describe the legacy mapping behavior at the pinned commit. A
current bind uses the caller's existing host buffers as its
orchestration views, so it performs no `halHostRegister` calls and reports
`host_view_close count=0 bytes=0`. On Qwen3-14B this makes the close marker
20.12–24.73 us instead of the 0.25 s shown above. The old `args` figure included
20 registrations in addition to staging 19 tensors H2D; current `args` retains
20 registrations in addition to copying 19 tensors in H2D; current `args` retains
the H2D work but removes that registration side.

Three of these deserve reading together. `host_orch` is the whole story on dsv4 —
839 `submit_task`, 743 `record_in_graph_task` and 272 `alloc_tensors` per bind against qwen's
5, 277 and 2 — and its 2.3 ms of scatter is why a claim about it needs a
sub-counter rather than a stopwatch. At the pinned commit, `args` plus
`host_view_close` are two orders of magnitude above everything else while being
excluded from the control plane: they are staging and legacy mapping costs over
excluded from the control plane: they are copy-in and legacy mapping costs over
the ~41–46 GB of weights, not graph dispatch. Current qwen runs retain the
staging cost in `args` but close no mappings; moving dsv4's parameters to child
memory left its bind staging one 8-byte tensor, whose caller-buffer view also
copy-in cost in `args` but close no mappings; moving dsv4's parameters to child
memory left its bind copying in one 8-byte tensor, whose caller-buffer view also
needs no mapping. And dsv4's device wall is absent because the case did not
complete on device at the pinned commit — it is a completion case with no golden
whose host path is what these numbers describe, which is also why
Expand Down
4 changes: 2 additions & 2 deletions docs/dfx/l2-timing.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,8 +163,8 @@ at case setup, outside `Worker.run` and its round markers. Include setup and
final validation readback when reporting total case time; do not label the
round table alone as end-to-end case latency.

A child-memory argument skips the per-round staging path entirely, so `bind.args`
reports fewer staged tensors and fewer staged bytes for it. Numbers taken
A child-memory argument skips the per-round copy-in path entirely, so `bind.args`
reports a smaller `h2d=` count and fewer bytes for it. Numbers taken
before and after a case declares child memory are therefore not comparable on the
host/bind component; re-measure both arms with identical fixtures, hardware,
round counts and validation settings.
Expand Down
10 changes: 5 additions & 5 deletions docs/task-flow.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,7 +192,7 @@ snapshot for that purpose.
④ ChipStorageTaskArgs (ChipTensor records + scalars)
│ native run or prepare/launch/poll/finalize lifecycle
▼
chip runtime stages host-backed data or uses owned device memory
chip runtime copies host-backed data in or uses owned device memory
```

A public L2 `Worker.run()` performs the same materialization in its own
Expand Down Expand Up @@ -248,7 +248,7 @@ run token for the staged prepare/launch/poll/finalize path.

This boundary is a descriptor resolution and materialization step, not a
memcpy from the mailbox tensor array into `ChipTensor[]`. Host-backed arguments
may need device staging and output copy-back; device-backed arguments must
may need a device copy-in and output copy-back; device-backed arguments must
resolve to allocations owned by the target chip. The native `ChipWorker`
consumes the resulting POD and invokes the runtime's execution lifecycle.

Expand Down Expand Up @@ -433,7 +433,7 @@ uncertain.

#### TRB temporary buffer

`tensormap_and_ringbuffer` stages ordinary non-child tensor arguments through a
`tensormap_and_ringbuffer` copies ordinary non-child tensor arguments in through a
retained temporary buffer owned per pipeline slot, instead of a per-run `device_malloc()` /
`device_free()` pair. This is always on for TRB — an internal allocation
optimization with no user-facing switch. It is not serialized in task mailboxes
Expand All @@ -445,7 +445,7 @@ non-child tensors, growing it (free old + malloc new) only when a run needs
more than is currently retained, and bump-slices each tensor from it. The
buffer lives on the `DeviceRunner` across runs (freed once at finalize); the
platform only stores its `{addr, size}` slot. If a grow allocation fails the
run fails before device argument staging. See the runtime's `RUNTIME_LOGIC.md`
run fails before the device arguments are copied in. See the runtime's `RUNTIME_LOGIC.md`
§2.4 for the grow/reuse mechanics.

### SUB-type child loop (Python callable leaf)
Expand Down Expand Up @@ -753,7 +753,7 @@ Step-by-step (one chip worker):
| 5 | WT_chip_0 parent side | encode one leased task frame: write `config`, digest prefix, and the args blob; publish `TASK_READY` for the active lane or `PREPARE_READY` for a staged successor |
| 6 | chip_0 child process | validate the frame and resolve its digest; ordinary HBG with an active predecessor also prepares the leased inactive arena bank before publishing `FRAME_STAGED`, while a frame with no active predecessor, diagnostic HBG, and TMR publish after validation and defer native prepare |
| 7 | chip_0 native-run path | after activation and the predecessor's finalization fence, launch an already-prepared HBG run or finish deferred native preparation and then launch; poll it to completion and finalize it before another staged frame may launch. Compatibility endpoints perform the equivalent operation through blocking `ChipWorker::run` |
| 8 | runtime.so | stage resolved host-backed tensors on the device; dispatch AICPU / AICore; copy output back to `c` during finalization |
| 8 | runtime.so | copy resolved host-backed tensors in to the device; dispatch AICPU / AICore; copy output back to `c` during finalization |
| 9 | chip_0 child | native finalization returns; write `TASK_DONE` |
| 10 | WT_chip_0 parent | observe `TASK_DONE`; push success completion |
| 11 | Scheduler | mark slot COMPLETED; fanout release (none in this DAG); scope_end will release scope ref |
Expand Down
10 changes: 5 additions & 5 deletions docs/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -907,17 +907,17 @@ When kernels themselves differ (e.g., templated tile sizes tuned for device), se

`TensorArg(name, value, child_memory=True)` keeps a case-owned device buffer
across all rounds, including `--rounds 1`. `TaskArgsBuilder.add_tensor` accepts
the same keyword. The default remains host staging on every round.
the same keyword. The default remains host memory; IN and INOUT tensors are copied in on every round, while pure OUT buffers skip the copy.

| Declaration / direction | Setup | Between rounds | Validation |
| ----------------------- | ----- | -------------- | ---------- |
| Host-staged (default) | Existing path | Restore OUT/INOUT host fixtures | Existing per-round copy-back |
| Host memory (default) | Existing path | Restore OUT/INOUT host fixtures | Existing per-round copy-back |
| Child-memory IN | Allocate and upload once | Keep device address and input contents | No output readback |
| Child-memory OUT | Allocate without upload | Keep device contents; the case must define all compared elements | Final readback |
| Child-memory INOUT | Allocate and upload once | Keep device state | Final readback |

Golden evaluation follows the same state evolution: child-memory outputs retain
state and host-staged outputs reset. Cases with child-memory outputs compare after
state and host-memory outputs reset. Cases with child-memory outputs compare after
the final round; other cases continue comparing every round.

A tensor whose contents the HBG host orchestration reads (`get_tensor_data`) or
Expand All @@ -933,7 +933,7 @@ touches costs nothing.

On the second row the cost is per access, not per tensor, so a tensor the
orchestration reads thousands of times — `paged_attention`'s `block_table` is
read once per (batch, block) pair — is better left host-staged there. The
read once per (batch, block) pair — is better left in host memory there. The
declaration is per argument, so a data-dependent case can mix freely. The bind's
`BindHostViewClose` phase attributes report `devcopy=N` when this path was taken.

Expand All @@ -955,4 +955,4 @@ The HBG `paged_attention_unroll_manual_scope` examples include matched manual
`HostStaged` and `ChildMemory` cases, the latter declaring every tensor —
including the two the orchestration reads. The HBG `paged_attention` scene tests
carry the same pairing as non-manual cases, so CI covers an orchestration
reading child memory on both arches. Existing default cases retain host staging.
reading child memory on both arches. Existing default cases retain host memory.
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ class TestBenchmarkBgemmHostBuildGraph(SceneTestCase):
"function_name": "aicpu_orchestration_entry",
# C is a zero-initialized accumulator: the AIV add kernel reads C
# from GM, adds the matmul result, and stores it back across grid_k
# iterations. Its host-provided zeros must be staged H2D, so C is
# iterations. Its host-provided zeros must be copied in H2D, so C is
# INOUT (read-before-write), not a pure OUT.
"signature": [D.IN, D.IN, D.INOUT, D.IN],
},
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ class TestBenchmarkBgemm(SceneTestCase):
"function_name": "aicpu_orchestration_entry",
# C is a zero-initialized accumulator: the AIV add kernel reads C
# from GM, adds the matmul result, and stores it back across grid_k
# iterations. Its host-provided zeros must be staged H2D, so C is
# iterations. Its host-provided zeros must be copied in H2D, so C is
# INOUT (read-before-write), not a pure OUT.
"signature": [D.IN, D.IN, D.INOUT, D.IN],
},
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_t_inline9086__ssa_v0_tensor->buffer.addr) +
comb_t_inline9086__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_ffn_inline9269__ssa_v0_tensor->buffer.addr) +
comb_ffn_inline9269__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_t_inline9848__ssa_v0_tensor->buffer.addr) +
comb_t_inline9848__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_ffn_inline10031__ssa_v0_tensor->buffer.addr) +
comb_ffn_inline10031__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_t_inline10315__ssa_v0_tensor->buffer.addr) +
comb_t_inline10315__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_ffn_inline11076__ssa_v0_tensor->buffer.addr) +
comb_ffn_inline11076__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_t_inline11617__ssa_v0_tensor->buffer.addr) +
comb_t_inline11617__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1999,7 +1999,7 @@ extern "C" __aicore__ __attribute__((always_inline)) void kernel_entry(__gm__ in
reinterpret_cast<__gm__ float *>(comb_ffn_inline11955__ssa_v0_tensor->buffer.addr) +
comb_ffn_inline11955__ssa_v0_tensor->start_offset;

// Unpack tensor: hc_scale (scale2 read from GM instead of a host-staged scalar)
// Unpack tensor: hc_scale (scale2 read from GM instead of a host-passed scalar)
__gm__ Tensor *hc_scale_tensor = reinterpret_cast<__gm__ Tensor *>(args[4]);
__gm__ float *hc_scale =
reinterpret_cast<__gm__ float *>(hc_scale_tensor->buffer.addr) + hc_scale_tensor->start_offset;
Expand Down
Loading
Loading