Skip to content

Refactor to allow example to write kernels and test - #11

Merged
ChaoWao merged 3 commits into
mainfrom
runtime
Jan 29, 2026
Merged

ChaoWao merged 3 commits into
mainfrom
runtime

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Jan 29, 2026

Copy link
Copy Markdown
Collaborator
  • Refactor kernel_add to use PTO tile-based operations
  • Rename compile_kernel to compile_incore and add integration tests
  • Add RuntimeBuilder and reorganize runtime code structure

ChaoWao and others added 3 commits January 28, 2026 21:40
- Add python/runtime_builder.py for building runtime from orchestration functions
- Add orchestration example (example_orch.cpp/h) demonstrating task graph building
- Reorganize runtime code into src/runtime/host_build_graph/ directory
- Update CI script to run pytest tests before examples
- Add tests/test_runtime_builder.py for RuntimeBuilder testing

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Rename PTOCompiler.compile_kernel() to compile_incore() for clarity
- Update example/main.py to use the new method name
- Add integration tests for RuntimeBuilder with real Ascend compilation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Replace simple for-loop implementation with tile-based approach using
GlobalTensor, Tile, and synchronization events for optimized tensor
addition on AI accelerators.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
@ChaoWao
ChaoWao merged commit 8877206 into main Jan 29, 2026
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph:

- scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset).
  BitStates::Storage is unsigned __int128 because a cluster spans 3 bits,
  so the shift reaches 72 (a2a3) / 108 (a5) — well past 64.

- scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default
  left it indeterminate; a never-init'd tracker returns garbage core counts.

- scheduler_completion.cpp: guard slot_state after pending_task acquire.
  The coordinator clears pending_task before resetting drain state; a thread
  exiting the ack barrier in that window loads nullptr and dereferences it.

- pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas.

- aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive,
  so two threads cannot both tear down shared state.

- pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before
  using it as a multiplier that could overflow int16_t.

- aicore_executor.cpp: bound speculative arg fill loops with
  n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write
  past the args[] array into the SPMD context fields.

Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab
layout (count|error_code|entries[]) is contiguous, so the flush covers
the correct range.
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph:

- scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset).
  BitStates::Storage is unsigned __int128 because a cluster spans 3 bits,
  so the shift reaches 72 (a2a3) / 108 (a5) — well past 64.

- scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default
  left it indeterminate; a never-init'd tracker returns garbage core counts.

- scheduler_completion.cpp: guard slot_state after pending_task acquire.
  The coordinator clears pending_task before resetting drain state; a thread
  exiting the ack barrier in that window loads nullptr and dereferences it.

- pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas.

- aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive,
  so two threads cannot both tear down shared state.

- pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before
  using it as a multiplier that could overflow int16_t.

- aicore_executor.cpp: bound speculative arg fill loops with
  n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write
  past the args[] array into the SPMD context fields.

Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab
layout (count|error_code|entries[]) is contiguous, so the flush covers
the correct range.
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph:

- scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset).
  BitStates::Storage is unsigned __int128 because a cluster spans 3 bits,
  so the shift reaches 72 (a2a3) / 108 (a5) — well past 64.

- scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default
  left it indeterminate; a never-init'd tracker returns garbage core counts.

- scheduler_completion.cpp: guard slot_state after pending_task acquire.
  The coordinator clears pending_task before resetting drain state; a thread
  exiting the ack barrier in that window loads nullptr and dereferences it.

- pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas.

- aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive,
  so two threads cannot both tear down shared state.

- pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before
  using it as a multiplier that could overflow int16_t.

- aicore_executor.cpp: bound speculative arg fill loops with
  n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write
  past the args[] array into the SPMD context fields.

Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab
layout (count|error_code|entries[]) is contiguous, so the flush covers
the correct range.
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph:

- scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset).
  BitStates::Storage is unsigned __int128 because a cluster spans 3 bits,
  so the shift reaches 72 (a2a3) / 108 (a5) — well past 64.

- scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default
  left it indeterminate; a never-init'd tracker returns garbage core counts.

- scheduler_completion.cpp: guard slot_state after pending_task acquire.
  The coordinator clears pending_task before resetting drain state; a thread
  exiting the ack barrier in that window loads nullptr and dereferences it.

- pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas.

- aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive,
  so two threads cannot both tear down shared state.

- pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before
  using it as a multiplier that could overflow int16_t.

- aicore_executor.cpp: bound speculative arg fill loops with
  n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write
  past the args[] array into the SPMD context fields.

Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab
layout (count|error_code|entries[]) is contiguous, so the flush covers
the correct range.
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph:

- scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset).
  BitStates::Storage is unsigned __int128 because a cluster spans 3 bits,
  so the shift reaches 72 (a2a3) / 108 (a5) — well past 64.

- scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default
  left it indeterminate; a never-init'd tracker returns garbage core counts.

- scheduler_completion.cpp: guard slot_state after pending_task acquire.
  The coordinator clears pending_task before resetting drain state; a thread
  exiting the ack barrier in that window loads nullptr and dereferences it.

- pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas.

- aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive,
  so two threads cannot both tear down shared state.

- pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before
  using it as a multiplier that could overflow int16_t.

- aicore_executor.cpp: bound speculative arg fill loops with
  n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write
  past the args[] array into the SPMD context fields.

Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab
layout (count|error_code|entries[]) is contiguous, so the flush covers
the correct range.
@zhusy54 zhusy54 mentioned this pull request Sep 8, 2026
4 of 5 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant