Conversation
ChaoWao
commented
Jan 29, 2026
Collaborator
- Refactor kernel_add to use PTO tile-based operations
- Rename compile_kernel to compile_incore and add integration tests
- Add RuntimeBuilder and reorganize runtime code structure
- Add python/runtime_builder.py for building runtime from orchestration functions - Add orchestration example (example_orch.cpp/h) demonstrating task graph building - Reorganize runtime code into src/runtime/host_build_graph/ directory - Update CI script to run pytest tests before examples - Add tests/test_runtime_builder.py for RuntimeBuilder testing Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Rename PTOCompiler.compile_kernel() to compile_incore() for clarity - Update example/main.py to use the new method name - Add integration tests for RuntimeBuilder with real Ascend compilation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Replace simple for-loop implementation with tile-based approach using GlobalTensor, Tile, and synchronization events for optimized tensor addition on AI accelerators. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph: - scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset). BitStates::Storage is unsigned __int128 because a cluster spans 3 bits, so the shift reaches 72 (a2a3) / 108 (a5) — well past 64. - scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default left it indeterminate; a never-init'd tracker returns garbage core counts. - scheduler_completion.cpp: guard slot_state after pending_task acquire. The coordinator clears pending_task before resetting drain state; a thread exiting the ack barrier in that window loads nullptr and dereferences it. - pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas. - aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive, so two threads cannot both tear down shared state. - pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before using it as a multiplier that could overflow int16_t. - aicore_executor.cpp: bound speculative arg fill loops with n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write past the args[] array into the SPMD context fields. Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab layout (count|error_code|entries[]) is contiguous, so the flush covers the correct range.
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph: - scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset). BitStates::Storage is unsigned __int128 because a cluster spans 3 bits, so the shift reaches 72 (a2a3) / 108 (a5) — well past 64. - scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default left it indeterminate; a never-init'd tracker returns garbage core counts. - scheduler_completion.cpp: guard slot_state after pending_task acquire. The coordinator clears pending_task before resetting drain state; a thread exiting the ack barrier in that window loads nullptr and dereferences it. - pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas. - aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive, so two threads cannot both tear down shared state. - pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before using it as a multiplier that could overflow int16_t. - aicore_executor.cpp: bound speculative arg fill loops with n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write past the args[] array into the SPMD context fields. Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab layout (count|error_code|entries[]) is contiguous, so the flush covers the correct range.
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph: - scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset). BitStates::Storage is unsigned __int128 because a cluster spans 3 bits, so the shift reaches 72 (a2a3) / 108 (a5) — well past 64. - scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default left it indeterminate; a never-init'd tracker returns garbage core counts. - scheduler_completion.cpp: guard slot_state after pending_task acquire. The coordinator clears pending_task before resetting drain state; a thread exiting the ack barrier in that window loads nullptr and dereferences it. - pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas. - aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive, so two threads cannot both tear down shared state. - pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before using it as a multiplier that could overflow int16_t. - aicore_executor.cpp: bound speculative arg fill loops with n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write past the args[] array into the SPMD context fields. Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab layout (count|error_code|entries[]) is contiguous, so the flush covers the correct range.
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph: - scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset). BitStates::Storage is unsigned __int128 because a cluster spans 3 bits, so the shift reaches 72 (a2a3) / 108 (a5) — well past 64. - scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default left it indeterminate; a never-init'd tracker returns garbage core counts. - scheduler_completion.cpp: guard slot_state after pending_task acquire. The coordinator clears pending_task before resetting drain state; a thread exiting the ack barrier in that window loads nullptr and dereferences it. - pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas. - aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive, so two threads cannot both tear down shared state. - pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before using it as a multiplier that could overflow int16_t. - aicore_executor.cpp: bound speculative arg fill loops with n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write past the args[] array into the SPMD context fields. Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab layout (count|error_code|entries[]) is contiguous, so the flush covers the correct range.
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
Aug 4, 2026
Applied to both a2a3 and a5 host_build_graph: - scheduler_dispatch.cpp: 1ULL << cluster_offset → BitStates::bit(offset). BitStates::Storage is unsigned __int128 because a cluster spans 3 bits, so the shift reaches 72 (a2a3) / 108 (a5) — well past 64. - scheduler_types.h: zero-initialize cluster_count_. CoreTracker() = default left it indeterminate; a never-init'd tracker returns garbage core counts. - scheduler_completion.cpp: guard slot_state after pending_task acquire. The coordinator clears pending_task before resetting drain state; a thread exiting the ack barrier in that window loads nullptr and dereferences it. - pto_runtime2.cpp: guard slot.task before dereference in both wait lambdas. - aicpu_executor.cpp: CAS on finished_ makes the deinit claim exclusive, so two threads cannot both tear down shared state. - pto_orchestrator.cpp: validate block_num <= PLATFORM_MAX_BLOCKDIM before using it as a multiplier that could overflow int16_t. - aicore_executor.cpp: bound speculative arg fill loops with n < SPMD_LOCAL_CONTEXT_INDEX so corrupt payload counters cannot write past the args[] array into the SPMD context fields. Found by CodeRabbit on hw-native-sys#1661. defer_flush (hw-native-sys#11) was verified — the slab layout (count|error_code|entries[]) is contiguous, so the flush covers the correct range.
4 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.