Conversation
Owner
Author
|
独立 review 记录(GLM-5.3,独立新会话):
两次 review 的所有意见(4×P2)均已修复并经 1874/1874 lit、144/144 tilelib 单测、9/9 VMI 脚本回归验证。 |
added 2 commits
September 7, 2026 18:38
…ernels Fixes all six compile-time failure classes from Issue #24 on the merged-device VPTO route (--vpto-emit-merged-device-only, level3): 1. pto.trsqrt 3-operand form: register template_trsqrt_tmp / template_trsqrt_tmp_1d (src, dst, tmp) via register_unary(has_tmp=) and vmi_trsqrt_with_tmp_narrow for compact rows below the 128-byte elementwise minimum (UB slots are 256-byte planned so a full-row VMI access stays in-slot). 2. pto.tgather 4-operand index form: register template_tgather_tmp whose tmp workspace is not read by the A5 vgather2 sequence, keyed to the verifier contract (tmp dtype == indices dtype, index width follows data width). Fix _axis_is_row/_axis_is_col parameter names (the context key is 'axis', the old names made every mask-form candidate reject) and add equal-storage-width cross-dtype signatures for the packed-sort-key mask gather. Export axis/cmp_mode/offset context attrs from InsertTemplateAttributes. 3. pto.tmax / pto.tcolexpand / pto.trowexpand / trowexpandmul(tadd) constraint gaps: template_tmax_dst_valid_prefix accepts the PyPTO row-reduce accumulate form (src1 fills a strict prefix of the dst row; strict so it stays mutually exclusive with template_tmax); _valid_column_expand accepts src valid cols <= dst valid cols; trowexpand accepts single-column col-major sources; valid-row expand-binary accepts src1 valid rows >= dst. 4. pto.tsetval / pto.tgetval VPTO lowering: FoldTileBufIntrinsics now rewrites tile scalar IO to tile_buf_addr + store/load_scalar before addr folding, so the ops no longer survive to LLVM translation as unrealized_conversion_cast. 5. pto.initialize_l2l_pipe peer pairing across compilation units: PTOResolveReservedBuffers accepts incomplete peer groups when every init belongs to a cross-module pipe, records unresolvable imports under a synthetic peer key, and materializes the import base to the shared pipe-contract constant. The relaxation is gated on the VPTO single-kernel-per-file shape (new isSingleKernelVptoUnit helper counts authored non-private functions), so multi-function units keep rejecting bad peer_func symbols (3 existing guard tests still pass). 6. One-entry-per-ELF vs AIC+AIV groups: VPTOSplitCVModule stamps every entry candidate when each kernel kind contributes at most one candidate; ImportReservedBufferOp::verify consults a driver-installed parse-time backend hint so single-kernel units parse before the module attribute is stamped. The remaining 13 of 27 DeepSeek-V4 Pro cases are blocked on the tpush/tpop/tstore VPTO data-plane lowering (loop-carried fixpipe slot rotation), tracked as a follow-up; this PR unblocks everything up to that boundary. 14 of 27 cases now compile end-to-end (was 1). lit: 1874/1874 pass including the new import_reserved_buffer_single_kernel_unit_resolves_a5 regression test; ptodsl tilelib suites and VMI scripts pass.
…pipe fallback Review follow-ups on 49fb751 (all P2, no behavior change on the Issue #24 reproducer set): - narrow_full_row_vmi_constraint: add the <128-byte upper bound so the narrow candidate stays mutually exclusive with the min_128b_row / full_physical_row candidates and wide rows keep selecting the standard elementwise forms. - trowexpand _row_major_or_single_column: require the dst to be row-major, matching what the op verifier can produce; the col-major allowance now applies to the src vector only. - cross-module import materialization: emit a remark when the pipe-contract base-0 fallback fires, so a future divergent PyPTO contract is diagnosable instead of silently disagreeing with the peer unit's reserve-side materialization. Document the one-input per-process assumption of the parse-time backend hint. - tmax: reuse _ub_or_vec_row_major from _elementwise instead of a third copy. The tgather 3-/4-operand bodies stay duplicated (the template tracer evaluates loop bounds through the traced function's scope, so a shared helper breaks range() tracing); cross-reference comments now tie the two copies together.
|
Warning: @liuzidi, ci-sim exceeded its soft runtime budget.
This warning is advisory only and does not affect required checks. Please inspect the step timings for an unexpected regression. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix Issue #24: VPTO backend lowering failures for PyPTO merged-device A5 kernels
Fixes #24
Summary
On the PyPTO merged-device VPTO route (
--pto-backend=vpto --pto-level=level3 --vpto-emit-merged-device-only, one kernel per compilation unit), 26 of the 27 DeepSeek-V4 Pro operator test cases (pypto-lib/models/deepseek_v4_pro/*.py, platform=a5) failed at PTOAS compile time. This PR closes all six compile-time failure classes from the issue analysis.Result: 16 of 27 cases now compile end-to-end (previously 1). The remaining 11 are blocked on the
tpush/tpop/tstoreVPTO data-plane lowering (loop-carried fixpipe slot rotation), which is a separate architectural gap documented as a follow-up below — everything up to that boundary now works.Root causes fixed
pto.trsqrt3-operand form (src, tmp) → dst had no TileLib template (11 cases)template_trsqrt_tmp/template_trsqrt_tmp_1dvia a newregister_unary(has_tmp=True)flag;vmi_trsqrt_with_tmp_narrowfor compact rows below the 128-byte elementwise minimum (UB slots are planned 256-byte aligned, so a full-row VMI access stays in-slot)pto.tgather4-operand index form had no template;_axis_is_row/_axis_is_colparam-name bug made every mask-form candidate reject; packed-sort-key gather needs cross-dtype signatures (20 cases)template_tgather_tmp(tmp not read by the A5 vgather2 sequence; tmp dtype == indices dtype; index width follows data width, mirroring the op verifier); fixed the axis param names; equal-storage-width cross-dtype mask signatures withpto.vbitcaston the result;InsertTemplateAttributesexportsaxis/cmp_mode/offsetcontext attrspto.tmax/tcolexpand/trowexpand/trowexpandmulconstraints rejected real PyPTO tile shapes (2 cases)template_tmax_dst_valid_prefix(src1 valid cols strictly less than dst — strict so it stays mutually exclusive withtemplate_tmaxand selection stays unambiguous);_valid_column_expandaccepts src valid cols ≤ dst valid cols; trowexpand accepts[rows,1]col-major single-column sources; expand-binary valid-row check eq → gepto.tsetval/pto.tgetvalhad no VPTO lowering →unrealized_conversion_cast(11 cases)FoldTileBufIntrinsicsrewrites tile scalar IO totile_buf_addr+store_scalar/load_scalarbefore addr folding (VEC space + element-type match enforced; hard error otherwise)pto.initialize_l2l_pipepeer-init pairing failed when the peer kernel lives in another compilation unit (4 cases)PTOResolveReservedBuffersaccepts incomplete peer groups when all inits are cross-module pipes, records unresolvable imports under a synthetic peer key, and materializes the import base to the shared pipe-contract constant (PyPTO plans both sides from base=0). The relaxation is gated on the single-kernel-per-file shape (isSingleKernelVptoUnit: exactly one authored, non-privatefunc.funcin the unit), so multi-function units still reject badpeer_funcsymbols — the three existing guard tests keep passingVPTOSplitCVModulestamps every entry candidate when each kernel kind contributes at most one candidate;ImportReservedBufferOp::verifyconsults a driver-installed parse-time backend hint so single-kernel units parse before the module attribute is stampedVMI SPEC compliance
The new VMI template (
vmi_trsqrt_with_tmp_narrow) and the VMI-related constraint helpers follow the VMI spec: thepto.vmilayer only emits vector ops on logical vector/mask values (scalar-pipe ops stay in thepto.milayer);requires_full_physical_row=Falseis only used where memory planning guarantees the slot is padded to 256 bytes, so a full-row access cannot over-read the slot.Testing
pto+vpto+vmi_new): 1874/1874 pass, including the three pre-existing peer-lookup guard tests and the new regression testtest/lit/pto/import_reserved_buffer_single_kernel_unit_resolves_a5.pto(verifies: parse-time peer check relaxed for the single-kernel shape → resolve pass materializes the import to the pipe-contract constant → pipeline proceeds past peer resolution into VPTO LLVM emission, failing only at the documented tpop data-plane gap).test_vmi_*.pyscripts pass._tmptemplates use distinct candidate IDs, and the tgather_tmpsignature mirrors the op-verifier dtype coupling (f32→i32, f16/bf16→i16).Known follow-up (out of scope)
The remaining 13 cases stop at the
tpush/tpop(fixpipe FIFO slot rotation, loop-carried tile index) and onetstoreconstraint-form VPTO data-plane lowering gap. That requires loweringtpopto TASSIGN atCONSUMER_BUF + (tileIndex % SLOT_NUM) * SLOT_SIZEwith loop-carried slot rotation — large enough to warrant its own PR. This PR unblocks everything up to that boundary (peer resolution, template selection, scalar IO, entry stamping all work).