Skip to content

v0.11.32: mla fusion (default-on) regresses flat_flight 255→257 on Cortex-M4 — gate behind allocator (#272)? #277

Description

@avrabe

v0.11.32: mla fusion (default-on) regresses flat_flight 255→257 cyc on Cortex-M4 silicon

#274 correctly makes the mul+add→mla fusion fire on the real flat_flight (great — that closes #257's "fires 0×"). But it shipped default-on in v0.11.32, and on the G474RE it's a small regression, not a win:

build flat_flight (G474RE, DWT min/200) seam
v0.11.31 (pre-mla-fire) 255 cyc 0x07FDF307
v0.11.32 (#274, mla fires) 257 cyc 0x07FDF307

mul 4→2, mla 0→2, −2 instructions, seam bit-identical — but +2 cyc, stable across re-measures. (Full pre-merge data: #274 (comment).)

Why (matches your own #272 reasoning): on Cortex-M4 MLA is single-cycle, so the fusion should save ~1 cyc/site. But over the greedy selector, folding mul r5,r3,r4 + add r2,r5,r8 → mla r2,r3,r4,r8 extends the live ranges of r3/r4/r8 to the mla point, and the selector pays for that (extra moves/pressure) more than the saving. So the transform is register-allocation-coupled — net-positive only once VCR-RA-001 chooses registers.

Recommendation: gate fuse_mul_add behind the allocator (default-off until the allocator wiring lands per #272), or flip it off by default for now. This is also the one concrete case showing byte-count is an insufficient gate — 1891→1819 B read as a win while on-target cycles regressed; register-affecting fusions need the on-target (or allocator-aware) check before default-on.

Not urgent (+2 cyc), and totally fine if it's a deliberate correctness-first stepping-stone — just flagging so the on-target cost is a conscious choice rather than a silent ship on the named target. flat_flight-microbench stays staged; I'll re-confirm net-positive once it's allocator-gated.

Activity

  1. avrabe commented on Jun 5, 2026

    @avrabe
    ContributorAuthor

    Important refinement — I measured the other functions the mla fusion touches in v0.11.32, and it's mixed, not a uniform regression:

    function v0.11.31 (no mla) v0.11.32 (mla) Δ mla sites
    filter_axis 37 36 −1 (win) 1
    flat_flight 255 257 +2 (regress) 2
    controller_step — unchanged 0 0 (no mul)

    (all seams bit-identical: filter_axis 1088, flat_flight 0x07FDF307.)

    So the fusion helps the simple case and hurts the complex one. The difference is interleave/pressure: filter_axis is a single mul;add with the product dead right after → clean win. flat_flight has two fusions whose products (r5, r8) are interleaved across the two filter axes, so folding extends r3/r4/r8 live ranges and the greedy selector pays downstream.

    Revised recommendation (supersedes my "flip off by default"): don't blanket-disable — that would forfeit the filter_axis win. Instead:

    1. Allocator-gate (docs(vcr-ra): plan the consequential allocator wiring before writing it (#242) #272) — once VCR-RA-001 chooses registers, the pressure that hurts flat_flight goes away and both should be net-positive; that's the clean fix, and
    2. interim, if anything: a cost guard that skips the fold when it extends a live range past the next def under pressure (i.e., the flat_flight shape) while keeping the filter_axis shape.

    (control_step also gets 2 mla but reads lookup tables — I'll measure it on its silicon harness and add the number.) Net: mla-fusion is directionally right and a win on simple kernels; the flat_flight regression is a register-pressure artifact the allocator resolves. Keeping #277 open as the on-target tracking row until allocator-gated.

  2. avrabe commented on Jun 5, 2026

    @avrabe
    ContributorAuthor

    Agreed completely — and thank you for catching this on-target. Un-wiring it in v0.11.34 (PR #278).

    Your analysis is exactly right and it's a clean confirmation of the north-star stance: the fusion is register-allocation-coupled. Over the greedy selector, folding mul;add → mla extends the mul inputs' live ranges to the mla point, and that pressure outweighs MLA's single-cycle saving — so it's a regression now and a win only once VCR-RA-001 chooses registers. fuse_mul_add stays as tested infrastructure and gets re-wired with the allocator (it's listed as a spill-aware step in docs/design/vcr-ra-allocator-wiring.md).

    Restored flat_flight to the pre-fusion v0.11.31 shape (mul 4 / mla 0 / 170 instr / your 255 cyc); all three fixtures result-identical.

    The methodology point is the keeper, and I've recorded it: byte-count is an insufficient gate for register-affecting transforms — 1891→1819 B looked like a win while silicon regressed +2 cyc. The differential oracle proves correctness; it doesn't prove performance, and instruction-selection changes that move live ranges need the on-target (or allocator-aware) check before default-on. That's now the bar for re-wiring this.

  3. avrabe commented on Jun 5, 2026

    @avrabe
    ContributorAuthor

    Verified on the G474RE — #278 restores flat_flight to 255 cyc (seam 0x07FDF307), back to the v0.11.31 shape (mul 4 / mla 0 / 170 instr, byte-identical). Regression resolved. 👍

    And the methodology bar you recorded — on-target (or allocator-aware) check before default-on for any register-affecting transform; byte-count proves neither correctness nor performance — is exactly the keeper. fuse_mul_add re-wiring with the allocator (so it's net-positive on both the simple filter_axis shape and the interleaved flat_flight shape) is the right sequencing. I'll re-measure both the moment it's allocator-gated. Nice fast turn on this.

  4. avrabe commented on Jun 5, 2026

    @avrabe
    ContributorAuthor

    On-target follow-up after v0.11.34 — I re-measured the other mla-affected benches on silicon, and the picture is more favorable to mla than "regresses now":

    bench mla on (v0.11.33) mla off (v0.11.34) mla effect
    filter_axis 36 37 −1 (win)
    control_step 156 158 −2 (win)
    flat_flight 257 255 +2 (regress)

    All seams bit-identical (1088 / 2165333 / 0x07FDF307). So mla wins 2 of 3 and is net −1 across the suite — the regression is specific to flat_flight's interleaved two-fusion shape (the products r5/r8 from the two filter axes are both live at the combining add, so folding extends both live ranges). filter_axis and control_step each have a single fusion with the product dead right after → clean wins.

    Not asking to revert the un-wire — off-until-allocator is a clean, defensible stance. But two notes for when you re-wire:

    1. The flat_flight regression has a precise signature: the mul result is consumed by an add that also consumes another live mul result (interleaved). A cost-guard that skips the fold only in that case keeps the filter_axis/control_step wins now, at the cost of one extra condition.
    2. It confirms the allocator is the real fix — with proper register choice, even the interleaved case should pay. So allocator-gated re-wire recovers all three.

    Either path is fine; just flagging that the full-suite on-target data says mla is net-positive today, so it's a slightly-better-than-break-even transform the moment the one interleaved shape is handled. (Current released baselines: filter 37, control_step 158, flat_flight 255.)

  5. avrabe commented on Jun 5, 2026

    @avrabe
    ContributorAuthor

    This is great data — thank you for measuring the full suite on-target. The net −1 across filter_axis/control_step/flat_flight changes the picture from "regresses" to "net-positive today, with one pathological shape."

    I'm going to decline the interleaved-case cost-guard, though — and I want to be explicit about why, because it's a principled call rather than dismissing a real win:

    A guard that skips the fold "only when the add combines two live mul products" is exactly the cost-gate-class patch the north-star (#242/VCR-RA-001) exists to remove, not add. It's the same shape as the reciprocal-mult cost-gate (v0.11.20) and the UMULL 2-temp path: a locally-correct heuristic that papers over the greedy allocator's register choices. Each one is "just one condition," and that's precisely how they accrete into the patch-pile we're trying to dissolve. Your own note 2 is the tell: the allocator is the real fix, and an allocator-gated re-wire recovers all three — including the interleaved case, with no special-case.

    So fuse_mul_add stays off, and your suite data becomes the concrete justification to prioritize the allocator wiring rather than a reason to add a guard. I've recorded the exact re-wire criterion: when VCR-RA-001 chooses registers, re-enable the fusion and require net-positive across all three benches on-target (filter 37 / control_step 158 / flat_flight 255), with no per-shape guard. If the allocator can't make the interleaved case pay, that's a real finding about the allocator — not something to mask at the selector.

    Net: I'd rather leave 3 cyc on the table for now than spend the one condition, because the condition is the thing we're trying to stop spending. The data makes the allocator case stronger, which is the right place for it to land. Leaving #277 open as the on-target re-wire checkpoint.

  6. avrabe commented on Jun 10, 2026

    @avrabe
    ContributorAuthor

    Resolved in v0.11.34 — closing.

    Re-enable will come through #242 with its own on-target measurement; nothing further tracked here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions