Repository navigation
v0.11.32: mla fusion (default-on) regresses flat_flight 255→257 on Cortex-M4 — gate behind allocator (#272)? #277
Description
Activity
Important refinement — I measured the other functions the mla fusion touches in v0.11.32, and it's mixed, not a uniform regression:
function v0.11.31 (no mla) v0.11.32 (mla) Δ mla sites filter_axis37 36 −1 (win) 1 flat_flight255 257 +2 (regress) 2 controller_step— unchanged 0 0 (no mul) (all seams bit-identical: filter_axis
1088, flat_flight0x07FDF307.)So the fusion helps the simple case and hurts the complex one. The difference is interleave/pressure:
filter_axisis a singlemul;addwith the product dead right after → clean win.flat_flighthas two fusions whose products (r5, r8) are interleaved across the two filter axes, so folding extends r3/r4/r8 live ranges and the greedy selector pays downstream.Revised recommendation (supersedes my "flip off by default"): don't blanket-disable — that would forfeit the
filter_axiswin. Instead:- Allocator-gate (docs(vcr-ra): plan the consequential allocator wiring before writing it (#242) #272) — once VCR-RA-001 chooses registers, the pressure that hurts
flat_flightgoes away and both should be net-positive; that's the clean fix, and - interim, if anything: a cost guard that skips the fold when it extends a live range past the next def under pressure (i.e., the
flat_flightshape) while keeping thefilter_axisshape.
(control_step also gets 2 mla but reads lookup tables — I'll measure it on its silicon harness and add the number.) Net: mla-fusion is directionally right and a win on simple kernels; the flat_flight regression is a register-pressure artifact the allocator resolves. Keeping #277 open as the on-target tracking row until allocator-gated.
- Allocator-gate (docs(vcr-ra): plan the consequential allocator wiring before writing it (#242) #272) — once VCR-RA-001 chooses registers, the pressure that hurts
Agreed completely — and thank you for catching this on-target. Un-wiring it in v0.11.34 (PR #278).
Your analysis is exactly right and it's a clean confirmation of the north-star stance: the fusion is register-allocation-coupled. Over the greedy selector, folding
mul;add → mlaextends the mul inputs' live ranges to the mla point, and that pressure outweighs MLA's single-cycle saving — so it's a regression now and a win only once VCR-RA-001 chooses registers.fuse_mul_addstays as tested infrastructure and gets re-wired with the allocator (it's listed as a spill-aware step indocs/design/vcr-ra-allocator-wiring.md).Restored
flat_flightto the pre-fusion v0.11.31 shape (mul 4 / mla 0 / 170 instr / your 255 cyc); all three fixtures result-identical.The methodology point is the keeper, and I've recorded it: byte-count is an insufficient gate for register-affecting transforms —
1891→1819 Blooked like a win while silicon regressed +2 cyc. The differential oracle proves correctness; it doesn't prove performance, and instruction-selection changes that move live ranges need the on-target (or allocator-aware) check before default-on. That's now the bar for re-wiring this.Verified on the G474RE — #278 restores
flat_flightto 255 cyc (seam0x07FDF307), back to the v0.11.31 shape (mul 4 / mla 0 / 170 instr, byte-identical). Regression resolved. 👍And the methodology bar you recorded — on-target (or allocator-aware) check before default-on for any register-affecting transform; byte-count proves neither correctness nor performance — is exactly the keeper.
fuse_mul_addre-wiring with the allocator (so it's net-positive on both the simplefilter_axisshape and the interleavedflat_flightshape) is the right sequencing. I'll re-measure both the moment it's allocator-gated. Nice fast turn on this.- added a commit that references this issue
on Jun 5, 2026 On-target follow-up after v0.11.34 — I re-measured the other mla-affected benches on silicon, and the picture is more favorable to mla than "regresses now":
bench mla on (v0.11.33) mla off (v0.11.34) mla effect filter_axis36 37 −1 (win) control_step156 158 −2 (win) flat_flight257 255 +2 (regress) All seams bit-identical (
1088/2165333/0x07FDF307). So mla wins 2 of 3 and is net −1 across the suite — the regression is specific toflat_flight's interleaved two-fusion shape (the products r5/r8 from the two filter axes are both live at the combining add, so folding extends both live ranges).filter_axisandcontrol_stepeach have a single fusion with the product dead right after → clean wins.Not asking to revert the un-wire — off-until-allocator is a clean, defensible stance. But two notes for when you re-wire:
- The flat_flight regression has a precise signature: the mul result is consumed by an add that also consumes another live mul result (interleaved). A cost-guard that skips the fold only in that case keeps the
filter_axis/control_stepwins now, at the cost of one extra condition. - It confirms the allocator is the real fix — with proper register choice, even the interleaved case should pay. So allocator-gated re-wire recovers all three.
Either path is fine; just flagging that the full-suite on-target data says mla is net-positive today, so it's a slightly-better-than-break-even transform the moment the one interleaved shape is handled. (Current released baselines: filter 37, control_step 158, flat_flight 255.)
- The flat_flight regression has a precise signature: the mul result is consumed by an add that also consumes another live mul result (interleaved). A cost-guard that skips the fold only in that case keeps the
This is great data — thank you for measuring the full suite on-target. The net −1 across filter_axis/control_step/flat_flight changes the picture from "regresses" to "net-positive today, with one pathological shape."
I'm going to decline the interleaved-case cost-guard, though — and I want to be explicit about why, because it's a principled call rather than dismissing a real win:
A guard that skips the fold "only when the add combines two live mul products" is exactly the cost-gate-class patch the north-star (#242/VCR-RA-001) exists to remove, not add. It's the same shape as the reciprocal-mult cost-gate (v0.11.20) and the UMULL 2-temp path: a locally-correct heuristic that papers over the greedy allocator's register choices. Each one is "just one condition," and that's precisely how they accrete into the patch-pile we're trying to dissolve. Your own note 2 is the tell: the allocator is the real fix, and an allocator-gated re-wire recovers all three — including the interleaved case, with no special-case.
So
fuse_mul_addstays off, and your suite data becomes the concrete justification to prioritize the allocator wiring rather than a reason to add a guard. I've recorded the exact re-wire criterion: when VCR-RA-001 chooses registers, re-enable the fusion and require net-positive across all three benches on-target (filter 37 / control_step 158 / flat_flight 255), with no per-shape guard. If the allocator can't make the interleaved case pay, that's a real finding about the allocator — not something to mask at the selector.Net: I'd rather leave 3 cyc on the table for now than spend the one condition, because the condition is the thing we're trying to stop spending. The data makes the allocator case stronger, which is the right place for it to land. Leaving #277 open as the on-target re-wire checkpoint.
Resolved in v0.11.34 — closing.
- The default-on fusion was un-wired in v0.11.34:
flat_flightrestored to the pre-fusion 255-cyc codegen (mul 4 / mla 0 / 170 instr), all three differential fixtures result-identical. - The answer to the title question is yes — gated behind the allocator:
fuse_mul_addstays as fully-tested infrastructure insynth_synthesis::liveness, and the re-wiring is tracked as part of VCR-RA-001 under the Epic: reach and beat native codegen — proof-carrying specialization as the mechanism (VCR-*) #242 epic. The substrate it needs just landed: the flag-gated range re-allocation pass (feat(vcr-ra): flag-gated range re-allocation pass — first consequential allocator step (wiring 3a) #288) is the allocator context in which the fusion becomes net-positive (live-range-aware register choice). - Lesson recorded in the epic: a register-pressure-affecting transform needs an on-target/allocator-aware gate, not a byte-count gate, before it defaults on.
Re-enable will come through #242 with its own on-target measurement; nothing further tracked here.
- The default-on fusion was un-wired in v0.11.34:
v0.11.32: mla fusion (default-on) regresses
flat_flight255→257 cyc on Cortex-M4 silicon#274 correctly makes the
mul+add→mlafusion fire on the realflat_flight(great — that closes #257's "fires 0×"). But it shipped default-on in v0.11.32, and on the G474RE it's a small regression, not a win:0x07FDF3070x07FDF307mul 4→2, mla 0→2, −2 instructions, seam bit-identical — but +2 cyc, stable across re-measures. (Full pre-merge data: #274 (comment).)Why (matches your own #272 reasoning): on Cortex-M4
MLAis single-cycle, so the fusion should save ~1 cyc/site. But over the greedy selector, foldingmul r5,r3,r4+add r2,r5,r8→mla r2,r3,r4,r8extends the live ranges of r3/r4/r8 to the mla point, and the selector pays for that (extra moves/pressure) more than the saving. So the transform is register-allocation-coupled — net-positive only once VCR-RA-001 chooses registers.Recommendation: gate
fuse_mul_addbehind the allocator (default-off until the allocator wiring lands per #272), or flip it off by default for now. This is also the one concrete case showing byte-count is an insufficient gate —1891→1819 Bread as a win while on-target cycles regressed; register-affecting fusions need the on-target (or allocator-aware) check before default-on.Not urgent (+2 cyc), and totally fine if it's a deliberate correctness-first stepping-stone — just flagging so the on-target cost is a conscious choice rather than a silent ship on the named target.
flat_flight-microbenchstays staged; I'll re-confirm net-positive once it's allocator-gated.