Repository navigation
Optimized path silently miscompiles register-exhausting i32 folds (no spill, no decline) #496
Description
Activity
- added a commit that references this issue
on Jun 25, 2026 [issue-hunt loop] Root cause located + a silicon-validated fixture is a victim.
The miscompile is the optimized path borrowing R12/IP as a pressure-relief scratch under register exhaustion.
optimizer_bridge.rsalloc_i32_scratch(lines ~2657-2661):// Pressure-relief fallback. R12 is acceptable here because // the call sites that use this helper write the destination // BEFORE using R12 as scratch (e.g. MemLoad emits the LDR // last, after the address math). Reg::R12
When R4-R8 are all live the helper returns R12 unconditionally. The "acceptable because the dest is written before R12 is used" assumption holds for a single immediately-consumed value but breaks in two ways, both observed:
- Two allocations both fall back to R12 → the second clobbers the first. The helper has no view of another concurrently-live value already in R12.
- An R12-allocated value must survive ACROSS a later R12 use — R12 is also the encoder's indexed-load base scratch (
MOVW/MOVT R12,#base; ADD R12,R12,Rindex), so the held value is destroyed by the address math.
Evidence
(a) The original
highfixture — two products land in R12, the first is lost:be: mul.w ip, r0, r2 ; product A -> R12 c2: mul.w ip, r1, r3 ; product B -> R12 (A clobbered) c6: add.w ip, ip, ip ; = B+B, not A+B(b)
control_step.wasm(the #209 / 0x00210A55 silicon-validated fixture) MISCOMPILES on the optimized path — case (2): a dynamic-indexi32.load16_ukeeps its index in R12, the base materialization overwrites it:13c: lsl.w ip, r4, ip ; index -> R12 144: movw ip, #0x100 148: movt ip, #0x2000 ; ip = 0x20000100 (base) — index clobbered 14c: add.w ip, ip, ip ; = 2*base = 0x40000200 150: ldrh.w ip, [ip] ; READ_UNMAPPED (vs correct base+index)Under unicorn
control_step_decide(3000,50,40,0)faults reading0x40000200where wasmtime gives0x00210A55. So this is not only contrived exhaustion folds — a core module doing an indexed load is hit.Minimal repro (
idxload_hp: a dynamic-index load + enough live products to exhaust R4-R8) reproduces the same R12 collisions.Scope
Pre-existing (present at
d79f6ad, independent of #490/#483). NOT in shipped firmware — gale ships--relocatable, whose direct selector spills on exhaustion (#242 step 3b) and never borrows R12; this is optimized-path-only.Fix direction (byte-changing → separate gated step)
The R12 borrow is unsound. Under exhaustion the optimized path should spill to the frame (mirror the direct selector's spill-on-exhaustion) or decline to the direct selector, not borrow R12 — the #212 lesson ("R12 is encoder scratch, never allocate it") applied to
ir_to_arm. A liveness-gated borrow (only when the value dies before the next R12 use) is a narrower but fragile alternative. Gate: re-freeze + the existinghigh_pressure_i32/ a new control_step-optimized execution differential.[issue-hunt loop] Severity escalation — BOTH silicon-reference fixtures miscompile on the optimized path, not just control_step.
Swept the other reference fixtures through the optimized path. The same R12-borrow clobber signature appears in both:
flight_seam.wasm(the0x07FDF307reference): 5×add.w ip, ip, ip, e.g.29e: mul.w ip, ip, r9 ; product A -> R12 2b6: mul.w ip, ip, r9 ; product B -> R12 (A clobbered) 2ba: add.w ip, ip, ip ; = B+B, not A+Bflight_seam_flat.wasm: 10×add.w ip, ip, ip.
add ip,ip,ipsumming a value with itself where the source wasm isa*b + c*dis a definitive clobber (two distinct sub-products both homed in R12). Combined with control_step's execution-verified fault, this means the default (non---relocatable)synth compilepath miscompiles every realistic fixture with enough register pressure — control_step (0x00210A55), flight_seam (0x07FDF307), and flight_seam_flat all hit it. The frozen byte gate masks this because it compiles these fixtures--relocatable(the direct selector, which spills on exhaustion and never borrows R12).This is the broad consequence of the R12 pressure-relief fallback in
alloc_i32_scratch(root cause above): it is not a fold-corner-case but the optimized path's default behavior under any non-trivial register pressure. Re-prioritizes #496 to the front of the optimizer_bridge frontier — the gate for the fix should include an optimized-path execution differential on control_step + flight_seam (both currently fail), and the fix is "spill-on-exhaustion or decline, never R12-borrow". Still not in shipped firmware (gale ships--relocatable).- added a commit that references this issue
on Jun 26, 2026 Fixed in #502 (squashed to `988ea3b` on main).
Fix: the optimized path's last-resort scratch fallbacks (`alloc_i32_scratch` → R12/IP, and the i64-pair → (R4,R5) alias) now flag pool exhaustion via a shared `Cell` instead of silently borrowing R12. After the build, `ir_to_arm` returns `Err(UnsupportedInstruction)`, so `arm_backend` retries via the direct selector, which has real spill-on-exhaustion support. The declined output is byte-identical to `--no-optimize`.
This is a correctness retreat, stated as such: the affected functions lose optimized-path codegen until proper spilling lands in the VCR-RA verified selector (#242). The point was to stop the default `synth compile` from shipping a miscompile.
Gate (new CI job `r12-spill-496-oracle`, `scripts/repro/r12_spill_496_differential.py`): both silicon fixtures, compiled via the default optimized path (no `--relocatable`), execute bit-identically to wasmtime under unicorn — control_step → `0x00210A55`, flight_algo → `0x07FDF307`. Non-vacuity: pre-fix the harness faults on every vector (READ_UNMAPPED from `add ip,ip,ip` = 2×base), and the fix changed both functions' bytes (control_step 626→464, flat 2096→1070).
Re-verified on main after merge: r12-spill-496 PASS, frozen byte gate bit-identical, control_step (relocatable) 13/13, callee-saved-490 + block-brif-483 PASS, all 20 PR checks green.
Remaining optimized-path frontier: #498 (byte-size estimator drift), #499 (spill-frame SP imbalance), #500 (broader br_if). The structural fix for this whole class is the VCR-RA verified selector + verified selector DSL (#242).
- added 7 commits that reference this issue
on Jun 26, 2026 - added a commit that references this issue
on Jul 17, 2026 - added a commit that references this issue
on Sep 15, 2026 - added a commit that references this issue
on Sep 30, 2026
Summary
The optimized ARM selection path (default, non-
--relocatable) produces wrong code for an i32 expression that keeps more values simultaneously live than the R0-R8 allocation pool holds. Unlike the direct selector — which gained spill-on-exhaustion (#242, VCR-RA-001 step 3b) — the optimized path neither spills nor declines to the direct selector; it silently emits an incorrect instruction stream.This is pre-existing (independent of #490, found while building that fix's fixture).
Repro
Run
highunder unicorn (Thumb) vs wasmtime ground truth:0x800x340x2880x8c5d60Compiling the same module with
--relocatable(direct selector) computes correctly.Expected
Either spill on exhaustion (as the direct path does) or decline to the direct selector (the existing fallback for declined modules). Silent wrong-code is the worst outcome — a
Result::Err/decline would be caught by the existing fallback inarm_backend.rs.Notes