Skip to content

Optimized path silently miscompiles register-exhausting i32 folds (no spill, no decline) #496

Description

@avrabe

Summary

The optimized ARM selection path (default, non---relocatable) produces wrong code for an i32 expression that keeps more values simultaneously live than the R0-R8 allocation pool holds. Unlike the direct selector — which gained spill-on-exhaustion (#242, VCR-RA-001 step 3b) — the optimized path neither spills nor declines to the direct selector; it silently emits an incorrect instruction stream.

This is pre-existing (independent of #490, found while building that fix's fixture).

Repro

(module
  (memory 1)
  (func $high (export "high") (param i32 i32 i32 i32) (result i32)
    (i32.add
      (i32.add
        (i32.add (i32.mul (local.get 0) (local.get 1))
                 (i32.mul (local.get 2) (local.get 3)))
        (i32.add (i32.mul (local.get 0) (local.get 3))
                 (i32.mul (local.get 1) (local.get 2))))
      (i32.add
        (i32.add (i32.mul (local.get 0) (local.get 2))
                 (i32.mul (local.get 1) (local.get 3)))
        (i32.add (i32.mul (local.get 0) (local.get 0))
                 (i32.mul (local.get 3) (local.get 3)))))))
synth compile high.wat -o /tmp/high.elf --target cortex-m4 --all-exports

Run high under unicorn (Thumb) vs wasmtime ground truth:

args synth (optimized) wasmtime
(1,2,3,4) 0x80 0x34
(3000,50,7,9) 0x288 0x8c5d60

Compiling the same module with --relocatable (direct selector) computes correctly.

Expected

Either spill on exhaustion (as the direct path does) or decline to the direct selector (the existing fallback for declined modules). Silent wrong-code is the worst outcome — a Result::Err/decline would be caught by the existing fallback in arm_backend.rs.

Notes

Activity

  1. avrabe commented on Jun 25, 2026

    @avrabe
    ContributorAuthor

    [issue-hunt loop] Root cause located + a silicon-validated fixture is a victim.

    The miscompile is the optimized path borrowing R12/IP as a pressure-relief scratch under register exhaustion. optimizer_bridge.rs alloc_i32_scratch (lines ~2657-2661):

    // Pressure-relief fallback. R12 is acceptable here because
    // the call sites that use this helper write the destination
    // BEFORE using R12 as scratch (e.g. MemLoad emits the LDR
    // last, after the address math).
    Reg::R12

    When R4-R8 are all live the helper returns R12 unconditionally. The "acceptable because the dest is written before R12 is used" assumption holds for a single immediately-consumed value but breaks in two ways, both observed:

    1. Two allocations both fall back to R12 → the second clobbers the first. The helper has no view of another concurrently-live value already in R12.
    2. An R12-allocated value must survive ACROSS a later R12 use — R12 is also the encoder's indexed-load base scratch (MOVW/MOVT R12,#base; ADD R12,R12,Rindex), so the held value is destroyed by the address math.

    Evidence

    (a) The original high fixture — two products land in R12, the first is lost:

    be: mul.w ip, r0, r2     ; product A -> R12
    c2: mul.w ip, r1, r3     ; product B -> R12  (A clobbered)
    c6: add.w ip, ip, ip     ; = B+B, not A+B
    

    (b) control_step.wasm (the #209 / 0x00210A55 silicon-validated fixture) MISCOMPILES on the optimized path — case (2): a dynamic-index i32.load16_u keeps its index in R12, the base materialization overwrites it:

    13c: lsl.w ip, r4, ip    ; index -> R12
    144: movw ip, #0x100
    148: movt ip, #0x2000     ; ip = 0x20000100 (base) — index clobbered
    14c: add.w ip, ip, ip     ; = 2*base = 0x40000200
    150: ldrh.w ip, [ip]      ; READ_UNMAPPED (vs correct base+index)
    

    Under unicorn control_step_decide(3000,50,40,0) faults reading 0x40000200 where wasmtime gives 0x00210A55. So this is not only contrived exhaustion folds — a core module doing an indexed load is hit.

    Minimal repro (idxload_hp: a dynamic-index load + enough live products to exhaust R4-R8) reproduces the same R12 collisions.

    Scope

    Pre-existing (present at d79f6ad, independent of #490/#483). NOT in shipped firmware — gale ships --relocatable, whose direct selector spills on exhaustion (#242 step 3b) and never borrows R12; this is optimized-path-only.

    Fix direction (byte-changing → separate gated step)

    The R12 borrow is unsound. Under exhaustion the optimized path should spill to the frame (mirror the direct selector's spill-on-exhaustion) or decline to the direct selector, not borrow R12 — the #212 lesson ("R12 is encoder scratch, never allocate it") applied to ir_to_arm. A liveness-gated borrow (only when the value dies before the next R12 use) is a narrower but fragile alternative. Gate: re-freeze + the existing high_pressure_i32 / a new control_step-optimized execution differential.

  2. avrabe commented on Jun 26, 2026

    @avrabe
    ContributorAuthor

    [issue-hunt loop] Severity escalation — BOTH silicon-reference fixtures miscompile on the optimized path, not just control_step.

    Swept the other reference fixtures through the optimized path. The same R12-borrow clobber signature appears in both:

    • flight_seam.wasm (the 0x07FDF307 reference): 5× add.w ip, ip, ip, e.g.
      29e: mul.w ip, ip, r9     ; product A -> R12
      2b6: mul.w ip, ip, r9     ; product B -> R12 (A clobbered)
      2ba: add.w ip, ip, ip     ; = B+B, not A+B
      
    • flight_seam_flat.wasm: 10× add.w ip, ip, ip.

    add ip,ip,ip summing a value with itself where the source wasm is a*b + c*d is a definitive clobber (two distinct sub-products both homed in R12). Combined with control_step's execution-verified fault, this means the default (non---relocatable) synth compile path miscompiles every realistic fixture with enough register pressure — control_step (0x00210A55), flight_seam (0x07FDF307), and flight_seam_flat all hit it. The frozen byte gate masks this because it compiles these fixtures --relocatable (the direct selector, which spills on exhaustion and never borrows R12).

    This is the broad consequence of the R12 pressure-relief fallback in alloc_i32_scratch (root cause above): it is not a fold-corner-case but the optimized path's default behavior under any non-trivial register pressure. Re-prioritizes #496 to the front of the optimizer_bridge frontier — the gate for the fix should include an optimized-path execution differential on control_step + flight_seam (both currently fail), and the fix is "spill-on-exhaustion or decline, never R12-borrow". Still not in shipped firmware (gale ships --relocatable).

  3. avrabe commented on Jun 26, 2026

    @avrabe
    ContributorAuthor

    Fixed in #502 (squashed to `988ea3b` on main).

    Fix: the optimized path's last-resort scratch fallbacks (`alloc_i32_scratch` → R12/IP, and the i64-pair → (R4,R5) alias) now flag pool exhaustion via a shared `Cell` instead of silently borrowing R12. After the build, `ir_to_arm` returns `Err(UnsupportedInstruction)`, so `arm_backend` retries via the direct selector, which has real spill-on-exhaustion support. The declined output is byte-identical to `--no-optimize`.

    This is a correctness retreat, stated as such: the affected functions lose optimized-path codegen until proper spilling lands in the VCR-RA verified selector (#242). The point was to stop the default `synth compile` from shipping a miscompile.

    Gate (new CI job `r12-spill-496-oracle`, `scripts/repro/r12_spill_496_differential.py`): both silicon fixtures, compiled via the default optimized path (no `--relocatable`), execute bit-identically to wasmtime under unicorn — control_step → `0x00210A55`, flight_algo → `0x07FDF307`. Non-vacuity: pre-fix the harness faults on every vector (READ_UNMAPPED from `add ip,ip,ip` = 2×base), and the fix changed both functions' bytes (control_step 626→464, flat 2096→1070).

    Re-verified on main after merge: r12-spill-496 PASS, frozen byte gate bit-identical, control_step (relocatable) 13/13, callee-saved-490 + block-brif-483 PASS, all 20 PR checks green.

    Remaining optimized-path frontier: #498 (byte-size estimator drift), #499 (spill-frame SP imbalance), #500 (broader br_if). The structural fix for this whole class is the VCR-RA verified selector + verified selector DSL (#242).

  4. added a commit that references this issue on Jul 17, 2026
  5. added a commit that references this issue on Sep 15, 2026
  6. added a commit that references this issue on Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions