Skip to content

perf: consecutive const-address stores re-materialize the linear-memory base every time #468

Description

@avrabe

Pattern

On the absolute / native-pointer store path, i32.store* (i32.const ADDR) V lowers to:

movw ip, #<base_lo>     ; e.g. 0x0100
movt ip, #<base_hi>     ; e.g. 0x2000   -> ip = 0x20000100 (linear-memory base)
add.w ip, ip, <addr>
str V, [ip]

The base is a loop-invariant compile-time constant, but it's re-materialized (movw+movt = 8 B, ~2 cyc) before every store. A straight-line run of N stores pays N base-materializations where 1 would do. This is common in struct/field initializers (a function that sets up several adjacent fields).

Repro (generic)

scripts/repro/redundant_base_materialization.wat — 7 consecutive const-address stores.

$ synth compile scripts/repro/redundant_base_materialization.wat -o /tmp/rbm.elf --target cortex-m4 --all-exports
$ arm-none-eabi-objdump -d -M force-thumb /tmp/rbm.elf | grep -c 'movw\s\+ip, #256'
7        # 7 stores -> 7 identical base materializations; 6 are redundant

Cost

~(N-1) × 8 B + (N-1) × ~2 cyc per straight-line store run. For a 7-store initializer that's ~48 B and ~12 cyc recoverable in one function; it scales linearly with field count, so field-heavy init code pays the most.

Fix direction

Hoist the constant base into a stable register once per straight-line region (const-CSE of the base — ip/R12 is encoder scratch and clobbered per-access, so the hoist target must be a preserved reg or the already-reserved memory-base register), then address each store as add ip, base_reg, <addr>; str. Same const-CSE family as the existing redundant-const elimination, specialized to the 2-instruction MOVW/MOVT base.

Byte-changing → would land flag-off (frozen-safe), execution-differential + on-target cycle gate, then default-on — the established lever path. Tracked under #390 / VCR-RA (epic #242).

Related

A DWARF .debug_line emitter already exists (--debug-line, shipped v0.12.0) for source-line stepping into generated code; richer variable-location info is the gated Tier-2 follow-on (VCR-DBG-001, blocked on the register allocator).

Activity

  1. avrabe commented on Jun 24, 2026

    @avrabe
    ContributorAuthor

    [issue-hunt loop] Scoped — spike in PR #469 (no codegen change, frozen-safe).

    Confirmed empirically + found the load-bearing result: the two codegen paths diverge, and the relocatable path already implements the target shape.

    Base-materialization count (movw ip,#<base_lo>), cortex-m4:

    fixture optimized (non-reloc) relocatable
    7-store repro 7 (6 redundant) 0
    flight_seam 21 0
    flight_seam_flat 42 0
    • Relocatable (select_with_stack): base is a relocated symbol → pinned in fp once, each store [fp,#off]. Already optimal.
    • Optimized (optimizer_bridge.rs): base is an absolute constant → re-materialized into ip/r12 per store. ip is encoder scratch (can't persist across the indexed-load expansion), so the path has no persistent base register today.

    ⇒ fix is allocator-aware (reserve a callee-saved base reg + hoist movw/movt once per const-base store-run, mirroring the relocatable [base,#off] shape) — not a post-pass peephole. ~48 B + ~12 cyc on the 7-store fixture; scales with run length.

    Spike (PR #469) commits the generic fixture + the measurement/fix-shape note, frozen-safe. The byte-changing CSE (SYNTH_BASE_CSE flag-off → differential → on-target gate → flip) is the separate next gated step.

    — issue-hunt loop (avrabe automation), not a maintainer decision.

  2. avrabe commented on Jul 15, 2026

    @avrabe
    ContributorAuthor

    Verified fixed on v0.42.0: three consecutive i32.store (i32.const …) on the direct/self-contained path now emit one movw/movt base materialization (2 instructions total), reused for all three stores — the re-materialization pattern this issue reported is gone (landed across the const-CSE #514/#516 + base-CSE #482 levers). Repro used: 3-store .wat, --cortex-m -t cortex-m3, objdump count. Closing; reopen with a .wat if a shape still re-materializes. (— autonomous issue-hunt/release loop, posting under the shared avrabe account.)

  3. added a commit that references this issue on Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions