Skip to content

RQ-62-REACH increment 1 (#242): reachable-corpus acceptance census — denominator first, ranked blockers, honest non-comparability to the 805 - #1156

Merged
avrabe merged 4 commits into
mainfrom
feat/reach-census-1017
Sep 5, 2026
Merged

avrabe merged 4 commits into
mainfrom
feat/reach-census-1017

Conversation

@avrabe

@avrabe avrabe commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

RQ-62-REACH increment 1 — measurement only, report and stop

Refs #242, #1017

The artifact's increment 1 asks for the 805-module acceptance census re-run on v0.61.0. The 805 denominator is not reproducible on this machine — the census script's own # ci-status header says its input is a local corpus CI does not carry, and the wasm.directory component stratum plus the external toolchain-output modules are absent here. gale owns the full corpus and was asked for a re-run on #1017. So this PR does the honest subset: census what exists, and state the denominator as prominently as the rates. No blocker is fixed, no capability implemented.

The denominator (stated before any rate)

243 unique modules by sha256 (131 core + 112 components, 447 MB): every *.wasm reachable on this machine across .rivet/repos/{gale,kiln,loom,meld,scry,sigil}, the sibling ../loom and ../kiln checkouts, and synth itself. Per source: kiln 124, meld 56, synth 19, scry 16, loom sibling 15, gale 6, loom rivet-mirror 6, kiln rivet-mirror 1. This corpus is org-built fixtures and composites. Rates over it are NOT comparable to the 805 figures and no delta against 66%/14%/1.6% is computed anywhere in this PR. The done-when's "805-module census" leg is explicitly not satisfied; the artifact says so and stays proposed.

Per-backend, accepted / partial / declined / errored (synth 0.61.0, --all-exports --relocatable, zero timeouts)

backend core (131) components (112) total accepted
arm 21 / 0 / 110 / 0 6 / 0 / 106 / 0 27/243 (11%)
riscv 35 / 0 / 96 / 0 5 / 0 / 107 / 0 40/243 (16%)
aarch64 41 / 0 / 90 / 0 5 / 0 / 107 / 0 46/243 (19%)

Ranked blocker histogram (one primary blocker per non-accepted module; #952/#1102 policy refusals attributed to the modal per-function skip reason behind them; each histogram sums to its stratum's decline count)

What moved, only where comparison is honest

The one comparable stratum is v0.59's 2026-08-25 run (same script, same flags, same org-repos-on-this-machine method; snapshot differs 307→243, overlap unquantifiable): arm 81% → 11% accepted — and the fall is the v0.59–v0.61 honesty hardening doing its job. The #1041 active-data refusal, #1052 global-init refusal, and #1102 dangling-reloc refusal converted silent-drop "accepts" (125 partial objects, 77 accepts with silently-dropped data in the v0.59 run) into loud declines. An ACCEPT today excludes every known silent-drop class. Reading the fall as a capability regression compares an honest number to a flattered one — the compliance-envelope rule exists for exactly this.

Also worth naming: aarch64 (1.6% on the 805) is now the highest-accepting backend on this corpus post-RQ-60-A64IMPORT, and multi-memory blocks 50/112 components on riscv+aarch64 but zero on arm (VCR-MEM-002 phase 1 is ARM-only) — those same components then decline on arm for deeper causes, so no single capability un-blocks the component stratum.

Changes

  1. scripts/repro/partial_census_1017.py — four-bucket summary + ranked blocker histogram (instance lists collapsed so one cause is one bucket; A declined export exits 0 with the symbol simply absent — the decline is honest to stderr but invisible to the build system #952/rv32: a retained function relocating against a declined INTERNAL emits a dangling synth_func_N and exits 0 — the #851 guard #1013 gave aarch64 is not on the RISC-V path #1102 refusals attributed to their root-cause skip reasons). Still ci-status: manual: no expected values, no verdict.
  2. artifacts/release-v0.62/RQ-62-REACH.yaml — the census result published in the artifact, denominator first.

Gates

  • cargo fmt --check 0; cargo clippy --workspace --all-targets -- -D warnings 0 (no Rust touched)
  • claim_check 0, oracle_wiring_check 0
  • status_evidence_check: R4 asks the artifact to acknowledge the increment — fields.landed is filled with this PR's number in the follow-up commit on this branch
  • rivet validate: 40 errors / 298 warnings identical before and after this diff (pre-existing on main; none introduced)

🤖 Generated with Claude Code

https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

avrabe and others added 4 commits September 5, 2026 16:58
…mary and a ranked blocker histogram

partial_census_1017.py gains (e): per stratum, an accepted / partial /
declined / errored line with the denominator printed beside it, and a
RANKED BLOCKER HISTOGRAM — one PRIMARY blocker per non-accepted module
(the normalized module-level error for a module-level decline; the modal
per-function skip reason for a skip-only decline, ties broken
lexicographically so the ranking is deterministic). primary_blocker is
also attached to the --json records so downstream ranking needs no
re-parse.

Still a MEASUREMENT (ci-status: manual preserved): no expected values,
no verdict, nothing for CI to fail on. The behaviour it measures stays
gated by the wired decline-honesty oracles.

Refs #242, #1017
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
…apse instance lists, attribute #952/#1102 refusals to the skip reason behind them

First run over the org corpus fragmented the histogram: the #1102
dangling-reloc symbol lists, #952 skipped-export name lists, and global-
initializer dumps differ per module while naming the same class, so one
cause read as forty singleton buckets. And the #952/#1102 rows are POLICY
refusals — the capability actually missing is the per-function skip
reason synth's own stderr names behind them (e.g. 70 of arm's core
declines said "#952: _start skipped"; the modal cause behind them names
the real blocker).

Changes, all measurement-side (ci-status: manual preserved, no verdict):
- collapse_instance_lists(): symbol/export/initializer payloads collapse
  so one CAUSE buckets as one row (full text kept in --json records)
- DECLINE_MODULE_LEVEL records now carry the plain run's per-function
  skip reasons; primary_blocker attributes the #952/#1102 refusal
  classes to the modal skip reason when stderr names one
- _modal(): deterministic tie-break shared by both attribution paths

Refs #242, #1017
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
… denominator first, rates second

The 805-module denominator is NOT reproducible on this machine (the
census script's own header says its input is a local corpus CI does not
carry; the wasm.directory stratum and external toolchain outputs are
absent here), so the honest increment is: census what exists and state
the denominator as prominently as the rates. The full 805 re-run remains
gale's (asked on #1017).

Published into the artifact:
- DENOMINATOR: 243 unique modules by sha256 (131 core + 112 components,
  447 MB), every *.wasm reachable across the org repos on this machine,
  with per-source provenance and the missing strata named. No delta
  against 66%/14%/1.6% is computed — different corpus, and the artifact
  says so instead of flattering.
- PER-BACKEND: arm 27/243 (11%), riscv 40/243 (16%), aarch64 46/243
  (19%); four buckets per stratum, zero timeouts.
- RANKED BLOCKERS per backend and stratum (histograms sum to their
  decline counts). Headlines: arm's top blockers are its own honesty
  refusals' root causes (i64 imm12 offsets 46, global-init 26,
  active-data 45 across strata, AAPCS pairs 40); riscv's is GlobalGet
  (132 across strata) then multi-memory (51); aarch64's are MemoryCopy
  (45), call value-stack discipline (77), multi-memory (51).
- WHAT MOVED, only where comparable (v0.59's same-stratum arm run):
  81% -> 11% accepted, and the fall is the v0.59-v0.61 silent-drop-to-
  loud-decline hardening — an ACCEPT today excludes every known
  silent-drop class; the old number included 125 partial objects and 77
  silently-dropped-data accepts. Reading it as a capability regression
  compares an honest number to a flattered one.
- INVERSION: aarch64 (1.6% on the 805) is the HIGHEST-accepting backend
  on this corpus post-RQ-60-A64IMPORT; multi-memory blocks 50/112
  components on riscv+aarch64 but zero on arm (VCR-MEM-002 phase 1 is
  ARM-only) — and those components still decline on arm for deeper
  causes, so no single capability un-blocks the component stratum.

Status stays proposed: the done-when's "805-module census" leg is
explicitly NOT satisfied by this subset and the artifact says so.
Measurement only — no blocker fixed, no capability implemented.

Refs #242, #1017
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
…1156)

Status stays proposed — the done-when's 805-module leg is not satisfied
by the reachable subset and the landed note says so. Greens the R4
status-evidence acknowledgment for the increment's delivery commits.

Refs #242, #1017
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
@codecov

codecov Bot commented Sep 5, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@avrabe
avrabe merged commit 36ae130 into main Sep 5, 2026
60 checks passed
@avrabe
avrabe deleted the feat/reach-census-1017 branch September 5, 2026 19:49
avrabe added a commit that referenced this pull request Sep 6, 2026
…, and make the NOTE 5 carry a reference

Two consistency fixes, both prompted by the cold review's "unverifiable as
stated" section rather than by a finding.

1. THE 17-HOUR FIGURE. The review correctly listed "the federated-graph job
   validated nothing for 17 hours" as a CI-history claim not derivable from
   the repo. I went looking: ci.yml run history shows the 09-03 main runs and
   the fix landing, but the federated job is ADVISORY, so per-job validation
   history is not reconstructable from run conclusions. The figure is my own
   contemporaneous observation recorded on #1143 — real, but not checkable by
   a later reader.

   That is the SAME epistemic status as the v0.59 partial/silent-drop
   subcounts, which this release already handles by ATTRIBUTING them rather
   than asserting them (NOTE 3). Treating the two differently would be
   arbitrary, so the federated claim is now attributed the same way: the
   window is named as something #1143 recorded, not as something the notes
   assert. The FIX itself is verified and unchanged — only the duration was
   ever taken on faith.

2. NOTE 5's CARRY IS NOW A REFERENCE, NOT A PROMISE. Filed as #1159. The
   disposition previously said "carried to v0.63", which is exactly the kind
   of claim this project does not accept from anyone else. The issue also
   records why it is more than cosmetic: the histogram's job is RANKING, and
   a cause fragmented across N buckets is systematically under-ranked against
   a cause with no varying payload — which is the input v0.63 plans to
   prioritise from.

Refs #1143, #1156, #1159

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
avrabe added a commit that referenced this pull request Sep 6, 2026
…lemented (#1158)

* release(v0.62.0): version bump, pin sweep, and release notes derived from merged code

Theme: "Reach is part of correctness."

Version 0.61.0 -> 0.62.0 across Cargo.toml (workspace + 10 path-dep pins),
MODULE.bazel, npm/package.json, artifacts/status.json and Cargo.lock. Pin
sweep green.

CHANGELOG [0.62.0] written against MERGED CODE and claims.yaml ON THE TREE,
not against PR bodies — the v0.57 cold review found four errors in
same-author release prose derived the other way. Every load-bearing number
re-derived in this session:

  - census table + ranked blockers: artifacts/release-v0.62/RQ-62-REACH.yaml
  - ratchet deltas: claims.yaml at v0.61.0 tag vs HEAD (selector_lines_code
    19213 -> 19227, +14, one new waiver bound to the exact value; every
    other pin flat)
  - 630 Qed / 2 Admitted: coq/STATUS.md
  - 324,646 emulations / 158 wired scripts: oracle_wiring_check.py
  - MIN_DERIVED_SLOTS=4, MIN_ONLINE=2, 30 loop-conformance unit tests,
    `runs-on: [self-hosted, linux, x64, light]`: read out of the shipped
    scripts and ci.yml, not from the commit prose that claimed them

The notes state two things the release would rather not say: it SUBTRACTED
NOTHING (6,871 insertions / 22 deletions) and the subtraction ratchet moved
the wrong way, waived, with the reason printed. A reach-and-gates release
should look like one in the metric that exists to detect it.

docs/status/FEATURE_MATRIX.md + artifacts/status.json regenerated via
`claim_check.py claims.yaml --emit-status`; claim gate 58/58.

Refs #242, #1017, #1062, #1102, #1131, #1132, #1133, #1136, #1143, #1145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* fix(ci): RQ-62-FLOORTIGHT (#910) — the summed emulation ratchet enforced 1,892 below what it declared

Found while cutting v0.62, on THIS RELEASE'S OWN NEW ORACLE, and confirmed
independently by the release's cold review.

THE DEFECT. `oracle_wiring_check.py --min-emulation-floor N` is the #910
ratchet; its own failure text states its purpose — "An oracle lost
execution, or a floor was lowered." It enforced 322754 against a DECLARED
total of 324640 at v0.61.0 and 324646 here: 1,892 emulations of slack,
more than most single oracles declare. So an oracle could stay present,
stay wired, stay referenced, and drop to ZERO executions while the gate
built to notice exactly that reported green.

DEMONSTRATED in an isolated worktree, not reasoned about:

  mutation 1  delete mem_isolation_red_1145.py whole
              -> exit 1, but by the DANGLING-CI-REFERENCE check, NOT the
                 floor (total fell only to 324640, still above 322754).
                 Two mechanisms, two different failures; only one is the
                 ratchet's job and it was the one that could not fire.
  mutation 2  oracle stays wired, declares `emulations >= 0`
              -> exit 0. GREEN. The ratchet's stated failure mode, blind.
  after fix   same mutation -> exit 1, "RATCHET BROKEN: summed floors
                 324640 < recorded minimum 324646"; unmutated tree exit 0,
                 so the floor is CORRECT, not merely stricter.

This is #1113 one level up: that was INTRA-oracle (a decline half
contributing zero to its own floor, fixed per-oracle in #1112); this is
INTER-oracle (the summed ratchet below the summed declarations). Same
sentence — the floor cannot see part of what it asserts — at two scales,
one release apart. The per-oracle fix was correct and did not imply the
aggregate was tight; nobody checked, because every gate reporting on this
floor reports it PRESENT AND WIRED, never TIGHT.

The slack was INHERITED, never introduced — so the fix is an INVARIANT,
not a number: enforced floor EQUALS declared total, pinned verbatim, a
landing oracle bumps both in the same PR. claims.yaml's
SYNTH-ORACLE-CHECK-FLOORS-910-CI pin caught the edit at 57/58 before it
could ship silently.

ALSO IN THIS COMMIT — two cold-review corrections:

1. A FALSE CLAIM I WROTE, propagated to three files. Release prose said
   the dynamic index in the #1102-residual fixture was load-bearing
   because "a constant index devirtualizes and the table never
   materializes." FALSE. Compiling a const-index call_indirect on ARM
   --relocatable emits `movw r2, #0` / bounds check / `ldr.w ip, [fp, ip]`
   / `blx ip` — a runtime table load and an indirect call, same shape as
   the dynamic case; synth has NO devirtualization pass at all (verified
   against emitted bytes, and no such code exists in the tree). The guard
   is index-insensitive and refuses the const-index dangling shape too, so
   the fix and its red-first evidence are UNAFFECTED — what was wrong is
   the RATIONALE, which is what a future contributor reasons from.
   Corrected in CHANGELOG.md, RQ-62-TABLEDANGLE.yaml and
   tabledangle_1102_elem.rs, each stating the correction rather than
   quietly deleting the sentence.

2. R10 ATTRIBUTION. The floor fix first landed with no artifact or issue
   anchor in its subject, and status_evidence_check R10 — the gate added
   in #1124 for exactly this — failed it: "delivery-shaped commit in the
   release window is attributable to NO release artifact." It was right.
   Fixed properly rather than by loosening: RQ-62-FLOORTIGHT is now a real
   v0.62 artifact carrying the finding, and this subject names it.

Refs #910, #1102, #1113, #1124, #1145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* docs(review): commit the v0.62 independent cold review + disposition, and apply its four prose corrections

The clean-room review record loop_conformance_check step 7 requires, written
by an independent reviewer with no inherited framing, dispatched against
0a8c1e1. Its text is committed UNEDITED; the disposition table is additive.

WHAT IT FOUND: 1 MUST-FIX, 1 SHOULD-FIX, 5 NOTEs, 4 unverifiable-as-stated,
~20 claims independently verified by execution — including a full re-run of
the acceptance census from a freshly built binary (27/243, 40/243, 46/243
and every ranked blocker reproduced exactly), and a rebuild of the PRE-FIX
compiler at 858ff8d^ to reproduce the #1102-residual hole end-to-end
(object ships exit 0, links clean, zero UNDEF symbols).

BOTH ACTIONABLE FINDINGS WERE MINE, AND BOTH ARE FIXED:

  MUST-FIX 1  the tip failed its own required Claim Check (R10, the rule
              added in #1124 for exactly this shape). Fixed by giving the
              work the artifact it deserved — RQ-62-FLOORTIGHT — not by
              loosening the rule. Now exit 0, 2 delivery commits, 2
              attributed.

  SHOULD-FIX 1  a FALSE claim I wrote, propagated to three files: that a
              constant call_indirect index devirtualizes so the table never
              materializes. Re-verified against emitted bytes before acting
              rather than taking the review's word — `movw r2,#0` / bounds
              check / `ldr.w ip,[fp,ip]` / `blx ip`. It does not
              devirtualize; synth has no such pass. Corrected in all three
              places AS A STATED CORRECTION, so the wrong rationale cannot
              be re-derived from history. The fix and its red-first evidence
              are unaffected — only the rationale was wrong, which is the
              dangerous shape: a correct artifact resting on a reason a
              future contributor would reason from.

NOTES 1, 3, 4 applied to the CHANGELOG, each narrowing a sentence that
claimed more than the code or the data supports:
  - the diff-stat's scope is now stated (feature merges; release commit
    +213/-51 named separately)
  - the 81%->11% comparison now says the denominator differs (307 vs 243),
    that the subcounts come from the v0.59 artifact rather than this
    release, that per-module records were not preserved, and that the two
    subcounts may overlap
  - "fail-closed on unknown mnemonics" now says what the code does:
    undecodable instructions and zero-instruction scans always refuse; an
    unknown MNEMONIC refuses only when its operand text names a reserved
    register

NOTE 5 (census hex-immediate collapse) is deliberately NOT fixed here and
the disposition says why: no number in the release is wrong, and changing a
measurement script mid-cut, under the oracle it feeds, risks the census for
no gain. Carried to v0.63, where re-measurement is increment 1.

Step 8's PR-head-vs-merge attestation slot is stubbed in the record with
its rationale; it is filled at merge time. The gate labels it ATTESTED in
both modes, so filling it is a deliberate act — an attestation with nothing
behind it is the vacuity this release spent its scope finding.

Refs #910, #1102, #1124, #1136

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* docs: attribute the one release-notes number that is not re-derivable, and make the NOTE 5 carry a reference

Two consistency fixes, both prompted by the cold review's "unverifiable as
stated" section rather than by a finding.

1. THE 17-HOUR FIGURE. The review correctly listed "the federated-graph job
   validated nothing for 17 hours" as a CI-history claim not derivable from
   the repo. I went looking: ci.yml run history shows the 09-03 main runs and
   the fix landing, but the federated job is ADVISORY, so per-job validation
   history is not reconstructable from run conclusions. The figure is my own
   contemporaneous observation recorded on #1143 — real, but not checkable by
   a later reader.

   That is the SAME epistemic status as the v0.59 partial/silent-drop
   subcounts, which this release already handles by ATTRIBUTING them rather
   than asserting them (NOTE 3). Treating the two differently would be
   arbitrary, so the federated claim is now attributed the same way: the
   window is named as something #1143 recorded, not as something the notes
   assert. The FIX itself is verified and unchanged — only the duration was
   ever taken on faith.

2. NOTE 5's CARRY IS NOW A REFERENCE, NOT A PROMISE. Filed as #1159. The
   disposition previously said "carried to v0.63", which is exactly the kind
   of claim this project does not accept from anyone else. The issue also
   records why it is more than cosmetic: the histogram's job is RANKING, and
   a cause fragmented across N buckets is systematically under-ranked against
   a cause with no varying payload — which is the input v0.63 plans to
   prioritise from.

Refs #1143, #1156, #1159

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
avrabe added a commit that referenced this pull request Sep 7, 2026
…ng — arm's non-rotated-immediate class was rank 5 at 4, is rank 3 at 9

Measurement (full 243-module corpus, MANIFEST-verified, synth at 8862ee2 =
the v0.63 compiler, BEFORE script-on-main vs AFTER, verdicts/rungs identical
between runs, 0 TIMEOUT): the fix CHANGED the arm ranking and nothing else.

  ladder NEVER (the table v0.63 was scoped from, 141 of 243, rungs 27/72/3):
    encode_operand2 non-rotated immediate  4 @ rank 5  ->  9 @ rank 3
    (above start-section 7 and rule_i32_rotl 5); GI-FPU-002 3 -> 1; the top
    two (register exhaustion 71->70, #929 41->40) keep their ranks.
  plain census: core #929 16->15 loses its tie for 3rd; components gain an
    encode_operand2 row (0 -> 3) above GI-FPU-002 (4 -> 2).
  riscv, aarch64: identical before/after in both modes — their records carry
    no hex payload at all (surveyed), so the no-op is by construction.

Why 9 and not 5: `_modal()` is a per-module plurality vote. In four modules
the fragments each lost that vote to a cause with no varying payload
(yolo_inference_{release,debug}: 84 functions across ~20 immediates, largest
fragment 6, vs 58 for GI-FPU-002; sockets-tcp-connect 10 vs 8; stat-dev-ino
5 vs 4). That is the under-ranking one level down from the row split.

The ranked blocker histogram had no cut to rise from; the census's SECOND
ranking ("module-level decline reasons (top 10)") ranked RAW text and its
top-10 cut silently dropped 56/110, 65/106, 60/78, 35/103, 58/86, 42/107
modules per stratum — routed through the one module_reason pipeline now,
with a remainder row.

Spec census: RESULT PASS, PINS / FAMILY_OK_PINS / AT_LEAST_ONE_EXPORT
unchanged (it emits exact-pinned bucket counts, no histogram).

Wiring (coordinator review): the fix would have shipped a self-test CI never
runs. The claims.yaml note beside this file's `manual` slot already said
"if the census ever grows an expected value or a pass/fail verdict ... it
must be wired" — so `# ci-status: wired` with `# ci-checks: stdout` binding
the assertion COUNT (14), a step in the required claim-check job via
oracle_run plus the per-job evidence ledger; manual ceiling 8 -> 7 in
claims.yaml and ORACLE_WIRING.md. Red-first both ways: hex mask reverted ->
assertion 2 fails (exit 1); evidence line renamed -> script exits 0 but
oracle_run refuses on the floor (measured 0 of 14). Assertion 1 asserts the
PRE-#1159 rule still keeps the two wild shapes apart, so the discriminator
cannot go vacuous silently. The census MODES are unchanged: still a local
measurement, no verdict about synth. Emulation floor re-derived: 324845.

Also: reference identifiers (#1102, §4.5.5, VCR-MEM-002) are kept by the
mask — format-string constants cannot vary per instance, masking them only
loses identity; the published tables were restoring them by hand.

docs/status/ACCEPTANCE_LADDER.md gains the re-attributed arm table under
the v0.63 one it corrects; the shipped measurement is not rewritten.

Refs #1159, #242, #1156

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
avrabe added a commit that referenced this pull request Sep 7, 2026
…sert the sum invariant, wire the self-test — arm's non-rotated-immediate gap is rank 3 at 9, not rank 5 at 4 (#1181)

* fix(census): RQ-64-HISTOGRAM (#1159) — mask HEX numeric payloads so one cause is one row, and assert every ranked view sums to what it ranks

WIP checkpoint committed by the coordinator to preserve the lane's work while
its arm ladder quarters finish. The lane owns the remaining measurement.

`\b\d+\b` left `0x5dc` intact (the `x` glues the digits into one word), so
`immediate 0x624 (1500)` normalized to `immediate 0x624 (N)` — the decimal
collapsed, the hex did not, and one cause fragmented into one row per distinct
immediate. A cause with a varying payload was systematically UNDER-RANKED
against one without, in the very histogram v0.63 was scoped from.

Hex is now masked first, to a distinct token (`0xN` vs `N`) so the message
shape stays readable. Identifiers that merely contain digits (`i32`, `R11`,
`func_25`, `RV32`) are not word-bounded numbers and survive; reference
identifiers carried as literal format-string text (`#1102`, `§4.5.5`,
`VCR-MEM-002`) are kept whole, since they cannot vary per instance and masking
them could only lose identity. The mask is idempotent, so a `--json` record's
already-normalized `skip_reasons` can be re-ranked offline without
re-fragmenting.

Every ranked view now asserts its printed rows sum to the modules it ranks — a
collapse that LOSES rows is worse than one that fragments them. `--self-test`
exercises both wild shapes and proves the sum check can fail (negative
control).

Refs #1159

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* RQ-64-HISTOGRAM (#1159): wire the census self-test, measure the ranking — arm's non-rotated-immediate class was rank 5 at 4, is rank 3 at 9

Measurement (full 243-module corpus, MANIFEST-verified, synth at 8862ee2 =
the v0.63 compiler, BEFORE script-on-main vs AFTER, verdicts/rungs identical
between runs, 0 TIMEOUT): the fix CHANGED the arm ranking and nothing else.

  ladder NEVER (the table v0.63 was scoped from, 141 of 243, rungs 27/72/3):
    encode_operand2 non-rotated immediate  4 @ rank 5  ->  9 @ rank 3
    (above start-section 7 and rule_i32_rotl 5); GI-FPU-002 3 -> 1; the top
    two (register exhaustion 71->70, #929 41->40) keep their ranks.
  plain census: core #929 16->15 loses its tie for 3rd; components gain an
    encode_operand2 row (0 -> 3) above GI-FPU-002 (4 -> 2).
  riscv, aarch64: identical before/after in both modes — their records carry
    no hex payload at all (surveyed), so the no-op is by construction.

Why 9 and not 5: `_modal()` is a per-module plurality vote. In four modules
the fragments each lost that vote to a cause with no varying payload
(yolo_inference_{release,debug}: 84 functions across ~20 immediates, largest
fragment 6, vs 58 for GI-FPU-002; sockets-tcp-connect 10 vs 8; stat-dev-ino
5 vs 4). That is the under-ranking one level down from the row split.

The ranked blocker histogram had no cut to rise from; the census's SECOND
ranking ("module-level decline reasons (top 10)") ranked RAW text and its
top-10 cut silently dropped 56/110, 65/106, 60/78, 35/103, 58/86, 42/107
modules per stratum — routed through the one module_reason pipeline now,
with a remainder row.

Spec census: RESULT PASS, PINS / FAMILY_OK_PINS / AT_LEAST_ONE_EXPORT
unchanged (it emits exact-pinned bucket counts, no histogram).

Wiring (coordinator review): the fix would have shipped a self-test CI never
runs. The claims.yaml note beside this file's `manual` slot already said
"if the census ever grows an expected value or a pass/fail verdict ... it
must be wired" — so `# ci-status: wired` with `# ci-checks: stdout` binding
the assertion COUNT (14), a step in the required claim-check job via
oracle_run plus the per-job evidence ledger; manual ceiling 8 -> 7 in
claims.yaml and ORACLE_WIRING.md. Red-first both ways: hex mask reverted ->
assertion 2 fails (exit 1); evidence line renamed -> script exits 0 but
oracle_run refuses on the floor (measured 0 of 14). Assertion 1 asserts the
PRE-#1159 rule still keeps the two wild shapes apart, so the discriminator
cannot go vacuous silently. The census MODES are unchanged: still a local
measurement, no verdict about synth. Emulation floor re-derived: 324845.

Also: reference identifiers (#1102, §4.5.5, VCR-MEM-002) are kept by the
mask — format-string constants cannot vary per instance, masking them only
loses identity; the published tables were restoring them by hand.

docs/status/ACCEPTANCE_LADDER.md gains the re-attributed arm table under
the v0.63 one it corrects; the shipped measurement is not rewritten.

Refs #1159, #242, #1156

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* chore(rivet): RQ-64-HISTOGRAM landed: names PR #1181

Refs #1159

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

* RQ-64-HISTOGRAM (#1159): reconcile 135 / 138 / 141 — the pre-fix tool printed all twelve NEVER rows; 138 and 135 are transcription cuts, not tooling losses

Replayed the PRE-FIX script's own report_ladder over the exact BEFORE NEVER
set (each module's stored pre-fix blocker, no synth re-run): twelve rows
summing to 141. The held nine-row listing (138) is its first nine rows; the
three missing modules are rows 10-12 (LdrSym Thumb-N-only, encode_operand2
0x5dc, no exports). Both v0.63.0 and pre-fix main print most_common(12), so
the shipped tool lost nothing — the new sum assertion correctly does NOT go
red on the pre-fix bucketing; a cut forced to top=9 now prints a remainder
row and still closes at 141. No third in-tool loss mechanism exists: 0 of
141 NEVER modules have an empty primary blocker. Loss mechanisms named:
the CUT hides rows (closes now), FRAGMENTATION misranks (masked now),
TRANSCRIPTION drops rows outside the tool (135, 138) — checkable now
against the printed 'N rows sum to M' line.

Also: the full-corpus single-pass arm ladders finished and agree with the
quarter merge on all 243 modules' rung and blocker, both sides (0 differ);
the artifact's denominator note is upgraded from a caveat to a cross-check.

Refs #1159

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant