Repository navigation
RQ-62-REACH increment 1 (#242): reachable-corpus acceptance census — denominator first, ranked blockers, honest non-comparability to the 805 - #1156
Merged
Conversation
…mary and a ranked blocker histogram partial_census_1017.py gains (e): per stratum, an accepted / partial / declined / errored line with the denominator printed beside it, and a RANKED BLOCKER HISTOGRAM — one PRIMARY blocker per non-accepted module (the normalized module-level error for a module-level decline; the modal per-function skip reason for a skip-only decline, ties broken lexicographically so the ranking is deterministic). primary_blocker is also attached to the --json records so downstream ranking needs no re-parse. Still a MEASUREMENT (ci-status: manual preserved): no expected values, no verdict, nothing for CI to fail on. The behaviour it measures stays gated by the wired decline-honesty oracles. Refs #242, #1017 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
…apse instance lists, attribute #952/#1102 refusals to the skip reason behind them First run over the org corpus fragmented the histogram: the #1102 dangling-reloc symbol lists, #952 skipped-export name lists, and global- initializer dumps differ per module while naming the same class, so one cause read as forty singleton buckets. And the #952/#1102 rows are POLICY refusals — the capability actually missing is the per-function skip reason synth's own stderr names behind them (e.g. 70 of arm's core declines said "#952: _start skipped"; the modal cause behind them names the real blocker). Changes, all measurement-side (ci-status: manual preserved, no verdict): - collapse_instance_lists(): symbol/export/initializer payloads collapse so one CAUSE buckets as one row (full text kept in --json records) - DECLINE_MODULE_LEVEL records now carry the plain run's per-function skip reasons; primary_blocker attributes the #952/#1102 refusal classes to the modal skip reason when stderr names one - _modal(): deterministic tie-break shared by both attribution paths Refs #242, #1017 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
… denominator first, rates second The 805-module denominator is NOT reproducible on this machine (the census script's own header says its input is a local corpus CI does not carry; the wasm.directory stratum and external toolchain outputs are absent here), so the honest increment is: census what exists and state the denominator as prominently as the rates. The full 805 re-run remains gale's (asked on #1017). Published into the artifact: - DENOMINATOR: 243 unique modules by sha256 (131 core + 112 components, 447 MB), every *.wasm reachable across the org repos on this machine, with per-source provenance and the missing strata named. No delta against 66%/14%/1.6% is computed — different corpus, and the artifact says so instead of flattering. - PER-BACKEND: arm 27/243 (11%), riscv 40/243 (16%), aarch64 46/243 (19%); four buckets per stratum, zero timeouts. - RANKED BLOCKERS per backend and stratum (histograms sum to their decline counts). Headlines: arm's top blockers are its own honesty refusals' root causes (i64 imm12 offsets 46, global-init 26, active-data 45 across strata, AAPCS pairs 40); riscv's is GlobalGet (132 across strata) then multi-memory (51); aarch64's are MemoryCopy (45), call value-stack discipline (77), multi-memory (51). - WHAT MOVED, only where comparable (v0.59's same-stratum arm run): 81% -> 11% accepted, and the fall is the v0.59-v0.61 silent-drop-to- loud-decline hardening — an ACCEPT today excludes every known silent-drop class; the old number included 125 partial objects and 77 silently-dropped-data accepts. Reading it as a capability regression compares an honest number to a flattered one. - INVERSION: aarch64 (1.6% on the 805) is the HIGHEST-accepting backend on this corpus post-RQ-60-A64IMPORT; multi-memory blocks 50/112 components on riscv+aarch64 but zero on arm (VCR-MEM-002 phase 1 is ARM-only) — and those components still decline on arm for deeper causes, so no single capability un-blocks the component stratum. Status stays proposed: the done-when's "805-module census" leg is explicitly NOT satisfied by this subset and the artifact says so. Measurement only — no blocker fixed, no capability implemented. Refs #242, #1017 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
…1156) Status stays proposed — the done-when's 805-module leg is not satisfied by the reachable subset and the landed note says so. Greens the R4 status-evidence acknowledgment for the increment's delivery commits. Refs #242, #1017 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
avrabe
added a commit
that referenced
this pull request
Sep 6, 2026
…, and make the NOTE 5 carry a reference Two consistency fixes, both prompted by the cold review's "unverifiable as stated" section rather than by a finding. 1. THE 17-HOUR FIGURE. The review correctly listed "the federated-graph job validated nothing for 17 hours" as a CI-history claim not derivable from the repo. I went looking: ci.yml run history shows the 09-03 main runs and the fix landing, but the federated job is ADVISORY, so per-job validation history is not reconstructable from run conclusions. The figure is my own contemporaneous observation recorded on #1143 — real, but not checkable by a later reader. That is the SAME epistemic status as the v0.59 partial/silent-drop subcounts, which this release already handles by ATTRIBUTING them rather than asserting them (NOTE 3). Treating the two differently would be arbitrary, so the federated claim is now attributed the same way: the window is named as something #1143 recorded, not as something the notes assert. The FIX itself is verified and unchanged — only the duration was ever taken on faith. 2. NOTE 5's CARRY IS NOW A REFERENCE, NOT A PROMISE. Filed as #1159. The disposition previously said "carried to v0.63", which is exactly the kind of claim this project does not accept from anyone else. The issue also records why it is more than cosmetic: the histogram's job is RANKING, and a cause fragmented across N buckets is systematically under-ranked against a cause with no varying payload — which is the input v0.63 plans to prioritise from. Refs #1143, #1156, #1159 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
avrabe
added a commit
that referenced
this pull request
Sep 6, 2026
…lemented (#1158) * release(v0.62.0): version bump, pin sweep, and release notes derived from merged code Theme: "Reach is part of correctness." Version 0.61.0 -> 0.62.0 across Cargo.toml (workspace + 10 path-dep pins), MODULE.bazel, npm/package.json, artifacts/status.json and Cargo.lock. Pin sweep green. CHANGELOG [0.62.0] written against MERGED CODE and claims.yaml ON THE TREE, not against PR bodies — the v0.57 cold review found four errors in same-author release prose derived the other way. Every load-bearing number re-derived in this session: - census table + ranked blockers: artifacts/release-v0.62/RQ-62-REACH.yaml - ratchet deltas: claims.yaml at v0.61.0 tag vs HEAD (selector_lines_code 19213 -> 19227, +14, one new waiver bound to the exact value; every other pin flat) - 630 Qed / 2 Admitted: coq/STATUS.md - 324,646 emulations / 158 wired scripts: oracle_wiring_check.py - MIN_DERIVED_SLOTS=4, MIN_ONLINE=2, 30 loop-conformance unit tests, `runs-on: [self-hosted, linux, x64, light]`: read out of the shipped scripts and ci.yml, not from the commit prose that claimed them The notes state two things the release would rather not say: it SUBTRACTED NOTHING (6,871 insertions / 22 deletions) and the subtraction ratchet moved the wrong way, waived, with the reason printed. A reach-and-gates release should look like one in the metric that exists to detect it. docs/status/FEATURE_MATRIX.md + artifacts/status.json regenerated via `claim_check.py claims.yaml --emit-status`; claim gate 58/58. Refs #242, #1017, #1062, #1102, #1131, #1132, #1133, #1136, #1143, #1145 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * fix(ci): RQ-62-FLOORTIGHT (#910) — the summed emulation ratchet enforced 1,892 below what it declared Found while cutting v0.62, on THIS RELEASE'S OWN NEW ORACLE, and confirmed independently by the release's cold review. THE DEFECT. `oracle_wiring_check.py --min-emulation-floor N` is the #910 ratchet; its own failure text states its purpose — "An oracle lost execution, or a floor was lowered." It enforced 322754 against a DECLARED total of 324640 at v0.61.0 and 324646 here: 1,892 emulations of slack, more than most single oracles declare. So an oracle could stay present, stay wired, stay referenced, and drop to ZERO executions while the gate built to notice exactly that reported green. DEMONSTRATED in an isolated worktree, not reasoned about: mutation 1 delete mem_isolation_red_1145.py whole -> exit 1, but by the DANGLING-CI-REFERENCE check, NOT the floor (total fell only to 324640, still above 322754). Two mechanisms, two different failures; only one is the ratchet's job and it was the one that could not fire. mutation 2 oracle stays wired, declares `emulations >= 0` -> exit 0. GREEN. The ratchet's stated failure mode, blind. after fix same mutation -> exit 1, "RATCHET BROKEN: summed floors 324640 < recorded minimum 324646"; unmutated tree exit 0, so the floor is CORRECT, not merely stricter. This is #1113 one level up: that was INTRA-oracle (a decline half contributing zero to its own floor, fixed per-oracle in #1112); this is INTER-oracle (the summed ratchet below the summed declarations). Same sentence — the floor cannot see part of what it asserts — at two scales, one release apart. The per-oracle fix was correct and did not imply the aggregate was tight; nobody checked, because every gate reporting on this floor reports it PRESENT AND WIRED, never TIGHT. The slack was INHERITED, never introduced — so the fix is an INVARIANT, not a number: enforced floor EQUALS declared total, pinned verbatim, a landing oracle bumps both in the same PR. claims.yaml's SYNTH-ORACLE-CHECK-FLOORS-910-CI pin caught the edit at 57/58 before it could ship silently. ALSO IN THIS COMMIT — two cold-review corrections: 1. A FALSE CLAIM I WROTE, propagated to three files. Release prose said the dynamic index in the #1102-residual fixture was load-bearing because "a constant index devirtualizes and the table never materializes." FALSE. Compiling a const-index call_indirect on ARM --relocatable emits `movw r2, #0` / bounds check / `ldr.w ip, [fp, ip]` / `blx ip` — a runtime table load and an indirect call, same shape as the dynamic case; synth has NO devirtualization pass at all (verified against emitted bytes, and no such code exists in the tree). The guard is index-insensitive and refuses the const-index dangling shape too, so the fix and its red-first evidence are UNAFFECTED — what was wrong is the RATIONALE, which is what a future contributor reasons from. Corrected in CHANGELOG.md, RQ-62-TABLEDANGLE.yaml and tabledangle_1102_elem.rs, each stating the correction rather than quietly deleting the sentence. 2. R10 ATTRIBUTION. The floor fix first landed with no artifact or issue anchor in its subject, and status_evidence_check R10 — the gate added in #1124 for exactly this — failed it: "delivery-shaped commit in the release window is attributable to NO release artifact." It was right. Fixed properly rather than by loosening: RQ-62-FLOORTIGHT is now a real v0.62 artifact carrying the finding, and this subject names it. Refs #910, #1102, #1113, #1124, #1145 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * docs(review): commit the v0.62 independent cold review + disposition, and apply its four prose corrections The clean-room review record loop_conformance_check step 7 requires, written by an independent reviewer with no inherited framing, dispatched against 0a8c1e1. Its text is committed UNEDITED; the disposition table is additive. WHAT IT FOUND: 1 MUST-FIX, 1 SHOULD-FIX, 5 NOTEs, 4 unverifiable-as-stated, ~20 claims independently verified by execution — including a full re-run of the acceptance census from a freshly built binary (27/243, 40/243, 46/243 and every ranked blocker reproduced exactly), and a rebuild of the PRE-FIX compiler at 858ff8d^ to reproduce the #1102-residual hole end-to-end (object ships exit 0, links clean, zero UNDEF symbols). BOTH ACTIONABLE FINDINGS WERE MINE, AND BOTH ARE FIXED: MUST-FIX 1 the tip failed its own required Claim Check (R10, the rule added in #1124 for exactly this shape). Fixed by giving the work the artifact it deserved — RQ-62-FLOORTIGHT — not by loosening the rule. Now exit 0, 2 delivery commits, 2 attributed. SHOULD-FIX 1 a FALSE claim I wrote, propagated to three files: that a constant call_indirect index devirtualizes so the table never materializes. Re-verified against emitted bytes before acting rather than taking the review's word — `movw r2,#0` / bounds check / `ldr.w ip,[fp,ip]` / `blx ip`. It does not devirtualize; synth has no such pass. Corrected in all three places AS A STATED CORRECTION, so the wrong rationale cannot be re-derived from history. The fix and its red-first evidence are unaffected — only the rationale was wrong, which is the dangerous shape: a correct artifact resting on a reason a future contributor would reason from. NOTES 1, 3, 4 applied to the CHANGELOG, each narrowing a sentence that claimed more than the code or the data supports: - the diff-stat's scope is now stated (feature merges; release commit +213/-51 named separately) - the 81%->11% comparison now says the denominator differs (307 vs 243), that the subcounts come from the v0.59 artifact rather than this release, that per-module records were not preserved, and that the two subcounts may overlap - "fail-closed on unknown mnemonics" now says what the code does: undecodable instructions and zero-instruction scans always refuse; an unknown MNEMONIC refuses only when its operand text names a reserved register NOTE 5 (census hex-immediate collapse) is deliberately NOT fixed here and the disposition says why: no number in the release is wrong, and changing a measurement script mid-cut, under the oracle it feeds, risks the census for no gain. Carried to v0.63, where re-measurement is increment 1. Step 8's PR-head-vs-merge attestation slot is stubbed in the record with its rationale; it is filled at merge time. The gate labels it ATTESTED in both modes, so filling it is a deliberate act — an attestation with nothing behind it is the vacuity this release spent its scope finding. Refs #910, #1102, #1124, #1136 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * docs: attribute the one release-notes number that is not re-derivable, and make the NOTE 5 carry a reference Two consistency fixes, both prompted by the cold review's "unverifiable as stated" section rather than by a finding. 1. THE 17-HOUR FIGURE. The review correctly listed "the federated-graph job validated nothing for 17 hours" as a CI-history claim not derivable from the repo. I went looking: ci.yml run history shows the 09-03 main runs and the fix landing, but the federated job is ADVISORY, so per-job validation history is not reconstructable from run conclusions. The figure is my own contemporaneous observation recorded on #1143 — real, but not checkable by a later reader. That is the SAME epistemic status as the v0.59 partial/silent-drop subcounts, which this release already handles by ATTRIBUTING them rather than asserting them (NOTE 3). Treating the two differently would be arbitrary, so the federated claim is now attributed the same way: the window is named as something #1143 recorded, not as something the notes assert. The FIX itself is verified and unchanged — only the duration was ever taken on faith. 2. NOTE 5's CARRY IS NOW A REFERENCE, NOT A PROMISE. Filed as #1159. The disposition previously said "carried to v0.63", which is exactly the kind of claim this project does not accept from anyone else. The issue also records why it is more than cosmetic: the histogram's job is RANKING, and a cause fragmented across N buckets is systematically under-ranked against a cause with no varying payload — which is the input v0.63 plans to prioritise from. Refs #1143, #1156, #1159 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
avrabe
added a commit
that referenced
this pull request
Sep 7, 2026
…ng — arm's non-rotated-immediate class was rank 5 at 4, is rank 3 at 9 Measurement (full 243-module corpus, MANIFEST-verified, synth at 8862ee2 = the v0.63 compiler, BEFORE script-on-main vs AFTER, verdicts/rungs identical between runs, 0 TIMEOUT): the fix CHANGED the arm ranking and nothing else. ladder NEVER (the table v0.63 was scoped from, 141 of 243, rungs 27/72/3): encode_operand2 non-rotated immediate 4 @ rank 5 -> 9 @ rank 3 (above start-section 7 and rule_i32_rotl 5); GI-FPU-002 3 -> 1; the top two (register exhaustion 71->70, #929 41->40) keep their ranks. plain census: core #929 16->15 loses its tie for 3rd; components gain an encode_operand2 row (0 -> 3) above GI-FPU-002 (4 -> 2). riscv, aarch64: identical before/after in both modes — their records carry no hex payload at all (surveyed), so the no-op is by construction. Why 9 and not 5: `_modal()` is a per-module plurality vote. In four modules the fragments each lost that vote to a cause with no varying payload (yolo_inference_{release,debug}: 84 functions across ~20 immediates, largest fragment 6, vs 58 for GI-FPU-002; sockets-tcp-connect 10 vs 8; stat-dev-ino 5 vs 4). That is the under-ranking one level down from the row split. The ranked blocker histogram had no cut to rise from; the census's SECOND ranking ("module-level decline reasons (top 10)") ranked RAW text and its top-10 cut silently dropped 56/110, 65/106, 60/78, 35/103, 58/86, 42/107 modules per stratum — routed through the one module_reason pipeline now, with a remainder row. Spec census: RESULT PASS, PINS / FAMILY_OK_PINS / AT_LEAST_ONE_EXPORT unchanged (it emits exact-pinned bucket counts, no histogram). Wiring (coordinator review): the fix would have shipped a self-test CI never runs. The claims.yaml note beside this file's `manual` slot already said "if the census ever grows an expected value or a pass/fail verdict ... it must be wired" — so `# ci-status: wired` with `# ci-checks: stdout` binding the assertion COUNT (14), a step in the required claim-check job via oracle_run plus the per-job evidence ledger; manual ceiling 8 -> 7 in claims.yaml and ORACLE_WIRING.md. Red-first both ways: hex mask reverted -> assertion 2 fails (exit 1); evidence line renamed -> script exits 0 but oracle_run refuses on the floor (measured 0 of 14). Assertion 1 asserts the PRE-#1159 rule still keeps the two wild shapes apart, so the discriminator cannot go vacuous silently. The census MODES are unchanged: still a local measurement, no verdict about synth. Emulation floor re-derived: 324845. Also: reference identifiers (#1102, §4.5.5, VCR-MEM-002) are kept by the mask — format-string constants cannot vary per instance, masking them only loses identity; the published tables were restoring them by hand. docs/status/ACCEPTANCE_LADDER.md gains the re-attributed arm table under the v0.63 one it corrects; the shipped measurement is not rewritten. Refs #1159, #242, #1156 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
avrabe
added a commit
that referenced
this pull request
Sep 7, 2026
…sert the sum invariant, wire the self-test — arm's non-rotated-immediate gap is rank 3 at 9, not rank 5 at 4 (#1181) * fix(census): RQ-64-HISTOGRAM (#1159) — mask HEX numeric payloads so one cause is one row, and assert every ranked view sums to what it ranks WIP checkpoint committed by the coordinator to preserve the lane's work while its arm ladder quarters finish. The lane owns the remaining measurement. `\b\d+\b` left `0x5dc` intact (the `x` glues the digits into one word), so `immediate 0x624 (1500)` normalized to `immediate 0x624 (N)` — the decimal collapsed, the hex did not, and one cause fragmented into one row per distinct immediate. A cause with a varying payload was systematically UNDER-RANKED against one without, in the very histogram v0.63 was scoped from. Hex is now masked first, to a distinct token (`0xN` vs `N`) so the message shape stays readable. Identifiers that merely contain digits (`i32`, `R11`, `func_25`, `RV32`) are not word-bounded numbers and survive; reference identifiers carried as literal format-string text (`#1102`, `§4.5.5`, `VCR-MEM-002`) are kept whole, since they cannot vary per instance and masking them could only lose identity. The mask is idempotent, so a `--json` record's already-normalized `skip_reasons` can be re-ranked offline without re-fragmenting. Every ranked view now asserts its printed rows sum to the modules it ranks — a collapse that LOSES rows is worse than one that fragments them. `--self-test` exercises both wild shapes and proves the sum check can fail (negative control). Refs #1159 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * RQ-64-HISTOGRAM (#1159): wire the census self-test, measure the ranking — arm's non-rotated-immediate class was rank 5 at 4, is rank 3 at 9 Measurement (full 243-module corpus, MANIFEST-verified, synth at 8862ee2 = the v0.63 compiler, BEFORE script-on-main vs AFTER, verdicts/rungs identical between runs, 0 TIMEOUT): the fix CHANGED the arm ranking and nothing else. ladder NEVER (the table v0.63 was scoped from, 141 of 243, rungs 27/72/3): encode_operand2 non-rotated immediate 4 @ rank 5 -> 9 @ rank 3 (above start-section 7 and rule_i32_rotl 5); GI-FPU-002 3 -> 1; the top two (register exhaustion 71->70, #929 41->40) keep their ranks. plain census: core #929 16->15 loses its tie for 3rd; components gain an encode_operand2 row (0 -> 3) above GI-FPU-002 (4 -> 2). riscv, aarch64: identical before/after in both modes — their records carry no hex payload at all (surveyed), so the no-op is by construction. Why 9 and not 5: `_modal()` is a per-module plurality vote. In four modules the fragments each lost that vote to a cause with no varying payload (yolo_inference_{release,debug}: 84 functions across ~20 immediates, largest fragment 6, vs 58 for GI-FPU-002; sockets-tcp-connect 10 vs 8; stat-dev-ino 5 vs 4). That is the under-ranking one level down from the row split. The ranked blocker histogram had no cut to rise from; the census's SECOND ranking ("module-level decline reasons (top 10)") ranked RAW text and its top-10 cut silently dropped 56/110, 65/106, 60/78, 35/103, 58/86, 42/107 modules per stratum — routed through the one module_reason pipeline now, with a remainder row. Spec census: RESULT PASS, PINS / FAMILY_OK_PINS / AT_LEAST_ONE_EXPORT unchanged (it emits exact-pinned bucket counts, no histogram). Wiring (coordinator review): the fix would have shipped a self-test CI never runs. The claims.yaml note beside this file's `manual` slot already said "if the census ever grows an expected value or a pass/fail verdict ... it must be wired" — so `# ci-status: wired` with `# ci-checks: stdout` binding the assertion COUNT (14), a step in the required claim-check job via oracle_run plus the per-job evidence ledger; manual ceiling 8 -> 7 in claims.yaml and ORACLE_WIRING.md. Red-first both ways: hex mask reverted -> assertion 2 fails (exit 1); evidence line renamed -> script exits 0 but oracle_run refuses on the floor (measured 0 of 14). Assertion 1 asserts the PRE-#1159 rule still keeps the two wild shapes apart, so the discriminator cannot go vacuous silently. The census MODES are unchanged: still a local measurement, no verdict about synth. Emulation floor re-derived: 324845. Also: reference identifiers (#1102, §4.5.5, VCR-MEM-002) are kept by the mask — format-string constants cannot vary per instance, masking them only loses identity; the published tables were restoring them by hand. docs/status/ACCEPTANCE_LADDER.md gains the re-attributed arm table under the v0.63 one it corrects; the shipped measurement is not rewritten. Refs #1159, #242, #1156 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * chore(rivet): RQ-64-HISTOGRAM landed: names PR #1181 Refs #1159 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L * RQ-64-HISTOGRAM (#1159): reconcile 135 / 138 / 141 — the pre-fix tool printed all twelve NEVER rows; 138 and 135 are transcription cuts, not tooling losses Replayed the PRE-FIX script's own report_ladder over the exact BEFORE NEVER set (each module's stored pre-fix blocker, no synth re-run): twelve rows summing to 141. The held nine-row listing (138) is its first nine rows; the three missing modules are rows 10-12 (LdrSym Thumb-N-only, encode_operand2 0x5dc, no exports). Both v0.63.0 and pre-fix main print most_common(12), so the shipped tool lost nothing — the new sum assertion correctly does NOT go red on the pre-fix bucketing; a cut forced to top=9 now prints a remainder row and still closes at 141. No third in-tool loss mechanism exists: 0 of 141 NEVER modules have an empty primary blocker. Loss mechanisms named: the CUT hides rows (closes now), FRAGMENTATION misranks (masked now), TRANSCRIPTION drops rows outside the tool (135, 138) — checkable now against the printed 'N rows sum to M' line. Also: the full-corpus single-pass arm ladders finished and agree with the quarter merge on all 243 modules' rung and blocker, both sides (0 differ); the artifact's denominator note is upgraded from a caveat to a cross-check. Refs #1159 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RQ-62-REACH increment 1 — measurement only, report and stop
Refs #242, #1017
The artifact's increment 1 asks for the 805-module acceptance census re-run on v0.61.0. The 805 denominator is not reproducible on this machine — the census script's own
# ci-statusheader says its input is a local corpus CI does not carry, and the wasm.directory component stratum plus the external toolchain-output modules are absent here. gale owns the full corpus and was asked for a re-run on #1017. So this PR does the honest subset: census what exists, and state the denominator as prominently as the rates. No blocker is fixed, no capability implemented.The denominator (stated before any rate)
243 unique modules by sha256 (131 core + 112 components, 447 MB): every
*.wasmreachable on this machine across.rivet/repos/{gale,kiln,loom,meld,scry,sigil}, the sibling../loomand../kilncheckouts, and synth itself. Per source: kiln 124, meld 56, synth 19, scry 16, loom sibling 15, gale 6, loom rivet-mirror 6, kiln rivet-mirror 1. This corpus is org-built fixtures and composites. Rates over it are NOT comparable to the 805 figures and no delta against 66%/14%/1.6% is computed anywhere in this PR. The done-when's "805-module census" leg is explicitly not satisfied; the artifact says so and staysproposed.Per-backend, accepted / partial / declined / errored (synth 0.61.0,
--all-exports --relocatable, zero timeouts)Ranked blocker histogram (one primary blocker per non-accepted module; #952/#1102 policy refusals attributed to the modal per-function skip reason behind them; each histogram sums to its stratum's decline count)
What moved, only where comparison is honest
The one comparable stratum is v0.59's 2026-08-25 run (same script, same flags, same org-repos-on-this-machine method; snapshot differs 307→243, overlap unquantifiable): arm 81% → 11% accepted — and the fall is the v0.59–v0.61 honesty hardening doing its job. The #1041 active-data refusal, #1052 global-init refusal, and #1102 dangling-reloc refusal converted silent-drop "accepts" (125 partial objects, 77 accepts with silently-dropped data in the v0.59 run) into loud declines. An ACCEPT today excludes every known silent-drop class. Reading the fall as a capability regression compares an honest number to a flattered one — the compliance-envelope rule exists for exactly this.
Also worth naming: aarch64 (1.6% on the 805) is now the highest-accepting backend on this corpus post-RQ-60-A64IMPORT, and multi-memory blocks 50/112 components on riscv+aarch64 but zero on arm (VCR-MEM-002 phase 1 is ARM-only) — those same components then decline on arm for deeper causes, so no single capability un-blocks the component stratum.
Changes
scripts/repro/partial_census_1017.py— four-bucket summary + ranked blocker histogram (instance lists collapsed so one cause is one bucket; A declined export exits 0 with the symbol simply absent — the decline is honest to stderr but invisible to the build system #952/rv32: a retained function relocating against a declined INTERNAL emits a danglingsynth_func_Nand exits 0 — the #851 guard #1013 gave aarch64 is not on the RISC-V path #1102 refusals attributed to their root-cause skip reasons). Stillci-status: manual: no expected values, no verdict.artifacts/release-v0.62/RQ-62-REACH.yaml— the census result published in the artifact, denominator first.Gates
cargo fmt --check0;cargo clippy --workspace --all-targets -- -D warnings0 (no Rust touched)claim_check0,oracle_wiring_check0status_evidence_check: R4 asks the artifact to acknowledge the increment —fields.landedis filled with this PR's number in the follow-up commit on this branchrivet validate: 40 errors / 298 warnings identical before and after this diff (pre-existing on main; none introduced)🤖 Generated with Claude Code
https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L