Repository navigation
Guide residuals from guide-pins-v2: route closure, constants provenance, fixture churn, gate-semantics reform #18
Description
Activity
Owner triage (2026-09-07): F7 = promote — one Pass-A clause (course-declared constants that don't participate in contracts land in CTX as provenance; contract-participating constants stay binding) + one §14 audit line + a CTX name-map line in the agent-memory spec. F8 = will be presented for rejection/deferral in the pilot promotion pass (churn is by-design; no rule proposed). F6 and F12 ride the same pilot pass. Cross-course validation deferred: the next new course's /generate-spec is the de-facto test — any principle-miss there triggers a promotion pass.
🤖 Generated with Claude Code
Pass update from PR #21 (the guide-scope-and-residuals shakedown):
Done in #21
- F6 (route closure): implemented as the "no route admits a non-decision" precondition (195a94d), confirmed across the confirmation set (its known coin-flip failure mode was re-tested at higher N per the skill).
- F7 (constants provenance): Pass-A now mines contract-participating constants' provenance into CTX (195a94d); the adopted spec's §3 constants all carry CTX anchors.
Deferred by owner decision
- F8 (fixture churn): presented at triage, deferred — no rule shipped; proposed fix text stays here, unapplied (fix and test travel together).
- F12 (gate-semantics reform) remains the next pass: "recommended baseline build" rename, two-label precedence, deviation walkthrough, ban bare "(default)".
New residuals from the pass
- Ledger row-set ±1 variance: across 9 regenerations over three guide hashes, the row set wobbles at exactly the judgment-boundary subjects (topology row folded into the binding contract ~2/9; one partitioned-context row promoted via the structural route). Two targeted wording fixes each reduced but did not eliminate it; assessed as generation variance, not rule ambiguity — full analysis in docs/research/guide-scope-and-residuals-matrix.md. Adopted specs are matrix-verified individually, so exposure is limited to future regenerations.
- Unpinned model-role split: one regen pinned a single model for all LLM roles while another kept the course's loop/memory-ops split as the model-row default. Candidate for a future deterministic model-role pinning rule.
🤖 Generated with Claude Code
F12 (gate-semantics reform) is implemented via PR #23 — the last gate-semantics residual from guide-pins-v2. Validated by 3 clean-room Opus-class regenerations (template-conformance scoring with quoted evidence) and 2 gate probes, including one where the subject, told directly to 'skip the paperwork and get coding', still printed the deviation-marked checklist and walked the substituted rows before building.
Residual re-characterization (supersedes the earlier ±1 framing): the Ledger row-set variance is a class of judgment-boundary subjects, not one subject — across 12 total regenerations (9 Fable-class over the earlier passes, 3 Opus-class this pass) the flickering subjects are memory-core topology, summary thread-scoping, docstring augmentation, and initialization semantics, while the ~24 core properties stay stable in every sample. Observed on both model classes → generation variance at genuinely debatable boundaries, not model behavior.
Owner decision recorded: summary thread-scoping is ADMITTED to the canonical subject list (both sides are course-demonstrated working code, never warned against; the right answer depends on learner context — the profile of a Ledger row). Implementation queued per fix-and-test-travel-together: the spec's row rides the owner's upcoming revisions batch; the guide example-sentence and the canon count update ride the next promotion pass with its validation round.
New candidate findings queued for the next pass:
- Identifier-format mining, alphabet included — all 3 Opus-class regens pinned summary-id length but dropped the hex alphabet (one even floored minLength:6 under a stated 8 chars), while both Fable-class regens pinned hex: the first clean cross-model pinning gap. Proposed direction: Pass-A names identifier formats (alphabet included) as working parameters. Not applied — recorded only.
- Importable fixture module names — two independent build agents (run-09, gate-probe-2) tripped on the spec's fixture literal
fixtures/llm-stubs.pybeing unimportable in Python. Proposed direction: the guide requires fixture module filenames to be importable identifiers. Not applied — recorded only. - Summary-scoping guide example sentence (per the owner decision above).
🤖 Generated with Claude Code
Via PR #25: the two queued technical findings are implemented and validated — identifier-format mining (alphabet included) is now a Pass-A rule (the hex contract reproduced 3/3 on Opus-class regens, closing the 0/3 cross-model gap), and importable fixture module filenames is now a §5 anatomy rule (all regens produced importable names). The summary-scoping canon admission is fully implemented: D13 row in the spec, the contradicted-route worked example in the guide, canonical subject list now 13 (one regen produced the first exact 13/13 Ledger). F8 remains parked by owner decision.
New candidate queued from the live-AC run (recorded in the guide-revisions-batch matrix, unapplied): harness-known identities are injected, never model-guessed — AC21 caught the build letting the model supply its own thread id, which it cannot know.
🤖 Generated with Claude Code
Closing: all four named residuals are resolved or explicitly parked.
Item Status Where F6 route closure Done PR #21 — the "no route admits a non-decision" precondition, confirmed at raised N (its failure mode was a known coin-flip) F7 constants provenance Done PR #21 — Pass-A mines contract-participating constants as binding, the rest as CTX provenance F12 gate-semantics reform Done PR #23 — recommended-baseline rename, two-label timing frames, deviation-marked checklist, bare-"(default)" ban; validated 3/3 on clean-room regens plus two gate probes F8 fixture churn Parked by owner decision Carried forward to the next-pass slate so it isn't lost. Note: with the Option-B pattern (surgical spec edits rather than regen adoption) the churn isn't currently occurring Findings added to this issue after it was filed, also delivered in PR #25:
- Identifier-format mining, alphabet included — the hex contract reproduced 3/3 on Opus-class regens, closing a 0/3 cross-model gap.
- Importable fixture module filenames — all regens now produce importable names; two independent builds had tripped on
llm-stubs.py. - Summary-scoping canon admission — D13 row in the spec, the contradicted-route worked example in the guide, canonical subject list at 13; one regen produced the first exact 13/13 Ledger.
Residual left deliberately open (documented, not fixed): the Ledger row-set variance is a class of judgment-boundary subjects (topology, summary-scoping, augmentation, initialization semantics), stable ~24 core properties across 12 regenerations on two model classes. Accepted as generation variance; adopted specs are matrix-verified individually.
The one candidate this issue's work generated but did not implement — the guide-level rule that harness-known identities are injected, never model-guessed — moves to the next-pass slate with its evidence.
Residuals from the guide-pins-v2 pass (implemented #8 + #11; matrix:
docs/research/guide-pins-v2-property-matrix.md). Each rides the next guide pass; per the fix-and-test-travel-together rule, none of the proposed texts below is applied yet.F6 — Route closure: a baked-in subject re-enters the Ledger through alternate routes
Across six clean-room regenerations the tool-exposure subject re-entered the Ledger three ways: a sanitized
design-arguedrow (regen-A1), a plaindesign-arguedrow (regen-B2), and — after the degenerate-case clause sealed that route — adesign-structuralrow (regen-A3, file-verified: "the toolset passed to the model per turn MUST be bounded" reads as a guarantee and passes the structural litmus). Each patch sealed one door; the subject found another. Root cause: routes test properties of a subject (argued? structural?) but never the precondition that a decision exists.Proposed rule (§5.5, above the routes):
F7 — Course-declared constants should land in CTX provenance
The adopted spec (regen-B3) carries no mention of the course's store-name constants (
CONVERSATIONAL_MEMORY,SEMANTIC_MEMORY, …). Adjudicated: not a fidelity regression (no contract, AC, or eval-rubric dependency — the eval rubric has no naming criterion; the decision surface is intact), but a mining-completeness miss: the guide's audit says mined course facts land somewhere. The principled line: constants that participate in contracts (e.g. the PERSON/PLACE/SYSTEM enum) stay binding; non-contract constants land as a CTX provenance line ("the course's notebooks name these stores …") so a learner can map spec concepts back to the lessons.Proposed: a Pass-A clause + §14 audit: course-declared identifiers/constants that don't participate in contracts are recorded in CTX as provenance, never as binding values.
F8 — Fixture-corpus churn across regenerations
Each regeneration authors a fresh synthetic fixture corpus (by design; no course data may be copied). Consequence: fixture identities churn per regeneration, adding noise to cross-generation comparisons. No rule proposed yet — logged for a future think (e.g., a fixture-stability preference when regenerating an existing course's spec).
F12 — Gate-semantics reform (headline item; absorbs F10/F11)
The gate's express lane currently says "build the course-default takeaway as-is" — but on substitution rows the default is not the course's choice (SQLite vs Oracle; omit vs Tavily), so the lane's own name softly commits the misattribution the labeling rule exists to prevent. The adopted spec's §0 carries this wording (known, accepted for this window). Two further leaks: bare "(default)" markers inside Options cells (unsanctioned third label; present in the previous spec too), and no rule for how a build agent earns a "(Recommended)" flag.
Agreed design (owner-specified):
Requires its own regeneration/convergence validation (it edits the §6.0 gate template that specs emit verbatim).
Process note (F9, feeds the planned /evolve-guide skill, not a guide change): scoring claims must carry verified-vs-reported labels; N=1 regen to iterate, N=2 only to confirm a pass; fix and test travel together.
🤖 Generated with Claude Code