Skip to content

Guide residuals from guide-pins-v2: route closure, constants provenance, fixture churn, gate-semantics reform #18

Description

@Carr1005

Residuals from the guide-pins-v2 pass (implemented #8 + #11; matrix: docs/research/guide-pins-v2-property-matrix.md). Each rides the next guide pass; per the fix-and-test-travel-together rule, none of the proposed texts below is applied yet.

F6 — Route closure: a baked-in subject re-enters the Ledger through alternate routes

Across six clean-room regenerations the tool-exposure subject re-entered the Ledger three ways: a sanitized design-argued row (regen-A1), a plain design-argued row (regen-B2), and — after the degenerate-case clause sealed that route — a design-structural row (regen-A3, file-verified: "the toolset passed to the model per turn MUST be bounded" reads as a guarantee and passes the structural litmus). Each patch sealed one door; the subject found another. Root cause: routes test properties of a subject (argued? structural?) but never the precondition that a decision exists.

Proposed rule (§5.5, above the routes):

No route admits a non-decision. Every route admits decisions, and a decision requires at least two legitimate options. When every alternative to the taught side is a course-warned anti-pattern or its degenerate case, there is no decision to surface — the subject enters the Ledger through no route, design-structural included, even where its guarantee would pass the structural litmus: the guarantee lives in a business rule and its AC; the warned side lives in the Trade-off/CTX narrative.

F7 — Course-declared constants should land in CTX provenance

The adopted spec (regen-B3) carries no mention of the course's store-name constants (CONVERSATIONAL_MEMORY, SEMANTIC_MEMORY, …). Adjudicated: not a fidelity regression (no contract, AC, or eval-rubric dependency — the eval rubric has no naming criterion; the decision surface is intact), but a mining-completeness miss: the guide's audit says mined course facts land somewhere. The principled line: constants that participate in contracts (e.g. the PERSON/PLACE/SYSTEM enum) stay binding; non-contract constants land as a CTX provenance line ("the course's notebooks name these stores …") so a learner can map spec concepts back to the lessons.

Proposed: a Pass-A clause + §14 audit: course-declared identifiers/constants that don't participate in contracts are recorded in CTX as provenance, never as binding values.

F8 — Fixture-corpus churn across regenerations

Each regeneration authors a fresh synthetic fixture corpus (by design; no course data may be copied). Consequence: fixture identities churn per regeneration, adding noise to cross-generation comparisons. No rule proposed yet — logged for a future think (e.g., a fixture-stability preference when regenerating an existing course's spec).

F12 — Gate-semantics reform (headline item; absorbs F10/F11)

The gate's express lane currently says "build the course-default takeaway as-is" — but on substitution rows the default is not the course's choice (SQLite vs Oracle; omit vs Tavily), so the lane's own name softly commits the misattribution the labeling rule exists to prevent. The adopted spec's §0 carries this wording (known, accepted for this window). Two further leaks: bare "(default)" markers inside Options cells (unsanctioned third label; present in the previous spec too), and no rule for how a build agent earns a "(Recommended)" flag.

Agreed design (owner-specified):

  1. Express lane renamed "recommended baseline build", described in the gate question itself: "Most decisions on this path are the course's own defaults; the few that deviate do so only because the course's choice needs an API key, a large download, or admin setup — each deviation is flagged and explained before the build starts."
  2. Exactly two gate labels. "(course default)" = provenance, unchanged. "(recommended — reason)" = advice, from two sources with fixed precedence: the build agent's learner-contextual recommendation (including environment detection, e.g. a sandbox-provisioned key) overrides the generator's baked zero-setup conditional ("recommended for zero-setup: the course's choice needs an API key"). A recommendation must state its reason; with no learner signal, surface the generator's conditional as-is. Never relabel the baseline value "(Recommended)" merely because it is the baseline. Bare "(default)" is banned from Options cells and the gate.
  3. Baseline-build walkthrough: after the learner picks the baseline build, the required resolved-decision checklist flags each row deviating from the course's own choice, with its reason (informational — no re-asking; the silent-default and partial-present traps stay closed), and the learner may bail into customize.
  4. The Ledger's Default field survives untouched as machinery (one buildable target per row; determinism doctrine) — only its gate presentation changes; the word "default" never surfaces as a label.

Requires its own regeneration/convergence validation (it edits the §6.0 gate template that specs emit verbatim).


Process note (F9, feeds the planned /evolve-guide skill, not a guide change): scoring claims must carry verified-vs-reported labels; N=1 regen to iterate, N=2 only to confirm a pass; fix and test travel together.

🤖 Generated with Claude Code

Activity

  1. Carr1005 commented on Sep 7, 2026

    @Carr1005
    CollaboratorAuthor

    Owner triage (2026-09-07): F7 = promote — one Pass-A clause (course-declared constants that don't participate in contracts land in CTX as provenance; contract-participating constants stay binding) + one §14 audit line + a CTX name-map line in the agent-memory spec. F8 = will be presented for rejection/deferral in the pilot promotion pass (churn is by-design; no rule proposed). F6 and F12 ride the same pilot pass. Cross-course validation deferred: the next new course's /generate-spec is the de-facto test — any principle-miss there triggers a promotion pass.

    🤖 Generated with Claude Code

  2. Carr1005 commented on Sep 8, 2026

    @Carr1005
    CollaboratorAuthor

    Pass update from PR #21 (the guide-scope-and-residuals shakedown):

    Done in #21

    • F6 (route closure): implemented as the "no route admits a non-decision" precondition (195a94d), confirmed across the confirmation set (its known coin-flip failure mode was re-tested at higher N per the skill).
    • F7 (constants provenance): Pass-A now mines contract-participating constants' provenance into CTX (195a94d); the adopted spec's §3 constants all carry CTX anchors.

    Deferred by owner decision

    • F8 (fixture churn): presented at triage, deferred — no rule shipped; proposed fix text stays here, unapplied (fix and test travel together).
    • F12 (gate-semantics reform) remains the next pass: "recommended baseline build" rename, two-label precedence, deviation walkthrough, ban bare "(default)".

    New residuals from the pass

    1. Ledger row-set ±1 variance: across 9 regenerations over three guide hashes, the row set wobbles at exactly the judgment-boundary subjects (topology row folded into the binding contract ~2/9; one partitioned-context row promoted via the structural route). Two targeted wording fixes each reduced but did not eliminate it; assessed as generation variance, not rule ambiguity — full analysis in docs/research/guide-scope-and-residuals-matrix.md. Adopted specs are matrix-verified individually, so exposure is limited to future regenerations.
    2. Unpinned model-role split: one regen pinned a single model for all LLM roles while another kept the course's loop/memory-ops split as the model-row default. Candidate for a future deterministic model-role pinning rule.

    🤖 Generated with Claude Code

  3. Carr1005 commented on Sep 8, 2026

    @Carr1005
    CollaboratorAuthor

    F12 (gate-semantics reform) is implemented via PR #23 — the last gate-semantics residual from guide-pins-v2. Validated by 3 clean-room Opus-class regenerations (template-conformance scoring with quoted evidence) and 2 gate probes, including one where the subject, told directly to 'skip the paperwork and get coding', still printed the deviation-marked checklist and walked the substituted rows before building.

    Residual re-characterization (supersedes the earlier ±1 framing): the Ledger row-set variance is a class of judgment-boundary subjects, not one subject — across 12 total regenerations (9 Fable-class over the earlier passes, 3 Opus-class this pass) the flickering subjects are memory-core topology, summary thread-scoping, docstring augmentation, and initialization semantics, while the ~24 core properties stay stable in every sample. Observed on both model classes → generation variance at genuinely debatable boundaries, not model behavior.

    Owner decision recorded: summary thread-scoping is ADMITTED to the canonical subject list (both sides are course-demonstrated working code, never warned against; the right answer depends on learner context — the profile of a Ledger row). Implementation queued per fix-and-test-travel-together: the spec's row rides the owner's upcoming revisions batch; the guide example-sentence and the canon count update ride the next promotion pass with its validation round.

    New candidate findings queued for the next pass:

    1. Identifier-format mining, alphabet included — all 3 Opus-class regens pinned summary-id length but dropped the hex alphabet (one even floored minLength:6 under a stated 8 chars), while both Fable-class regens pinned hex: the first clean cross-model pinning gap. Proposed direction: Pass-A names identifier formats (alphabet included) as working parameters. Not applied — recorded only.
    2. Importable fixture module names — two independent build agents (run-09, gate-probe-2) tripped on the spec's fixture literal fixtures/llm-stubs.py being unimportable in Python. Proposed direction: the guide requires fixture module filenames to be importable identifiers. Not applied — recorded only.
    3. Summary-scoping guide example sentence (per the owner decision above).

    🤖 Generated with Claude Code

  4. Carr1005 commented on Sep 9, 2026

    @Carr1005
    CollaboratorAuthor

    Via PR #25: the two queued technical findings are implemented and validated — identifier-format mining (alphabet included) is now a Pass-A rule (the hex contract reproduced 3/3 on Opus-class regens, closing the 0/3 cross-model gap), and importable fixture module filenames is now a §5 anatomy rule (all regens produced importable names). The summary-scoping canon admission is fully implemented: D13 row in the spec, the contradicted-route worked example in the guide, canonical subject list now 13 (one regen produced the first exact 13/13 Ledger). F8 remains parked by owner decision.

    New candidate queued from the live-AC run (recorded in the guide-revisions-batch matrix, unapplied): harness-known identities are injected, never model-guessed — AC21 caught the build letting the model supply its own thread id, which it cannot know.

    🤖 Generated with Claude Code

  5. Carr1005 commented on Sep 10, 2026

    @Carr1005
    CollaboratorAuthor

    Closing: all four named residuals are resolved or explicitly parked.

    Item Status Where
    F6 route closure Done PR #21 — the "no route admits a non-decision" precondition, confirmed at raised N (its failure mode was a known coin-flip)
    F7 constants provenance Done PR #21 — Pass-A mines contract-participating constants as binding, the rest as CTX provenance
    F12 gate-semantics reform Done PR #23 — recommended-baseline rename, two-label timing frames, deviation-marked checklist, bare-"(default)" ban; validated 3/3 on clean-room regens plus two gate probes
    F8 fixture churn Parked by owner decision Carried forward to the next-pass slate so it isn't lost. Note: with the Option-B pattern (surgical spec edits rather than regen adoption) the churn isn't currently occurring

    Findings added to this issue after it was filed, also delivered in PR #25:

    • Identifier-format mining, alphabet included — the hex contract reproduced 3/3 on Opus-class regens, closing a 0/3 cross-model gap.
    • Importable fixture module filenames — all regens now produce importable names; two independent builds had tripped on llm-stubs.py.
    • Summary-scoping canon admission — D13 row in the spec, the contradicted-route worked example in the guide, canonical subject list at 13; one regen produced the first exact 13/13 Ledger.

    Residual left deliberately open (documented, not fixed): the Ledger row-set variance is a class of judgment-boundary subjects (topology, summary-scoping, augmentation, initialization semantics), stable ~24 core properties across 12 regenerations on two model classes. Accepted as generation variance; adopted specs are matrix-verified individually.

    The one candidate this issue's work generated but did not implement — the guide-level rule that harness-known identities are injected, never model-guessed — moves to the next-pass slate with its evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions