Skip to content

feat(budget): name a budget set after its consumer, and make check enforce it - #263

Merged
wenzowski merged 3 commits into
mainfrom
claude/phase-3-sequential-landing-h26kx0
Aug 11, 2026
Merged

wenzowski merged 3 commits into
mainfrom
claude/phase-3-sequential-landing-h26kx0

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

Closes the two clauses the 2026-08-11 DoD-conformance audit demoted CLOUD-50 from
Done for. PR #237 landed budget.rs and policy budget; these are the two places
the landed code disagreed with its own Definition of Done.

1. The consumer's set name was baked into the engine

Budget was a struct with an instructions field. "instructions" is this
repository's
name for its always-loaded context — a consumer-specific identifier
living in crates/batten, which non-negotiable rule 1 forbids. A second consumer
budgeting a different surface needed an engine change to do it.

[budget.<name>] is now a map, and the key is files rather than paths as
the DoD specifies. Any number of sets, under any names the engine has never heard
of, with no code change.

2. The gate never fired

The DoD: "batten check evaluates every declared budget as a non-spawning read
gate: over either bound → a deny finding, exit 2." It didn't. policy budget was
the only surface, so an over-budget set was visible only to whoever thought to run
the report — and a budget that reports when asked is not a gate.

Budgets are now evaluated in run_rules, the one funnel check and enforce
share, and an over-budget set produces an ordinary Finding rather than a
private verdict path. That choice is the point: budgets inherit waivers, the -J
shape, the exit contract and the findings store for free, all of which a bespoke
channel would have had to re-implement. FindingKind::Scope is the honest kind —
a budget is a whole-repo condition — and the identity is over the set, so a
bigger overrun is the same finding rather than a new one each time the count moves.

Counting reads files and sums them, so check's declared read effect is
preserved.

The absent [budget] table reads two ways, on purpose

check measures nothing (a repository that declares no budget has no budget to
fail). policy budget is exit 1 (a report that measured nothing must not print
0 — the false green the engine exists to catch). Two callers asking different
questions, and both readings are honest. Asserted in both directions.

Blast radius worth reviewing

Wiring budgets into check means CLOUD-298's per-entry dead-glob refusal now
reaches the main gate: a budget entry matching no file is exit 1, where before it
was only reachable through policy budget. That is the stronger and intended
reading — a gate that could not run must say so — but it did surface in three
fixture suites (tests/cli.rs, tests/config-lint.bats,
tests/prebuilt-lint.bats) which copy the committed batten.toml into a repo
with no AGENTS.md. Each now supplies the file its budget names.

One detail found on the way: the fixtures symlink most of the tree, but
rules::tree_files counts regular files only, so a symlinked AGENTS.md is
invisible to the walk and the entry reads as dead. The bats fixtures copy it.

Tests

Unit (budget.rs): a set name is the consumer's and any name validates; an
over-budget report is a finding naming its set, with identity stable across
overrun size and distinct per set; a set within budget produces none;
measure_all(None) is empty rather than an error.

E2E (tests/cli.rs): an over-budget set denies through check and rides the
normal -J findings channel; a set within budget leaves check silent; two named
sets are each measured, only the over one is named, and policy budget -J reports
an array in name order so the shape does not change as a consumer adds a budget.

Consumer #1's batten.toml adopts files, with both thresholds unchanged and
still pinned by the existing test.

Refs: CLOUD-50

@linear-code

linear-code Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
CLOUD-50 Token-budget enforcement for the instruction-file set (`[budget]` + `policy budget`)

Why

The always-loaded context budget is live policy today as mise-tasks/context-budget, wired into the hk gate: it sums AGENTS.md plus .serena/memories/always/*.md after stripping what the loader drops (YAML frontmatter, block-level HTML comments), estimates tokens at chars/4, and fails over a 200-line target or a 3500-token budget. Both thresholds live as env-var defaults inside a shell script — outside the committed policy authority, invisible to config show, and outside the raise-only override contract (house style §8). The predicate itself sits squarely in the engine's decidable fragment (§0.3: a monotone sum against a constant guard over declared files), so the policy-engine-first answer is to declare it in batten.toml and evaluate it in the engine. House style §2 already names the surface: policy budget (read).

The threshold is a convention-level bound, not a literature-backed performance claim: pinned deliberately in a test, moved by a config edit, never tuned on a paper.

Rejected alternatives

  • Count lines instead of tokens. Lines proxy layout, not the cost every agent pays every turn. A line target stays as an optional secondary bound (max_lines), but tokens are the primary predicate.
  • An exact tokenizer. Needs a model-specific vocabulary and, for some, a network fetch; a budget gate that can fail because a download failed is worse than one ~10% out. chars/4 is stable, offline, and monotonic — all a budget needs. (The shell gate's own recorded rationale; it carries over.)
  • A [[rule]] row or new rule kind. A budget is a property of a declared file set with its own threshold keys, not a per-file banned shape; it sits beside scope and protected as the third policy table, which is exactly how §2 groups the verb (policy scope | protect | budget).

Definition of done

  • [budget.<name>] tables in batten.toml: files (repo-relative globs), max_tokens (required), max_lines (optional). The set name is the consumer's ([budget.instructions] here); the engine carries no file names (non-negotiable rule 1).
  • The counting convention is defined once in the engine: per matched file, drop YAML frontmatter and block-level HTML comments (bytes the loader never injects), then lines, chars, and tokens = chars/4; the verdict is over the summed set.
  • batten check evaluates every declared budget as a non-spawning read gate: over either bound → a deny finding, exit 2. Enforcement in this repo needs no new wiring — the hk gate already runs mise run batten-check.
  • batten policy budget (new §2 rows policy and policy budget, both read) prints the measurement: one row per counted file (path, lines, chars, ~tokens) plus the set total against its declared bounds; exit 2 over budget, 0 otherwise.
  • Overrides only tighten: batten.local.toml and env may lower max_tokens/max_lines, never raise them (§8 raise-only, in the budget direction — smaller is stricter).
  • Consumer ci: check in the main-branch protection ruleset #1 adopts: batten.toml declares [budget.instructions] over the always-loaded set with the current 3500/200 bounds; the shell gate's file-set counting retires in favour of the engine gate, and any remnant surface (see assumptions) carries no second threshold.

Acceptance

  • A fixture set over max_tokens fails check with a finding naming the budget set and the measured count — never file content.
  • The same tree under budget: this gate prints nothing and contributes exit 0.
  • A file that is only frontmatter and HTML comments contributes zero.
  • An empty or absent files list is an empty set (counts 0) — absence is never "everything".
  • policy budget output is byte-identical across two runs; -J likewise.
  • The threshold stays pinned in a fixture test, per the recorded convention-not-literature stance.

Refinement — Ready (a declared file-set token budget: config table, check gate, policy budget introspection)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). batten.toml's [budget.<name>] tables own every threshold; the engine owns the counting convention, defined once in crates/batten. mise-tasks/context-budget's env-var defaults stop being an authority: after this lands, the committed config is the only place a budget is declared.
  • Computable predicate (§2). In the decidable fragment (§0.3): a monotone sum of per-file token estimates against a constant guard. Not expressible as a [[rule]] row today — no kind counts, and the thresholds are set-level, not per-file — so this issue grows the engine: a budget gate on check's read surface (inspection-only, spawns nothing) plus the policy budget verb. Gated by mise run test:cargo inside the hk gate and CI.
  • Effect (§3). Two new SURFACE rows, both read: policy and policy budget — inspection only, structurally incapable of mutation, so both join the derived read-only allowlist. check's declared read is preserved: counting reads files and spawns nothing.
  • Generated artifacts (§4). The batten.toml JSON Schema gains [budget] through its generator (batten generate schema, asserted by the existing schema-check gate step); completions regenerate for the new verb (batten generate completions, asserted by completions-check). No hand-edited copies.
  • Output & exit (§5). Pointer-only: findings and rows carry paths, counts, and bounds — never file content. A clean check prints nothing and exits 0; over-budget is a policy verdict, exit 2; -J is byte-identical for identical input.
  • Commit / bump (§6). feat → patch until 0.1.0.
  • Test obligation (§7). E2E over the compiled binary (crates/batten/tests/cli.rs): (a) over-budget fixture → check exit 2, finding names the set and the measured count; (b) under-budget → silent 0; (c) strip convention — a frontmatter/comment-only file contributes zero; (d) empty files → empty set; (e) policy budget per-file rows plus total, byte-identical across runs, -J stable; (f) a local override raising max_tokens is refused, lowering is honoured. The counting function is unit-tested in-module.
  • Blockers (§8). blockedBy CLOUD-298 — it settles the counted set this budget inherits, and retiring the shell gate before it lands would silently drop the surface it adds. relatedTo CLOUD-82 — the same token-estimate convention on a different plane (drain output budgeting, not file sets).

Stated assumptions

  1. Whole files are the unit. A budget set counts files matched by globs. The always-given prompt key inside the project configuration (CLOUD-298's second surface) is a fragment of a file, not a file; it stays with the shell gate until it is retired or the mechanism learns extraction. If a remnant shell gate survives for that one surface, it derives its threshold from the committed config (batten config show), never from a second constant — asserted in tests/context-budget.bats.
  2. chars/4, integer division, matching the shell gate — the migration changes the authority, not the measurement. A fixture asserts the engine and the shell gate agree on the current tree before the shell step is removed.

Review in Linear

@wenzowski
wenzowski marked this pull request as ready for review August 11, 2026 09:06
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit dcf58df into main Aug 11, 2026
11 checks passed
@wenzowski
wenzowski deleted the claude/phase-3-sequential-landing-h26kx0 branch August 11, 2026 09:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants