Repository navigation
feat(mcp): dispatch declared MCP calls and return reductions (CLOUD-1260, CLOUD-1122, CLOUD-1251) - #799
Conversation
CLOUD-1260 Tracker round trips are 73% of a session's tool output and no gate sees them: Batten should dispatch the MCP call itself and return a reduction
Why Measured over one session's own transcript (79 MB, 22,377 lines) on 2026-08-31: 43.5 MB of tool content, of which tracker round trips are 13.2 MB — 73% of all tool output, across 973 calls against 208 rows. Bash + Grep + Read together are 1.9 MB. Reading the tree is nearly free; the connector is the entire cost.
Two measurements from the same transcript rule out the cheap fixes, and both are recorded here so the next session prices them instead of re-deriving them:
This is the successor to CLOUD-782's unremovable half. That row measured the same gap from the gate side, landed the transcript harvest that removed the second copy, and recorded the rest as structural: "the ~3x is structural and stays until the tracker offers a projection." The tracker never has to offer one. MCP is JSON-RPC 2.0, so Batten can be the client, own the response, and give the tracker an API it does not ship — a projection, a diff against the prior version, an acknowledgement instead of an echo. CLOUD-782 is Done and stays Done; this is the direction it could not see. Not CLOUD-204. That decision record declined shipping an MCP server in v0.1. This is Batten as an MCP client, which is the opposite direction and touches nothing that record decided. Refinement — Ready (dispatch the call, reduce the response, and close the raw path) Refinement gate: Definition of Ready & Done. This body carries only specializations.
Measured 2026-09-01 while landing this: the rule-1 grep is ALREADY RED on Two deviations from this Ready block, taken deliberately and recorded here rather than absorbed. (1) §5 specifies exit Acceptance
Not in this issue Shipping an MCP server — CLOUD-204 declined that and this does not reopen it. The doctrine amendment AGENTS.md needs is real and named below, not deferred silently. Scope note, stated rather than absorbed. AGENTS.md says Batten is "not a hook runner, file-shape linter, secret scanner, AST linter, or reference monitor." A component that dispatches calls and reduces responses is a real expansion of the core. Per non-negotiable rule 2 the doctrine amendment lands in the same change, or the next session refuses this work on sight. Refs: CLOUD-782, CLOUD-1122, CLOUD-685, CLOUD-1251, CLOUD-204, CLOUD-918, CLOUD-919 CLOUD-1122 Every `save_issue` echoes the whole stored body back into context, and the capture store already holds it — the write path pays a read it never asked for
Why A Measured 2026-08-28. A grooming session made roughly 45 Re-measured 2026-08-31, and it is an order of magnitude larger than the first reading. Over one session's full transcript (79 MB, 22,377 lines): **435 ** CLOUD-415's body recorded this in passing — " Why this is not simply "ask for less"The echo is not useless, and a naive "make it quieter" would break two live consumers:
So the body must keep arriving at the boundary. What is redundant is the copy that arrives in context: CLOUD-919 already persists every That makes this the write-side twin of CLOUD-1121, which measured the same waste on the read side. Same store, same redundancy, opposite direction — and stated separately because the remedies differ: CLOUD-1121's is a read route, this one's is a request and response shape. What this repository can and cannot decideStated plainly rather than discovered later, because it bounds the row:
Refinement — Ready (measure the echo, bound it, and stop the agent re-reading what the recorder already has) Refinement gate: Definition of Ready & Done. This body carries only specializations.
Acceptance
CLOUD-1251 `[[rule.external]]` declares ONE path under ONE root, so a set of out-of-root files discovered at runtime is unspellable — and the id cannot be written in advance because the name is minted per session
Why CLOUD-1167 landed That bound is correct AND it makes one real consumer unreachable, which nothing currently records.
Each of the three would be fatal alone. This is not "a path outside the root", which CLOUD-1167 solved; it is a set discovered at runtime, which is a different question and was never asked. A SECOND named consumer, found 2026-08-31 — and it is not a variation on the firstCLOUD-1260 needs the same file for the opposite reason. Why a second consumer changes this row rather than merely lengthening it. With one consumer, answer 3 ("unit 4 does not migrate") is cheap and arguably correct — a launcher-specific discovery loop staying in a launcher-specific program. With two, answer 3 stops being a bounded verdict: it would decline a fact that a second, unrelated subsystem also requires, and the same globbing loop then gets written twice outside the engine. That is the argument FOR answer 2 getting stronger and answer 3 getting weaker, and it is evidence rather than preference. The other installs are already spellable and that bounds the row usefully. Claude Code local is How this was found, and why the reason matters more than the instanceCLOUD-1163 recorded its unit 4 ( The decision, which comes before any designDo not assume this should be built. Widening
Answer 2 is the likeliest and answer 1 is the one to argue against, because it is the one that looks like a small change. Refinement — Ready (decide the shape; build nothing until it is decided) Refinement gate: Definition of Ready & Done. This body carries only specializations.
DECIDED 2026-09-01 — answer 2, and a correction to this row's own premiseAnswer 2 is taken: a producer resolves the runtime-discovered set outside the engine, and a module reads what was written. The engine does not grow a glob. Recorded in the tree beside the rows it bounds — Cost of rejecting answer 1 (a declared glob under a declared root): it is the one that looks like a small change and is not. An id's arity goes from one node to many, every module reading it must handle the empty match as could-not-look rather than as absent, and what is spent is the schema's own line — "the difference between a fact and a filesystem scanner". That line is the whole safety property Cost of rejecting answer 3 (decline outright): CLOUD-1260 arrived as a second, unrelated consumer of the same file, so declining would leave one globbing loop written twice outside the engine — the argument this row already makes, and it holds. The correction, and it changes what answer 2 can deliverThis row expects the second consumer to STRENGTHEN answer 2. It cannot be served by answer 2 at all, and the reason is structural rather than a matter of effort. A count answers a module deciding a predicate. It cannot carry an endpoint and headers, which is what DISPATCH needs. So:
What deliberately did NOT landNo Because nothing mechanical landed, acceptance clauses 2 and 3 are answered by their own precondition ("if anything lands") rather than skipped: the engine gained no glob, so a path no row's declaration reaches is still unreadable by any module for exactly the reason it was before — asserted by Acceptance
Found while pressure-testing whether the retirement campaign was 100% unblocked: CLOUD-1163 named unit 4's blocker as out-of-root, that blocker is Done, and the unit is still blocked for a reason no row held. CLOUD-745 Stop shelling out to `curl` permanently: the measurement that justified it tested three `reqwest` presets and generalised past them
Decision, taken 2026-08-20: Why the standing verdict does not hold CLOUD-320 records the row as measured, and the measurement is a three-row table: That is not a proof that no Rust HTTP client can link under Corroborating evidence the bar is narrower than assumed. The workspace Two facts that make this cheaper than it looks
The measurement to run, past the three presets Both chokepoints must clear at once, and neither has been tried:
Acceptance: If every links-free combination still fails, that is a measured result and it changes the shape of the answer, not the goal: the row becomes blocked on CLOUD-737 with the measurement recorded, and the escalation is a decision about the macOS build. Do not restore a "stays, with a measured reason" verdict. Scope of the tokio adoption — narrow, deliberately
What this buys, stated so nobody over-claims it later. There is exactly one network call site in the crate. Every other IO path is local files, the deliberately-serial walk, or subprocess pipes staying threaded. This delivers idiom and kills That last clause is scoped to today, and the fact model is what would falsify it (noted 2026-08-20, so it is not read later as a durable property of the crate). Forge state — PR and check-run facts — is a candidate second network call site, and the first with genuinely concurrent work behind it: 18 tasks call It does not change this issue's scope or its conclusion — the narrow adoption is right either way, and forge facts are not filed as work. It changes only what a future reader may infer from the sentence above. Hardening item 5 becomes more load-bearing rather than less: a forge fact is bounded and cheap, yet still barred from the mediated path because it needs a runtime, which is a cost-class rung CLOUD-757 does not currently have. Hardening — six items, because a runtime in a process that owns process groups is not neutral
Acceptance
Found while auditing the crate's subprocess and string-boundary sites against existing ticket coverage. Probe plan (the measurement §2 below turns into a verdict)
Refinement — Ready (the link gates decide, and either outcome is a result) Refinement gate: Definition of Ready & Done. This body carries only specializations.
CLOUD-919 Persist every PostToolUse response as a local capture
Why
The bytes are then dropped. Measured against so a structured tool — every MCP call, every Where the capture happens, and why thereIn
Not through
|
| Response member | Capture | Provenance row |
|---|---|---|
| absent | none | absence recorded |
| present, empty | a real record of zero bytes | digest of the empty record |
| present, non-empty | the bytes at declared fidelity | digest + fidelity |
facts::rows_in and the Sourced receipt are unchanged and remain a separate consumer reading the same envelope field. They do not read the capture, and the capture does not read them.
Capture failure
CLOUD-917's decision, applied here: hook execution continues, the exit code is unchanged and never 2, a degraded provenance row is written with fidelity = Unavailable and a stable reason id, and the observable signal is that reason id on the doctor and advisory channels — never bytes, never a path. This is the opposite of capture::store's "never a silent skip", and deliberately so: on this surface no Batten failure may block a tool call.
Refinement — Ready
Refinement gate: Definition of Ready & Done. This body carries only specializations.
- Source of truth (§1).
crates/batten/src/lib.rs's post-tool arm inrun_hook— one call site, besiderecord_agent_factat:1881and ahead of it. The store, the handle shape and the fidelity vocabulary are CLOUD-918's and CLOUD-917's; nothing about them is re-decided here. - Computable predicate (§2).
mise run verifygreen with a payload fixture per row of the table above, plus the security predicate end to end: a planted secret is absent from stdout, stderr, every-Jdocument and every file under$GIT_DIR/batten-receipts/, and byte-identical throughbatten capture show <handle> --raw. Failure case: a build that keeps the!command.is_empty()conjunct passes the Bash fixture and reds the structured one — which is why both are required rather than one standing in for the other. - Effect (§3).
hookis already classified. The write is the only new cost and it is a self-declared write on the storecaptureowns; the engine still spawns nothing and builds no runtime, and house-style §5's read promise is untouched because reading a buffer the harness handed over is not execution. - Generated artifacts (§4). None. The wiring is already derived and drift-gated, and this adds no registration.
- Output & exit (§5). Nothing is emitted on the capture path. A post-tool event is not a deny channel on any surveyed host and stays exit
0; a capture failure does not change that. Pointer-only holds structurally: the only new bytes leave the process throughcapture show --raw, which a caller has to name. - Commit / bump (§6).
feat(hook)→ patch until0.1.0. - Test obligation (§7). Over the compiled binary, in
tests/cli.rs's existing idiom beside the CLOUD-776 block at:9114, because the halves live in different processes:-
a Bash fixture and an MCP content-block fixture each persist the authoritative bytes at the declared fidelity;
-
all three aliases —
tool_response,toolResponse,tool_result— reach the capture; -
a present-but-empty response yields a record of zero bytes, and an absent member yields no record and an absence row; a test fails if the two collapse;
-
the planted secret case, both halves — absent from four channels, byte-identical through
--raw; -
spilled, truncated and unavailable responses each produce the specified degraded record rather than a silent full-capture claim, one case per fidelity value;
-
a store that cannot be written leaves the hook at exit
0with a degraded row and a reason id, driven through the state-root seam the suites already use rather than by permission bits (.claude/rules/rust.md: this sandbox runs as root); -
A NEW POST-TOOL BENCH ARM, and this is the row's load-bearing measurement obligation.
mise-tasks/perf.shhas five arms —noop,check,hook,passthrough,wired— andhookandwiredboth feedcrates/batten/tests/fixtures/hooks/claude-code.json, a canned PreToolUse payload, withwired_cmddrawn from.hooks.PreToolUse[]. There is no PostToolUse arm. So the write this row adds — on every post-tool event, therefore on every tool call — is invisible toperf,perf-compare,perf-pairandperf-assert, and runningperf-gateproves nothing about it. This row adds the arm, with its own canned PostToolUse fixture besideclaude-code.jsonandclaude-code-passthrough.json, and publishes the number.Without it the cost is unmeasurable now and unmeasurable later, which is the CLOUD-851 shape exactly: its store acquisition regressed
checkp50 4.76ms → 10.01ms (2.103x) with 2134 cargo tests green across it, because none of them measures invocation cost. CLOUD-875 is the same class.For the arms that already exist:
REGRESSION_RATIOis 1.30, the measured noise floor 1.102, andwiredcarries an exemption to 1.60 until 2026-11-30 (CLOUD-843).perf-pair's two arms are measured sequentially, so concurrent load does not divide out — run it with nothing else in flight. -
perf-assertholds, and a call carrying no response still does less work than--help. -
Shown able to fail per CLOUD-418: restoring the
commandconjunct, and moving the capture after the projection, each turn a named case red.
-
- Blockers (§8).
blockedByCLOUD-918 — the store has noStream::Responseand no byte-exact read until it lands, so the security predicate's second half is unassertable before then. Ordered after CLOUD-917 for the fidelity vocabulary this records.relatedToCLOUD-776 (the projection this must leave intact) and CLOUD-777 (the registration this rides, landed — not a blocker).
CLOUD-1121 The capture spine recovers a payload the agent has already paid for: a board sweep spends ~234k tokens of context so a gate can read the same bytes off disk
Why
Measured 2026-08-28. Re-verifying the Stage 2/3 Todo column meant re-reading 52 rows with get_issue(includeRelations: true), harvesting them with board-payloads, and linting each. At CLOUD-782's measured ~4.5k tokens per payload that is ~234k tokens of context, spent so a shell script could read the same bytes out of .git/batten-payloads/.
Not one of those bytes needed to be in context. The agent's only role was to cause the call. The consumer was ready-lint, reading a file. And context is re-sent every turn, so each payload was paid again for the rest of the session.
The spine already exists, and every piece is Done
| row | what it built |
|---|---|
| CLOUD-918 | Stream::Response — a tool response the harness handed over, content-addressed in the capture store |
| CLOUD-919 | **every **PostToolUse response is persisted as a local capture, automatically |
| CLOUD-121 | handles: batten capture list / show <handle>, expand without re-running |
| CLOUD-990 | the board gates' remedies already name that store |
[[mint]] issue-read |
fires on the get_issue result and keeps {digest:description} — the boundary already reads the whole payload and retains a pointer |
So on every get_issue this session the bytes were written to the capture store automatically, and the boundary had already parsed them. The agent paid for them anyway.
The gap: everything built is RECOVERY, nothing is SUBSTITUTION
- The capture fires at
PostToolUse, which is after the result has entered context. CLOUD-685 already establishes why it cannot be otherwise — that event is non-blocking and carries noupdatedOutput, so a rule there "could only ever be a sensor". board-payloadsreads the transcript, and the transcript holds those bytes only because they entered context. It removes the second copy — the agent re-typing, CLOUD-526 — and leaves the first one untouched.- **That is why this stayed invisible. **
board-payloadsreads as an optimisation ("byte-perfect, nothing re-typed, the fetch is paid once"), and its own header calls the ~3x "structural and this task does not touch it". Both true, and together they disguise that the expensive half was never anyone's subject. - The host's own oversize spill does not reach this: CLOUD-685 measured its threshold between ~43K and ~74K characters, and a
get_issueat ~18K sails under it.
Nothing in the tree removes the payment. Every mechanism either measures it (CLOUD-415, CLOUD-417), refuses an oversize one (CLOUD-685), or recovers the bytes afterwards (CLOUD-919, CLOUD-990).
The route is open, and one premise is unmeasured
CLOUD-790 (Done, PR #593) overturned CLOUD-673's "a task cannot authenticate to the session's MCP endpoint": the container's session-ingress token does authenticate it, and the 401 was a missing Authorization header. Measured there:
| route | result |
|---|---|
POST /v2/ccr-sessions/<id>/github/mcp |
200 — 55 tools |
POST /v2/ccr-sessions/<id>/mcp?toolbox_mcp_server_id=<toolbox> |
403 MCP server not allowed |
So a task can make tool calls the agent never sees, and land already does exactly that to drop a PR subscription. What is unmeasured is whether the tracker connector's server id is allowed on that route — the 403 above is a live possibility, not a formality. That measurement is this row's first step, and both outcomes are results: allowed, and the fetch moves off the agent entirely; refused, and the row falls back to the resolve-side half, which is worth landing alone.
The same shape elsewhere, censused
| class | owner | state |
|---|---|---|
get_issue has no field projection — ~4.5k to decide on ~1.6k |
CLOUD-782 | Done — the structural third that sets the size of this waste |
save_issue echoes the entire description on every write |
recorded in CLOUD-415's body, owned by nothing | paid ~40 times this session; filed separately, since its remedy is a request shape and this row's is a read route |
Read / Skill unmatched by the PreToolUse matcher |
CLOUD-685 | Todo — the refusal half of the same question |
| hook output is 20% of a long session | CLOUD-417 | Todo |
| refusal prose spends context every fire | CLOUD-1053 | Done — the dereference pattern this generalises |
| a per-CALL ceiling was inexpressible | CLOUD-925 | Done — the mechanism CLOUD-685 needs now exists |
| session cost has no sensor at all | CLOUD-415 | Backlog |
Refinement — Ready (a payload a gate reads off disk must never have to enter context first)
Refinement gate: Definition of Ready & Done. This body carries only specializations.
- Authority boundary (§1). The capture store stays the one home for a recovered payload —
crates/batten/src/capture.rs,Stream::Response, addressed byidentity::capture_fingerprint. This row adds no second store and no second addressing scheme; it adds a way to populate it without the agent, and a way for a gate to resolve from it without a pipe.board-payloadsstays the transcript route and is not deleted: a host with no ingress token still needs it. - Computable predicate (§2), in two independent halves. Either lands alone; together they close the class.
- Resolve-side, unconditional. A board gate resolves its payload from the capture store by issue key, so a caller names an id and pipes nothing. Today's remedy is
batten capture show <handle> --raw | mise run <gate>, which still needs the agent to find a handle — itself a read. Keying the lookup on the id the gate already holds removes the search and the pipe together. Decidable: the same gate over the same row reaches the same verdict with nothing on stdin. - Fetch-side, conditional on the measurement above. A task issues the
tools/callitself with the session-ingress token, writes the response asStream::Response, and prints a handle and a count, never a body. Per-row context cost becomes one line. If the tracker's server id is refused, this half is recorded as unreachable with its HTTP status — the way CLOUD-673 recorded its 401 — and half 1 ships alone.
- Resolve-side, unconditional. A board gate resolves its payload from the capture store by issue key, so a caller names an id and pipes nothing. Today's remedy is
- Effect (§3). Half 1 is
read. Half 2 is awrite, declared on the command surface rather than smuggled into a read verb, into the same storebatten execalready writes. Neither runs onSurface::Hook, so CLOUD-689's 100 ms per-call ceiling is untouched and no runtime is built on the mediated path. - Output & exit (§5). Pointer-only, and here that is the substance rather than a formality: the emitted bytes are an id, a handle and a count — never a body, never a token, never a session or server identifier. A verb existing to keep a payload out of context that printed the payload would defeat itself. Exit follows the one table:
0resolved,1a usage error or unreadable store,3the endpoint could not be reached. It renders no policy verdict, so it never exits2. - **Commit / bump (§6). **
feat(board)— patch until0.1.0, since below that release-plz bumps the patch whatever the type says. Not!for the consumer surface: both halves are additive, the transcript route and the existingcapture showpipe keep working byte-identically, and no exit code or output shape moves.mise run semverdecides the library half, since a new resolution path oncaptureis apubAPI change and that must be asked, not assumed. - Test obligation (§7). Over the compiled binary and the gate suites, shown able to fail per CLOUD-418. The discriminator for half 1 is a negative case, because the easy implementation quietly keeps reading stdin: a gate invoked with an id and an empty stdin must reach the same verdict as the same gate fed the payload. Its partner is the anti-vacuity case — a fixture whose capture store is empty must report could-not-look rather than passing, since an absent payload reading as a clean row is the false green this class produces. For half 2, the token file unset exits
3and writes nothing, mirroringpr-unsubscribed drop's fail-open shape. And the property that makes the row worth landing, asserted rather than hoped: the emitted bytes contain no substring of any body, thepointer_onlycorpus re-run on this path. - Blockers (§8). None.
relatedToCLOUD-790 (which opened the route and supplies the token mechanism), CLOUD-919 and CLOUD-918 (the store this populates and resolves from), CLOUD-121 (handles), CLOUD-990 (the remedy text this replaces), CLOUD-782 (the structural third), CLOUD-685 (the refusal half — that row stops an oversize read, this one removes a right-sized one), CLOUD-526 (why the payload must stay the tracker's own bytes), CLOUD-1118 (the staleness bug in the route this supersedes for the fetch case) and CLOUD-1100 (which retires the Ready grammar into the engine and is the largest consumer of a payload that never enters context).
Acceptance
- A board gate decides over a row with no payload on stdin and none in context, proven by a case where stdin is empty.
- An empty or absent capture reports could-not-look; it never reads as a clean row.
- The tracker connector's reachability on the session MCP route is measured and recorded with its HTTP status, whichever way it comes out.
- A 52-row sweep costs a bounded number of lines rather than ~234k tokens, reported against today's figure.
- No emitted byte is a substring of any issue body.
|
Warning Review limit reachedNext included review available in 55 minutes. View limit detailsLimit details: You’ve used the included review currently available. Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. Review configuration: ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Free Run ID: ⛔ Files ignored due to path filters (2)
📒 Files selected for processing (29)
Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Essentials by visiting https://app.coderabbit.ai/settings/billing. Comment |
Measured over one session's own transcript on 2026-08-31: 43.5 MB of tool content, of which tracker round trips were 13.2 MB — 73% of all tool output, across 973 calls against 208 rows. Bash, Grep and Read together were 1.9 MB. Reading the tree is nearly free; the connector was the entire cost, and a whole document moved every time to convey a ~2 KB delta. That is non-negotiable rule 4 unenforced at the tool boundary. The connector never has to offer a projection. MCP is JSON-RPC 2.0, so Batten can be the CLIENT: dispatch the request, own the response, store it whole, and return a shape the connector does not ship. A client, never a server — CLOUD-204 is untouched and this is the opposite direction. `batten mcp call <server> <method> [params]` resolves the wiring a declared `[[mcp.source]]` names, dispatches, stores the response, and prints a pointer, a delta and the declared reduction. Every tracker identifier lives in `batten.toml`; the crate knows only "dispatch a declared method, reduce by a declared projection", asserted by a rule-1 grep in `tests/mcp_dispatch.rs`. THE TRANSPORT WAS ALREADY VENDORED. `fetch.rs` gains POST, response headers and `spend`, which runs a sequence on one current-thread runtime. The MCP handshake cannot be one batch — `initialize` mints the session id the rest carry back — so it is one exchange to mint and the remaining two together. No feature was widened; `rt-multi-thread` is still not in the graph. THE STORE HOLDS THE UNFRAMED PAYLOAD, and that is a fidelity requirement rather than a preference. `capture::find` resolves a stored response by a key at `id`, which is how `ready lint --issue`, `claim check --issue` and the board gates reach a payload without its bytes entering context. Measured against this repository's own store: the harness files the decoded content with `id` at the top level. Filing the JSON-RPC envelope would have put `id` two levels down and every one of those lookups would have silently resolved nothing over a full store. `payload` takes MCP's own content-block framing off — protocol vocabulary, not a tracker's schema — and is three-valued, so an unrecognised framing keeps its bytes. CLOSING THE RAW PATH IS A GATE, NOT A CONVENTION. `policy/connector-not-granted.rego` refuses a `permissions.allow` entry granting a tool a `[[mcp.result]]` row reduces, and the grants are dropped. What the gate does NOT claim is registration: that happens where the launcher writes its wiring, outside this repository and outside every gate here, and asserting it would be an authority this repo does not hold — `harness-grant` records the same boundary one file over. `--raw` IS RECORDED WHEN SPENT. `capture::record_escape` appends a pointer-only row on both routes, so the invariant's true form — no unreduced route BY DEFAULT — is measurable as a count rather than asserted. TWO DEVIATIONS FROM THE ROW'S READY BLOCK, stated rather than absorbed: * §5 says exit 2 on a dispatch error. §7's table reserves 2 for a policy verdict with no per-verb exception, and every host with a pre-tool hook reads 2 as deny, so a dropped network would read as a refusal. Dispatch failure and an unreadable config are exit 3; malformed params are exit 1. AGENTS.md: the spec wins. * `batten.toml` is a protected path. Its own refusal names "change it in a pull request", which is this branch, so BATTEN_HOOK_BYPASS was spent for the edit. AGENTS.md's scope amendment lands here per non-negotiable rule 2, paid for by compressing rules 7-8 and the memory paragraph: the file sits on its budget ceiling, so the prose is net-zero lines rather than a raised threshold. The compiled-binary tier covers everything up to the socket. A completed dispatch is not hermetically testable — `fetch` is https_only and nothing signs a loopback CA — which `fetch.rs`'s header already records for its own case. Refs: CLOUD-1260, CLOUD-745, CLOUD-919, CLOUD-1121, CLOUD-418, CLOUD-204 BREAKING CHANGE: `fetch::Response` gains a `headers` field, so a caller constructing one by literal no longer compiles. `mise run semver` names the lint (`constructible_struct_adds_field`) and the honest type is declared rather than worked around: the field is what carries a session-bearing protocol's session id back, and hiding it behind an accessor would move the break rather than remove it. Below 0.1.0 release-plz still bumps the patch whatever the type says; the marker is what the changelog and the history depend on.
… body
`save_issue` was 435 calls returning 4.98 MB on the measured session — 38% of the
13.2 MB the tracker cost, which was itself 73% of all tool output. A write echoes
the entire stored description back seconds after the author sent it: a `patch`
sends a few hundred bytes of anchors and gets a full body back, and that body then
sits in context, re-sent every turn, for the rest of the session. It is the one
operation whose input the caller already knows.
One row on the table CLOUD-1260 created: `reduce = "acknowledge"` over
`save_issue`, carrying `{id, status, url}`. The handle is not a field because it
is not a tracker field — it travels in the verb's own pointer line, which is where
a pointer belongs.
THE BOUND IS LENGTH AND SCALAR, NOT `facts::Reduction::Token`'S, and the
difference is this row's own near-miss caught before it shipped. That reduction
additionally refuses any value carrying whitespace, and it is right to: its
product reaches the POLICY INPUT, where a module could lift a sentence into a
`subjects` pointer. This arm's product reaches the CALLER, and the thing being
kept out is a 15k-character description — which the length bound stops on its own.
Copying the whitespace clause across would have silently dropped `status = "In
Progress"`, the one field this row's remedy names, with every test still green. A
bound borrowed from a surface with a different threat model is how a reduction
quietly stops answering the question it was built for. Pinned by an assertion so
it cannot come back.
BOTH LIVE CONSUMERS ARE PROVEN, NOT ASSUMED, because a fix that starved them would
trade one defect for two. CLOUD-815 compares the stored body against the sent one
— that comparison IS the detection — and CLOUD-1118 wants the post-write body so a
lint straight after a write decides over what was actually stored. Both read the
CAPTURE STORE, and every response is stored WHOLE before anything is reduced. The
case drives the engine's own `PostToolUse` event to file the response, then
resolves it back through `capture find --raw` and asserts the body survives
entire. An earlier draft of that case round-tripped JSON through serde and
asserted nothing about Batten — CLOUD-418's non-discriminating coverage — and was
thrown away rather than kept.
There is no `updatedOutput` and none was looked for: re-verified twice on
2026-08-31, absent from the runtime and from the published hook reference. That
limit is real and SCOPED — a POST-tool hook cannot rewrite a result — and it says
nothing about a client that owns the response, which is what this reduces.
Refs: CLOUD-1122, CLOUD-1260, CLOUD-815, CLOUD-1118, CLOUD-919, CLOUD-685, CLOUD-418
…ct its premise `[[rule.external]]` declares ONE path under ONE root. A Claude Code remote session's launcher mints its MCP config per session, so the path is a glob, the match arity is not one, and the id a declaration needs cannot be authored for a name that does not yet exist. Each of the three would be fatal alone. ANSWER 2 IS TAKEN: a producer resolves such a set outside the engine and a module reads what was written. The engine grows no glob. Cost of rejecting answer 1 (a declared glob under a declared root): it is the one that looks like a small change and is not. An id's arity goes from one node to many, every reader must handle the empty match as could-not-look rather than as absent, and what is spent is the schema's own line — "the difference between a fact and a filesystem scanner", which is the whole safety property the family was built to state. Cost of rejecting answer 3 (decline outright): CLOUD-1260 arrived as a second, unrelated consumer of the same file, so declining would leave one globbing loop written twice outside the engine. AND A CORRECTION TO THE ROW'S OWN PREMISE, measured rather than argued. That row expects the second consumer to strengthen answer 2. It cannot be served by answer 2 at all: `facts.rs`'s `Sourced` states that "no byte of it reaches THIS record — `rows_in` reduces the buffer to a COUNT at the boundary", so the whole record family is payload-free BY CONSTRUCTION. That is exactly right for a module deciding a predicate, and structurally incapable of carrying an endpoint and headers, which is what DISPATCH needs. So answer 2 serves the permission question CLOUD-1163's unit 4 asks — recorded there as buildable — and remote dispatch is now a known gap with a stated cause rather than an assumed solution. WHAT DELIBERATELY DID NOT LAND: no `[[fact]]` or `[[recorder]]` row. Its only reader would be a module that has not been written, and unit 4's migration is CLOUD-1163's and out of this row's scope. A declaration nothing reads is the dead gate this repository refuses everywhere else, so the shape stays available and is declared when a reader exists. The decision is recorded where the reader who hits the limit is looking — beside the `[[mcp.source]]` rows it bounds, and in the module header — rather than only on the tracker. `docs` and no bump: the engine is unchanged, and the negative half this row cares about holds for exactly the reason it did before, asserted by `tests/mcp_dispatch.rs`. Refs: CLOUD-1251, CLOUD-1260, CLOUD-1163, CLOUD-1167, CLOUD-1154, CLOUD-1170
68d9058 to
1a79d8a
Compare
|
❌ The last analysis has failed. |
|
/fast-forward |
Batten becomes an MCP client: it dispatches the call the session was going to
make anyway, stores the response whole, and hands back a pointer, a delta and a
declared reduction.
Closes CLOUD-1260
Closes CLOUD-1122
Closes CLOUD-1251
One PR because all three edit
batten.tomland regenerate the same derivedschema, and CLOUD-1260 creates the
[[mcp.result]]table CLOUD-1122 adds a rowto. Three commits, one per row.
Why
Measured over one session's own transcript on 2026-08-31: 43.5 MB of tool
content, of which tracker round trips were 13.2 MB — 73% of all tool output,
across 973 calls against 208 rows. Bash, Grep and Read together were 1.9 MB.
Reading the tree is nearly free; the connector was the entire cost, and a whole
document moved every time to convey a ~2 KB delta. That is non-negotiable rule 4
unenforced at the tool boundary.
The two cheap fixes were already priced and both fail: only 8% of issue bytes are
byte-identical repeats, and the connector's schema is
additionalProperties: false, so there is no projection to ask for. What is left is to stop being abystander to the call.
What landed
feat(mcp)!— CLOUD-1260.batten mcp call <server> <method> [params]resolves the wiring a declared
[[mcp.source]]names, runs the JSON-RPC session(
initialize→ initialized →tools/call), stores the response, and prints thedeclared reduction. The transport was already vendored (CLOUD-745):
fetch.rsgains POST, response headers and
spend, which runs a sequence on onecurrent-thread runtime. No feature was widened —
rt-multi-threadis still not inthe graph.
feat(mcp)— CLOUD-1122.reduce = "acknowledge"onsave_issue, which was435 calls returning 4.98 MB on the measured session — 38% of the tracker's cost —
because a write echoes the whole stored description back seconds after the author
sent it.
docs(facts)— CLOUD-1251. The decision, in writing, with the cost of the tworejected answers, plus a correction to the row's own premise.
Three things a reviewer should look at first
The stored shape is a fidelity requirement, not a preference.
capture::findresolves a stored response by a key at
id, and that is howready lint --issue,claim check --issueand the board gates reach a payload without its bytesentering context. Measured against this repository's own store: the harness files
the decoded content with
idat the top level. Filing the JSON-RPC envelopewould have put
idtwo levels down, every one of those lookups would havesilently resolved nothing over a full store, and nothing would have gone red.
mcp::payloadtakes MCP's own content-block framing off — protocol vocabulary,not a tracker's schema — and is three-valued, so an unrecognised framing keeps its
bytes.
Closing the raw path is a gate, not a convention.
policy/connector-not-granted.regorefuses apermissions.allowentry granting atool a
[[mcp.result]]row reduces, and the grants are dropped. What it doesnot claim is registration: that happens where the launcher writes its wiring,
outside this repository and outside every gate here. Asserting it would be an
authority this repo does not hold, and
harness-grantrecords the same boundaryone file over.
hook -> mcpis forbidden, not justhook -> fetch. The existing row keeps aruntime off the mediated path (CLOUD-689's ceiling, CLOUD-747's bound).
mcpreaches
fetch, andmodule-layeringdecides over direct edges — so withoutthe new row the guarantee was routable around in one hop.
Deviations from the Ready blocks, stated rather than absorbed
Both are recorded on CLOUD-1260 itself.
2on a dispatcherror; the house-style table reserves
2for a policy verdict with no per-verbexception, and every host with a pre-tool hook reads
2as deny — so a droppednetwork would have read as a refusal. Dispatch failure and an unreadable config
are
3; malformed params are1. AGENTS.md's "where they disagree the specwins" is the tiebreak.
-Jdata channel on the verb. That contract is that a document isemitted unconditionally, and this verb's document is a server's answer, so
there is none when the exchange did not happen. It takes
exec's split: thereduction on stdout, the pointer and delta on stderr.
Two findings reported rather than repaired
Both are on their rows; neither is fixed here, because repairing landed code this
bundle did not scope would widen the PR.
main. It asks for zerohits under
crates/batten/for any tracker method name. Run over code withcomments stripped,
lib.rsdeclaresREAD_TOOLandWRITE_TOOLas crateconstants naming two tracker methods. This change adds none of its own, and the
compiled-binary tier asserts exactly that, scoped to what the change introduced.
.claude/rules/rust.md's concurrency section is stale — it still saystokioappears nowhere inCargo.lock, which CLOUD-745 falsified.Bounds
A completed dispatch is not hermetically testable:
fetchishttps_onlyandnothing signs a loopback CA, which
fetch.rs's own header already records for itsown case. The compiled-binary tier covers everything up to the socket — config
loading, wiring resolution, each refusal's exit class, and the pointer discipline
over every message. What a successful exchange returns is the module tier's, over
a response value. Claiming a live dispatch case would be coverage rather than a
test.
Scope amendment
Per non-negotiable rule 2, AGENTS.md's scope reminder is amended in the same
change: a component that dispatches calls and reduces responses is a real
expansion of a core that says it is "not a hook runner… or reference monitor".
The file sits on its budget ceiling, so the prose is net-zero lines — paid for by
compressing rules 7–8 and the memory paragraph rather than by raising the
threshold.
Verification
mise run verify—fast-forward-green, rebased on latest main, ci + cross +commit-lint all pass.
mise run policy-test— 419 passed across 37 bundles,including four new cases pinning both directions of the layering edges.
Refs: CLOUD-1260, CLOUD-1122, CLOUD-1251, CLOUD-745, CLOUD-919, CLOUD-1121,
CLOUD-418, CLOUD-204