Repository navigation
fix(hooks): a comment records the issue key, never the comment's own uuid - #418
Conversation
CLOUD-491 A live `plan-hold` did not stop a container restart, and nothing records that it was live — the hold ships with no sensor
Why CLOUD-451 landed Measured 2026-08-12, ~22:12 UTC. A hold was armed ( So the container went down with four live tracked tasks, one of them a full Reproduced 2026-08-12, ~23:45 UTC, in the session grooming this issue: a hold was armed, and the container was restarted roughly a minute later, killing it. Two independent instances now, and the second one left artifacts the first did not — see the measurements below. The finding is not "the hold is wrong". It is that the hold cannot be graded. A mechanism that occupies the container leaves no record of having done so, so from inside there is no way to distinguish the hold having been live and reclaimed anyway from the hold having already exited from a platform event no occupancy could defer. All of them produce the same observation — a fresh container — and the issue's acceptance is written in terms nobody can check after the fact. That is the sensor-without-a-gate shape inverted: a mechanism with no sensor. It is why an instance can be reported but not diagnosed, and why the honest first move is evidence rather than a bigger hammer. Measured 2026-08-12 ~23:30–23:46, and it changes the mechanism The first draft of this issue assumed one kind of container replacement. There are two, they are indistinguishable from inside without a sensor, and they have opposite consequences for any sensor built from a local file. 1. The 23:30 boot destroyed everything. A session was demonstrably alive at 23:28 — this issue's own last edit — and nothing it wrote survived anywhere writable. 2. The 23:45 boot preserved the disk. Same appearance from inside, opposite consequences. A heartbeat file under 3. The surviving hold directory was empty. 4. The last 182 seconds of writes did not survive, and this falsifies the predicate this issue shipped with. Every surviving pre-boot write stops at 23:42:15 — the MCP logs of four different servers, Serena's log, So the first draft's predicate — The heartbeat question, stated so it can be settled rather than argued. Two different things are called a heartbeat here and only one is cheap:
The first is a prerequisite for deciding the second. Ship the sensor, read it, then decide. The activity-versus-existence question — every wait in this repo that survives does I/O, and this one does not — is CLOUD-500, deliberately blocked on this issue for that reason. What the sensor must distinguish, and how The hold records why it stopped, not merely when it last ran — absence of an intentional-exit record is the signal, and absence is the one reading robust to losing the tail. Two line kinds in one appended file:
The residual error is bounded and in the conservative direction: a hold released inside the lost-write window reads as "was live". The residue probe of the first draft does not work either, and is replaced. It separated the last two rows by whether prior-container residue exists, naming Refinement — Ready
Acceptance
Measured in the session that also filed CLOUD-488; the plan whose approval button was destroyed was the one grooming CLOUD-427. The 23:30–23:46 measurements and the empty-hold-directory artifact were added in the grooming session, which was itself restarted twice while doing it. CLOUD-514 Nothing prices filing over fixing, so spinning off a defect in the PR's own diff is arithmetically cheaper than finishing it
Why Every gate in this repo prices failing to record something. Nothing anywhere prices the opposite: recording something instead of doing it. Filing satisfies every one of those gates at once and costs a few seconds, while finishing costs a diff, a suite and a landing. For an agent under pressure that is not a temptation, it is arithmetic — and the board becomes the escape hatch every guardrail points at. AGENTS.md already names the behaviour: "A punt is any deferral you could have closed … offering an action you are already authorized to take." That rule is prose, and prose is feedforward only. Nor is the substitution a fair trade. Across studies of admitted technical debt only 26.3–63.5% of it is ever removed, with median lifespans of 18–172 days and instances surviving more than ten years; in trackers specifically the repayment distribution is severely skewed, median 25 hours against a mean of 872 hours. A ~35× median/mean gap is the signature of a long tail never repaid at all. Filing does not defer a fix, it converts one into a weighted coin-flip. Measured 2026-08-13, PR #390. CLOUD-513 is a defect in code written in that PR: two new fixture suites read ambient git config, passed No reviewer is present at the moment of the choice, so the cost has to land on the author. Landing here is trunk-based: a branch fast-forwards onto Two mechanisms are ruled out before any is proposed 1. Judging the spin-off is forbidden. "Is this issue related enough to the PR to belong in it?" and "should this have been fixed instead?" are both model verdicts, which non-negotiable 3 refuses: a gate resolves to a command and an exit code over an object it decides. CLOUD-505 hit the identical wall, and its resolution is the template — do not judge the content, price the action. 2. A time window is measured, and rejected. The obvious credential-free proxy for "issues this branch filed" is "issues created between the PR opening and its merge". Measured over the last 40 merged PRs against every It fires on essentially everything, and structurally rather than tunably: this repo runs a fleet, so a window captures every session's filings. PR #373's window alone holds 20 issues, none of them its author's. A gate at 99.5% is switched off the first day. That measurement is what makes a receipt load-bearing rather than convenient. Which branch filed which row has never been recorded anywhere — not in git, not on the PR, not on the board. Without a record the predicate is not merely noisy, it is not computable at all. Refinement — Ready Refinement gate: Definition of Ready & Done. This body carries only specializations.
Chat is not a fourth sink, and that is a memory property rather than a style rule A finding stated in prose and never written durably is context-resident state. Long-horizon agents lose exactly that: unbounded context forces truncation and summarisation, attentional dilution degrades recall of any single item as context grows, and material in the middle of a long context is recalled least reliably — which is why an external store holding raw records at full fidelity is the standing recommendation. So a finding left in an aside is not merely undocumented, it is predictably lost, and the container reclaim this session survived is only the abrupt version of the same thing. How this gets satisfied without doing the work Three, found by attacking the mechanism rather than reasoning about it, each with local evidence. Two are closed above by changing the design; the third is a residual, stated with its failure direction. A ready-lint verdict is forgeable if the agent supplies the payload — closed by moving the lint into the recorder, which lints the tracker's response instead. The finding stands and is why §1 reads as it does; the original remedy (a second receipt, made trustworthy elsewhere) was the wrong fix for it. The measurement that made the hole real. An unfiltered listing launders sink 2 — closed by dropping the requirement, because it bought nothing. The record survives a branch restart — and that is fail-closed here, which is why it is a residual and not a blocker. What this cannot do It cannot judge the reason, and a minimal-compliance path still exists — a Ready block written to satisfy Gating creates pushes pressure toward commenting instead, and that is accepted rather than fixed. A comment on the row that already owns a finding is a legitimate durable home — cheaper than filing by design, since the friction is meant to sit on the impulsive path and not on honest recording. What makes it acceptable is that the pressure runs toward recording in the right place, not toward silence. It is still recorded, so the ratio is observable: CLOUD-475 records that a comment already counts as a durable home. Re-open predicate: re-open if the observation window shows spin-off comments rising while spin-off creates fall and fixes do not, which is the signature of the toll being routed around rather than paid or avoided. Acceptance
Not in this issue Deciding whether a given spin-off was legitimate — the judgement the gate must never make. The In Review transition gate, which is CLOUD-512's. And retrofitting receipts for branches predating the recorder, which is why the gate fails open on their absence. |
42cb84e to
db556f6
Compare
db556f6 to
966d7f6
Compare
966d7f6 to
85afe6f
Compare
85afe6f to
3d7c250
Compare
…uuid
Found by the recorder's first five live rows, which came out as
comment 4d16245a-43ea-49ae-b67d-c2ee0b64b96e 2026-08-13T06:46:26.338Z -
A `save_comment` response is the COMMENT object: its `.id` is the comment,
and it carries no reference to the row the comment landed on. Taking `.id`
for both kinds therefore filled an issue-key column with uuids. Every count
stayed correct, so the create-versus-comment ratio survived — but which row
a comment landed on did not, and that is sink 2's entire definition.
The key is in the input instead, as the parent reference. A reply carries
`parentId` and no `issueId`, and a comment on a project, document or
milestone is not a board row at all; both record `-`, the same "could not
look" this file already draws for a verdict. Never the uuid: a uuid in an
issue-key column is a wrong answer wearing a right answer's shape, and it
reads as data rather than as a gap, which is strictly worse than the gap.
Two defects in the new rows, both found by running them:
`${2:-CLOUD-42}` substitutes the default over an explicitly EMPTY argument,
so the reply fixture silently exercised the issue-id case instead. `${2-…}`
is the form that distinguishes unset from empty.
A uuid-shape assertion over the whole line matched a correct row, because
`CLOUD-42` contains `-4`. The regression row now reads the id COLUMN and
holds it to `CLOUD-<digits>` or `-`.
Refs: CLOUD-514
3d7c250 to
43f8242
Compare
|
|
/fast-forward |



A defect in what PR #399 shipped, found by that PR's own first live output twenty
minutes after it merged.
What was wrong
The recorder's first five rows came out as:
Comment uuids, in a column that is meant to hold an issue key. A
save_commentresponse is the comment object — its
.idis the comment, and it carries noreference to the row the comment landed on at all. Taking
.idfor both kindstherefore filled the key column with uuids.
Every count stayed correct, so CLOUD-514's create-versus-comment ratio survived. What
did not survive is which row a comment landed on — and "a durable comment on the
existing row that owns it" is sink 2's entire definition. The residual that issue records
("re-open if comments on rows the search did not surface start appearing") was
unobservable.
The fix
The key is in the input, as the parent reference.
issueIdis the only parent that is aboard row. A reply carries
parentIdand noissueId— the thread determines the issue,which a hook cannot see — and a comment on a project, document or milestone is not a row
at all. Both record
-, the same "could not look" this file already draws for a verdict.Never the uuid. A uuid in an issue-key column is a wrong answer wearing a right
answer's shape: it reads as data rather than as a gap, so nothing downstream can tell the
two apart. That is strictly worse than the gap.
Verified live rather than only in fixtures — the newest row in this clone reads
comment CLOUD-520 …, against the five uuid rows above it.Two defects in the new test rows, both found by running them
${2:-CLOUD-42}substitutes the default over an explicitly empty argument, so thereply fixture silently exercised the issue-id case instead.
${2-…}is the form thatdistinguishes unset from empty.
A uuid-shape assertion over the whole line matched a correct row, because
CLOUD-42contains
-4. The regression row now reads the id column and holds it toCLOUD-<digits>or-.Coverage
tests/board-write-record.batsat 16 rows; three#MUTANTrows now, the new onerestoring the shipped defect so the regression case reddens.
mutantgreen at 20/20.Full suite 1658/1658, run with
GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/nulland no
BATTEN_*_BYPASSexported. That second condition matters: an earlier sweepreported 1489/1489 with
BATTEN_CLAIM_GUARD_BYPASS=1in the environment, which disabledclaim-guardand would have failed five of its rows.test:batsinherits ambient bypassvariables exactly as it inherits ambient git config — recorded on CLOUD-513, whose
[tasks."test:bats".env]block is where both belong.Why this closes nothing
DO-NOT-CLOSE
Three issues are named and none is completed here:
filed-here-check,the gate) is unbuilt and blocked on the observation window the sensor produces. The row
stays In Review.
comment recording that
BATTEN_*_BYPASSleaks in the same way. No code here touches it.claim-checkrefusing an author their own work.This PR is a consequence of that gap, not a fix for it; it was authored with
claim-guardbypassed because no minting path exists for follow-up work on an In Reviewrow.
Refs: CLOUD-514