fix(pull): make the pilot flag a true either/or + soak measurement set (#1766) - #1982
Merged
Merged
Conversation
…ased rows (#1766) Prep for the #1766 production soak. Two defects made a flag-flip soak unable to produce the evidence its ACs ask for. 1. The pilot flag was purely ADDITIVE, so push and pull ran in parallel INSIDE one agent. `CapacityManager.acquire` had no pilot branch: a free slot still meant a push, so a row only reached the queue on capacity overflow. When one did queue, the backend's own `backlog_service` `drain_next` raced the agent's worker through the SAME `db.claim_next_queued` — no double-run (the claim is one atomic UPDATE) but the winner was whoever polled first, and the backend was structurally favoured (it drains on slot release plus a 60s orphan sweep; a worker idles up to 15s). Pull coverage was therefore near-zero and drain-biased. Worse, the two paths were bounded by two independent counters — Redis ZSET vs the container's worker pool, both sized max_parallel_tasks — so a pilot could run up to 2x its configured concurrency, invisible to canary S-02 (ZCARD only) and S-01 (excludes leased rows by design). `pull_owns_dispatch` now makes a pilot's autonomous work queue-ONLY: never admitted, never drained by the backend, so its worker pool IS its capacity. This is the pilot-scoped slice of #1081 Phase 5 ("capacity becomes physical"). Interactive triggers are deliberately EXCLUDED and keep the synchronous push path — the scope cut in TARGET_ARCHITECTURE.md Open Question 7, load-bearing because one FIFO ordered by queued_at would park a human turn behind N batch tasks and let N workers claim two turns of one session concurrently. Fails safe to push; an empty allowlist (the default) leaves every path byte-for-byte unchanged. 2. E-05 fired on every pull turn older than 60s. A pull-claimed row is `running` with a NULL claude_session_id BY DESIGN, and `mark_no_session_executions_failed` — the very sweep E-05 watches — already carries `lease_expires_at IS NULL` for that reason. E-05 was flagging the rows that sweep deliberately leaves alone, which would have buried a real #106 regression under noise for the whole soak window and broken the "canary green throughout" AC. Mirrors S-01's exclusion; a NULL-lease control still fires. Closes the T3.6 gap in docs/testing/PULL_MIGRATION_TESTING.md §8. Full unit suite green (7199 passed).
… in §8 (#1766) Adds §9 — the queries that turn "we ran it for a few days" into a verdict, each mapped to a #1766 acceptance criterion, plus a pre-flight baseline step, abort criteria, and the rollback procedure. Motivated by a gap worth stating plainly: `instance-monitor` has zero references to leases, claim_token, redelivery_count or queue depth, so it reports generic health and stays green whether pull carries the fleet or does nothing at all. AC "instance monitor active" is therefore not sufficient evidence on its own; M1/M3/M5 are the substitute (and are the natural candidates to fold into its deep-probe later). M1 is the load-bearing one: if `pulled` is ~0 the soak measured nothing and every downstream conclusion is void — the exact failure mode the pre-#1766 additive flag produced on a healthy instance. PostgreSQL flavour with explicit ::timestamptz casts, since every timestamp column is Text holding ISO-Z (Invariant #16). §8 updated: E-05 lease-awareness struck (closed by this branch) with E-01 left open, and the push/pull coexistence defect recorded alongside it.
…t + import fix) Two problems, one root cause. `tests/lint_sys_modules.py` failed on a bare `sys.modules.pop(...)` in the new test's bootstrap — copied from test_capacity_manager.py, but this file imports every backend module lazily so it never needed the preamble. Removed rather than papered over with a _STUBBED_MODULE_NAMES declaration: test_1081_physical_meter.py warns that popping `utils` actively unbinds conftest's canonical importlib registration, so the idiom is harmful here, not merely unnecessary. Removing it exposed the real defect. `services/agent_service/__init__.py` eagerly imports helpers/lifecycle/crud/deploy/terminal, so the guard's `from services.agent_service.pull_mode import is_pull_pilot_agent` inside `drain_next` dragged the whole agent-lifecycle stack (→ models) into backlog_service. Under a test that stubs models/database that surfaced as `ImportError: cannot import name 'AgentGitConfig' from '<unknown module name>'` — naming nothing to do with the actual dependency. Verified against clean dev: tests/unit/test_backlog.py passes 34 there and failed 7 on this branch, so this was introduced here, not pre-existing. The predicates now live in `services/pull_pilot.py` — stdlib-only, no package side effects. `pull_mode` re-exports them so crud/lifecycle and the existing #1081 tests are untouched; the three dispatch-path callers (capacity_manager, backlog_service, routers/internal) import the leaf directly. This also retires a latent instance of the same hazard: capacity_manager's pre-existing shadow-meter imports had the identical shape and were simply never exercised by a stubbed test. A capacity/backlog module should not need the agent CRUD stack to answer "is this name in an env var". test_backlog isolation 34 passed; lint clean; combined pull/capacity/canary suites 284 passed. NOTE: the full unit suite is order-sensitive on dev independently of this branch — clean dev fails 2 in test_1081_physical_meter under full-suite ordering while passing in isolation. Both files pass in isolation and together here; deferring to CI's seeded base/head runs.
This was referenced Aug 4, 2026
vybe
approved these changes
Aug 4, 2026
vybe
left a comment
Contributor
There was a problem hiding this comment.
Validated via /validate-pr — clean pass: fail-safe direction is push, empty-allowlist inertness tested both sides, E-05 lease exclusion mirrors the sweep predicate with a control test. Merging in sequence. Follow-ups noted: stale §Capacity & Backlog paragraph in architecture.md (acquire/drain now have a pilot branch), and #1766 status bump is manual (no closing keyword by design).
vybe
pushed a commit
that referenced
this pull request
Aug 6, 2026
…g critical (#1990) (#2046) * fix(canary): make E-01 lease-aware so a healthy pull turn stops paging critical (#1990) E-01 ("no running row older than execution_timeout_seconds + 300s") was the last invariant reading the running-row set with no lease-awareness, after S-01 shipped with the pull phases and E-05 got it in #1982. A #1081 Phase-3 pull-CLAIMED row (lease_expires_at IS NOT NULL) is status='running' but owned exclusively by the lease-reaper, which re-queues it (redelivery_count++, started_at reset) or poison-parks it terminal at MAX_REDELIVERY. Its age is not evidence of the stuck-execution class E-01 detects: the deadline governing a leased row is its lease, and a different component owns the recovery. The overlap is exact, not merely awkward. claim_next_queued stamps lease_expires_at = started_at + (execution_timeout_seconds + SLOT_TTL_BUFFER) (pull_coordination_service.py) -- the same threshold E-01 builds. So age > threshold became true at the precise instant the lease expired, which is the instant the reaper's recovery window OPENS. The 300s buffer exists so E-01 fires only after the owning component has had its window; against a leased row that head-room was zero, and E-01 paged critical for the whole gap between lease expiry and the reaper's next sweep -- on every re-delivery, per execution, while the machinery worked as designed. The db layer already encoded this ownership split, including on the very sweep whose failure E-01 exists to detect: mark_stale_executions_failed carries lease_expires_at IS NULL, as do get_running_executions, get_running_executions_with_agent_info, fail_stale_slot_execution and mark_no_session_executions_failed. The canary was the layer that hadn't caught up. This blocks the #1766 soak, whose AC is "canary green throughout" and whose abort criterion is a critical canary violation on the pilot -- so a healthy soak would abort itself, or worse, train the operator to ignore the one invariant that matters most during it. The exclusion is keyed on the lease, not a blanket silencing: a NULL-lease (push) row of identical age still fires, the threshold boundary is unchanged, and an absent snapshot key fails OPEN (an older image that never populated the column keeps being checked). Closes T3.7 in docs/testing/PULL_MIGRATION_TESTING.md section 3. E-02 assessed and deliberately NOT excluded ------------------------------------------- E-01, E-05, S-01 and E-02 are the four readers of running_exec_ids; E-02 is the one that must keep seeing leased rows. It asks a different question -- was this id terminal in a previous cycle and is it non-terminal now -- whose answer does not depend on who owns the row. The reaper cannot produce that transition (requeue_expired_lease / park_expired_lease both CAS on status='running' with a past lease, so a terminal row is unreachable to them), and re-delivery preserves the execution_id by construction (#1084/#525 are execution_id-scoped), so a terminal id reappearing as running/queued is exactly the corruption E-02 exists to catch. Pull is MORE exposed to it than push -- a late worker result and a reaper pass race for the same row -- so copying this exclusion into E-02 "for consistency" would blind it on the very path #1081 introduces. Recorded in the E-01 docstring, snapshot.py, the catalog and architecture.md, and guarded executably by a test. Testing ------- tests/unit/test_1990_e01_lease_awareness.py -- 13 tests. Under tests/unit/ because CI runs `cd tests && pytest unit/` only; the pre-existing E-01 tests sit in tests/test_canary_invariants.py, which no workflow executes (trinity#2037), so a guard placed beside them would never go red. Verified red without the fix: 5 failed, 8 passed -- the 8 being the NULL-lease controls, which must pass either way. Load-bearing: a NULL-lease control row of identical age still fires, asserted both alone and alongside a leased row in one snapshot. Every other test in the module would still pass if someone "fixed" E-01 by deleting it. Boundary asserted exactly (threshold silent / threshold+1 fires) off integer offsets from one literal T0, per the #1909 clock discipline. The E-02 verdict is a test, not a comment. Two mirrored cases added to the in-place suite for locality beside the rest of E-01's coverage, with a pointer to the enforcing module. Full unit suite: 7814 passed, 18 skipped (7801 baseline + 13). Also: the E-01 Slack runbook now states a firing row is always NULL-lease, so on-call is not sent after the lease-reaper for a push-path stall. Closes #1990 * docs(canary): five functions carry the six lease selectors, not five * fix(canary): bound E-01's lease skip instead of silencing leased rows forever Review finding (dolho, #2046): the unconditional `continue` on any non-NULL lease is wider than the bug requires, and it locks in permanent silence for a real failure — a stuck or dead lease-reaper. `requeue_expired_lease` sets `started_at=now` and clears `lease_expires_at` in one atomic UPDATE, and the reaper runs on `cleanup_service`'s cycle, so under a healthy reaper a leased row's observable overdue-ness is bounded by one interval. The false-positive window this PR targets is small and bounded; the blanket skip was not. Nothing else covered the failure case. E-01, E-05 and S-01 all exclude leased rows; E-02 only catches terminal→non-terminal reversals; there is no lease-overdue invariant; and /deep-probe v1.3 (trinity-ops-agent#297) ships M1/M3/M5, not M4. Meanwhile PULL_MIGRATION_TESTING.md §9 M4 — "a running row past its lease_expires_at that the reaper has not touched is a reaper failure" — is a #1766 soak ABORT criterion. It would have had no automated owner. So the skip is now bounded by LEASE_REAPER_GRACE_SECONDS (600s = 2 x cleanup_service.CLEANUP_INTERVAL_SECONDS): one reaper cycle of worst case plus a full spare cycle. A healthy reaper never fires; one that has merely missed a cycle never fires; one that has STOPPED fires within grace + one canary cadence. The constant carries its derivation and is asserted against the live interval so the coupling cannot drift silently. A firing leased row is a DIFFERENT diagnosis — the lease-reaper failed, not cleanup_service's stale watchdog — so `lease_expires_at`, `lease_overdue_seconds` and the grace ride in observed_state, the signal_query names it, and the runbook hint now branches on `lease_expires_at` instead of asserting "a firing row is always NULL-lease". Two states that now fire deliberately, both on-call-actionable and documented: the #1085 re-delivery governor hold (opt-in, 300s pause TTL, absorbed by the spare cycle unless sustained) and reaper saturation (`find_expired_leases` takes limit=500/pass). Also from the review: `_parse_lease` collapses None/""/unparseable to "no lease", so a future collector coercing NULL to "" widens nothing — a bare `is not None` would have read that as a live lease and silenced the row.
7 tasks
obasilakis
added a commit
that referenced
this pull request
Sep 9, 2026
…igger is stranded (#2524) Completes #1081 Phase 4. After this, every autonomous trigger can run on the durable queue; only the interactive ones stay on push, which is the deliberate Open Question 7 scope cut (#1982/#1989). FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]` inside one `asyncio.gather`, so the batch existed only while the request that started it did. A pull-claimed subtask returns nothing to collect (the row is queued and the turn runs later in the agent's worker), and nothing could answer about a batch afterwards — no `async_mode`, no status endpoint, a disconnect lost it. The batch now lives on `schedule_executions`: every subtask row carries `fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used to be a dict key no async batch could reach), and `build_aggregate` rebuilds the result from the rows. Adds `async_mode` and `GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch belongs to that agent: `fan_out_id` is opaque but not secret. Two decisions the issue asked for. `max_concurrency` keeps its meaning and needed no branch. The semaphore stays around the `execute_task` call: on push that call spans the whole turn so it paces dispatch as before, and under pull it returns in milliseconds so the worker pool becomes the cap — Phase 5's "capacity becomes physical", by construction. Deleting it, as first planned, would have fired N dispatches at an agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull. The outer deadline bounds the WAIT, not the work — a contract change. A still-open subtask now reports `running`, not `failed`; the batch still reports `deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and on push the old cancellation was half-illusory anyway (it abandoned the HTTP call while the agent kept running and billing the turn). After a deadline the status endpoint is the source of truth. A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a receipt to poll" — true and beside the point: it does not need a receipt, it needs to BLOCK CORRECTLY while the turn happens elsewhere, which `sync_waiter.wait_for_sync_terminal` already did for `/task`. `dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED return, wait for that row's terminal and rebuild from it. Nothing signals that waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail latency on an a2a call against a pilot, deliberately not worth a second signalling path. `PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an enumerated allow-list on purpose: a structural test forbids deriving it, because that would hand reach to the next autonomous trigger with nobody checking dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger` is kept and still tested, against a synthetic narrowing. The loop advance (#2523) and fan-out join share one `_terminal_side_effects` shim off `spawn_task_terminal_event`, with separate guards so one raising cannot cost the other its terminal. Migration 0051 + the SQLite twin: `fan_out_task_id`, plus `idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one batch on every fan-out terminal. Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here: `effect_guard` still fails open when the execution id is absent, and a fan-out multiplies that by N. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa
obasilakis
added a commit
that referenced
this pull request
Sep 9, 2026
…igger is stranded (#2524) Completes #1081 Phase 4. After this, every autonomous trigger can run on the durable queue; only the interactive ones stay on push, which is the deliberate Open Question 7 scope cut (#1982/#1989). FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]` inside one `asyncio.gather`, so the batch existed only while the request that started it did. A pull-claimed subtask returns nothing to collect (the row is queued and the turn runs later in the agent's worker), and nothing could answer about a batch afterwards — no `async_mode`, no status endpoint, a disconnect lost it. The batch now lives on `schedule_executions`: every subtask row carries `fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used to be a dict key no async batch could reach), and `build_aggregate` rebuilds the result from the rows. Adds `async_mode` and `GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch belongs to that agent: `fan_out_id` is opaque but not secret. Two decisions the issue asked for. `max_concurrency` keeps its meaning and needed no branch. The semaphore stays around the `execute_task` call: on push that call spans the whole turn so it paces dispatch as before, and under pull it returns in milliseconds so the worker pool becomes the cap — Phase 5's "capacity becomes physical", by construction. Deleting it, as first planned, would have fired N dispatches at an agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull. The outer deadline bounds the WAIT, not the work — a contract change. A still-open subtask now reports `running`, not `failed`; the batch still reports `deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and on push the old cancellation was half-illusory anyway (it abandoned the HTTP call while the agent kept running and billing the turn). After a deadline the status endpoint is the source of truth. A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a receipt to poll" — true and beside the point: it does not need a receipt, it needs to BLOCK CORRECTLY while the turn happens elsewhere, which `sync_waiter.wait_for_sync_terminal` already did for `/task`. `dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED return, wait for that row's terminal and rebuild from it. Nothing signals that waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail latency on an a2a call against a pilot, deliberately not worth a second signalling path. `PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an enumerated allow-list on purpose: a structural test forbids deriving it, because that would hand reach to the next autonomous trigger with nobody checking dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger` is kept and still tested, against a synthetic narrowing. The loop advance (#2523) and fan-out join share one `_terminal_side_effects` shim off `spawn_task_terminal_event`, with separate guards so one raising cannot cost the other its terminal. Migration 0051 + the SQLite twin: `fan_out_task_id`, plus `idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one batch on every fan-out terminal. Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here: `effect_guard` still fails open when the execution id is absent, and a fan-out multiplies that by N. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa
vybe
pushed a commit
that referenced
this pull request
Sep 15, 2026
…igger is stranded (#2524) (#2532) Completes #1081 Phase 4. After this, every autonomous trigger can run on the durable queue; only the interactive ones stay on push, which is the deliberate Open Question 7 scope cut (#1982/#1989). FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]` inside one `asyncio.gather`, so the batch existed only while the request that started it did. A pull-claimed subtask returns nothing to collect (the row is queued and the turn runs later in the agent's worker), and nothing could answer about a batch afterwards — no `async_mode`, no status endpoint, a disconnect lost it. The batch now lives on `schedule_executions`: every subtask row carries `fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used to be a dict key no async batch could reach), and `build_aggregate` rebuilds the result from the rows. Adds `async_mode` and `GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch belongs to that agent: `fan_out_id` is opaque but not secret. Two decisions the issue asked for. `max_concurrency` keeps its meaning and needed no branch. The semaphore stays around the `execute_task` call: on push that call spans the whole turn so it paces dispatch as before, and under pull it returns in milliseconds so the worker pool becomes the cap — Phase 5's "capacity becomes physical", by construction. Deleting it, as first planned, would have fired N dispatches at an agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull. The outer deadline bounds the WAIT, not the work — a contract change. A still-open subtask now reports `running`, not `failed`; the batch still reports `deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and on push the old cancellation was half-illusory anyway (it abandoned the HTTP call while the agent kept running and billing the turn). After a deadline the status endpoint is the source of truth. A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a receipt to poll" — true and beside the point: it does not need a receipt, it needs to BLOCK CORRECTLY while the turn happens elsewhere, which `sync_waiter.wait_for_sync_terminal` already did for `/task`. `dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED return, wait for that row's terminal and rebuild from it. Nothing signals that waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail latency on an a2a call against a pilot, deliberately not worth a second signalling path. `PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an enumerated allow-list on purpose: a structural test forbids deriving it, because that would hand reach to the next autonomous trigger with nobody checking dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger` is kept and still tested, against a synthetic narrowing. The loop advance (#2523) and fan-out join share one `_terminal_side_effects` shim off `spawn_task_terminal_event`, with separate guards so one raising cannot cost the other its terminal. Migration 0051 + the SQLite twin: `fan_out_task_id`, plus `idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one batch on every fan-out terminal. Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here: `effect_guard` still fails open when the execution id is absent, and a fan-out multiplies that by N. Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Sep 16, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Prep for the #1766 production soak. Two defects made a flag-flip soak unable to produce the evidence its acceptance criteria ask for. Part of #1766 — deliberately does not close it; the soak itself is the ticket.
1. The pilot flag was purely additive — push and pull ran in parallel inside one agent
CapacityManager.acquirehad no pilot branch, so a free slot still meant a push and a row only reached the queue on capacity overflow. When one did queue, the backend's ownbacklog_service.drain_nextraced the agent's worker through the samedb.claim_next_queued— no double-run (the claim is one atomic UPDATE), but the winner was whoever polled first, and the backend was structurally favoured: it drains on slot release plus a 60s orphan sweep, while a worker idles up to 15s between polls.Two consequences:
max_parallel_tasks) meant a pilot could run up to 2x its configured concurrency — invisible to canary S-02 (countsZCARDonly) and S-01 (excludes leased rows by design).pull_owns_dispatchmakes a pilot's autonomous work queue-only: never admitted, never drained by the backend, so its worker pool is its capacity. This is the pilot-scoped slice of #1081 Phase 5 ("capacity becomes physical").Interactive triggers are deliberately excluded and keep the synchronous push path — the scope cut in
TARGET_ARCHITECTURE.mdOpen Question 7, load-bearing because one FIFO ordered byqueued_atwould park a human turn behind N batch tasks until the held connection timed out, and N competing workers could claim two turns of the same session concurrently (the concurrent--resumeon one JSONL the Redis session lock exists to prevent).Fails safe to push. Empty allowlist — the default — leaves every path byte-for-byte unchanged.
2. E-05 fired on every pull turn past 60s
A pull-claimed row is
runningwith a NULLclaude_session_idby design, andmark_no_session_executions_failed— the very sweep E-05 watches — already carrieslease_expires_at IS NULLfor exactly that reason. E-05 was flagging the rows that sweep deliberately leaves alone.Left unfixed it would have broken the "canary green throughout" AC and, worse, trained everyone to ignore canary alerts during precisely the window where a real S-01 or L-03 matters. Mirrors S-01's exclusion; a NULL-lease control still fires. Closes the T3.6 gap in
docs/testing/PULL_MIGRATION_TESTING.md§8.3. Soak measurement set (docs §9)
Eight queries mapped to #1766's ACs, plus a pre-flight baseline step, abort criteria and rollback. Motivated by a gap worth stating plainly:
instance-monitorhas zero references to leases,claim_token,redelivery_countor queue depth — it reports generic health and stays green whether pull carries the fleet or does nothing at all, so AC "instance monitor active" is not sufficient evidence on its own.M1 is load-bearing: if
pulledis ~0 the soak measured nothing and every downstream conclusion is void.Testing
tests/unit/test_1766_pull_pilot_exclusive.py— 23 tests: the predicate (every autonomous trigger, every interactive trigger, empty allowlist, unresolvable-trigger-set fail-safe), the producer gate (pilot bypasses admission / interactive still pushes / non-pilot unchanged /queue_in_memoryuntouched), and the consumer gate (drain short-circuits ahead of the queued COUNT).tests/test_canary_invariants.py.Deployment note
PULL_MODE_PILOT_AGENTSis process env (backend restart) and the pilot agent must be recreated — there is no config-drift predicate for pull mode, soTRINITY_PULL_MODEonly bakes in at create/recreate. Restart alone leaves the agent not pulling.🤖 Generated with Claude Code