Skip to content

fix(pull): make the pilot flag a true either/or + soak measurement set (#1766) - #1982

Merged
vybe merged 3 commits into
devfrom
feature/1766-pull-pilot-exclusive
Aug 4, 2026
Merged

vybe merged 3 commits into
devfrom
feature/1766-pull-pilot-exclusive

Conversation

@obasilakis

Copy link
Copy Markdown
Contributor

Prep for the #1766 production soak. Two defects made a flag-flip soak unable to produce the evidence its acceptance criteria ask for. Part of #1766 — deliberately does not close it; the soak itself is the ticket.

1. The pilot flag was purely additive — push and pull ran in parallel inside one agent

CapacityManager.acquire had no pilot branch, so a free slot still meant a push and a row only reached the queue on capacity overflow. When one did queue, the backend's own backlog_service.drain_next raced the agent's worker through the same db.claim_next_queued — no double-run (the claim is one atomic UPDATE), but the winner was whoever polled first, and the backend was structurally favoured: it drains on slot release plus a 60s orphan sweep, while a worker idles up to 15s between polls.

Two consequences:

  • Pull coverage was near-zero and drain-biased on a healthy instance — the soak would have come back green having exercised almost nothing.
  • Two independent capacity counters (Redis ZSET vs the container's worker pool, both sized max_parallel_tasks) meant a pilot could run up to 2x its configured concurrency — invisible to canary S-02 (counts ZCARD only) and S-01 (excludes leased rows by design).

pull_owns_dispatch makes a pilot's autonomous work queue-only: never admitted, never drained by the backend, so its worker pool is its capacity. This is the pilot-scoped slice of #1081 Phase 5 ("capacity becomes physical").

Interactive triggers are deliberately excluded and keep the synchronous push path — the scope cut in TARGET_ARCHITECTURE.md Open Question 7, load-bearing because one FIFO ordered by queued_at would park a human turn behind N batch tasks until the held connection timed out, and N competing workers could claim two turns of the same session concurrently (the concurrent --resume on one JSONL the Redis session lock exists to prevent).

Fails safe to push. Empty allowlist — the default — leaves every path byte-for-byte unchanged.

2. E-05 fired on every pull turn past 60s

A pull-claimed row is running with a NULL claude_session_id by design, and mark_no_session_executions_failed — the very sweep E-05 watches — already carries lease_expires_at IS NULL for exactly that reason. E-05 was flagging the rows that sweep deliberately leaves alone.

Left unfixed it would have broken the "canary green throughout" AC and, worse, trained everyone to ignore canary alerts during precisely the window where a real S-01 or L-03 matters. Mirrors S-01's exclusion; a NULL-lease control still fires. Closes the T3.6 gap in docs/testing/PULL_MIGRATION_TESTING.md §8.

3. Soak measurement set (docs §9)

Eight queries mapped to #1766's ACs, plus a pre-flight baseline step, abort criteria and rollback. Motivated by a gap worth stating plainly: instance-monitor has zero references to leases, claim_token, redelivery_count or queue depth — it reports generic health and stays green whether pull carries the fleet or does nothing at all, so AC "instance monitor active" is not sufficient evidence on its own.

M1 is load-bearing: if pulled is ~0 the soak measured nothing and every downstream conclusion is void.

Testing

  • New tests/unit/test_1766_pull_pilot_exclusive.py — 23 tests: the predicate (every autonomous trigger, every interactive trigger, empty allowlist, unresolvable-trigger-set fail-safe), the producer gate (pilot bypasses admission / interactive still pushes / non-pilot unchanged / queue_in_memory untouched), and the consumer gate (drain short-circuits ahead of the queued COUNT).
  • E-05 lease-exclusion + non-leased control added to tests/test_canary_invariants.py.
  • Full unit suite: 7199 passed, 18 skipped, 0 failures.

Deployment note

PULL_MODE_PILOT_AGENTS is process env (backend restart) and the pilot agent must be recreated — there is no config-drift predicate for pull mode, so TRINITY_PULL_MODE only bakes in at create/recreate. Restart alone leaves the agent not pulling.

🤖 Generated with Claude Code

…ased rows (#1766)

Prep for the #1766 production soak. Two defects made a flag-flip soak
unable to produce the evidence its ACs ask for.

1. The pilot flag was purely ADDITIVE, so push and pull ran in parallel
   INSIDE one agent. `CapacityManager.acquire` had no pilot branch: a free
   slot still meant a push, so a row only reached the queue on capacity
   overflow. When one did queue, the backend's own `backlog_service`
   `drain_next` raced the agent's worker through the SAME
   `db.claim_next_queued` — no double-run (the claim is one atomic UPDATE)
   but the winner was whoever polled first, and the backend was
   structurally favoured (it drains on slot release plus a 60s orphan
   sweep; a worker idles up to 15s). Pull coverage was therefore near-zero
   and drain-biased. Worse, the two paths were bounded by two independent
   counters — Redis ZSET vs the container's worker pool, both sized
   max_parallel_tasks — so a pilot could run up to 2x its configured
   concurrency, invisible to canary S-02 (ZCARD only) and S-01 (excludes
   leased rows by design).

   `pull_owns_dispatch` now makes a pilot's autonomous work queue-ONLY:
   never admitted, never drained by the backend, so its worker pool IS its
   capacity. This is the pilot-scoped slice of #1081 Phase 5 ("capacity
   becomes physical"). Interactive triggers are deliberately EXCLUDED and
   keep the synchronous push path — the scope cut in TARGET_ARCHITECTURE.md
   Open Question 7, load-bearing because one FIFO ordered by queued_at
   would park a human turn behind N batch tasks and let N workers claim two
   turns of one session concurrently. Fails safe to push; an empty
   allowlist (the default) leaves every path byte-for-byte unchanged.

2. E-05 fired on every pull turn older than 60s. A pull-claimed row is
   `running` with a NULL claude_session_id BY DESIGN, and
   `mark_no_session_executions_failed` — the very sweep E-05 watches —
   already carries `lease_expires_at IS NULL` for that reason. E-05 was
   flagging the rows that sweep deliberately leaves alone, which would have
   buried a real #106 regression under noise for the whole soak window and
   broken the "canary green throughout" AC. Mirrors S-01's exclusion; a
   NULL-lease control still fires. Closes the T3.6 gap in
   docs/testing/PULL_MIGRATION_TESTING.md §8.

Full unit suite green (7199 passed).
… in §8 (#1766)

Adds §9 — the queries that turn "we ran it for a few days" into a verdict,
each mapped to a #1766 acceptance criterion, plus a pre-flight baseline
step, abort criteria, and the rollback procedure.

Motivated by a gap worth stating plainly: `instance-monitor` has zero
references to leases, claim_token, redelivery_count or queue depth, so it
reports generic health and stays green whether pull carries the fleet or
does nothing at all. AC "instance monitor active" is therefore not
sufficient evidence on its own; M1/M3/M5 are the substitute (and are the
natural candidates to fold into its deep-probe later).

M1 is the load-bearing one: if `pulled` is ~0 the soak measured nothing
and every downstream conclusion is void — the exact failure mode the
pre-#1766 additive flag produced on a healthy instance.

PostgreSQL flavour with explicit ::timestamptz casts, since every
timestamp column is Text holding ISO-Z (Invariant #16).

§8 updated: E-05 lease-awareness struck (closed by this branch) with E-01
left open, and the push/pull coexistence defect recorded alongside it.
…t + import fix)

Two problems, one root cause.

`tests/lint_sys_modules.py` failed on a bare `sys.modules.pop(...)` in the new
test's bootstrap — copied from test_capacity_manager.py, but this file imports
every backend module lazily so it never needed the preamble. Removed rather than
papered over with a _STUBBED_MODULE_NAMES declaration: test_1081_physical_meter.py
warns that popping `utils` actively unbinds conftest's canonical importlib
registration, so the idiom is harmful here, not merely unnecessary.

Removing it exposed the real defect. `services/agent_service/__init__.py` eagerly
imports helpers/lifecycle/crud/deploy/terminal, so the guard's
`from services.agent_service.pull_mode import is_pull_pilot_agent` inside
`drain_next` dragged the whole agent-lifecycle stack (→ models) into
backlog_service. Under a test that stubs models/database that surfaced as
`ImportError: cannot import name 'AgentGitConfig' from '<unknown module name>'`
— naming nothing to do with the actual dependency. Verified against clean dev:
tests/unit/test_backlog.py passes 34 there and failed 7 on this branch, so this
was introduced here, not pre-existing.

The predicates now live in `services/pull_pilot.py` — stdlib-only, no package
side effects. `pull_mode` re-exports them so crud/lifecycle and the existing
#1081 tests are untouched; the three dispatch-path callers (capacity_manager,
backlog_service, routers/internal) import the leaf directly. This also retires
a latent instance of the same hazard: capacity_manager's pre-existing
shadow-meter imports had the identical shape and were simply never exercised by
a stubbed test.

A capacity/backlog module should not need the agent CRUD stack to answer "is
this name in an env var".

test_backlog isolation 34 passed; lint clean; combined pull/capacity/canary
suites 284 passed. NOTE: the full unit suite is order-sensitive on dev
independently of this branch — clean dev fails 2 in test_1081_physical_meter
under full-suite ordering while passing in isolation. Both files pass in
isolation and together here; deferring to CI's seeded base/head runs.

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validated via /validate-pr — clean pass: fail-safe direction is push, empty-allowlist inertness tested both sides, E-05 lease exclusion mirrors the sweep predicate with a control test. Merging in sequence. Follow-ups noted: stale §Capacity & Backlog paragraph in architecture.md (acquire/drain now have a pilot branch), and #1766 status bump is manual (no closing keyword by design).

@vybe
vybe merged commit c455aa5 into dev Aug 4, 2026
22 checks passed
vybe pushed a commit that referenced this pull request Aug 6, 2026
…g critical (#1990) (#2046)

* fix(canary): make E-01 lease-aware so a healthy pull turn stops paging critical (#1990)

E-01 ("no running row older than execution_timeout_seconds + 300s") was the last
invariant reading the running-row set with no lease-awareness, after S-01 shipped
with the pull phases and E-05 got it in #1982.

A #1081 Phase-3 pull-CLAIMED row (lease_expires_at IS NOT NULL) is status='running'
but owned exclusively by the lease-reaper, which re-queues it (redelivery_count++,
started_at reset) or poison-parks it terminal at MAX_REDELIVERY. Its age is not
evidence of the stuck-execution class E-01 detects: the deadline governing a leased
row is its lease, and a different component owns the recovery.

The overlap is exact, not merely awkward. claim_next_queued stamps
lease_expires_at = started_at + (execution_timeout_seconds + SLOT_TTL_BUFFER)
(pull_coordination_service.py) -- the same threshold E-01 builds. So age > threshold
became true at the precise instant the lease expired, which is the instant the
reaper's recovery window OPENS. The 300s buffer exists so E-01 fires only after the
owning component has had its window; against a leased row that head-room was zero,
and E-01 paged critical for the whole gap between lease expiry and the reaper's next
sweep -- on every re-delivery, per execution, while the machinery worked as designed.

The db layer already encoded this ownership split, including on the very sweep whose
failure E-01 exists to detect: mark_stale_executions_failed carries
lease_expires_at IS NULL, as do get_running_executions,
get_running_executions_with_agent_info, fail_stale_slot_execution and
mark_no_session_executions_failed. The canary was the layer that hadn't caught up.

This blocks the #1766 soak, whose AC is "canary green throughout" and whose abort
criterion is a critical canary violation on the pilot -- so a healthy soak would
abort itself, or worse, train the operator to ignore the one invariant that matters
most during it.

The exclusion is keyed on the lease, not a blanket silencing: a NULL-lease (push) row
of identical age still fires, the threshold boundary is unchanged, and an absent
snapshot key fails OPEN (an older image that never populated the column keeps being
checked). Closes T3.7 in docs/testing/PULL_MIGRATION_TESTING.md section 3.

E-02 assessed and deliberately NOT excluded
-------------------------------------------

E-01, E-05, S-01 and E-02 are the four readers of running_exec_ids; E-02 is the one
that must keep seeing leased rows. It asks a different question -- was this id
terminal in a previous cycle and is it non-terminal now -- whose answer does not
depend on who owns the row. The reaper cannot produce that transition
(requeue_expired_lease / park_expired_lease both CAS on status='running' with a past
lease, so a terminal row is unreachable to them), and re-delivery preserves the
execution_id by construction (#1084/#525 are execution_id-scoped), so a terminal id
reappearing as running/queued is exactly the corruption E-02 exists to catch. Pull is
MORE exposed to it than push -- a late worker result and a reaper pass race for the
same row -- so copying this exclusion into E-02 "for consistency" would blind it on
the very path #1081 introduces. Recorded in the E-01 docstring, snapshot.py, the
catalog and architecture.md, and guarded executably by a test.

Testing
-------

tests/unit/test_1990_e01_lease_awareness.py -- 13 tests. Under tests/unit/ because CI
runs `cd tests && pytest unit/` only; the pre-existing E-01 tests sit in
tests/test_canary_invariants.py, which no workflow executes (trinity#2037), so a guard
placed beside them would never go red. Verified red without the fix: 5 failed, 8
passed -- the 8 being the NULL-lease controls, which must pass either way.

Load-bearing: a NULL-lease control row of identical age still fires, asserted both
alone and alongside a leased row in one snapshot. Every other test in the module would
still pass if someone "fixed" E-01 by deleting it. Boundary asserted exactly
(threshold silent / threshold+1 fires) off integer offsets from one literal T0, per
the #1909 clock discipline. The E-02 verdict is a test, not a comment.

Two mirrored cases added to the in-place suite for locality beside the rest of E-01's
coverage, with a pointer to the enforcing module.

Full unit suite: 7814 passed, 18 skipped (7801 baseline + 13).

Also: the E-01 Slack runbook now states a firing row is always NULL-lease, so on-call
is not sent after the lease-reaper for a push-path stall.

Closes #1990

* docs(canary): five functions carry the six lease selectors, not five

* fix(canary): bound E-01's lease skip instead of silencing leased rows forever

Review finding (dolho, #2046): the unconditional `continue` on any non-NULL
lease is wider than the bug requires, and it locks in permanent silence for a
real failure — a stuck or dead lease-reaper.

`requeue_expired_lease` sets `started_at=now` and clears `lease_expires_at` in
one atomic UPDATE, and the reaper runs on `cleanup_service`'s cycle, so under a
healthy reaper a leased row's observable overdue-ness is bounded by one
interval. The false-positive window this PR targets is small and bounded; the
blanket skip was not.

Nothing else covered the failure case. E-01, E-05 and S-01 all exclude leased
rows; E-02 only catches terminal→non-terminal reversals; there is no
lease-overdue invariant; and /deep-probe v1.3 (trinity-ops-agent#297) ships
M1/M3/M5, not M4. Meanwhile PULL_MIGRATION_TESTING.md §9 M4 — "a running row
past its lease_expires_at that the reaper has not touched is a reaper failure"
— is a #1766 soak ABORT criterion. It would have had no automated owner.

So the skip is now bounded by LEASE_REAPER_GRACE_SECONDS (600s = 2 x
cleanup_service.CLEANUP_INTERVAL_SECONDS): one reaper cycle of worst case plus
a full spare cycle. A healthy reaper never fires; one that has merely missed a
cycle never fires; one that has STOPPED fires within grace + one canary
cadence. The constant carries its derivation and is asserted against the live
interval so the coupling cannot drift silently.

A firing leased row is a DIFFERENT diagnosis — the lease-reaper failed, not
cleanup_service's stale watchdog — so `lease_expires_at`,
`lease_overdue_seconds` and the grace ride in observed_state, the signal_query
names it, and the runbook hint now branches on `lease_expires_at` instead of
asserting "a firing row is always NULL-lease".

Two states that now fire deliberately, both on-call-actionable and documented:
the #1085 re-delivery governor hold (opt-in, 300s pause TTL, absorbed by the
spare cycle unless sustained) and reaper saturation (`find_expired_leases`
takes limit=500/pass).

Also from the review: `_parse_lease` collapses None/""/unparseable to "no
lease", so a future collector coercing NULL to "" widens nothing — a bare
`is not None` would have read that as a live lease and silenced the row.
obasilakis added a commit that referenced this pull request Sep 9, 2026
…igger is stranded (#2524)

Completes #1081 Phase 4. After this, every autonomous trigger can run on the
durable queue; only the interactive ones stay on push, which is the deliberate
Open Question 7 scope cut (#1982/#1989).

FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]`
inside one `asyncio.gather`, so the batch existed only while the request that
started it did. A pull-claimed subtask returns nothing to collect (the row is
queued and the turn runs later in the agent's worker), and nothing could answer
about a batch afterwards — no `async_mode`, no status endpoint, a disconnect
lost it. The batch now lives on `schedule_executions`: every subtask row carries
`fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used
to be a dict key no async batch could reach), and `build_aggregate` rebuilds the
result from the rows. Adds `async_mode` and
`GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch
belongs to that agent: `fan_out_id` is opaque but not secret.

Two decisions the issue asked for.

`max_concurrency` keeps its meaning and needed no branch. The semaphore stays
around the `execute_task` call: on push that call spans the whole turn so it
paces dispatch as before, and under pull it returns in milliseconds so the
worker pool becomes the cap — Phase 5's "capacity becomes physical", by
construction. Deleting it, as first planned, would have fired N dispatches at an
agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull.

The outer deadline bounds the WAIT, not the work — a contract change. A
still-open subtask now reports `running`, not `failed`; the batch still reports
`deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and
on push the old cancellation was half-illusory anyway (it abandoned the HTTP
call while the agent kept running and billing the turn). After a deadline the
status endpoint is the source of truth.

A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a
receipt to poll" — true and beside the point: it does not need a receipt, it
needs to BLOCK CORRECTLY while the turn happens elsewhere, which
`sync_waiter.wait_for_sync_terminal` already did for `/task`.
`dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED
return, wait for that row's terminal and rebuild from it. Nothing signals that
waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail
latency on an a2a call against a pilot, deliberately not worth a second
signalling path.

`PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an
enumerated allow-list on purpose: a structural test forbids deriving it, because
that would hand reach to the next autonomous trigger with nobody checking
dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger`
is kept and still tested, against a synthetic narrowing.

The loop advance (#2523) and fan-out join share one `_terminal_side_effects`
shim off `spawn_task_terminal_event`, with separate guards so one raising cannot
cost the other its terminal.

Migration 0051 + the SQLite twin: `fan_out_task_id`, plus
`idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one
batch on every fan-out terminal.

Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here:
`effect_guard` still fails open when the execution id is absent, and a fan-out
multiplies that by N.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa
obasilakis added a commit that referenced this pull request Sep 9, 2026
…igger is stranded (#2524)

Completes #1081 Phase 4. After this, every autonomous trigger can run on the
durable queue; only the interactive ones stay on push, which is the deliberate
Open Question 7 scope cut (#1982/#1989).

FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]`
inside one `asyncio.gather`, so the batch existed only while the request that
started it did. A pull-claimed subtask returns nothing to collect (the row is
queued and the turn runs later in the agent's worker), and nothing could answer
about a batch afterwards — no `async_mode`, no status endpoint, a disconnect
lost it. The batch now lives on `schedule_executions`: every subtask row carries
`fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used
to be a dict key no async batch could reach), and `build_aggregate` rebuilds the
result from the rows. Adds `async_mode` and
`GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch
belongs to that agent: `fan_out_id` is opaque but not secret.

Two decisions the issue asked for.

`max_concurrency` keeps its meaning and needed no branch. The semaphore stays
around the `execute_task` call: on push that call spans the whole turn so it
paces dispatch as before, and under pull it returns in milliseconds so the
worker pool becomes the cap — Phase 5's "capacity becomes physical", by
construction. Deleting it, as first planned, would have fired N dispatches at an
agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull.

The outer deadline bounds the WAIT, not the work — a contract change. A
still-open subtask now reports `running`, not `failed`; the batch still reports
`deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and
on push the old cancellation was half-illusory anyway (it abandoned the HTTP
call while the agent kept running and billing the turn). After a deadline the
status endpoint is the source of truth.

A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a
receipt to poll" — true and beside the point: it does not need a receipt, it
needs to BLOCK CORRECTLY while the turn happens elsewhere, which
`sync_waiter.wait_for_sync_terminal` already did for `/task`.
`dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED
return, wait for that row's terminal and rebuild from it. Nothing signals that
waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail
latency on an a2a call against a pilot, deliberately not worth a second
signalling path.

`PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an
enumerated allow-list on purpose: a structural test forbids deriving it, because
that would hand reach to the next autonomous trigger with nobody checking
dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger`
is kept and still tested, against a synthetic narrowing.

The loop advance (#2523) and fan-out join share one `_terminal_side_effects`
shim off `spawn_task_terminal_event`, with separate guards so one raising cannot
cost the other its terminal.

Migration 0051 + the SQLite twin: `fan_out_task_id`, plus
`idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one
batch on every fan-out terminal.

Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here:
`effect_guard` still fails open when the execution id is absent, and a fan-out
multiplies that by N.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa
vybe pushed a commit that referenced this pull request Sep 15, 2026
…igger is stranded (#2524) (#2532)

Completes #1081 Phase 4. After this, every autonomous trigger can run on the
durable queue; only the interactive ones stay on push, which is the deliberate
Open Question 7 scope cut (#1982/#1989).

FAN-OUT. `FanOutService.execute` built a `dict[task_id, FanOutTaskResult]`
inside one `asyncio.gather`, so the batch existed only while the request that
started it did. A pull-claimed subtask returns nothing to collect (the row is
queued and the turn runs later in the agent's worker), and nothing could answer
about a batch afterwards — no `async_mode`, no status endpoint, a disconnect
lost it. The batch now lives on `schedule_executions`: every subtask row carries
`fan_out_id` plus the caller's own `fan_out_task_id` (new column — the id used
to be a dict key no async batch could reach), and `build_aggregate` rebuilds the
result from the rows. Adds `async_mode` and
`GET /api/agents/{name}/fan-out/{fan_out_id}`, which also checks the batch
belongs to that agent: `fan_out_id` is opaque but not secret.

Two decisions the issue asked for.

`max_concurrency` keeps its meaning and needed no branch. The semaphore stays
around the `execute_task` call: on push that call spans the whole turn so it
paces dispatch as before, and under pull it returns in milliseconds so the
worker pool becomes the cap — Phase 5's "capacity becomes physical", by
construction. Deleting it, as first planned, would have fired N dispatches at an
agent whose `max_parallel_tasks` is 3 and turned the excess into CapacityFull.

The outer deadline bounds the WAIT, not the work — a contract change. A
still-open subtask now reports `running`, not `failed`; the batch still reports
`deadline_exceeded`. A queued or claimed row is not the backend's to cancel, and
on push the old cancellation was half-illusory anyway (it abandoned the HTTP
call while the agent kept running and billing the turn). After a deadline the
status endpoint is the source of truth.

A2A + OPERATOR_RESPONSE. These were deferred with "a2a cannot hand back a
receipt to poll" — true and beside the point: it does not need a receipt, it
needs to BLOCK CORRECTLY while the turn happens elsewhere, which
`sync_waiter.wait_for_sync_terminal` already did for `/task`.
`dispatch_and_await_terminal` is the adapter: `execute_task`, and on a QUEUED
return, wait for that row's terminal and rebuild from it. Nothing signals that
waiter on the pull path, so the wake is the 5s DB poll — up to ~5s of extra tail
latency on an a2a call against a pilot, deliberately not worth a second
signalling path.

`PULL_REACHABLE_TRIGGERS` now equals `_AUTONOMOUS_TRIGGERS` and stays an
enumerated allow-list on purpose: a structural test forbids deriving it, because
that would hand reach to the next autonomous trigger with nobody checking
dispatch can deliver it — precisely #2048's defect. `note_unreachable_pull_trigger`
is kept and still tested, against a synthetic narrowing.

The loop advance (#2523) and fan-out join share one `_terminal_side_effects`
shim off `spawn_task_terminal_event`, with separate guards so one raising cannot
cost the other its terminal.

Migration 0051 + the SQLite twin: `fan_out_task_id`, plus
`idx_executions_fan_out_status` — the join COUNTs non-terminal rows for one
batch on every fan-out terminal.

Refs #1081, #2048, #2391, #2523, ent#157, ent#329. #2392 gets worse from here:
`effect_guard` still fails open when the execution id is absent, and a fan-out
multiplies that by N.


Claude-Session: https://claude.ai/code/session_019Tm7UEkd4G9KQZD5oeLRSa

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants