Skip to content

fix(watchdog): stop false-orphaning executions parked before the agent spawns them (abilityai/trinity#2433) - #2435

Merged
vybe merged 3 commits into
devfrom
feature/2433-watchdog-parked-executions
Aug 31, 2026
Merged

vybe merged 3 commits into
devfrom
feature/2433-watchdog-parked-executions

Conversation

@webmixgamer

Copy link
Copy Markdown
Contributor

Summary

  • The cleanup watchdog false-failed executions that were admitted but parked — in the backend's global agent-call queue (agent_call_limiter), in the agent's CPU-sized default thread pool, behind the agent /api/chat lock, or in the post-exit drain before unregister() — because its proof-of-life (GET agent/api/executions/running) could not see any of those. It wrote failed ("completed on agent but status not reported"), released the slot, and the parked call then ran anyway: billed, overbooked, its late 200 silently overwriting the row (bug: Execution shows Failed then Success after cleanup (race condition) #378).
  • Orphan now = the agent does not know it AND no live backend dispatcher owns it. Agent side: /api/executions/running gains pending_ids (accepted at /api/task, /api/chat and the refactor: fire-and-forget dispatch — a hung turn holds zero backend resource #1083 async spawn but not yet spawned) and recently_completed_ids covers exited-but-registered handles. Backend side: every outbound agent call is registered for its whole lifetime (track_inflight_dispatch, in-process registry + a cross-worker Redis liveness marker execution:inflight:{id}, 60s TTL / 15s refresh, ONE refresher task per process); the watchdog reads a tri-state verdict (alive / absent / unknown) and withholds recovery accordingly (CleanupReport.dispatch_inflight_skipped).
  • A park no longer spends the run's budget: at grant, a park ≥ 5s re-stamps started_at (admission kept in queued_at, the drained-backlog shape) and renews the slot lease, so the registry-blind Phase-1 sweep, the slot TTL, canary E-01 and duration_ms all measure the run, not the wait.
  • Parked rows are cancellable — and agent-scoped: terminate consults the in-process registry, then the cross-worker marker; a parked phase is finalized CANCELLED and the grant raises BackendAgentCallCancelled, where the dispatcher writes CANCELLED itself (never FAILED). Cancel-while-pending on the agent is consumed by register() (SIGKILL at spawn), closing the check-then-Popen window.
  • Gemini runtime now registers its subprocess (it never did — every Gemini run longer than the grace was false-orphaned); headless runs use a dedicated 32-thread pool pinned to MAX_PARALLEL_TASKS_CEILING_MAX.
  • Packaging: BACKEND_AGENT_CALL_LIMIT / BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S forwarded in prod + hosted compose and documented in .env.example (they lived only in docker-compose.yml, the #1039 class); the >5s queue-wait warning fires on both acquire branches.
  • Security (from /cso --diff, fixed here): terminate_execution wrote CANCELLED keyed on the caller-supplied task_execution_id while only execution_id was scoped by the agent's 404 — a caller authorised on agent A could flip agent B's running row. One agent-scope gate at the function entry now covers all three arms (uniform 404, fail-closed 503). Report: docs/security-reports/cso-diff-2026-08-28-2433-watchdog-parked-executions.md.

Changes

  • docker/base-image/agent_server/: services/process_registry.py (pending registry, cancel-consumed-at-register, exited-but-registered ids), routers/chat.py (/api/task, /api/chat, /api/executions/running), services/result_callback.py, services/headless_executor.py (_HEADLESS_EXECUTOR, pre-spawn 409), services/gemini_runtime.py (registers)
  • src/backend/services/: agent_call_limiter.py (in-flight registry + marker refresher + cancel + on_granted), task_execution_service.py (whole-call tracking, re-stamp + slot renew at grant, CANCELLED terminal), chat_execution_service.py (entry agent-scope gate, parked-cancel arm, 409/503 split), cleanup_service.py (tri-state skip, honest orphan string, dispatch_inflight_skipped), slot_service.py (renew_slot), agent_runtime_state.py (exempt-by-construction note)
  • src/backend/db/schedules/executions.py + database.py: restamp_execution_dispatch (CAS on RUNNING + NULL lease)
  • docker-compose.prod.yml, docker-compose.hosted.yml, .env.example
  • Docs: architecture.md (new "In-Flight Dispatch Proof-of-Life (bug: watchdog fails admitted-but-undispatched executions as "completed on agent but status not reported", releasing their slots and masking the real terminal #2433)" block + Cleanup/Redis rows), requirements/infrastructure.md, feature flows (cleanup-service, task-execution-service, capacity-management, parallel-headless-execution, chat-turn-cancellation), learnings.md (2 entries), security report
  • Tests: 10 new tests/unit/test_2433_*.py files (registry, task handler, headless pool, Gemini, limiter, slot renew, DB restamp, dispatch wiring incl. the exploit replay, watchdog verdicts, packaging parity); 6 existing files adjusted (mock hygiene, _EXPECTED_UPDATE_SITES)

Test Plan

  • Full unit suite under CI conditions (branch on a clean origin/dev worktree, no submodules): 12969 passed, 0 failed (baseline origin/dev: 12863 passed, 0 failed)
  • python tests/lint_sys_modules.py — no new violations; secret scan on the diff clean
  • Repro A (backend-queue parking: 10 async tasks over 3 agents, global cap 8): 10/10 success — 2 parked 485s, withheld by the watchdog at both 5-min cycles, re-anchored at dispatch (started_at restamped, slot lease renewed), duration_ms ≈ the 486s run; 0 orphan recoveries, 0 lost CAS, 0 active slots after
  • Repro B (8 tasks on one agent, 2-CPU pin to recreate the old 6-thread default-pool bottleneck): 8/8 success — 5 parked (485s and 972s waves), every cycle withheld, every park re-anchored; 0 orphan recoveries, 0 lost CAS
  • Live pending_ids probe: two concurrent /api/chat turns on one agent — the second reported as pending while waiting on the chat lock, then running, with the first in recently_completed_ids
  • Rebuilt base image required for the agent-side half; the backend half alone already covers old images through the whole-call marker

Fixes #2433

Generated with Claude Code

…t spawns them

The cleanup watchdog's proof-of-life (GET agent/api/executions/running:
running ∪ recently-completed) could not see an admitted execution that was
waiting in the backend's global agent-call queue, in the agent's CPU-sized
default thread pool, behind the agent chat lock, or in the post-exit drain
before unregister(). After the 60s grace it wrote a false `failed`
("completed on agent but status not reported"), released the slot, and the
parked call then ran anyway — billed, overbooked, its late 200 silently
overwriting the row (#378). Reproduced twice locally; three mechanisms, one
string.

Orphan now means: the agent does not know the execution AND no live backend
dispatcher owns it.

- agent server: /api/executions/running gains `pending_ids` (accepted at
  /api/task, /api/chat and the #1083 async spawn but not yet spawned; lazily
  expired) and `recently_completed_ids` covers exited-but-registered handles.
  Cancel-while-pending is consumed by register() (SIGKILL at spawn, #679 marker
  kept); the pre-spawn 409 is only an optimisation. Headless runs use a
  dedicated 32-thread pool pinned to MAX_PARALLEL_TASKS_CEILING_MAX; the Gemini
  runtime now registers its subprocess at both Popen sites (it never did).
- backend: every outbound agent call is registered for its whole lifetime
  (track_inflight_dispatch — queue wait, connect retries, POST) in an
  in-process registry plus a cross-worker Redis liveness marker
  execution:inflight:{id} (60s TTL, one refresher task per process, 15s tick).
  The watchdog reads a tri-state verdict (alive / absent / unknown) and
  withholds recovery on `alive`, and on `unknown` only while a dispatcher could
  still own the row; a process with no Redis reads `absent` (its own registry
  is the whole truth). CleanupReport.dispatch_inflight_skipped counts withheld
  rows; the orphan error string states what was observed.
- a park no longer spends the run's budget: at grant, a park ≥ 5s restamps
  started_at (admission kept in queued_at, the drained-backlog shape, CAS on
  RUNNING + NULL lease) and renews the slot lease (ZADD XX + EXPIRE together);
  the refresher renews the slot every tick while parked.
- parked rows are cancellable and agent-scoped: terminate consults the
  in-process registry, then the cross-worker cancel key; a parked phase is
  finalized CANCELLED and the grant raises BackendAgentCallCancelled, where the
  dispatcher writes CANCELLED itself (never FAILED; the /chat arm answers 409).
- terminate_execution gains ONE agent-scope gate at its entry for all three
  arms: the row behind the caller-supplied task_execution_id must belong to the
  agent the route proved (uniform 404; an unreadable row fails closed with
  503). The proxy arm's 404 scoped only execution_id while the CANCELLED CAS
  was keyed on task_execution_id, so a caller authorised on agent A could flip
  agent B's running row (found by the /cso --diff verifier; report under
  docs/security-reports/).
- packaging: BACKEND_AGENT_CALL_LIMIT / BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S
  forwarded in prod + hosted compose and documented in .env.example; the >5s
  queue-wait warning fires on both acquire branches.

Verified: full unit suite under CI conditions 12969 passed / 0 failed
(baseline origin/dev 12863 / 0); Repro A 10/10 success (2 parked 485s,
withheld at both watchdog cycles, re-anchored at dispatch); Repro B 8/8
success (5 parked, two waves); live pending_ids probe on the agent. The
agent-side half needs a rebuilt base image; the backend half alone covers old
images through the whole-call marker.

Fixes #2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@dolho dolho left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review — #2435 (fix #2433)

Read the whole diff, built a worktree off cfc2cfef, and ran the suites locally.

Verification I did

  • pytest tests/unit/test_2433_*.py → 106 passed
  • pytest tests/unit -k "cleanup or watchdog or slot or limiter or terminate or 679 or 921 or 1332 or 1094 or 2127 or schedule_status or 226 or 1804" → 534 passed, 2 skipped
  • Verified the duration_ms claim end-to-end: it is computed in update_execution_status from the DB started_at (db/schedules/executions.py:439), not from the in-coroutine start_time (task_execution_service.py:1164, which is stamped before capacity acquire and still spans the park). So the restamp genuinely fixes duration_ms — the PR body is right, and the reason is worth keeping in the docstring because the local execution_time_ms still measures the park.
  • Redis ACL: both ~*, so the new execution:* keyspace is not blocked. agent_runtime_state exemption note is correct — the parity test only greps agent:*.

Overall: the diagnosis is right and the two-sided design (pending_ids on the agent + a live-dispatcher marker on the backend) is the correct shape — proof-of-life had exactly one half. The tri-state alive/absent/unknown and the "a read that could not be asked ≠ a read that said no" rule are applied consistently, and the terminate_execution agent-scope gate is a real fix on a real hole. Three things below I'd want addressed before merge; the first is a new race, the second turns a pre-existing leak into a permanent one.


1. Cross-worker cancel reads a marker phase that is up to 15s stale → cancel-without-terminate on a running turn

acquire_agent_call_slot flips entry.phase = "calling" in memory at grant, but nothing writes the marker at that transition — the payload is only rewritten by _refresher_loop on its next INFLIGHT_TICK_SECONDS (15s) tick. So for up to 15s after a park→dispatch, execution:inflight:{id} still says "phase": "parked".

_cancel_inflight_if_parked trusts that value:

phase = agent_call_limiter.cancel_inflight(eid, agent_name=name)     # in-process: accurate
if phase is None:
    phase = await agent_call_limiter.request_cross_worker_cancel(...)  # marker: up to 15s stale
if phase != "parked":
    return None

With --workers 2 (prod), roughly half of cancels land on the worker that does not own the coroutine, so they take the marker path. If the cancel arrives in that 15s window:

  • terminate_execution returns cancelled_while_parked and never reaches the proxy arm — the agent is never asked to terminate;
  • the row is written CANCELLED and the capacity slot is released via release_if_matches;
  • the other worker is already inside client.post(...) — its grant-time cancel check has already passed — so the agent runs the full turn, billed, with its slot given away (overbooking);
  • when the 200 comes back, the SUCCESS write loses the CAS to the standing CANCELLED, so the work is silently discarded.

That is the #378 symptom this PR exists to remove, reproduced in a narrower window. The in-process arm is fine — cancel_inflight reads the live entry.phase; it is only the marker arm that can lie.

Cheapest fix: make the parked→calling flip write the marker synchronously (one SET in a to_thread at grant, alongside the existing _cancel_requested_cross_worker_sync round-trip you already pay for on a ≥5s park), so the marker can only ever be stale in the safe direction. Belt-and-braces alternative: on the cross-worker path, set the cancel key and then fall through to the proxy arm regardless of reported phase — a genuinely parked execution 404s on the agent, so falling through costs nothing and the grant-time check still catches the true park.

2. list_recently_completed_ids now reports exited-but-registered ids with no age bound

The new union is right in principle — an exited process whose owner hasn't reached finally: unregister() is semantically "completed, not yet reported". But _recently_completed has RECENTLY_COMPLETED_TTL_SECONDS = 300, and this new set has no clock at all:

for eid, entry in self._processes.items():
    if entry["process"].poll() is not None:
        ids.add(eid)

PoC against this branch:

running ids: []
recently_completed (t=0): ['leaked-eid']
recently_completed after clearing the 300s TTL buffer: ['leaked-eid']

A registry entry that leaks — an exception between register() and the finally: unregister() — is now reported to the backend as agent-known forever, so the orphan watchdog never recovers that row. Before this change the same leak was self-healing: list_running() filters on poll() is None, so the row was still orphan-eligible. The registry-blind Phase-1 stale sweep is still the backstop, so this is not unbounded — but it means a leak now costs the full 120-min sweep with a fabricated duration_ms instead of one watchdog cycle.

And the PR makes that leak newly reachable rather than theoretical. In both gemini_runtime.execute() and execute_headless() the new register() sits before process.stdin.write(prompt), while the finally: unregister() only wraps the run_in_executor further down:

get_process_registry().register(_registered_id, process, metadata={...})

process.stdin.write(prompt)     # <-- exception here leaks the entry permanently
process.stdin.close()

The new kill-at-spawn path makes an exception there likely: register() SIGKILLs the group, control returns, and the very next statement writes to a pipe whose reader is gone → BrokenPipeError. claude_code.py has the same register-before-stdin-write shape (pre-existing, its try/finally starts after the write), so it inherits the same newly-permanent consequence.

Two small changes cover it: bound the exited-but-registered set by entry["started_at"] against RECENTLY_COMPLETED_TTL_SECONDS (the entry already carries the timestamp), and move Gemini's register() inside the block the finally: unregister() guards.

3. restamp_execution_dispatch runs sync sqlite on the event loop

In _on_dispatch_granted:

restamped = db.restamp_execution_dispatch(execution_id)          # sync DB, on the loop
...
renewed = await asyncio.to_thread(get_slot_service().renew_slot, ...)   # correctly off-loop

agent_call_limiter's own module docstring exists because sync sqlite inside an async coroutine stalls the loop, and this write happens while both semaphores are held, at the moment the queue is by definition congested. The line immediately below already does the right thing. Wrap it the same way.


Smaller things

  • /api/chat pending entries expire well before a chat can (chat.py) — execute_task passes timeout_seconds=request.timeout_seconds or 900, but the chat handler passes none, so the entry gets the 900s default → a 960s deadline. ChatRequest carries no timeout_seconds field at all (models.py:19-24), and a chat waiting on the execution lock behind a long turn can legitimately wait up to the agent's execution_timeout_seconds (7200s ceiling). The entry is then evicted mid-wait and pending_ids stops covering it. The backend in-flight marker covers this case today, so it is not a live regression — but the defence-in-depth layer silently stops defending, which is worth a comment at minimum and a plumbed timeout ideally.
  • /api/chat finally: discard_pending is inside the lock, register_pending is outside it — a request cancelled while waiting on get_execution_lock() (client disconnect) never discards. Backstopped by the deadline; still, moving register_pending under the same try would make the pairing structural.
  • _process_stale_slot_reclaims does one _inflight_verdict_map([execution_id]) per row inside the loop — i.e. one Redis MGET round-trip per candidate — while _reconcile_orphaned_executions deliberately batches into one. Under a backlog those are the same cycle. Worth batching for symmetry.
  • register_pending logs at INFO unconditionally, so every /api/task and /api/chat now emits an extra line into Vector. DEBUG seems right for the happy path; the eviction WARNING is the line that matters.
  • renew_slot moves the ZSET score before re-EXPIREing the metadata hash. If the hash has already expired, zadd XX still succeeds and the function returns True while expire is a no-op — leaving exactly the ZSET-without-hash state canary S-03 reports as missing. Not introduced here (the hash can expire first today), but renew_slot can now perpetuate it and reports success. A hset-if-missing, or returning False when the hash is gone, would keep the return value honest.
  • prod/hosted run maxmemory-policy allkeys-lru, so an in-flight marker can be evicted before its TTL under memory pressure → absent → false orphan. Same exposure the slot ZSETs already have, so not a blocker; worth one line in the architecture block so the next reader knows the marker is not eviction-proof.

Things I checked and am happy with

  • terminate_execution's entry gate is correctly placed above all three arms, fails closed on an unreadable row (503), uniform-404s a foreign row, and covers getattr(..., None) != name so a row object missing the attribute also refuses. All three routers (chat.py, public.py, client_portal/service.py) funnel through it, so the fix is complete, and test_proxy_arm_cannot_flip_a_foreign_row_via_task_execution_id replays the actual exploit.
  • restamp_execution_dispatch's func.coalesce(queued_at, started_at) in the SET clause reads the pre-update row on both dialects — correct, and the CAS on RUNNING + lease_expires_at IS NULL correctly leaves pull-mode rows to the lease reaper.
  • _inflight_verdict_map's "anything that is not a dict of known verdicts collapses to absent" guard, and the eager module-level import with the stub-leak rationale, are the right lesson applied — a leaked sys.modules MagicMock here would have silently disabled the whole watchdog.
  • BackendAgentCallCancelled as a subclass so every existing except BackendAgentCallBudgetExhausted keeps working, with the terminal branched to CANCELLED — good, and emit_task_terminal_event already maps every non-SUCCESS terminal to agent.task.failed with the precise status in the payload, so no event-vocabulary change is needed.
  • register_inflight doing _PENDING_DELETES.discard(execution_id) — that is what keeps the #678/SUB-003 same-execution-id retries from having their fresh marker deleted by the previous attempt's queued delete. Easy to miss, correctly handled.
  • _get_client(use_negative_cache=False) on the watchdog read only, with the reason stated — right call; the negative cache exists for the 15s refresher, not for a 5-min sweep.
  • Packaging parity (prod + hosted + .env.example) with a test pinning it — the #1039 class closed properly rather than only in the file that was noticed.

Nice work on the write-up and the repro evidence; the reasoning in the docstrings is genuinely the useful kind. The stale-marker race (1) is the one I'd insist on before merge.

…d-but-registered set (#2435 review)

Review of the #2433 fix found that it reintroduced the #378 symptom in a
narrower window and turned a pre-existing registry leak into a permanent one.

1. Cross-worker cancel acted on a marker phase that predated its own write.
   `entry.phase` flipped parked->calling in memory only; the marker was
   rewritten by the 15s refresher, so `execution:inflight:{id}` advertised
   `parked` for up to a full tick after the POST had begun. Under --workers 2
   about half of all cancels are served by the worker that does NOT own the
   coroutine and therefore read it: the row was finalized CANCELLED and its
   slot released while the agent ran the turn to a billed completion whose
   SUCCESS then lost the CAS. Closed by ordering, not by narrowing — the owner
   publishes the transition in the SAME round-trip that reads the cancel key
   (`_publish_calling_and_check_cancel_sync`), and the remote sets the cancel
   key BEFORE re-reading the phase (`_set_cancel_then_reread_phase_sync`), so
   an observed `parked` gives W_remote(cancel) < R_remote(marker) <
   W_owner(marker) < R_owner(cancel) and the grant is guaranteed to see the
   key. Neither side pays an extra round-trip. The owner gates the publish on
   the ENTRY's age rather than this attempt's park, because
   `track_inflight_dispatch` wraps the whole retry loop and a retry can grant
   instantly under a marker a tick left saying `parked`; the remote's scope
   check stays on its first read, so no key is written for a foreign agent.

2. `list_recently_completed_ids` reported exited-but-registered ids with no
   age bound, so a leaked entry was agent-known forever and the watchdog never
   recovered that row — a regression against pre-#2433, where `list_running()`
   self-healed it. Now bounded by the same 300s TTL as the buffer, measured
   from when the exit was first OBSERVED (not `started_at`, which would drop a
   long turn the moment it entered its drain). The leak is also closed at
   source: `register()` SIGKILLs the group for a cancel that arrived while
   pending, so the following `stdin.write` can raise BrokenPipeError — all
   three prompt-writing runtimes (claude_code, gemini x2) now pair that write
   with `unregister()` on failure.

3. `restamp_execution_dispatch` is a sync sqlite write and ran on the event
   loop, while both semaphores are held and the queue is by definition
   congested. Now `asyncio.to_thread`, like the slot renewal beside it.

Smaller items from the same review:
- /api/chat sizes its pending entry to PENDING_CHAT_TIMEOUT_SECONDS (7200s):
  `ChatRequest` carries no timeout and a chat can wait on the execution lock
  for the agent's whole budget, so the /api/task default evicted the entry
  mid-wait. Its discard now wraps the lock acquisition, so a request cancelled
  while waiting (client disconnect) cannot leak one.
- Phase 3 batches its in-flight verdict read (one MGET per cycle, not per row),
  matching Phase 0.
- `renew_slot` refuses, score untouched, when the metadata hash has already
  expired: `ZADD XX` succeeds while `EXPIRE` no-ops, so it used to report a
  renewal it had not performed and re-anchor exactly the ZSET-without-hash
  state canary S-03 calls `missing`.
- `register_pending` logs at DEBUG (it fires on every /api/task and /api/chat).
- Documented that the in-flight marker is not eviction-proof under the prod
  `allkeys-lru` policy.

Tests: tests/unit/test_2433_review_fixes.py (15) — 11 of them fail against
cfc2cfe, verified in a worktree. Full unit suite under CI conditions
(clean origin/dev worktree, no submodules): 12985 passed, 0 failed.

Refs #2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@webmixgamer

Copy link
Copy Markdown
Contributor Author

Thanks — the verification you did (worktree off cfc2cfef, the duration_ms trace to db/schedules/executions.py:439, the Redis ACL check) made every finding actionable without re-deriving it. All three blocking items and all six smaller ones are addressed in da1c0b07.

1. Stale marker phase → cancel-without-terminate — fixed, and closed rather than narrowed

Confirmed exactly as you traced it. I took the cheap fix you suggested but ordered both sides, because the synchronous write alone still leaves a window: if the remote reads the marker before the owner's write and then writes the cancel key after the owner's read, the owner dispatches and the remote finalizes CANCELLED anyway. It is microseconds instead of 15s, but it is the same bug.

The fix is store-then-load on both sides:

owner : W(marker=calling) -> R(cancel)    _publish_calling_and_check_cancel_sync
remote: W(cancel)         -> R(marker)    _set_cancel_then_reread_phase_sync

If the remote observes parked, then R_remote(marker) < W_owner(marker), and since W_remote(cancel) < R_remote(marker) and W_owner(marker) < R_owner(cancel), the chain gives W_remote(cancel) < R_owner(cancel) — the grant is guaranteed to see the key and raise before any POST. Both are single pipelines, so neither side pays a round-trip it wasn't already paying (the owner's replaces the bare cancel read).

Two details worth flagging:

  • The publish gates on the ENTRY's age, not this attempt's park. track_inflight_dispatch wraps the whole retry loop, so a Async chat_with_agent: long execution silently fails with null response (reader-thread) #678/SUB-003 retry can grant in ~0s under a marker a tick left saying parked — gating on waited_s would have left that path exposed. entry_age >= INFLIGHT_MARKER_GRACE_SECONDS is _tick's own filter, so it publishes exactly when a marker can exist, and the fast-acquire hot path still never touches Redis (pinned by a test).
  • The scope check stays on the remote's FIRST read. Setting the key before verifying the agent would have let a caller authorised on agent A poison agent B's cancel key — reintroducing the class the entry gate closed. So the remote reads (scope), writes, re-reads (phase).

I did not take the fall-through-to-proxy alternative: for a genuine park the agent 404s, so the user gets Execution not found in agent for a cancel that will in fact succeed at grant. The ordering fix keeps cancelled_while_parked honest.

2. Unbounded exited-but-registered set — fixed, plus the source of the leak

Bounded by the same RECENTLY_COMPLETED_TTL_SECONDS, but measured from when the exit was first observed, not started_at: a started_at bound would drop a legitimately long turn the moment it entered its drain, which is the hole the widening exists to close. list_recently_completed_ids is the only reader, so stamping there is sufficient, and a first observation that lags the true exit only extends the window conservatively.

You were right that this PR made the leak reachable rather than theoretical, so I also closed it at source: stdin.write is now paired with unregister() on failure in both Gemini sites and in claude_code.py (same shape, same newly-permanent consequence). There's an AST guard so a fourth runtime can't ship the unpaired shape. codex_runtime.py writes no prompt to stdin, so it isn't exposed.

3. Sync sqlite on the event loop — fixed

asyncio.to_thread, like the renew_slot line below it.

Smaller things

  • /api/chat pending window — plumbed rather than commented: PENDING_CHAT_TIMEOUT_SECONDS = 7200 (the agent-timeout ceiling), since ChatRequest genuinely has no timeout field to thread through.
  • register_pending / discard_pending pairing — made structural: the finally now wraps the lock acquisition, so a client disconnect during the wait can't leak an entry. Costs a re-indent of the handler body; git diff --ignore-all-space shows only the intended lines.
  • Phase 3 batching — one _inflight_verdict_map for every reclaimed id before the loop, matching Phase 0.
  • register_pending log level — DEBUG.
  • renew_slot — reads the metadata hash first and refuses (score untouched) when it is gone, so it neither reports a renewal it didn't perform nor re-anchors the S-03 missing state. The default-TTL fallback still applies to a hash that is present with an unreadable timeout_seconds.
  • allkeys-lru — documented in the architecture block: the marker is not eviction-proof, blast radius is one row, and the agent-side half is unaffected.

Verification

  • tests/unit/test_2433_review_fixes.py — 15 tests; 11 fail against cfc2cfef (checked in a worktree), so they're real regression tests rather than restatements. The 4 that pass there are controls: scope-before-write, hot path stays Redis-free, fail-soft at grant, and the long-turn-in-drain case that must keep passing.
  • Full unit suite under CI conditions (branch patch on a clean origin/dev worktree, no submodules): 12985 passed, 0 failed.
  • origin/dev has moved 3 commits (ent#451) since the merge-base; test-merged clean, only architecture.md auto-merges, no overlapping code files.

Docs updated in the same commit: architecture.md plus the five affected flow docs, and two learnings.md entries — one for the cached-state/remote-decision class, one for "widening a liveness set without a clock un-heals a leak".

@dolho dolho left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review — da1c0b07

Built a worktree off the new head, and a second one off cfc2cfef to check the claims against the code they replace rather than against their own description.

All three blockers are closed, and the reply's claims hold where I checked them.

Claim Verified
"11 of 15 fail against cfc2cfef" Ran the new file in a cfc2cfef worktree: 11 failed, 4 passed — exact. Real regression tests, not restatements
Ordering closes the race, not narrows it _marker_payload reads entry.phase, and the flip to "calling" precedes the pipeline, so the published value is genuinely calling
"gates on the ENTRY's age, _tick's own filter" _tick: (now - e.registered_at) >= INFLIGHT_MARKER_GRACE_SECONDS; grant: entry_age >= INFLIGHT_MARKER_GRACE_SECONDS. Same constant, and both on time.monotonic() — the equivalence is real, not approximate
Exit bound measured from first observation Stamped under self._lock, and register() builds a fresh dict so a re-registration cannot inherit a stale exited_seen_at
Codex not exposed stdin=subprocess.DEVNULL, no write
"--ignore-all-space shows only the intended lines" 156 changed lines → 13 insertions, 5 deletions. Exactly the re-indent claimed
restamp off the loop await asyncio.to_thread(db.restamp_execution_dispatch, …)

The store-then-load argument is right, and I like that the scope check stays on the first read — writing before verifying would have let a caller authorised on agent A poison agent B's cancel key, which is a fresh hole in the fix for a different one.

Two things left. Neither blocks; the first is a claim that is wider than the code.


1. The AST guard enumerates two files — it does not discover, so a fourth runtime can ship the unpaired shape

"There's an AST guard so a fourth runtime can't ship the unpaired shape."

_stdin_write_is_guarded is sound, but its caller is a hardcoded list:

@pytest.mark.parametrize("rel", [
    "docker/base-image/agent_server/services/gemini_runtime.py",
    "docker/base-image/agent_server/services/claude_code.py",
])

Proven — dropped a mistral_runtime.py into agent_server/services/ with register() followed by a bare process.stdin.write(prompt):

2 passed, 13 deselected

The guard did not notice. That is the class this repo keeps re-learning: #1677's caller-parity guard walks every call site for exactly this reason, and Invariant #5 records "a guard that walks only one of the two trees is not a guard". Here it walks two of N.

Fix is small — glob agent_server/services/*_runtime.py plus claude_code.py and assert over what it finds. One adjustment when you do: _stdin_write_is_guarded returns total > 0 and guarded == total, so a discovered file with no stdin write would read as a failure; that needs to become "vacuously true when total == 0".

The three real sites are correctly paired (except BaseException: unregister(); raise — and BaseException is the right width here).

2. The publish half fails soft in the unsafe direction, and the docstring only describes the other half

"Fails soft — a Redis error reads as 'no cancel', exactly as the bare read it replaces did."

True of the read. The write is new, and its failure has a consequence the bare read never had:

except Exception as e:
    _note_redis_failure(e)
    return False          # -> `cancelled` stays False -> the POST proceeds

On a transient pipeline error the marker is never republished, so it keeps saying parked — and if the remote worker's own pipeline succeeds where the owner's failed, the remote reads parked, finalizes CANCELLED, releases the slot, and the owner POSTs anyway. That is the original symptom, on the Redis-error path.

Narrow, and worth saying why: a hard unavailability is safe by construction — _get_client() is None in the owning process means its refresher never wrote a marker either, so the remote's first read returns None and routes through the agent. It needs a per-connection transient failure on one worker while the other is healthy.

Not asking for a mechanism — a retry cannot close it either. Asking for the docstring to say it, because the reply is otherwise scrupulous about exactly this shape ("true of the path it tested and false of the other one"), and a future reader will take "fails soft" as covering both halves.


Also verified

  • renew_slot reads the metadata hash first and refuses when it is gone — so it neither reports a renewal it did not perform nor re-anchors S-03's missing state, and it deliberately does not rebuild the hash (its timeout_seconds is unknowable from there). Correct on both counts.
  • PENDING_CHAT_TIMEOUT_SECONDS = 7200 is plumbed through to the register_pending call, not just defined.
  • The finally now wraps the lock acquisition, so a client disconnect during the wait cannot leak a pending entry — structural, as claimed.
  • da1c0b07 still test-merges clean against current dev (3 commits ahead; architecture.md the only overlap).
  • tests/unit/test_2433_*.py → 122 passed on this head; the wider blast radius (cleanup / watchdog / slot / limiter / terminate / 679 / 921 / 1332 / 1804 / process_registry / schedule_status) → 554 passed, 2 skipped.

Recommendation

Approve once #1 is addressed — it is a five-line change to the guard, and the guard is the only thing standing between this fix and its own recurrence. #2 is a docstring sentence. Everything substantive in the three blockers is genuinely closed, and closed with tests that fail against the commit they fix.

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validated: root-caused watchdog false-orphaning fix (#2433) with extensive per-surface tests (14 new test files incl. cross-worker cancel race and slot-renew coverage); pg-migrations green; full pytest matrix green. Merged origin/dev to resolve the learnings.md append conflict (kept both sides).

@vybe
vybe enabled auto-merge (squash) August 31, 2026 11:46
@vybe
vybe merged commit 743a28f into dev Aug 31, 2026
27 checks passed
webmixgamer added a commit that referenced this pull request Aug 31, 2026
…d two residuals

Follow-up to #2435, whose re-review landed after the PR had already
auto-merged. Three things, all of them corrections to claims that shipped
wider than the code.

1. The guard enumerated two files, so it did not guard. Proven by dropping a
   `mistral_runtime.py` into the tree with an unpaired `stdin.write`:
   "2 passed, 13 deselected". It now DISCOVERS — every
   `agent_server/**/*.py` that drives the process registry — and requires each
   `stdin.write` to be covered.

   A `*_runtime.py` glob would not have been enough either: it misses
   `claude_code.py` AND `headless_executor.py`, and the latter is a real
   fourth site the enumeration had never counted. `headless_executor` is safe
   by a DIFFERENT shape — `_run_headless_subprocess` registers and writes
   stdin with no local guard, but it is handed to `run_in_executor` inside
   `execute_headless_task`'s guarded try, so an exception reaches that
   `finally: unregister()`. The guard models both shapes, and a companion test
   pins WHICH shape covers each known site, because caller-pairing is
   name-based and a collision would otherwise mask a real offender.

   Also fixes the vacuity bug that blocked this move: the old helper returned
   `total > 0 and guarded == total`, so a discovered file with no stdin write
   (codex, `stdin=DEVNULL`) would have read as a FAILURE the moment
   enumeration stopped. It now returns an offender list, and a separate test
   pins the discovery floor so a broken walk fails loudly instead of passing
   on an empty set.

2. `_publish_calling_and_check_cancel_sync` said "fails soft", which is true
   of the read half it replaced and NOT of the write half it adds. If the
   publish pipeline raises transiently while the other worker's connection is
   healthy, the marker keeps saying `parked`, so the remote can still finalize
   CANCELLED under a live POST — the original #378 symptom on the Redis-error
   path. Documented rather than mechanised: bounded by the 30s negative cache
   plus one tick, a hard outage is safe by construction (a process whose
   client is None never wrote a marker, so the remote routes through the
   agent), and failing closed was rejected because every other Redis touch
   here is fail-open. architecture.md and chat-turn-cancellation.md had
   inherited the too-wide "closed" claim and are qualified to match.

3. `restamp_execution_dispatch` records WHY the re-stamp reaches `duration_ms`
   — it is computed from that DB column, not from the in-coroutine
   `start_time`, which is taken before the capacity acquire and is why the
   sibling `execution_time_ms` still spans the park and is not a bug.

No behaviour change outside the docstrings; the guard is the substantive part.

Tests: tests/unit/test_2433_review_fixes.py 15 -> 20, verified against the
PoC above by putting a real file in the tree (caught by file:line, then
removed). Full unit suite under CI conditions on this base: 13098 passed,
0 failed.

Fixes #2448
Refs #2433, #2435

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@webmixgamer

Copy link
Copy Markdown
Contributor Author

Heads-up: this merged (11:52 UTC) while I was still working through the re-review, so items 1 and 2 from your re-review did not land here. dev currently carries the enumerating guard — the one your mistral_runtime.py PoC walks straight past.

Both are in #2450 (issue #2448), plus the duration_ms docstring note from your first review's verification section.

Two things worth flagging from doing it, since one changes the shape of the fix you suggested:

A *_runtime.py glob would not have been enough. It misses claude_code.py and headless_executor.py — and headless_executor is a real fourth site the enumeration never counted. It is safe, but by a different shape: _run_headless_subprocess registers and writes stdin with no local guard, yet it is handed to run_in_executor inside execute_headless_task's try whose finally unregisters. So the guard walks every agent_server/**/*.py that drives the process registry and models both shapes (local, caller-paired), with a companion test pinning which shape covers each site — caller-pairing is name-based, so a name collision could otherwise mask a real offender.

Your vacuity note was load-bearing, exactly as you said: total > 0 and guarded == total would have failed codex_runtime.py the moment enumeration stopped. It returns an offender list now, and the discovery floor is pinned so a broken walk fails loudly rather than passing on an empty set.

On residual #2 — the docstring now says the fail-soft applies to the read half only, that the write half's failure leaves the marker saying parked (so a remote with a healthy connection can still finalize CANCELLED under a live POST), that it is bounded by the 30s negative cache plus one tick, that a hard outage is safe by construction, and why failing closed was rejected. Same qualification added to architecture.md and chat-turn-cancellation.md, which had both inherited the too-wide "closed" claim from my reply.

Thanks for catching the guard — "the guard is the only thing standing between this fix and its own recurrence" was the right call, and the fourth site is the proof.

webmixgamer added a commit that referenced this pull request Aug 31, 2026
…d two residuals

Follow-up to #2435, whose re-review landed after the PR had already
auto-merged. Three things, all of them corrections to claims that shipped
wider than the code.

1. The guard enumerated two files, so it did not guard. Proven by dropping a
   `mistral_runtime.py` into the tree with an unpaired `stdin.write`:
   "2 passed, 13 deselected". It now DISCOVERS — every
   `agent_server/**/*.py` that drives the process registry — and requires each
   `stdin.write` to be covered.

   A `*_runtime.py` glob would not have been enough either: it misses
   `claude_code.py` AND `headless_executor.py`, and the latter is a real
   fourth site the enumeration had never counted. `headless_executor` is safe
   by a DIFFERENT shape — `_run_headless_subprocess` registers and writes
   stdin with no local guard, but it is handed to `run_in_executor` inside
   `execute_headless_task`'s guarded try, so an exception reaches that
   `finally: unregister()`. The guard models both shapes, and a companion test
   pins WHICH shape covers each known site, because caller-pairing is
   name-based and a collision would otherwise mask a real offender.

   Also fixes the vacuity bug that blocked this move: the old helper returned
   `total > 0 and guarded == total`, so a discovered file with no stdin write
   (codex, `stdin=DEVNULL`) would have read as a FAILURE the moment
   enumeration stopped. It now returns an offender list, and a separate test
   pins the discovery floor so a broken walk fails loudly instead of passing
   on an empty set.

2. `_publish_calling_and_check_cancel_sync` said "fails soft", which is true
   of the read half it replaced and NOT of the write half it adds. If the
   publish pipeline raises transiently while the other worker's connection is
   healthy, the marker keeps saying `parked`, so the remote can still finalize
   CANCELLED under a live POST — the original #378 symptom on the Redis-error
   path. Documented rather than mechanised: bounded by the 30s negative cache
   plus one tick, a hard outage is safe by construction (a process whose
   client is None never wrote a marker, so the remote routes through the
   agent), and failing closed was rejected because every other Redis touch
   here is fail-open. architecture.md and chat-turn-cancellation.md had
   inherited the too-wide "closed" claim and are qualified to match.

3. `restamp_execution_dispatch` records WHY the re-stamp reaches `duration_ms`
   — it is computed from that DB column, not from the in-coroutine
   `start_time`, which is taken before the capacity acquire and is why the
   sibling `execution_time_ms` still spans the park and is not a bug.

No behaviour change outside the docstrings; the guard is the substantive part.

Tests: tests/unit/test_2433_review_fixes.py 15 -> 20, verified against the
PoC above by putting a real file in the tree (caught by file:line, then
removed). Full unit suite under CI conditions on this base: 13098 passed,
0 failed.

Fixes #2448
Refs #2433, #2435

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
webmixgamer added a commit that referenced this pull request Aug 31, 2026
…d two residuals

Follow-up to #2435, whose re-review landed after the PR had already
auto-merged. Three things, all of them corrections to claims that shipped
wider than the code.

1. The guard enumerated two files, so it did not guard. Proven by dropping a
   `mistral_runtime.py` into the tree with an unpaired `stdin.write`:
   "2 passed, 13 deselected". It now DISCOVERS — every
   `agent_server/**/*.py` that drives the process registry — and requires each
   `stdin.write` to be covered.

   A `*_runtime.py` glob would not have been enough either: it misses
   `claude_code.py` AND `headless_executor.py`, and the latter is a real
   fourth site the enumeration had never counted. `headless_executor` is safe
   by a DIFFERENT shape — `_run_headless_subprocess` registers and writes
   stdin with no local guard, but it is handed to `run_in_executor` inside
   `execute_headless_task`'s guarded try, so an exception reaches that
   `finally: unregister()`. The guard models both shapes, and a companion test
   pins WHICH shape covers each known site, because caller-pairing is
   name-based and a collision would otherwise mask a real offender.

   Also fixes the vacuity bug that blocked this move: the old helper returned
   `total > 0 and guarded == total`, so a discovered file with no stdin write
   (codex, `stdin=DEVNULL`) would have read as a FAILURE the moment
   enumeration stopped. It now returns an offender list, and a separate test
   pins the discovery floor so a broken walk fails loudly instead of passing
   on an empty set.

2. `_publish_calling_and_check_cancel_sync` said "fails soft", which is true
   of the read half it replaced and NOT of the write half it adds. If the
   publish pipeline raises transiently while the other worker's connection is
   healthy, the marker keeps saying `parked`, so the remote can still finalize
   CANCELLED under a live POST — the original #378 symptom on the Redis-error
   path. Documented rather than mechanised: bounded by the 30s negative cache
   plus one tick, a hard outage is safe by construction (a process whose
   client is None never wrote a marker, so the remote routes through the
   agent), and failing closed was rejected because every other Redis touch
   here is fail-open. architecture.md and chat-turn-cancellation.md had
   inherited the too-wide "closed" claim and are qualified to match.

3. `restamp_execution_dispatch` records WHY the re-stamp reaches `duration_ms`
   — it is computed from that DB column, not from the in-coroutine
   `start_time`, which is taken before the capacity acquire and is why the
   sibling `execution_time_ms` still spans the park and is not a bug.

No behaviour change outside the docstrings; the guard is the substantive part.

Tests: tests/unit/test_2433_review_fixes.py 15 -> 20, verified against the
PoC above by putting a real file in the tree (caught by file:line, then
removed). Full unit suite under CI conditions on this base: 13098 passed,
0 failed.

Fixes #2448
Refs #2433, #2435

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
webmixgamer added a commit that referenced this pull request Aug 31, 2026
…2450 review)

`_unguarded_stdin_writes` collected both `ast.Name` ids and `ast.Attribute`
attrs into the caller-pairing set, so ANY method call inside ANY guarded try
exempted every same-named function in the module. An ordinary dispatch shape

    try:
        return runtime.execute(prompt, process, execution_id)
    finally:
        get_process_registry().unregister(execution_id)

therefore exempted a `def execute(...)` that registered and then wrote stdin
unpaired — a real leak, invisible. `execute` / `run` / `send` are exactly the
names a dispatch try calls, so this was reachable rather than theoretical, and
it landed on the one case discovery exists for: the four known sites are
pinned by the mechanism test, but a NEW module has no such pin.

Verified both directions before and after: the planted shape is a false
negative with the attribute half and is caught without it, while the real tree
is unchanged — the single legitimate caller-paired site passes a bare name
(`run_in_executor(_HEADLESS_EXECUTOR, _run_headless_subprocess, ctx)`), so the
attribute half bought nothing. Removed from both copies of the logic
(`_unguarded_stdin_writes` and `_pairing_mechanisms`), with the boundary and
its remedy stated in the docstring: a future attribute-paired site is reported
rather than silently exempted, and the fix is to reference the function by
name or add a justified allowlist entry — never to re-add the attribute half.

Also records why scope keys on `register(` and not `register_pending(` — a
chosen boundary (no such module exists today), not an oversight.

Mutation battery, all as expected: real tree 21 passed; a new unpaired runtime,
gemini with its local pairing stripped, and claude_code with its local pairing
stripped each fail.

Refs #2448, #2435

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
webmixgamer added a commit that referenced this pull request Aug 31, 2026
…d two residuals (#2448) (#2450)

Follow-up to #2435, whose re-review landed after that PR had already auto-merged.

- The guard enumerated two files, so it did not guard. It now DISCOVERS every `agent_server/**/*.py` that drives the process registry, and models both accepted pairing shapes (local try/except, and caller-paired via a bare name reference). A `*_runtime.py` glob would not have sufficed either: it misses `claude_code.py` and `headless_executor.py`, the latter a live fourth site the enumeration never counted.
- The caller-pairing exemption no longer collects `ast.Attribute` attrs — that made any method call in any guarded try exempt every same-named function in the module, hiding a real leak behind an ordinary dispatch shape.
- `_publish_calling_and_check_cancel_sync` records that its fail-soft covers the read half only; the write half's failure leaves a stale `parked` marker, bounded by the 30s negative cache plus one tick. architecture.md and chat-turn-cancellation.md are qualified to match.
- `restamp_execution_dispatch` records why the re-stamp reaches `duration_ms` and why `execution_time_ms` still spans the park.

Source changes are docstrings only (verified by AST comparison with docstrings stripped). Reviewed and approved by dolho after a mutation battery.

Fixes #2448
vybe added a commit that referenced this pull request Sep 17, 2026
* fix(workspace): the ask badge said 2 and gave you no way to find them (#2424)

The sidebar advertised "2 asks are waiting on your answer" and then stranded
you: the agent that raised them carried no badge, its tooltip did not mention
them, and it could be collapsed out of the roster entirely. The only way to
locate a blocked agent was to open agents one at a time.

Observed on a 12-agent roster with two asks on ws-sage (11th of 12), so on a
fresh load the one row that mattered was behind the "show more" toggle.

Three failures, fixed together because separately each is a half-measure — a
badge with no destination, or a destination nobody can see.

1. The unit. `askCount` is `openAsks.length`, and the tooltip said "agents":
   two asks on ONE agent rendered as "2 agents are waiting on your answer". The
   number was right, the noun was wrong, and they only diverge when a single
   agent raises more than one ask — which is why it went unnoticed. Resolved
   toward ASKS rather than agents, because the row badges added here now answer
   "which agent", leaving the header to answer "how many decisions".

2. The row. `PortalSidebar.vue:139` renders a per-agent badge from
   `unreadByAgent` — unread REPLIES. Keeping asks out of that count is
   deliberate and documented at line 9 ("one is waiting on you to decide, the
   other on you to read"), and is preserved: the ask gets the *own badge* that
   comment promised, in `status-urgent` — the token the operator NavBar's
   pending-operator-queue badge already uses, so the two surfaces agree — and
   visually distinct from the indigo unread pill beside it. `agentRowTitle` had
   the same hole, so this is an accessibility fix too: a blocked agent's
   accessible name was the bare "Open ws-sage".

3. The collapse. #2159 capped the roster at five for a good reason (a long
   fleet pushed chats below the fold), but the slice is plain roster order with
   no ask weighting. Ask-bearing agents are now never hidden — appended, NOT
   floated to the top, because re-sorting on a transient count moves rows under
   the cursor between refreshes, the same reason the roster is not re-sorted by
   availability.

Not a regression: every piece shipped in its intended form; the gap was between
them.

Everything decidable moved into `portalUtils` (`asksByAgent`, `askBadgeTitle`,
`agentRowTitle`, `visibleAgentRows`, `AGENT_COLLAPSE_LIMIT`) because vitest runs
`environment: 'node'` with no mount harness — a rule inside the SFC is one no
test can reach, which is how all three of these shipped. Mutation-checked:
reverting the noun, dropping asks from the title, and restoring the plain slice
each turn the suite red.

`bg-amber-500` -> `bg-status-urgent-500` is required, not drive-by: new code must
be at zero raw palette classes, so the new badge needed a token, and the header
had to match it or the two ask indicators would differ. Amber maps to
`state-autonomous` (an operating mode), which is the wrong claim. PortalSidebar
is now at zero non-gray raw classes.

Two pre-existing guards asserted the moved expressions as source strings and are
rewritten to assert the properties behaviourally — strictly stronger, since they
now fail on a broken bound or a dropped chip title, not only on a reworded one:
- portalRosterRow #2159 "shows a fixed number by default"
- portalAvailabilityChip #2196 "row title carries the state"

Verification: 1518/1518 frontend unit tests, raw-color ratchet exit 0,
production build clean.

Closes #2424

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(workspace): the sync portal turn never carried its session, so report-back could not fire (#2426)

ent#457 gave the Workspace a report-back: an agent that delegates during a chat
turn gets the completion posted into that thread. It could not fire on the
SYNCHRONOUS path, because the parent execution never received the session
binding the report needs. `report_completion` gates on
`if not source_channel_chat_id`, and there the field was NULL.

Measured on a dev instance — 5 of 8 portal rows NULL, split exactly by path:

    07:55 -> 09:04   chat=7d27744d...   browser, streaming path
    09:06 -> 09:09   chat=NULL          POST .../chat, synchronous path

TWO CORRECT CHANGES THAT COLLIDE. ent#457 passes the binding down, and
`execute_task` persists it — but only inside `if not execution_id:`. ent#365's
`_precreate_sync_execution` has already created the row and handed the id over,
so that branch never runs, and the pre-create stamped only `source_channel`.
Its own docstring named the invariant it broke: "Mirrors `start_portal_turn`'s
creation exactly ... so the two paths produce indistinguishable rows and a
report published from either can be joined back to its chat."

The sibling comment in `start_portal_turn` says "both creation sites or the
stamp is a coin flip depending on which path made the row" — ent#457 covered the
two sites that existed when it was written; ent#365 had added a third.

Fix: stamp `source_channel_chat_id` + `source_channel_client` in the pre-create.
`session_id` is a REQUIRED parameter, not an optional one — the value is in
scope at the only call site, and a default would let a future caller silently
reintroduce the inert row. Rejected: teaching `execute_task` to UPDATE an
adopted row, which widens a hot path used by every trigger to repair one
caller's omission.

ALSO REPAIRS TWO GUARDS THAT WERE RED ON `dev`. `backend-unit-test` is failing
on dev right now; both failures are in this feature area and both are guards
that had gone inert, so they are fixed here rather than left for the next PR to
trip over. Frontend-only PRs pass because the `changes` job path-filters the
backend suite away, which is why this went unnoticed.

  * `test_both_portal_row_creation_sites_name_the_chat` asserted a literal
    census of `== 2` sites. It went red the moment the third site appeared —
    the guard WORKING — and the bug it names shipped anyway. Now asserts the
    rule instead of the count: every site that stamps the surface must also
    stamp the destination. Census-proof.

  * `test_portal_turn_kwargs_bind_against_execute_task` parsed `portal_chat`
    for a literal `run_resumable_turn(...)` call. That call had moved into
    `_run_sync_turn_and_clear_marker`, where it is `run_resumable_turn(**kwargs)`
    — a splat, which names nothing — so the walk found no keywords and the
    guard asserted itself dead. Now reads the keywords where they are actually
    named (the wrapper's call site), scanning both entry names and subtracting
    the wrapper's own consumed parameters.

Neither rewrite loses coverage; both now fail for the reason their docstring
gives rather than because a number or a call site moved.

WHY THE BUG SURVIVED ITS TESTS. ent#457's mock the engine and assert the kwargs
are passed (they are). ent#365's assert no orphan `running` row (still true).
Nothing asserted the PERSISTED ROW, which is the only place the two meet — the
same lesson `test_ent457_portal_turn_kwargs.py` states about itself. The new
suite asserts at that layer, and adds a derived parity check so a fourth
channel field added to one writer and forgotten in another fails here instead
of shipping as another silently-inert report path.

Verification: 402 passed on the portal/ent457/ent365 selection (was 2 failed
before this branch). Mutation-checked: removing the stamp turns 4 red; feeding
`execute_task` an unknown kwarg turns the repaired binding guard red.

Closes #2426

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(subscriptions): auto-switch ranks alternatives by cached headroom, never by load alone (#2409) (#2422)

## Summary
- `select_best_alternative_subscription` returned the **first** survivor of the 2h failure filter in `agent_count ASC` order and read no headroom — SUB-003 could move an agent onto a subscription at 99% of its weekly window, and an *unused dead-token* subscription (no agents ⇒ no failure rows) sorted **first**.
- Now: **filter in the db, rank in the service, never a probe.** The db lists survivors (kind-blind 2h filter unchanged and first, #444/#2352, `agent_count ASC, name ASC`); the service ranks them over the cached provider snapshot (one `MGET`) furthest-from-the-nearest-wall first (the fuller of the 5h/7d windows — the #792 retry lands on the destination immediately), in 10-point bands so load still spreads a storm; a **fresh** provider refusal is dropped; anything unusable sorts in today's order; any failure of the ranking half falls back to today's pick **with a warning**.
- `classify_headroom` (ent#434) and the ranker share one usability gate (`headroom_reading`) — verdicts byte-identical, pinned by a differential test against a frozen copy. New-agent auto-assign (#74) rides the same ranker. The switch now records **why** (`destination_headroom` + one notification clause).
- Approved deviations from the literal AC, recorded on the issue: nearest-wall key instead of 7d-only; fresh refusals filtered instead of ranked last.

## Changes
- `src/backend/services/subscription_headroom_service.py` — gate, MGET reader, threshold-free ranker, `MAX_READING_AGE_SECONDS` (owned here now); `classify_headroom` becomes policy over the gate
- `src/backend/services/subscription_auto_switch.py` — service-layer selector (`asyncio.to_thread` under the agent lock), `destination_headroom` on activity / notification / result
- `src/backend/services/subscription_service.py` — `select_subscription_for_new_agent`
- `src/backend/db/subscriptions.py` + `database.py` — `list_viable_alternative_subscriptions` / `list_assignable_subscriptions` (filter only); first-match selectors retired
- `src/backend/services/agent_service/crud.py` — call site; `subscription_headroom_alerts.py` — constant re-export + docstring
- Tests: new `tests/unit/test_2409_headroom_ranked_switch.py`; pingpong / 2352 / concurrency / 1484 / 1759 adapted to the list form (assertions kept)
- Docs: `architecture.md`, `subscription-auto-switch.md` (+ management, usage-tracking), requirements §20.4, `learnings.md` (2 entries), CSO diff report

## Test Plan
- [x] New suite: `pytest tests/unit/test_2409_headroom_ranked_switch.py` — 83 passed; **80/81 fail on the unmodified source**
- [x] Full `tests/unit`: 12,718 passed / 30 skipped / 1 pre-existing failure (`test_1920`, private submodule, untouched)
- [x] API integration (`test_subscription_auto_switch`, `test_subscriptions`, `test_subscription_usage`): 36 passed
- [x] Live: a real switch chose the 18%/9% subscription over the 0-agent 88%/60% one; every-survivor-refused → no switch + WARNING; no snapshot → today's order
- [x] `/review` clean (informational findings fixed in-review); `/cso --diff` no findings
- Follow-ups filed while testing: #2419 (parser overage), #2420 (destructive integration suite), #2421 (subscription audit gap)

Fixes #2409

🤖 Generated with [Claude Code](https://claude.com/claude-code)

* feat(workspace): an answer given in the Workspace resumes the agent (ent#430)

Slice 5 of ent#364, and the gate: until now the client route recorded an answer
and returned. The operator route called `spawn_resume_dispatch`; this one did
not. So an ask addressed to a Workspace client — the entire point of
ent#364/#428/#429 — was recorded, reached the agent's queue file in about three
seconds, and re-triggered nothing.

Measured on a live instance before this change: answered from the Workspace,
`operator-queue.json` flipped to `responded` with the answer in under 3s, and no
execution followed.

Unblocked because ent#329 is in dev.

WHAT THIS ADDS: one call. ent#430's body rules out the alternative — "a second
dispatch surface for the same event is how the cost, trigger-label and
loop-prevention questions get answered twice, differently" — so the per-agent
opt-in, the idempotency key, the audit row and the failure handling all stay
inside `maybe_dispatch_resume`. AC #2 and AC #3 are satisfied by REUSE rather
than by re-implementation, and the tests assert the CALL for that reason.

Four properties, each load-bearing:

* Hung off the CAS WIN only, like the operator route. The 409 above already
  returned for a lost race, so reaching the dispatch means this answer is the
  one that landed — two people answering at once produce one resume.
* `updated`, never `item`. The pre-answer read still says `pending`; a resume
  handed that row acts on an ask that does not yet carry its answer. Looks
  identical in a green test, which is why there is one for it.
* The spawn is wrapped. It is fire-and-forget, but a raise ON THE CALLING LINE
  would still propagate, and a 500 after the CAS landed would tell the client
  their answer failed while it is committed and already on its way to the agent.
  The answer is the thing that must not be lost.
* #2376's choice validator runs first, so an answer that was never offered
  cannot spend.

AC #5 — `resume_requested` on the answer response, read from the SAME accessor
the dispatch gates on, so the two cannot disagree about what is about to happen.
It reports INTENT, not success: the dispatch is backgrounded, so at that moment
the only honest claim is whether it will be attempted. Fails CLOSED — an
unreadable flag claims nothing, because over-claiming is exactly the failure
AC #5 names ("the ask does not read as resolved while nothing happened").

RESIDUAL, stated rather than implied: a dispatch that fails AFTER this point
surfaces as a FAILED execution row plus an `operator_resume_dispatch` audit
entry (ent#329) — operator-visible, and a client cannot see either. The client
half of AC #5 is satisfied negatively for now: the ask surface says nothing
about work starting, so it cannot mis-claim. `resume_requested` is the field a
surface needs to say something true; consuming it is an ent#429 UI change and is
deliberately not in this PR.

The per-agent flag DEFAULT IS UNCHANGED (`operator_resume_enabled`, OFF,
owner-only). "Turn the flag on" is an operator action per agent, not a code
default: flipping it would hand every shared agent's client a spend button,
which is the one thing AC #3 rules out.

Verification: 145 passed across the asks/ent#329/ent#364/#428/#429/#2376
selection. Mutation-checked — removing the dispatch (4 red), passing the
pre-answer row (1 red), and making the opt-in read fail open (1 red).

Closes ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(workspace): New chat means a new chat (ent#451 — the fresh-thread slice)

Reported: pressing New chat in the Workspace drops you back into the existing
conversation with that agent. Decided at the 2026-08-21 weekly.

ONE VALUE CARRYING TWO MEANINGS. An absent `session_id` meant both "I don't know
which thread" and "I want a fresh one", and the platform resolved it as the
first, in both readers:

    _resolve_session_id(..., None)  -> resume the client's latest
    get_history(..., None)          -> return the most-recent thread

Both readings are RIGHT for the case they were written for — a deep link, a
refresh, an API caller that never held a session id — so neither could be
inverted. The intent had to become sayable: `new_thread` on the request,
`newChat` on the component, checked before the resume.

The frontend tell was an asymmetry: New chat with the agent you were ALREADY on
started fresh, while New chat with a different agent resumed. The watcher read a
changed agent as "load that agent's history" and called `fetchHistory(name,
null)`, discarding the `pendingSession = null` that `newChatWithAgent` had just
set to mean the opposite.

MOST OF ent#451 TURNED OUT TO BE BUILT. Recorded because the issue is
complexity-high and this PR is not:

* the data model already allows many sessions per (agent, client) — no UNIQUE
  constraint, a `title` column, an index on
  `(agent_name, client_email, last_message_at)`, and auto-titling. AC #4's
  "migrates cleanly" is nothing to migrate.
* AC #2's list is the existing sidebar: titles, recency, starred lifted out,
  search, per-agent avatars.
* AC #3's landing rule is already decided and documented in
  `ensure_thread_for_ask` — reuse the latest thread so asks do not accumulate
  beside the conversation. UNCHANGED here, and pinned by a test so this cannot
  move it silently. It matters MORE once several chats exist, not less.

So what was missing is AC #1, and it is two bits rather than a data model.

Four properties:

* An explicit `session_id` WINS over the flag. A caller sending both contradicts
  itself; the id is a fact, the flag an intent, and abandoning a named thread
  would strand a turn meant for a conversation the caller could see.
* The ownership check runs first either way — the flag is never a route past it.
* BOTH turn entry points carry it. The Workspace uses the streaming path and
  falls back to the synchronous one, so a flag honoured by only one brings the
  bug back exactly when streaming fails.
* The intent is spent on adoption. The send guard already ANDs on "no session
  yet", so a second turn was never going to open a third thread; clearing it in
  `onSessionAdopted` keeps the two bits from disagreeing after a navigation.

Test doubles updated, not worked around: seven `_resolve_session_id` lambdas and
four `_fake_chat` stubs did not accept the new keyword. They take `**kw` now — a
stub that must be edited for every new parameter is a second signature — and one
hand-rolled `_Body` model double gained the field. All are stale stubs rather
than behaviour changes.

Verification: 392 passed across the portal/ent#286/#287/#358/#429/#430/#451
selection; 1497 frontend unit tests. Mutation-checked: making the flag inert, and
letting it override an explicit session id, each turn the suite red. The full
backend suite exceeds a local foreground run and is left to CI.

Pre-existing and NOT from this branch: `test_ent457_portal_turn_kwargs` and
`test_both_portal_row_creation_sites_name_the_chat` fail on `dev` today; both are
fixed in #2427.

Related to ent#451

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the ?new=1 deep link, the missing frontend test, and three latent desyncs (ent#451)

Blocker 1 was real and I had not seen it. `resolveAgentQuery` passed `forceNew`
to `resolveAgentLanding` and set `pendingSession = null`, but never raised
`startingNewChat` — so `/workspace?agent=X&new=1` rendered an empty conversation
and then sent `new_thread: false`, resuming the thread the user asked to leave.
The reported bug, intact on the documented `?new=1` contract, in the PR that
exists to fix it.

The cause is the one this PR is about, one level up: `route.query.new` was read
in two places for two different decisions — WHICH THREAD to land on and WHAT THE
FIRST SEND ASKS FOR — and only the first honoured it. Now read ONCE into a local
that feeds both, so they cannot drift again. AND-ed with the landing result, so
a `?new=1` that still resolved a thread never claims a fresh start.

Blocker 2: a frontend test, which the change genuinely had none of — the
`1497 passed` in the body was the pre-existing suite, as the review says.
`workspaceNewChat.spec.js` (9 tests) covers the deep link, the watcher branch
ORDER, the first-paint guard, both send conjunctions, and the settle-everywhere
rule, using the two established patterns (pure function + source assertion in
the `portalLeaveSpecificRoute.spec.js` shape) since vitest runs
`environment: 'node'` with no mount harness. Mutation-checked, and M1 is the
reviewer's own blocker: reverting it turns the suite red.

Blocker 3: `test_history_without_a_session_is_unchanged` cited "the spec in
tests/unit/... frontend suite" — a dangling reference asserting coverage that
did not exist. It now names the real file.

Comments addressed:

* Three more sites nulled `pendingSession` without settling the intent — the
  deep-link watcher (the commonest way in), `openRoom`, `openAgentPage`, plus
  the unreachable-agent branch. Latent because both consumers AND on "no session
  yet", but a flag that is only correct because of a second variable is one
  refactor from being wrong, and the declaration claims it is cleared the moment
  a real thread exists. Now true.
* `test_both_turn_entry_points_forward_it` was `getsource` + a substring, so a
  comment or a misspelled kwarg satisfied it. It now BINDS the keyword against
  each service signature and asserts the routes forward `body.new_thread`
  through a comment-stripped source — verified by mutation.
* `workspace-absorbs-session.md` updated at both seams the change touches
  (`resolveAgentLanding`'s landing rule and `_resolve_session_id`'s three
  states), and `architecture.md`'s Workspace section documents the new public
  `new_thread` field on the ent#83 headless surface.
* Gating stated rather than inferred: "OSS-core by decision (ent#451)", matching
  the ent#326/#384/#392 convention.

ONE CORRECTION, offered with evidence rather than silently applied. The review
says "`test_ent457_portal_turn_kwargs.py` doesn't exist on `dev`, #2427
introduces it". It does exist on `dev` — added by d6a4bc10 (ent#457) — and #2427
modifies it. `git cat-file -e origin/dev:tests/unit/test_ent457_portal_turn_kwargs.py`
succeeds, and `backend-unit-test` is failing on `dev` independently of any PR.
So the body's "fails on dev today" stands. Everything else in the review is
accepted as written.

Verification: frontend 1497 -> 1506 (+9). Backend 392 passed on the portal
selection, the same 2 pre-existing dev failures unchanged.

Related to ent#451

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(resume): the respond→resume dispatch never ran — bad import, masked by its own stub (ent#329)

Found by testing this PR's feature against a live local instance. ent#430 wires
a Workspace answer to `spawn_resume_dispatch`, so this PR is dead on arrival
without it — the client path would have hit the same wall the operator path has
been hitting since ent#329 merged.

THE BUG. `operator_resume_service.maybe_dispatch_resume` did:

    from services.task_execution_service import task_execution_service

That name has never existed on that module; it exports
`get_task_execution_service()`. The import sits on the FIRST line of the
function, above the try, so every dispatch raised ImportError before it even
read the opt-in.

WHY NOBODY NOTICED, twice over:

* the call is fire-and-forget, so the traceback surfaces only as asyncio's
  "Task exception was never retrieved" — nothing fails, nothing 500s, the
  answer is recorded and the config audit row is written. It looks like it
  worked.
* the ent#329 unit test stubbed `services.task_execution_service` with
  `SimpleNamespace(task_execution_service=recorder)` — MANUFACTURING the very
  symbol whose absence was the bug. 21 tests green, feature dead.

MEASURED on a live instance, opt-in ON:

  before: answer 200, audit row written, executions 0->0, log carries
          "cannot import name 'task_execution_service'"
  after : answer 200, executions 0->1, triggered_by=operator_response,
          audit `operator_resume_dispatch` with the execution id, 0 ImportErrors

(The dispatched run then failed on a missing AGENT_AUTH_SECRET — a limitation of
the test box, and correctly recorded as an honest FAILED row, which is ent#329's
"never silent" requirement doing its job.)

THE GUARD is the durable part, because the stub is the real lesson: a stub that
invents an API the real module lacks converts a production crash into a green
suite. `test_the_names_this_service_imports_actually_exist_on_the_real_modules`
parses the REAL module source with `ast` — never the stubbed `sys.modules`
entry, which is what made this invisible — and asserts every
`from services.X import Y` resolves. Mutation-checked: reverting the import
turns 11 tests red.

Related to ent#430, ent#329

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the dispatch could not run, the race loser spent, the flag over-claimed (ent#430)

All three blockers from the review, each verified rather than argued.

1. THE FEATURE WAS INERT. `client_portal/asks/router.py` declares `answer_ask`
as a plain `def`, so FastAPI runs it through `run_in_threadpool` — a worker
thread with no event loop — and `asyncio.create_task` raises
`RuntimeError: no running event loop` there. The `except` swallowed it, so every
client answer recorded the answer and dispatched nothing: byte-for-byte the
behaviour this PR exists to remove.

Fixed in `spawn_resume_dispatch` rather than by flipping the route to
`async def`, for the two reasons the review names: the route does blocking DB
I/O, so `async def` alone would move it onto the loop; and ent#430's stated
shape is ONE dispatch site, which moving the spawn back out to the caller would
undo. It now detects the absence of a loop and hops back via
`anyio.from_thread.run_sync` — Starlette's threadpool is anyio's, so the portal
is always there on this path. Any future sync caller inherits the fix.

A thread anyio does not own reaches neither branch. That is not a production
shape, but it must not become the silent no-op this change removes, so it raises
with the cause named instead.

2. THE RACE LOSER SPENT MONEY. `respond_to_operator_queue_item` returns None
only when the row is GONE; when the row exists and has left `pending` — the race
that actually happens — it returns a TRUTHY dict carrying `_status_conflict`,
having written nothing. `if not updated` fell straight through it. The loser
then dispatched a paid execution for an answer not in the database, and because
the idempotency key hashes the response text, the loser's differing text yields
a different digest: one queue item, two paid dispatches. `routers/operator_queue.py`
already pops that flag before its own spawn; this is that rule, not a new one.
Popped, not read, so the sentinel cannot serialize to the client.

3. `resume_requested` OVER-CLAIMED. It was computed after the swallowed spawn
from the opt-in flag alone, so a spawn that raised still answered `true` — the
exact failure AC #5 names, and given (1) that was EVERY production answer on an
opted-in agent. It now reports what was actually scheduled.

TESTS — the reason all three survived 24 green checks is that every existing test
replaced `spawn_resume_dispatch` with a synchronous lambda, stubbing out the one
call whose runtime context was the defect. `test_ent430_dispatch_actually_runs.py`
drives the REAL spawn from a REAL anyio worker thread (the production context,
not an approximation) and asserts the premise before the behaviour. The lost-race
test uses the truthy `_status_conflict` shape that actually occurs, not the
`None` shape that does not. Mutation-checked: reverting fix 1 turns 1 red, fix 2
turns 3 red, fix 3 turns 2 red.

Writing those tests also caught a stubbing bug of my own, worth recording because
it is the trap that hid the original: patching only `sys.modules` leaves
`from services import operator_resume_service` resolving the PACKAGE ATTRIBUTE,
so the real function ran anyway. Both paths are patched now.

Related to ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the answered ask said pending, and the docs described one caller (ent#430)

Non-blocking findings from review pass 2. The three blockers landed in d11956a8.

STATUS. `_project` mapped every row to pending/expired, so the response to a
just-recorded answer read `status: "pending"` beside `resume_requested: true` —
one row reporting both that nobody has answered it and that answering it started
work. Harmless while the second field did not exist; contradictory once it did.
`_status_of` adds `answered` (`responded`/`acknowledged`), reachable only from
the answer response since the listing carries neither. Answered is checked
BEFORE expiry — an answer that landed is a fact, and an `expires_at` that has
since passed does not un-answer it; the obvious refactor is to test expiry first,
which would make a slow client's own answer vanish, so the ordering is pinned.

The existing test asserted `out.status in ("pending", "expired")` with the
comment 'the point is it returned at all' — it was papering over exactly this.
It now asserts `answered` and, on the spawn-failure path it covers, that
`resume_requested` is False.

THE TWO READS. `_resume_requested`'s docstring claimed it read 'the SAME
accessor … so the two cannot disagree'. True of the accessor, false of the
instant: it is a second read a task hop earlier, and an owner disabling the
opt-in in between gets `true` and no resume. Collapsing them is not the fix —
they answer different questions (one must produce a value for THIS response, the
other is the authority at the moment it would spend), so the window is stated,
with AC #5's own remedy named, rather than described away.

DOCS. architecture.md's ent#329 section described a single caller and stated the
CAS-win property the second caller broke. It now carries the second caller, the
truthy-`_status_conflict` shape that defeated `if not updated`, the
sync-endpoint/no-loop defect and its `anyio.from_thread.run_sync` fix, and what
`resume_requested` actually reports.

Related to abilityai/trinity-enterprise#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(enterprise): bump the submodule pointer to main (a419812 -> 90f2f2c) (#2440)

dev's pointer was OLDER than main's — an inversion, not just staleness. The
next dev -> main release merge would have carried it backwards and undone
ent#443's enterprise-side removal:

  OSS main -> 2a5def3   (ent#443: shared_sessions removed from enterprise)
  OSS dev  -> a419812   (4 behind enterprise main, 2026-08-19)
  ENT main -> 90f2f2c

90f2f2c is a fast-forward from BOTH (a419812...main = ahead 0 / behind 4;
2a5def3...main = ahead 3 / behind 0), so nothing is being rewound.

WHAT THE FOUR COMMITS ARE

  2a5def3  refactor(rooms): remove shared_sessions — it now lives in OSS core (ent#443)
  ff0a4f1  fix(security): guard the enterprise system_settings sinks against
           cleartext credentials (ent#435) — the private twin of the OSS sink
           guard architecture.md already records as "the private submodule owns
           its twin"
  6d82a3f  docs(workspace): agent-initiated asks — design of record
  90f2f2c  feat(credential-vault): governed system credential vault module (ent#279)

WHY IT MATTERS RATHER THAN BEING HOUSEKEEPING. ent#443 moved rooms into OSS
core, and dev has that. With the stale pin an entitled dev install mounts the
OSS rooms routers AND the enterprise shared_sessions module, and relies on
main.py's include-order (OSS before register_enterprise) to decide which one
serves. architecture.md documents that ordering as the transition safety net —
this bump is the follow-through that ends the transition.

VERIFIED BY BOOTING BOTH POINTERS against dev, same box, same DB shape:

  a419812 (today):  17 modules | shared_sessions registered: True  | 6 room paths | 0 errors
  90f2f2c (this):   17 modules | shared_sessions registered: False | 6 room paths | 0 errors

Both boot clean and log "Trinity Enterprise modules registered" — the line
deploy-dev greps. Module count is unchanged because shared_sessions leaves as
credential_vault arrives. No duplicate room paths in either, confirming the
ordering net held; after the bump there is nothing to net.

Gitlink only — no OSS source changes, so public CI (which never checks the
submodule out) is unaffected.

Related to ent#443, ent#435, ent#279

* chore(metrics): code-health dashboard 2026-08-31 @ 135248e9 (#2438)

Co-authored-by: Trinity Agent (trinity) <trinity-agent@ability.ai>

* chore(deps): bump node (#2400)

Bumps the docker-base-images group with 1 update in the /docker/frontend directory: node.


Updates `node` from 24-alpine to 26-alpine

---
updated-dependencies:
- dependency-name: node
  dependency-version: 26-alpine
  dependency-type: direct:production
  dependency-group: docker-base-images
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix(files): shared links were unopenable on mobile — Range, disposition, MIME, CORP (trinity-enterprise#461) (#2439)

* fix(files): shared links were unopenable on mobile — Range, disposition, MIME, CORP (trinity-enterprise#461)

The bytes were never wrong. Verified from the Cloudflare edge, the object
returned HTTP 200 with correct content-length and correct WAV bytes for every
user-agent tried, and the signature check worked. The RESPONSE SHAPE was wrong
in four ways at once, and each one alone is enough to break playback in an iOS
in-app browser:

* no Range support — `Range: bytes=0-1023` returned 200 with the whole 2 MB
  body and no `accept-ranges`. iOS Safari and Telegram's player require a 206
  to start audio at all, so this alone made the file unplayable.
* `content-disposition: attachment` — a forced 2 MB download inside Telegram's
  iOS browser is a blank screen.
* `audio/x-wav` under `nosniff` — unregistered type, so a strict player
  declines it and the browser is forbidden from guessing better.
* `cross-origin-resource-policy: same-origin` on a link whose entire purpose is
  to be opened from another platform.

Plus `cache-control: no-store`, which forbids the in-app browser from buffering
media it will not play without buffering.

THE INLINE CHANGE IS A NARROWING, NOT A REVERSAL. The old code forced
`attachment` on everything with the note 'defense against XSS via
agent-uploaded HTML', and that reasoning is still correct — this route serves
agent-authored bytes from the same origin as public chat. So inline is an
ALLOWLIST (`_INLINE_SAFE_TYPES`: audio, video, image, PDF) and `text/html`,
`application/xhtml+xml` and `image/svg+xml` stay attachments. SVG is called out
because it is the one a reviewer waves through: it is an image by name and a
script host in fact. The type is python-magic-detected from the file's own bytes
at share time, never agent-supplied, and its unavailable-fallback
(`application/octet-stream`) sits outside the allowlist, so the failure
direction is `attachment`. `nosniff` is kept and matters more now, not less.

TWO THINGS THE ISSUE DID NOT ASK FOR, both found while implementing:

* `Content-Length` came from the DB's `size_bytes`, written at share time. Any
  drift from the file on disk is unrecoverable for the client — too small
  truncates, too large hangs — and Range math against a wrong total produces a
  `Content-Range` that contradicts the body. It now comes from
  `os.path.getsize`, with a WARNING on divergence.
* a media player fetches one file as MANY ranged requests. Counting each as a
  download would turn one play into dozens and write an audit row per chunk, so
  the counter and the audit fire only on the transfer START (a plain GET, or a
  range beginning at byte 0).

VERIFIED end-to-end against the real route, not just the parsers:

  full GET      : 200 | type audio/wav | disp inline | ranges bytes
                | corp cross-origin | cc private, max-age=3600
  range 0-1023  : 206 | body 1024 | bytes 0-1023/2048000 | bytes ok
  suffix -500   : 206 | bytes 2047500-2047999/2048000 | bytes ok
  unsatisfiable : 416 | bytes */2048000
  HEAD          : 200 | accept-ranges bytes | content-length 2048000
  no sig / bad sig / unknown id / expired : 401 / 401 / 404 / 410
  html file / svg file : attachment

That covers the issue's Definition of Done line by line, including that the
signature check still rejects unsigned and expired requests.

46 new unit tests, weighted to the allowlist and to the range parser's
silent-corruption case (`bytes=-500` is the LAST 500 bytes; reading it as
start=0 serves the wrong bytes under a 206, which no client can detect).

Related to trinity-enterprise#461

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(files): the CORP header was inert — the security middleware clobbered it (trinity-enterprise#461)

Found by testing the PR against a real local instance rather than a TestClient.

`main.add_security_headers` runs after EVERY route and set
`Cross-Origin-Resource-Policy` with a plain assignment. So the `cross-origin`
policy the file-download route sets — one of the four fixes in this PR, and the
one that decides whether Telegram, Slack or WhatsApp can embed or preview the
link at all — was silently overwritten back to `same-origin` on its way out.

The fix shipped INERT and every test passed, because a bare `FastAPI()` +
router harness has no middleware. Measured on the running server:

  before: cross-origin-resource-policy: same-origin
  after : cross-origin-resource-policy: cross-origin   (file route)
          cross-origin-resource-policy: same-origin    (/health, unchanged)

`setdefault` rather than a route allowlist: absence still resolves to the strict
default, so every other route keeps today's behaviour and a new route has to opt
out deliberately rather than inherit an exception.

Pinned by a source assertion — asserting it end-to-end needs a live stack, and
what must not regress is the `setdefault`; an edit back to `=` would re-break it
invisibly.

Related to trinity-enterprise#461

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(main): lifespan is an orchestrator, not a 580-line procedure (#1028) (#2437)

* refactor(main): lifespan is an orchestrator, not a 580-line procedure (#1028)

`main.py::lifespan` was 580 lines at cyclomatic complexity 109 — the longest
function in the backend and the first item the 2026-06-02 refactor audit named.
It is now 25 lines at CC 1: a flat list of `await _phase()` calls over twelve
startup helpers and four shutdown helpers.

WHAT MOVED, AND THE PROOF THAT NOTHING ELSE DID. Every body is verbatim. That
is asserted mechanically rather than claimed: extracting all non-blank lines
from the sixteen helpers in call order and diffing against the original
`lifespan` body gives 518 = 518, identical. The only relocated line is `yield`.
No behaviour change, no logic touched, no try/except reshaped — each phase keeps
its own guard, because a failing phase must not take the boot down, which is
what the original did.

WHY THE TEST IS THE POINT. Splitting the function is easy; keeping it split is
not, and the thing worth guarding is not the line count — it is THE ORDER.
Boot ordering is load-bearing in ways invisible at the call site: a reviewer
looking at sixteen await lines cannot see that moving one breaks something,
because the coupling lives in the bodies. Before this it was implicit in a
function nobody could read in one sitting; after it, it is a list — which is an
improvement only if something enforces the list.

So `tests/unit/test_1028_lifespan_phases.py` pins the sequence WITH the reason
for each constrained pair recorded beside it: logging first so a later hang
cannot swallow the boot log (#858); the event bus before any WebSocket client
needs a live dispatcher (#306); Docker/system-agent before the fleet sweepers;
startup recovery before the channel transports, so inbound traffic cannot create
an execution that races the reconcile; the event-bus drain LAST on shutdown so
late broadcasts still land. A reorder now fails with the reason attached instead
of surfacing weeks later as a boot bug nobody connects to this commit.

It also pins the `yield` split (a phase appended after it silently becomes
shutdown work), the per-helper thresholds the issue asked for (<100 lines,
CC <20), and orphan/double calls. Mutation-checked five ways — recovery moved
after the transports, event bus after Docker, a dropped shutdown phase, the
drain no longer last, a phase pushed past `yield` — all caught.

ONE DEFECT THIS FOUND IN ITSELF, worth recording because it is the failure this
refactor's shape invites. The extraction moved `@asynccontextmanager` by one
definition: it landed on the first phase helper and `lifespan` was left a bare
async generator, which FastAPI cannot use as a lifespan. Boot-breaking — and the
entire 12,900-test unit suite stayed green, because nothing in it imports `main`
and asks what shape `lifespan` is. It surfaced only from an explicit import
check (`iscoroutinefunction` on each helper returned False for one, with
`co_filename` pointing into contextlib). Fixed, and pinned by its own test.

Docs: the `main.py` row in architecture.md now records that the order is the
contract and where the constraints are, so the next person to add a startup step
knows it belongs in a phase helper.

Scope: one file per the issue's own recommendation. The remaining ACs
(`routers/settings.py`, `routers/ops.py`, `services/git_service.py`,
`services/agent_client.py`, `routers/public.py::public_chat`) stay open on #1028.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): three phase helpers used locals the split left behind (#1028)

/review's scope pass caught a real runtime break in my own extraction, and the
interesting part is that three separate verifications had already passed on it.

`main.py` never imports `database` at module level. The old `lifespan` did
`from database import db as _db` once near the top, and three blocks 100+ lines
later used it through the enclosing function scope: the system-agent
`setup_completed` gate, the Telegram transport and the WhatsApp transport.
`message_router` was the same shape — imported in the Slack block, used by the
Telegram one. After the split all four are NameError.

The severity is in the swallow. Every one of those use sites sits inside
`try/except Exception`, so the boot SUCCEEDS: the log carries "Error starting
Telegram transport: name '_db' is not defined" and the Telegram and WhatsApp
integrations are simply never wired. A silently dead integration, not a failed
boot — and the per-phase guard that makes each phase fail-open is exactly what
hides the extraction bug.

Why the existing checks missed it, all three: the AST equivalence proof compares
LINES and the lines are identical; the structural pin asserts order, thresholds
and decorators, not name resolution; and the `import main` smoke never RUNS
`lifespan`, so nothing resolves those names at import time.

Fixed by re-materialising each import in the helper that needs it — the
original's own idiom — with a comment saying why it is there, so a later reader
does not "tidy" it back out.

CORRECTION TO THE CLAIM: the bodies are no longer byte-identical. They are
verbatim EXCEPT these four re-materialised imports, which is now what the PR
body and the docstrings say. A verbatim claim stops being true the moment a
leaked name has to be restored, and quietly keeping the claim is worse than the
bug.

Also pinned, because this gets more likely with every future split of the same
function: test_no_phase_helper_depends_on_another_phases_locals asserts, per
helper, that `names_loaded - names_bound - module_globals` is empty. Mutation-
checked by deleting the restored `_db` import — reproduces the shipped bug and
turns the suite red.

Also: the phase-count docstrings said "of 10" in 9 helpers; the transports were
split into three after that text was written, so it is 12.

learnings.md gains the class: extract-method has a failure mode the diff cannot
show and an import smoke cannot reach.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(review): the verbatim claim survived in three docstrings (#1028)

Re-review finding. The previous commit corrected "bodies are byte-identical"
in the PR body but left the same claim standing in the code, where it is more
likely to be believed: the three helpers that gained a re-materialised import
still said "the body below is unchanged".

A stale claim next to the exact line that falsifies it is worse than no claim —
it is the thing a future reader checks against before deciding the import looks
redundant. Now each says verbatim EXCEPT the restored import, and points at the
comment explaining why it is there.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): two pre-existing lifespan guards read a function the split emptied (#1028)

CI's regression diff caught 4 new failures, deterministic across all three
seeds. Both guards assert properties of `lifespan`'s SOURCE, and #1028 moved
that source into phase helpers — so they were asserting things about a function
that is now 25 lines of `await` calls.

The properties still hold. The guards had stopped being able to see them, which
is the worse failure: a guard that silently stops covering its subject reads
identically to one that passes.

test_1267_lifespan_db_alias — and this one is pointed: #1267 IS the bug class
/review caught in this branch. It fired when the transport blocks called a bare
`db` while only `_db` was in scope, NameError swallowed by the surrounding
try/except and surfaced as a misleading "Error starting Telegram transport".
The split re-introduced the same class in a new form (`_db` bound in phase 1,
read in three later helpers), and this guard could not see it because it only
ever looked inside `lifespan`.

So it now follows the calls: `_lifespan_surface()` returns `lifespan` plus the
helpers it awaits, and every check scans all of them. The alias check is
STRENGTHENED rather than merely relocated — binding `_db` somewhere on the
surface is no longer sufficient, because after the split each helper is its own
scope, so every function that READS `_db` must bind it. That assertion fails on
the exact defect this branch shipped.

test_858_dockerfile_unbuffered — the #858 invariant is an ORDERING one
(setup_logging -> first-run notice -> event_bus.start), and after the split
those three sit in three different functions. `_lifespan_body()` now flattens
the phases inline in call order, so the existing index comparisons keep meaning
what they meant. An unresolvable helper is left as the bare `await` rather than
skipped, so a phase this cannot expand can never silently drop the statements
it contains.

No production code changed.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(retention): every install writes its own retention rows, not just fresh ones (#2085) (#2432)

#1645 closed #1638 by reverting OPS_SETTINGS_DEFAULTS to the wide historical
values and applying the #1039 community floor through explicit system_settings
rows seeded on FRESH installs only. Every install that has ever upgraded rather
than been created fresh therefore had no rows at all, so cleanup_service
resolved all 11 windows at prune time from a dict that ships inside the backend
image and is replaced on every rebuild. The only thing between a future edit to
that dict and the #1638 failure mode — a silent hard-DELETE of existing data
seconds after the next boot, green /health, no error — was a code comment.

database._seed_retention_windows{,_engine} now writes an explicit row for every
RETENTION_OPS_KEYS member that has none, at the value already in force, on every
boot and regardless of install age. Behaviourally inert: it writes the number the
prune already used, so nothing prunes differently the day it runs.

Three properties are load-bearing:

* The key set is DERIVED from RETENTION_OPS_KEYS, never a second hand-written
  list, so a window added later is covered the day it ships instead of quietly
  inheriting the image default forever (the issue text said "eight windows"; it
  was 11 by the time this landed — ent#433 added two, #2216 a third).

* Ordering. It MUST run after _seed_fresh_install_retention. Both writers are
  insert-or-ignore, so the first to reach a key wins: reversed, a fresh install
  silently gets the wide defaults instead of the #1039 floor — the community
  floor deleted by the change meant to protect retention. Pinned behaviourally
  and by a source-order guard on both the SQLite and engine arms.

* It must actually run. The first cut imported the constants from
  services.settings_service, which has a module-level `from database import db`.
  init_database() is called from DatabaseManager.__init__, i.e. while database.py
  is still executing its own module body, so database.db does not exist yet and
  that import raises ImportError — which this seed's fail-safe contract then
  SWALLOWS. The feature was dead on every boot with a fully green unit suite,
  because every in-process test calls the function after database has finished
  importing. RETENTION_OPS_KEYS, OPS_SETTINGS_DEFAULTS and
  NON_ROW_RETENTION_OPS_KEYS therefore move to config.py (a leaf, already home to
  OPS_SETTINGS_VALIDATION and validate_ops_setting for these same keys) and are
  re-exported from settings_service, extending the pattern
  COMMUNITY_FRESH_INSTALL_SEED already used for exactly this reason. Two tests
  now pay for a real subprocess import; both fail if the import is reverted.

Stated tradeoff: a seeded install stops inheriting later changes to the code
default in EITHER direction, so widening a window for existing installs becomes a
deliberate migration rather than something that arrives silently with an image.
That is the intended consequence — retention becomes explicit per-install config
instead of implicit inheritance from whatever image happens to be running,
symmetric with the rule the OPS_SETTINGS_DEFAULTS comment already imposes on
narrowing.

backup_retention_days is seeded too. That makes OPS_SETTINGS_DEFAULTS' value the
one that lands in the DB for a key whose private reader
(db_backup_service.effective_backup_retention_days, inverted coercion) falls back
to its own module constant; the two are now parity-tested.

No schema change, no migration — row inserts at boot, same as the #1638 seed.

Not fixed here: generic DELETE /api/settings/{key} carries no RETENTION_OPS_KEYS
guard (only PUT does), so an admin can still delete a window row. After this it is
transient — the next boot re-seeds it — but the asymmetry with PUT remains.

Unblocks ops#300 once deployed: with rows on every install, /update step 8e can
drop its source-text guessing for a plain assertion over stored values.

Closes #2085

* fix(watchdog): stop false-orphaning executions parked before the agent spawns them (abilityai/trinity#2433) (#2435)

* fix(watchdog): stop false-orphaning executions parked before the agent spawns them

The cleanup watchdog's proof-of-life (GET agent/api/executions/running:
running ∪ recently-completed) could not see an admitted execution that was
waiting in the backend's global agent-call queue, in the agent's CPU-sized
default thread pool, behind the agent chat lock, or in the post-exit drain
before unregister(). After the 60s grace it wrote a false `failed`
("completed on agent but status not reported"), released the slot, and the
parked call then ran anyway — billed, overbooked, its late 200 silently
overwriting the row (#378). Reproduced twice locally; three mechanisms, one
string.

Orphan now means: the agent does not know the execution AND no live backend
dispatcher owns it.

- agent server: /api/executions/running gains `pending_ids` (accepted at
  /api/task, /api/chat and the #1083 async spawn but not yet spawned; lazily
  expired) and `recently_completed_ids` covers exited-but-registered handles.
  Cancel-while-pending is consumed by register() (SIGKILL at spawn, #679 marker
  kept); the pre-spawn 409 is only an optimisation. Headless runs use a
  dedicated 32-thread pool pinned to MAX_PARALLEL_TASKS_CEILING_MAX; the Gemini
  runtime now registers its subprocess at both Popen sites (it never did).
- backend: every outbound agent call is registered for its whole lifetime
  (track_inflight_dispatch — queue wait, connect retries, POST) in an
  in-process registry plus a cross-worker Redis liveness marker
  execution:inflight:{id} (60s TTL, one refresher task per process, 15s tick).
  The watchdog reads a tri-state verdict (alive / absent / unknown) and
  withholds recovery on `alive`, and on `unknown` only while a dispatcher could
  still own the row; a process with no Redis reads `absent` (its own registry
  is the whole truth). CleanupReport.dispatch_inflight_skipped counts withheld
  rows; the orphan error string states what was observed.
- a park no longer spends the run's budget: at grant, a park ≥ 5s restamps
  started_at (admission kept in queued_at, the drained-backlog shape, CAS on
  RUNNING + NULL lease) and renews the slot lease (ZADD XX + EXPIRE together);
  the refresher renews the slot every tick while parked.
- parked rows are cancellable and agent-scoped: terminate consults the
  in-process registry, then the cross-worker cancel key; a parked phase is
  finalized CANCELLED and the grant raises BackendAgentCallCancelled, where the
  dispatcher writes CANCELLED itself (never FAILED; the /chat arm answers 409).
- terminate_execution gains ONE agent-scope gate at its entry for all three
  arms: the row behind the caller-supplied task_execution_id must belong to the
  agent the route proved (uniform 404; an unreadable row fails closed with
  503). The proxy arm's 404 scoped only execution_id while the CANCELLED CAS
  was keyed on task_execution_id, so a caller authorised on agent A could flip
  agent B's running row (found by the /cso --diff verifier; report under
  docs/security-reports/).
- packaging: BACKEND_AGENT_CALL_LIMIT / BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S
  forwarded in prod + hosted compose and documented in .env.example; the >5s
  queue-wait warning fires on both acquire branches.

Verified: full unit suite under CI conditions 12969 passed / 0 failed
(baseline origin/dev 12863 / 0); Repro A 10/10 success (2 parked 485s,
withheld at both watchdog cycles, re-anchored at dispatch); Repro B 8/8
success (5 parked, two waves); live pending_ids probe on the agent. The
agent-side half needs a rebuilt base image; the backend half alone covers old
images through the whole-call marker.

Fixes abilityai/trinity#2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(watchdog): close the cross-worker cancel race and bound the exited-but-registered set (#2435 review)

Review of the #2433 fix found that it reintroduced the #378 symptom in a
narrower window and turned a pre-existing registry leak into a permanent one.

1. Cross-worker cancel acted on a marker phase that predated its own write.
   `entry.phase` flipped parked->calling in memory only; the marker was
   rewritten by the 15s refresher, so `execution:inflight:{id}` advertised
   `parked` for up to a full tick after the POST had begun. Under --workers 2
   about half of all cancels are served by the worker that does NOT own the
   coroutine and therefore read it: the row was finalized CANCELLED and its
   slot released while the agent ran the turn to a billed completion whose
   SUCCESS then lost the CAS. Closed by ordering, not by narrowing — the owner
   publishes the transition in the SAME round-trip that reads the cancel key
   (`_publish_calling_and_check_cancel_sync`), and the remote sets the cancel
   key BEFORE re-reading the phase (`_set_cancel_then_reread_phase_sync`), so
   an observed `parked` gives W_remote(cancel) < R_remote(marker) <
   W_owner(marker) < R_owner(cancel) and the grant is guaranteed to see the
   key. Neither side pays an extra round-trip. The owner gates the publish on
   the ENTRY's age rather than this attempt's park, because
   `track_inflight_dispatch` wraps the whole retry loop and a retry can grant
   instantly under a marker a tick left saying `parked`; the remote's scope
   check stays on its first read, so no key is written for a foreign agent.

2. `list_recently_completed_ids` reported exited-but-registered ids with no
   age bound, so a leaked entry was agent-known forever and the watchdog never
   recovered that row — a regression against pre-#2433, where `list_running()`
   self-healed it. Now bounded by the same 300s TTL as the buffer, measured
   from when the exit was first OBSERVED (not `started_at`, which would drop a
   long turn the moment it entered its drain). The leak is also closed at
   source: `register()` SIGKILLs the group for a cancel that arrived while
   pending, so the following `stdin.write` can raise BrokenPipeError — all
   three prompt-writing runtimes (claude_code, gemini x2) now pair that write
   with `unregister()` on failure.

3. `restamp_execution_dispatch` is a sync sqlite write and ran on the event
   loop, while both semaphores are held and the queue is by definition
   congested. Now `asyncio.to_thread`, like the slot renewal beside it.

Smaller items from the same review:
- /api/chat sizes its pending entry to PENDING_CHAT_TIMEOUT_SECONDS (7200s):
  `ChatRequest` carries no timeout and a chat can wait on the execution lock
  for the agent's whole budget, so the /api/task default evicted the entry
  mid-wait. Its discard now wraps the lock acquisition, so a request cancelled
  while waiting (client disconnect) cannot leak one.
- Phase 3 batches its in-flight verdict read (one MGET per cycle, not per row),
  matching Phase 0.
- `renew_slot` refuses, score untouched, when the metadata hash has already
  expired: `ZADD XX` succeeds while `EXPIRE` no-ops, so it used to report a
  renewal it had not performed and re-anchor exactly the ZSET-without-hash
  state canary S-03 calls `missing`.
- `register_pending` logs at DEBUG (it fires on every /api/task and /api/chat).
- Documented that the in-flight marker is not eviction-proof under the prod
  `allkeys-lru` policy.

Tests: tests/unit/test_2433_review_fixes.py (15) — 11 of them fail against
cfc2cfef, verified in a worktree. Full unit suite under CI conditions
(clean origin/dev worktree, no submodules): 12985 passed, 0 failed.

Refs abilityai/trinity#2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* test(resume): the guard architecture.md promised did not exist (ent#430 review)

The reviewer's one condition before merge. `architecture.md`'s 'Two callers,
one rule' bullet said the CAS-win rule was 'guarded now by enumerating every
caller rather than the one route ent#329 knew about, so a third site inherits
the rule instead of re-losing it'. No such guard existed:
`test_dispatch_hangs_off_the_cas_win_only` read exactly one hardcoded file,
`routers/operator_queue.py` — so `client_portal/asks/service.py`, the caller
this PR adds and the one that LOST the rule, was outside its reach.

A sentence claiming protection that is not there is worse than no sentence: the
next person adding a dispatch site reads it and stops looking. This is the shape
#2428 filed a learnings entry about this morning — a comment that names a
failure mode is a request for a guard — so it lands the same way.

DISCOVERED, NOT LISTED. `_dispatch_call_sites` walks the backend tree for
callers, because a hardcoded list structurally cannot catch the case that
matters: the file it would need to check is the one being added.

ASSERTED AGAINST CODE, NOT FILE TEXT — and this is the part I got wrong first.
The initial version tested `"_status_conflict" in source` against the raw file
and MUTATION PROVED IT BLIND: deleting the check from the `if` still passed,
because the long comment above it explaining the race still contained the
string. A source-substring guard cannot tell a check from a paragraph about the
check — the same defect the guard exists to prevent, inside the guard. It now
parses each dispatching function and compares `ast.unparse` output, where
comments do not survive.

Verified by three mutations, each caught:
  1. delete the check in asks/service.py, keep the comment  -> FAIL
  2. neuter the check in routers/operator_queue.py          -> FAIL
  3. add a brand-new third caller with no check at all      -> FAIL
and all 23 pass on the real tree.

`test_the_discovery_walk_finds_both_known_callers` pins the floor, so a rename
of the helper cannot leave the loop iterating an empty list and passing in
silence — the failure a discovery guard trades for the one it fixes.

ALSO (non-blocking, from the same review): `WorkspaceAsk.status`'s comment still
read 'pending | expired (terminal ones are not listed)' after `_status_of`
gained a third value. Corrected to say where each value is reachable from.

The remaining non-blocking item — `resume_requested` and the new `answered`
status are unconsumed by any surface — is deliberately NOT in this commit. It is
a product decision about where a transient confirmation lives, and it is filed
so it stays a decision rather than becoming an oversight.

Related to abilityai/trinity-enterprise#430
Related to abilityai/trinity-enterprise#329

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(systems): the four post-deploy endpoints — a nonexistent DB call, an ungated restart, a broken export round-trip, and prefix-collision membership (#2373)

The deploy half has been hardened by every commit since ent#124; the four
post-deploy endpoints were essentially untouched since 2025.

## Membership is now ONE predicate

`get_system`, `restart_system` and `export_manifest` each matched
`startswith(f"{system_name}-")`, so an operation on `acme` also captured every
agent of a system named `acme-extra` — including `restart`, which stops and
starts containers. Three copies of a wrong rule.

`system_service.system_member_names` is the one rule, and it prefers TAGS:
`configure_tags` already applies the system name to every member, so a tag is a
RECORD of membership where a prefix is an inference from a naming convention.
The prefix survives only as a fallback for pre-tag deployments, narrowed so an
agent claimed by another system's own tag is excluded — a tagged `acme-extra`
agent is never captured by `acme` even there. A failing tag read degrades to the
prefix rather than 500ing.

Residual, stated rather than hidden: two systems deployed BEFORE tagging where
one name is a prefix of the other remain ambiguous, because nothing distinguishes
them. This is also the prerequisite for the teardown verb, where the same
collision would delete rather than restart.

## GET /{name} returns real schedules

It called `db.get_agent_schedules`, which does not exist — the facade exposes
`list_agent_schedules` and `database.py` deliberately has no `__getattr__`
fallback. The AttributeError was swallowed by the surrounding `except
Exception`, so every response omitted `schedules` for every agent and logged one
warning each, while `tests/test_systems.py` never asserted on the key. Exactly
the failure mode the db facade's own comment warns about — so the test also pins
that the fallback stays absent, since adding one would turn the next typo into a
silent Mock.

## POST /{name}/restart is creator-gated

It was bare `get_current_user` — below `POST /deploy` and below even the
READ-ONLY bundled-catalog routes — so any authenticated principal, including
`role: user`, could stop and start every container in a system whose agents it
could see. A mutating fleet-wide verb under a lighter gate than the catalog it
reads is an oversight, not a decision. `require_role` also rejects agent
principals (#1890), which matters because an agent-scoped MCP key resolves to
its owner carrying the owner's role.

## Export round-trips

The non-full-mesh permissions branch sliced `target_agent[len(name)+1:]` with no
membership filter — the sibling branch had one — so an edge pointing outside the
system exported as a blind-sliced garbage short name that then failed
`validate_manifest`'s unknown-agent check on re-deploy. The export broke its own
round trip. Both branches now test membership.

And the export no longer embeds the instance-global `trinity_prompt` as the
manifest's `prompt:`. Deploying that manifest elsewhere overwrote THAT
instance's platform-wide prompt — a fleet-wide side effect from what reads like
a copy of one system. Nothing records whether the source system ever set a
prompt, so there is no honest way to distinguish it from whatever the instance
happens to have configured, and the only correct export of an unknown is to
omit it.

## Two preview hardenings

Unknown PER-AGENT keys now warn like top-level ones (ent#126): `credentials:`,
`skills:` and `display_label:` are the fields people try first and they vanished
in silence.

Preview and deploy now resolve the identical resource default. Deploy hardcoded
`{"cpu": "2", "memory": "4g"}` while `_preflight_template` validated against the
admin-configurable `get_agent_default_resources()`, so the two disagreed the
moment an admin moved the fleet default — the one spot that escaped ent#126's
pure-resolver no-drift pattern.

## Verification

14 unit tests, one per defect plus the exempt shapes. Two mutation-checked: the
restart gate and the tag-first membership each turn a test red when reverted.
414 pass across the system/manifest/ent#126/#1884 suites.

`tests/test_systems.py` is live-backend tier and…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants