Skip to content

Setup improvements - #5

Merged
vybe merged 1 commit into
mainfrom
setting-up-project-improvements
Jan 16, 2026
Merged

vybe merged 1 commit into
mainfrom
setting-up-project-improvements

Conversation

@dolho

@dolho dolho commented Jan 15, 2026

Copy link
Copy Markdown
Contributor

Utilize the modern docker compose instead of legacy docker-compose. Allow configuring frontend port via env variables

d of legacy docker-compose. Allow configuring frontend por via env variables
@dolho
dolho requested a review from vybe January 15, 2026 14:11
@vybe
vybe merged commit 8310a83 into main Jan 16, 2026
oleksandr-korin added a commit that referenced this pull request Jan 19, 2026
Test Results:
- T3.1: Human approval (approved) ✅
- T3.2: Human approval (rejected) with on_error:skip_step ✅
- T3.3: Approval timeout ⚠️ (not implemented - needs scheduler)
- T3.4: Approval with artifacts/context ✅

Key Findings:
- Approval API works: GET /api/approvals, POST .../approve, POST .../reject
- Rejection causes step to FAIL by design (use on_error:skip_step)
- Timeout enforcement NOT implemented (deadline set but not checked)
- Approval description can include dynamic context

Known Limitations:
- Issue #5: Approval timeout requires scheduler to enforce
- Issue #6: Rejection = failure (by design, documented)

Running Total: 11/22 tests passing (50%)
Refs: PROCESS_ENGINE_ROADMAP.md Phase 1
oleksandr-korin added a commit that referenced this pull request Jan 19, 2026
Test Results:
- T3.1: Human approval (approved) ✅
- T3.2: Human approval (rejected) with on_error:skip_step ✅
- T3.3: Approval timeout ⚠️ (not implemented - needs scheduler)
- T3.4: Approval with artifacts/context ✅

Key Findings:
- Approval API works: GET /api/approvals, POST .../approve, POST .../reject
- Rejection causes step to FAIL by design (use on_error:skip_step)
- Timeout enforcement NOT implemented (deadline set but not checked)
- Approval description can include dynamic context

Known Limitations:
- Issue #5: Approval timeout requires scheduler to enforce
- Issue #6: Rejection = failure (by design, documented)

Running Total: 11/22 tests passing (50%)
Refs: PROCESS_ENGINE_ROADMAP.md Phase 1
vybe pushed a commit that referenced this pull request May 4, 2026
* fix(security): split Docker compose into platform and agent networks (#589)

Redis at 172.28.0.0/16 was reachable from any agent container. AISEC scan
3aad5469 demonstrated end-to-end exfiltration / cross-user task injection
from a legitimately deployed agent. Network segmentation is the strongest
control — agents now physically cannot route to Redis.

Topology:
- trinity-platform (172.29.0.0/16, NEW) — Redis, scheduler, vector
- trinity-agent (172.28.0.0/16, name preserved) — frontend, agents
- Backend / mcp-server / otel-collector / cloudflared straddle both

Agent-creation sites in services/agent_service/* and system_agent_service.py
need zero changes because the agent-network external name is preserved.

Dev: bind Redis host port to 127.0.0.1:6379 (was 0.0.0.0). Tests connect
from the dev machine; LAN cannot. Auth lands in the next commit.

Refs #589 — acceptance criterion #3 (network segment separation).

* fix(security): mandatory Redis auth, ACL users, auth-aware healthcheck (#589)

Both compose files now enforce two passwords (REDIS_PASSWORD admin /
REDIS_BACKEND_PASSWORD runtime) with the fail-on-missing :? form.
docker compose refuses to render without them.

Per-user ACL via inline --user flags. Additive (start from zero, allow
only what the runtime needs) — never +@ALL -X, which lets newly added
dangerous commands through. backend + scheduler get standard data
families plus scripting/transactions/pubsub minus -@dangerous, which
covers FLUSHALL, CONFIG, SHUTDOWN, MIGRATE, REPLICAOF, MONITOR.

Verified at runtime against redis:7-alpine: PING/SET/GET work for the
backend user, FLUSHALL and CONFIG GET return NOPERM, unauth requests
return NOAUTH.

REDIS_URL on backend + scheduler now embeds the backend ACL user.
mcp-server: REDIS_URL and depends_on:redis dropped in prod compose
(zero Redis imports in src/mcp-server/).

Healthcheck pings as the backend ACL user so a typo'd ACL keeps redis
unhealthy and gates dependent services. depends_on:redis switches to
service_healthy so backend/scheduler don't race the ACL load.

Refs #589 — acceptance criteria #1, #2, #5.

* fix(scheduler-test-rig): mirror Redis auth posture (#589)

Without this, scheduler container fails fast on startup against the rig
because src/scheduler/config.py requires creds in REDIS_URL after #589.
No ACL or network split here — this is a 2-service standalone debugging
rig, not the production posture.

* fix(security): fail-fast on REDIS_URL missing credentials (#589)

Backend (src/backend/config.py) and scheduler (src/scheduler/config.py)
now raise RuntimeError at import time if REDIS_URL is unset or lacks
credentials.

Removed the splicing fallback in backend config that papered over an
unauth REDIS_URL by joining REDIS_PASSWORD into the URL — single source
of truth (compose) eliminates silent drift.

Tests that import backend modules need a creds-bearing REDIS_URL in their
environment; tests/conftest.py will set a dummy one in the test commit.

Refs #589 — acceptance criterion #5.

* fix(webhooks): use REDIS_URL for rate-limit client (#589)

Webhooks rate-limit was the one Redis client that bypassed REDIS_URL —
it used redis.Redis(host="redis", port=6379) and would silently
fail-open under requirepass. Switching to redis.from_url(REDIS_URL)
picks up the credentialed URL like every other client.

Also: distinguish auth/ACL errors (logged at ERROR with exception class)
from transient errors (WARN). Fail-open behavior preserved so a Redis
blip doesn't 500 legitimate webhooks, but a misconfigured deploy now
surfaces in alerts instead of via a webhook abuse incident.

Drops the now-unused REDIS_HOST/REDIS_PORT env reads.

* feat(deploy): auto-generate Redis passwords on fresh installs (#589)

start.sh ensure_redis_passwords matches the existing
CREDENTIAL_ENCRYPTION_KEY pattern, with one safety guard:

- Fresh install (no redis-data volume) → generate both passwords with
  openssl rand -hex 24 and append to .env. One-command boot keeps
  working.
- Existing volume + missing password → refuse with a loud error pointing
  at docs/migrations/REDIS_AUTH.md. Re-keying a populated Redis would
  lock the backend out of its own data; ops needs to follow the explicit
  upgrade path.

Idempotent — second run is a no-op when both passwords are already set.

* docs(security): add Redis auth migration guide + architecture notes (#589)

- docs/migrations/REDIS_AUTH.md: operator upgrade guide. Covers fresh
  installs (auto-generated by start.sh), live upgrades (down
  --remove-orphans + docker network rm + add passwords), production,
  and verification commands.
- docs/memory/architecture.md: new "Network Topology (Issue #589)"
  section above Container Security. Documents the two-network split,
  service membership table, the "agents NEVER on platform network"
  rule, and the three Redis ACL users + their access patterns.

* test(security): network isolation, ACL, fail-fast, webhook rate-limit (#589)

tests/conftest.py: top-level autouse env stub for backend imports.
Backend config now raises at import-time if REDIS_URL lacks credentials;
without this, every test that transitively imports backend modules
breaks. Real Redis tests under tests/security/ override via their own
conftest from .env. Adds the `integration` marker.

tests/unit/test_config_fail_fast.py (new): backend refuses to import
without creds-bearing REDIS_URL. 3 cases — missing env, unauth URL,
URL with creds.

tests/security/test_redis_network_isolation.py (new): 5 integration
tests covering acceptance criteria #1-#3:
  - agent-network container has no route to redis (BLOCKED)
  - unauth client gets NOAUTH on platform network
  - backend ACL user can PING with creds
  - backend ACL user FLUSHALL → NOPERM (no admin)
  - backend ACL user CONFIG GET → NOPERM (no requirepass leak)

tests/security/conftest.py (new): session-scoped fixture loads real
.env values for the integration tests; skips the suite if missing.

tests/integration/test_webhook_rate_limit.py (new): regression for the
from_url switch in webhooks.py. Self-contained — creates agent +
schedule + webhook token inline, hits 11×, expects 429 on the 11th.
Catches the silent fail-open if Redis auth ever regresses.

tests/run-integration.sh (new): pytest -m integration runner. Excluded
from run-smoke.sh per the smoke runner's ~30s no-Docker contract.

* docs(security): detach agents before network rm (#589)

Trinity-managed agent containers are created via the Docker SDK
outside compose, so they store the agent network's UUID, not its
name. After `docker network rm trinity-agent-network` (step 3 of
the upgrade procedure), any later `docker start <agent>` fails:

    Error response from daemon: failed to set up container
    networking: network <old-uuid> not found

Compose-managed services don't hit this — they're recreated with
fresh network refs on `up`. Agent containers aren't, so they keep
the stale UUID until disconnected.

Add an explicit detach loop as step 2, before the network removal.
Verified against a populated install with one running and four
stopped agents: all five reattach cleanly to the new network on
next start.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): CSO OBS-1/2/3 follow-ups — webhook rate-limit + healthcheck hardening (#589)

Resolves three observations from the CSO audit
(docs/security-reports/cso-2026-05-04-589-diff.md):

OBS-1 — webhook rate-limit fail-open + connection-per-request DoS amplifier:
* Added in-process secondary rate limiter (3x primary, per-worker) in
  src/backend/routers/webhooks.py. Bounds blast radius during a Redis
  outage without breaking the documented fail-open philosophy.
* Cached the Redis client at module level under threading.Lock with
  double-checked init. _check_webhook_rate_limit resets the cache on
  inner exceptions so stale connections rebuild cleanly. Without
  caching, a flood would open a fresh TCP per request and exhaust
  Redis maxclients — turning the rate limiter into the DoS amplifier.

OBS-2 — tightened _TOKEN_RE from {20,60} to {43} matching
secrets.token_urlsafe(32) (verified against db/schedules.py:524).

OBS-3 — switched all three compose healthchecks from
`redis-cli -a $$PASS` to `REDISCLI_AUTH="$$PASS" redis-cli` so the
password no longer appears in /proc/<pid>/cmdline.

Additional #589 hardening (caught while resolving OBS-1):
* src/backend/config.py + src/scheduler/config.py: tightened the
  REDIS_URL credential check from `"@" in url` substring to urlparse
  validation. Catches redis://@redis:6379, redis://user@redis:6379, etc.
* src/scheduler/main.py: redact password from REDIS_URL before logging
  (was leaking via Vector log aggregator).

Tests:
* tests/unit/test_webhook_rate_limit_inprocess.py — 7 new tests covering
  cap, window expiry, token isolation, runtime-error fallback, regex
  shape, cache hit, cache reset.
* tests/unit/test_config_fail_fast.py — 4 new parametrized cases for
  malformed-credential URL rejection.
* 15/15 unit tests pass.
* Live Redis healthcheck verified — trinity-redis reports healthy with
  the new REDISCLI_AUTH form; `redis-cli ping` returns PONG.

Also adds .gstack/ to .gitignore so future skill artifacts stay local.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request May 5, 2026
* docs(planning): add Session tab design — --resume-default chat surface

Adds docs/planning/SESSION_TAB_2026-04.md, the comprehensive plan for a
new "Session" tab living alongside Chat. Sessions reattach to their own
Claude Code JSONL via --resume, preserving tool memory, mid-skill state,
and reasoning state across turns.

Plan covers:
- UI design (tab placement, multi-session model, +New Session, Reset memory)
- Data model (agent_sessions / agent_session_messages — parallel to chat)
- Backend architecture (separate router, single shared change to
  task_execution_service for persist_session plumbing)
- Phased rollout (foundation → backend → frontend → hardening → GA)
- Edge cases & failure-mode lessons baked in from a prior local spike
  (parser bug, --no-session-persistence dependency, cold-turn detection,
  port allocation)
- Test plan including the cross-session contamination test for
  Anthropic claude-code#26964
- Retention/cleanup policy, observability, security checklist
- Local-first workflow: implementation runs entirely on this branch
  until validation passes; only then does the standard SDLC engage
  (issue, push, PR)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add agent_sessions + agent_session_messages tables

Phase 1.1 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).
Schema definitions go in db/schema.py for fresh installs; the matching
idempotent migration agent_sessions_tables in db/migrations.py upgrades
existing databases.

The schema mirrors chat_sessions / chat_messages but is strictly parallel
— no foreign keys, no shared columns, separate index namespace. Three
fields are unique to the session model:

- agent_sessions.cached_claude_session_id — the Claude Code session UUID
  the next turn will pass to ``--resume``
- agent_sessions.consecutive_resume_failures — drives the resume-failure
  fallback (Phase 2.2)
- agent_session_messages.cache_read_tokens — observability for whether
  Anthropic's prompt cache engaged

CASCADE on session delete cleans up message rows automatically.

Verified locally: backend restart applies the migration cleanly, tables
have 15 columns each with correct types/defaults/PKs, all four indexes
created, second restart confirms idempotency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add SessionOperations for Session tab persistence

Phase 1.2 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).

- Adds AgentSession and AgentSessionMessage Pydantic models in db_models.py
  with the new fields the Session tab needs beyond ChatSession/ChatMessage:
  cached_claude_session_id, last_resume_at, consecutive_resume_failures on
  the session row, and cache_read_tokens + claude_session_id on each message.

- Creates db/sessions.py with a SessionOperations class mirroring the
  ChatOperations shape: create_session, get_session, list_sessions,
  delete_session, add_session_message, get_session_messages, plus the
  Claude UUID cache helpers (get/update/clear_cached_claude_session_id)
  and resume health helpers (mark_resume_failure, mark_resume_success).

- Wires the new ops into the DatabaseManager facade alongside the
  existing _chat_ops, with one delegating method per public operation.

No router, no agent-server change, no frontend yet — those land in later
phases. Tables agent_sessions and agent_session_messages already exist
from the prior schema commit.

* feat(session-tab): backend foundation for --resume-default Session surface

Phases 1.3 through 1.7 of the Session tab plan
(docs/planning/SESSION_TAB_2026-04.md). Pure backend / agent-server work
behind a flag — no UI surface yet, no behavior change to Chat or any
existing /task caller.

Agent server (base image):

- Stream-json parser fix (Appendix B). Both parse_stream_json_output and
  process_stream_line now recognize {"type":"system","subtype":"init"}
  for session_id capture, with the result event as a fallback when init
  was missed (truncated streams). The legacy bare-init shape is
  intentionally rejected. This is the same bug that would have made
  Session caching corrupt on every cold turn.

- Same bug in execute_headless_task's permission-mode validation site:
  the check matched the wrong shape, so permission_mode_validated never
  flipped to True and the protective kill-on-misconfigured-permission
  path silently failed open. Now uses type=system + subtype=init.

- New persist_session flag threaded through ParallelTaskRequest →
  routers/chat.py → AgentRuntime ABC → ClaudeCodeRuntime.execute_headless
  → execute_headless_task. When True, --no-session-persistence is
  omitted so the JSONL is written and the next turn's --resume can find
  it. --session-id is still passed for unique cold-turn namespace.
  Default False keeps every existing caller stateless.

- gemini_runtime accepts the parameter for ABC parity and ignores it
  (Gemini CLI has no resume).

Backend:

- task_execution_service.execute_task now accepts persist_session: bool
  = False and threads it into the agent payload. All existing callers
  (Chat, schedules, MCP, fan-out, webhooks) keep today's behavior; only
  the future routers/sessions.py (Phase 2) opts in.

- settings_service.is_session_tab_enabled() — feature flag resolving
  system_settings.session_tab_enabled → SESSION_TAB_ENABLED env →
  False. Module-level convenience function exposed.

Tests (run inside trinity-backend container — Python 3.11):

- tests/unit/test_session_operations.py — 9 tests against an isolated
  SQLite DB exercising the full SessionOperations CRUD plus the cached
  claude session UUID lifecycle and resume failure / success counters.

- tests/unit/test_claude_code_session_id_parser.py — 8 tests covering
  both parsers (batch + streaming): system/init recognition, result
  fallback, init-wins-over-result, legacy bare-init rejection, and a
  source-level regression guard for the permission-mode validation
  fix.

- tests/unit/test_session_persistence_flag.py — 8 tests pinning the
  contract: signatures across the runtime ABC, ParallelTaskRequest,
  agent chat router, execute_headless_task, and
  task_execution_service.execute_task. Includes the gating regex check
  on --no-session-persistence and a live signature import to catch
  drift AST parsing alone would miss.

Total: 25 passing tests covering every touchpoint of Phase 1.

Base image (trinity-agent-base) rebuilt to embed the agent-server
changes; existing agent containers will pick them up on next recreate.

* feat(session-tab): backend turn endpoint for --resume-default Session surface

Phase 2 of docs/planning/SESSION_TAB_2026-04.md. Six endpoints under
/api/agents/{name}/session{s,...} that mirror routers/chat.py's auth
model and TaskExecutionService usage but persist to the parallel
agent_sessions / agent_session_messages tables and request
persist_session=True on every turn so each call reattaches via
`claude --print --resume <uuid>`.

Surface gated on is_session_tab_enabled() — flag-off default returns
404 from every endpoint.

  POST   /api/agents/{name}/session                  create row
  GET    /api/agents/{name}/sessions                 list (per-user)
  GET    /api/agents/{name}/sessions/{id}            session + messages
  POST   /api/agents/{name}/sessions/{id}/message    THE turn
  POST   /api/agents/{name}/sessions/{id}/reset      clear cached uuid
  DELETE /api/agents/{name}/sessions/{id}            delete row + msgs

Spike-pitfall defenses baked into the turn endpoint:

- L3 (first-turn-has-no-session-id): the agent_sessions row is created
  server-side via POST /session BEFORE the turn endpoint ever calls
  execute_task. No frontend-first model.
- L2 (cold turn writes empty JSONL): persist_session=True is passed
  unconditionally — Phase 1.4 already wired the flag through the agent
  stack; Phase 2 just promises to set it on every turn.
- L1 (parser misses system/init): trust result.session_id directly —
  Phase 1.3 fixed the parser. Scenario A confirms the captured UUID is
  the real Claude UUID end-to-end.

Phase 2.2 resume-failure fallback: when execute_task returns "no
conversation found" on a turn that had a cached UUID, clear the cache,
mark_resume_failure, and retry once with resume_session_id=None. Logs
event=session_resume_fallback with the stale UUID and consecutive
failure count. Anthropic #39667 (cleanupPeriodDays) and #53417 (CLI
upgrade) both produce this signal.

Phase 2.3 Redis lock: SET NX EX per (agent, claude_uuid) with 5-min TTL
and Lua-script release. Async poll loop (250ms tick) so the event loop
stays free during contention. Cold turns skip the lock (no JSONL to
corrupt). Hard 30s wait ceiling — beyond that the contender gets HTTP
429 with retry hint. Mitigation for Anthropic #20992 (concurrent
--resume JSONL writes corrupt the file).

Per-user ownership at the row layer: even agent owners cannot read or
send into another user's session (E6 isolation in the design doc).
Returns 404 for ownership failures so we don't leak session-id existence.

Tests (tests/integration/test_session_turns.py, run inside
trinity-backend container with docker.sock mounted for testfix
recreation + JSONL surgery in Scenario C):

  Scenario A: 3-turn happy path — same Claude UUID across turns
  Scenario B: turn 2 recalls a secret from turn 1, no text-replay
  Scenario C: JSONL deletion mid-session triggers fallback + recovery
  Scenario D: concurrent POSTs serialise via Redis lock
              (asserts finish_gap ≈ winner_work_time, NOT total wall)
  Scenario E: switching sessions A → B → A preserves A's UUID

5 passed in 54.5s against the live agent-testfix container (recreated
onto the rebuilt base image first per L4 in the plan). Phase 1's 25
unit tests still pass — no regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): frontend Session surface

Phase 3 of docs/planning/SESSION_TAB_2026-04.md. Adds the new "Session"
tab in AgentDetail, gated on the is_session_tab_enabled() platform flag
so it stays invisible until explicit opt-in (default off).

Backend prerequisite — routers/settings.py:

- GET /api/settings/feature-flags exposes a curated allowlist of UI-
  relevant flags to any authed user. The existing /api/settings/{key}
  endpoint is admin-only and would block non-admin frontends from even
  knowing whether to render the Session tab. The new endpoint reads
  through services.settings_service.is_session_tab_enabled() so the
  resolution order (DB → env → False) stays in one place.

Frontend:

- src/frontend/src/stores/sessions.js — Pinia store wrapping the six
  /api/agents/{name}/sessions* endpoints with per-agent state isolation
  and the feature-flag cache. Optimistic user-message insert with
  rollback on send failure.

- src/frontend/src/components/SessionPanel.vue — structural copy of
  ChatPanel reusing ChatMessages + ChatInput + ModelSelector. Differs
  from Chat in three places per the design doc:
    * Sends bare user_message to POST .../sessions/{id}/message — no
      buildContextPrompt text-replay (the agent already has working
      memory via --resume).
    * "Reset memory" button + confirm modal that clears the cached
      Claude UUID without deleting the message log (Phase 3.4).
    * Per-session selector subtitle: turn count, context % used,
      cached-memory dot (emerald/gray), and consecutive_resume_failures
      indicator (Phase 3.5).
  Lean cut for first-visible-surface: voice mic, file upload, and SSE
  dynamic status labels are deferred — those need backend extensions
  (file payload on the turn endpoint, async_mode + SSE on the same).

- src/frontend/src/views/AgentDetail.vue — new Session tab inserted
  between Chat and Dashboard/Schedules, gated on
  sessionsStore.sessionTabEnabled. Layout sites that previously
  branched on activeTab === 'chat' now use a shared isFullscreenTab
  computed so Chat and Session both get the input-pinned-to-bottom flex
  layout. ?tab=session deep-link allowlist updated.

- src/frontend/e2e/session-tab.spec.js — Phase 3.6 Playwright spec.
  Marked @Interactive (not @smoke) because each run makes one real
  Claude API call (~10–60s). Snapshots the prior flag value in
  beforeAll, force-enables for the run, restores in afterAll so a
  failed run doesn't leave the platform with the flag dirty. Three
  cases:
    * tab is hidden when flag is off
    * tab appears, "+ New Session" → send turn → reply visible →
      Reset memory modal opens + closes
    * Chat tab still works after Session interaction; switching back
      preserves Session state

Visually verified in the live dev server: tab renders in correct
position, header layout matches Chat's structure, empty state and
placeholder copy match the design doc, "Reset memory" only shown when
an active session exists, full-viewport flex layout pins input to
bottom.

Phase 1 + Phase 2 work behind this change is unchanged: 25 unit tests
+ 5 integration tests still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): hardening + observability — cleanup service, contamination gate, docs

Phase 4 of docs/planning/SESSION_TAB_2026-04.md. Closes the JSONL
disk-growth loop, validates the GA-blocking cross-session contamination
hypothesis empirically, and lands architecture.md / feature-flows
documentation so the surface is discoverable.

Phase 4.3 — cross-session contamination GA gate (the load-bearing one):

- tests/integration/test_session_cross_contamination.py exercises the
  Anthropic #26964 hypothesis end-to-end. Plants a randomly-generated
  secret token in session A with explicit "do not echo" framing, asks
  session B (different UUID, same agent, same cwd) to recall the token.
  Hard-fails if the exact token leaks; soft-fails on partial-prefix
  recall (PURPLE-DRAGON without the random suffix would only be
  knowable from A's JSONL, not from training).
- PASSED in 9.5s on the current Claude Code version → shared-cwd model
  is safe → Phase 5 rollout unblocked. Test stays in the suite as the
  per-version regression guard.

Phase 4.2 — JSONL cleanup service:

- services/session_cleanup_service.py runs a 6h periodic sweep that
  diffs every running agent's
  ~/.claude/projects/-home-developer/<uuid>.jsonl set against
  db.list_active_claude_session_ids(agent) and reaps orphans whose
  mtime is older than the 1h race guard. Race guard prevents the
  cold-turn-vs-cleanup window where a brand-new JSONL exists on disk
  before the backend has updated cached_claude_session_id.
- Same service exposes a synchronous reap_jsonl(agent, uuid) helper
  called best-effort from routers/sessions.py reset/delete handlers so
  the user-perceived disk-reclaim latency is sub-second. Never raises;
  failures are logged and the periodic sweep is the safety net.
- Implementation uses execute_command_in_container — the same primitive
  git_service / ssh_service / scheduler pre-check / agent terminal use.
  No new agent-server endpoint, no base-image rebuild.
- New db.list_active_claude_session_ids(agent) facade method backed by
  SessionOperations.list_active_claude_session_ids querying every
  agent_sessions row whose cached_claude_session_id is non-null for the
  agent.
- main.py wires startup (staggered +7.5s after cleanup_service to
  offset Docker hits) and clean shutdown.
- tests/integration/test_session_cleanup.py: reset reaps synchronously,
  delete reaps synchronously, periodic sweep keeps the active JSONL,
  reaps an aged orphan, respects the 1h race guard for fresh orphans.

Phase 4.4 — architecture.md updates:

- Background Services table gets a Session Cleanup row.
- New "Session Tab" subsection in API Endpoints documenting all six
  /api/agents/{name}/sessions* routes including the per-user ownership
  rule (404 not 403) and the resume-failure fallback / Redis lock.
- New /api/settings/feature-flags row.
- New agent_sessions / agent_session_messages DDL block in Database
  Schema, with the three Session-specific fields called out
  (cached_claude_session_id, consecutive_resume_failures,
  cache_read_tokens, claude_session_id audit).

Phase 4.5 — feature-flows/session-tab.md vertical slice:

- Full path from UI → API → DB → Side Effects with the JSONL lifecycle
  table, the spike-pitfall defense map (L1/L2/L3/#20992/#26964), the
  error-handling matrix, and the complete test catalog with the docker
  run command for the integration suite.
- feature-flows.md index updated (Recent Updates row + Chat & Sessions
  section entry).

Test totals: 25 unit + 9 integration = 34 tests, all green. Phase 4.3
serves as both the GA gate and the per-Claude-version regression guard.

Phase 4.1 (cache_read_tokens UI surfacing) deferred — the column is
already populated by the Phase 2 turn endpoint; surfacing is a minor
observability follow-up that doesn't block Phase 5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): tag Session-tab turns with triggered_by="session" and a gold badge

Previously Session-tab turns went into schedule_executions with
triggered_by="chat", so the Tasks tab couldn't tell them apart from
the Chat tab. The user-visible signal was that every Session turn
showed up under the sky-blue "chat" badge.

Backend (routers/sessions.py): both call sites that invoke
task_execution_service.execute_task — the cold/resume turn and the
resume-failure fallback retry — now pass triggered_by="session".
Existing rows are unchanged; the cutover is per-write.

Frontend (TasksPanel.vue): adds a "Session" option between "Chat" and
"Manual" in the trigger filter dropdown, plus an amber/gold badge
branch (bg-amber-100 dark:bg-amber-900/30 text-amber-700
dark:text-amber-300) — visually distinct from "paid" (bright yellow)
and from the sky-blue "chat" badge.

triggered_by is a free-form TEXT column (no enum constraint at the DB
or service layer), so adding "session" as a new value doesn't require
any migration or downstream consumer updates. Filter, badge, audit
log, activity stream, and dashboards all just see another value and
display it; nothing has to know about it explicitly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(session-tab): correct context-window accounting + raise frontend turn timeout

Five interrelated fixes from manual testing — all about the per-turn
"context %" metric being misleading and the browser timing out before
long-running session turns finished.

1) Agent server (docker/base-image/agent_server/services/claude_code.py)
   process_stream_line's `result` event handler used to overwrite
   metadata.input_tokens, cache_read_tokens, and cache_creation_tokens
   with the values from result.usage. Those values are CUMULATIVE
   across every internal API call the turn made (Claude Code packs
   tool-use loops into a single user turn that maps to N internal API
   calls). For an 18-iteration turn each reading the same 70K cached
   prefix, result.usage.cache_read_input_tokens = 18 * 70K = 1.26M
   tokens — billing-cumulative, not the prompt size of any single call.
   Overwriting per-message values with that aggregate made our
   context-window-pressure metric grow far beyond the 200K limit even
   when no individual API call was anywhere close to the wall.

   Fix: result handler now only extracts model-level facts (cost,
   duration, num_turns, session_id, error info, modelUsage.contextWindow).
   Per-API-call usage stays in the per-assistant-message handler, where
   the LATEST message's values represent the FINAL API call's prompt
   size — exactly what determines whether the next turn will fit.

   Also added a per-message usage-extraction block to the assistant
   branch of process_stream_line (it previously had no usage extraction
   at all, relying entirely on the result handler — which made my
   first attempt at this fix produce zero values). parse_stream_json_output
   already had the equivalent block (lines 211-215).

   Base image rebuilt; agent-testfix recreated onto the new image
   (image sha 0a1e20b40da1).

2) Backend (services/task_execution_service.py)
   Replaced `context_used = metadata.input_tokens` with
   `cache_read + cache_creation` (with input_tokens fallback when
   caching isn't engaged). input_tokens is sometimes the disjoint
   fresh value and sometimes inflated by the agent server's
   modelUsage.inputTokens override on tool-call turns. cache_read
   and cache_creation come straight from Anthropic's usage object
   and (post agent-server fix) are reliable per-call values that
   monotonically reflect the cached conversation prefix.

3) DB (db/sessions.py)
   total_context_used is now a HIGH-WATERMARK (MAX of prior + new),
   not the latest value. Per-turn context naturally oscillates by ~2x
   between text-only and tool-call turns; the watermark gives users
   a stable monotonic upper bound on session pressure that only goes
   up.

   Capped the watermark at total_context_max as a safety belt against
   any future agent-server bug that emits cumulative-billing token
   counts. Genuine per-call peaks should never exceed the model's
   context window — if they do, that's an accounting error not a
   real overflow, and the UI should display 100% rather than 648%.

4) Frontend (stores/sessions.js)
   Bumped the Axios timeout on the session turn endpoint from 305s
   (~5 min) to 7260s (= TIMEOUT-001 cap of 7200s + 60s slack). The
   session turn endpoint is synchronous and may legitimately run for
   the agent's full execution timeout. With the previous 305s ceiling
   the browser threw a misleading "failed" toast on tool-heavy turns
   that ran longer; the response still landed in the DB and the UI
   recovered after a page refresh, but the user saw a phantom error.

Verified end-to-end with a 6-turn mixed sequence (text + tool-call):
per-call cache_read now reports ~11636 on text-only turns and ~18000
on tool-call turns (one extra round-trip's worth) instead of the
previous bogus 1,257,915 on tool-heavy turns. Watermark grows from
18073 to 18429 across 6 turns — monotonic, no oscillation, real
per-call peaks.

Existing inflated session rows (the bogus 648% / 100% sessions from
before this fix) stay as-is — the watermark cap stops them growing
further but the historical max is permanently stored. New sessions
created after this commit are accurate.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): Phase 5.1 — context warnings, stdout-race recovery, limitations doc

Bundles the round-of-decisions outcomes from the Phase 5 open-questions
discussion. Three code changes, three deferrals documented, one item
gated on a manual test before shipping.

Frontend (SessionPanel.vue) — context-window pressure warnings.
Buckets driven by the active session row's watermark divided by
total_context_max:
  < 75%  : nothing
  75–89% : subtle gray "Session context: X% — heavy" hint
  90–99% : amber banner with "Reset memory" suggestion
  ≥ 100% : red banner warning the next turn may fail or trigger
           memory-loss fallback
No hard send-block at 100% — fallback path stays as the safety net.
Computed values gracefully handle a brand-new session (no row yet).

Agent server (claude_code.py) — stdout pipe race soft recovery. When
a child subprocess inherits Claude Code's stdout, the final result
event line can be lost ("I/O operation on closed file" + "Reader
thread(s) stuck after process exit"). Previously surfaced as 502 with
the misleading "infrastructure failure; retry the task" message even
though the assistant had completed its work and accumulated text into
response_parts. The reply was sitting in the JSONL on disk untouched;
only the closing-stats line was lost in the pipe.

  Fix: at the empty-result classification site, if response_parts has
  accumulated assistant text, log a warning and fall through to the
  success path instead of raising. cost_usd / duration_ms stay None
  for these recovered turns (we don't know what they were); the backend
  records the execution as success with null cost rather than a
  misleading FAILED. Hard failure path stays for the truly empty case
  (no response_parts content).

  Base image rebuilt; agent-testfix recreated onto image
  305731fc34e0. The user's last "comprehensive report" turn that
  produced 67KB of content but threw a phantom error in the UI would
  now succeed cleanly.

Docs (feature-flows/session-tab.md) — Known limitations section. Five
entries that need to make it into the user-facing docs at Phase 5.3:
  1. Voice mic not wired into Session (deferred — Chat tab for voice)
  2. Per-message file upload not yet supported on Session (Phase 5.2)
  3. Agent restore from backup may require fresh sessions (separate
     platform-level issue — workspace volumes not in backup script)
  4. Long Session turns may surface phantom errors in browsers (Axios
     7260s ceiling vs. browser/OS sleep — refresh recovers)
  5. Stdout pipe race recovery is best-effort (recovered turns will
     show null cost / duration in Tasks tab)

Round-of-decisions outcomes:
  ✅ #1 — UI thresholds shipped here, stdout race fixed here
  ⏳ #2 — subscription-change cache clear (E7) — gated on manual
         §5.5 test before shipping the proactive clear
  ⏭ #3 — backup/restore (E9) — separate platform-level issue
         later; documented in §Known limitations
  ⏭ #4 — admin JSONL UI access — deferred follow-up
  📋 #5 — file upload parity — Phase 5.2
  ⏭ #6 — voice on Session — deferred + documented limitation

The §5.5 E7 verification test was added to tests/manual/session-tab/README.md
which is tracked locally only (per the testing-kit precedent). Decision
on the proactive cache-clear hinges on its outcome.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): Phase 5.2 — file-upload parity with Chat

The Session tab's input had drag-drop and the paperclip button (inherited
from ChatInput), but the second arg from ChatInput's submit event
(the files array) was being parameter-discarded in SessionPanel.vue
and never reached the backend. Result: users uploaded a file, sent
their message, and the agent reported "I can't see any attached file
in this conversation" — the upload silently dropped.

Closes the parity gap by mirroring routers/chat.py's exact upload
handling. No new behavior — same limits (3 files, 5 MB each, 5 MB per
image, 20 MB total image budget per WEB_MAX_*), same image-vs-non-image
split, same prompt-line append.

Backend (routers/sessions.py)
  - SessionMessageRequest: new optional `files` list (same shape as
    ParallelTaskRequest.files / WebFileUpload — name, mimetype, size,
    data_base64).
  - Turn endpoint: when body.files is set, decode each via
    services.upload_service.decode_web_file, then run the same
    process_file_uploads() helper Chat uses. Non-images get written
    into the agent workspace via Docker put_archive; images come back
    as already-decoded vision-block dicts in image_data.
  - Append the file_descs lines to a new `effective_message` so the
    agent prompt contains "[File uploaded by X]: name (size) saved to
    path" references inline. Persisted user message stays as the
    original body.message — visible chat log reads naturally; the
    agent sees the file references.
  - Pass image_data to execute_task as `images=` so vision blocks land
    on the next API call. Both the resume call and the cold-retry
    fallback receive the same effective_message + images (files were
    already written to the workspace before the first attempt; cold
    retry just re-references them).
  - 502 from process_file_uploads' all_writes_failed bubbles up
    cleanly; agent-not-found pre-check returns 503.

Frontend (stores/sessions.js)
  - sendMessage() accepts a `files` array in opts; included in the
    POST body when non-empty. Optimistic user-message insert stays
    text-only (the chat log preview doesn't need to render
    file chips).

Frontend (SessionPanel.vue)
  - onSubmit() now takes both args from ChatInput's submit event.
    Allow-empty-text-with-files is permitted (matches Chat's UX).
  - Forwards files into sessionsStore.sendMessage's opts.

Verified end-to-end on a fresh session: uploaded a 66-byte text file
with a unique sentinel embedded; agent read the file out of the
workspace and recalled the sentinel verbatim. No fallback fired.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(session-tab): drop watermark, store last-cache reading directly

agent_sessions.total_context_used was a high-water mark (MAX of all
per-turn cache readings, capped at total_context_max). Claude Code's
auto-compact (~85% of the model window) silently resets the cache
mid-turn, so the watermark asymptoted near the compact threshold and
stopped conveying useful information — every heavy session ended up
visually stuck around 67-69% regardless of actual recent activity.

Switch to direct assignment of the most recent assistant turn's cache
size. The total_context_max cap (cc5c37b) stays as defence-in-depth
against agent-server accounting bugs reporting impossible token
counts. Existing inflated rows on agent-testfix self-heal on next
turn — no migration, no backfill.

Two new unit tests pin the contract: a lower new value overwrites a
higher prior value, and a None reading preserves the prior value
(stdout-pipe-race recovered turns).

* fix(session-tab): rename "% context" to "% last cache" and remove pressure banners

The 75/90/100% banners never fired empirically — Claude Code auto-
compacts mid-turn at ~85% of the model window, which prevents the
underlying metric from ever crossing 75%. A pressure warning that
silently never warns is worse than none.

The subtitle label "% context" reads as session memory pressure when
it's actually the most recent assistant turn's cache size. Relabel to
"% last cache" so the meaning matches what the column actually stores
(after the prior commit dropped the watermark semantics).

Reset modal copy: drop the "context-window pressure" framing for the
same reason, replace with "start a clean line of thought."

Removes orphaned contextPct / contextPctBucket / currentSession
script helpers — no remaining users after the banner deletion.

* docs(user-docs): add Session tab user guide

Covers what the Session tab is, when to use it vs Chat, and the three
limits/behaviours users will hit on a long heavy session:

- The "last cache" metric — explicitly per-turn cache size, NOT
  session memory pressure. Bounces because of auto-compact.
- Auto-compact at ~85% of the model window — what users see (sudden
  drop in % last cache, ~2 min added latency), what survives (working
  memory in compressed form), what doesn't (verbatim history).
- The 50-turn agentic-loop cap (max_turns_task) — per-turn iteration
  budget, NOT session message count. Heavy 12-step tasks routinely
  hit it. Includes a curl recipe for raising the cap per agent via
  PUT /api/agents/{name}/guardrails.

Also documents Reset memory, file attachments, the Session API
endpoint surface, and the four known limitations carried over from
docs/memory/feature-flows/session-tab.md (voice not wired, DR
backup forces fresh sessions, suspended-browser phantom errors,
stdout-pipe-race best-effort recovery).

Modeled after docs/user-docs/agents/agent-chat.md for tone and
structure.

* feat(db): persist auto-compact events on session messages and executions

Claude Code's auto-compact (~85% of the model window) silently summarizes
~170k tokens of conversation into a ~10k summary mid-turn, takes ~2 min,
and previously left no trail in our data — users saw only an unexplained
long execution and a sudden % drop on the next turn.

This change adds three additive nullable columns plus the persistence
plumbing across the standard Trinity router → service → db pipeline:

  - agent_session_messages.compact_metadata (TEXT, JSON list of events)
  - agent_sessions.compact_count            (INTEGER, running tally)
  - schedule_executions.compact_metadata    (TEXT, denormalized for Tasks)

The session-level compact_count drives a future inline "consider starting
fresh" hint without scanning per-message rows. The denormalized copy on
schedule_executions lets the Tasks tab render the badge without joining
session messages — Trinity invariant 1 (router → service → db) preserved
on both the Session router and the task_execution_service.

Migration is idempotent via the existing _safe_add_column helper. Empty
columns on existing rows; populates from the next turn forward once the
agent server is rebuilt with the parser branch (separate commit).

One new unit test pins compact_metadata persistence + compact_count
accumulation across turns; the existing test_session_operations DDL
fixture mirrors the new columns.

* feat(agent-server): parse compact_boundary events and emit structured log

Claude Code emits {"type":"system","subtype":"compact_boundary","compactMetadata":{trigger,preTokens,postTokens,durationMs}}
on its stream-json output when it auto-compacts mid-turn. Both parsers
(parse_stream_json_output for batch, process_stream_line for live) now
recognise the event, append a CompactEvent to metadata.compact_events,
and emit a structured INFO log line per event:

  event=session_auto_compact claude_session_id=... trigger=auto
  pre_tokens=170325 post_tokens=12691 duration_ms=110361

Vector picks the line up via Docker stdout — no infrastructure change.
The compact_events list rides through ExecutionMetadata.model_dump()
into the agent server HTTP response, where the backend extracts and
persists it (separate commit).

Five new unit tests exercise: single-event capture in batch parser,
multi-event ordering within one turn, empty-events on a normal turn,
streaming-parser capture, and model_dump round-trip for the wire shape.

Live behaviour requires rebuilding trinity-agent-base:latest and
recreating each agent container — already done locally at the new SHA.

* feat(session-tab): surface auto-compact events in Session + Tasks UI

Two surfaces, one signal:

TasksPanel — adds a small violet "compacted" badge next to the trigger
badge for any execution that fired one or more compact_boundary events.
Multi-compact turns show `compacted ×N`. Hover tooltip lists each event
with pre→post tokens and duration so latency anomalies become legible
at a glance. Reads task.compact_metadata (denormalized JSON column)
without needing to follow the relation back to agent_session_messages.

SessionPanel — adds an inline italic hint adjacent to the Reset memory
button when the active session has compacted more than 5 times. Reads
session.compact_count (running tally maintained by the DB layer). Quiet
under normal use; visible only when stacked compacts have meaningfully
degraded summary fidelity. Threshold lives as a script-local constant
(COMPACT_HINT_THRESHOLD) for easy tuning from telemetry later.

The currentSession computed (deleted in Bundle A when the pressure
banners went away) is reinstated here since the inline hint needs it.

Pinia store sessions.js needs no change — compact_count and
compact_metadata ride through on the existing API responses without
new state slots.

* fix(db): propagate compact_metadata through DatabaseManager facade

b12307c extended ScheduleOperations.update_execution_status with the
new compact_metadata kwarg but missed the thin facade wrapper in
database.py. The session turn endpoint hit:

  TypeError: DatabaseManager.update_execution_status() got an
  unexpected keyword argument 'compact_metadata'

Trivial fix — the facade just forwards. Caught at runtime on the first
Session turn after Bundle B landed.

* fix(db): propagate compact_metadata through DatabaseManager.add_session_message

Same facade gap as the prior fix for update_execution_status — the thin
wrapper in database.py was missed. Surfaced as:

  TypeError: DatabaseManager.add_session_message() got an unexpected
  keyword argument 'compact_metadata'

…on the first Session-tab turn after the previous facade fix unblocked
the schedule_executions write path.

Adds the same two pass-through kwargs (compact_metadata,
compact_event_count) so the Session router can persist auto-compact
events on session messages and bump the running compact_count tally.

* fix(agent-server): recover from stdout pipe race via JSONL fallback

When a tool subprocess (or MCP grandchild) inherits Claude Code's
stdout fd and wedges the agent server's reader thread, the stream-json
result event is lost. The Phase 5.1 soft-recovery (response_parts !=
[] → synthesize success) only fires when stdout managed to deliver at
least one assistant text block before the wedge. For races that fire
mid-tool-call (zero text emitted to stdout), response_parts is empty
and the soft-recovery falls through to a hard 502.

Symptom users saw: heavy multi-step tasks (e.g. /session-context-pressure
running `python3 -c "..."` via Bash to generate 32KB of synthetic text)
sometimes returned "Execution completed without a result message after
0 tool calls / 1 turns (raw_messages=4)" — even though the JSONL on
disk contained the fully-completed turn. Probabilistic; retry usually
worked; root cause was the kernel pipe buffer exhausting before the
reader thread could drain after grandchildren died.

Fix: when the existing soft-recovery path falls through (no
response_parts text), read the session's JSONL via a side-channel that
is INDEPENDENT of stdout (Claude Code's own session record). Walk
backward to the most recent user-input boundary (string content,
distinguished from tool_results which have list-of-dicts content),
collect every assistant.text block emitted after, and synthesize a
soft-success response. Sets metadata.recovered_from_jsonl=True for
observability. If no text was emitted between boundary and EOF, the
turn is genuinely incomplete and the original 502 surfaces.

11 unit tests cover happy path (matches the actual failure shape with
2× Bash tool_use + tool_result + final text), boundary discrimination
(tool_result must not terminate the user-input search), multi-turn
isolation (only the LAST turn's text recovers), genuine-incomplete
detection (thinking-only or tool-only turns return None and surface
the 502), and robustness (malformed/truncated tail lines, blank lines).

Bounded read at 10MB to defend against pathological JSONL sizes; uses
existing /home/developer/.claude/projects/-home-developer/ path
already known to the cleanup service.

Verified end-to-end on the rebuilt agent-base: 7 turns of
/session-context-pressure executed cleanly with 3 reader-thread races
that drained naturally (none required the new fallback). Recovery is
a safety net only; happy-path behavior unchanged.

Backend reload not required. Agent-base rebuild + container recreate
required to deploy. Backward compat: old images returning no
recovered_from_jsonl field default to False.

* fix(agent-server): pull compact event detail from JSONL after the turn

Bundle B's stdout-side parser detected compact_boundary events but
captured them with all-None pre/post/duration/trigger fields — Claude
Code's --output-format stream-json strips the compactMetadata envelope
on the way out (the JSONL on disk has it, stdout doesn't). Confirmed
on the live agent: stored compact_metadata blobs were
[{"trigger":null,"pre_tokens":null,"post_tokens":null,...}] while the
JSONL contained the canonical shape with real numbers.

Add _extract_compact_events_from_jsonl(session_id, since_iso=None) — a
short helper modeled after the recovery extractor — that scans the
session JSONL post-turn and returns CompactEvent records with the
real detail fields. Called in execute_headless_task right before
returning, scoped via since_iso to records emitted from a turn-start
anchor onward (the JSONL accumulates across the resumed session, so
unfiltered would include compacts from prior turns).

Detection-only stdout branch is retained as a safety net (count
preserved if JSONL read fails) but its log line is removed — the
authoritative log line now fires from the JSONL extract with full
fields.

Effect on existing data: stored compact_metadata blobs from before
this commit keep their nulls (no backfill); new turns from now on
land with full pre_tokens / post_tokens / duration_ms / trigger /
timestamp. Tasks-tab tooltip on the violet "compacted" badge now
shows real numbers.

9 new unit tests (20 total in the suite): canonical JSONL shape,
since_iso scoping with exclusive and inclusive boundaries, multi-
compact ordering, missing/null compactMetadata defensive handling,
malformed-line robustness, empty-result paths.

Live smoke verified: rebuilt base image, recreated testfix, turn
endpoint returns compact_events:[] for non-compacting turns; the
existing session a2trktlTnXA9KD- already shows compact_count=1
unchanged.

Requires agent-base rebuild + container recreate to deploy.

* fix(credentials): strip auto-injected mcpServers.trinity before validating

Commit b474520 (sec(mcp): structure-validate .mcp.json content) added
RESERVED_SERVER_NAMES = {"trinity"} as a defense against attacker-
controlled redefinition of the auto-injected Trinity MCP server entry.
But the agent server's inject_trinity_mcp_if_configured() writes
mcpServers.trinity into .mcp.json on every agent start where
TRINITY_MCP_API_KEY is set — so the file the user loads in the
credentials editor always contains it, and every save trips the
validator with "MCP server name 'trinity' is reserved by Trinity".

Effect on the user: the credentials Save button silently failed for
every agent that had been started at least once. Confirmed regression
introduced by b474520 — pre-this-branch, .mcp.json went through the
inject endpoint with no content validation and the same auto-injected
trinity entry was accepted.

Fix at the backend: strip mcpServers.trinity from the submitted JSON
before it reaches validate_mcp_config. The agent re-injects the
canonical trinity entry from env vars on next startup, so the user
can't lose it by leaving it out of the saved file. The defense-in-
depth value of the reserved-name rule is preserved — an attacker still
can't substitute a different shape under the trinity name, because
their substitution is dropped before validation rather than rejected.

Side benefit: this self-heals corrupted bearer values in the existing
trinity entry (e.g. the historical "Bearer sqlite3.IntegrityError: ..."
strings written by some long-removed prior code path). The agent
overwrites the entry on next start with a fresh value from env.

11 unit tests cover happy path, mixed user+trinity entries, the exact
historical corruption shape, no-op cases (no trinity / no servers /
no mcpServers), robustness on malformed input (passes through to the
validator's own JSON error), and output-shape preservation.

Verified end-to-end on testfix: POST /api/agents/testfix/credentials/inject
with the corrupted-bearer trinity entry + a context7 entry returns
HTTP 200, backend logs the strip, on-disk .mcp.json contains only
context7 with the corrupted bearer cleanly removed.

Backend reload only — no agent rebuild required.

* fix(credentials): allow canonical trinity entry through validator instead of stripping

Replaces the strip approach from 705de47 (Option B) with a shape-
equivalence check in the validator (Option A) — the strip silently
dropped the user's edits to the trinity entry, which broke the
legitimate "rotate my MCP API key" flow on agents that don't have the
auto-inject env var set to recover the entry on next start.

Symptom 705de47 introduced: opening .mcp.json, editing one digit in
the trinity Bearer token, clicking Save → backend stripped the entry
before validation → on-disk file became {"mcpServers": {}} → user's
trinity MCP entry was lost.

Fix: keep RESERVED_SERVER_NAMES = {"trinity"} as the closed-shape
defense, but special-case it in `_validate_entry`: if the entry under
the trinity name matches the canonical Trinity-MCP shape exactly
(only `type`, `url`, `headers` keys; `type=http`; `url` matches the
configured TRINITY_MCP_URL or the documented default; `headers`
contains only `Authorization` with a `Bearer trinity_mcp_…` token),
accept it. Otherwise hit the existing reserved-name reject.

Strict allowlist on the canonical shape:
  - exact key set: {type, url, headers} — no extras
  - type == "http"
  - url ∈ {TRINITY_MCP_URL env, "http://mcp-server:8080/mcp"}
  - headers has only Authorization
  - Authorization matches /^Bearer\s+trinity_mcp_[A-Za-z0-9_-]{1,200}$/

Attacker scenarios still rejected:
  - stdio redefinition (npx + args) under trinity → reserved
  - http with evil URL (https://evil.com/mcp) → reserved
  - canonical url + extra header (X-Custom) → reserved
  - canonical url + non-Bearer auth → reserved
  - canonical url + Bearer with non-trinity_mcp_ token → reserved

Strip helper and its caller in routers/credentials.py reverted.
test_strip_reserved_trinity.py removed (its scenarios now covered by
the validator's own tests).

10 new unit tests in test_mcp_validator.py cover the canonical-shape
allowance + each rejection variant. 102 total tests in the validator
suite — all green.

Live verified end-to-end on testfix:
  - POST /credentials/inject with edited bearer → 200, file updated
  - POST /credentials/inject with stdio shape → 400 reserved-name
  - User's original bearer restored after the test

* chore(gitignore): exclude saved-conversations/ and tests/manual/ from future stages

Both directories are local-only working artifacts (conversation
transcripts and the hand-driven testing kit) that should never be
pushed. They were already gitignored implicitly by being untracked,
but a future ``git add .`` could inadvertently stage them. Listing
them in .gitignore makes the exclusion explicit and accident-proof.

* feat(session-tab): flip feature flag default to True for GA (#651)

Phase 5.3 GA. Session tab is now exposed to users on every fresh
install without an admin opt-in step.

Resolution order is unchanged:
  1. system_settings row 'session_tab_enabled' if present (admin override)
  2. SESSION_TAB_ENABLED env var (only honored as opt-out: false/0/no)
  3. Default: True

Admins who want to keep it hidden can set
``session_tab_enabled = false`` in system_settings or export
``SESSION_TAB_ENABLED=false``.

Closes the Phase 5.3 step in #651.

* docs: address PR #652 review feedback — requirements, GA status, security section

- Add §5.8 Session Tab entry to requirements.md (Rule 4 — new P1 capability
  must be registered in the single source of truth)
- Flip three stale "default off / Phase 5 rollout pending" references in
  architecture.md, feature-flows.md, and feature-flows/session-tab.md to
  reflect the GA default-on flag flip from PR #651
- Add ## Security Considerations section to feature-flows/session-tab.md
  covering 404-not-403 ownership isolation (E6), Redis lock for concurrent
  --resume (Anthropic #20992), cross-session contamination empirical gate
  (Anthropic #26964), and JSONL prompt-injection persistence with the
  Reset memory mitigation

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
vybe added a commit that referenced this pull request Jun 1, 2026
* refactor(frontend): sweep auth + channels + top-level views (#554) (#609)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554) (#623)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554)

Slice 7 stacked on #609. Migrates 24 files spanning the small-domain pool:

  CHAT (4 files)
    ChatBubble  — copy-success checkmark, self-task accent panel (purple)
    ChatPanel   — agent-not-running warning state, error banner
    ChatInput   — file-remove button, voice-active recording indicator
    ChatHistoryDropdown — error state

  FILE MANAGER (4 files)
    FileManager        — notification toast (success/error), no-agents warning,
                         loading error, delete button + modal + confirm action
    FileTreeNode       — search-matched row highlight (file-type icons stay raw
                         decorative; need their own accent palette later)
    FilePreview        — preview-error icon + text
    FileSharingPanel   — Revoke button

  PROCESS (3 files)
    TrendChart        — completed/failed/cost bars (chart series + legend),
                         success-rate threshold ladder
    RoleMatrix        — no-executor row + badge
    TemplateSelector  — category badges (business/devops/support → status-info /
                         accent-purple / status-urgent)

  CROSS-CUTTING / MODALS (10 files)
    NavBar               — Ops critical-pulse + high indicator, WS connected dot
    GitConflictModal     — yellow warning header (×2), all destructive (red) options
    ReplayTimeline       — system-agent purple panel + badge, schedule-marker arrow,
                            live-feed dot, activity-state success rate ladder
    UnifiedActivityPanel — running/success/fail indicators (live + modal)
    OnboardingChecklist  — completed-state styling (ring, bg, indicator, text)
    CreateAgentModal     — templates-error + general error
    ConfirmDialog        — danger/warning variant icons, text, confirm buttons
    ResourceModal        — (no migrations — amber notice, deferred)
    AvatarGenerateModal  — error text, remove-avatar button
    HelpChatWidget       — error banner + retry button

  MISC (3 files)
    YamlEditor          — error and warning banners + counts + success checkmark
    EditorHelpPanel     — required-field indicator
    TerminalPanelContent — restart-required notice, start-agent button
    TagsEditor          — error message

Net diff: 24 files, +106 / -106 (1:1 palette aliases, byte-identical CSS).

Deferred (existing pattern):
  - Indigo / blue primary action buttons (Login & elsewhere)
  - Blue selected-state (NavBar tabs, OperatingRoom tabs)
  - Amber notices (ResourceModal, RoleMatrix amber missing-role marker —
    these still use the amber palette which differs from yellow)
  - File-type icon colors in FileTreeNode (decorative, need accent-yellow /
    accent-purple-blue / etc. — folder ≠ warning, video ≠ accent, etc.)
  - Slack/Telegram/WhatsApp logo brand colors

These map to the pending `action-primary`, `state-selected`, and accent-color-
expansion tickets.

Verified locally:
  - npm run check:tokens                         passes (10 tokens valid)
  - npm run build                                passes
  - npm run test:e2e:smoke (7 tests, 7.9s)       all green

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep agent surfaces — 22 files (#554) (#614)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554)

Slice 7 stacked on #609. Migrates 24 files spanning the small-domain pool:

  CHAT (4 files)
    ChatBubble  — copy-success checkmark, self-task accent panel (purple)
    ChatPanel   — agent-not-running warning state, error banner
    ChatInput   — file-remove button, voice-active recording indicator
    ChatHistoryDropdown — error state

  FILE MANAGER (4 files)
    FileManager        — notification toast (success/error), no-agents warning,
                         loading error, delete button + modal + confirm action
    FileTreeNode       — search-matched row highlight (file-type icons stay raw
                         decorative; need their own accent palette later)
    FilePreview        — preview-error icon + text
    FileSharingPanel   — Revoke button

  PROCESS (3 files)
    TrendChart        — completed/failed/cost bars (chart series + legend),
                         success-rate threshold ladder
    RoleMatrix        — no-executor row + badge
    TemplateSelector  — category badges (business/devops/support → status-info /
                         accent-purple / status-urgent)

  CROSS-CUTTING / MODALS (10 files)
    NavBar               — Ops critical-pulse + high indicator, WS connected dot
    GitConflictModal     — yellow warning header (×2), all destructive (red) options
    ReplayTimeline       — system-agent purple panel + badge, schedule-marker arrow,
                            live-feed dot, activity-state success rate ladder
    UnifiedActivityPanel — running/success/fail indicators (live + modal)
    OnboardingChecklist  — completed-state styling (ring, bg, indicator, text)
    CreateAgentModal     — templates-error + general error
    ConfirmDialog        — danger/warning variant icons, text, confirm buttons
    ResourceModal        — (no migrations — amber notice, deferred)
    AvatarGenerateModal  — error text, remove-avatar button
    HelpChatWidget       — error banner + retry button

  MISC (3 files)
    YamlEditor          — error and warning banners + counts + success checkmark
    EditorHelpPanel     — required-field indicator
    TerminalPanelContent — restart-required notice, start-agent button
    TagsEditor          — error message

Net diff: 24 files, +106 / -106 (1:1 palette aliases, byte-identical CSS).

Deferred (existing pattern):
  - Indigo / blue primary action buttons (Login & elsewhere)
  - Blue selected-state (NavBar tabs, OperatingRoom tabs)
  - Amber notices (ResourceModal, RoleMatrix amber missing-role marker —
    these still use the amber palette which differs from yellow)
  - File-type icon colors in FileTreeNode (decorative, need accent-yellow /
    accent-purple-blue / etc. — folder ≠ warning, video ≠ accent, etc.)
  - Slack/Telegram/WhatsApp logo brand colors

These map to the pending `action-primary`, `state-selected`, and accent-color-
expansion tickets.

Verified locally:
  - npm run check:tokens                         passes (10 tokens valid)
  - npm run build                                passes
  - npm run test:e2e:smoke (7 tests, 7.9s)       all green

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep agent surfaces — 22 files (#554)

Slice 8 of the design-system migration. Replaces semantic palette refs
with status/state/accent tokens across all per-agent panels, the
agents list, and agent-detail surfaces.

Files migrated:
- 16 panels (Tasks, Git, Dashboard, Playbooks, Nevermined, Info,
  Credentials, Schedules, SystemViewEditor, HostTelemetry, Folders,
  Files, Skills, Observability, Permissions, Metrics)
- 5 agent surfaces (Agents, AgentNode, AgentHeader, AgentTerminal,
  AgentDetail)
- SystemAgentNode

Mappings (palette-equivalent, no visual change):
- yellow → status-warning
- green  → status-success
- red    → status-danger
- orange → status-urgent
- amber  → state-autonomous (token name slightly stretched for tool-call
  / queued-task amber; same palette, future cleanup may rename)
- rose   → state-locked
- purple → accent-purple

Deferred (no token family yet):
- indigo (action-primary)
- blue/sky/cyan (selected-state, category labels)
- teal (category labels)

SystemViewsSidebar untouched — only contains deferred blue/indigo
selected-state references.

Verification:
- npm run check:tokens → 10 tokens equivalent, all references resolve
- npm run build → clean
- npm run test:e2e:smoke → 7/7 passed against live Trinity (HMR)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(agent-runtime): kill npx MCP orphans outside claude pgid that hold stdout pipe open (#618) (#620)

* fix(voice): orb animation loop dies before voice session starts

The canvas is inside v-if="voice.isActive.value" so canvasEl.value is
null when onMounted fires. renderFrame() exits early without scheduling
the next frame, killing the loop permanently.

Replace onMounted initialization with watch(canvasEl) so the RAF loop
starts when the canvas enters the DOM and stops when it leaves.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent-runtime): kill npx MCP orphans outside claude pgid that hold stdout pipe open (#618)

After terminate_process_group kills claude's pgid, npm→node MCP server chains
spawned via npx call setsid() and land in a new session — they survive the
pgid kill and keep the stdout pipe write FD open indefinitely.  The kernel
cannot deliver EOF to our reader thread while any writer FD remains open, so
drain_reader_threads blocked for the full 30s post_kill_grace, then lost the
buffered result line via force-close (HTTP 502).

Add _kill_orphan_pipe_writers(): after terminate_process_group, scan /proc/*/fd
for any process outside our pgid that holds the pipe's write end (detected via
fdinfo flags), and SIGKILL it.  Killing the orphan releases all its FDs
(stdout AND stderr write ends) simultaneously, delivering EOF to both reader
threads so they drain naturally before the post_kill_grace window.

New tests (Linux-only, skipped on macOS — /proc required): verify that a
setsid() grandchild is detected and killed, that our own read-end process is
not touched, and that the end-to-end drain path preserves buffered data.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(slack): replace slackify-markdown with own renderer (#293) (#622)

Replaces slackify-markdown with a custom renderer fixing 5 compounding bugs (nested lists, headings, blockquotes, tables, horizontal rules). Includes 35 unit tests and updated feature flow doc.

Closes #293

Co-Authored-By: pavshulin <pavshulin@users.noreply.github.com>

* feat(agents): per-agent token usage display in AgentHeader (#250) (#632)

Adds a token usage row to AgentHeader showing 7-day cost sparkline,
today's cost vs 7-day daily average (with trend arrow), and lifetime
totals. Data sourced from schedule_executions in the DB so it persists
across agent restarts.

- New GET /api/agents/{name}/token-stats endpoint
- ScheduleOperations.get_agent_token_stats(): single-pass 24h/7d/lifetime
  aggregation + 7-day daily breakdown with gap-filling
- agentsStore.getAgentTokenStats() action
- TOKEN USAGE ROW in AgentHeader.vue: SparklineChart (amber, 56x16),
  trend indicator (warning/success/gray), lifetime summary
- Hidden for agents with no runs (lifetime_executions == 0)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(site): agent website proxy via /site/{token} endpoint (SITE-001) (#634)

Adds live HTTP reverse-proxy so agents can serve public websites from
their container. A new `type='site'` public link routes requests through
`GET /site/{token}/{path}` → httpx streaming proxy → agent web server at
`http://agent-{name}:3000`. Includes DB migration, nginx routing, rate
limiting (per-IP + per-token), SSRF guard, security header stripping, and
audit event `site_link_visit`. UI adds Chat/Website selector in the link
create modal with a "Website" badge on site links.

Fixes #633

Co-authored-by: Claude <noreply@anthropic.com>

* fix(site): centralize SITE_PORT, atomic rate limit, fire-and-forget audit log, update docs (SITE-001)

- Move SITE_PORT to config.py; import in site.py and public_links.py
- Fix TOCTOU race in _check_site_rate_limit: pipeline INCR+check-after
- Audit log is now asyncio.create_task() so streaming is not delayed
- Add SITE-001 to requirements.md (section 15.1a-3)
- Add site.py to architecture.md router listing + /site/ endpoint table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent-runtime): bound _kill_orphan_pipe_writers to 10s to prevent drain stall (#649) (#650)

/proc scanning can block indefinitely when a process is in D state
(uninterruptible sleep), causing drain_reader_threads to stall for
tens of minutes instead of the expected ~30 seconds.

Run _kill_orphan_pipe_writers in a daemon thread with a 10s cap so
a blocked /proc entry cannot push the drain past its deadline.

Also use wall-clock accounting for the post-kill join timeout so time
spent in terminate + orphan scan doesn't silently erode the budget,
and log actual elapsed time instead of the expected value so future
incidents are easier to diagnose.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): voice WebSocket + stop endpoint missing ownership check (#600) (#638)

The /ws/voice/{voice_session_id} handler decoded the JWT but threw the
payload away — only the signature was checked. Any authenticated user
holding a valid JWT who learned the 128-bit session id (logs, browser
inspection, XSS) could attach to the audio stream, eavesdrop on the
victim's transcript, and trigger tool calls audit-logged under the
victim's identity.

POST /api/agents/{name}/voice/stop had the same gap: the path agent was
gated via get_authorized_agent, but request.voice_session_id was never
cross-checked against the path agent or the caller's user_id, so the
caller could end and persist a transcript onto another user's session.

Fix:
- WS: extract sub from the decoded JWT, look up the user, and close 4003
  if user.id != session.user_id (admin role bypasses, for support).
- voice_stop: load the session via get_session before mutating, assert
  agent_name == path name AND user_id == current_user.id (admin bypasses),
  raise 403 otherwise.

Added tests/unit/test_voice_auth.py covering: missing token, invalid
token, missing sub claim, unknown user, owner happy path, admin bypass,
attacker rejected, plus voice_stop variants. Loads voice.py via
importlib to avoid pulling in the full routers/__init__.py chain.

Reported by /security-review on PR #599 (2026-04-30); origin commit
7d8abe8 (#581 voice tool calls).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): split login rate limit into per-account + per-IP buckets (#591) (#621)

Pentest finding AISEC-H2 (CVSS 7.5, CWE-307): the previous design used a
single per-IP bucket at 5 fails / 10 min. Any user behind a corporate
NAT, VPN, or CDN locked out everyone else at the same egress IP after
just four bad attempts. A rotating-proxy attacker could keep an
organisation locked out continuously, so the protection doubled as a
platform-wide DoS primitive.

Replace with two independent buckets:

  * Per-account (tight) — 5 fails / 15 min: limits credential stuffing
    on one targeted account; never affects other accounts.
  * Per-IP (loose)      — 30 fails / 5 min: catches single-source abuse
    but stays well above the legitimate-traffic threshold for users
    sharing a NAT/VPN/CDN egress.

Both buckets are checked on every attempt; 429 fires when either is
exhausted. Successful login clears both. Account names are normalised
(lowercase + strip) before keying. Endpoints without an account context
(public access-request) skip the per-account bucket and rely on the
per-IP one only.

Lockout state-changes log a structured WARNING (visible via Vector) so
operators can see when buckets are being exercised.

Live verification on the running backend:
  attempts 1-5 → 401 (counter ticking)
  attempt  6   → 429 "Too many failed attempts for this account..."
  valid pwd    → 429 (account stays locked even with right password)
  other account from same IP → 401 (per-account isolation works)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): require creator role on /api/systems/deploy (#592) (#624)

Auditing all entry points that ultimately call `create_agent_internal`
(per #592 AC #2) turned up a real bypass on `POST /api/systems/deploy`.
The system-manifest deployment route gated on `Depends(get_current_user)`
without a role check, so any authenticated user-role account could spawn
an entire fleet of agents through that path — a strictly stronger
privilege than the single-agent bypass the AISEC-H1 finding originally
named (which `POST /api/agents/deploy-local` had already closed via #150).

Add `Depends(require_role("creator"))` to `deploy_system`, matching the
existing dependencies on `POST /api/agents` and `POST /api/agents/deploy-local`.

Regression test (`tests/unit/test_agent_creation_role_gates.py`) walks
the FastAPI router source AST and asserts that every agent-creation
route uses `Depends(require_role("creator"))`. AST-level so the check is
fast, stable across formatting changes, and fires the moment someone
removes the dependency. Confirmed the test catches the regression by
reverting the change and observing the test fail before re-applying.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): stop mirroring JWT into document.cookie (#188) (#642)

* fix(security): stop mirroring JWT into document.cookie (#188)

UnderDefense pentest 3.3.5 flagged the frontend for mirroring the
authentication token into a `token` cookie without `Secure` or
`HttpOnly` flags. The cookie was set in setupAxiosAuth via
`document.cookie =` so it was readable from JS (HttpOnly is impossible
on JS-set cookies), transmitted over HTTP without the Secure flag, and
auto-attached to every outbound request as a CSRF vector.

The cookie's stated purpose was "for nginx auth_request to validate
agent UI access" — but that nginx directive was never configured in
any committed deployment (`grep -r auth_request -- *.conf` is empty,
git log -S confirms it never existed). The cookie was pure attack
surface with zero functional value.

Per the issue's "Best" remediation, drop the cookie mirror entirely.
API authentication uses the `Authorization: Bearer` header
exclusively; nothing else needs the cookie.

The cookie-clear on logout is intentionally kept so users carrying a
stale cookie from the pre-fix version get cleaned up on their next
logout cycle. The cookie's `max-age=1800` also naturally expires it
within 30 minutes of the upgrade.

The backend's `/api/auth/validate` endpoint still accepts a cookie as
one of three token sources — left untouched as out-of-scope. With the
frontend no longer setting the cookie nothing legitimate sends one,
but the fallback path remains available if a future nginx
auth_request setup is wired up properly (with Secure + HttpOnly flags
set server-side via Set-Cookie, not via document.cookie).

Closes #188.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(admin-login): remove stale references to JWT cookie mirror (#188)

PR #642 removed the document.cookie set in setupAxiosAuth, but the
admin-login feature flow still documented the cookie as live. Update
the code snippet and the storage table to reflect current behaviour.

Note in the snippet describes why the cookie was removed so readers
who see the diff history can find the rationale without reading the
PR.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): override git User-Agent on skills library sync (#184) (#646)

UnderDefense pentest 3.3.1 flagged the backend for leaking the
underlying tech stack via outbound User-Agent. Skills sync uses git
subprocess (not httpx), so the leaked UA is `git/<version>
(libcurl/<version> ...)` — verified live with GIT_TRACE_CURL=1.

Add `-c http.useragent=Trinity-Skills-Sync` (positioned correctly
before the subcommand, as git's `-c` requires) to the two HTTP-bearing
git invocations: `_git_clone` and `_git_pull`'s fetch. The local-only
`git reset --hard` and `git rev-parse HEAD` calls intentionally do
not get the flag — they make no HTTP and threading the flag through
would suggest otherwise.

The SSRF allowlist (#179) already locks the destination to github.com
so the practical exposure is small (GitHub already knows what we are),
but defense-in-depth: even if the allowlist is ever loosened the UA
stays generic.

The constant has no version suffix to avoid yet another version string
drifting against VERSION / package.json / pyproject.

Tests in tests/unit/test_skill_service_user_agent.py mock subprocess
and assert the flag is present at the right argv position for clone
and fetch, and absent for the local-only reset and rev-parse calls.

Live verification with `GIT_TRACE_CURL=1 git -c http.useragent=... ls-remote ...`
confirms the wire UA changes from `git/2.43.0` to `Trinity-Skills-Sync`.

Closes #184.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(deploy): document stale-image symptom + recovery in start.sh and DEPLOYMENT.md (#557) (#626)

Self-hosted developers tracking `dev` occasionally pull a commit that
adds a new Python or Node dependency to one of the platform Dockerfiles,
re-run `start.sh`, and end up with new source running against an old
image's Python env. Uvicorn crashes with `ModuleNotFoundError`, compose
keeps respawning the worker, and `start.sh` reports success — leaving
the UI "Disconnected" with no obvious diagnosis.

Adopting Option B from #557 discussion (PR #625 closed): treat this as
a documentation problem rather than building auto-detection into the
critical-path startup script. Auto-detection has clear cost (Python
subprocess + Docker inspect on every cold start) and unclear benefit
(production deploys use `compose pull` and don't hit this; the affected
population is self-hosted devs whose recovery is one command).

Two changes:
- `scripts/deploy/start.sh`: append a 4-line hint after the "Ready!"
  banner naming the symptom (`ModuleNotFoundError`, "Disconnected" UI)
  and the exact recovery command.
- `docs/DEPLOYMENT.md`: add a Troubleshooting entry with full diagnosis
  walkthrough, root-cause explanation, and the rationale for not
  auto-detecting (links to #557).

A future `scripts/deploy/upgrade.sh` is the right place to bundle
backup + rebuild + start + verify for the explicit upgrade path; that
is bigger-than-#557 scope.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): lock down Redis — auth + ACL + network split (#589)

* fix(security): split Docker compose into platform and agent networks (#589)

Redis at 172.28.0.0/16 was reachable from any agent container. AISEC scan
3aad5469 demonstrated end-to-end exfiltration / cross-user task injection
from a legitimately deployed agent. Network segmentation is the strongest
control — agents now physically cannot route to Redis.

Topology:
- trinity-platform (172.29.0.0/16, NEW) — Redis, scheduler, vector
- trinity-agent (172.28.0.0/16, name preserved) — frontend, agents
- Backend / mcp-server / otel-collector / cloudflared straddle both

Agent-creation sites in services/agent_service/* and system_agent_service.py
need zero changes because the agent-network external name is preserved.

Dev: bind Redis host port to 127.0.0.1:6379 (was 0.0.0.0). Tests connect
from the dev machine; LAN cannot. Auth lands in the next commit.

Refs #589 — acceptance criterion #3 (network segment separation).

* fix(security): mandatory Redis auth, ACL users, auth-aware healthcheck (#589)

Both compose files now enforce two passwords (REDIS_PASSWORD admin /
REDIS_BACKEND_PASSWORD runtime) with the fail-on-missing :? form.
docker compose refuses to render without them.

Per-user ACL via inline --user flags. Additive (start from zero, allow
only what the runtime needs) — never +@all -X, which lets newly added
dangerous commands through. backend + scheduler get standard data
families plus scripting/transactions/pubsub minus -@dangerous, which
covers FLUSHALL, CONFIG, SHUTDOWN, MIGRATE, REPLICAOF, MONITOR.

Verified at runtime against redis:7-alpine: PING/SET/GET work for the
backend user, FLUSHALL and CONFIG GET return NOPERM, unauth requests
return NOAUTH.

REDIS_URL on backend + scheduler now embeds the backend ACL user.
mcp-server: REDIS_URL and depends_on:redis dropped in prod compose
(zero Redis imports in src/mcp-server/).

Healthcheck pings as the backend ACL user so a typo'd ACL keeps redis
unhealthy and gates dependent services. depends_on:redis switches to
service_healthy so backend/scheduler don't race the ACL load.

Refs #589 — acceptance criteria #1, #2, #5.

* fix(scheduler-test-rig): mirror Redis auth posture (#589)

Without this, scheduler container fails fast on startup against the rig
because src/scheduler/config.py requires creds in REDIS_URL after #589.
No ACL or network split here — this is a 2-service standalone debugging
rig, not the production posture.

* fix(security): fail-fast on REDIS_URL missing credentials (#589)

Backend (src/backend/config.py) and scheduler (src/scheduler/config.py)
now raise RuntimeError at import time if REDIS_URL is unset or lacks
credentials.

Removed the splicing fallback in backend config that papered over an
unauth REDIS_URL by joining REDIS_PASSWORD into the URL — single source
of truth (compose) eliminates silent drift.

Tests that import backend modules need a creds-bearing REDIS_URL in their
environment; tests/conftest.py will set a dummy one in the test commit.

Refs #589 — acceptance criterion #5.

* fix(webhooks): use REDIS_URL for rate-limit client (#589)

Webhooks rate-limit was the one Redis client that bypassed REDIS_URL —
it used redis.Redis(host="redis", port=6379) and would silently
fail-open under requirepass. Switching to redis.from_url(REDIS_URL)
picks up the credentialed URL like every other client.

Also: distinguish auth/ACL errors (logged at ERROR with exception class)
from transient errors (WARN). Fail-open behavior preserved so a Redis
blip doesn't 500 legitimate webhooks, but a misconfigured deploy now
surfaces in alerts instead of via a webhook abuse incident.

Drops the now-unused REDIS_HOST/REDIS_PORT env reads.

* feat(deploy): auto-generate Redis passwords on fresh installs (#589)

start.sh ensure_redis_passwords matches the existing
CREDENTIAL_ENCRYPTION_KEY pattern, with one safety guard:

- Fresh install (no redis-data volume) → generate both passwords with
  openssl rand -hex 24 and append to .env. One-command boot keeps
  working.
- Existing volume + missing password → refuse with a loud error pointing
  at docs/migrations/REDIS_AUTH.md. Re-keying a populated Redis would
  lock the backend out of its own data; ops needs to follow the explicit
  upgrade path.

Idempotent — second run is a no-op when both passwords are already set.

* docs(security): add Redis auth migration guide + architecture notes (#589)

- docs/migrations/REDIS_AUTH.md: operator upgrade guide. Covers fresh
  installs (auto-generated by start.sh), live upgrades (down
  --remove-orphans + docker network rm + add passwords), production,
  and verification commands.
- docs/memory/architecture.md: new "Network Topology (Issue #589)"
  section above Container Security. Documents the two-network split,
  service membership table, the "agents NEVER on platform network"
  rule, and the three Redis ACL users + their access patterns.

* test(security): network isolation, ACL, fail-fast, webhook rate-limit (#589)

tests/conftest.py: top-level autouse env stub for backend imports.
Backend config now raises at import-time if REDIS_URL lacks credentials;
without this, every test that transitively imports backend modules
breaks. Real Redis tests under tests/security/ override via their own
conftest from .env. Adds the `integration` marker.

tests/unit/test_config_fail_fast.py (new): backend refuses to import
without creds-bearing REDIS_URL. 3 cases — missing env, unauth URL,
URL with creds.

tests/security/test_redis_network_isolation.py (new): 5 integration
tests covering acceptance criteria #1-#3:
  - agent-network container has no route to redis (BLOCKED)
  - unauth client gets NOAUTH on platform network
  - backend ACL user can PING with creds
  - backend ACL user FLUSHALL → NOPERM (no admin)
  - backend ACL user CONFIG GET → NOPERM (no requirepass leak)

tests/security/conftest.py (new): session-scoped fixture loads real
.env values for the integration tests; skips the suite if missing.

tests/integration/test_webhook_rate_limit.py (new): regression for the
from_url switch in webhooks.py. Self-contained — creates agent +
schedule + webhook token inline, hits 11×, expects 429 on the 11th.
Catches the silent fail-open if Redis auth ever regresses.

tests/run-integration.sh (new): pytest -m integration runner. Excluded
from run-smoke.sh per the smoke runner's ~30s no-Docker contract.

* docs(security): detach agents before network rm (#589)

Trinity-managed agent containers are created via the Docker SDK
outside compose, so they store the agent network's UUID, not its
name. After `docker network rm trinity-agent-network` (step 3 of
the upgrade procedure), any later `docker start <agent>` fails:

    Error response from daemon: failed to set up container
    networking: network <old-uuid> not found

Compose-managed services don't hit this — they're recreated with
fresh network refs on `up`. Agent containers aren't, so they keep
the stale UUID until disconnected.

Add an explicit detach loop as step 2, before the network removal.
Verified against a populated install with one running and four
stopped agents: all five reattach cleanly to the new network on
next start.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): CSO OBS-1/2/3 follow-ups — webhook rate-limit + healthcheck hardening (#589)

Resolves three observations from the CSO audit
(docs/security-reports/cso-2026-05-04-589-diff.md):

OBS-1 — webhook rate-limit fail-open + connection-per-request DoS amplifier:
* Added in-process secondary rate limiter (3x primary, per-worker) in
  src/backend/routers/webhooks.py. Bounds blast radius during a Redis
  outage without breaking the documented fail-open philosophy.
* Cached the Redis client at module level under threading.Lock with
  double-checked init. _check_webhook_rate_limit resets the cache on
  inner exceptions so stale connections rebuild cleanly. Without
  caching, a flood would open a fresh TCP per request and exhaust
  Redis maxclients — turning the rate limiter into the DoS amplifier.

OBS-2 — tightened _TOKEN_RE from {20,60} to {43} matching
secrets.token_urlsafe(32) (verified against db/schedules.py:524).

OBS-3 — switched all three compose healthchecks from
`redis-cli -a $$PASS` to `REDISCLI_AUTH="$$PASS" redis-cli` so the
password no longer appears in /proc/<pid>/cmdline.

Additional #589 hardening (caught while resolving OBS-1):
* src/backend/config.py + src/scheduler/config.py: tightened the
  REDIS_URL credential check from `"@" in url` substring to urlparse
  validation. Catches redis://@redis:6379, redis://user@redis:6379, etc.
* src/scheduler/main.py: redact password from REDIS_URL before logging
  (was leaking via Vector log aggregator).

Tests:
* tests/unit/test_webhook_rate_limit_inprocess.py — 7 new tests covering
  cap, window expiry, token isolation, runtime-error fallback, regex
  shape, cache hit, cache reset.
* tests/unit/test_config_fail_fast.py — 4 new parametrized cases for
  malformed-credential URL rejection.
* 15/15 unit tests pass.
* Live Redis healthcheck verified — trinity-redis reports healthy with
  the new REDISCLI_AUTH form; `redis-cli ping` returns PONG.

Also adds .gstack/ to .gitignore so future skill artifacts stay local.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(dev): update gitea overlay network name after #589 split

trinity-network no longer exists; gitea dev overlay must attach to
trinity-agent-network (the preserved agent-network name).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tests): use credentialed Redis URL in scheduler test_config fixture (#661)

After #589 hardened Redis auth, SchedulerConfig raises on bare redis:// URLs.
The test_config fixture bypassed the env-level patch in tests/conftest.py by
passing redis_url="redis://localhost:6379" directly.

Fixes #659

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* Session tab — `--resume`-default chat surface (Closes #651) (#652)

* docs(planning): add Session tab design — --resume-default chat surface

Adds docs/planning/SESSION_TAB_2026-04.md, the comprehensive plan for a
new "Session" tab living alongside Chat. Sessions reattach to their own
Claude Code JSONL via --resume, preserving tool memory, mid-skill state,
and reasoning state across turns.

Plan covers:
- UI design (tab placement, multi-session model, +New Session, Reset memory)
- Data model (agent_sessions / agent_session_messages — parallel to chat)
- Backend architecture (separate router, single shared change to
  task_execution_service for persist_session plumbing)
- Phased rollout (foundation → backend → frontend → hardening → GA)
- Edge cases & failure-mode lessons baked in from a prior local spike
  (parser bug, --no-session-persistence dependency, cold-turn detection,
  port allocation)
- Test plan including the cross-session contamination test for
  Anthropic claude-code#26964
- Retention/cleanup policy, observability, security checklist
- Local-first workflow: implementation runs entirely on this branch
  until validation passes; only then does the standard SDLC engage
  (issue, push, PR)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add agent_sessions + agent_session_messages tables

Phase 1.1 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).
Schema definitions go in db/schema.py for fresh installs; the matching
idempotent migration agent_sessions_tables in db/migrations.py upgrades
existing databases.

The schema mirrors chat_sessions / chat_messages but is strictly parallel
— no foreign keys, no shared columns, separate index namespace. Three
fields are unique to the session model:

- agent_sessions.cached_claude_session_id — the Claude Code session UUID
  the next turn will pass to ``--resume``
- agent_sessions.consecutive_resume_failures — drives the resume-failure
  fallback (Phase 2.2)
- agent_session_messages.cache_read_tokens — observability for whether
  Anthropic's prompt cache engaged

CASCADE on session delete cleans up message rows automatically.

Verified locally: backend restart applies the migration cleanly, tables
have 15 columns each with correct types/defaults/PKs, all four indexes
created, second restart confirms idempotency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add SessionOperations for Session tab persistence

Phase 1.2 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).

- Adds AgentSession and AgentSessionMessage Pydantic models in db_models.py
  with the new fields the Session tab needs beyond ChatSession/ChatMessage:
  cached_claude_session_id, last_resume_at, consecutive_resume_failures on
  the session row, and cache_read_tokens + claude_session_id on each message.

- Creates db/sessions.py with a SessionOperations class mirroring the
  ChatOperations shape: create_session, get_session, list_sessions,
  delete_session, add_session_message, get_session_messages, plus the
  Claude UUID cache helpers (get/update/clear_cached_claude_session_id)
  and resume health helpers (mark_resume_failure, mark_resume_success).

- Wires the new ops into the DatabaseManager facade alongside the
  existing _chat_ops, with one delegating method per public operation.

No router, no agent-server change, no frontend yet — those land in later
phases. Tables agent_sessions and agent_session_messages already exist
from the prior schema commit.

* feat(session-tab): backend foundation for --resume-default Session surface

Phases 1.3 through 1.7 of the Session tab plan
(docs/planning/SESSION_TAB_2026-04.md). Pure backend / agent-server work
behind a flag — no UI surface yet, no behavior change to Chat or any
existing /task caller.

Agent server (base image):

- Stream-json parser fix (Appendix B). Both parse_stream_json_output and
  process_stream_line now recognize {"type":"system","subtype":"init"}
  for session_id capture, with the result event as a fallback when init
  was missed (truncated streams). The legacy bare-init shape is
  intentionally rejected. This is the same bug that would have made
  Session caching corrupt on every cold turn.

- Same bug in execute_headless_task's permission-mode validation site:
  the check matched the wrong shape, so permission_mode_validated never
  flipped to True and the protective kill-on-misconfigured-permission
  path silently failed open. Now uses type=system + subtype=init.

- New persist_session flag threaded through ParallelTaskRequest →
  routers/chat.py → AgentRuntime ABC → ClaudeCodeRuntime.execute_headless
  → execute_headless_task. When True, --no-session-persistence is
  omitted so the JSONL is written and the next turn's --resume can find
  it. --session-id is still passed for unique cold-turn namespace.
  Default False keeps every existing caller stateless.

- gemini_runtime accepts the parameter for ABC parity and ignores it
  (Gemini CLI has no resume).

Backend:

- task_execution_service.execute_task now accepts persist_session: bool
  = False and threads it into the agent payload. All existing callers
  (Chat, schedules, MCP, fan-out, webhooks) keep today's behavior; only
  the future routers/sessions.py (Phase 2) opts in.

- settings_service.is_session_tab_enabled() — feature flag resolving
  system_settings.session_tab_enabled → SESSION_TAB_ENABLED env →
  False. Module-level convenience function exposed.

Tests (run inside trinity-backend container — Python 3.11):

- tests/unit/test_session_operations.py — 9 tests against an isolated
  SQLite DB exercising the full SessionOperations CRUD plus the cached
  claude session UUID lifecycle and resume failure / success counters.

- tests/unit/test_claude_code_session_id_parser.py — 8 tests covering
  both parsers (batch + streaming): system/init recognition, result
  fallback, init-wins-over-result, legacy bare-init rejection, and a
  source-level regression guard for the permission-mode validation
  fix.

- tests/unit/test_session_persistence_flag.py — 8 tests pinning the
  contract: signatures across the runtime ABC, ParallelTaskRequest,
  agent chat router, execute_headless_task, and
  task_execution_service.execute_task. Includes the gating regex check
  on --no-session-persistence and a live signature import to catch
  drift AST parsing alone would miss.

Total: 25 passing tests covering every touchpoint of Phase 1.

Base image (trinity-agent-base) rebuilt to embed the agent-server
changes; existing agent containers will pick them up on next recreate.

* feat(session-tab): backend turn endpoint for --resume-default Session surface

Phase 2 of docs/planning/SESSION_TAB_2026-04.md. Six endpoints under
/api/agents/{name}/session{s,...} that mirror routers/chat.py's auth
model and TaskExecutionService usage but persist to the parallel
agent_sessions / agent_session_messages tables and request
persist_session=True on every turn so each call reattaches via
`claude --print --resume <uuid>`.

Surface gated on is_session_tab_enabled() — flag-off default returns
404 from every endpoint.

  POST   /api/agents/{name}/session                  create row
  GET    /api/agents/{name}/sessions                 list (per-user)
  GET    /api/agents/{name}/sessions/{id}            session + messages
  POST   /api/agents/{name}/sessions/{id}/message    THE turn
  POST   /api/agents/{name}/sessions/{id}/reset      clear cached uuid
  DELETE /api/agents/{name}/sessions/{id}            delete row + msgs

Spike-pitfall defenses baked into the turn endpoint:

- L3 (first-turn-has-no-session-id): the agent_sessions row is created
  server-side via POST /session BEFORE the turn endpoint ever calls
  execute_task. No frontend-first model.
- L2 (cold turn writes empty JSONL): persist_session=True is passed
  unconditionally — Phase 1.4 already wired the flag through the agent
  stack; Phase 2 just promises to set it on every turn.
- L1 (parser misses system/init): trust result.session_id directly —
  Phase 1.3 fixed the parser. Scenario A confirms the captured UUID is
  the real Claude UUID end-to-end.

Phase 2.2 resume-failure fallback: when execute_task returns "no
conversation found" on a turn that had a cached UUID, clear the cache,
mark_resume_failure, and retry once with resume_session_id=None. Logs
event=session_resume_fallback with the stale UUID and consecutive
failure count. Anthropic #39667 (cleanupPeriodDays) and #53417 (CLI
upgrade) both produce this signal.

Phase 2.3 Redis lock: SET NX EX per (agent, claude_uuid) with 5-min TTL
and Lua-script release. Async poll loop (250ms tick) so the event loop
stays free during contention. Cold turns skip the lock (no JSONL to
corrupt). Hard 30s wait ceiling — beyond that the contender gets HTTP
429 with retry hint. Mitigation for Anthropic #20992 (concurrent
--resume JSONL writes corrupt the file).

Per-user ownership at the row layer: even agent owners cannot read or
send into another user's session (E6 isolation in the design doc).
Returns 404 for ownership failures so we don't leak session-id existence.

Tests (tests/integration/test_session_turns.py, run inside
trinity-backend container with docker.sock mounted for testfix
recreation + JSONL surgery in Scenario C):

  Scenario A: 3-turn happy path — same Claude UUID across turns
  Scenario B: turn 2 recalls a secret from turn 1, no text-replay
  Scenario C: JSONL deletion mid-session triggers fallback + recovery
  Scenario D: concurrent POSTs serialise via Redis lock
              (asserts finish_gap ≈ winner_work_time, NOT total wall)
  Scenario E: switching sessions A → B → A preserves A's UUID

5 passed in 54.5s against the live agent-testfix container (recreated
onto the rebuilt base image first per L4 in the plan). Phase 1's 25
unit tests still pass — no regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): frontend Session surface

Phase 3 of docs/planning/SESSION_TAB_2026-04.md. Adds the new "Session"
tab in AgentDetail, gated on the is_session_tab_enabled() platform flag
so it stays invisible until explicit opt-in (default off).

Backend prerequisite — routers/settings.py:

- GET /api/settings/feature-flags exposes a curated allowlist of UI-
  relevant flags to any authed user. The existing /api/settings/{key}
  endpoint is admin-only and would block non-admin frontends from even
  knowing whether to render the Session tab. The new endpoint reads
  through services.settings_service.is_session_tab_enabled() so the
  resolution order (DB → env → False) stays in one place.

Frontend:

- src/frontend/src/stores/sessions.js — Pinia store wrapping the six
  /api/agents/{name}/sessions* endpoints with per-agent state isolation
  and the feature-flag cache. Optimistic user-message insert with
  rollback on send failure.

- src/frontend/src/components/SessionPanel.vue — structural copy of
  ChatPanel reusing ChatMessages + ChatInput + ModelSelector. Differs
  from Chat in three places per the design doc:
    * Sends bare user_message to POST .../sessions/{id}/message — no
      buildContextPrompt text-replay (the agent already has working
      memory via --resume).
    * "Reset memory" button + confirm modal that clears the cached
      Claude UUID without deleting the message log (Phase 3.4).
    * Per-session selector subtitle: turn count, context % used,
      cached-memory dot (emerald/gray), and consecutive_resume_failures
      indicator (Phase 3.5).
  Lean cut for first-visible-surface: voice mic, file upload, and SSE
  dynamic status labels are deferred — those need backend extensions
  (file payload on the turn endpoint, async_mode + SSE on the same).

- src/frontend/src/views/AgentDetail.vue — new Session tab inserted
  between Chat and Dashboard/Schedules, gated on
  sessionsStore.sessionTabEnabled. Layout sites that previously
  branched on activeTab === 'chat' now use a shared isFullscreenTab
  computed so Chat and Session both get the input-pinned-to-bottom flex
  layout. ?tab=session deep-link allowlist updated.

- src/frontend/e2e/session-tab.spec.js — Phase 3.6 Playwright spec.
  Marked @interactive (not @smoke) because each run makes one real
  Claude API call (~10–60s). Snapshots the prior flag value in
  beforeAll, force-enables for the run, restores in afterAll so a
  failed run doesn't leave the platform with the flag dirty. Three
  cases:
    * tab is hidden when flag is off
    * tab appears, "+ New Session" → send turn → reply visible →
      Reset memory modal opens + closes
    * Chat tab still works after Session interaction; switching back
      preserves Session state

Visually verified in the live dev server: tab renders in correct
position, header layout matches Chat's structure, empty state and
placeholder copy match the design doc, "Reset memory" only shown when
an active session exists, full-viewport flex layout pins input to
bottom.

Phase 1 + Phase 2 work behind this change is unchanged: 25 unit tests
+ 5 integration tests still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): hardening + observability — cleanup service, contamination gate, docs

Phase 4 of docs/planning/SESSION_TAB_2026-04.md. Closes the JSONL
disk-growth loop, validates the GA-blocking cross-session contamination
hypothesis empirically, and lands architecture.md / feature-flows
documentation so the surface is discoverable.

Phase 4.3 — cross-session contamination GA gate (the load-bearing one):

- tests/integration/test_session_cross_contamination.py exercises the
  Anthropic #26964 hypothesis end-to-end. Plants a randomly-generated
  secret token in session A with explicit "do not echo" framing, asks
  session B (different UUID, same agent, same cwd) to recall the token.
  Hard-fails if the exact token leaks; soft-fails on partial-prefix
  recall (PURPLE-DRAGON without the random suffix would only be
  knowable from A's JSONL, not from training).
- PASSED in 9.5s on the current Claude Code version → shared-cwd model
  is safe → Phase 5 rollout unblocked. Test stays in the suite as the
  per-version regression guard.

Phase 4.2 — JSONL cleanup service:

- services/session_cleanup_service.py runs a 6h periodic sweep that
  diffs every running agent's
  ~/.claude/projects/-home-developer/<uuid>.jsonl set against
  db.list_active_claude_session_ids(agent) and reaps orphans whose
  mtime is older than the 1h race guard. Race guard prevents the
  cold-turn-vs-cleanup window where a brand-new JSONL exists on disk
  before the backend has updated cached_claude_session_id.
- Same service exposes a synchronous reap_jsonl(agent, uuid) helper
  called best-effort from routers/sessions.py reset/delete handlers so
  the user-perceived disk-reclaim latency is sub-second. Never raises;
  failures are logged and the periodic sweep is the safety net.
- Implementation uses execute_command_in_container — the same primitive
  git_service / ssh_service / scheduler pre-check / agent terminal use.
  No new agent-server endpoint, no base-image rebuild.
- New db.list_active_claude_session_ids(agent) facade method backed by
  SessionOperations.list_active_claude_session_ids querying every
  agent_sessions row whose cached_claude_session_id is non-null for the
  agent.
- main.py wires startup (staggered +7.5s after cleanup_service to
  offset Docker hits) and clean shutdown.
- tests/integration/test_session_cleanup.py: reset reaps synchronously,
  delete reaps synchronously, periodic sweep keeps the active JSONL,
  reaps an aged orphan, respects the 1h race guard for fresh orphans.

Phase 4.4 — architecture.md updates:

- Background Services table gets a Session Cleanup row.
- New "Session Tab" subsection in API Endpoints documenting all six
  /api/agents/{name}/sessions* routes including the per-user ownership
  rule (404 not 403) and the resume-failure fallback / Redis lock.
- New /api/settings/feature-flags row.
- New agent_sessions / agent_session_messages DDL block in Database
  Schema, with the three Session-specific fields called out
  (cached_claude_session_id, consecutive_resume_failures,
  cache_read_tokens, claude_session_id audit).

Phase 4.5 — feature-flows/session-tab.md vertical slice:

- Full path from UI → API → DB → Side Effects with the JSONL lifecycle
  table, the spike-pitfall defense map (L1/L2/L3/#20992/#26964), the
  error-handling matrix, and the complete test catalog with the docker
  run command for the integration suite.
- feature-flows.md index updated (Recent Updates row + Chat & Sessions
  section entry).

Test totals: 25 unit + 9 integration = 34 tests, all green. Phase 4.3
serves as both the GA gate and the per-Claude-version regression guard.

Phase 4.1 (cache_read_tokens UI surfacing) deferred — the column is
already populated by the Phase 2 turn endpoint; surfacing is a minor
observability follow-up that doesn't block Phase 5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): tag Session-tab turns with triggered_by="session" and a gold badge

Previously Session-tab turns went into schedule_executions with
triggered_by="chat", so the Tasks tab couldn't tell them apart from
the Chat tab. The user-visible signal was that every Session turn
showed up under the sky-blue "chat" badge.

Backend (routers/sessions.py): both call sites that invoke
task_execution_service.execute_task — the cold/resume turn and the
resume-failure fallback retry — now pass triggered_by="session".
Existing rows are unchanged; the cutover is per-write.

Frontend (TasksPanel.vue): adds a "Session" option between "Chat" and
"Manual" in the trigger filter dropdown, plus an amber/gold badge
branch (bg-amber-100 dark:bg-amber-900/30 text-amber-700
dark:text-amber-300) — visually distinct from "paid" (bright yellow)
and from the sky-blue "chat" badge.

triggered_by is a free-form TEXT column (no enum constraint at the DB
or service layer), so adding "session" as a new value doesn't require
any migration or downstream consumer updates. Filter, badge, audit
log, activity stream, and dashboards all just see another value and
display it; nothing has to know about it explicitly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(session-tab): correct context-window accounting + raise frontend turn timeout

Five interrelated fixes from manual testing — all about the per-turn
"context %" metric being misleading and the browser timing out before
long-running session turns finished.

1) Agent server (docker/base-image/agent_server/services/claude_code.py)
   process_stream_line's `result` event handler used to overwrite
   metadata.input_tokens, cache_read_tokens, and cache_creation_tokens
   with the values from result.usage. Those values are CUMULATIVE
   across every internal API call the turn made (Claude Code packs
   tool-use loops into a single user turn that maps to N internal API
   calls). For an 18-iteration turn each reading the same 70K cached
   prefix, result.usage.cache_read_input_tokens = 18 * 70K = 1.26M
   tokens — billing-cumulative, not the prompt size of any single call.
   Overwriting per-message values with that aggregate made our
   context-window-pressure metric grow far beyond the 200K limit even
   when no individual API call was anywhere close to the wall.

   Fix: result handler now only extracts model-level facts (cost,
   duration, num_turns, session_id, error info, modelUsage.contextWindow).
   Per-API-call usage stays in the per-assistant-message handler, where
   the LATEST message's values represent the FINAL API call's prompt
   size — exactly what determines whether the next turn will fit.

   Also added a per-message usage-extraction block to the assistant
   branch of process_stream_line (it previously had no usage extraction
   at all, relying entirely on the result handler — which made my
   first attempt at this fix produce zero values). parse_stream_json_output
   already had the equivalent block (lines 211-215).

   Base image rebuilt; agent-testfix recreated onto the new image
   (image sha 0a1e20b40da1).

2) Backend (services/task_execution_service.py)
   Replaced `context_used = metadata.input_tokens` with
   `cache_read + cache_creation` (with input_tokens fallback when
   caching isn't engaged). input_tokens is sometimes the disjoint
   fresh value and sometimes inflated by the agent server's
   modelUsage.inputTokens override on tool-call turns. cache_read
   and cache_creation come straight from Anthropic's usage object
   and (post agent-server fix) are reliable per-call values that
   monotonically reflect the cached conversation prefix.

3) DB (db/sessions.py)
   total_context_used is now a HIGH-WATERMARK (MAX of prior + new),
   not the latest value. Per-turn context naturally oscillates by ~2x
   between text-only and tool-call turns; the watermark gives users
   a stable monotonic upper bound on session pressure that only goes
   up.

   Capped the watermark at total_context_max as a safety belt against
   any future agent-server bug that emits cumulative-billing token
   counts. Genuine per-call peaks should never exceed the model's
   context window — if they do, that's an accounting error not a
   real overflow, and the UI should display 100% rather than 648%.

4) Frontend (stores/sessions.js)
   Bumped the Axios timeout on the session turn endpoint from 305s
   (~5 min) to 7260s (= TIMEOUT-001 cap of 7200s + 60s slack). The
   session turn endpoint is synchronous and may legitimately run for
   the agent's full execution timeout. With the previous 305s ceiling
   the browser threw a misleading "failed" toast on tool-heavy turns
   that ran longer; the response still landed in the DB and the UI
   recovered after a page refresh, but the user saw a phantom error.

Verified end-to-end with a 6-turn mixed sequence (text + tool-call):
per-call cache_read now reports ~11636 on text-only turns and ~18000
on tool…
vybe added a commit that referenced this pull request Jun 1, 2026
* refactor(frontend): sweep auth + channels + top-level views (#554) (#609)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554) (#623)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554)

Slice 7 stacked on #609. Migrates 24 files spanning the small-domain pool:

  CHAT (4 files)
    ChatBubble  — copy-success checkmark, self-task accent panel (purple)
    ChatPanel   — agent-not-running warning state, error banner
    ChatInput   — file-remove button, voice-active recording indicator
    ChatHistoryDropdown — error state

  FILE MANAGER (4 files)
    FileManager        — notification toast (success/error), no-agents warning,
                         loading error, delete button + modal + confirm action
    FileTreeNode       — search-matched row highlight (file-type icons stay raw
                         decorative; need their own accent palette later)
    FilePreview        — preview-error icon + text
    FileSharingPanel   — Revoke button

  PROCESS (3 files)
    TrendChart        — completed/failed/cost bars (chart series + legend),
                         success-rate threshold ladder
    RoleMatrix        — no-executor row + badge
    TemplateSelector  — category badges (business/devops/support → status-info /
                         accent-purple / status-urgent)

  CROSS-CUTTING / MODALS (10 files)
    NavBar               — Ops critical-pulse + high indicator, WS connected dot
    GitConflictModal     — yellow warning header (×2), all destructive (red) options
    ReplayTimeline       — system-agent purple panel + badge, schedule-marker arrow,
                            live-feed dot, activity-state success rate ladder
    UnifiedActivityPanel — running/success/fail indicators (live + modal)
    OnboardingChecklist  — completed-state styling (ring, bg, indicator, text)
    CreateAgentModal     — templates-error + general error
    ConfirmDialog        — danger/warning variant icons, text, confirm buttons
    ResourceModal        — (no migrations — amber notice, deferred)
    AvatarGenerateModal  — error text, remove-avatar button
    HelpChatWidget       — error banner + retry button

  MISC (3 files)
    YamlEditor          — error and warning banners + counts + success checkmark
    EditorHelpPanel     — required-field indicator
    TerminalPanelContent — restart-required notice, start-agent button
    TagsEditor          — error message

Net diff: 24 files, +106 / -106 (1:1 palette aliases, byte-identical CSS).

Deferred (existing pattern):
  - Indigo / blue primary action buttons (Login & elsewhere)
  - Blue selected-state (NavBar tabs, OperatingRoom tabs)
  - Amber notices (ResourceModal, RoleMatrix amber missing-role marker —
    these still use the amber palette which differs from yellow)
  - File-type icon colors in FileTreeNode (decorative, need accent-yellow /
    accent-purple-blue / etc. — folder ≠ warning, video ≠ accent, etc.)
  - Slack/Telegram/WhatsApp logo brand colors

These map to the pending `action-primary`, `state-selected`, and accent-color-
expansion tickets.

Verified locally:
  - npm run check:tokens                         passes (10 tokens valid)
  - npm run build                                passes
  - npm run test:e2e:smoke (7 tests, 7.9s)       all green

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep agent surfaces — 22 files (#554) (#614)

* refactor(frontend): sweep auth + channels + top-level views (#554)

Slice 4 — bundles three subdomains into one PR per request:

  AUTH (3 files, 24 migrations)
    - Login.vue           — error icon
    - SetupPassword.vue   — password match indicator, requirement checklist,
                            error banner, 6-tier strength visualization
                            (red/red/orange/yellow/green/green)
    - MobileAdmin.vue     — logout, status dot, fleet running/high-context
                            counts, chat thinking status

  CHANNELS (5 files, 23 migrations)
    - PublicLinksPanel    — Active badge, expired text, Slack
                            connected/enable/disable/delete icons,
                            form error, delete-confirm modal,
                            success toast
    - WhatsAppChannelPanel — connected dot, sandbox badge, disconnect btn,
                             webhook warning, success/error message
    - TelegramChannelPanel — connected dot, disconnect btn, webhook
                             warning, group-remove, success/error message
    - SlackChannelPanel   — connected dot, disconnect btn, success/error
    - SharingPanel        — Approve button, success/error message,
                            remove button

  TOP-LEVEL VIEWS (4 files, 33 migrations)
    - Dashboard.vue       — running count + dot, message count, clear-tags,
                            connection status dot, history badge,
                            live-feed indicator, message arrow icon
    - PublicChat.vue      — agent online dot, invalid-link icon + bg,
                            agent-unavailable icon + bg, verify error,
                            chat error
    - OperatingRoom.vue   — empty-state success indicator
    - Templates.vue       — error icon

80 token replacements; net diff +80 / -80 (all 1:1 palette aliases).

Deferred (30 raw refs remain; all need new token families):
  - Login: 8 blue (primary action buttons + focus rings)
  - Dashboard: 10 (blue actions, purple tag-cloud button, blue selected-tab)
  - OperatingRoom: 5 (blue selected-tab indicators)
  - PublicChat: 5 (amber AUTO badge, rose READ-ONLY badge, indigo loading)
  - WhatsApp + Sharing: 2 (amber "deployment prerequisite" notices)

These map to pending `action-primary`, `state-selected`, and an accent
expansion that's tracked under #555 follow-up territory.

Tests:
  No new specs — these routes are either auth pages (covered by
  auth.setup), already smoke-tested (Templates), or require fixtures
  (PublicChat needs a public link token; MobileAdmin lives at /m).
  Existing 6 @smoke tests cover the high-traffic routes.

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(vite): tighten /api proxy prefix to /api/

`/api` (no trailing slash) is a path-prefix match in http-proxy-middleware,
so the SPA route `/api-keys` was being captured by the proxy and forwarded
to the backend, returning 404 in dev mode.

This silently broke the `/api-keys` @smoke e2e test on every PR since the
test was added in #597 — that PR's frontend-e2e check failed on merge but
wasn't required, so the failure was missed.

Backend endpoints all live under `/api/...` (with slash), so the tighter
prefix preserves all real proxy traffic and only excludes the SPA route.

Verified locally: 7/7 @smoke tests pass after this change (was 6/7 with
/api-keys failing).

Refs #554 #556

* refactor(frontend): sweep cross-cutting + chat + file-mgr + process + misc (#554)

Slice 7 stacked on #609. Migrates 24 files spanning the small-domain pool:

  CHAT (4 files)
    ChatBubble  — copy-success checkmark, self-task accent panel (purple)
    ChatPanel   — agent-not-running warning state, error banner
    ChatInput   — file-remove button, voice-active recording indicator
    ChatHistoryDropdown — error state

  FILE MANAGER (4 files)
    FileManager        — notification toast (success/error), no-agents warning,
                         loading error, delete button + modal + confirm action
    FileTreeNode       — search-matched row highlight (file-type icons stay raw
                         decorative; need their own accent palette later)
    FilePreview        — preview-error icon + text
    FileSharingPanel   — Revoke button

  PROCESS (3 files)
    TrendChart        — completed/failed/cost bars (chart series + legend),
                         success-rate threshold ladder
    RoleMatrix        — no-executor row + badge
    TemplateSelector  — category badges (business/devops/support → status-info /
                         accent-purple / status-urgent)

  CROSS-CUTTING / MODALS (10 files)
    NavBar               — Ops critical-pulse + high indicator, WS connected dot
    GitConflictModal     — yellow warning header (×2), all destructive (red) options
    ReplayTimeline       — system-agent purple panel + badge, schedule-marker arrow,
                            live-feed dot, activity-state success rate ladder
    UnifiedActivityPanel — running/success/fail indicators (live + modal)
    OnboardingChecklist  — completed-state styling (ring, bg, indicator, text)
    CreateAgentModal     — templates-error + general error
    ConfirmDialog        — danger/warning variant icons, text, confirm buttons
    ResourceModal        — (no migrations — amber notice, deferred)
    AvatarGenerateModal  — error text, remove-avatar button
    HelpChatWidget       — error banner + retry button

  MISC (3 files)
    YamlEditor          — error and warning banners + counts + success checkmark
    EditorHelpPanel     — required-field indicator
    TerminalPanelContent — restart-required notice, start-agent button
    TagsEditor          — error message

Net diff: 24 files, +106 / -106 (1:1 palette aliases, byte-identical CSS).

Deferred (existing pattern):
  - Indigo / blue primary action buttons (Login & elsewhere)
  - Blue selected-state (NavBar tabs, OperatingRoom tabs)
  - Amber notices (ResourceModal, RoleMatrix amber missing-role marker —
    these still use the amber palette which differs from yellow)
  - File-type icon colors in FileTreeNode (decorative, need accent-yellow /
    accent-purple-blue / etc. — folder ≠ warning, video ≠ accent, etc.)
  - Slack/Telegram/WhatsApp logo brand colors

These map to the pending `action-primary`, `state-selected`, and accent-color-
expansion tickets.

Verified locally:
  - npm run check:tokens                         passes (10 tokens valid)
  - npm run build                                passes
  - npm run test:e2e:smoke (7 tests, 7.9s)       all green

Refs #554

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(frontend): sweep agent surfaces — 22 files (#554)

Slice 8 of the design-system migration. Replaces semantic palette refs
with status/state/accent tokens across all per-agent panels, the
agents list, and agent-detail surfaces.

Files migrated:
- 16 panels (Tasks, Git, Dashboard, Playbooks, Nevermined, Info,
  Credentials, Schedules, SystemViewEditor, HostTelemetry, Folders,
  Files, Skills, Observability, Permissions, Metrics)
- 5 agent surfaces (Agents, AgentNode, AgentHeader, AgentTerminal,
  AgentDetail)
- SystemAgentNode

Mappings (palette-equivalent, no visual change):
- yellow → status-warning
- green  → status-success
- red    → status-danger
- orange → status-urgent
- amber  → state-autonomous (token name slightly stretched for tool-call
  / queued-task amber; same palette, future cleanup may rename)
- rose   → state-locked
- purple → accent-purple

Deferred (no token family yet):
- indigo (action-primary)
- blue/sky/cyan (selected-state, category labels)
- teal (category labels)

SystemViewsSidebar untouched — only contains deferred blue/indigo
selected-state references.

Verification:
- npm run check:tokens → 10 tokens equivalent, all references resolve
- npm run build → clean
- npm run test:e2e:smoke → 7/7 passed against live Trinity (HMR)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(agent-runtime): kill npx MCP orphans outside claude pgid that hold stdout pipe open (#618) (#620)

* fix(voice): orb animation loop dies before voice session starts

The canvas is inside v-if="voice.isActive.value" so canvasEl.value is
null when onMounted fires. renderFrame() exits early without scheduling
the next frame, killing the loop permanently.

Replace onMounted initialization with watch(canvasEl) so the RAF loop
starts when the canvas enters the DOM and stops when it leaves.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent-runtime): kill npx MCP orphans outside claude pgid that hold stdout pipe open (#618)

After terminate_process_group kills claude's pgid, npm→node MCP server chains
spawned via npx call setsid() and land in a new session — they survive the
pgid kill and keep the stdout pipe write FD open indefinitely.  The kernel
cannot deliver EOF to our reader thread while any writer FD remains open, so
drain_reader_threads blocked for the full 30s post_kill_grace, then lost the
buffered result line via force-close (HTTP 502).

Add _kill_orphan_pipe_writers(): after terminate_process_group, scan /proc/*/fd
for any process outside our pgid that holds the pipe's write end (detected via
fdinfo flags), and SIGKILL it.  Killing the orphan releases all its FDs
(stdout AND stderr write ends) simultaneously, delivering EOF to both reader
threads so they drain naturally before the post_kill_grace window.

New tests (Linux-only, skipped on macOS — /proc required): verify that a
setsid() grandchild is detected and killed, that our own read-end process is
not touched, and that the end-to-end drain path preserves buffered data.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(slack): replace slackify-markdown with own renderer (#293) (#622)

Replaces slackify-markdown with a custom renderer fixing 5 compounding bugs (nested lists, headings, blockquotes, tables, horizontal rules). Includes 35 unit tests and updated feature flow doc.

Closes #293

Co-Authored-By: pavshulin <pavshulin@users.noreply.github.com>

* feat(agents): per-agent token usage display in AgentHeader (#250) (#632)

Adds a token usage row to AgentHeader showing 7-day cost sparkline,
today's cost vs 7-day daily average (with trend arrow), and lifetime
totals. Data sourced from schedule_executions in the DB so it persists
across agent restarts.

- New GET /api/agents/{name}/token-stats endpoint
- ScheduleOperations.get_agent_token_stats(): single-pass 24h/7d/lifetime
  aggregation + 7-day daily breakdown with gap-filling
- agentsStore.getAgentTokenStats() action
- TOKEN USAGE ROW in AgentHeader.vue: SparklineChart (amber, 56x16),
  trend indicator (warning/success/gray), lifetime summary
- Hidden for agents with no runs (lifetime_executions == 0)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(site): agent website proxy via /site/{token} endpoint (SITE-001) (#634)

Adds live HTTP reverse-proxy so agents can serve public websites from
their container. A new `type='site'` public link routes requests through
`GET /site/{token}/{path}` → httpx streaming proxy → agent web server at
`http://agent-{name}:3000`. Includes DB migration, nginx routing, rate
limiting (per-IP + per-token), SSRF guard, security header stripping, and
audit event `site_link_visit`. UI adds Chat/Website selector in the link
create modal with a "Website" badge on site links.

Fixes #633

Co-authored-by: Claude <noreply@anthropic.com>

* fix(site): centralize SITE_PORT, atomic rate limit, fire-and-forget audit log, update docs (SITE-001)

- Move SITE_PORT to config.py; import in site.py and public_links.py
- Fix TOCTOU race in _check_site_rate_limit: pipeline INCR+check-after
- Audit log is now asyncio.create_task() so streaming is not delayed
- Add SITE-001 to requirements.md (section 15.1a-3)
- Add site.py to architecture.md router listing + /site/ endpoint table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent-runtime): bound _kill_orphan_pipe_writers to 10s to prevent drain stall (#649) (#650)

/proc scanning can block indefinitely when a process is in D state
(uninterruptible sleep), causing drain_reader_threads to stall for
tens of minutes instead of the expected ~30 seconds.

Run _kill_orphan_pipe_writers in a daemon thread with a 10s cap so
a blocked /proc entry cannot push the drain past its deadline.

Also use wall-clock accounting for the post-kill join timeout so time
spent in terminate + orphan scan doesn't silently erode the budget,
and log actual elapsed time instead of the expected value so future
incidents are easier to diagnose.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): voice WebSocket + stop endpoint missing ownership check (#600) (#638)

The /ws/voice/{voice_session_id} handler decoded the JWT but threw the
payload away — only the signature was checked. Any authenticated user
holding a valid JWT who learned the 128-bit session id (logs, browser
inspection, XSS) could attach to the audio stream, eavesdrop on the
victim's transcript, and trigger tool calls audit-logged under the
victim's identity.

POST /api/agents/{name}/voice/stop had the same gap: the path agent was
gated via get_authorized_agent, but request.voice_session_id was never
cross-checked against the path agent or the caller's user_id, so the
caller could end and persist a transcript onto another user's session.

Fix:
- WS: extract sub from the decoded JWT, look up the user, and close 4003
  if user.id != session.user_id (admin role bypasses, for support).
- voice_stop: load the session via get_session before mutating, assert
  agent_name == path name AND user_id == current_user.id (admin bypasses),
  raise 403 otherwise.

Added tests/unit/test_voice_auth.py covering: missing token, invalid
token, missing sub claim, unknown user, owner happy path, admin bypass,
attacker rejected, plus voice_stop variants. Loads voice.py via
importlib to avoid pulling in the full routers/__init__.py chain.

Reported by /security-review on PR #599 (2026-04-30); origin commit
7d8abe8 (#581 voice tool calls).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): split login rate limit into per-account + per-IP buckets (#591) (#621)

Pentest finding AISEC-H2 (CVSS 7.5, CWE-307): the previous design used a
single per-IP bucket at 5 fails / 10 min. Any user behind a corporate
NAT, VPN, or CDN locked out everyone else at the same egress IP after
just four bad attempts. A rotating-proxy attacker could keep an
organisation locked out continuously, so the protection doubled as a
platform-wide DoS primitive.

Replace with two independent buckets:

  * Per-account (tight) — 5 fails / 15 min: limits credential stuffing
    on one targeted account; never affects other accounts.
  * Per-IP (loose)      — 30 fails / 5 min: catches single-source abuse
    but stays well above the legitimate-traffic threshold for users
    sharing a NAT/VPN/CDN egress.

Both buckets are checked on every attempt; 429 fires when either is
exhausted. Successful login clears both. Account names are normalised
(lowercase + strip) before keying. Endpoints without an account context
(public access-request) skip the per-account bucket and rely on the
per-IP one only.

Lockout state-changes log a structured WARNING (visible via Vector) so
operators can see when buckets are being exercised.

Live verification on the running backend:
  attempts 1-5 → 401 (counter ticking)
  attempt  6   → 429 "Too many failed attempts for this account..."
  valid pwd    → 429 (account stays locked even with right password)
  other account from same IP → 401 (per-account isolation works)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): require creator role on /api/systems/deploy (#592) (#624)

Auditing all entry points that ultimately call `create_agent_internal`
(per #592 AC #2) turned up a real bypass on `POST /api/systems/deploy`.
The system-manifest deployment route gated on `Depends(get_current_user)`
without a role check, so any authenticated user-role account could spawn
an entire fleet of agents through that path — a strictly stronger
privilege than the single-agent bypass the AISEC-H1 finding originally
named (which `POST /api/agents/deploy-local` had already closed via #150).

Add `Depends(require_role("creator"))` to `deploy_system`, matching the
existing dependencies on `POST /api/agents` and `POST /api/agents/deploy-local`.

Regression test (`tests/unit/test_agent_creation_role_gates.py`) walks
the FastAPI router source AST and asserts that every agent-creation
route uses `Depends(require_role("creator"))`. AST-level so the check is
fast, stable across formatting changes, and fires the moment someone
removes the dependency. Confirmed the test catches the regression by
reverting the change and observing the test fail before re-applying.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): stop mirroring JWT into document.cookie (#188) (#642)

* fix(security): stop mirroring JWT into document.cookie (#188)

UnderDefense pentest 3.3.5 flagged the frontend for mirroring the
authentication token into a `token` cookie without `Secure` or
`HttpOnly` flags. The cookie was set in setupAxiosAuth via
`document.cookie =` so it was readable from JS (HttpOnly is impossible
on JS-set cookies), transmitted over HTTP without the Secure flag, and
auto-attached to every outbound request as a CSRF vector.

The cookie's stated purpose was "for nginx auth_request to validate
agent UI access" — but that nginx directive was never configured in
any committed deployment (`grep -r auth_request -- *.conf` is empty,
git log -S confirms it never existed). The cookie was pure attack
surface with zero functional value.

Per the issue's "Best" remediation, drop the cookie mirror entirely.
API authentication uses the `Authorization: Bearer` header
exclusively; nothing else needs the cookie.

The cookie-clear on logout is intentionally kept so users carrying a
stale cookie from the pre-fix version get cleaned up on their next
logout cycle. The cookie's `max-age=1800` also naturally expires it
within 30 minutes of the upgrade.

The backend's `/api/auth/validate` endpoint still accepts a cookie as
one of three token sources — left untouched as out-of-scope. With the
frontend no longer setting the cookie nothing legitimate sends one,
but the fallback path remains available if a future nginx
auth_request setup is wired up properly (with Secure + HttpOnly flags
set server-side via Set-Cookie, not via document.cookie).

Closes #188.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(admin-login): remove stale references to JWT cookie mirror (#188)

PR #642 removed the document.cookie set in setupAxiosAuth, but the
admin-login feature flow still documented the cookie as live. Update
the code snippet and the storage table to reflect current behaviour.

Note in the snippet describes why the cookie was removed so readers
who see the diff history can find the rationale without reading the
PR.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): override git User-Agent on skills library sync (#184) (#646)

UnderDefense pentest 3.3.1 flagged the backend for leaking the
underlying tech stack via outbound User-Agent. Skills sync uses git
subprocess (not httpx), so the leaked UA is `git/<version>
(libcurl/<version> ...)` — verified live with GIT_TRACE_CURL=1.

Add `-c http.useragent=Trinity-Skills-Sync` (positioned correctly
before the subcommand, as git's `-c` requires) to the two HTTP-bearing
git invocations: `_git_clone` and `_git_pull`'s fetch. The local-only
`git reset --hard` and `git rev-parse HEAD` calls intentionally do
not get the flag — they make no HTTP and threading the flag through
would suggest otherwise.

The SSRF allowlist (#179) already locks the destination to github.com
so the practical exposure is small (GitHub already knows what we are),
but defense-in-depth: even if the allowlist is ever loosened the UA
stays generic.

The constant has no version suffix to avoid yet another version string
drifting against VERSION / package.json / pyproject.

Tests in tests/unit/test_skill_service_user_agent.py mock subprocess
and assert the flag is present at the right argv position for clone
and fetch, and absent for the local-only reset and rev-parse calls.

Live verification with `GIT_TRACE_CURL=1 git -c http.useragent=... ls-remote ...`
confirms the wire UA changes from `git/2.43.0` to `Trinity-Skills-Sync`.

Closes #184.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(deploy): document stale-image symptom + recovery in start.sh and DEPLOYMENT.md (#557) (#626)

Self-hosted developers tracking `dev` occasionally pull a commit that
adds a new Python or Node dependency to one of the platform Dockerfiles,
re-run `start.sh`, and end up with new source running against an old
image's Python env. Uvicorn crashes with `ModuleNotFoundError`, compose
keeps respawning the worker, and `start.sh` reports success — leaving
the UI "Disconnected" with no obvious diagnosis.

Adopting Option B from #557 discussion (PR #625 closed): treat this as
a documentation problem rather than building auto-detection into the
critical-path startup script. Auto-detection has clear cost (Python
subprocess + Docker inspect on every cold start) and unclear benefit
(production deploys use `compose pull` and don't hit this; the affected
population is self-hosted devs whose recovery is one command).

Two changes:
- `scripts/deploy/start.sh`: append a 4-line hint after the "Ready!"
  banner naming the symptom (`ModuleNotFoundError`, "Disconnected" UI)
  and the exact recovery command.
- `docs/DEPLOYMENT.md`: add a Troubleshooting entry with full diagnosis
  walkthrough, root-cause explanation, and the rationale for not
  auto-detecting (links to #557).

A future `scripts/deploy/upgrade.sh` is the right place to bundle
backup + rebuild + start + verify for the explicit upgrade path; that
is bigger-than-#557 scope.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): lock down Redis — auth + ACL + network split (#589)

* fix(security): split Docker compose into platform and agent networks (#589)

Redis at 172.28.0.0/16 was reachable from any agent container. AISEC scan
3aad5469 demonstrated end-to-end exfiltration / cross-user task injection
from a legitimately deployed agent. Network segmentation is the strongest
control — agents now physically cannot route to Redis.

Topology:
- trinity-platform (172.29.0.0/16, NEW) — Redis, scheduler, vector
- trinity-agent (172.28.0.0/16, name preserved) — frontend, agents
- Backend / mcp-server / otel-collector / cloudflared straddle both

Agent-creation sites in services/agent_service/* and system_agent_service.py
need zero changes because the agent-network external name is preserved.

Dev: bind Redis host port to 127.0.0.1:6379 (was 0.0.0.0). Tests connect
from the dev machine; LAN cannot. Auth lands in the next commit.

Refs #589 — acceptance criterion #3 (network segment separation).

* fix(security): mandatory Redis auth, ACL users, auth-aware healthcheck (#589)

Both compose files now enforce two passwords (REDIS_PASSWORD admin /
REDIS_BACKEND_PASSWORD runtime) with the fail-on-missing :? form.
docker compose refuses to render without them.

Per-user ACL via inline --user flags. Additive (start from zero, allow
only what the runtime needs) — never +@all -X, which lets newly added
dangerous commands through. backend + scheduler get standard data
families plus scripting/transactions/pubsub minus -@dangerous, which
covers FLUSHALL, CONFIG, SHUTDOWN, MIGRATE, REPLICAOF, MONITOR.

Verified at runtime against redis:7-alpine: PING/SET/GET work for the
backend user, FLUSHALL and CONFIG GET return NOPERM, unauth requests
return NOAUTH.

REDIS_URL on backend + scheduler now embeds the backend ACL user.
mcp-server: REDIS_URL and depends_on:redis dropped in prod compose
(zero Redis imports in src/mcp-server/).

Healthcheck pings as the backend ACL user so a typo'd ACL keeps redis
unhealthy and gates dependent services. depends_on:redis switches to
service_healthy so backend/scheduler don't race the ACL load.

Refs #589 — acceptance criteria #1, #2, #5.

* fix(scheduler-test-rig): mirror Redis auth posture (#589)

Without this, scheduler container fails fast on startup against the rig
because src/scheduler/config.py requires creds in REDIS_URL after #589.
No ACL or network split here — this is a 2-service standalone debugging
rig, not the production posture.

* fix(security): fail-fast on REDIS_URL missing credentials (#589)

Backend (src/backend/config.py) and scheduler (src/scheduler/config.py)
now raise RuntimeError at import time if REDIS_URL is unset or lacks
credentials.

Removed the splicing fallback in backend config that papered over an
unauth REDIS_URL by joining REDIS_PASSWORD into the URL — single source
of truth (compose) eliminates silent drift.

Tests that import backend modules need a creds-bearing REDIS_URL in their
environment; tests/conftest.py will set a dummy one in the test commit.

Refs #589 — acceptance criterion #5.

* fix(webhooks): use REDIS_URL for rate-limit client (#589)

Webhooks rate-limit was the one Redis client that bypassed REDIS_URL —
it used redis.Redis(host="redis", port=6379) and would silently
fail-open under requirepass. Switching to redis.from_url(REDIS_URL)
picks up the credentialed URL like every other client.

Also: distinguish auth/ACL errors (logged at ERROR with exception class)
from transient errors (WARN). Fail-open behavior preserved so a Redis
blip doesn't 500 legitimate webhooks, but a misconfigured deploy now
surfaces in alerts instead of via a webhook abuse incident.

Drops the now-unused REDIS_HOST/REDIS_PORT env reads.

* feat(deploy): auto-generate Redis passwords on fresh installs (#589)

start.sh ensure_redis_passwords matches the existing
CREDENTIAL_ENCRYPTION_KEY pattern, with one safety guard:

- Fresh install (no redis-data volume) → generate both passwords with
  openssl rand -hex 24 and append to .env. One-command boot keeps
  working.
- Existing volume + missing password → refuse with a loud error pointing
  at docs/migrations/REDIS_AUTH.md. Re-keying a populated Redis would
  lock the backend out of its own data; ops needs to follow the explicit
  upgrade path.

Idempotent — second run is a no-op when both passwords are already set.

* docs(security): add Redis auth migration guide + architecture notes (#589)

- docs/migrations/REDIS_AUTH.md: operator upgrade guide. Covers fresh
  installs (auto-generated by start.sh), live upgrades (down
  --remove-orphans + docker network rm + add passwords), production,
  and verification commands.
- docs/memory/architecture.md: new "Network Topology (Issue #589)"
  section above Container Security. Documents the two-network split,
  service membership table, the "agents NEVER on platform network"
  rule, and the three Redis ACL users + their access patterns.

* test(security): network isolation, ACL, fail-fast, webhook rate-limit (#589)

tests/conftest.py: top-level autouse env stub for backend imports.
Backend config now raises at import-time if REDIS_URL lacks credentials;
without this, every test that transitively imports backend modules
breaks. Real Redis tests under tests/security/ override via their own
conftest from .env. Adds the `integration` marker.

tests/unit/test_config_fail_fast.py (new): backend refuses to import
without creds-bearing REDIS_URL. 3 cases — missing env, unauth URL,
URL with creds.

tests/security/test_redis_network_isolation.py (new): 5 integration
tests covering acceptance criteria #1-#3:
  - agent-network container has no route to redis (BLOCKED)
  - unauth client gets NOAUTH on platform network
  - backend ACL user can PING with creds
  - backend ACL user FLUSHALL → NOPERM (no admin)
  - backend ACL user CONFIG GET → NOPERM (no requirepass leak)

tests/security/conftest.py (new): session-scoped fixture loads real
.env values for the integration tests; skips the suite if missing.

tests/integration/test_webhook_rate_limit.py (new): regression for the
from_url switch in webhooks.py. Self-contained — creates agent +
schedule + webhook token inline, hits 11×, expects 429 on the 11th.
Catches the silent fail-open if Redis auth ever regresses.

tests/run-integration.sh (new): pytest -m integration runner. Excluded
from run-smoke.sh per the smoke runner's ~30s no-Docker contract.

* docs(security): detach agents before network rm (#589)

Trinity-managed agent containers are created via the Docker SDK
outside compose, so they store the agent network's UUID, not its
name. After `docker network rm trinity-agent-network` (step 3 of
the upgrade procedure), any later `docker start <agent>` fails:

    Error response from daemon: failed to set up container
    networking: network <old-uuid> not found

Compose-managed services don't hit this — they're recreated with
fresh network refs on `up`. Agent containers aren't, so they keep
the stale UUID until disconnected.

Add an explicit detach loop as step 2, before the network removal.
Verified against a populated install with one running and four
stopped agents: all five reattach cleanly to the new network on
next start.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(security): CSO OBS-1/2/3 follow-ups — webhook rate-limit + healthcheck hardening (#589)

Resolves three observations from the CSO audit
(docs/security-reports/cso-2026-05-04-589-diff.md):

OBS-1 — webhook rate-limit fail-open + connection-per-request DoS amplifier:
* Added in-process secondary rate limiter (3x primary, per-worker) in
  src/backend/routers/webhooks.py. Bounds blast radius during a Redis
  outage without breaking the documented fail-open philosophy.
* Cached the Redis client at module level under threading.Lock with
  double-checked init. _check_webhook_rate_limit resets the cache on
  inner exceptions so stale connections rebuild cleanly. Without
  caching, a flood would open a fresh TCP per request and exhaust
  Redis maxclients — turning the rate limiter into the DoS amplifier.

OBS-2 — tightened _TOKEN_RE from {20,60} to {43} matching
secrets.token_urlsafe(32) (verified against db/schedules.py:524).

OBS-3 — switched all three compose healthchecks from
`redis-cli -a $$PASS` to `REDISCLI_AUTH="$$PASS" redis-cli` so the
password no longer appears in /proc/<pid>/cmdline.

Additional #589 hardening (caught while resolving OBS-1):
* src/backend/config.py + src/scheduler/config.py: tightened the
  REDIS_URL credential check from `"@" in url` substring to urlparse
  validation. Catches redis://@redis:6379, redis://user@redis:6379, etc.
* src/scheduler/main.py: redact password from REDIS_URL before logging
  (was leaking via Vector log aggregator).

Tests:
* tests/unit/test_webhook_rate_limit_inprocess.py — 7 new tests covering
  cap, window expiry, token isolation, runtime-error fallback, regex
  shape, cache hit, cache reset.
* tests/unit/test_config_fail_fast.py — 4 new parametrized cases for
  malformed-credential URL rejection.
* 15/15 unit tests pass.
* Live Redis healthcheck verified — trinity-redis reports healthy with
  the new REDISCLI_AUTH form; `redis-cli ping` returns PONG.

Also adds .gstack/ to .gitignore so future skill artifacts stay local.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(dev): update gitea overlay network name after #589 split

trinity-network no longer exists; gitea dev overlay must attach to
trinity-agent-network (the preserved agent-network name).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tests): use credentialed Redis URL in scheduler test_config fixture (#661)

After #589 hardened Redis auth, SchedulerConfig raises on bare redis:// URLs.
The test_config fixture bypassed the env-level patch in tests/conftest.py by
passing redis_url="redis://localhost:6379" directly.

Fixes #659

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* Session tab — `--resume`-default chat surface (Closes #651) (#652)

* docs(planning): add Session tab design — --resume-default chat surface

Adds docs/planning/SESSION_TAB_2026-04.md, the comprehensive plan for a
new "Session" tab living alongside Chat. Sessions reattach to their own
Claude Code JSONL via --resume, preserving tool memory, mid-skill state,
and reasoning state across turns.

Plan covers:
- UI design (tab placement, multi-session model, +New Session, Reset memory)
- Data model (agent_sessions / agent_session_messages — parallel to chat)
- Backend architecture (separate router, single shared change to
  task_execution_service for persist_session plumbing)
- Phased rollout (foundation → backend → frontend → hardening → GA)
- Edge cases & failure-mode lessons baked in from a prior local spike
  (parser bug, --no-session-persistence dependency, cold-turn detection,
  port allocation)
- Test plan including the cross-session contamination test for
  Anthropic claude-code#26964
- Retention/cleanup policy, observability, security checklist
- Local-first workflow: implementation runs entirely on this branch
  until validation passes; only then does the standard SDLC engage
  (issue, push, PR)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add agent_sessions + agent_session_messages tables

Phase 1.1 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).
Schema definitions go in db/schema.py for fresh installs; the matching
idempotent migration agent_sessions_tables in db/migrations.py upgrades
existing databases.

The schema mirrors chat_sessions / chat_messages but is strictly parallel
— no foreign keys, no shared columns, separate index namespace. Three
fields are unique to the session model:

- agent_sessions.cached_claude_session_id — the Claude Code session UUID
  the next turn will pass to ``--resume``
- agent_sessions.consecutive_resume_failures — drives the resume-failure
  fallback (Phase 2.2)
- agent_session_messages.cache_read_tokens — observability for whether
  Anthropic's prompt cache engaged

CASCADE on session delete cleans up message rows automatically.

Verified locally: backend restart applies the migration cleanly, tables
have 15 columns each with correct types/defaults/PKs, all four indexes
created, second restart confirms idempotency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(db): add SessionOperations for Session tab persistence

Phase 1.2 of the Session tab plan (docs/planning/SESSION_TAB_2026-04.md).

- Adds AgentSession and AgentSessionMessage Pydantic models in db_models.py
  with the new fields the Session tab needs beyond ChatSession/ChatMessage:
  cached_claude_session_id, last_resume_at, consecutive_resume_failures on
  the session row, and cache_read_tokens + claude_session_id on each message.

- Creates db/sessions.py with a SessionOperations class mirroring the
  ChatOperations shape: create_session, get_session, list_sessions,
  delete_session, add_session_message, get_session_messages, plus the
  Claude UUID cache helpers (get/update/clear_cached_claude_session_id)
  and resume health helpers (mark_resume_failure, mark_resume_success).

- Wires the new ops into the DatabaseManager facade alongside the
  existing _chat_ops, with one delegating method per public operation.

No router, no agent-server change, no frontend yet — those land in later
phases. Tables agent_sessions and agent_session_messages already exist
from the prior schema commit.

* feat(session-tab): backend foundation for --resume-default Session surface

Phases 1.3 through 1.7 of the Session tab plan
(docs/planning/SESSION_TAB_2026-04.md). Pure backend / agent-server work
behind a flag — no UI surface yet, no behavior change to Chat or any
existing /task caller.

Agent server (base image):

- Stream-json parser fix (Appendix B). Both parse_stream_json_output and
  process_stream_line now recognize {"type":"system","subtype":"init"}
  for session_id capture, with the result event as a fallback when init
  was missed (truncated streams). The legacy bare-init shape is
  intentionally rejected. This is the same bug that would have made
  Session caching corrupt on every cold turn.

- Same bug in execute_headless_task's permission-mode validation site:
  the check matched the wrong shape, so permission_mode_validated never
  flipped to True and the protective kill-on-misconfigured-permission
  path silently failed open. Now uses type=system + subtype=init.

- New persist_session flag threaded through ParallelTaskRequest →
  routers/chat.py → AgentRuntime ABC → ClaudeCodeRuntime.execute_headless
  → execute_headless_task. When True, --no-session-persistence is
  omitted so the JSONL is written and the next turn's --resume can find
  it. --session-id is still passed for unique cold-turn namespace.
  Default False keeps every existing caller stateless.

- gemini_runtime accepts the parameter for ABC parity and ignores it
  (Gemini CLI has no resume).

Backend:

- task_execution_service.execute_task now accepts persist_session: bool
  = False and threads it into the agent payload. All existing callers
  (Chat, schedules, MCP, fan-out, webhooks) keep today's behavior; only
  the future routers/sessions.py (Phase 2) opts in.

- settings_service.is_session_tab_enabled() — feature flag resolving
  system_settings.session_tab_enabled → SESSION_TAB_ENABLED env →
  False. Module-level convenience function exposed.

Tests (run inside trinity-backend container — Python 3.11):

- tests/unit/test_session_operations.py — 9 tests against an isolated
  SQLite DB exercising the full SessionOperations CRUD plus the cached
  claude session UUID lifecycle and resume failure / success counters.

- tests/unit/test_claude_code_session_id_parser.py — 8 tests covering
  both parsers (batch + streaming): system/init recognition, result
  fallback, init-wins-over-result, legacy bare-init rejection, and a
  source-level regression guard for the permission-mode validation
  fix.

- tests/unit/test_session_persistence_flag.py — 8 tests pinning the
  contract: signatures across the runtime ABC, ParallelTaskRequest,
  agent chat router, execute_headless_task, and
  task_execution_service.execute_task. Includes the gating regex check
  on --no-session-persistence and a live signature import to catch
  drift AST parsing alone would miss.

Total: 25 passing tests covering every touchpoint of Phase 1.

Base image (trinity-agent-base) rebuilt to embed the agent-server
changes; existing agent containers will pick them up on next recreate.

* feat(session-tab): backend turn endpoint for --resume-default Session surface

Phase 2 of docs/planning/SESSION_TAB_2026-04.md. Six endpoints under
/api/agents/{name}/session{s,...} that mirror routers/chat.py's auth
model and TaskExecutionService usage but persist to the parallel
agent_sessions / agent_session_messages tables and request
persist_session=True on every turn so each call reattaches via
`claude --print --resume <uuid>`.

Surface gated on is_session_tab_enabled() — flag-off default returns
404 from every endpoint.

  POST   /api/agents/{name}/session                  create row
  GET    /api/agents/{name}/sessions                 list (per-user)
  GET    /api/agents/{name}/sessions/{id}            session + messages
  POST   /api/agents/{name}/sessions/{id}/message    THE turn
  POST   /api/agents/{name}/sessions/{id}/reset      clear cached uuid
  DELETE /api/agents/{name}/sessions/{id}            delete row + msgs

Spike-pitfall defenses baked into the turn endpoint:

- L3 (first-turn-has-no-session-id): the agent_sessions row is created
  server-side via POST /session BEFORE the turn endpoint ever calls
  execute_task. No frontend-first model.
- L2 (cold turn writes empty JSONL): persist_session=True is passed
  unconditionally — Phase 1.4 already wired the flag through the agent
  stack; Phase 2 just promises to set it on every turn.
- L1 (parser misses system/init): trust result.session_id directly —
  Phase 1.3 fixed the parser. Scenario A confirms the captured UUID is
  the real Claude UUID end-to-end.

Phase 2.2 resume-failure fallback: when execute_task returns "no
conversation found" on a turn that had a cached UUID, clear the cache,
mark_resume_failure, and retry once with resume_session_id=None. Logs
event=session_resume_fallback with the stale UUID and consecutive
failure count. Anthropic #39667 (cleanupPeriodDays) and #53417 (CLI
upgrade) both produce this signal.

Phase 2.3 Redis lock: SET NX EX per (agent, claude_uuid) with 5-min TTL
and Lua-script release. Async poll loop (250ms tick) so the event loop
stays free during contention. Cold turns skip the lock (no JSONL to
corrupt). Hard 30s wait ceiling — beyond that the contender gets HTTP
429 with retry hint. Mitigation for Anthropic #20992 (concurrent
--resume JSONL writes corrupt the file).

Per-user ownership at the row layer: even agent owners cannot read or
send into another user's session (E6 isolation in the design doc).
Returns 404 for ownership failures so we don't leak session-id existence.

Tests (tests/integration/test_session_turns.py, run inside
trinity-backend container with docker.sock mounted for testfix
recreation + JSONL surgery in Scenario C):

  Scenario A: 3-turn happy path — same Claude UUID across turns
  Scenario B: turn 2 recalls a secret from turn 1, no text-replay
  Scenario C: JSONL deletion mid-session triggers fallback + recovery
  Scenario D: concurrent POSTs serialise via Redis lock
              (asserts finish_gap ≈ winner_work_time, NOT total wall)
  Scenario E: switching sessions A → B → A preserves A's UUID

5 passed in 54.5s against the live agent-testfix container (recreated
onto the rebuilt base image first per L4 in the plan). Phase 1's 25
unit tests still pass — no regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): frontend Session surface

Phase 3 of docs/planning/SESSION_TAB_2026-04.md. Adds the new "Session"
tab in AgentDetail, gated on the is_session_tab_enabled() platform flag
so it stays invisible until explicit opt-in (default off).

Backend prerequisite — routers/settings.py:

- GET /api/settings/feature-flags exposes a curated allowlist of UI-
  relevant flags to any authed user. The existing /api/settings/{key}
  endpoint is admin-only and would block non-admin frontends from even
  knowing whether to render the Session tab. The new endpoint reads
  through services.settings_service.is_session_tab_enabled() so the
  resolution order (DB → env → False) stays in one place.

Frontend:

- src/frontend/src/stores/sessions.js — Pinia store wrapping the six
  /api/agents/{name}/sessions* endpoints with per-agent state isolation
  and the feature-flag cache. Optimistic user-message insert with
  rollback on send failure.

- src/frontend/src/components/SessionPanel.vue — structural copy of
  ChatPanel reusing ChatMessages + ChatInput + ModelSelector. Differs
  from Chat in three places per the design doc:
    * Sends bare user_message to POST .../sessions/{id}/message — no
      buildContextPrompt text-replay (the agent already has working
      memory via --resume).
    * "Reset memory" button + confirm modal that clears the cached
      Claude UUID without deleting the message log (Phase 3.4).
    * Per-session selector subtitle: turn count, context % used,
      cached-memory dot (emerald/gray), and consecutive_resume_failures
      indicator (Phase 3.5).
  Lean cut for first-visible-surface: voice mic, file upload, and SSE
  dynamic status labels are deferred — those need backend extensions
  (file payload on the turn endpoint, async_mode + SSE on the same).

- src/frontend/src/views/AgentDetail.vue — new Session tab inserted
  between Chat and Dashboard/Schedules, gated on
  sessionsStore.sessionTabEnabled. Layout sites that previously
  branched on activeTab === 'chat' now use a shared isFullscreenTab
  computed so Chat and Session both get the input-pinned-to-bottom flex
  layout. ?tab=session deep-link allowlist updated.

- src/frontend/e2e/session-tab.spec.js — Phase 3.6 Playwright spec.
  Marked @interactive (not @smoke) because each run makes one real
  Claude API call (~10–60s). Snapshots the prior flag value in
  beforeAll, force-enables for the run, restores in afterAll so a
  failed run doesn't leave the platform with the flag dirty. Three
  cases:
    * tab is hidden when flag is off
    * tab appears, "+ New Session" → send turn → reply visible →
      Reset memory modal opens + closes
    * Chat tab still works after Session interaction; switching back
      preserves Session state

Visually verified in the live dev server: tab renders in correct
position, header layout matches Chat's structure, empty state and
placeholder copy match the design doc, "Reset memory" only shown when
an active session exists, full-viewport flex layout pins input to
bottom.

Phase 1 + Phase 2 work behind this change is unchanged: 25 unit tests
+ 5 integration tests still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): hardening + observability — cleanup service, contamination gate, docs

Phase 4 of docs/planning/SESSION_TAB_2026-04.md. Closes the JSONL
disk-growth loop, validates the GA-blocking cross-session contamination
hypothesis empirically, and lands architecture.md / feature-flows
documentation so the surface is discoverable.

Phase 4.3 — cross-session contamination GA gate (the load-bearing one):

- tests/integration/test_session_cross_contamination.py exercises the
  Anthropic #26964 hypothesis end-to-end. Plants a randomly-generated
  secret token in session A with explicit "do not echo" framing, asks
  session B (different UUID, same agent, same cwd) to recall the token.
  Hard-fails if the exact token leaks; soft-fails on partial-prefix
  recall (PURPLE-DRAGON without the random suffix would only be
  knowable from A's JSONL, not from training).
- PASSED in 9.5s on the current Claude Code version → shared-cwd model
  is safe → Phase 5 rollout unblocked. Test stays in the suite as the
  per-version regression guard.

Phase 4.2 — JSONL cleanup service:

- services/session_cleanup_service.py runs a 6h periodic sweep that
  diffs every running agent's
  ~/.claude/projects/-home-developer/<uuid>.jsonl set against
  db.list_active_claude_session_ids(agent) and reaps orphans whose
  mtime is older than the 1h race guard. Race guard prevents the
  cold-turn-vs-cleanup window where a brand-new JSONL exists on disk
  before the backend has updated cached_claude_session_id.
- Same service exposes a synchronous reap_jsonl(agent, uuid) helper
  called best-effort from routers/sessions.py reset/delete handlers so
  the user-perceived disk-reclaim latency is sub-second. Never raises;
  failures are logged and the periodic sweep is the safety net.
- Implementation uses execute_command_in_container — the same primitive
  git_service / ssh_service / scheduler pre-check / agent terminal use.
  No new agent-server endpoint, no base-image rebuild.
- New db.list_active_claude_session_ids(agent) facade method backed by
  SessionOperations.list_active_claude_session_ids querying every
  agent_sessions row whose cached_claude_session_id is non-null for the
  agent.
- main.py wires startup (staggered +7.5s after cleanup_service to
  offset Docker hits) and clean shutdown.
- tests/integration/test_session_cleanup.py: reset reaps synchronously,
  delete reaps synchronously, periodic sweep keeps the active JSONL,
  reaps an aged orphan, respects the 1h race guard for fresh orphans.

Phase 4.4 — architecture.md updates:

- Background Services table gets a Session Cleanup row.
- New "Session Tab" subsection in API Endpoints documenting all six
  /api/agents/{name}/sessions* routes including the per-user ownership
  rule (404 not 403) and the resume-failure fallback / Redis lock.
- New /api/settings/feature-flags row.
- New agent_sessions / agent_session_messages DDL block in Database
  Schema, with the three Session-specific fields called out
  (cached_claude_session_id, consecutive_resume_failures,
  cache_read_tokens, claude_session_id audit).

Phase 4.5 — feature-flows/session-tab.md vertical slice:

- Full path from UI → API → DB → Side Effects with the JSONL lifecycle
  table, the spike-pitfall defense map (L1/L2/L3/#20992/#26964), the
  error-handling matrix, and the complete test catalog with the docker
  run command for the integration suite.
- feature-flows.md index updated (Recent Updates row + Chat & Sessions
  section entry).

Test totals: 25 unit + 9 integration = 34 tests, all green. Phase 4.3
serves as both the GA gate and the per-Claude-version regression guard.

Phase 4.1 (cache_read_tokens UI surfacing) deferred — the column is
already populated by the Phase 2 turn endpoint; surfacing is a minor
observability follow-up that doesn't block Phase 5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(session-tab): tag Session-tab turns with triggered_by="session" and a gold badge

Previously Session-tab turns went into schedule_executions with
triggered_by="chat", so the Tasks tab couldn't tell them apart from
the Chat tab. The user-visible signal was that every Session turn
showed up under the sky-blue "chat" badge.

Backend (routers/sessions.py): both call sites that invoke
task_execution_service.execute_task — the cold/resume turn and the
resume-failure fallback retry — now pass triggered_by="session".
Existing rows are unchanged; the cutover is per-write.

Frontend (TasksPanel.vue): adds a "Session" option between "Chat" and
"Manual" in the trigger filter dropdown, plus an amber/gold badge
branch (bg-amber-100 dark:bg-amber-900/30 text-amber-700
dark:text-amber-300) — visually distinct from "paid" (bright yellow)
and from the sky-blue "chat" badge.

triggered_by is a free-form TEXT column (no enum constraint at the DB
or service layer), so adding "session" as a new value doesn't require
any migration or downstream consumer updates. Filter, badge, audit
log, activity stream, and dashboards all just see another value and
display it; nothing has to know about it explicitly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(session-tab): correct context-window accounting + raise frontend turn timeout

Five interrelated fixes from manual testing — all about the per-turn
"context %" metric being misleading and the browser timing out before
long-running session turns finished.

1) Agent server (docker/base-image/agent_server/services/claude_code.py)
   process_stream_line's `result` event handler used to overwrite
   metadata.input_tokens, cache_read_tokens, and cache_creation_tokens
   with the values from result.usage. Those values are CUMULATIVE
   across every internal API call the turn made (Claude Code packs
   tool-use loops into a single user turn that maps to N internal API
   calls). For an 18-iteration turn each reading the same 70K cached
   prefix, result.usage.cache_read_input_tokens = 18 * 70K = 1.26M
   tokens — billing-cumulative, not the prompt size of any single call.
   Overwriting per-message values with that aggregate made our
   context-window-pressure metric grow far beyond the 200K limit even
   when no individual API call was anywhere close to the wall.

   Fix: result handler now only extracts model-level facts (cost,
   duration, num_turns, session_id, error info, modelUsage.contextWindow).
   Per-API-call usage stays in the per-assistant-message handler, where
   the LATEST message's values represent the FINAL API call's prompt
   size — exactly what determines whether the next turn will fit.

   Also added a per-message usage-extraction block to the assistant
   branch of process_stream_line (it previously had no usage extraction
   at all, relying entirely on the result handler — which made my
   first attempt at this fix produce zero values). parse_stream_json_output
   already had the equivalent block (lines 211-215).

   Base image rebuilt; agent-testfix recreated onto the new image
   (image sha 0a1e20b40da1).

2) Backend (services/task_execution_service.py)
   Replaced `context_used = metadata.input_tokens` with
   `cache_read + cache_creation` (with input_tokens fallback when
   caching isn't engaged). input_tokens is sometimes the disjoint
   fresh value and sometimes inflated by the agent server's
   modelUsage.inputTokens override on tool-call turns. cache_read
   and cache_creation come straight from Anthropic's usage object
   and (post agent-server fix) are reliable per-call values that
   monotonically reflect the cached conversation prefix.

3) DB (db/sessions.py)
   total_context_used is now a HIGH-WATERMARK (MAX of prior + new),
   not the latest value. Per-turn context naturally oscillates by ~2x
   between text-only and tool-call turns; the watermark gives users
   a stable monotonic upper bound on session pressure that only goes
   up.

   Capped the watermark at total_context_max as a safety belt against
   any future agent-server bug that emits cumulative-billing token
   counts. Genuine per-call peaks should never exceed the model's
   context window — if they do, that's an accounting error not a
   real overflow, and the UI should display 100% rather than 648%.

4) Frontend (stores/sessions.js)
   Bumped the Axios timeout on the session turn endpoint from 305s
   (~5 min) to 7260s (= TIMEOUT-001 cap of 7200s + 60s slack). The
   session turn endpoint is synchronous and may legitimately run for
   the agent's full execution timeout. With the previous 305s ceiling
   the browser threw a misleading "failed" toast on tool-heavy turns
   that ran longer; the response still landed in the DB and the UI
   recovered after a page refresh, but the user saw a phantom error.

Verified end-to-end with a 6-turn mixed sequence (text + tool-call):
per-call cache_read now reports ~11636 on text…
vybe added a commit that referenced this pull request Jun 5, 2026
…1081) (#1086)

Replace the push actor-model with pull / work-stealing in
TARGET_ARCHITECTURE.md, per the 2026-06-05 design review. The backend
owns one durable per-agent queue (schedule_executions); agents pull when
they have free capacity; capacity is physical (worker count); recovery is
lease-expiry re-delivery of the same execution_id.

- Governing principle #5, data layer (queue→Postgres, Redis loses the
  mailbox/slot-ZSET), agent runtime, and observability updated to match.
- Adds the side-effect idempotency contract (the hardest open problem,
  #1084) and the Postgres-before-the-queue sequencing constraint (#300).
- Reconciles the tracking table + open questions with the new issues:
  #1081 umbrella, #1082 status-as-projection, #1083 fire-and-forget,
  #1084 effect-idempotency, #1085 herd controls (all under Epic #1045);
  #946 reframed as the pilot, #307/#526 as gate→alert.
- Includes the team announcement record.

Co-authored-by: Eugene Vyborov <eugene@beingluminous.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
dolho added a commit that referenced this pull request Jun 19, 2026
OSS side of two-factor auth (enterprise issue #5). All IP (TOTP verify,
secret store, recovery codes, policy) lives in the private trinity-enterprise
submodule; this PR adds only the edition-agnostic seam + the entitlement-gated
Vue surface, consistent with the open-core pattern (#847).

Backend seam:
- services/mfa_gate.py — single hook the enterprise provider registers into.
  No provider (OSS-only build) → gate_login() returns None → login unchanged.
  Fail-open on provider error (never lock everyone out of login).
- dependencies.py — create_mfa_challenge_token / decode_mfa_challenge; a
  challenge-scoped token is rejected as a session token in get_current_user
  and decode_token.
- routers/auth.py — /token and /api/auth/email/verify consult mfa_gate after
  the first factor; when a second factor is required they return a short-lived
  challenge token instead of an access token (audited as mfa_challenge_issued).
- models.py — Token gains optional mfa_required / challenge_token fields.
- Dockerfile — pyotp (used by the enterprise module running in this image).

Frontend (gated by `2fa` in GET /api/settings/feature-flags):
- Settings → Security tab (TwoFactorPanel): self-service enroll/confirm/
  disable/recovery; admin-only org policy.
- Login.vue: second-factor step (verify + forced-enroll) after the password/
  email factor; QrCode.vue renders the otpauth URI (graceful manual-key
  fallback). auth store carries the challenge through to the real token.

OSS-only builds: no Security tab, no login step, /api/enterprise/2fa/* → 404.

Related to Abilityai/trinity-enterprise#5

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
dolho added a commit that referenced this pull request Aug 31, 2026
…ag over-claimed (ent#430)

All three blockers from the review, each verified rather than argued.

1. THE FEATURE WAS INERT. `client_portal/asks/router.py` declares `answer_ask`
as a plain `def`, so FastAPI runs it through `run_in_threadpool` — a worker
thread with no event loop — and `asyncio.create_task` raises
`RuntimeError: no running event loop` there. The `except` swallowed it, so every
client answer recorded the answer and dispatched nothing: byte-for-byte the
behaviour this PR exists to remove.

Fixed in `spawn_resume_dispatch` rather than by flipping the route to
`async def`, for the two reasons the review names: the route does blocking DB
I/O, so `async def` alone would move it onto the loop; and ent#430's stated
shape is ONE dispatch site, which moving the spawn back out to the caller would
undo. It now detects the absence of a loop and hops back via
`anyio.from_thread.run_sync` — Starlette's threadpool is anyio's, so the portal
is always there on this path. Any future sync caller inherits the fix.

A thread anyio does not own reaches neither branch. That is not a production
shape, but it must not become the silent no-op this change removes, so it raises
with the cause named instead.

2. THE RACE LOSER SPENT MONEY. `respond_to_operator_queue_item` returns None
only when the row is GONE; when the row exists and has left `pending` — the race
that actually happens — it returns a TRUTHY dict carrying `_status_conflict`,
having written nothing. `if not updated` fell straight through it. The loser
then dispatched a paid execution for an answer not in the database, and because
the idempotency key hashes the response text, the loser's differing text yields
a different digest: one queue item, two paid dispatches. `routers/operator_queue.py`
already pops that flag before its own spawn; this is that rule, not a new one.
Popped, not read, so the sentinel cannot serialize to the client.

3. `resume_requested` OVER-CLAIMED. It was computed after the swallowed spawn
from the opt-in flag alone, so a spawn that raised still answered `true` — the
exact failure AC #5 names, and given (1) that was EVERY production answer on an
opted-in agent. It now reports what was actually scheduled.

TESTS — the reason all three survived 24 green checks is that every existing test
replaced `spawn_resume_dispatch` with a synchronous lambda, stubbing out the one
call whose runtime context was the defect. `test_ent430_dispatch_actually_runs.py`
drives the REAL spawn from a REAL anyio worker thread (the production context,
not an approximation) and asserts the premise before the behaviour. The lost-race
test uses the truthy `_status_conflict` shape that actually occurs, not the
`None` shape that does not. Mutation-checked: reverting fix 1 turns 1 red, fix 2
turns 3 red, fix 3 turns 2 red.

Writing those tests also caught a stubbing bug of my own, worth recording because
it is the trap that hid the original: patching only `sys.modules` leaves
`from services import operator_resume_service` resolving the PACKAGE ATTRIBUTE,
so the real function ran anyway. Both paths are patched now.

Related to ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dolho added a commit that referenced this pull request Aug 31, 2026
…e caller (ent#430)

Non-blocking findings from review pass 2. The three blockers landed in d11956a.

STATUS. `_project` mapped every row to pending/expired, so the response to a
just-recorded answer read `status: "pending"` beside `resume_requested: true` —
one row reporting both that nobody has answered it and that answering it started
work. Harmless while the second field did not exist; contradictory once it did.
`_status_of` adds `answered` (`responded`/`acknowledged`), reachable only from
the answer response since the listing carries neither. Answered is checked
BEFORE expiry — an answer that landed is a fact, and an `expires_at` that has
since passed does not un-answer it; the obvious refactor is to test expiry first,
which would make a slow client's own answer vanish, so the ordering is pinned.

The existing test asserted `out.status in ("pending", "expired")` with the
comment 'the point is it returned at all' — it was papering over exactly this.
It now asserts `answered` and, on the spawn-failure path it covers, that
`resume_requested` is False.

THE TWO READS. `_resume_requested`'s docstring claimed it read 'the SAME
accessor … so the two cannot disagree'. True of the accessor, false of the
instant: it is a second read a task hop earlier, and an owner disabling the
opt-in in between gets `true` and no resume. Collapsing them is not the fix —
they answer different questions (one must produce a value for THIS response, the
other is the authority at the moment it would spend), so the window is stated,
with AC #5's own remedy named, rather than described away.

DOCS. architecture.md's ent#329 section described a single caller and stated the
CAS-win property the second caller broke. It now carries the second caller, the
truthy-`_status_conflict` shape that defeated `if not updated`, the
sync-endpoint/no-loop defect and its `anyio.from_thread.run_sync` fix, and what
`resume_requested` actually reports.

Related to Abilityai/trinity-enterprise#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Sep 8, 2026
…not hold (#2434) (#2594)

* docs(requirements): unmeasurable durations are NULL (#2434)

Rule #1 — requirements before implementation. Adds §10.15.1 under the #1804
terminal-activity contract it amends: a fabricated duration_ms overflows the
PostgreSQL int4 column past 24.855 days, and because every sweep batches its
UPDATEs in one transaction the first overflow rolls back the whole batch, so no
stale row anywhere is ever closed again.

Records the rule (NULL at fabrication sites, one representability helper at
measured writers), why the column is deliberately NOT widened, and that no
schema change means Invariant #9's dual-track migration is not engaged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(helpers): one representability helper for duration_ms (#2434)

duration_ms_between(started_at, completed_at) folds both ends of the range into
one rule: max(0, ...) at the low end (#1832, unchanged) and None above the
PostgreSQL int4 ceiling, where a value that large means started_at is stale, not
that the work ran 24.8+ days.

Takes both datetimes as parameters and never calls now() — the backend passes
aware datetimes (parse_iso_timestamp) and the scheduler naive ones
(parse_scheduler_ts), so resolving "now" internally would mix the two and raise.

Vendored into src/scheduler/utils.py as a behavioural mirror, not an import
(Invariant #16): the scheduler is a separate package that cannot reach
src/backend, and it DOES run against PostgreSQL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cleanup): a fabricated duration_ms is NULL, not a number int4 cannot hold (#2434)

Six fabrication sites now record NULL unconditionally — the four sweeps in
db/schedules/cleanup.py and the two bulk closers in db/activities.py. The sweep
is inventing the end time (`now` is when the watchdog noticed, not when the work
stopped), so there is no measurement to record at any magnitude. Past 24.855
days the fiction was not even storable: duration_ms is a PostgreSQL INTEGER, and
because each sweep batches its SELECT and its whole per-row UPDATE loop in ONE
transaction, the first overflow aborted the batch and rolled back rows already
updated — so one 33-day row froze every stale execution and activity in the
fleet, forever, while each sweep logged itself complete.

The three measured writers route through duration_ms_between instead:
db/schedules/executions.py::update_execution_status and both
src/scheduler/database.py finalizers already had max(0, ...) from #1832, which
guards only the low end — max(0, 2_850_000_000) is still 2.85 bn.
db/activities.py::complete_activity had no guard at either end. Each logs a
warning on the high branch so the number survives where an operator can find it;
NULL alone collides three meanings (swept, unrepresentable, bulk-terminated).

finalize_orphaned_skipped_executions keeps duration_ms=0 — a skipped row ran for
zero time, and that IS a measurement.

Also bounds models.TaskResultPayload.execution_time_ms to [0, 2**31-1]: it is
agent-supplied, was unbounded, and lands in two Integer columns, so the #1083
async result callback was the same class on a hostile input path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(guards): the duration guard sees keyword sinks, and max(0,..) no longer passes (#2434)

The #1832 AST guard would not have caught this fix's own shape. It yielded only
ast.Assign nodes bound to a bare `duration_ms` name, but the sink at all six
watchdog-sweep sites is an ast.keyword — `.values(duration_ms=...)`. Adding those
files to _WRITERS without the extension would "work" only while a redundant
`duration_ms = None` local survived at each site: a line whose sole purpose is to
feed a scanner, which the next reviewer or formatter inlines, silently neutering
the guard on all six at once.

So: the walker yields ast.keyword sinks too; acceptance is classified per site
(fabricated -> None, skipped -> literal 0, measured -> duration_ms_between)
rather than blanket-accepting None, which would let a measured writer silently
discard a real datum; and a bare max(...) is REJECTED — it guards only the low
end, and max(0, 2_850_000_000) is still 2.85 bn.

Pass-throughs are skipped narrowly so the guard stays fail-closed: a bare Name
(judged at its Assign) and `row["duration_ms"]` (a row->model read-back that
cannot produce a value the column did not already hold). `row["anything_else"]`
and any computation still have to satisfy the guard.

Non-vacuity proven, not assumed — run against origin/dev source the extended
guard flags exactly the ten census sites (cleanup 138/226/295/403, activities
229/305/504, executions 451, scheduler 617/1282) and zero post-fix.

The scheduler parity test gains duration_ms_between across both tz shapes
(the backend feeds it aware datetimes, the scheduler naive) with the int4
boundary pinned on both sides, plus a check that neither copy resolves "now"
itself — which would subtract mixed shapes and raise for one of the two callers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(2434): the regression proof — batch integrity on real PostgreSQL

Red-proof run against a pristine origin/dev worktree + postgres:16-alpine:
19 failed / 6 passed pre-fix, RED on BOTH backends for DIFFERENT reasons —
[postgres] raises psycopg2.errors.NumericValueOutOfRange (the reported defect
verbatim), [sqlite] fails the IS NULL assertion because SQLite stores an 8-byte
int and happily persists 2851200002. Post-fix: 25 passed, 12 of them [postgres].

A direct probe on that same baseline confirms the property the test exists for:
seed one 33-day row and one healthy 10-minute row, call the sweep, and after the
raise BOTH are still `running` — begin() rolled back the row it had already
updated. That is why case 1 seeds two rows: a single-row seed would pass against
a "fix" that merely caught the exception per row.

Three things are load-bearing and stated in the docstring: it calls the DB-ops
layer (cleanup_service catches Exception and returns falsy, so a service-layer
test passes pre-fix); every case asserts the writer's return value before
touching duration_ms (IS NULL is trivially true for a row never selected, and
two sweeps have narrow predicates — claude_session_id and the dispatch
activity_type — that make that easy to hit by accident); and it does not reuse
test_1771c's insert_execution, whose #2243 anchor is deliberately relative to
now so durations stay bounded.

Covers all four execution sweeps, both activity sweeps, both measured writers
(including that they still record a real duration), sweep idempotence, and E-01
clearability — built on AgentSnapshot.running_exec_ids, which is what E-01
actually reads, not terminal_rows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(1804): the bulk close now writes NULL — flip the one collateral assertion (#2434)

test_1804_recovery_closes_activity.py asserted `duration_ms is not None` on the
bulk closer. That site FABRICATES its end time, so it now records NULL.

Verified as the only collateral flip in the repo: the other #1804 assertions are
complete_activity on rows seeded 60s / 900s ago (both still green — the measured
path keeps its number) or non-clobber checks on pre-seeded values, and of the 15
other test files calling a fabrication site, 12 contain no `duration` reference
at all and the remaining three reference constructed values, not swept ones.

Also corrects two comments this decision made wrong:
  - E-03's parenthetical said a future duration_ms-coverage check could scope to
    dispatched rows; a dispatched row swept as stale is now legitimately NULL
    too, so it is unwritable even scoped.
  - test_1771c's #2243 note deferred the column width to a product decision and
    assumed widening. #2434 settled it the other way — the column stays int4 and
    the helper returns None above the ceiling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: the sweeps record NULL, and a wedged instance has a runbook (#2434)

architecture/execution.md gains "Unmeasurable Durations Are NULL" directly under
the #1804 close contract it amends — the core architecture.md is untouched
(editorial rule 4 + the 500-line budget).

cleanup-service.md: the four sweep SQL blocks now show `duration_ms = NULL`
(the `= 0` on finalize_orphaned_skipped stays — a skipped row ran for zero time),
an error-table row stating that the blast radius is the batch and not the row,
and a "Wedged pre-upgrade instance" runbook: the two log lines that identify it,
that the remedy is the upgrade RESTART (boot recovery self-heals; no manual SQL
post-upgrade), and the pre-upgrade workaround SQL for an instance that cannot be
upgraded yet.

activity-stream.md: the literal duration code block was stale — replaced with the
helper call plus a table of what it returns at each end of the range.
dashboard-timeline-view.md: the one user-visible change, stated as such — a swept
activity now draws a 30s bar flagged `estimated` instead of a fabricated
~120-minute one, through the existing `|| 30000` path at ReplayTimeline.vue:668.
No frontend change was needed or made.

task-execution-service.md: the terminal write's max(0, ...) -> the helper, and
why the clamp was never coverage at the high end.

Plus the Recent Updates row in the feature-flows index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(2434): restore the evicted db.* modules — the lint's rule was load-bearing

tests/lint_sys_modules.py flagged four bare sys.modules.pop() calls and named the
fix: the _STUBBED_MODULE_NAMES + autouse _restore_sys_modules pair (precedent
test_1832_duration_clamp.py). Adding it did not just satisfy the linter — it
fixed a real order-dependence. Without the restore, the fixtures left db.* popped
after teardown, and under pytest-randomly (which CI runs across 3 seeds) a later
test in another file inherited the eviction: `relation "schedule_executions" does
not exist`. Reproduced before, gone after; verified 25/25 on 6 seeds standalone
and 235/235 on 3 seeds against real PostgreSQL alongside test_1832 / test_1713 /
test_1771c / both test_1804 files.

`models` is deliberately NOT in the list — ActivityCloseOutcome is compared by
identity, and it lives in models precisely so an evicted db.* cannot make that
comparison silently go False at full-suite scale (#1804).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci(schema-parity): a required PostgreSQL tier, scoped by marker (#2434)

TEST_POSTGRES_URL appeared NOWHERE in .github/, so db_harness.available_backends()
returned ["sqlite"] in every CI run and every db_backend-parametrized test ran
SQLite-only — where an int4 overflow cannot fail. That is how #2434 shipped, and
how #2243 hit the same overflow in CI without the product fix landing.

The tier goes in schema-parity (required, unconditional, already path-matching
db/** and utils/helpers.py) rather than pg-migrations, whose own header says
"NOT a required check: keep it advisory" and which runs zero pytest — the same
reasoning architecture.md records for check_alembic_heads.py.

Selection is `-m requires_postgres`, NOT the plan's `-k "postgres or
dual_backend"`. Measured before wiring it rather than assumed: the wide
selection is 75 failed / 26 errors against a real postgres:16-alpine, and the
same five worst files produce an IDENTICAL "2 failed, 153 passed, 4 errors" on a
pristine origin/dev worktree — so those are pre-existing (those tests have never
run on PostgreSQL), not regressions from this change, and making them a REQUIRED
gate would brick every schema-touching PR on arrival. The marker is an opt-in
seam: a future test whose assertion can only fail on PostgreSQL adds one line
instead of editing a workflow. Both limits of that choice are written into the
workflow header rather than left implicit.

Verified by extracting the shipped `run:` block and executing it verbatim against
a disposable postgres:16-alpine — exit 0, 25 passed under whole-tree collection
(the real CI condition, which an isolated file run does not reproduce). The
"fail-loud on empty selection" claim in the comment is verified too: a typo'd
marker exits 5, not 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: /sync-feature-flows finds two flows the manual pass missed (#2434)

Both carried a literal, now-stale code block: tasks-tab.md documents
update_execution_status (the bare subtraction is now duration_ms_between) and
replay-timeline.md documents the bar-width behaviour (a swept row now takes the
`|| 30000` estimate path and is flagged isEstimated, instead of snapping to a
fabricated ~120-minute width).

activity-monitoring.md is annotated as reviewed-NOT-overlooked rather than
edited: its duration lives in agent_state.session_activity, an in-process dict on
the agent server, so there is no int4 column to overflow and no shared
transaction to abort. The unguarded subtraction is correct there, and saying so
is what lets the next reader tell a decision from an omission.

Both new flows added to the index row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(guards): the None acceptance is scoped to the site, not the shape (#2434)

/review found the guard extension had loosened the guard for the writers it
already covered. Two mutation-confirmed holes in tests/unit/test_1832_duration_clamp.py:

- `_is_guarded` accepted `None` globally, so replacing `duration_ms_between(...)`
  with `None` inside `update_execution_status` or either scheduler finalizer —
  discarding a real measurement — passed. Acceptance now takes the enclosing
  FUNCTION name and allows `None` only inside the six fabrication sweeps.
  Function-level, not file-level: db/activities.py holds a fabricator and the
  measured `complete_activity` in the same file.
- `_is_pass_through` skipped ANY `ast.Name`, but the scanner only judges an
  `Assign` whose target is literally `duration_ms` — so `elapsed_ms = int((c - s)
  .total_seconds() * 1000)` + `.values(duration_ms=elapsed_ms)` bypassed the
  guard entirely, a shape services/fan_out_service.py already writes.

Also corrects the models.py comment on the `execution_time_ms` bound: the value
does not reach an int4 column today (schedule_executions.duration_ms is
recomputed; the chat columns are written from the backend's own measurement), so
this is a boundary check that keeps the path from becoming one — not a live
overflow path. The bound itself is unchanged.

Learning appended: extending a source guard to new sites can loosen it for the
old ones — pass the site into the predicate, and mutate at a new site AND an old
one before believing it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(2434): the riding-along execution_time_ms bound had no test

`ExecutionResultEnvelope.execution_time_ms` gained `ge=0, le=2**31-1` in
f554e6e with zero coverage — a validation constraint on an agent-supplied
field that nothing would catch being reverted.

Enforcement is at the contract, so FastAPI rejects the body before
`agent_execution_result` runs. The cases assert `ValidationError` / a wire-level
422 and that the route body never ran, which is stronger than asserting anything
inside the handler. Both ends of the bound (#1832 was the NEGATIVE twin), the
ceiling not off-by-one, and `None` still accepted — an agent that never measured
a turn omits the field, and rejecting that would fail every callback rather than
the malformed ones.

Two cases beyond the obvious:

- A vacuity guard: a representable value MUST reach the handler. Without it a
  bound that rejected everything, or a route wired to the wrong model, would
  satisfy every rejection assertion in the file.
- The cross-surface half (Invariant #5): the bound necessarily introduces a 422
  on a surface whose own docstring says the size caps avoid 422 because "the
  agent's retry logic special-cases status codes". That is safe only because 422
  is already in the agent's `_PERMANENT_STATUSES` — dropping it there would
  silently convert this hardening into a retry storm until the lease deadline,
  and the two files live in different images so nothing else would catch it.
  Read by AST, never imported: the backend test island cannot import
  `agent_server`.

Red-proof on pristine origin/dev (c04750e): 7 failed / 6 passed. On the branch:
13/13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(2434): register requires_postgres where pytest actually reads it

`tests/unit/pytest.ini` wins the rootdir for `pytest tests/unit`, so the
marker registered in `pyproject.toml` was inert — dead config that reads as
live. CI proved it: the new PostgreSQL tier passed 25/25 while emitting
`PytestUnknownMarkWarning: Unknown pytest.mark.requires_postgres`.

Selection was never broken (25 selected before and after, verified against a
real postgres:16-alpine both ways), so this is not a fix to the gate — it is
removing a permanently-warning line from the one tier whose safety property is
that an empty selection fails loud. A warning stream that is always noisy is
where a typo'd marker would hide.

No strict-markers: the repo's widespread unregistered `pytest.mark.unit` would
turn that into a mass failure, and that is not this issue's scope.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M3oj1NXVXmt9hMqBirHjb5

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: sim <sim@example.com>
louisss1016 pushed a commit to louisss1016/trinity that referenced this pull request Sep 13, 2026
…o it inverted the operator's intent (Abilityai#2411)

`GET /api/settings/ops/config` nests settings TWICE, and `Settings.vue` read
them flat:

    { settings: { ssh_access_enabled: { value: "true", default: "false", … } } }

`response.data.ssh_access_enabled` is `undefined`, `undefined === 'true'` is
`false`, and the switch rendered OFF whatever was stored. Because the click
handler computes `!sshAccessEnabled.value`, the FIRST click then always sent
`true` — an operator with ephemeral SSH enabled, clicking to DISABLE it,
re-enabled it and watched it render as off. The write path was always correct,
so the mismatch was one-sided: a change appeared to stick for the session and
silently reverted on reload, which is why it went unnoticed.

## The issue's own suggested fix is not sufficient

Abilityai#2411 proposes `response.data.settings?.ssh_access_enabled`. That is still
wrong, for the same underlying reason: `get_ops_settings` builds a DESCRIPTOR
per key (`value`/`default`/`description`/`is_default`), so it resolves to an
object and `object === 'true'` is `false` — the toggle would still have rendered
OFF while looking fixed. The `.value` hop is the one that matters. Pinned by its
own test so the half-fix cannot be reintroduced as a "simplification".

## Why a pure module rather than a one-line edit

`Settings.vue` cannot be mounted in this project's test setup — `@vue/test-utils`
is not a dependency and vitest runs `environment: 'node'` — so a rule kept
inline is a rule no test can reach. That is exactly how a one-line read bug
survived inside a security control, and it is the reason the issue says the
value is in the test rather than the fix.

`utils/opsSettings.js` owns both spellings: `readOpsBool` for the read and
`opsBoolValue` for the write, together in one file instead of at opposite ends
of a 3500-line SFC, which is what let them drift in the first place. It accepts
the descriptor form and a bare string, since the only thing separating them is a
wrapper the reader does not need.

Everything unreadable degrades to `false` (AC Abilityai#4) — absent key, absent
`settings`, null payload, a number, `settings` that is not an object. `false` is
the SAFE direction: `ssh_access_enabled` defaults to `"false"` server-side, and
a security control that cannot read its own state must not claim the permissive
one.

The endpoint is untouched, per the issue's explicit instruction: `PUT
/ops/config` takes the nested shape and the other ops readers depend on the
current contract. The endpoint is the side that is right.

## Verification

18 tests, and each side of the fix mutation-checked rather than assumed:

- reader reverted to `payload?.[key]`            → 4 failed
- SFC reverted to the flat read that shipped     → 2 failed
- restored                                        → 18 passed

AC Abilityai#5 is pinned rather than eyeballed once: a test asserts `ops/config` is read
in exactly one place, so a second GET added later has to come through the same
reader. Full frontend suite 64 files / 1454 tests; `vite build` clean.

Closes Abilityai#2411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
louisss1016 pushed a commit to louisss1016/trinity that referenced this pull request Sep 13, 2026
…ent#430)

Slice 5 of ent#364, and the gate: until now the client route recorded an answer
and returned. The operator route called `spawn_resume_dispatch`; this one did
not. So an ask addressed to a Workspace client — the entire point of
ent#364/Abilityai#428/Abilityai#429 — was recorded, reached the agent's queue file in about three
seconds, and re-triggered nothing.

Measured on a live instance before this change: answered from the Workspace,
`operator-queue.json` flipped to `responded` with the answer in under 3s, and no
execution followed.

Unblocked because ent#329 is in dev.

WHAT THIS ADDS: one call. ent#430's body rules out the alternative — "a second
dispatch surface for the same event is how the cost, trigger-label and
loop-prevention questions get answered twice, differently" — so the per-agent
opt-in, the idempotency key, the audit row and the failure handling all stay
inside `maybe_dispatch_resume`. AC Abilityai#2 and AC Abilityai#3 are satisfied by REUSE rather
than by re-implementation, and the tests assert the CALL for that reason.

Four properties, each load-bearing:

* Hung off the CAS WIN only, like the operator route. The 409 above already
  returned for a lost race, so reaching the dispatch means this answer is the
  one that landed — two people answering at once produce one resume.
* `updated`, never `item`. The pre-answer read still says `pending`; a resume
  handed that row acts on an ask that does not yet carry its answer. Looks
  identical in a green test, which is why there is one for it.
* The spawn is wrapped. It is fire-and-forget, but a raise ON THE CALLING LINE
  would still propagate, and a 500 after the CAS landed would tell the client
  their answer failed while it is committed and already on its way to the agent.
  The answer is the thing that must not be lost.
* Abilityai#2376's choice validator runs first, so an answer that was never offered
  cannot spend.

AC Abilityai#5 — `resume_requested` on the answer response, read from the SAME accessor
the dispatch gates on, so the two cannot disagree about what is about to happen.
It reports INTENT, not success: the dispatch is backgrounded, so at that moment
the only honest claim is whether it will be attempted. Fails CLOSED — an
unreadable flag claims nothing, because over-claiming is exactly the failure
AC Abilityai#5 names ("the ask does not read as resolved while nothing happened").

RESIDUAL, stated rather than implied: a dispatch that fails AFTER this point
surfaces as a FAILED execution row plus an `operator_resume_dispatch` audit
entry (ent#329) — operator-visible, and a client cannot see either. The client
half of AC Abilityai#5 is satisfied negatively for now: the ask surface says nothing
about work starting, so it cannot mis-claim. `resume_requested` is the field a
surface needs to say something true; consuming it is an ent#429 UI change and is
deliberately not in this PR.

The per-agent flag DEFAULT IS UNCHANGED (`operator_resume_enabled`, OFF,
owner-only). "Turn the flag on" is an operator action per agent, not a code
default: flipping it would hand every shared agent's client a spend button,
which is the one thing AC Abilityai#3 rules out.

Verification: 145 passed across the asks/ent#329/ent#364/Abilityai#428/Abilityai#429/Abilityai#2376
selection. Mutation-checked — removing the dispatch (4 red), passing the
pre-answer row (1 red), and making the opt-in read fail open (1 red).

Closes ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
louisss1016 pushed a commit to louisss1016/trinity that referenced this pull request Sep 13, 2026
…otfix

`main` carried five commits that never came back to `dev`, and the one that
matters is the Abilityai#2398 hotfix (Abilityai#2406): `_is_sensitive_kv_key` was still
`re.search(r'(?:.*TOKEN.*)\Z', key)`, quadratic in key length, on a thread
holding the GIL. Measured 0.32s at 4 KB and 40s at 44 KB; a py-spy watchdog
caught it 14 times across three agents. `key` is whatever preceded an `=` in
`_KV_LINE_RE` and is unbounded, and recursive `sanitize_dict` reaches it from
stream-json tool results — so multi-KB keys are the normal case, not an
adversarial one.

Found by /release pre-flight while cutting v0.9.5-rc1. That RC becomes the
DigitalOcean Marketplace snapshot (Abilityai#2281/Abilityai#2471), so tagging `dev` as-is would
have shipped a known, production-observed DoS in a public image.

Each branch had a sanitizer fix the other lacked, which is why this is a merge
and not a cherry-pick in either direction: `dev` has the Abilityai#2208 `sk-` family
consolidation (`sk-svcacct-` and anything after it), `main` has the Abilityai#2398
containment rewrite. Both are present on both copies after the merge, verified
by grep and by the 195 tests across the five sanitizer suites.

Conflict resolution, for the record:

- `src/backend/utils/credential_sanitizer.py` and its agent-server sibling
  both auto-merged cleanly. They are NOT byte-identical vendored twins — the
  Invariant Abilityai#5 mirror set is credential_paths / model_context / safe_yaml /
  mcp_validator — so no parity obligation attaches and the independent merges
  are legitimate.
- `src/backend/enterprise` resolved to **dev's** pin (90f2f2cb), which is three
  commits AHEAD of main's (2a5def35) and has it as an ancestor — verified, not
  assumed. Taking main's would have rolled the paid tree backwards.

Verified: 195 tests pass across test_2398 / test_2208 / test_1661 /
test_credential_sanitizer_{agent,backend}; alembic version line still resolves
to exactly 1 head (0049); the merge touched no migration on either track.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K9bL6DZ8QVyrSvPybp5ZuK
vybe pushed a commit that referenced this pull request Sep 14, 2026
…L, not one per sync (#2744) (#2777)

* docs(skills): the terminal legacy-adoption refusal is bounded by class (#2744)

Rule #1: the requirements delta lands before the code.

§21.1.3 described adoption as "idempotent and fail-soft" and said nothing
about the refusal's ALERTING. `_adopt_legacy_clone` runs as the first
statement of every `sync_library()`, so on an install that is past migration
but still carries a non-matching `skills_library_url` the terminal refusal
files a fresh `priority: "high"`, `expires_at: None` operator-queue item on
every sync — unattended under the ent#236 auto-sync loop (300s floor ⇒ 288
rows/day), never expiring, and un-dismissable because each row carries a new
timestamped `request_id`.

The rule this states: that refusal is the designed resting state of a
migrated install, not a failure, so it is `low` + `logger.info` with a stable
URL-derived id whose family prefix is reserved; the two genuine failure
branches keep `high` and their repeat-visible ids by product decision.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

* test(skills): pin the legacy-adoption alert's cadence, severity and echo (#2744)

TDD — RED on this commit, green on the next. Proven red for the right reason,
not merely red:

  test_n_syncs_..._exactly_one_item          5 distinct timestamped ids
  test_a_different_refused_url_...           frozen clock ⇒ both URLs share one id
  test_the_terminal_refusal_is_not_high...   priority "high", logger.error
  test_the_stable_id_is_reserved_...         'skills-legacy-adoption-<ts>' unreserved
  test_a_pat_bearing_url_is_never_echoed...  the PAT is in context.url AND the log

  test_the_actionable_branches_keep_high...  GREEN on base, by design (AC 4)

The last one is the anti-regression half: it pins behaviour the fix must NOT
change, and goes red only if the low/stable-id treatment is applied to all
three call sites instead of the one the issue names.

Harness notes, both load-bearing. `_record_adoption_failure` imports
`utc_now_iso` INSIDE its body, so the patch target is `utils.helpers` —
patching `services.skill_service.utc_now_iso` binds nothing and yields a
vacuous test. And `validate_skills_library_url` does a live
`socket.getaddrinfo`: on a sandboxed resolver a terminal-branch test would
silently drive the validation-reject branch and fail as "5 distinct ids" /
"priority is high", reading exactly like the fix regressing — so DNS is
stubbed to the `gaierror` that function already tolerates, and every
terminal-branch test additionally pins which branch produced its item.

New file rather than an append to test_ent346_skills_source_injection.py:

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

* fix(skills): the legacy-adoption refusal files one row per URL, not one per sync (#2744)

`_adopt_legacy_clone()` is the first statement of every `sync_library()`. On an
install that is past migration but still carries a `skills_library_url` matching
no configured source, the terminal "already has sources" branch called
`_record_adoption_failure`, which minted a TIMESTAMPED `request_id` at
`priority: "high"` with `expires_at: None`. One permanent, high-priority,
operator-unclearable row per sync, forever — 17 of them (~17% of everything
pending) on the reporting install, and 288/day at the ent#236 auto-sync floor.

Two behaviour changes, on ONE branch:

  * a STABLE, URL-keyed id (`skills-legacy-adoption-refused-{sha256(url)[:12]}`)
    so `create_item`'s `(agent_name, request_id)` ON CONFLICT DO NOTHING
    collapses N syncs to exactly one row — and, since that conflict target
    ignores `status`, an operator's dismissal finally sticks;
  * `priority: "low"` + `logger.info`, because this is the designed resting
    state of a migrated install, not a failure.

Shaped as a keyword-only `steady_state` flag on the existing emitter rather
than a second method: one #1677 `_ALLOWED_CALLERS` key, and a `False` default
that leaves the two actionable call sites LITERALLY UNCHANGED lines — the
strongest available proof of AC 4. The emitter keeps its name despite now
serving a non-failure; renaming costs the allowlist key and churns a file two
people are editing this week.

The hash is over the RAW `url.strip()`, and is computed INSIDE the try: a
non-str setting value must degrade to a warning and no alarm, not turn a
decorative alarm into a raiser. Normalising the input instead would re-enter
`validate_skills_library_url`, which does a live `socket.getaddrinfo` and can
raise — a network call and a raise path inside a fail-soft alarm.

Two things the stable id makes mandatory, both included:

  * `skills-legacy-adoption-` joins `_RESERVED_ID_PREFIXES`. An id derived from
    an admin-visible URL is guessable, so an agent could pre-create it and
    silence the alarm through the sink's ON CONFLICT (the #1632 C2 class); and
    `is_platform_minted` reads the same tuple to gate the ent#499 responded
    write-back and the ent#329 respond→resume dispatch, which this change makes
    an expected operator action. The FAMILY prefix, so all three call sites and
    the 17 historical rows classify correctly.
  * the URL echo is `strip_url_credentials`-scrubbed. The emitter's docstring
    claimed the credential case was handled, and that was true of `message` and
    of nothing else: `EmbeddedCredentialError` is a `ValueError` subclass, so
    the validation-reject branch is exactly the one a PAT-bearing URL reaches,
    and the raw value landed at ERROR in the Vector-captured log and durably in
    `operator_queue.context` — SQLite, every backup, rendered in the Operating
    Room (Invariant #12, Rule #5). The hash still keys on the raw value;
    scrubbing first would collide two different tokens on one repo.

The #1677 justification is corrected in the same commit: "admin-driven sync
cadence" is false (ent#236's loop is unattended), and the real bound — the only
input is a setting blocked on the generic settings PUT — is co-located as a
comment at the emitter, where it is likelier to stay true.

Unchanged and deliberately so: the `count_skill_sources() > 0` guard itself,
both validators, the grant branch, `expires_at: None`, the `title`/`question`
copy, and `context["alert_type"]`. No schema change — `request_id` and its
unique index shipped in #1631 — so Invariant #9 is not triggered: no
`db/migrations.py` entry and no Alembic revision. Clearing the lingering
`skills_library_url` key stays out of scope pending a separate investigation.

Tests: tests/unit/test_2744_skills_adoption_alert_idempotency.py

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

* docs(operator-queue): record the reserved prefix and the alarm's real bound (#2744)

The reservation of `skills-legacy-adoption-` is enumerated in three live places
and all three now carry it: `requirements/security.md` §26.7's reserved-id
guard, and `operating-room.md`'s two enumerations (the ingestion-guard list and
the #1632 ingestion-caps paragraph).

`operating-room.md`'s "Platform exemption & emitter budget (#1677)" bundled the
skills alarm into a disjunction that includes "operator-driven". That was the
same false claim the `_ALLOWED_CALLERS` justification made — ent#236's
auto-sync drives `sync_library()` unattended on a 300s-86400s timer, so nothing
admin- or operator-driven bounds it. The paragraph now names this emitter's
actual bound, per branch: the terminal refusal is idempotent by a URL-keyed id
(≤1 row per refused URL) and what makes it platform-only is that its only input
is a setting blocked on the generic settings PUT.

Plus the two dated rows (`operating-room.md` Revision History, the
`feature-flows.md` change log) and the Operating Room catalog row.

All three enumerations were ALREADY stale — each omits prefixes the live tuple
carries. Ours is added; their pre-existing drift is deliberately not swept here
(Rule #2) and is named as a follow-up instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

* test(skills): pin the steady-state discriminator and sweep the whole item for the PAT (#2744)

Two gaps the review found in the new file, both in tests only.

`context["reason"] = "already_migrated"` is the discriminator the emitter grew
because one `alert_type` and one title now span both `low` (the benign resting
state) and `high` (a URL that failed validation — the signature of an attempted
injection). Nothing asserted it, so the field the fix added to be read by a
machine could be dropped by a later edit in silence.

The credential test asserted the PAT is absent from `context["url"]` and from
the captured log, but not from `question`. `question` is credential-free only
because neither ent#346 validator echoes the URL in its `ValueError`
(`validate_skills_library_url` names the hostname or the resolved IP;
`reject_embedded_credentials` names neither) — the scrub does not reach it. A
validator message that starts echoing the URL would reopen the leak durably in
`operator_queue.question` with no guard. Sweeping the serialized item covers
every field the emitter writes, not the two that were remembered.

Both were verified to bite: dropping the `reason` key fails the first, and
reverting `strip_url_credentials` to the raw url fails the second.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

* fix(skills): drop the already_migrated context key the owner declined (#2744)

`QueueItemDetail.vue` renders every `context` key, so the discriminator the
emitter grew would have surfaced to operators as a `reason | already_migrated`
row. Put to the product owner as keep / drop / rename; the answer was drop —
nothing reads the key today, so removing it is non-breaking.

The steady-state branch is now discriminated by `priority: low` plus the
`logger.info` level alone. The assertion added in 69274de7 to pin the key goes
with it; it lived inside an existing test, so no test is removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0176XEqK8PTCAK5K6yZURQLv

---------

Co-authored-by: trinity-ability <309458136+trinity-ability@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Sep 15, 2026
… the agent's index.lock (#2742) (#2797)

* docs(git-sync): requirements, architecture and flow for lock-free sync-health polling (#2742)

Trinity Rule #1 — the requirements/architecture delta lands before the code.

What the docs now say:

- requirements/github.md §11.9 gains "Lock-free status read" and "Stuck-lock
  report, never a runtime delete", and the stale-lock hygiene bullet records
  that the boot reap now speaks. §11.8's SyncHealthService bullet names the
  leader lease AND its honest side effect (time-to-sync_failing ~90s → ~180s).
- architecture/agent-lifecycle.md is the single home: why the status read took
  index.lock ~2x/min in every workspace, what replaced it, and why the runtime
  delete was cut on measurement rather than hardened.
- architecture/agent-runtime.md gets a one-line pointer (no second home).
- architecture/background-services.md's Sync Health row states the lease, its
  fail direction and the reason (a feed that raises sync_failing must not go
  dark when Redis does) — per that file's header rule.
- feature-flows/git-sync-health.md: §1a announced boot reap + the observe-only
  runtime report with the three measurements that decide it; a new §2a for the
  agent status handler with the two distinct bounds (35s caller / 90s
  computation) and the ~130s child-budget arithmetic that replaces the doc's
  wrong "~30s worst case"; Files Touched, Testing, Operator Controls; and a
  Known Limitations rewrite that retires the "still blocks the event loop"
  bullet and replaces it with the residuals this change does NOT close.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sync-health): leader lease across uvicorn workers, fail-open (#2742)

main.py starts SyncHealthService in EVERY uvicorn worker and prod runs
`--workers 2`, with no lease — unlike opqueue:leader (#1632), monitoring:leader
(#1464), skills:sync:leader and canary:leader. So every git-enabled agent was
asked for /api/git/status twice a minute, and each ask runs a 30s credentialed
`git fetch origin` inside the container.

`synchealth:leader` in the #1464 shape: SET NX EX, own-lease-only refresh,
release from stop(), acquire + transition logs at the top of _poll_cycle with an
early return for non-leaders.

Two deliberate divergences from the copied shape:

- Release is a single compare-and-delete EVAL, not GET-then-DEL. The non-atomic
  form lets a worker whose lease lapsed between the two calls delete a sibling's
  FRESH grant — two pollers for a whole cycle. EVAL is outside the -@dangerous
  categories denied to the backend/scheduler ACL users.
- Fail direction is OPEN and stated in the docstring: Redis down means every
  worker polls, i.e. today's behaviour. Failing closed would darken the only
  feed that ever raises sync_failing, exactly when infra is already degraded.

The lease has an honest side effect and it is named, not buried.
db.upsert_sync_state INCREMENTS consecutive_failures per failed upsert, so it is
NOT idempotent: two unleased workers reached sync_failing in ~90s, one leader
takes ~180s. Asserted in TestLeaderLeaseAlertTiming — including the load-bearing
fact itself (the same payload polled twice increments twice), so the test cannot
pass against an idempotent upsert and prove nothing.

Also here (same file, same subject): SYNC_HEALTH_POLL_INTERVAL_SECONDS, read at
CALL time in the _maintenance_timeout_seconds shape, parse-guarded and
positive-clamped, with the default DELIBERATELY unchanged at 60s — the AC is
written as "the 60s poll" and the sampler evidence was taken there. poll_interval
becomes a property so the module-level singleton (built at import) still honours
the env; `is not None` rather than truthiness so the tests' poll_interval=0 keeps
meaning "one cycle then exit".

The `service` fixture now stubs get_breaker_redis -> None: with a local Redis it
would otherwise leave a real 30s lease behind and the NEXT test's fresh service
would lose the election and silently poll nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(agent-server): status read is lock-free and sweep-registered (#2742)

`git status --porcelain` takes .git/index.lock on EVERY invocation to refresh
the index — whether or not a rewrite follows — for a window that scales with
index size (0.6ms at one file, 12.8-15.4ms at 20k; 334-473ms measured in a real
container). The backend polls /api/git/status every 60s for every git-enabled
agent, so the platform was taking that lock ~2x/min in every workspace, outside
_REPO_LOCK, racing the agent's own `git add`.

`git --no-optional-locks status --porcelain` on exactly that one call site.
Scoped deliberately: the auto-sync cycle's status (under the repo lock, right
after `git add -A`, proceeding to commit) and the sync/pull bodies keep the plain
form — they are lock-serialized and they WANT the refreshed stat cache. Flag
form rather than GIT_OPTIONAL_LOCKS=0 because run_registered has no env= kwarg,
the env form would silently change the mutating sites too, and only the argv is
assertable in a test. Verified across a clean index, a stat-dirty index,
core.untrackedCache, --untracked-files=all, an fsmonitor hook and a repo with no
.git/index: zero lock sightings and zero rewrites in every case.

The flag closes only the index-lock half, so every status child now routes
through run_registered (#1595's seam): a sweep tick straddling the 30s
`git fetch` still SIGKILLs it, and a killed fetch orphans FETCH_HEAD.lock /
packed-refs.lock — which NO reaper covers, not even startup.sh's (its find is
scoped to refs/ and logs/). run_registered accepts neither capture_output nor
text, so both kwargs are DELETED at all ten sites; a name-only substitution
would be a TypeError ten times over.

Three of those ten are the shared helpers _compute_ahead_behind,
_get_pull_branch and _persist_last_remote_sha, which are also reached from
_conflict_response (the 409 arm of every locked endpoint), sync_to_github and
pull_from_github. So the conversion sweep-registers three children on the
MUTATING paths too. That is desirable — those are the children a repack-length
operation exposes longest — but it is stated and pinned by a test rather than
left for a reviewer to find.

Also here, because the diff was already moving these lines:

- remote_url is now unconditionally redact_url_userinfo'd. The old shape
  special-cased @github.com and returned everything else VERBATIM, so any
  non-github.life-white.uk remote (GHES, GitLab, a host-rewritten origin) put a live
  `https://oauth2:<PAT>@host/...` in the response body — proxied unmodified by
  git_service.get_git_status to the UI and the MCP tool. ent#615 owns the
  broader class; this is the one line of it in this diff.
- _read_sync_state_file gates on st_size (64 KiB) BEFORE reading. That file is
  fully agent-authored and merged wholesale, and the backend reads every agent
  concurrently once a minute, so an unbounded read_text() OOMs both sides.
- The comment beside _REPO_LOCK now says what it does NOT exclude: the agent's
  own git and the backend's docker exec sites. Misreading that is what produced
  a plan to treat a successful non-blocking acquire as evidence of quiescence.

_compute_git_status is extracted verbatim as one blocking callable (the handler
calls it directly for now) so the next commit can put the coalescing in front of
it without also moving the body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(agent-server): coalesce /api/git/status on the event loop (#2742)

Three callers converge on this one route — the backend's 60s sync-health poll,
the UI git panel's own 60s poll while open, and the MCP get_git_status tool —
and each computation runs a 30s `git fetch origin` against one repo. The issue
observed two overlapping fetches 0.9s apart. A guard on the poller alone would
not have satisfied AC2.

Now the first caller starts the computation and every other caller awaits the
SAME future. The slot is keyed on the resolved home path, check-and-set is
atomic because there is no `await` between the test and the assignment, and the
agent server is single-process (agent_server/main.py calls uvicorn.run(app, ...)
with no workers=) — the BACKEND is not, which is what the Redis lease is for.

Coalescing happens ON THE LOOP, and only the computation takes a thread. This
is the load-bearing shape choice, not a style preference: asyncio.to_thread uses
the loop's DEFAULT executor — min(32, cpu+4) = 6 threads on a 2-vCPU agent — and
services/headless_executor.py records that ctx.terminate, auto-sync and
pipe-close deliberately stay on that same pool. Parking followers there is the
#2433 starvation class, where a burst of status callers stalls EXECUTION
TERMINATION. Five callers must cost exactly one to_thread, and a test asserts it
so the rejected shape cannot come back.

Two bounds, deliberately distinct numbers for distinct things:

- _STATUS_FOLLOWER_WAIT_SECONDS = 35 is a CALLER bound, at or just past the
  point every real client has already given up (poller 10s, git_service 30s);
  a longer wait can only produce work nobody awaits. Timeout is a 504.
- _STATUS_LEADER_DEADLINE_SECONDS = 90 is a COMPUTATION bound. The child
  timeouts sum to ~130s nominal before run_registered's post-killpg drain, and
  a slow leader costs no follower threads but DOES hold the slot, so every
  caller in that window 504s. The docstring carries the arithmetic; the flow
  doc's old "~30s worst case" (which is where a 60s follower bound came from)
  is corrected in the docs commit.

asyncio.shield is load-bearing: a follower that times out or disconnects must
not cancel the computation everyone else is waiting on. And the future is
ALWAYS resolved — to_thread -> run_in_executor -> _WorkItem.run catches
BaseException and calls set_exception — so there is no hand-rolled
set_result/set_exception pair and no "leader vanished" fallback. The comment
says not to add one.

`computed_at` admits what coalescing costs: a follower arriving at t=29s of a
30s leader run is served a 29-second-old snapshot, visible as "1 ahead" right
after a successful push. Stamping the age is cheaper and more honest than
claiming coalescing changes nothing. TTL caching was rejected — serving late
followers a fresh run reintroduces the overlapping fetch AC2 forbids.

No 409 on status by design: it is a read, and _with_repo_lock on it would make
every poll a contended write and flap the agent `unreachable`.

D5: test_1920's walk now covers docker/base-image/agent_server/ as well as
src/backend — Invariant #5, "a guard that walks only one of the two trees is
not a guard" (ent#314). Proven by planting an nx=True in an agent-server module
and watching it fail. That tree gets no allowlist row on purpose: it issues no
nx=True set and cannot (agents are not on the platform network, so Redis is
unreachable from one), and #2742's coalescing is a different class entirely.
Its one-home property is asserted directly instead, beside a meta-assertion
that fails loudly if the agent-server root ever moves out from under the walk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(agent-server): announce the startup lock reap and report a stuck lock (#2742)

AC3 has two halves and only one of them was missing. The RECOVERY already
existed and is provably safe: startup.sh reaps index.lock at container start,
where "no git process is running" is definitional because the PID namespace is
empty. What was missing was the OBSERVABLE half — the block was `rm -f`, silent
whether or not it removed anything, so the one moment the platform reliably
heals a wedge produced no evidence that it had.

startup.sh now tests-then-removes, echoes a line naming each lock it cleared
(Vector captures it), and drops ~/.trinity/lock-recovery.json. It also reaps
index.lock under .git/modules/* and .git/worktrees/*, which NOTHING covered
before — its find is scoped to refs/ and logs/ — so a submodule or
linked-worktree wedge used to survive every restart. Proven by extracting the
shipped block and running it against fixtures: /verify-local's agent stage boots
local:test-echo, so startup.sh never executes there and trusting the stage would
prove nothing.

_record_lock_recovery folds that marker into sync_state.last_lock_recovery on
the next status read, then deletes it so the episode is reported once rather
than every minute forever. It goes through a dedicated metrics-only writer:
_write_sync_state_file unconditionally stamps last_sync_at = now, which would
mark a never-synced agent as freshly synced and turn its dashboard dot green.

_index_lock_stuck REPORTS a currently-present lock and never unlinks it. The
runtime delete was cut on measurement, not on taste:

- st_size == 0 is the signature of a LIVE writer, not an abandoned one — git
  creates the lock with O_EXCL before walking the worktree and writes the new
  index into it only at the very end. Measured 100% of a healthy `git add -A`'s
  life: 3.9s on 450MB, 29s at 60k files, 155s under a clean filter.
- st_mtime is stamped at create and never advances, so age measures the
  IN-FLIGHT OPERATION. At t=+130s a healthy add reads size=0 age=130s — both
  naive gates satisfied.
- A wrong unlink is permanent and strictly worse than the wedge: git renames by
  PATH, so a second git's in-flight file gets promoted onto .git/index, the
  corrupting process exits rc=0 with empty stderr, and the 0-byte index is
  cleared by nothing — not the boot reap, not _reap_stale_git_litter, not
  `git reset`.
- _REPO_LOCK would not have helped: it excludes this server's own auto-sync
  cycle and nothing else, not the agent's own `git add`, which is the premise
  of this issue.

So detection is two-point inode stability — the same (st_ino, st_mtime_ns,
st_size) unchanged across >=3 status reads spanning >=15 min — on
time.monotonic(), which a forward NTP step or a live migration cannot move.
The tunable is the sighting COUNT, not a wall-clock age.

Three properties that are each a defect if dropped:

- It resolves the REAL gitdir. `.git` is a FILE for a linked worktree and for a
  submodule, both creatable by the agent in one command, and assuming a
  directory makes the observer look where the lock provably is not. Candidates
  cover <gitdir>/index.lock, modules/*/index.lock and worktrees/*/index.lock.
  A symlinked .git is skipped outright.
- It takes NO repo lock. An lstat needs no mutual exclusion, and holding
  _REPO_LOCK across the observation would make a status poll a brand-new source
  of 409 agent_busy on an operator's POST /api/git/sync.
- It is wrapped end to end in its own except OSError. It runs inside
  _compute_git_status's try, whose tail is HTTPException(500), and
  _fetch_git_status treats any non-200 as None and writes nothing — so one
  EACCES would silently stop the agent's sync-health row advancing and
  sync_failing would never fire either. An observability path must never be
  able to darken the feed it feeds.

The sighting ledger is in memory, deliberately NOT in sync-state (a deviation
from the plan, on safety grounds): a per-tick read-modify-write into the
agent-authored document would race the auto-sync writer and could drop
consecutive_failures or last_sync_status — the observability path corrupting
the feed. It needs no durability either, since the only thing that clears a
genuinely wedged lock is the container restart that also clears the ledger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sync-health): log self-healed index.lock recoveries once (#2742)

The backend half of AC3, plus the AC4 concurrency proofs.

_coerce_lock_recovery / _coerce_lock_stuck rebuild both new fields from values
the backend has checked, and never pass the agent's dict through. That is not
belt-and-braces: sync-state.json is agent-authored, the agent server merges it
wholesale (merged.update(data)), so last_lock_recovery is fully agent-controlled
even on an agent where nothing ever reaped anything — and
git_service.get_git_status proxies response.json() UNMODIFIED to the UI and the
MCP tool.

Three guards, each for a named failure:

- isinstance(value, str) before parsing: parse_iso_timestamp raises
  AttributeError, not ValueError, on a non-str.
- except (ValueError, TypeError) around the parse AND the window comparison: a
  valid-but-naive ISO parses cleanly and then raises TypeError on aware > naive,
  and that raise lands in _sync_agent AFTER the upsert where
  gather(return_exceptions=True) swallows it — so the symptom is a lost alert
  with no traceback. The same guard with the same comment already exists in the
  file this record comes from.
- An explicit UTC offset is REQUIRED. startup.sh always stamps Z, so a naive
  value did not come from us, and the house posture over agent-authored JSON is
  to reject rather than guess.

The window is [now - 1h, now + 60s], with the hour deliberately a floor rather
than 2x the poll interval: startup.sh stamps `at` at container boot and the
agent server folds it in on its FIRST status read, and a cold boot can put
several minutes between those two. The clamp's job is to reject nonsense — the
far-future `at` that would otherwise be "newer" forever — and the DEDUP is what
stops a real record repeating.

Dedup is against the last OBSERVED value. The rejected alternative ("at newer
than the prior row's last_check_at") is not a dedup at all: last_check_at is
re-stamped to now on EVERY upsert, so 9999-01-01 would be newer on every tick,
per agent, for the life of the process — a WARNING asserting a platform action
that never happened. Both halves are tested, so neither can pass by rejecting
everything.

index_lock_stuck is edge-triggered and re-arms after the lock clears. No
operator-queue item and no DB column: the report is a diagnosis, not a decision.
The agent-supplied `path` is dropped at the boundary — it is composed from a
.git the agent can point anywhere and adds nothing to a fleet-level WARNING.

AC4 ships in three phases, strongest first:

- The GATE is deterministic AND genuinely concurrent: a real
  `git status --porcelain` child is SIGSTOP-frozen the instant it takes
  .git/index.lock, and the agent's own `git add` — a second real process — then
  fails rc!=0 with index.lock in stderr. 10/10 locally. SIGSTOP is the only way
  to hold that window open: git runs hooks and filters OUTSIDE the index lock
  (an fsmonitor hook sleeping 1s stretches status to 1279ms while the lock
  window stays 0.8ms). SIGCONT is in a `finally` — --timeout-method=signal
  raises inside the test and a leaked frozen child would hold the lock for the
  rest of the session. Its twin proves the flagged argv cannot be caught at all.
- The lock-sighting sampler (AC1, earlier commit) is the property.
- The threaded witness through the real get_git_status() route is
  self-validating and non-gating: its control arm must reproduce an index.lock
  failure IN THIS RUN or the test skips. Measured through this route the duty
  cycle is ~0.7% (about 95% of each iteration is git fetch), so the margin is
  roughly one failure per run — it skips ~2 runs in 5 here, which is exactly the
  point: it can never pass without having demonstrated it can fail.

Writer rounds write time.time_ns() and failures are classified by
"index.lock" in stderr, never by return code: `git commit` exits 1 with EMPTY
stderr when a round writes content identical to the previous one, which under
check=True is indistinguishable from a lock failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(git-sync): sync the adjacent flows and bound the leader (#2742)

Tail steps: /update-tests and /sync-feature-flows.

Two gaps the plan's test matrix listed and the implementation had not yet
covered, both now closed:

- test_leader_deadline_releases_the_slot — a wedged leader must release the
  in-flight slot rather than monopolise the read. It holds no follower threads
  on this design, but it does hold the slot, so every caller in that window
  504s.
- An AST guard pinning _STATUS_LEADER_DEADLINE_SECONDS to the status body's
  REAL child-timeout budget. They are both 90s today (10 rev-parse + 10 status
  + 10 log + 30 fetch + 10 merge-base + 10 log + 10 remote get-url). Adding a
  child, or widening one, now fails here instead of showing up as an
  unexplained 504 in production — and the flow doc's arithmetic is guarded with
  it. Plus the ordering invariant between the two bounds (caller < computation).

Flow-doc sync — one home per feature, so the two adjacent docs get pointers,
not copies:

- github-sync.md documents GET /api/git/status's own contract, so its endpoint
  row, its _persist_last_remote_sha row (whose child is now sweep-registered on
  the LOCKED sync path too), its stale "line 249-250" call-site reference, and
  its revision history all needed the #2742 delta.
- mcp-git-tools.md: the MCP tool proxies the agent payload verbatim, so it
  gains computed_at / lock_recovery / index_lock_stuck with no signature change
  — and its callers now coalesce with the poller and the UI panel rather than
  stacking a third overlapping git fetch.
- feature-flows.md's Git Sync Health category row and git-sync-health.md's
  index_lock_stuck field list corrected to match what the code actually emits.

The test-runner catalog entry lives in the private .claude submodule and is
committed THERE, on its own branch, deliberately WITHOUT bumping this repo's
gitlink — `git add -A` stages that gitlink silently under
`diff.ignoresubmodules=all`, which is what ejected #2606 from a merge train.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(git-sync): drop the ghp_-shaped placeholder from the redaction test (#2742)

Public repo. The fixture token was never real and is too short to match
GitHub's own secret-scanning format, but a `ghp_` prefix is exactly what the
repo's pre-commit checklist tells a reviewer to grep for, and the test's point
is that URL userinfo is stripped — not that it is a GitHub PAT specifically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(learnings): record the to_thread-deletes-the-accidental-exclusion pitfall (#2742)

Moving the status handler off the event loop removed a mutual exclusion that
sync_to_github/pull_from_github were providing by never awaiting. _REPO_LOCK
was added for the cycle, so it does not cover a newly threaded path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(git-sync): replace the 29x flag-cost figure with fleet measurements (#2742)

Re-measured inside a real agent container on the ext4 workspace volume. The
penalty tracks content bytes re-hashed, not index size: ~1x on 42491 files of
~1 MB, ~390x on 1500 files of 294 MB. The single 29x figure was taken outside
the fleet and sat between the two regimes, describing neither.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sync-health): wire the poll-cadence knob into both compose files (#2742)

`SYNC_HEALTH_POLL_INTERVAL_SECONDS` was read by `sync_health_service` at call
time — a property rather than an import-time copy, specifically so an operator
override is honoured — and then reached no container: it was absent from
`docker-compose.yml`, `docker-compose.prod.yml` and `.env.example` alike. The
knob shipped inert. Proven by rendering rather than grepping: before, the var
does not appear in `docker compose config`'s backend environment at all.

This is the #1056 packaging class (`VOIP_*`), which
`test_ent237_skill_source_env_packaging.py` records as having already recurred
seven times. Prod compose launches standalone — no base-compose merge and no
`env_file:` on the backend service — so the explicit `environment:` list is the
only route in, and wiring dev alone would not have carried over.

The `${VAR:-60}` pass-through form is the correct one here: a cadence has no
"disable" sentinel, so unset and empty must both land on the unchanged 60 s
default. `TestPollIntervalReachesTheContainer` asserts the FORM, not mere
presence, and pins the compose default against `DEFAULT_POLL_INTERVAL` rather
than a literal so the two cannot drift into a container that polls at a
different rate from a laptop. Verified to have teeth by deleting the prod
wiring and watching it go red.

Found by /validate-pr §4.9 on this branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sync-health): the cadence knob must reach the hosted compose too (#2742)

`b6757ba66` wired `SYNC_HEALTH_POLL_INTERVAL_SECONDS` into `docker-compose.yml`,
`docker-compose.prod.yml` and `.env.example` — and missed
`docker-compose.hosted.yml`, the pull-only twin every marketplace and managed-host
channel actually runs. `test_2280_hosted_compose_parity` caught it in CI:
deterministic, all three head seeds, clean in all three base seeds.

That is the bug reproducing inside its own repair. The commit being fixed exists
because a knob never reached the container; it then failed to reach the container
that ships to marketplace installs. The hosted file's own header names the class
and lists its five prior victims (#1039, #1056, #1707, #1871, #2381) — this is
the sixth, and the first to be caught by the guard rather than by an operator.

The line is byte-identical to prod's, since #2280 compares the backend
`environment` list wholesale. No hosted-specific assertion is added here: that
guard is strictly stronger than anything this file could restate, and two guards
over one fact drift apart. The reasoning is recorded in the test's docstring so
the next reader does not "helpfully" add the redundant one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: sim <eugene@beingluminous.com>
vybe pushed a commit that referenced this pull request Sep 15, 2026
…#2804)

Mechanical conflict resolution, per the merge-train note on the PR.

dev's #2742 replaced the single-tree `_iter_backend_py()` with the two-tree
`_iter_guarded_py()` (src/backend + docker/base-image/agent_server), yielding
`(path, root, prefix)`. This branch edited the old single-tree signature, so
git spliced dev's loop body under this branch's header — `root` undefined.

Neither side is correct alone: taking this branch's side reverts #2742's
two-tree walk (re-opening the Invariant #5 failure its own docstring cites);
taking dev's side drops the enterprise/ exclusion this PR exists to add; and a
naive port raises ValueError on every agent-server path, which does not live
under src/backend.

Resolved per-root: `rel = path.relative_to(root)`, exclusion applied inside
dev's iterator. Added `venv/` alongside, so the guard is green on a dev machine
with a local src/backend/venv (it was not before).

tests/unit/test_1920_no_hand_rolled_single_flight.py: 6 passed, including
#2742's test_both_trees_are_actually_walked and
test_agent_server_single_flight_has_exactly_one_home.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Sep 16, 2026
…2791) (#2811)

* fix(session): one platform credential, one 401 verdict, one handler (#2791)

Log out and log back in on the main app with a Workspace tab open from the
previous session, and the NEW session dies within seconds. That tab holds the
old JWT, its 20s poll 401s, and the handler calls `authStore.logout()` — which
removes `localStorage['token']`, i.e. the token the re-login had just written.
The handler never asked whether the credential that failed was still the current
one.

Underneath it, one browser held the platform JWT in two places that could
disagree (the in-memory `axios.defaults` copy vs localStorage re-read per
request), with no `storage` listener anywhere under `src/frontend/src`, and
three separate 401 implementations that had each drifted.

`utils/platformSession.js` makes all three singular.

**One source.** `readStoredToken()` is the only reader. The `axios.defaults`
copy is no longer written (`setupAxiosAuth` is a documented no-op); `main.js`
installs a global axios REQUEST interceptor that rebuilds the header per
request, so ~368 bare-`axios` call sites get the current credential without
being rewritten and a new one cannot forget to opt in. This is the AC's second
half ("or is provably never read in preference to the store") and it is the
stronger of the two.

An explicit header still wins, and exactly one caller needs that: the logout
revoke. #2258 clears local state BEFORE the revoke, so with the defaults copy
gone the revoke would have gone out unauthenticated and #187 would have silently
stopped revoking anything. The token is captured before the clear and passed
after it.

**One verdict.** `sessionLostVerdict()` → `ignore | stale | logout`, pure so a
node-env spec can reach it. `stale` — the failed token is not the stored one —
is the fix for the report: adopt the current session instead of destroying it.
The Workspace veto closes AC #5: a client whose browser holds a DEAD operator
JWT is no longer thrown onto the operator login by `initializeAuth`'s
`fetchUserProfile`. It stays scoped by path as well as by portal token, so an
expired operator JWT still bounces off an operator surface.

**One handler.** `setPlatformUnauthorizedHandler` / `notifyPlatformUnauthorized`.
`main.js` registers the reaction; `api.js`, the global interceptor and
`portalHttp` report to it. `api.js` no longer hard-reloads, no longer leaves
`auth0_user` behind, and carries no private predicate.

**Cross-tab sync.** A `storage` listener adopts a sibling's login and drops the
mirror on a sibling's logout — without a second server revoke and without
writing to storage, since N background tabs reacting to one event would each
clear it again. Neither branch navigates: a background tab pushing /login is the
noise this issue reports.

`workspaceSession.spec.js`'s predicate block asserted its own hand-copied
`shouldBounce` helper — which is why it stayed green while the three real
predicates drifted, and would have stayed green through this change too. It now
asserts the real function.

Verified: 131 files / 2937 tests pass. The three load-bearing guards were
mutation-checked — removing the `stale` arm reds 2, letting the interceptor
overwrite explicit headers reds 1, restoring `api.js`'s own logout reds 2.

Related to #2791

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf

* docs(http): boundedHttp explained itself with the mechanism this PR deletes (#2791)

Review finding on my own diff. The docblock's "why not `axios.create()`" argued
from `stores/auth.js` mutating `axios.defaults.headers.common.Authorization` at
login and deleting it at logout — the exact copy #2791 removes.

The conclusion survives the mechanism (the global is still the only thing
carrying a live credential, now because the request interceptor resolves it per
request and `create()` gives an instance its own chain the global never
reaches), which is precisely why the comment would have gone on reading as true.
A comment that describes a mechanism the code no longer has is the class this
repo's learnings ledger already records; the old reason is kept in parentheses
because it explains why the answer did not change.

Related to #2791

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf

* fix(session): the boot-time axios default is gone, the reactions execute under test, and the veto reads the per-tab store (#2791)

Merge-train review of #2811, all six ❌/W items:

C1  App.vue wrote `axios.defaults.headers.common['Authorization']` on every
    boot with a token. Axios merges that default into the request BEFORE the
    interceptor chain runs, so it arrived at the per-request rebuild looking
    explicit and won over storage for the life of the tab — the whole
    mechanism was inert, and 119 bare-axios sites kept the boot-time token
    after a sibling re-login. The write is removed (the store seed + WS
    connect stay); `adoptStoredSession` / `applySessionEndedElsewhere` delete
    any such copy as a belt, so a tab can never ride a credential storage no
    longer holds (W2).
C2  The "nobody writes the default" guard walked auth.js only. It now walks
    src/frontend/src/** and asserts zero writers; App.vue would have failed it.
C3  architecture/workspace.md + workspace-session-signout.md said the copy is
    written nowhere while App.vue wrote it. Both now describe what is true
    and why the guard is tree-wide.
C4  handlePlatformUnauthorized ended `router.push('/login')` without `return`,
    so notifyPlatformUnauthorized's absorber never engaged and a redundant
    navigation escaped as an unhandled rejection. It returns the navigation.
CI  The reaction, the storage listener and the request rebuild were inline in
    main.js and pinned by regex — restoring the reported bug on the `stale`
    branch and inverting the listener both stayed green. They are now
    `reactToPlatformUnauthorized`, `reactToStorageEvent` and
    `applyRequestCredential` in utils/platformSession.js, taking their
    collaborators as arguments; the spec EXECUTES them with fakes, and a
    five-mutation battery (stale→logout, inverted listener, header never
    applied, navigation dropped, App.vue writer back) is red on every one.
    main.js is wiring only, and the source guards assert exactly that.
W1  The Workspace veto read `portalTokenPresent` from shared localStorage
    while portalHttp gates on the per-tab store; a client signing in in
    another tab stranded an operator's Workspace tab on an expired JWT with
    `ignore`. It reads `useClientPortalStore().portalToken` now.
W3  The 22 direct `localStorage.getItem('token')` reads outside the reader
    (WS/EventSource included) go through `readStoredToken()`; a tree-wide
    guard keeps "one reader" true.
W4  The stale `installCrossTabSync()` reference is gone with the rewrite.
W6  The storage listener also hears `auth0_user`, and adopting an identical
    token refreshes the user from storage, so a sibling login's profile
    landing a tick after its token is not missed.

3066 frontend unit tests green; vite build green.

Fixes #2791

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf

* merge-train: the Workspace veto honours the tab's own client claim after expiry (#2811) — mechanical, per the merge-train note on the PR

`handlePlatformUnauthorized` read `portalToken` alone. `endSession({expired})`
nulls that token and sets `platformFallbackSuppressed` in the same breath, so
for a client tab holding a dead operator JWT the 5 s ticket retry's next 401
fell through to `logout` and threw the client from the OTP form onto the
operator login — the #2258/#2261 bounce, reopened for exactly that instant.
`portalTokenPresent` now folds in the suppression flag (per-tab sessionStorage,
so W1 cannot return). The wiring pin follows, and the tautological
"racing away" spec case (identical inputs both sides) is replaced by one that
can fail: client tab → ignore, live or expired; operator → logout.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
vybe added a commit that referenced this pull request Sep 17, 2026
* fix(workspace): the ask badge said 2 and gave you no way to find them (#2424)

The sidebar advertised "2 asks are waiting on your answer" and then stranded
you: the agent that raised them carried no badge, its tooltip did not mention
them, and it could be collapsed out of the roster entirely. The only way to
locate a blocked agent was to open agents one at a time.

Observed on a 12-agent roster with two asks on ws-sage (11th of 12), so on a
fresh load the one row that mattered was behind the "show more" toggle.

Three failures, fixed together because separately each is a half-measure — a
badge with no destination, or a destination nobody can see.

1. The unit. `askCount` is `openAsks.length`, and the tooltip said "agents":
   two asks on ONE agent rendered as "2 agents are waiting on your answer". The
   number was right, the noun was wrong, and they only diverge when a single
   agent raises more than one ask — which is why it went unnoticed. Resolved
   toward ASKS rather than agents, because the row badges added here now answer
   "which agent", leaving the header to answer "how many decisions".

2. The row. `PortalSidebar.vue:139` renders a per-agent badge from
   `unreadByAgent` — unread REPLIES. Keeping asks out of that count is
   deliberate and documented at line 9 ("one is waiting on you to decide, the
   other on you to read"), and is preserved: the ask gets the *own badge* that
   comment promised, in `status-urgent` — the token the operator NavBar's
   pending-operator-queue badge already uses, so the two surfaces agree — and
   visually distinct from the indigo unread pill beside it. `agentRowTitle` had
   the same hole, so this is an accessibility fix too: a blocked agent's
   accessible name was the bare "Open ws-sage".

3. The collapse. #2159 capped the roster at five for a good reason (a long
   fleet pushed chats below the fold), but the slice is plain roster order with
   no ask weighting. Ask-bearing agents are now never hidden — appended, NOT
   floated to the top, because re-sorting on a transient count moves rows under
   the cursor between refreshes, the same reason the roster is not re-sorted by
   availability.

Not a regression: every piece shipped in its intended form; the gap was between
them.

Everything decidable moved into `portalUtils` (`asksByAgent`, `askBadgeTitle`,
`agentRowTitle`, `visibleAgentRows`, `AGENT_COLLAPSE_LIMIT`) because vitest runs
`environment: 'node'` with no mount harness — a rule inside the SFC is one no
test can reach, which is how all three of these shipped. Mutation-checked:
reverting the noun, dropping asks from the title, and restoring the plain slice
each turn the suite red.

`bg-amber-500` -> `bg-status-urgent-500` is required, not drive-by: new code must
be at zero raw palette classes, so the new badge needed a token, and the header
had to match it or the two ask indicators would differ. Amber maps to
`state-autonomous` (an operating mode), which is the wrong claim. PortalSidebar
is now at zero non-gray raw classes.

Two pre-existing guards asserted the moved expressions as source strings and are
rewritten to assert the properties behaviourally — strictly stronger, since they
now fail on a broken bound or a dropped chip title, not only on a reworded one:
- portalRosterRow #2159 "shows a fixed number by default"
- portalAvailabilityChip #2196 "row title carries the state"

Verification: 1518/1518 frontend unit tests, raw-color ratchet exit 0,
production build clean.

Closes #2424

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(workspace): the sync portal turn never carried its session, so report-back could not fire (#2426)

ent#457 gave the Workspace a report-back: an agent that delegates during a chat
turn gets the completion posted into that thread. It could not fire on the
SYNCHRONOUS path, because the parent execution never received the session
binding the report needs. `report_completion` gates on
`if not source_channel_chat_id`, and there the field was NULL.

Measured on a dev instance — 5 of 8 portal rows NULL, split exactly by path:

    07:55 -> 09:04   chat=7d27744d...   browser, streaming path
    09:06 -> 09:09   chat=NULL          POST .../chat, synchronous path

TWO CORRECT CHANGES THAT COLLIDE. ent#457 passes the binding down, and
`execute_task` persists it — but only inside `if not execution_id:`. ent#365's
`_precreate_sync_execution` has already created the row and handed the id over,
so that branch never runs, and the pre-create stamped only `source_channel`.
Its own docstring named the invariant it broke: "Mirrors `start_portal_turn`'s
creation exactly ... so the two paths produce indistinguishable rows and a
report published from either can be joined back to its chat."

The sibling comment in `start_portal_turn` says "both creation sites or the
stamp is a coin flip depending on which path made the row" — ent#457 covered the
two sites that existed when it was written; ent#365 had added a third.

Fix: stamp `source_channel_chat_id` + `source_channel_client` in the pre-create.
`session_id` is a REQUIRED parameter, not an optional one — the value is in
scope at the only call site, and a default would let a future caller silently
reintroduce the inert row. Rejected: teaching `execute_task` to UPDATE an
adopted row, which widens a hot path used by every trigger to repair one
caller's omission.

ALSO REPAIRS TWO GUARDS THAT WERE RED ON `dev`. `backend-unit-test` is failing
on dev right now; both failures are in this feature area and both are guards
that had gone inert, so they are fixed here rather than left for the next PR to
trip over. Frontend-only PRs pass because the `changes` job path-filters the
backend suite away, which is why this went unnoticed.

  * `test_both_portal_row_creation_sites_name_the_chat` asserted a literal
    census of `== 2` sites. It went red the moment the third site appeared —
    the guard WORKING — and the bug it names shipped anyway. Now asserts the
    rule instead of the count: every site that stamps the surface must also
    stamp the destination. Census-proof.

  * `test_portal_turn_kwargs_bind_against_execute_task` parsed `portal_chat`
    for a literal `run_resumable_turn(...)` call. That call had moved into
    `_run_sync_turn_and_clear_marker`, where it is `run_resumable_turn(**kwargs)`
    — a splat, which names nothing — so the walk found no keywords and the
    guard asserted itself dead. Now reads the keywords where they are actually
    named (the wrapper's call site), scanning both entry names and subtracting
    the wrapper's own consumed parameters.

Neither rewrite loses coverage; both now fail for the reason their docstring
gives rather than because a number or a call site moved.

WHY THE BUG SURVIVED ITS TESTS. ent#457's mock the engine and assert the kwargs
are passed (they are). ent#365's assert no orphan `running` row (still true).
Nothing asserted the PERSISTED ROW, which is the only place the two meet — the
same lesson `test_ent457_portal_turn_kwargs.py` states about itself. The new
suite asserts at that layer, and adds a derived parity check so a fourth
channel field added to one writer and forgotten in another fails here instead
of shipping as another silently-inert report path.

Verification: 402 passed on the portal/ent457/ent365 selection (was 2 failed
before this branch). Mutation-checked: removing the stamp turns 4 red; feeding
`execute_task` an unknown kwarg turns the repaired binding guard red.

Closes #2426

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(subscriptions): auto-switch ranks alternatives by cached headroom, never by load alone (#2409) (#2422)

## Summary
- `select_best_alternative_subscription` returned the **first** survivor of the 2h failure filter in `agent_count ASC` order and read no headroom — SUB-003 could move an agent onto a subscription at 99% of its weekly window, and an *unused dead-token* subscription (no agents ⇒ no failure rows) sorted **first**.
- Now: **filter in the db, rank in the service, never a probe.** The db lists survivors (kind-blind 2h filter unchanged and first, #444/#2352, `agent_count ASC, name ASC`); the service ranks them over the cached provider snapshot (one `MGET`) furthest-from-the-nearest-wall first (the fuller of the 5h/7d windows — the #792 retry lands on the destination immediately), in 10-point bands so load still spreads a storm; a **fresh** provider refusal is dropped; anything unusable sorts in today's order; any failure of the ranking half falls back to today's pick **with a warning**.
- `classify_headroom` (ent#434) and the ranker share one usability gate (`headroom_reading`) — verdicts byte-identical, pinned by a differential test against a frozen copy. New-agent auto-assign (#74) rides the same ranker. The switch now records **why** (`destination_headroom` + one notification clause).
- Approved deviations from the literal AC, recorded on the issue: nearest-wall key instead of 7d-only; fresh refusals filtered instead of ranked last.

## Changes
- `src/backend/services/subscription_headroom_service.py` — gate, MGET reader, threshold-free ranker, `MAX_READING_AGE_SECONDS` (owned here now); `classify_headroom` becomes policy over the gate
- `src/backend/services/subscription_auto_switch.py` — service-layer selector (`asyncio.to_thread` under the agent lock), `destination_headroom` on activity / notification / result
- `src/backend/services/subscription_service.py` — `select_subscription_for_new_agent`
- `src/backend/db/subscriptions.py` + `database.py` — `list_viable_alternative_subscriptions` / `list_assignable_subscriptions` (filter only); first-match selectors retired
- `src/backend/services/agent_service/crud.py` — call site; `subscription_headroom_alerts.py` — constant re-export + docstring
- Tests: new `tests/unit/test_2409_headroom_ranked_switch.py`; pingpong / 2352 / concurrency / 1484 / 1759 adapted to the list form (assertions kept)
- Docs: `architecture.md`, `subscription-auto-switch.md` (+ management, usage-tracking), requirements §20.4, `learnings.md` (2 entries), CSO diff report

## Test Plan
- [x] New suite: `pytest tests/unit/test_2409_headroom_ranked_switch.py` — 83 passed; **80/81 fail on the unmodified source**
- [x] Full `tests/unit`: 12,718 passed / 30 skipped / 1 pre-existing failure (`test_1920`, private submodule, untouched)
- [x] API integration (`test_subscription_auto_switch`, `test_subscriptions`, `test_subscription_usage`): 36 passed
- [x] Live: a real switch chose the 18%/9% subscription over the 0-agent 88%/60% one; every-survivor-refused → no switch + WARNING; no snapshot → today's order
- [x] `/review` clean (informational findings fixed in-review); `/cso --diff` no findings
- Follow-ups filed while testing: #2419 (parser overage), #2420 (destructive integration suite), #2421 (subscription audit gap)

Fixes #2409

🤖 Generated with [Claude Code](https://claude.com/claude-code)

* feat(workspace): an answer given in the Workspace resumes the agent (ent#430)

Slice 5 of ent#364, and the gate: until now the client route recorded an answer
and returned. The operator route called `spawn_resume_dispatch`; this one did
not. So an ask addressed to a Workspace client — the entire point of
ent#364/#428/#429 — was recorded, reached the agent's queue file in about three
seconds, and re-triggered nothing.

Measured on a live instance before this change: answered from the Workspace,
`operator-queue.json` flipped to `responded` with the answer in under 3s, and no
execution followed.

Unblocked because ent#329 is in dev.

WHAT THIS ADDS: one call. ent#430's body rules out the alternative — "a second
dispatch surface for the same event is how the cost, trigger-label and
loop-prevention questions get answered twice, differently" — so the per-agent
opt-in, the idempotency key, the audit row and the failure handling all stay
inside `maybe_dispatch_resume`. AC #2 and AC #3 are satisfied by REUSE rather
than by re-implementation, and the tests assert the CALL for that reason.

Four properties, each load-bearing:

* Hung off the CAS WIN only, like the operator route. The 409 above already
  returned for a lost race, so reaching the dispatch means this answer is the
  one that landed — two people answering at once produce one resume.
* `updated`, never `item`. The pre-answer read still says `pending`; a resume
  handed that row acts on an ask that does not yet carry its answer. Looks
  identical in a green test, which is why there is one for it.
* The spawn is wrapped. It is fire-and-forget, but a raise ON THE CALLING LINE
  would still propagate, and a 500 after the CAS landed would tell the client
  their answer failed while it is committed and already on its way to the agent.
  The answer is the thing that must not be lost.
* #2376's choice validator runs first, so an answer that was never offered
  cannot spend.

AC #5 — `resume_requested` on the answer response, read from the SAME accessor
the dispatch gates on, so the two cannot disagree about what is about to happen.
It reports INTENT, not success: the dispatch is backgrounded, so at that moment
the only honest claim is whether it will be attempted. Fails CLOSED — an
unreadable flag claims nothing, because over-claiming is exactly the failure
AC #5 names ("the ask does not read as resolved while nothing happened").

RESIDUAL, stated rather than implied: a dispatch that fails AFTER this point
surfaces as a FAILED execution row plus an `operator_resume_dispatch` audit
entry (ent#329) — operator-visible, and a client cannot see either. The client
half of AC #5 is satisfied negatively for now: the ask surface says nothing
about work starting, so it cannot mis-claim. `resume_requested` is the field a
surface needs to say something true; consuming it is an ent#429 UI change and is
deliberately not in this PR.

The per-agent flag DEFAULT IS UNCHANGED (`operator_resume_enabled`, OFF,
owner-only). "Turn the flag on" is an operator action per agent, not a code
default: flipping it would hand every shared agent's client a spend button,
which is the one thing AC #3 rules out.

Verification: 145 passed across the asks/ent#329/ent#364/#428/#429/#2376
selection. Mutation-checked — removing the dispatch (4 red), passing the
pre-answer row (1 red), and making the opt-in read fail open (1 red).

Closes ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(workspace): New chat means a new chat (ent#451 — the fresh-thread slice)

Reported: pressing New chat in the Workspace drops you back into the existing
conversation with that agent. Decided at the 2026-08-21 weekly.

ONE VALUE CARRYING TWO MEANINGS. An absent `session_id` meant both "I don't know
which thread" and "I want a fresh one", and the platform resolved it as the
first, in both readers:

    _resolve_session_id(..., None)  -> resume the client's latest
    get_history(..., None)          -> return the most-recent thread

Both readings are RIGHT for the case they were written for — a deep link, a
refresh, an API caller that never held a session id — so neither could be
inverted. The intent had to become sayable: `new_thread` on the request,
`newChat` on the component, checked before the resume.

The frontend tell was an asymmetry: New chat with the agent you were ALREADY on
started fresh, while New chat with a different agent resumed. The watcher read a
changed agent as "load that agent's history" and called `fetchHistory(name,
null)`, discarding the `pendingSession = null` that `newChatWithAgent` had just
set to mean the opposite.

MOST OF ent#451 TURNED OUT TO BE BUILT. Recorded because the issue is
complexity-high and this PR is not:

* the data model already allows many sessions per (agent, client) — no UNIQUE
  constraint, a `title` column, an index on
  `(agent_name, client_email, last_message_at)`, and auto-titling. AC #4's
  "migrates cleanly" is nothing to migrate.
* AC #2's list is the existing sidebar: titles, recency, starred lifted out,
  search, per-agent avatars.
* AC #3's landing rule is already decided and documented in
  `ensure_thread_for_ask` — reuse the latest thread so asks do not accumulate
  beside the conversation. UNCHANGED here, and pinned by a test so this cannot
  move it silently. It matters MORE once several chats exist, not less.

So what was missing is AC #1, and it is two bits rather than a data model.

Four properties:

* An explicit `session_id` WINS over the flag. A caller sending both contradicts
  itself; the id is a fact, the flag an intent, and abandoning a named thread
  would strand a turn meant for a conversation the caller could see.
* The ownership check runs first either way — the flag is never a route past it.
* BOTH turn entry points carry it. The Workspace uses the streaming path and
  falls back to the synchronous one, so a flag honoured by only one brings the
  bug back exactly when streaming fails.
* The intent is spent on adoption. The send guard already ANDs on "no session
  yet", so a second turn was never going to open a third thread; clearing it in
  `onSessionAdopted` keeps the two bits from disagreeing after a navigation.

Test doubles updated, not worked around: seven `_resolve_session_id` lambdas and
four `_fake_chat` stubs did not accept the new keyword. They take `**kw` now — a
stub that must be edited for every new parameter is a second signature — and one
hand-rolled `_Body` model double gained the field. All are stale stubs rather
than behaviour changes.

Verification: 392 passed across the portal/ent#286/#287/#358/#429/#430/#451
selection; 1497 frontend unit tests. Mutation-checked: making the flag inert, and
letting it override an explicit session id, each turn the suite red. The full
backend suite exceeds a local foreground run and is left to CI.

Pre-existing and NOT from this branch: `test_ent457_portal_turn_kwargs` and
`test_both_portal_row_creation_sites_name_the_chat` fail on `dev` today; both are
fixed in #2427.

Related to ent#451

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the ?new=1 deep link, the missing frontend test, and three latent desyncs (ent#451)

Blocker 1 was real and I had not seen it. `resolveAgentQuery` passed `forceNew`
to `resolveAgentLanding` and set `pendingSession = null`, but never raised
`startingNewChat` — so `/workspace?agent=X&new=1` rendered an empty conversation
and then sent `new_thread: false`, resuming the thread the user asked to leave.
The reported bug, intact on the documented `?new=1` contract, in the PR that
exists to fix it.

The cause is the one this PR is about, one level up: `route.query.new` was read
in two places for two different decisions — WHICH THREAD to land on and WHAT THE
FIRST SEND ASKS FOR — and only the first honoured it. Now read ONCE into a local
that feeds both, so they cannot drift again. AND-ed with the landing result, so
a `?new=1` that still resolved a thread never claims a fresh start.

Blocker 2: a frontend test, which the change genuinely had none of — the
`1497 passed` in the body was the pre-existing suite, as the review says.
`workspaceNewChat.spec.js` (9 tests) covers the deep link, the watcher branch
ORDER, the first-paint guard, both send conjunctions, and the settle-everywhere
rule, using the two established patterns (pure function + source assertion in
the `portalLeaveSpecificRoute.spec.js` shape) since vitest runs
`environment: 'node'` with no mount harness. Mutation-checked, and M1 is the
reviewer's own blocker: reverting it turns the suite red.

Blocker 3: `test_history_without_a_session_is_unchanged` cited "the spec in
tests/unit/... frontend suite" — a dangling reference asserting coverage that
did not exist. It now names the real file.

Comments addressed:

* Three more sites nulled `pendingSession` without settling the intent — the
  deep-link watcher (the commonest way in), `openRoom`, `openAgentPage`, plus
  the unreachable-agent branch. Latent because both consumers AND on "no session
  yet", but a flag that is only correct because of a second variable is one
  refactor from being wrong, and the declaration claims it is cleared the moment
  a real thread exists. Now true.
* `test_both_turn_entry_points_forward_it` was `getsource` + a substring, so a
  comment or a misspelled kwarg satisfied it. It now BINDS the keyword against
  each service signature and asserts the routes forward `body.new_thread`
  through a comment-stripped source — verified by mutation.
* `workspace-absorbs-session.md` updated at both seams the change touches
  (`resolveAgentLanding`'s landing rule and `_resolve_session_id`'s three
  states), and `architecture.md`'s Workspace section documents the new public
  `new_thread` field on the ent#83 headless surface.
* Gating stated rather than inferred: "OSS-core by decision (ent#451)", matching
  the ent#326/#384/#392 convention.

ONE CORRECTION, offered with evidence rather than silently applied. The review
says "`test_ent457_portal_turn_kwargs.py` doesn't exist on `dev`, #2427
introduces it". It does exist on `dev` — added by d6a4bc10 (ent#457) — and #2427
modifies it. `git cat-file -e origin/dev:tests/unit/test_ent457_portal_turn_kwargs.py`
succeeds, and `backend-unit-test` is failing on `dev` independently of any PR.
So the body's "fails on dev today" stands. Everything else in the review is
accepted as written.

Verification: frontend 1497 -> 1506 (+9). Backend 392 passed on the portal
selection, the same 2 pre-existing dev failures unchanged.

Related to ent#451

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(resume): the respond→resume dispatch never ran — bad import, masked by its own stub (ent#329)

Found by testing this PR's feature against a live local instance. ent#430 wires
a Workspace answer to `spawn_resume_dispatch`, so this PR is dead on arrival
without it — the client path would have hit the same wall the operator path has
been hitting since ent#329 merged.

THE BUG. `operator_resume_service.maybe_dispatch_resume` did:

    from services.task_execution_service import task_execution_service

That name has never existed on that module; it exports
`get_task_execution_service()`. The import sits on the FIRST line of the
function, above the try, so every dispatch raised ImportError before it even
read the opt-in.

WHY NOBODY NOTICED, twice over:

* the call is fire-and-forget, so the traceback surfaces only as asyncio's
  "Task exception was never retrieved" — nothing fails, nothing 500s, the
  answer is recorded and the config audit row is written. It looks like it
  worked.
* the ent#329 unit test stubbed `services.task_execution_service` with
  `SimpleNamespace(task_execution_service=recorder)` — MANUFACTURING the very
  symbol whose absence was the bug. 21 tests green, feature dead.

MEASURED on a live instance, opt-in ON:

  before: answer 200, audit row written, executions 0->0, log carries
          "cannot import name 'task_execution_service'"
  after : answer 200, executions 0->1, triggered_by=operator_response,
          audit `operator_resume_dispatch` with the execution id, 0 ImportErrors

(The dispatched run then failed on a missing AGENT_AUTH_SECRET — a limitation of
the test box, and correctly recorded as an honest FAILED row, which is ent#329's
"never silent" requirement doing its job.)

THE GUARD is the durable part, because the stub is the real lesson: a stub that
invents an API the real module lacks converts a production crash into a green
suite. `test_the_names_this_service_imports_actually_exist_on_the_real_modules`
parses the REAL module source with `ast` — never the stubbed `sys.modules`
entry, which is what made this invisible — and asserts every
`from services.X import Y` resolves. Mutation-checked: reverting the import
turns 11 tests red.

Related to ent#430, ent#329

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the dispatch could not run, the race loser spent, the flag over-claimed (ent#430)

All three blockers from the review, each verified rather than argued.

1. THE FEATURE WAS INERT. `client_portal/asks/router.py` declares `answer_ask`
as a plain `def`, so FastAPI runs it through `run_in_threadpool` — a worker
thread with no event loop — and `asyncio.create_task` raises
`RuntimeError: no running event loop` there. The `except` swallowed it, so every
client answer recorded the answer and dispatched nothing: byte-for-byte the
behaviour this PR exists to remove.

Fixed in `spawn_resume_dispatch` rather than by flipping the route to
`async def`, for the two reasons the review names: the route does blocking DB
I/O, so `async def` alone would move it onto the loop; and ent#430's stated
shape is ONE dispatch site, which moving the spawn back out to the caller would
undo. It now detects the absence of a loop and hops back via
`anyio.from_thread.run_sync` — Starlette's threadpool is anyio's, so the portal
is always there on this path. Any future sync caller inherits the fix.

A thread anyio does not own reaches neither branch. That is not a production
shape, but it must not become the silent no-op this change removes, so it raises
with the cause named instead.

2. THE RACE LOSER SPENT MONEY. `respond_to_operator_queue_item` returns None
only when the row is GONE; when the row exists and has left `pending` — the race
that actually happens — it returns a TRUTHY dict carrying `_status_conflict`,
having written nothing. `if not updated` fell straight through it. The loser
then dispatched a paid execution for an answer not in the database, and because
the idempotency key hashes the response text, the loser's differing text yields
a different digest: one queue item, two paid dispatches. `routers/operator_queue.py`
already pops that flag before its own spawn; this is that rule, not a new one.
Popped, not read, so the sentinel cannot serialize to the client.

3. `resume_requested` OVER-CLAIMED. It was computed after the swallowed spawn
from the opt-in flag alone, so a spawn that raised still answered `true` — the
exact failure AC #5 names, and given (1) that was EVERY production answer on an
opted-in agent. It now reports what was actually scheduled.

TESTS — the reason all three survived 24 green checks is that every existing test
replaced `spawn_resume_dispatch` with a synchronous lambda, stubbing out the one
call whose runtime context was the defect. `test_ent430_dispatch_actually_runs.py`
drives the REAL spawn from a REAL anyio worker thread (the production context,
not an approximation) and asserts the premise before the behaviour. The lost-race
test uses the truthy `_status_conflict` shape that actually occurs, not the
`None` shape that does not. Mutation-checked: reverting fix 1 turns 1 red, fix 2
turns 3 red, fix 3 turns 2 red.

Writing those tests also caught a stubbing bug of my own, worth recording because
it is the trap that hid the original: patching only `sys.modules` leaves
`from services import operator_resume_service` resolving the PACKAGE ATTRIBUTE,
so the real function ran anyway. Both paths are patched now.

Related to ent#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): the answered ask said pending, and the docs described one caller (ent#430)

Non-blocking findings from review pass 2. The three blockers landed in d11956a8.

STATUS. `_project` mapped every row to pending/expired, so the response to a
just-recorded answer read `status: "pending"` beside `resume_requested: true` —
one row reporting both that nobody has answered it and that answering it started
work. Harmless while the second field did not exist; contradictory once it did.
`_status_of` adds `answered` (`responded`/`acknowledged`), reachable only from
the answer response since the listing carries neither. Answered is checked
BEFORE expiry — an answer that landed is a fact, and an `expires_at` that has
since passed does not un-answer it; the obvious refactor is to test expiry first,
which would make a slow client's own answer vanish, so the ordering is pinned.

The existing test asserted `out.status in ("pending", "expired")` with the
comment 'the point is it returned at all' — it was papering over exactly this.
It now asserts `answered` and, on the spawn-failure path it covers, that
`resume_requested` is False.

THE TWO READS. `_resume_requested`'s docstring claimed it read 'the SAME
accessor … so the two cannot disagree'. True of the accessor, false of the
instant: it is a second read a task hop earlier, and an owner disabling the
opt-in in between gets `true` and no resume. Collapsing them is not the fix —
they answer different questions (one must produce a value for THIS response, the
other is the authority at the moment it would spend), so the window is stated,
with AC #5's own remedy named, rather than described away.

DOCS. architecture.md's ent#329 section described a single caller and stated the
CAS-win property the second caller broke. It now carries the second caller, the
truthy-`_status_conflict` shape that defeated `if not updated`, the
sync-endpoint/no-loop defect and its `anyio.from_thread.run_sync` fix, and what
`resume_requested` actually reports.

Related to abilityai/trinity-enterprise#430

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(enterprise): bump the submodule pointer to main (a419812 -> 90f2f2c) (#2440)

dev's pointer was OLDER than main's — an inversion, not just staleness. The
next dev -> main release merge would have carried it backwards and undone
ent#443's enterprise-side removal:

  OSS main -> 2a5def3   (ent#443: shared_sessions removed from enterprise)
  OSS dev  -> a419812   (4 behind enterprise main, 2026-08-19)
  ENT main -> 90f2f2c

90f2f2c is a fast-forward from BOTH (a419812...main = ahead 0 / behind 4;
2a5def3...main = ahead 3 / behind 0), so nothing is being rewound.

WHAT THE FOUR COMMITS ARE

  2a5def3  refactor(rooms): remove shared_sessions — it now lives in OSS core (ent#443)
  ff0a4f1  fix(security): guard the enterprise system_settings sinks against
           cleartext credentials (ent#435) — the private twin of the OSS sink
           guard architecture.md already records as "the private submodule owns
           its twin"
  6d82a3f  docs(workspace): agent-initiated asks — design of record
  90f2f2c  feat(credential-vault): governed system credential vault module (ent#279)

WHY IT MATTERS RATHER THAN BEING HOUSEKEEPING. ent#443 moved rooms into OSS
core, and dev has that. With the stale pin an entitled dev install mounts the
OSS rooms routers AND the enterprise shared_sessions module, and relies on
main.py's include-order (OSS before register_enterprise) to decide which one
serves. architecture.md documents that ordering as the transition safety net —
this bump is the follow-through that ends the transition.

VERIFIED BY BOOTING BOTH POINTERS against dev, same box, same DB shape:

  a419812 (today):  17 modules | shared_sessions registered: True  | 6 room paths | 0 errors
  90f2f2c (this):   17 modules | shared_sessions registered: False | 6 room paths | 0 errors

Both boot clean and log "Trinity Enterprise modules registered" — the line
deploy-dev greps. Module count is unchanged because shared_sessions leaves as
credential_vault arrives. No duplicate room paths in either, confirming the
ordering net held; after the bump there is nothing to net.

Gitlink only — no OSS source changes, so public CI (which never checks the
submodule out) is unaffected.

Related to ent#443, ent#435, ent#279

* chore(metrics): code-health dashboard 2026-08-31 @ 135248e9 (#2438)

Co-authored-by: Trinity Agent (trinity) <trinity-agent@ability.ai>

* chore(deps): bump node (#2400)

Bumps the docker-base-images group with 1 update in the /docker/frontend directory: node.


Updates `node` from 24-alpine to 26-alpine

---
updated-dependencies:
- dependency-name: node
  dependency-version: 26-alpine
  dependency-type: direct:production
  dependency-group: docker-base-images
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix(files): shared links were unopenable on mobile — Range, disposition, MIME, CORP (trinity-enterprise#461) (#2439)

* fix(files): shared links were unopenable on mobile — Range, disposition, MIME, CORP (trinity-enterprise#461)

The bytes were never wrong. Verified from the Cloudflare edge, the object
returned HTTP 200 with correct content-length and correct WAV bytes for every
user-agent tried, and the signature check worked. The RESPONSE SHAPE was wrong
in four ways at once, and each one alone is enough to break playback in an iOS
in-app browser:

* no Range support — `Range: bytes=0-1023` returned 200 with the whole 2 MB
  body and no `accept-ranges`. iOS Safari and Telegram's player require a 206
  to start audio at all, so this alone made the file unplayable.
* `content-disposition: attachment` — a forced 2 MB download inside Telegram's
  iOS browser is a blank screen.
* `audio/x-wav` under `nosniff` — unregistered type, so a strict player
  declines it and the browser is forbidden from guessing better.
* `cross-origin-resource-policy: same-origin` on a link whose entire purpose is
  to be opened from another platform.

Plus `cache-control: no-store`, which forbids the in-app browser from buffering
media it will not play without buffering.

THE INLINE CHANGE IS A NARROWING, NOT A REVERSAL. The old code forced
`attachment` on everything with the note 'defense against XSS via
agent-uploaded HTML', and that reasoning is still correct — this route serves
agent-authored bytes from the same origin as public chat. So inline is an
ALLOWLIST (`_INLINE_SAFE_TYPES`: audio, video, image, PDF) and `text/html`,
`application/xhtml+xml` and `image/svg+xml` stay attachments. SVG is called out
because it is the one a reviewer waves through: it is an image by name and a
script host in fact. The type is python-magic-detected from the file's own bytes
at share time, never agent-supplied, and its unavailable-fallback
(`application/octet-stream`) sits outside the allowlist, so the failure
direction is `attachment`. `nosniff` is kept and matters more now, not less.

TWO THINGS THE ISSUE DID NOT ASK FOR, both found while implementing:

* `Content-Length` came from the DB's `size_bytes`, written at share time. Any
  drift from the file on disk is unrecoverable for the client — too small
  truncates, too large hangs — and Range math against a wrong total produces a
  `Content-Range` that contradicts the body. It now comes from
  `os.path.getsize`, with a WARNING on divergence.
* a media player fetches one file as MANY ranged requests. Counting each as a
  download would turn one play into dozens and write an audit row per chunk, so
  the counter and the audit fire only on the transfer START (a plain GET, or a
  range beginning at byte 0).

VERIFIED end-to-end against the real route, not just the parsers:

  full GET      : 200 | type audio/wav | disp inline | ranges bytes
                | corp cross-origin | cc private, max-age=3600
  range 0-1023  : 206 | body 1024 | bytes 0-1023/2048000 | bytes ok
  suffix -500   : 206 | bytes 2047500-2047999/2048000 | bytes ok
  unsatisfiable : 416 | bytes */2048000
  HEAD          : 200 | accept-ranges bytes | content-length 2048000
  no sig / bad sig / unknown id / expired : 401 / 401 / 404 / 410
  html file / svg file : attachment

That covers the issue's Definition of Done line by line, including that the
signature check still rejects unsigned and expired requests.

46 new unit tests, weighted to the allowlist and to the range parser's
silent-corruption case (`bytes=-500` is the LAST 500 bytes; reading it as
start=0 serves the wrong bytes under a 206, which no client can detect).

Related to trinity-enterprise#461

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(files): the CORP header was inert — the security middleware clobbered it (trinity-enterprise#461)

Found by testing the PR against a real local instance rather than a TestClient.

`main.add_security_headers` runs after EVERY route and set
`Cross-Origin-Resource-Policy` with a plain assignment. So the `cross-origin`
policy the file-download route sets — one of the four fixes in this PR, and the
one that decides whether Telegram, Slack or WhatsApp can embed or preview the
link at all — was silently overwritten back to `same-origin` on its way out.

The fix shipped INERT and every test passed, because a bare `FastAPI()` +
router harness has no middleware. Measured on the running server:

  before: cross-origin-resource-policy: same-origin
  after : cross-origin-resource-policy: cross-origin   (file route)
          cross-origin-resource-policy: same-origin    (/health, unchanged)

`setdefault` rather than a route allowlist: absence still resolves to the strict
default, so every other route keeps today's behaviour and a new route has to opt
out deliberately rather than inherit an exception.

Pinned by a source assertion — asserting it end-to-end needs a live stack, and
what must not regress is the `setdefault`; an edit back to `=` would re-break it
invisibly.

Related to trinity-enterprise#461

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(main): lifespan is an orchestrator, not a 580-line procedure (#1028) (#2437)

* refactor(main): lifespan is an orchestrator, not a 580-line procedure (#1028)

`main.py::lifespan` was 580 lines at cyclomatic complexity 109 — the longest
function in the backend and the first item the 2026-06-02 refactor audit named.
It is now 25 lines at CC 1: a flat list of `await _phase()` calls over twelve
startup helpers and four shutdown helpers.

WHAT MOVED, AND THE PROOF THAT NOTHING ELSE DID. Every body is verbatim. That
is asserted mechanically rather than claimed: extracting all non-blank lines
from the sixteen helpers in call order and diffing against the original
`lifespan` body gives 518 = 518, identical. The only relocated line is `yield`.
No behaviour change, no logic touched, no try/except reshaped — each phase keeps
its own guard, because a failing phase must not take the boot down, which is
what the original did.

WHY THE TEST IS THE POINT. Splitting the function is easy; keeping it split is
not, and the thing worth guarding is not the line count — it is THE ORDER.
Boot ordering is load-bearing in ways invisible at the call site: a reviewer
looking at sixteen await lines cannot see that moving one breaks something,
because the coupling lives in the bodies. Before this it was implicit in a
function nobody could read in one sitting; after it, it is a list — which is an
improvement only if something enforces the list.

So `tests/unit/test_1028_lifespan_phases.py` pins the sequence WITH the reason
for each constrained pair recorded beside it: logging first so a later hang
cannot swallow the boot log (#858); the event bus before any WebSocket client
needs a live dispatcher (#306); Docker/system-agent before the fleet sweepers;
startup recovery before the channel transports, so inbound traffic cannot create
an execution that races the reconcile; the event-bus drain LAST on shutdown so
late broadcasts still land. A reorder now fails with the reason attached instead
of surfacing weeks later as a boot bug nobody connects to this commit.

It also pins the `yield` split (a phase appended after it silently becomes
shutdown work), the per-helper thresholds the issue asked for (<100 lines,
CC <20), and orphan/double calls. Mutation-checked five ways — recovery moved
after the transports, event bus after Docker, a dropped shutdown phase, the
drain no longer last, a phase pushed past `yield` — all caught.

ONE DEFECT THIS FOUND IN ITSELF, worth recording because it is the failure this
refactor's shape invites. The extraction moved `@asynccontextmanager` by one
definition: it landed on the first phase helper and `lifespan` was left a bare
async generator, which FastAPI cannot use as a lifespan. Boot-breaking — and the
entire 12,900-test unit suite stayed green, because nothing in it imports `main`
and asks what shape `lifespan` is. It surfaced only from an explicit import
check (`iscoroutinefunction` on each helper returned False for one, with
`co_filename` pointing into contextlib). Fixed, and pinned by its own test.

Docs: the `main.py` row in architecture.md now records that the order is the
contract and where the constraints are, so the next person to add a startup step
knows it belongs in a phase helper.

Scope: one file per the issue's own recommendation. The remaining ACs
(`routers/settings.py`, `routers/ops.py`, `services/git_service.py`,
`services/agent_client.py`, `routers/public.py::public_chat`) stay open on #1028.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(review): three phase helpers used locals the split left behind (#1028)

/review's scope pass caught a real runtime break in my own extraction, and the
interesting part is that three separate verifications had already passed on it.

`main.py` never imports `database` at module level. The old `lifespan` did
`from database import db as _db` once near the top, and three blocks 100+ lines
later used it through the enclosing function scope: the system-agent
`setup_completed` gate, the Telegram transport and the WhatsApp transport.
`message_router` was the same shape — imported in the Slack block, used by the
Telegram one. After the split all four are NameError.

The severity is in the swallow. Every one of those use sites sits inside
`try/except Exception`, so the boot SUCCEEDS: the log carries "Error starting
Telegram transport: name '_db' is not defined" and the Telegram and WhatsApp
integrations are simply never wired. A silently dead integration, not a failed
boot — and the per-phase guard that makes each phase fail-open is exactly what
hides the extraction bug.

Why the existing checks missed it, all three: the AST equivalence proof compares
LINES and the lines are identical; the structural pin asserts order, thresholds
and decorators, not name resolution; and the `import main` smoke never RUNS
`lifespan`, so nothing resolves those names at import time.

Fixed by re-materialising each import in the helper that needs it — the
original's own idiom — with a comment saying why it is there, so a later reader
does not "tidy" it back out.

CORRECTION TO THE CLAIM: the bodies are no longer byte-identical. They are
verbatim EXCEPT these four re-materialised imports, which is now what the PR
body and the docstrings say. A verbatim claim stops being true the moment a
leaked name has to be restored, and quietly keeping the claim is worse than the
bug.

Also pinned, because this gets more likely with every future split of the same
function: test_no_phase_helper_depends_on_another_phases_locals asserts, per
helper, that `names_loaded - names_bound - module_globals` is empty. Mutation-
checked by deleting the restored `_db` import — reproduces the shipped bug and
turns the suite red.

Also: the phase-count docstrings said "of 10" in 9 helpers; the transports were
split into three after that text was written, so it is 12.

learnings.md gains the class: extract-method has a failure mode the diff cannot
show and an import smoke cannot reach.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(review): the verbatim claim survived in three docstrings (#1028)

Re-review finding. The previous commit corrected "bodies are byte-identical"
in the PR body but left the same claim standing in the code, where it is more
likely to be believed: the three helpers that gained a re-materialised import
still said "the body below is unchanged".

A stale claim next to the exact line that falsifies it is worse than no claim —
it is the thing a future reader checks against before deciding the import looks
redundant. Now each says verbatim EXCEPT the restored import, and points at the
comment explaining why it is there.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): two pre-existing lifespan guards read a function the split emptied (#1028)

CI's regression diff caught 4 new failures, deterministic across all three
seeds. Both guards assert properties of `lifespan`'s SOURCE, and #1028 moved
that source into phase helpers — so they were asserting things about a function
that is now 25 lines of `await` calls.

The properties still hold. The guards had stopped being able to see them, which
is the worse failure: a guard that silently stops covering its subject reads
identically to one that passes.

test_1267_lifespan_db_alias — and this one is pointed: #1267 IS the bug class
/review caught in this branch. It fired when the transport blocks called a bare
`db` while only `_db` was in scope, NameError swallowed by the surrounding
try/except and surfaced as a misleading "Error starting Telegram transport".
The split re-introduced the same class in a new form (`_db` bound in phase 1,
read in three later helpers), and this guard could not see it because it only
ever looked inside `lifespan`.

So it now follows the calls: `_lifespan_surface()` returns `lifespan` plus the
helpers it awaits, and every check scans all of them. The alias check is
STRENGTHENED rather than merely relocated — binding `_db` somewhere on the
surface is no longer sufficient, because after the split each helper is its own
scope, so every function that READS `_db` must bind it. That assertion fails on
the exact defect this branch shipped.

test_858_dockerfile_unbuffered — the #858 invariant is an ORDERING one
(setup_logging -> first-run notice -> event_bus.start), and after the split
those three sit in three different functions. `_lifespan_body()` now flattens
the phases inline in call order, so the existing index comparisons keep meaning
what they meant. An unresolvable helper is left as the bare `await` rather than
skipped, so a phase this cannot expand can never silently drop the statements
it contains.

No production code changed.

Related to #1028

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(retention): every install writes its own retention rows, not just fresh ones (#2085) (#2432)

#1645 closed #1638 by reverting OPS_SETTINGS_DEFAULTS to the wide historical
values and applying the #1039 community floor through explicit system_settings
rows seeded on FRESH installs only. Every install that has ever upgraded rather
than been created fresh therefore had no rows at all, so cleanup_service
resolved all 11 windows at prune time from a dict that ships inside the backend
image and is replaced on every rebuild. The only thing between a future edit to
that dict and the #1638 failure mode — a silent hard-DELETE of existing data
seconds after the next boot, green /health, no error — was a code comment.

database._seed_retention_windows{,_engine} now writes an explicit row for every
RETENTION_OPS_KEYS member that has none, at the value already in force, on every
boot and regardless of install age. Behaviourally inert: it writes the number the
prune already used, so nothing prunes differently the day it runs.

Three properties are load-bearing:

* The key set is DERIVED from RETENTION_OPS_KEYS, never a second hand-written
  list, so a window added later is covered the day it ships instead of quietly
  inheriting the image default forever (the issue text said "eight windows"; it
  was 11 by the time this landed — ent#433 added two, #2216 a third).

* Ordering. It MUST run after _seed_fresh_install_retention. Both writers are
  insert-or-ignore, so the first to reach a key wins: reversed, a fresh install
  silently gets the wide defaults instead of the #1039 floor — the community
  floor deleted by the change meant to protect retention. Pinned behaviourally
  and by a source-order guard on both the SQLite and engine arms.

* It must actually run. The first cut imported the constants from
  services.settings_service, which has a module-level `from database import db`.
  init_database() is called from DatabaseManager.__init__, i.e. while database.py
  is still executing its own module body, so database.db does not exist yet and
  that import raises ImportError — which this seed's fail-safe contract then
  SWALLOWS. The feature was dead on every boot with a fully green unit suite,
  because every in-process test calls the function after database has finished
  importing. RETENTION_OPS_KEYS, OPS_SETTINGS_DEFAULTS and
  NON_ROW_RETENTION_OPS_KEYS therefore move to config.py (a leaf, already home to
  OPS_SETTINGS_VALIDATION and validate_ops_setting for these same keys) and are
  re-exported from settings_service, extending the pattern
  COMMUNITY_FRESH_INSTALL_SEED already used for exactly this reason. Two tests
  now pay for a real subprocess import; both fail if the import is reverted.

Stated tradeoff: a seeded install stops inheriting later changes to the code
default in EITHER direction, so widening a window for existing installs becomes a
deliberate migration rather than something that arrives silently with an image.
That is the intended consequence — retention becomes explicit per-install config
instead of implicit inheritance from whatever image happens to be running,
symmetric with the rule the OPS_SETTINGS_DEFAULTS comment already imposes on
narrowing.

backup_retention_days is seeded too. That makes OPS_SETTINGS_DEFAULTS' value the
one that lands in the DB for a key whose private reader
(db_backup_service.effective_backup_retention_days, inverted coercion) falls back
to its own module constant; the two are now parity-tested.

No schema change, no migration — row inserts at boot, same as the #1638 seed.

Not fixed here: generic DELETE /api/settings/{key} carries no RETENTION_OPS_KEYS
guard (only PUT does), so an admin can still delete a window row. After this it is
transient — the next boot re-seeds it — but the asymmetry with PUT remains.

Unblocks ops#300 once deployed: with rows on every install, /update step 8e can
drop its source-text guessing for a plain assertion over stored values.

Closes #2085

* fix(watchdog): stop false-orphaning executions parked before the agent spawns them (abilityai/trinity#2433) (#2435)

* fix(watchdog): stop false-orphaning executions parked before the agent spawns them

The cleanup watchdog's proof-of-life (GET agent/api/executions/running:
running ∪ recently-completed) could not see an admitted execution that was
waiting in the backend's global agent-call queue, in the agent's CPU-sized
default thread pool, behind the agent chat lock, or in the post-exit drain
before unregister(). After the 60s grace it wrote a false `failed`
("completed on agent but status not reported"), released the slot, and the
parked call then ran anyway — billed, overbooked, its late 200 silently
overwriting the row (#378). Reproduced twice locally; three mechanisms, one
string.

Orphan now means: the agent does not know the execution AND no live backend
dispatcher owns it.

- agent server: /api/executions/running gains `pending_ids` (accepted at
  /api/task, /api/chat and the #1083 async spawn but not yet spawned; lazily
  expired) and `recently_completed_ids` covers exited-but-registered handles.
  Cancel-while-pending is consumed by register() (SIGKILL at spawn, #679 marker
  kept); the pre-spawn 409 is only an optimisation. Headless runs use a
  dedicated 32-thread pool pinned to MAX_PARALLEL_TASKS_CEILING_MAX; the Gemini
  runtime now registers its subprocess at both Popen sites (it never did).
- backend: every outbound agent call is registered for its whole lifetime
  (track_inflight_dispatch — queue wait, connect retries, POST) in an
  in-process registry plus a cross-worker Redis liveness marker
  execution:inflight:{id} (60s TTL, one refresher task per process, 15s tick).
  The watchdog reads a tri-state verdict (alive / absent / unknown) and
  withholds recovery on `alive`, and on `unknown` only while a dispatcher could
  still own the row; a process with no Redis reads `absent` (its own registry
  is the whole truth). CleanupReport.dispatch_inflight_skipped counts withheld
  rows; the orphan error string states what was observed.
- a park no longer spends the run's budget: at grant, a park ≥ 5s restamps
  started_at (admission kept in queued_at, the drained-backlog shape, CAS on
  RUNNING + NULL lease) and renews the slot lease (ZADD XX + EXPIRE together);
  the refresher renews the slot every tick while parked.
- parked rows are cancellable and agent-scoped: terminate consults the
  in-process registry, then the cross-worker cancel key; a parked phase is
  finalized CANCELLED and the grant raises BackendAgentCallCancelled, where the
  dispatcher writes CANCELLED itself (never FAILED; the /chat arm answers 409).
- terminate_execution gains ONE agent-scope gate at its entry for all three
  arms: the row behind the caller-supplied task_execution_id must belong to the
  agent the route proved (uniform 404; an unreadable row fails closed with
  503). The proxy arm's 404 scoped only execution_id while the CANCELLED CAS
  was keyed on task_execution_id, so a caller authorised on agent A could flip
  agent B's running row (found by the /cso --diff verifier; report under
  docs/security-reports/).
- packaging: BACKEND_AGENT_CALL_LIMIT / BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S
  forwarded in prod + hosted compose and documented in .env.example; the >5s
  queue-wait warning fires on both acquire branches.

Verified: full unit suite under CI conditions 12969 passed / 0 failed
(baseline origin/dev 12863 / 0); Repro A 10/10 success (2 parked 485s,
withheld at both watchdog cycles, re-anchored at dispatch); Repro B 8/8
success (5 parked, two waves); live pending_ids probe on the agent. The
agent-side half needs a rebuilt base image; the backend half alone covers old
images through the whole-call marker.

Fixes abilityai/trinity#2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(watchdog): close the cross-worker cancel race and bound the exited-but-registered set (#2435 review)

Review of the #2433 fix found that it reintroduced the #378 symptom in a
narrower window and turned a pre-existing registry leak into a permanent one.

1. Cross-worker cancel acted on a marker phase that predated its own write.
   `entry.phase` flipped parked->calling in memory only; the marker was
   rewritten by the 15s refresher, so `execution:inflight:{id}` advertised
   `parked` for up to a full tick after the POST had begun. Under --workers 2
   about half of all cancels are served by the worker that does NOT own the
   coroutine and therefore read it: the row was finalized CANCELLED and its
   slot released while the agent ran the turn to a billed completion whose
   SUCCESS then lost the CAS. Closed by ordering, not by narrowing — the owner
   publishes the transition in the SAME round-trip that reads the cancel key
   (`_publish_calling_and_check_cancel_sync`), and the remote sets the cancel
   key BEFORE re-reading the phase (`_set_cancel_then_reread_phase_sync`), so
   an observed `parked` gives W_remote(cancel) < R_remote(marker) <
   W_owner(marker) < R_owner(cancel) and the grant is guaranteed to see the
   key. Neither side pays an extra round-trip. The owner gates the publish on
   the ENTRY's age rather than this attempt's park, because
   `track_inflight_dispatch` wraps the whole retry loop and a retry can grant
   instantly under a marker a tick left saying `parked`; the remote's scope
   check stays on its first read, so no key is written for a foreign agent.

2. `list_recently_completed_ids` reported exited-but-registered ids with no
   age bound, so a leaked entry was agent-known forever and the watchdog never
   recovered that row — a regression against pre-#2433, where `list_running()`
   self-healed it. Now bounded by the same 300s TTL as the buffer, measured
   from when the exit was first OBSERVED (not `started_at`, which would drop a
   long turn the moment it entered its drain). The leak is also closed at
   source: `register()` SIGKILLs the group for a cancel that arrived while
   pending, so the following `stdin.write` can raise BrokenPipeError — all
   three prompt-writing runtimes (claude_code, gemini x2) now pair that write
   with `unregister()` on failure.

3. `restamp_execution_dispatch` is a sync sqlite write and ran on the event
   loop, while both semaphores are held and the queue is by definition
   congested. Now `asyncio.to_thread`, like the slot renewal beside it.

Smaller items from the same review:
- /api/chat sizes its pending entry to PENDING_CHAT_TIMEOUT_SECONDS (7200s):
  `ChatRequest` carries no timeout and a chat can wait on the execution lock
  for the agent's whole budget, so the /api/task default evicted the entry
  mid-wait. Its discard now wraps the lock acquisition, so a request cancelled
  while waiting (client disconnect) cannot leak one.
- Phase 3 batches its in-flight verdict read (one MGET per cycle, not per row),
  matching Phase 0.
- `renew_slot` refuses, score untouched, when the metadata hash has already
  expired: `ZADD XX` succeeds while `EXPIRE` no-ops, so it used to report a
  renewal it had not performed and re-anchor exactly the ZSET-without-hash
  state canary S-03 calls `missing`.
- `register_pending` logs at DEBUG (it fires on every /api/task and /api/chat).
- Documented that the in-flight marker is not eviction-proof under the prod
  `allkeys-lru` policy.

Tests: tests/unit/test_2433_review_fixes.py (15) — 11 of them fail against
cfc2cfef, verified in a worktree. Full unit suite under CI conditions
(clean origin/dev worktree, no submodules): 12985 passed, 0 failed.

Refs abilityai/trinity#2433

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* test(resume): the guard architecture.md promised did not exist (ent#430 review)

The reviewer's one condition before merge. `architecture.md`'s 'Two callers,
one rule' bullet said the CAS-win rule was 'guarded now by enumerating every
caller rather than the one route ent#329 knew about, so a third site inherits
the rule instead of re-losing it'. No such guard existed:
`test_dispatch_hangs_off_the_cas_win_only` read exactly one hardcoded file,
`routers/operator_queue.py` — so `client_portal/asks/service.py`, the caller
this PR adds and the one that LOST the rule, was outside its reach.

A sentence claiming protection that is not there is worse than no sentence: the
next person adding a dispatch site reads it and stops looking. This is the shape
#2428 filed a learnings entry about this morning — a comment that names a
failure mode is a request for a guard — so it lands the same way.

DISCOVERED, NOT LISTED. `_dispatch_call_sites` walks the backend tree for
callers, because a hardcoded list structurally cannot catch the case that
matters: the file it would need to check is the one being added.

ASSERTED AGAINST CODE, NOT FILE TEXT — and this is the part I got wrong first.
The initial version tested `"_status_conflict" in source` against the raw file
and MUTATION PROVED IT BLIND: deleting the check from the `if` still passed,
because the long comment above it explaining the race still contained the
string. A source-substring guard cannot tell a check from a paragraph about the
check — the same defect the guard exists to prevent, inside the guard. It now
parses each dispatching function and compares `ast.unparse` output, where
comments do not survive.

Verified by three mutations, each caught:
  1. delete the check in asks/service.py, keep the comment  -> FAIL
  2. neuter the check in routers/operator_queue.py          -> FAIL
  3. add a brand-new third caller with no check at all      -> FAIL
and all 23 pass on the real tree.

`test_the_discovery_walk_finds_both_known_callers` pins the floor, so a rename
of the helper cannot leave the loop iterating an empty list and passing in
silence — the failure a discovery guard trades for the one it fixes.

ALSO (non-blocking, from the same review): `WorkspaceAsk.status`'s comment still
read 'pending | expired (terminal ones are not listed)' after `_status_of`
gained a third value. Corrected to say where each value is reachable from.

The remaining non-blocking item — `resume_requested` and the new `answered`
status are unconsumed by any surface — is deliberately NOT in this commit. It is
a product decision about where a transient confirmation lives, and it is filed
so it stays a decision rather than becoming an oversight.

Related to abilityai/trinity-enterprise#430
Related to abilityai/trinity-enterprise#329

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(systems): the four post-deploy endpoints — a nonexistent DB call, an ungated restart, a broken export round-trip, and prefix-collision membership (#2373)

The deploy half has been hardened by every commit since ent#124; the four
post-deploy endpoints were essentially untouched since 2025.

## Membership is now ONE predicate

`get_system`, `restart_system` and `export_manifest` each matched
`startswith(f"{system_name}-")`, so an operation on `acme` also captured every
agent of a system named `acme-extra` — including `restart`, which stops and
starts containers. Three copies of a wrong rule.

`system_service.system_member_names` is the one rule, and it prefers TAGS:
`configure_tags` already applies the system name to every member, so a tag is a
RECORD of membership where a prefix is an inference from a naming convention.
The prefix survives only as a fallback for pre-tag deployments, narrowed so an
agent claimed by another system's own tag is excluded — a tagged `acme-extra`
agent is never captured by `acme` even there. A failing tag read degrades to the
prefix rather than 500ing.

Residual, stated rather than hidden: two systems deployed BEFORE tagging where
one name is a prefix of the other remain ambiguous, because nothing distinguishes
them. This is also the prerequisite for the teardown verb, where the same
collision would delete rather than restart.

## GET /{name} returns real schedules

It called `db.get_agent_schedules`, which does not exist — the facade exposes
`list_agent_schedules` and `database.py` deliberately has no `__getattr__`
fallback. The AttributeError was swallowed by the surrounding `except
Exception`, so every response omitted `schedules` for every agent and logged one
warning each, while `tests/test_systems.py` never asserted on the key. Exactly
the failure mode the db facade's own comment warns about — so the test also pins
that the fallback stays absent, since adding one would turn the next typo into a
silent Mock.

## POST /{name}/restart is creator-gated

It was bare `get_current_user` — below `POST /deploy` and below even the
READ-ONLY bundled-catalog routes — so any authenticated principal, including
`role: user`, could stop and start every container in a system whose agents it
could see. A mutating fleet-wide verb under a lighter gate than the catalog it
reads is an oversight, not a decision. `require_role` also rejects agent
principals (#1890), which matters because an agent-scoped MCP key resolves to
its owner carrying the owner's role.

## Export round-trips

The non-full-mesh permissions branch sliced `target_agent[len(name)+1:]` with no
membership filter — the sibling branch had one — so an edge pointing outside the
system exported as a blind-sliced garbage short name that then failed
`validate_manifest`'s unknown-agent check on re-deploy. The export broke its own
round trip. Both branches now test membership.

And the export no longer embeds the instance-global `trinity_prompt` as the
manifest's `prompt:`. Deploying that manifest elsewhere overwrote THAT
instance's platform-wide prompt — a fleet-wide side effect from what reads like
a copy of one system. Nothing records whether the source system ever set a
prompt, so there is no honest way to distinguish it from whatever the instance
happens to have configured, and the only correct export of an unknown is to
omit it.

## Two preview hardenings

Unknown PER-AGENT keys now warn like top-level ones (ent#126): `credentials:`,
`skills:` and `display_label:` are the fields people try first and they vanished
in silence.

Preview and deploy now resolve the identical resource default. Deploy hardcoded
`{"cpu": "2", "memory": "4g"}` while `_preflight_template` validated against the
admin-configurable `get_agent_default_resources()`, so the two disagreed the
moment an admin moved the fleet default — the one spot that escaped ent#126's
pure-resolver no-drift pattern.

## Verification

14 unit tests, one per defect plus the exempt shapes. Two mutation-checked: the
restart gate and the tag-first membership each turn a test red when reverted.
414 pass across the system/manifest/ent#126/#1884 suites.

`tests/test_systems.py` is live-backend tier and…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants