feat(auth): bound, attribute and fail-closed the machine credential for admin/ops APIs (#2323) - #2389
Conversation
…or admin/ops APIs (#2323) The issue asks for a service credential that survives enforced 2FA, on the premise that none exists. One does. Verified against a live instance: a `scope='user'` MCP key owned by an admin returns 200 on every admin/ops surface the issue lists, writes included; `mfa_gate.gate_login` has exactly two call sites, both login routes, so key validation has never passed through it; keys are revocable and rotate by minting a second one. What that credential is NOT is bounded, attributable, or expiring. It carries the owner's full role, `mcp_api_keys` has no expiry, and its use recorded byte-identically to the owner clicking in a browser. So the only way to enable 2FA and keep a dashboard alive was to hand it a permanent, invisible, unlimited admin key — a worse posture than the control it works around. The issue's proposed shape (a tier that *unlocks* admin) is therefore inverted: `user` already does that. The tier needed is one that *bounds* it. Three changes, smallest blast radius first. 1. The admin gate becomes an ALLOWLIST. ent#293/#297 closed it against `agent` and `connector` by NAME. `scope` is free-text with no CHECK constraint, so that is a denylist over an open space: any scope setting neither `agent_name` nor `connector_agent` walks both rejections and inherits the owner's role across ~163 admin-gated sites. `models.User`'s own docstring predicted it — "fail-closed against a sixth scope a future PR invents". This PR's `ops` scope is that sixth. `ADMIN_GATE_SCOPES = {None, "user", "system"}` is exactly the set that passes today, so this is zero behaviour change with one exception: `portal_delegate` passed both named guards and was contained *only* by its route fence, and now has a second layer. A principal lacking `mcp_scope` fails CLOSED via a sentinel — never `getattr(..., None)`, which would make an absent authorization discriminator the privileged JWT value. 2. Attribution comes from the credential, not from a header. `validate_mcp_api_key` has always returned `key_id`/`key_name`; `get_current_user` discarded them. Carrying them on `User` and deriving the three `mcp_*` audit columns from `actor_user` fixes ~70 call sites with no diff at any of them, and adds `mcp_key_id`/`mcp_scope` filters so "what did that leaked key touch?" is answerable. `actor_type` stays `user`: the owner is accountable, is the only branch yielding an email, and the enterprise user-activity view matches on it. The `X-MCP-Key-Id` header is removed from the audit path rather than out-ranked. It is `Header(None)` on six routes, validated nowhere, and persisted into the backlog replay blob — so honouring it let any authenticated caller forge the credential named in the two highest-volume audit events on the platform, with the forgery surfacing minutes later on queue drain. This half is a SECURITY FIX. `routers/a2a.py::_a2a_idem_scope` builds an idempotency scope from `mcp_key_id` and fell through to `username` because the field did not exist, so two agent-scoped keys of one owner shared a peer-controlled `messageId` namespace: caller B received caller A's full response text and B's task never ran. Its own docstring describes the failure it was not preventing. Reachable on an entitled install with >=2 `a2a_exposed` agents under one owner. The test asserting the distinction used a stub carrying a field the real principal lacked — green over a property production never had — and is now pointed at a real `User`. DEPLOY NOTE: the scope string moves from `a2a:{agent}:{username}` to `a2a:{agent}:{key_id}`, so a `messageId` replayed across the deploy re-executes instead of replaying. Bounded by the 24h TTL. 3. The `ops` scope: read-only, route-fenced, self-authorizing. Admin-minted and human-only (`reject_non_interactive_principal` — the guards used for `portal_delegate` are both no-ops for an ops principal, so an ops key could otherwise mint ops keys). Fenced at the single auth entry point beside the connector/ephemeral/portal_delegate fences. Every entry is a GET, asserted by a test that imports the constant: without that belt a prefix entry would admit a future `POST /api/ops/*`, and that is where `emergency-stop` lives. Kept OUT of `ADMIN_GATE_SCOPES`; ops reads opt in with `assert_admin(..., allow_scopes={"ops"})`. Authority therefore comes from being an ops key rather than from the owner's role — so it keeps working when that admin is offboarded, and a new ops route is inaccessible until granted rather than silently reachable. The fence set is the MEASURED read set of the real consumer (Trinity Control), not the issue's wording, which named only `/api/ops/*` and would have shipped a credential unable to run the dashboard it exists for. Never carries an `agent_name`: three sweeps find their work by filtering `scope IN ('agent','connector')`, so a non-agent scope holding one is invisible to all three. Excluded from the MCP tool surface by construction — `OPERATOR_SCOPES` is an allowlist pinned by its own test. Guards, mutation-tested three ways (add a write to the fence, drop a route the consumer needs, widen the admin allowlist — each detected): - fence is all-GET and fully anchored, asserted by importing the constant, never by grep, which would pass on the prose; - route parity resolved against live route objects, spanning both declaration families, plus the inverse direction (every non-allowlisted ops route denied), which is the half that catches a future destructive route; - admin gate rejects `ops`, `portal_delegate` and an invented `scope_from_2027`, and still admits the three that passed before. 44 test doubles updated: 38 principal stand-ins that did not carry `mcp_scope` and 6 stubs replacing the admin gate with a narrower signature. Every one claimed to stand for the real principal while not matching it — the same defect that hid both live bugs above, which is why the gate fails closed rather than defaulting. NOT delivered, stated rather than implied: key expiry; narrowing the existing `user` scope (would break the fleet); a write-capable ops tier — the ops toolkit's 24 writes stay on password auth, so the read fence does not retire the admin password. Two costs are inherited, not introduced: key validation writes per request, and fleet-status fans out per agent. The Observatory already polls exactly these endpoints on an admin JWT at the same cadence. Honest bound: this narrows the API surface only. Ops tooling that mutates containers over SSH never touches the API; SSH remains the real privilege boundary on those hosts. Full unit suite: 12313 passed, 21 failed — byte-identical to the failure set on an unmodified worktree at this branch point. Zero regressions. Fixes #2323
…d accent-blue `accent-*` is a closed set of one (`accent-purple`); `check:tokens` caught the invention in CI. An informational badge is what `status-info` is for, so the right fix is the semantic token rather than widening the palette for one badge.
Code review —
|
…ted, and no route accepts a credential name (#2389 review) Six findings from review on #2323, all verified against the code before fixing. **High — admitted by the fence, refused by the gate.** `GET /api/subscriptions/{id}/usage` is in `_OPS_ALLOWED_ROUTES` but its handler called bare `assert_admin`, and `ADMIN_GATE_SCOPES` excludes `ops` — so the subscription-pressure read the fence was measured for could not work. The admit-set test passed because it exercised only `_enforce_ops_key_fence`, never the endpoint's own gate. Fixed with `allow_scopes={"ops"}`, and guarded as a CLASS: a live-handler scan resolves every allowlisted GET from the router modules and reds if any calls `assert_admin` without the opt-in. **Medium — `/ws/events` was a second auth entry point.** That handler validates the key itself and never runs `get_current_user`, so none of its fences reached it and the ops fence's own docstring claim was false for exactly one surface — the broad one. The stream carries fleet-wide activity and execution events scoped by the OWNER's accessible agents. `WS_EVENT_STREAM_SCOPES = {None, "user", "agent", "system"}` gates it (close 4003), closing the same hole for `connector` and `portal_delegate` rather than preserving it out of politeness. **Medium — the forgeable header still reached durable storage.** #2323 removed `X-MCP-Key-Id` from the audit path only. Five routers still declared it and wrote it into durable provenance columns (`schedule_executions`, `agent_loops`, `agent_reminders`, fan-out rows), and `backlog_service` still persisted both names into `backlog_metadata` — the longest-lived copy of a request, the surface canary G-04 scans and #1449 scrubs — where `_spawn_drain` never read either key back. One request produced two provenance records that disagreed and the forged one outlived the honest one. Every writer now derives from the validated bearer on `User`; NO route declares either header, guarded by a router-tree scan so the class cannot return one endpoint at a time; the blob no longer carries them (a pre-existing queued row drains unchanged — the drain reads key-by-key with `.get()`). The dead parameters are removed from the whole `chat_execution_service` / `capacity_manager` / `backlog_service` chain rather than accepted-and-ignored, per the rule #2323 stated one function below the offending code. **Low — a guard that could never fire.** `_AGENTLESS_SCOPES` read `getattr(key_data, "agent_name", None)` on a `McpApiKeyCreate` that declares no such field, so Pydantic made it permanently falsy: protection in appearance only. The guard is removed and the invariant is now asserted where it is real — a test that drives the creator and reads the written row back, instead of asserting a tuple contains a string. **Low — dead forgeable parameters** on `admit_chat_request`, and a `_user()` helper in test_1310 that accepted `mcp_scope` and dropped it, letting a future test assert a pass the real principal would never get — the stand-in-does-not-match-production defect this work exists to eliminate. Verification: full unit suite 12383 passed / 21 failed, against a baseline run at the branch point of 12373 passed / 21 failed — byte-identical failure set (IPv6/SSRF, environment-dependent). Zero regressions, +10 new passing.
|
All six findings fixed in b996f3c. Each was verified against the code before touching it; three of them were bigger than reported, so here is what actually changed. 1 (High) — admitted by the fence, refused by the gate ✅
Your last line is the important one: the admit-set test exercised only 2 (Medium) —
|
| Writer | Column |
|---|---|
routers/chat.py (/chat, /task) |
schedule_executions.source_mcp_key_id |
routers/schedules.py (manual trigger) |
schedule_executions.source_mcp_key_id |
routers/loops.py |
agent_loops.source_mcp_key_id |
routers/reminders.py |
agent_reminders.source_mcp_key_id |
routers/fan_out.py |
fan-out execution rows |
plus backlog_service persisting both names into backlog_metadata — and _spawn_drain never reads either key back, so it was a forgeable value stored, never reconstructed, in the one blob canary G-04 scans and #1449 scrubs.
All five now derive from current_user.mcp_key_id/mcp_key_name, and the fix is asserted in the strongest available form: no router declares either header at all, checked by a scan over the whole router tree. Grepping for the use is what let this survive #2323; a value nothing declares cannot be forged, threaded, or picked up by a new consumer.
The dead parameters are removed from the entire chat_execution_service → capacity_manager → backlog_service chain (finding 5's rule, applied at scale — 17 signatures), not accepted-and-ignored. A pre-existing queued row still carrying the two blob keys drains unchanged, since the drain reads the blob key-by-key with .get().
The MCP server still sends the headers. Left alone deliberately: nothing on the backend reads them, and the value it sends is the same key the bearer already identifies — so removing it is an MCP rebuild for no behaviour change, and keeping it is one less rolling-deploy edge.
4 (Low) — the guard that could never fire ✅
Correct, and the right fix was deletion. getattr(key_data, "agent_name", None) on a model that declares no such field is permanently falsy — protection in appearance only, which is worse than none. The guard is gone; _AGENTLESS_SCOPES stays as the documented invariant, and test_ops_keys_must_never_carry_an_agent_name (which asserted a tuple contains a string) is replaced by a test that drives the real creator and SELECTs the written row.
5 (Low) — dead parameters on admit_chat_request ✅
Removed, along with the router's pass-through.
6 (Low) — _user(mcp_scope=...) swallowed ✅
Forwarded. Same defect in test_connector_auth.py::_user, fixed with it.
Verification: full unit suite 12383 passed / 21 failed, against a baseline run of the branch point in a clean worktree at 12373 passed / 21 failed — byte-identical failure set (IPv6/SSRF, environment-dependent). Zero regressions, +10 new passing. Docs updated in architecture.md (scope table + the /ws/events second-entry-point note) and requirements/security.md.
…ng the database singleton Two shards of the last CI run (head, seeds 12345/67890) hit the 25-minute job cap. The cause is runner-side, not code — per-test durations are unchanged between base and head (JUnit sums: head 777s vs base 750/794/841s; the two endemic stalls, `test_start_agent_skip_inject` at 2x60.5s and `test_1083_result_callback::TestPersistResend` at ~90s, are identical on both sides, and head seed 99999 finished in 810s). The failing shards lost 1050s and 959s to stalls *between* progress lines, on second-wave matrix runners. But the new temp-DB test had a real ordering-dependent defect that the failure made worth finding, and it is exactly the class this PR argues about — a test double that does not match production: - It set `TRINITY_DB_PATH` only. `db.engine.resolve_database_url()` prefers `DATABASE_URL`, so under any ordering where an earlier test left that set, this test silently created tables and INSERTed an admin user plus an ops key into someone else's database instead of its own temp file. - It never disposed the URL-keyed engine cache, so it could both miss its own engine and leave one holding a handle on a temp file about to be deleted. - It imported `database`, whose module-level singleton runs the entire SQLite migration chain inside `__init__` against whichever URL is active at first import — making the whole session's `database.db` a function of test order. Now follows the repo's established shape (`test_918`, `test_idempotency`): `DATABASE_URL` at the temp file, `dispose_engines()` on both sides, and `McpKeyOperations(UserOperations())` constructed directly so no migration chain runs at all. Same assertion, still reading the written row back.
Code review — re-review at
|
| # `assert_admin(..., allow_scopes={"ops"})`. | ||
| # | ||
| # Layer 2 is what makes this a machine identity rather than a human's proxy: | ||
| # authority comes from being an ops key, not from who owns it. It also flips the |
There was a problem hiding this comment.
MEDIUM — the documented "survives the minting admin being offboarded" property is not implemented.
Both this comment and docs/memory/requirements/security.md:314 claim authority comes from being an ops key, not from who owns it. It does not: assert_admin(..., allow_scopes={"ops"}) runs _reject_scope_at_admin_gate and then still enforces current_user.role != "admin", and get_current_user rejects the principal outright when the owner carries suspended_at.
So: admin Alice mints an ops key for the monitoring dashboard and later leaves. Offboarding demotes or suspends her account, and the ops key immediately 403s on all six /api/ops/* reads and on /api/subscriptions/{id}/usage — precisely the "every ops integration dies when that admin is offboarded" outcome the doc says this design prevents.
An operator will act on that sentence during an offboarding. Either drop the role check for the opted-in scope, or correct both claims.
| _reject_scope_at_admin_gate(current_user, allow_scopes) | ||
| if current_user.role != "admin": | ||
| raise HTTPException(status_code=status.HTTP_403_FORBIDDEN, detail=detail) | ||
|
|
There was a problem hiding this comment.
(anchored to the nearest line in this PR's diff; the finding is at src/backend/dependencies.py:1001)
MEDIUM — require_admin has no allow_scopes, and the new CI guard cannot see it.
require_admin is the Depends form and the dominant admin-gate spelling platform-wide, but it gained no allow_scopes parameter — so an ops-allowlisted route gated that way is permanently dead to ops keys with no opt-in available. And tests/unit/test_2323_machine_identities.py:365 only regex-matches assert_admin\( inside handler source, so it cannot see Depends(require_admin) in a signature.
Someone converts /api/ops/fleet/status to Depends(require_admin) for consistency with the enterprise routers: the ops dashboard silently 403s, and CI stays green.
That is exactly the two-gate false-assurance class the guard was written for after the /subscriptions/{id}/usage bug — the guard closed the instance and left the shape open. Either give require_admin a parameterised form, or have the guard reject any allowlisted handler whose source mentions require_admin.
| # them everything else — this closes a hole that predates #2323 rather than | ||
| # preserving it out of politeness; both are fenced to one or two routes, so | ||
| # nothing legitimate reaches this one. | ||
| WS_EVENT_STREAM_SCOPES = frozenset({None, "user", "agent", "system"}) |
There was a problem hiding this comment.
LOW — WS_EVENT_STREAM_SCOPES admits "agent" wholesale, including ephemeral ghost keys.
/ws/events authenticates the key itself and never calls get_current_user, so _enforce_ephemeral_key_fence — whose entire purpose is that a ghost's key on an untrusted workspace must not be a fleet skeleton key — does not reach it. A ghost's own TRINITY_MCP_API_KEY therefore opens the fleet-wide event stream, scoped by get_accessible_agent_names(owner_email, is_admin), i.e. everything on a default admin-owned install.
Pre-existing, but this PR is the one that enumerated the permitted scopes on that surface and closed the same hole for connector and portal_delegate; ghost keys are the one bounded principal left in. Gating "agent" on the row's is_ephemeral (the predicate the fence already uses) would finish it.
| </span> | ||
| <span v-else-if="key.scope === 'system'" class="ml-2 inline-flex items-center px-2 py-0.5 rounded-full text-xs font-medium bg-status-urgent-100 dark:bg-status-urgent-900/50 text-status-urgent-800 dark:text-status-urgent-300"> | ||
| System | ||
| </span> |
There was a problem hiding this comment.
(anchored to the nearest line in this PR's diff; the finding is at src/frontend/src/components/settings/McpKeysTab.vue:496)
LOW — the UI advertises the tier but cannot mint it.
The tab renders an "Ops (read-only)" badge, but the create modal only ever sends scope: 'portal_delegate' — there is no ops option in newKey. An operator reads the release note, opens Settings → MCP Keys, and finds no way to create the credential the badge describes. The tier is API-only.
| source_mcp_key_name=getattr(current_user, "mcp_key_name", None), | ||
| ) | ||
| return StartLoopResponse( | ||
| loop_id=loop_row["id"], |
There was a problem hiding this comment.
(anchored to the nearest line in this PR's diff; the finding is at src/backend/routers/loops.py:1)
MINOR — doubled comment marker.
The provenance comment here (and the matching ones in routers/fan_out.py, routers/reminders.py and routers/schedules.py) starts # # #2389:.
| # review; the constant stays as the documented invariant and | ||
| # `test_2323_machine_identities.py` asserts the WRITTEN ROW rather than | ||
| # membership in this tuple. | ||
| _AGENTLESS_SCOPES = ("user", "portal_delegate", "ops") |
There was a problem hiding this comment.
MINOR, no action — _AGENTLESS_SCOPES is defined and referenced nowhere.
Deliberate per its own comment ("holds HERE BY CONSTRUCTION"). Noted only so the next reader does not mistake it for an enforced guard — a second minting path could write an agent_name for a non-agent scope with nothing to stop it.
…t-in form, and fence ghost keys off the event stream (#2389 re-review) Four findings from the second review round. 1. The "survives the minting admin being offboarded" property was documented in two places and implemented in neither: `assert_admin(..., allow_scopes={"ops"})` admits the scope and then still enforces `role == "admin"`, and `get_current_user` rejects a principal whose owner carries `suspended_at` (#995) one layer above. An operator would have acted on that sentence during a real offboarding. Corrected rather than implemented. Dropping the role check for an opted-in scope would make this bounded tier harder to revoke than the unbounded `user`-scoped key it exists to displace — that one loses admin the moment its owner is demoted — and it still could not deliver the claim, because suspension kills the key before any of this runs. #2323 asked for a credential that survives enforced 2FA, which it does; not one that survives its owner. The operator consequence (mint under a service admin account; revoke-and-re-mint belongs in the offboarding runbook) is now stated where it will be read, and a test pins the behaviour so the claim cannot drift back. 2. `require_admin` — the dominant admin-gate spelling — took no `allow_scopes`, so an ops-allowlisted route gated that way was permanently dead to ops keys with no opt-in available, and the new CI guard regex-matched `assert_admin(` only. Adds `require_admin_allowing("ops")`, which delegates the whole ladder to `assert_admin` rather than restating it (two admin gates that must stay identical is how the `require_role("admin")` third spelling happened), plus a sibling guard that reds on bare `require_admin` in any allowlisted handler. This closes the shape; the earlier fix closed only the instance. 3. `/ws/events` admitted the `agent` scope wholesale, including an ephemeral ghost's own key. That handler never runs `get_current_user`, so it skips `_enforce_ephemeral_key_fence` — whose entire purpose is that a ghost on an untrusted workspace must not hold a fleet skeleton key — and the stream is scoped by the owner's accessible agents, i.e. everything on a default admin-owned install. The gate now takes the key's `agent_name` and refuses an `is_ephemeral` row, using the fence's own predicate. That sub-check fails closed, deliberately inverting the ephemeral fence's fail-open: that fence guards heartbeats and result callbacks where a DB blip must not take the fleet down, while losing this stream costs an observability client a reconnect. Presence of the `is_ephemeral` key is what makes the answer real — the accessor coalesces the column for every live row, so a dict without it is no row at all, and a bare `.get()` would map that onto the same falsy value a genuine durable agent gives. 4. The Settings tab rendered an "Ops (read-only)" badge with no way to mint the tier — the create modal only ever sent `portal_delegate`. Replaced the boolean with a three-option scope selector (mutually exclusive, because the column holds one value and a checkbox pair can express a state the backend cannot store). The shared chrome is hoisted into two scoped classes so the file's raw-color ratchet count stays flat at its baseline of 97. Minor: the four doubled `# # #2389:` provenance markers are single again, and the now-inert `X-MCP-Key-*` sends in the MCP client carry a note saying why they are kept and that no backend reader may return. Full unit suite: 12392 passed / 21 failed, against 12383 / 21 at the previous commit — identical failure set (IPv6/SSRF, environment-dependent), +9 net new passing tests.
|
All four fixed in 1f4b3ea, plus the three minors. One of them I fixed by correcting the claim rather than implementing it — reasoning below, because it is a design decision and not a typo. 1 (Medium) — the offboarding property ✅ corrected, deliberately not implementedYou are right on the facts: I took the second of your two options. Dropping the role check for the opted-in scope was considered and refused for two reasons:
What the opt-in actually buys stands unchanged and is now stated as such: the grant is per route, so a new ops route is inaccessible until someone adds it. The opt-in is an additional gate, never a substitute one — an ops key is a narrowing of its owner, not a decoupling from them. The operator consequence is now written where it will be read (code comment, Worth noting the scope of what #2323 asked for: a credential that survives enforced 2FA — which it does, key validation never passing through the MFA gate. Surviving its owner was an embellishment I added, not a requirement. 2 (Medium) —
|
Code review — re-review at
|
Correction — my previous comment reviewed a stale treeApologies: my re-review above was produced against a checkout that predates Withdrawn — already fixed at
Still open, verified at
Sorry for the noise — the withdrawn four were real work already done, and the comment should not have implied otherwise. 🤖 Generated with Claude Code |
… the last gate-spelling gap, drop the dead constant (#2389 re-review) Three of the four still-open re-review items. The fourth is refuted below with evidence rather than fixed. **The flow doc still carried the retracted claim.** `dependencies.py` was corrected to say plainly that an ops key does NOT survive its owner's offboarding; `docs/memory/feature-flows/mcp-api-keys.md` still said it does, so code and doc contradicted each other and the doc is the surface someone reads while writing an offboarding runbook. It now states the same thing, with the same two reasons (demotion 403s at the gate, `suspended_at` kills the key one layer up per #995), why dropping the role check was refused, the 2FA-vs-owner distinction, and the mint-under-a-service-admin consequence. **One gate spelling was still unguarded.** `1f4b3ea` already added a sibling test for bare `Depends(require_admin)`, so the reported blind spot was mostly closed — but neither guard caught `require_admin_allowing("telemetry")`, i.e. the opt-in form parameterised with the wrong scope. All three spellings now go through one `_ops_gate_offences` predicate: assert_admin(user, allow_scopes={"ops"}) imperative Depends(require_admin_allowing("ops")) declarative Depends(require_admin) declarative, no opt-in exists The sibling test delegates to that predicate instead of carrying a second copy of the regex — two guards over one policy, each with its own pattern, is how the spellings diverged in the first place. `test_the_ops_gate_guard_detects_every_ spelling` pins each offending form AND each correct form, so the guard can neither go blind again nor be satisfied by rejecting everything. Verified by live mutation: reverting the `/api/subscriptions/{id}/usage` opt-in reds it by name. **`_AGENTLESS_SCOPES` is deleted, not left unreferenced.** Removing the guard that read it left a tuple no code consults, which reads as enforcement to the next person — the same defect one step smaller. The invariant it named holds by construction (this creator writes the literal `agent_name=None`, and `McpApiKeyCreate` declares no field for a caller to supply) and is stated in the comment plus asserted against the written row. **Not fixed, because it is not a defect: `HEAD` on an allowlisted route.** The premise is Starlette's `Route`, which does add `HEAD` alongside `GET`; FastAPI's `APIRoute` does not. Measured on this stack: `GET /api/version` -> 200, `HEAD /api/version` -> **405**, for every principal including an admin JWT. Routing rejects the method before dependencies are solved, so the ops fence is never reached and cannot be the thing refusing it. Adding `HEAD` to `_OPS_ALLOWED_ROUTES` would widen the method belt for a method the app does not serve — and that belt is what stops a future `POST /api/ops/*`, where `emergency-stop` lives.
|
Thanks for going back and checking by hand — the correction is more useful than the review.
1 — the flow doc carried the retracted claim ✅Right, and it was the worse half of the pair: the code comment was corrected and the doc is what someone reads while writing an offboarding runbook. 2 — the guard's blind spot ✅, though smaller than reported
What neither guard caught was The sibling test delegates to that predicate rather than keeping a second copy of the regex — two guards over one policy, each with its own pattern, is how the spellings diverged in the first place. 3 —
|
…ack under its own cap /sync-feature-flows over the last five dev commits (#2638, #2703, #2571, two left the feature's *home* doc stale while the change was documented elsewhere. fan-out.md (FANOUT-001 home, untouched since April) — #2670 only landed in mcp-orchestration.md, so the doc that owns routers/fan_out.py still said "POST only". Now documents GET /api/agents/{name}/fan-out/{fan_out_id}, build_fan_out_batch_status status derivation, the on_started hook, FanOutBatchTask/FanOutBatchStatus, get_fan_out_executions and its dual-scope reason, the bounded client.ts::fanOut() + get_fan_out_result MCP tool (receipt rationale summarized, pointer to mcp-orchestration.md). Drift repaired while there: db/schedules.py:NNN paths -> the #1481 db/schedules/ package, migration ordinal #30 -> #33 (verified against MIGRATIONS), tool registration via addAllTools, X-MCP-Key-* headers documented as inert per #2389. subscription-auto-switch.md — the #2638 prose was complete but its three catalog tables were not reconciled: Files (subscription_headroom_service, execution_envelope, client_portal/service, the #2638 test; the toggle lives in SubscriptionsPanel.vue + stores/subscriptions.js, not views/Settings.vue), System Setting (subscription_api_key_fallback), API Endpoints (GET/PUT /settings/api-key-fallback). feature-flows.md — 547 -> 444 lines. Recent Updates trimmed 137 -> 20 rows, matching its own "newest ~20" header (#1360); one-line rows added for #2638, #2703, #2670, which had none. 17 flow docs had no Documented Flows category row — subscription-auto-switch.md among them, reachable only via a June Recent Updates row the trim would have removed — so every one of the 190 flow docs now has a category row. Skill Injection row refreshed for delivery-on-assign (#2703). Three links that were already broken in HEAD (AUDIT-001-execution-origin-tracking, skills-crud, mcp-skill-tools) are left as-is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01168enK9QL4DNuVSN6taw2E
…ack under its own cap (#2753) /sync-feature-flows over the last five dev commits (#2638, #2703, #2571, two left the feature's *home* doc stale while the change was documented elsewhere. fan-out.md (FANOUT-001 home, untouched since April) — #2670 only landed in mcp-orchestration.md, so the doc that owns routers/fan_out.py still said "POST only". Now documents GET /api/agents/{name}/fan-out/{fan_out_id}, build_fan_out_batch_status status derivation, the on_started hook, FanOutBatchTask/FanOutBatchStatus, get_fan_out_executions and its dual-scope reason, the bounded client.ts::fanOut() + get_fan_out_result MCP tool (receipt rationale summarized, pointer to mcp-orchestration.md). Drift repaired while there: db/schedules.py:NNN paths -> the #1481 db/schedules/ package, migration ordinal #30 -> #33 (verified against MIGRATIONS), tool registration via addAllTools, X-MCP-Key-* headers documented as inert per #2389. subscription-auto-switch.md — the #2638 prose was complete but its three catalog tables were not reconciled: Files (subscription_headroom_service, execution_envelope, client_portal/service, the #2638 test; the toggle lives in SubscriptionsPanel.vue + stores/subscriptions.js, not views/Settings.vue), System Setting (subscription_api_key_fallback), API Endpoints (GET/PUT /settings/api-key-fallback). feature-flows.md — 547 -> 444 lines. Recent Updates trimmed 137 -> 20 rows, matching its own "newest ~20" header (#1360); one-line rows added for #2638, #2703, #2670, which had none. 17 flow docs had no Documented Flows category row — subscription-auto-switch.md among them, reachable only via a June Recent Updates row the trim would have removed — so every one of the 190 flow docs now has a category row. Skill Injection row refreshed for delivery-on-assign (#2703). Three links that were already broken in HEAD (AUDIT-001-execution-origin-tracking, skills-crud, mcp-skill-tools) are left as-is. Claude-Session: https://claude.ai/code/session_01168enK9QL4DNuVSN6taw2E Co-authored-by: sim <sim@example.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…2487) * refactor(settings): routers/settings.py becomes a ten-module package (#1028) 3,529 lines — the largest file in the backend, 4.4x the 800-line critical threshold — split into ten domain modules composed onto ONE router, so the mounted API is byte-identical and `from routers.settings import router` is unchanged. Largest resulting module: credentials.py at 778 lines. Inclusion order is load-bearing (Invariant #4): `generic` owns the GET/PUT/DELETE /{key} catch-alls, which match any single segment, so it is included LAST — before its siblings it would swallow /ops/config, /brain-orb, /api-keys/anthropic and answer 'setting not found' for routes that exist. test_1028_settings_package.py pins: - the mounted route SET equals the pre-split module's, compared against the real blob out of git (60/60, none lost, none invented) - no route is shadowed by an earlier registration (the property that actually matters — literal order is deliberately NOT pinned, since regrouping specific routes relative to each other is inert) - the catch-all include stays last, named at the include line a human edits - every module stays under the 800-line threshold - the import surface callers depend on still resolves (resolve_mcp_url, the key sets, _REPO_PATTERN) Collaborators (db, platform_audit_service, settings_service) are deliberately NOT re-exported on the package __init__: ~20 tests patch them as module attributes, and after a move such a patch would apply cleanly to a module nobody reads — a test asserting nothing while hitting the real accessor. Absent attributes make every stale patch raise AttributeError instead, which is exactly how the 14 affected test files were found and repointed to the modules that own their handlers. GET '' (the root listing) is registered on the parent router because a prefix-less sub-router cannot carry an empty path (FastAPI refuses). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLfHNPtB5UCMk4LonZiJux * refactor(git): services/git_service.py becomes a six-module package (#1028) 2,322 lines split by responsibility — conflicts, gitignore, remotes, trinity_files, sync, provisioning — with the full public surface re-exported from the package __init__, so `from services.git_service import sync_to_github` and `git_service.<name>` callers are unchanged. Largest resulting module: gitignore.py at 649 lines. Cross-module calls go THROUGH the sibling module object (`gitignore._detect_git_dir(...)`), never a from-import of the function: a from-import freezes the binding, so a test patching the owning module would silently stop reaching the caller. Pinned structurally by test_1028_git_service_package.py, alongside the size threshold and the import surface. Private names are re-exported ONLY where another backend module imports them or a test reads them as data. A private function mirrored on both the package and its owning module can be monkeypatched on the wrong one and silently detach — which is exactly what happened to test_2069's readiness probes mid-split (the multiline setattr sites patched the package's re-exported copies while merge_gitignore_after_clone read the module's own), so the collaborator-shaped names are deliberately not mirrored: a stale patch raises AttributeError instead of testing nothing. ~15 test files repointed to the modules that own their handlers, including the three sys.modules-isolated file loaders and test_github_init_push, whose exec fake must now land on every module binding the driven function awaits through (provisioning + gitignore + remotes — patched via the loaded package instance, since its harness purges and reloads the package). The #2069 merge-caller guard now walks the whole package and matches qualified calls, so a caller cannot fall out of its census by moving between modules. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLfHNPtB5UCMk4LonZiJux * refactor(client): services/agent_client.py becomes a three-module package (#1028) 1,294 lines split by responsibility — circuit (the #631 transport breaker: constants, Lua, CircuitState, dormant alerting, admin read/reset), http_pool (the per-agent httpx pool + drop-grace stamps), client (AgentClient, typed errors, get_agent_client) — public surface re-exported from the package __init__, so every existing import is unchanged. Largest module: client.py at 698 lines. Same discipline as the git_service split, pinned by test_1028_agent_client_package.py: cross-module calls go through the sibling module object, collaborators are not mirrored on the package, and no module may from-import a sibling's function (a frozen binding silently detaches monkeypatches on the owning module). test_circuit_breaker.py's direct file-load gains package plumbing (submodule_search_locations + a sys.modules registration before exec — the __init__'s relative imports cannot resolve their parent otherwise), and its patches land on the owning modules. The #1677 caller-parity allowlist entry for _emit_dormant_alert follows the file to services/agent_client/circuit.py — that guard firing on the move is exactly what it is for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLfHNPtB5UCMk4LonZiJux * refactor(ops,public): the heavyweight handlers move behind their routes (#1028) The last two ACs. `public_chat` — 289 lines of session identity, access gating, rate accounting, upload decoding, memory injection and dispatch, inside routers/public.py — moves to services/public_chat_service.py in the #1483 shape (service raises PublicChatError, the thin route maps it 1:1); client-IP extraction, the per-IP limit and token resolution stay router-side because they are HTTP concerns. agent_requires_email / agent_allows_open_access move with it and the router re-imports them — one definition, not a copy. public.py: 1,239 → 901 lines. routers/ops.py's five heavyweights — fleet health, the #1860 locked fleet restart, fleet stop, emergency stop, the cost rollup — move to services/fleet_ops_service.py (704) and services/ops_costs_service.py (208). The auth gates stay IN the router deliberately: the #2389 fence-vs-gate scans read live handler source there, and a gate that moved with the body would satisfy auth while blinding the scan. ops.py: 1,304 → 506 lines. test_1028_extracted_services.py pins the thinness, the gates' location, the size class — and an unresolved-module-scope-name walk, added because the move surfaced exactly that class twice: PublicChatResponse was unresolved in the chat service while 623 tests passed (nothing drives the sync-success return), and utc_now_iso the same in the costs service. py_compile cannot see this; the walk can. test_1860 / test_1917 fixtures now hand back the SERVICE module with the route entry points attached, so collaborator patches land on the bindings the moved bodies actually read while the gate patch stays on the router. The #894 override-wiring census follows public_chat's two execute_task call sites to their new file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLfHNPtB5UCMk4LonZiJux * test(1028): repoint the settings source-guards at the package; satisfy the sys.modules lint Four guards read routers/settings.py as SOURCE TEXT (the ent#12 consent AuditEventType pin + generic-PUT block, the ent#434 catch-all window) and went FileNotFoundError when the module became a package — repointed at the package glob (or generic.py where the guard scopes a specific handler window). The new test files' own sys.modules registrations move onto monkeypatch.setitem / the _restore_sys_modules precedent, and the lint baseline is regenerated DOWNWARD (140 across 49 files — the patch migrations in the split commits removed ~66 stale entries). Related to #1028 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLfHNPtB5UCMk4LonZiJux * test(1028): the size guard counts blank lines out, comments in Self-review finding, and the more serious of the two: my previous commit changed the guard's METRIC so that my own edit would pass. That is the re-baselining this file exists to make hard, wearing a docstring as cover. Measured rather than argued. `credentials.py`: at 6d8c93a raw 778 non-blank 687 after raw 806 non-blank 687 (blank-line restoration only) non-blank-non-comment 643 Excluding blanks is exactly invariant under the change that prompted it. Excluding comments as well moves the calibration: these files carry 110-208 comment lines each, so a comment-blind count hands `credentials.py` ~160 lines of headroom the 800 ceiling never gave it, and `generic.py` 208. So the metric now excludes blank separators and nothing else. Restoring a PEP-8 blank line between two defs still does not read as a module growing, and the ceiling still means what it meant when these files were authored against it. Related to #1028 * test(1028): name the size metric for what it counts Re-review of my own fix. The metric was corrected to exclude blank lines only, and left named `_logical_lines` — which in Python means the opposite, since a logical line excludes comments. That is not a cosmetic mismatch. Reading the phrase "logical lines" in this test's docstring is precisely what talked the previous pass into excluding comments and re-baselining the guard by 160 lines. Leaving the name in place leaves the same trap armed, one identifier along, for the next reader who "corrects" the body to match it. `_non_blank_lines`, and the AC's docstring says "800 non-blank lines" in its own words rather than delegating the definition to a helper name. Related to #1028 * fix(1028): make the split's own safety nets actually run Addresses the `/validate-pr` CHANGES_REQUESTED on #2487. Every item is about a guard that is present and inert, which is why the refactor landed green. **C1 — the 60/60 route-set proof never ran in CI.** It read the pre-split module out of git, and every checkout in `backend-unit-test.yml` is `fetch-depth: 1`, so `git show dd91056:…` failed on every run and the test skipped — leaving "no route lost or invented" across a 3,529 → 10 module split proven nowhere. The fork-point set is now a frozen 60-tuple literal (a fork point is a historical fact, so freezing it costs no maintenance; a route added since goes in `_ADDED_SINCE_SPLIT`, one reviewed line at a time). The git read survives as a separate test that re-derives the literal wherever history is deep enough, so the transcription cannot drift. Its temp module is written under `tmp_path`, not `src/backend/` — an interrupted run there left a top-level module `Dockerfile:131` would bake into the image. **C2 — the integration suite broke at collection**, in FOUR files, not three: `test_monitoring_service.py` too. All of them `spec_from_file_location` on `services/agent_client.py`, which is now a package, so they raised FileNotFoundError before any test ran; this only stayed green because integration runs nightly rather than per-PR. Replaced with plain imports: the loaders existed to bypass `services/__init__.py`, and that has not been true since the module started importing `services.agent_auth` at import time (it is on `dev` too). Privates come from the module that owns them (`circuit._CIRCUIT_HASH_PREFIX`, `http_pool._client_pool`), per the package's own no-mirrored-collaborators rule, and the caplog assertions key on the parent logger name so they still capture from every submodule. Collection is back to 83 = `dev`'s 83. **Two invariant guards went blind on 3,529 lines.** `routers/settings/` is the first subdirectory ever created under `routers/`, and both `test_1310_auth_wiring.py` (Invariant #8) and `test_models_centralized.py` (Invariant #14) globbed one level deep — all ten modules escaped, and it fails open, so nothing showed. `rglob`, keyed by path relative to `routers/` so two packages cannot share an allowlist key (a top-level file's relative path is its bare name, so neither allowlist changes). 73 → 84 files scanned. **A dead constant with a live test guarding it.** `routers/public.py` still declared `MAX_CHAT_MESSAGES_PER_IP`/`_PER_TOKEN` while enforcement reads `public_chat_service`'s copy, so `test_ip_rate_limit_fix.py` was asserting a constant nothing enforces — equal values today, so only the guard had broken. Now re-exported from the enforcing module. **Split-detachment in `public.py`.** The two #311 gates were from-imported under private aliases while the service called its own module-locals: one function, two monkeypatch targets. Now called through the sibling module object, which is the rule both package `__init__` docstrings state; the two tests that patch or read it are repointed, and the `files.py` guard now bans both spellings so the retired alias cannot let a re-import through. Docs: `architecture.md` Invariant #1 gains the package paragraph (re-export the public surface only; reach siblings through the module object; guards use `rglob`), and the four stale `.py` references in the shards are corrected. Also: nine modules carried a duplicated module-level `logger`; and `fleet_ops_service` renames the fleet-restart log channel from `routers.ops`, which is now stated in the code rather than left for an operator to discover. Verified: unit suite 6269 passed, 1 failure — the pre-existing `::ffff:` IPv4-mapped parsing case (local 3.12 vs the repo's 3.13 target), byte-identical on clean `dev`. Related to #1028 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bd71qsYbFodvofba8P69eP * refactor(1028): split gitignore.py three ways after the dev merge The dev merge brought #2529's ~715 lines into `services/git_service/gitignore.py`, taking it to 1305 raw lines — past the 800-line threshold this PR exists to enforce, and its own guard (`test_1028_git_service_package::test_every_module_is_under_the_critical_threshold`) said so. Re-baselining the guard under cover of a fix is precisely what the settings-test docstring in this PR warns against, so the module is split instead: gitignore.py 686 the patterns, the regions, the command builders gitignore_sweep.py 418 what a sweep DID — tags, parse, alert, report gitignore_clone.py 273 the once-per-agent merge after clone The seam is "what the file CONTAINS" vs "what did that just do" vs "the one-shot at creation". Cross-module references go through the module object (`gitignore.<name>`, `gitignore_sweep.<name>`), never a from-import: a from-import freezes the binding and a monkeypatch then lands on a detached copy — which is how test_2069's readiness probes went dark mid-split. Three real defects surfaced while wiring it and are fixed here, not carried: - `_gitignore_merge_semaphore` and `_inflight_gitignore_merge_tasks` were referenced bare in `gitignore_clone` with no such globals — a NameError on the live clone-time merge path. The five merge constants + the semaphore + the in-flight set now live in `gitignore_clone`, their sole consumer. - `_shadowed_negations` read `_GITIGNORE_MANAGED_LINES` bare after the move. It now reads it off `gitignore` through a deliberately function-local import — `gitignore` imports this module at its top level, so a module-level one would close the cycle at import time. - `datetime` was left behind by `_augment_commit_message`. The sibling suites are repointed at the module that OWNS each name, so every symbol still has exactly one monkeypatch target: test_2529's sweep names to `gitignore_sweep` (new `_sweep()` accessor beside `_gs()`), and test_2069's readiness/merge collaborators to `gitignore_clone`. 741 passed, 3 skipped across the git/gitignore/1028/1310/models families. Related to #1028 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bd71qsYbFodvofba8P69eP * test: dev's post-fork tests reach the split modules the way every earlier one does (#1028) Five files landed on dev after the fork with patch targets and source reads on the monoliths. Re-pointed the same way the split re-pointed the rest: - test_ent582_platform_keys: `is_claude_auth_configured` / `connect_agents_to_first_credential` on `settings.credentials`, `platform_keys_service.check_resend_key` on `settings.provider_keys` - test_2695_stt_capability_probe: `_elevenlabs_settings_state_with_capability` on `settings.integrations` - test_2691_public_url_reachability: the save-path source read on `settings/generic.py`, the flag-surface read on `settings/flags.py`, `update_setting`/`db`/`platform_audit_service` on `generic` - test_github_init_push: the exec recorder also installed on `git_service.token_scrub`, which the ent#615 seed now runs through - routers/settings/generic.py: the #2572 hook reaches `credentials` through an absolute function-local import — `test_2216_backup_observability` and `test_2572` load this module in isolation via `spec_from_file_location`, where a module-level relative import raises at collection Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VpvcfgWkmQPD7DrDLmATTf * test: dev's two new settings tests reach the split modules (#1028) `test_workspace_flag_retired` patches `settings_service`, `telemetry_sharing_service` and `db` on the module it calls `get_public_feature_flags` from, and `test_2696_stt_provider_errors` calls `_elevenlabs_settings_state_with_capability`. Both imported the flat `routers.settings`. Now they import the `flags` and `integrations` submodules, like the earlier re-points in 9c34662, so the patches land on the globals the handlers actually read. Full unit suite on this tree: 16468 passed, 32 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: patch the modules that own the moved globals, and let #1917 follow the moved ops code (#1028) merge-train validation findings. Collection was fixed earlier; these are the runtime half. - `tests/integration/test_circuit_breaker.py`: 19 `monkeypatch.setattr(agent_client, ...)` calls and 34 `agent_client.CIRCUIT_*` reads now target `services.agent_client.circuit`, which reads its own globals. Patching the package re-export changed nothing, and `_get_circuit_redis` is not re-exported, so it raised AttributeError. Against a fakeredis server: 8 failed / 26 passed before, 34 passed after (dev: 34 passed). - `tests/git_sync/test_s5_conflict_classifier.py` loads `git_service/conflicts.py`, because the flat `git_service.py` no longer exists. - `tests/git_sync/test_s7_reserve_instance_id.py` patches `check_remote_branch_exists` and `db` on `git_service.provisioning`, where `reserve_and_generate_instance_id` looks them up. Both git_sync files: 33 passed. - `tests/unit/test_1917_stack_trace_exposure.py`: the raw `str(e)` ban now also scans `services/fleet_ops_service.py` and `services/ops_costs_service.py`, where the ops handler bodies moved. Neither file has any hits today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: sim <sim@example.com>
The issue's premise is false, and the correction inverts its proposed shape
#2323 asks for a service credential that survives enforced 2FA, on the premise that none exists. One does — verified against a live instance with a
scope='user'MCP key owned byadmin:devrequire_admin/assert_adminreject only agent and connector; a user-scoped key inherits the owner's roleGET /api/versionis admin-bearerget_current_usergate_loginhas 2 call sites, both login routesThe credential exists; it is unbounded, unattributable and non-expiring. So the only way to enable 2FA and keep a dashboard alive was to hand it a permanent, invisible, unlimited admin key — worse than the control it works around. The issue's "acceptable smaller shape" (a tier that unlocks admin) buys nothing, because
useralready does. The tier needed bounds it.Full premise-correction and consumer evidence: comment, corrections.
Changes
1 — the admin gate becomes an allowlist. ent#293/#297 closed it against two scopes by name.
scopeis free-text with no CHECK constraint, so that is a denylist over an open space: anything setting neitheragent_namenorconnector_agentinherits the owner's role across ~163 sites.models.User's docstring predicted it; this PR'sopsscope is that sixth scope.ADMIN_GATE_SCOPES = {None, "user", "system"}is exactly what passes today — zero behaviour change exceptportal_delegate, previously contained only by its route fence, which gains a second layer. An absentmcp_scopefails closed via a sentinel, never aNonedefault that would make the missing attribute the privileged JWT value.2 — attribution from the credential, not a header.
validate_mcp_api_keyalways returned the key id/name;get_current_userdiscarded them. Deriving the threemcp_*audit columns fromactor_userfixes ~70 call sites with no diff at any, plusmcp_key_id/mcp_scopequery filters.actor_typestaysuser— the owner is accountable, is the only branch yielding an email, and the enterprise user-activity view matches on it.X-MCP-Key-Idis removed from the audit path, not out-ranked:Header(None)on six routes, validated nowhere, persisted into the backlog replay blob — so honouring it let any authenticated caller forge the credential named in the platform's two highest-volume audit events, surfacing minutes later on queue drain.3 —
ops: read-only, route-fenced, self-authorizing. Admin-minted and human-only. Fenced at the single auth entry point. Every entry is aGET, asserted by importing the constant — without that belt a prefix entry admits a futurePOST /api/ops/*, and that is whereemergency-stoplives. Kept out of the admin allowlist; ops reads opt in withallow_scopes={"ops"}, so authority comes from being an ops key rather than from the owner's role — it survives that admin being offboarded, and a new ops route is inaccessible until granted rather than silently reachable.The fence set is the measured read set of Trinity Control, not the issue's wording (which named only
/api/ops/*and would have shipped a credential unable to run the dashboard it exists for).routers/a2a.py::_a2a_idem_scopebuilds an idempotency scope frommcp_key_idand fell through tousernamebecause the field did not exist. Two agent-scoped keys of one owner therefore shared a peer-controlledmessageIdnamespace — caller B received caller A's full response text and B's task never ran. Its own docstring describes the failure it was not preventing. Reachable on an entitled install with ≥2a2a_exposedagents under one owner (default OFF, OSS unaffected).The test asserting the distinction used a stub carrying a field the real principal lacked — green over a property production never had. Re-pointed at a real
User.Deploy note: the scope string moves from
a2a:{agent}:{username}toa2a:{agent}:{key_id}, so amessageIdreplayed across the deploy re-executes instead of replaying. Bounded by the 24h TTL. Worth a release note.Test plan
pytest tests/unit/test_2323_machine_identities.py— 41 passed, 0 skippedpytest tests/unit/test_293_admin_gate_rejects_agent_keys.py— 25 passedThe route-parity guard initially skipped (importing the app needs the full stack). A guard that skips is not a guard, so it resolves against the router modules instead. It then caught two bugs in this PR's own code on its first real execution: a doubled router prefix, and a gate using direct attribute access where a principal might not carry the field.
Not delivered — stated, not implied
mcp_api_keysstill has noexpires_at. A leaked ops key is bounded in reach, not in time.userstays unbounded. This adds a tier; narrowing the existing one would break the fleet.Note on 44 test doubles
38 principal stand-ins lacked
mcp_scope; 6 stubs replaced the admin gate with a narrower signature. Every one claimed to stand for the real principal while not matching it — the same defect that hid both live bugs above. That is why the gate fails closed rather than defaulting, and the stand-ins were fixed rather than the gate softened.Fixes #2323
Generated with Claude Code