Release/1.23.0 - #1206
Merged
Merged
Release/1.23.0#1206
Conversation
Renders the `user_question_required` event PR-1 emits: an inline picker with single- and multi-select, a pager, an "Other" free-text field and Skip, which resumes the paused turn in place. Scope change from the original plan. PR-2 was meant to ship single-question single-select and leave the carousel to PR-3, but measuring PR-1 against real models showed Haiku 4.5 and Sonnet 4.6 both routinely ask three or four questions in one call. A single-question renderer would have been broken for the common case, not a smaller version of it — so the pager and multi-select are here. PR-3 keeps reload rehydration. Two defects found by looking at it in the running app, both invisible to unit tests and to typecheck: * Every dark-mode rule in the component was dead. Angular's emulated encapsulation stamps its `[_ngcontent-…]` attribute onto each compound selector *inside* `:where()`, so `:where(.dark, .dark *)` compiled to `.dark[_ngcontent-…]` — and <html class="dark"> carries no such attribute. That shipped a near-white selected row under near-white text. Switched to `:host-context(.dark)`; measured back at rgba(255,255,255,0.04). * Utility classes set directly on an `<ng-icon>` element do not apply — even `text-gray-700` and `size-5` were ignored — leaving the header glyph on the inherited near-white at 1.10 contrast in light mode. The colour now lives on the wrapper and the icon inherits currentColor: 9.63. Verified end to end against real Bedrock through the local stack: the model paused the turn, the picker took a single-select, a two-value multi-select and a free-text answer, and all three reached the model on resume. Backend carries one fix from the same session: headers were sliced at 12 chars mid-word, rendering "DASHBOARD PU" for "Dashboard Purpose". Now trimmed on a word boundary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…wnload-links fix: make generated-document download links durable
…r-question-prompt-ui feat: clarifying-questions picker in the chat transcript (PR-2)
The `user_question_required` event fires once and never re-streams, so a refresh mid-prompt orphaned the turn: picker gone, agent still paused waiting for an answer. `GET /messages` already replayed the breadcrumb PR-1 persists; this reads it back into the picker. Also fixes a pre-existing bug in the same area, found while building this and affecting OAuth and tool-approval prompts too. A breadcrumb outlives its `pausedTurn` snapshot when the user abandons a paused turn by simply typing something else — nothing cleared it, because the two `remove_pending_interrupts` call sites are resume cleanup and the explicit dismiss endpoint. Once rehydration exists that resurrects a prompt the user can no longer answer: the resume route 400s on an interrupt id the rebuilt agent never saw, so Submit becomes a guaranteed error. `clear_pending_interrupts` now runs beside `clear_paused_turn` at the head of every non-resume turn. It is a separate function rather than an addition to `clear_paused_turn` on purpose. That one clears only the snapshot — `test_paused_turn_independent_of_ pending_interrupts` pins it deliberately — and it also runs on the resume-success and expired-snapshot paths, where the narrower cleanup is already correct. The supersede policy belongs at the call site that owns it, next to `clear_interrupted_turn` and `clear_truncated_turn`. Question-shape validation is extracted to `validateUserQuestions`, shared by the SSE path and the reload path. They arrive by different transports — a parsed SSE frame and a JSON string out of DynamoDB — but must agree on what is renderable; two copies would drift, and the failure mode is a prompt that renders on one path and vanishes on the other. Verified against the running stack, both halves: refreshed mid-prompt and the picker came back with its questions, pager and Other field intact, then answering it resumed the turn through the snapshot-rebuild path with the selection reaching the model. Separately, abandoned a prompt by typing something else, refreshed, and got no stale picker. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…r-question-rehydration feat: rehydrate the clarifying-questions picker after a refresh (PR-3)
…1104) Arming CDK_MANAGED_KB_NEW_DEFAULT in dev on 2026-09-14 made the handoff's flag table factually wrong, which is the one kind of staleness that actively misleads. Corrected, plus the closeout the specs were owed. - Flag table: dev NEW_DEFAULT is now `true`, with the point that matters — until this date born-managed had been merged, deployed and NEVER ONCE EXECUTED in any environment. Every test was a unit test against fakes, and a fake cannot tell you that AWS accepts CreateKnowledgeBase or that a work key the trigger writes is one the dispatcher's index actually returns. Also notes that dev now runs rungs 2 AND 3 at once, so an agent can reach managed by either path — needed when reading logs. - Records that setting the GitHub variable alone does nothing: it is threaded through platform.yml into CDK, so it takes effect only on the next Platform Stack deploy. Flipping the variable without deploying looks exactly like a broken feature. - "Do these first" reordered around reality: observing born-managed run in dev is now step 1, and it largely SATISFIES task 15.1, whose live create → ingest → retrieve is precisely born-managed's happy path. Record the result against 15.1 rather than running a separate probe. Also flags the 15-minute first-document latency as the thing most likely to be misread as a hang. - NEW: a "known open defects, none blocking" table, so three findings stop living in chat. The ingestion deleting-guard gap, with the correction that RETRIEVAL is not the hole (the status filter serves only `complete` and fails closed — the bug is that ingestion promotes the row back to a state the filter accepts) and that it is not reachable from a single-tab device upload but plainly is from an import or a crawl. Whether a large upload can be cancelled at all. And the chunking-strategy finding: managed KBs accept DEFAULT/FIXED_SIZE/NONE, semantic is refused, hierarchical is not offered, it is immutable after data-source creation, and we send nothing — which is why one image's description splits across chunks. - Status lines closed out on all three specs. born-managed: BUILT AND MERGED, armed in dev, dark in prod. kb-chunk-inspector requirements + design: BUILT, live in dev, task 6 outstanding, with the Requirement 2 amendment (managed-only) pointed at from the design doc so a reader of either lands on it. No archive directory was created: this repo keeps specs in place and uses tasks.md checkboxes as state, and inventing a convention for one feature would be worse than the mild untidiness. The superseded new-default-wiring.md already carries its banner. Documentation only. No code, no behaviour change.
The picker was first modelled on ToolApprovalPromptComponent, and inherited a visual language built for a different job. That component is a narrow pill with a hard 2px accent bar and hairline dividers, which is right for a two-button yes/no that should stay out of the way. This one is a form the reader has to think about, and at max-w-xl inside a `justify-start` wrapper it read as cramped and boxy. - Full width of the assistant column: the `flex justify-start` wrapper and `max-w-xl` are both gone. - Soft ground instead of chrome: rounded-2xl card, no border — a faint tint and a soft ring carry the edge. The left accent bar is dropped. - Options are discrete rounded rows with air between them rather than an edge-to-edge divided list, with roomier hit areas and larger type (text-sm/6 labels, text-xs/5 descriptions). - Rounder controls throughout; "Other" now shares the option row's shape so it reads as one more choice rather than a stray input. - The eyebrow loses its icon. It was pure decoration, and a little glyph on an AI prompt is the first thing that makes it look generated. Measured in the browser, light and dark: eyebrow 10.23, option label 17.75, description 7.56, pager/skip 7.30, selected-row label 14.31. Selection reads from the tinted surface and the filled marker together, never colour alone. `background: white` tripped the surface-literal guard; switched to var(--color-white), which is what that guard exists to enforce. Behaviour is untouched and re-verified end to end: single-select, a two-value multi-select and a free-text answer all reached the model on resume. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things were stacking a second outline around a selected row. A focus-outline fallback written for the "Other" row — a <label> whose focusable element is the input inside it — was scoped to every `.option`. On a plain option button `:focus-within` is also true after a MOUSE click, so clicking drew an outline outside the row. Now scoped to `.option--other`. And selection itself was adding a coloured ring on top of the row's existing hairline. Selection is now a fill change: every row carries exactly one hairline whatever its state, and a warm tint plus the filled marker carry the state together, so it is still never colour alone. Keyboard focus is unaffected and still draws a visible ring — the stroke that was removed is the redundant one, not the accessible one. Measured on the new tinted surface: label 15.63 light / 14.28 dark, description 6.66 / 10.13. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on-and-nightly-allowlist-guard fix(security): sanitize the remaining log-injection sinks, guard the nightly allowlist
…r-question-ui-polish style: full-width, softer clarifying-questions picker
…ost-drilldown-ui feat(admin-costs): conversations on the user page + session profile, diagnoses and trajectory
The tool worked end to end from PR-2, but the model rarely reached for it — 4/24 on deliberately ambiguous requests against a production-shaped tool set. A short clause in the system prompt takes that to 24/24, and leaves clear requests at 0/18 so it does not turn direct questions into interrogations. The clause is appended only when `ask_user_question` is in the turn's effective tool list, so a user without the tool never carries an instruction to call it. The check is on the POST-FILTER list rather than the request's `enabled_tools`: the two diverge — `ToolFilter` drops a catalog id the registry does not know, as `canvas_faculty` does in dev today — and keying on the request would advertise a tool absent from `toolConfig`. It also picks up the `ASK_USER_QUESTION_ENABLED` kill switch for free. It is applied to the prompt handed to the agent, never to `self.system_prompt`, which is snapshotted for resume and hashed into the agent cache key. The clause derives from `enabled_tools`, which that key already covers via `tools_hash`, so mutating the field would move resume onto a different cache slot for no benefit. `PrefixFingerprintHook` reads the prompt off the built agent, so `systemPromptHash` still reflects what was really sent. Three things the measurements ruled out, recorded so they are not retried: rewording the tool description (44-56%, within a baseline band of 17-44%), changing the tool's position in the list (25-38%), and removing the prompt's "Cost Awareness" clause (38%). Only the system-prompt clause escapes the noise. The text is therefore load-bearing in this position, and ships byte-identical to what was measured, with a test pinning that. Catalog seed flips to enabledByDefault: True. Cost: ~63 tokens, constant per configuration, inside the cacheable prefix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two user requests are blocked on the same gap: a Library agent that evaluates the accessibility of subscription databases it renews annually, and a VPAT agent that must test vendor demo tenants. Neither target is reachable by a logged-out browser, and there is no path today for a user to authenticate a browsing session. Specs the human-in-the-loop takeover flow: the agent calls take_control() (UpdateBrowserStream -> automation DISABLED, a service-side mutex, not a convention), raises an interrupt, the user signs in themselves in an embedded AWS DCV live view, and the turn resumes authenticated. Credentials never enter the prompt, the conversation, or AgentCore Memory. Nearly all of the capability already exists and is unused: take_control / release_control ship in the pinned bedrock-agentcore 1.21.0, and the Runtime role is already granted UpdateBrowserStream and ConnectBrowserLiveViewStream. Also covers browser profiles (so a login is once-per-vendor, not once-per-conversation), a second VPC-mode browser resource with session recording, and axe-core as an extension. Two findings worth flagging independently of whether this ships: - The existing `live_view` action is broken in the way PR #1101 diagnosed. generate_live_view_url signs with SigV4QueryAuth, so the signature is in the query string, and browse_tool.py:288 returns it as text in a tool result -- the model will re-emit it truncated at the `?`. Max expiry is 300s, which is also far too short for a human login. - RBAC granularity is exactly one tool_id, so takeover must be its own registered tool rather than a browse_web action (D1). As an action it would ship to everyone who can browse, which is backwards: students should browse and should not be able to drive a browser inside our AWS account. As separate tools, an ungranted user pays zero prefix tokens for them. Measured prefix cost: browse_web 493 tokens today, request_user_login ~206, accessibility_scan ~191. A full human login round trip costs about one short tool result. The real spend for these agents is screenshots and accumulated page text, which is why D9 requires a 40-target sweep to be 40 sessions rather than one turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…r-question-trigger-guidance feat: make the clarifying-questions tool actually fire (PR-4)
Fold KB storage usage into the documents-list response so the agent's knowledge-base card can show how much of the byte cap is in use. Backend: - Add KbUsage model (engine, storedBytes, reservedBytes, cap, elevated) to DocumentsListResponse. - _resolve_kb_usage reads the KB_Record once: managed KBs report their bytes and the binding effective_cap (min of owner tier and per-KB ceiling); legacy S3-Vectors KBs are uncapped (cap=null). Best-effort so a record-read failure never breaks the documents list. Frontend: - Usage bar on the KB card: 'X of Y used' for managed KBs, green/yellow/red at <75/75-90/>=90% of the cap; legacy KBs show 'X stored', always green, no denominator. Tests: backend route tests for managed/legacy/failure paths; frontend component specs for the thresholds and the uncapped legacy case.
feat(kb): storage usage bar with byte cap on the KB card
…gent (#1109) Born-managed treated the absence of a KB_Record as 'brand-new agent' and provisioned a managed KB on the first upload. But legacy KBs are not first-class: they share one S3-Vectors index and never write a KB_Record, so an established legacy agent looks identical to a new one. Its NEXT upload was therefore mistaken for a first upload, flipping retrieval to an empty managed KB and stranding the existing corpus on the legacy index. Guard the record-is-None branch on an existing-documents check (assistant_has_documents: cheap COUNT, Limit=1, no ownership check). Only provision when the agent has zero documents; fail toward legacy on any probe error. Adds mutation-guard + helper unit tests (31 pass).
Byte tracking is scoped to managed KBs (Req 12.11), so a legacy KB reports zeroed counters and cap=null. The usage bar therefore rendered '0 B stored' beside real documents on every Classic agent, which reads as a bug. Gate showUsageBar on the managed engine (kbUsage.engine === 'managed') instead of merely 'kbUsage is present'. Managed KBs are unaffected; legacy KBs show only per-document sizes. Adds a spec asserting the bar (and its progressbar element) is absent for a legacy KB in edit mode.
Run the backend suite with pytest-xdist -n auto on the PR gate so the ~3k tests fan across all runner cores instead of running single-threaded. Also drop -v from backend/pytest.ini (it produced thousands of PASSED lines with no diagnostic value). The nightly coverage run (scripts/backend/test.sh) is left serial on purpose.
Make ts-jest transpile-only (isolatedModules) so jest workers stop re-type-checking the whole project each — the dominant cost of the infra suite (the long pole once backend went parallel in #1111). Add a single 'tsc --noEmit' (npm run build) step to the infra CI job so type safety is preserved on PRs (previously only ts-jest enforced it; tsc ran only in teardown.yml). Workers stay at 2 (the --maxWorkers bump regressed 2.5x in #1112). No const enum in the tree, so isolatedModules is safe. Verified: tsc clean; 178 tests green transpile-only.
…icated-web-assessment-spec docs(spec): authenticated web assessment via browser takeover
… one turn
Mentioning an Agent ran that turn as the Agent and silently reverted the next
one. The thread still looked like the Agent's while its tools, skills and model
were gone, and nothing surfaced the change — not the UI, and not the model,
which cannot know its own toolset shrank. Asked to use a tool it had used a
moment earlier it got `Unknown tool: create_rubric`, and told the user to toggle
that tool in the picker: a confident wrong diagnosis sending them to fix a
setting that was already correct.
D11 chose the per-turn reading deliberately, so this reverses a decision rather
than repairing an oversight. What changed is the evidence. Measured on prod
sessions-metadata (`turnAgentId` on the `C#` rows): of **247 mentions, 247
started the conversation**. Zero were mid-thread consults; zero mentioned a
second Agent inside a thread bound to a first. (Dev: 60 of 61.) The borrow was
paying an invisible failure mode for a case that has never occurred.
A mention now *means* "talk to this Agent", with two outcomes and no third:
* empty thread -> the Agent binds the conversation, exactly like launching
it from its card. Safe because there is no history for the
binding to misrepresent.
* has messages -> the message opens a NEW conversation with that Agent, and
the SPA says so. The Agent cannot be bound to history
written under other instructions, and must not be borrowed.
Persisting `preferences.assistant_id` is necessary but NOT sufficient, which is
the part that is easy to miss: every turn's Agent is resolved from the request,
the SPA's only carrier is the `assistantId` query param, and the self-heal
effect that refills it from preferences runs on session *load*. So the SPA now
sets the param and stops sending `agent_mention` at all. The backend still
honours that flag for clients that predate this change, and `binds_conversation`
gains `thread_is_empty` so a stale tab mentioning into a fresh thread lands
where a current one does. Its thread lookup runs only for mention turns, so no
bound-Agent turn pays a query it cannot act on.
Also fixed, because it is the same invisible loss on the path users are *told*
to use: "Continue" after a max_tokens truncation skipped the whole assistant
block, so a properly launched Agent finished its reply with none of its tools,
skills, model or instructions (spec-acknowledged as a known edge). The SPA was
already resending `rag_assistant_id` there — `continueTruncatedTurn`'s own
comment says "so the backend rebuilds the same model/tools/assistant agent" —
and only the `not is_continuation` guard discarded it. The block now runs for a
continuation, with binding validation and persistence skipped (it binds nothing
new) and RAG skipped (the turn carries an empty message, so a KB search would
spend a query on "" and augment nothing).
A resume still skips the block and always did keep its tools: it rebuilds from
`PausedTurnSnapshot`, replaying the original turn's exact enabled_tools /
system_prompt / enabled_skills to reconstruct the same prompt-cache key.
Re-resolving there would risk a different effective set and orphan the paused
agent. Worth stating because resume rows carry no `turnAgentId`, so a census of
"turns with no Agent" reads them as losses and overcounts badly.
Two known costs retire with the borrow: the ~$0.12-per-mention prefix re-write
(a bound conversation swaps once and stays instead of swapping back), and the
history fork, where the mention agent and the plain agent were two cached
instances that never saw each other's turns.
The rule lives in one testable place at each end — `mention-routing.ts` on the
client, `agent_binding_policy.py` on the server — following the existing
`system_prompt_resolver` precedent: the rule is a handful of lines, the code
around it is a thousand, and a rule no test can reach is a rule that drifts.
Kept as one commit: the mention and continuation halves edit the same guards on
the same block, and splitting them would mean a first commit that knowingly
leaves the block wrong.
Backend 8585 passed / 3 skipped; frontend 3024 passed across 250 files; tsc
clean. Not exercised in a browser — this worktree's code is not what the local
stack serves, and that stack is down.
Spec: docs/specs/agent-marketplace.md (D11 + Phase 7 notes)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swaps heroPaperAirplaneSolid for heroArrowUpSolid on the chat composer's submit button. Button geometry, colors, the stop-icon branch while streaming, and aria-labels are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…poser-up-arrow-icon feat(chat): use up arrow instead of paper airplane for send button
…ention-binds-conversation fix(agents): an @-mention binds the conversation instead of borrowing one turn
The greeting heading sits in a 616px text column (720px column, 1rem side
padding, 56px logo, 1rem gap) at text-4xl/tight. Measured against that
column, 19 of 51 greetings wrapped to a second line — and because
AnimatedTextComponent types the greeting out a character at a time, the
wrap happens in full view and pushes the composer down mid-animation.
Rewrite the 20 offenders shorter, keeping the voice. Two of them are stock
DEFAULT_GREETING_TEMPLATES entries that brand.defaults.golden.spec.ts
pinned verbatim: "How can I help you today, {name}?" wrapped for any first
name of 8 characters or more, so the pin was preserving a bug. Update the
pin and say why. The worked example config and the rebranding README get
the same budget.
Add greeting-line-length.spec.ts to hold the line. jsdom has no font
metrics, so it sums per-character advance widths captured from the real
InterVariable woff2 in the app's own <h1>; summed advances track browser
layout to within ±7px across 455 name/greeting combinations, which the 8px
tolerance covers. Verified in a browser that all 57 committed greetings
hold one line with a 12-character first name substituted for {name}.
Narrow viewports are deliberately out of scope: below ~720px the column is
the viewport, and no greeting worth writing fits a phone on one line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…g-one-line fix(branding): keep every greeting on one line in the chat empty state
…evelop-1.22.0 # Conflicts: # backend/src/apis/app_api/documents/routes.py
…into-develop-1.22.0 Backmerge: main into develop (1.22.0)
The admin skill form capped `description` at 500 characters, rejecting valid Anthropic-authored skills on import. The Agent Skills spec allows 1,024 — so the cap was roughly half the spec and blocked real skills (`math-olympiad` at 704, `hook-development` at 525), including two of this repo's own `.claude/skills/`. The field was also enforced at three different numbers for one column: 500 on the admin path, 2,000 on the user-authored path, and no cap at all on the `Skill` record itself. Replace all of them with a single spec-derived constant. A bound is still warranted — `description` is the Level-1 catalog line injected into the cacheable system prompt for every enabled skill, on every turn — so this sets the right number rather than removing the cap. - Add `SKILL_DESCRIPTION_MAX_LENGTH = 1024` to `apis/shared/skills/models` - Apply it to SkillCreate/Update (was 500) and CreateMy/UpdateMySkill (was 2,000) - Mirror it client-side in `shared/skills/skill-field-limits.ts`, following the `skill-resource-types.ts` precedent, and use it in both skill forms - Interpolate the constant into the admin form's error text so the message cannot drift from the validator again Note: lowering the user-authored cap from 2,000 to 1,024 affects only writes. The `Skill` model has no read-side cap, so an existing longer record still loads; it would fail only on a PUT that resends the description. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ew (#1189) #1186 removed a real duplicate-connect bug but the viewer still never streamed: one `dcv.authenticate`, still failing with code 10 behind a socket that closed immediately. Measured against real AgentCore, one probe at a time: correct single query -> socket OPEN, held DOUBLED query -> HTTP 403 That is what we were sending. The SDK APPENDS `httpExtraSearchParams` to whatever URL it is given, and we handed `authenticate` the full signed URL with its query still attached AND the extras callback — so every SigV4 parameter went twice and the service refused it. `connect()` has always stripped the query for exactly this reason; its comment even says "the transport appends to whatever URL it is given". `authenticate()` was simply missed, and the failure mode (a socket that opens then closes, surfacing as "Failed to communicate with server") reads like a service problem rather than a malformed request. Also ruled out by probe, so the next person does not re-walk them: base-path SigV4 signing IS valid for the `/auth` sub-path (signing `/auth` itself 403s), the `dcv` subprotocol is requested correctly by the SDK, an `Origin` header changes nothing, and `take_control` is not implicated. Mutation-checked: restoring the signed URL fails the new assertion. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ovable (#1184) * perf(inference-api): decompose the preamble stage into five sub-marks `agent-state-feedback.md` PR-3 measured the pre-stream window into four stages, fixed the one that dominated a cold turn (`agent_build`), and left `preamble` closed. On a warm turn `preamble` is the largest stage in that table — 460-494ms of a 641-678ms handler total — and it is a single mark spanning ~500 lines, so nothing says which of its stages owns the time. Split it into `preamble.ownership` / `.skills` / `.files` / `.session_state` / `.quota`, and teach `TurnPrelude.emit` to sum dotted stages back into a `groups` field. Without that sum, decomposing the stage would discard the only measurement anyone has of the window — the four-turn dev baseline is stated in terms of the single coarse number, and nothing else in the log reproduces it. `groups.preamble` now does. Measurement only: no behavior changes, and the code reading suggests but does not prove where the time goes. The accompanying spec records the hypothesis (eight serialized reads of the same session row, five of them in `session_state`) *and* what would falsify it, so the next PR is decided by the log rather than by the reading. Nothing reaches the model, the conversation, or the cacheable prefix — four extra `perf_counter()` calls and one log field per turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * perf(observability): make a preamble fix provable, not just attributable PR-1 makes the preamble attributable. It does not make a fix provable, and those need different instruments: attribution needs one turn, proof needs a population. Three parts. **The field that would have lied to us.** `time_to_first_token` is the obvious instrument and the wrong one: it measures from the coordinator's `stream_start_time`, which since agent-state-feedback PR-3 runs after the preamble AND after the agent build. It is blind to the app-api hop, the whole preamble and the whole build — so a successful PR-2 moves it by exactly zero. Documented in the spec so nobody validates with it. Same trap this repo already paid for once with `latency.endToEndLatency` vs `turnDurationMs`. **Percentiles instead of grepped log lines.** `TurnPrelude.emit` now also writes one EMF record per turn into `AgentCoreStack/TurnLatency` through the existing helper — no new infrastructure, no new IAM. Metric names are derived from stage names rather than listed, so a new mark cannot silently go unmeasured. Dimension-less like every other EMF caller here; `isResume` and `deferredBuild` ride as log properties and Logs Insights does the slicing a dimension would have done badly. **Widgets on the EXISTING AgentCore dashboard, not a fourth one.** A dedicated construct was built and then removed: `observability-platform-dashboard.test.ts` pins the stack at three dashboards because CloudWatch charges $3/month beyond that. The ceiling is a deliberate cost decision with a test guarding it, and it is not worth breaking as a side effect of adding widgets. Folding in is the better design anyway — that dashboard already graphs AWS's own data-plane `Latency`, so the gap between it and `PreludeTotalMs` is an independent read on the routing overhead no server-side stage can see. Two of the widgets exist to catch our own errors rather than the platform's: a split by turn shape (a shift toward resumes would look exactly like a latency win) and an unaccounted-time panel (which says when the decomposition is incomplete). No alarms: a threshold needs a baseline and there is none yet. Inventing one is the guessing the spec exists to prevent. Kill switch `TURN_LATENCY_METRICS_ENABLED` (default on) — unlike PR-1's marks, which stay ungated, an EMF line costs log ingestion and custom-metric charges. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * perf(observability): give turn latency its own dashboard, and buy the fourth Reverses the placement decision in the previous commit, not its substance. PR-1b put the pre-stream stage widgets on the AgentCore Runtime dashboard to stay inside CloudWatch's free three. That version worked and was still worse: this board is read *while shipping a latency change*, side by side with load-test output, and burying the stage breakdown under runtime-health widgets answering an unrelated question made it harder to use for its one job. So the fourth dashboard is bought deliberately at $3/month, against a 450-900ms wait on every turn. The pinned count in `observability-platform-dashboard.test.ts` moves to four and now reads "three free + one bought", with the argument recorded next to it — the point of pinning the number was never that three is sacred, it was that crossing the tier should be a decision somebody made rather than a side effect of adding widgets. The next one should have to argue the same way. The AgentCore Runtime board stays the companion: it graphs AWS's own data-plane `Latency`, so the gap between that and `PreludeTotalMs` is the routing overhead no server-side stage can see. Both headers point at each other and the platform dashboard links to all three drill-downs. `inference-agentcore-construct.ts` is byte-identical to its pre-PR-1b state. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ket (#1190) * fix(browser): sign the live-view stream socket, not just the auth socket `connect` reads `httpExtraSearchParams` ONLY from `observers`. We passed it at the top level, where it is silently ignored — so the stream socket opened with no query string at all, unsigned, and the service refused it. Measured by running the shipped 1.13.2 bundle against a stubbed WebSocket: top-level -> wss://.../live-view/ws (UNSIGNED) observers -> wss://.../live-view/ws?X-Amz-Algorithm= (signed) That is exactly what the browser console showed: repeated failures to `/live-view/ws` carrying no parameters. `authenticate` is the opposite — it reads the callback from the top level — which is what made the top-level form look right. The two entry points genuinely differ. `firstFrame`/`disconnect` move into `observers` with it, because the SDK honours observers or callbacks and not both. This is also what AWS's own BrowserLiveView component does; its config was read from the published package rather than guessed. Mutation-checked: restoring the top-level form fails the new assertion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: remove a scratch probe file committed by mistake `.connprobe.mjs` was a throwaway jsdom harness used to capture the DCV SDK's WebSocket URL. Its cleanup `rm` sat in the same command line as a node process that never exited, so `git add -A` swept it into the branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Measured on dev after #1184 shipped the instrumentation (docs/specs/turn-latency-preamble.md): a warm turn spends ~455ms in the preamble, and ~445ms of that is DynamoDB reading ONE item eight times. The per-read cost is the finding. `preamble.ownership` is exactly one GSI query and nothing else, and it measures 53-55ms — dead stable across turns. `preamble.files` on a turn with no attachments is also exactly one read, also 53ms. So a GSI query from inside an AgentCore Runtime container costs ~53ms, not the ~12ms an in-region figure suggests, and the spec's earlier estimate that this work would "recover roughly a quarter of the stage" was low by 4x. `load_session_meta()` performs one `SessionLookupIndex` query and returns a `SessionMetaSnapshot` carrying both answers the preamble needs: this user's META row, and whether the session belongs to someone else. Both were always in that one response — `_get_session_by_gsi` discarded the half the ownership probe needed, which is why the probe was a second query of the same index for the same key. The route reads once and threads the snapshot into `pop_pending_attachments`, `ensure_session_metadata_exists` and the four `clear_*` helpers. Every one takes `snapshot=None` by default, so non-preamble callers are untouched. Two decisions worth naming: **Explicit, not a per-request memo.** Memoising inside `_get_session_by_gsi` would be a three-line diff and the wrong shape. CLAUDE.md's "never cache session state" rule has been paid for twice (#741, #751); a snapshot a caller opts into cannot leak into a caller that needs a fresh read, and a memo can. **The snapshot object, not the row dict.** `row is None` is a real answer — "no META row yet" — not "nothing prefetched". Taking the row alone would make a brand-new session indistinguishable from an absent prefetch, so the first turn of every conversation would silently fall back to re-reading. Pinned by test. Safety: the helpers were already best-effort and fail-open; the GSI is eventually consistent so re-reading never bought consistency; the single-flight lease excludes a concurrent turn on the same session; the SK is static so the conditional writes' key cannot drift. The atomic parts stay atomic — the `ReturnValues=UPDATED_OLD` writes are untouched, and the snapshot only replaces the gate read. Expected: preamble ~455ms -> ~125ms. To be validated on dev, not assumed. The quota session-notice read (62ms) is deliberately left for PR-2b: it goes through `get_session_metadata`, which lazily backfills cost aggregates, and skipping that would quietly stop session notices on legacy rows. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…coder (#1192) The live view died with "Display channel is not available" and a black frame. The browser said why: Connecting to 'data:application/octet-stream;base64,AGFzbQ...' violates the following Content Security Policy directive: "connect-src 'self' https://bedrock-agentcore.us-west-2.amazonaws.com ..." `AGFzbQ` is the WebAssembly magic number. The DCV Web Client SDK fetches its decoder from an inline `data:` URI, and `connect-src` was the ONLY directive here that did not allow `data:`/`blob:` — script-, style-, img-, font- and media-src all already do. This does not widen exfiltration. `data:` and `blob:` name bytes inside the document, not a network peer, so nothing can be sent anywhere through them; the remote origins a page may reach are still exactly `connectDomains`. The existing tests pinned the old directive string and failed on this change, which is what they are for — updated, keeping each one's intent. The malformed-input case asserted "nothing follows 'self'" as a proxy for "no injected domains"; it now pins the directive exactly, which tests the same property without depending on the constants being empty. Mutation-checked: reverting the directive fails 10 of 34. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…read (#1193) PR-2b of docs/specs/turn-latency-preamble.md, and the eighth and last read of the session META row in the preamble — a measured 62ms. `check_quota` / `_resolve_session_notice` take `session_total_cost: Optional[float] = None`. The inference-api route supplies it from the snapshot PR-2 introduced, but ONLY when that row actually carries `totalCost`. `None` means "not known", never "zero", and that distinction is the whole safety of this path. `get_session_metadata` lazily backfills the cost aggregates for legacy rows written before write-time aggregation; treating an absent attribute as 0.0 would skip the backfill and silence the session notice on exactly the long-lived conversations it exists to catch — a latency fix quietly becoming a feature regression. A genuinely free session reports 0.0, which is falsy, so the check is `is None` rather than truthiness: a truthy test would send that session back to the read this PR removes and the bug would be invisible, because the answer would still be correct — just expensive. app-api's converse route passes nothing and keeps its read unchanged. Expected: preamble.quota 62ms -> ~0, warm preamble ~163ms -> ~100ms. To be validated on dev. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#1194) This is why the sign-in viewer never streamed. AWS's own service reference lists the resource types for each action: GetBrowserSession -> ['browser', 'browser-custom'] UpdateBrowserStream -> ['browser', 'browser-custom'] ConnectBrowserLiveViewStream -> [] <- none An action with NO resource types never matches a resource-scoped statement, so granting all three on the browser ARN made the third an implicit deny. `iam simulate-principal-policy` confirms it: allowed for the first two and implicitDeny for the third, from the SAME statement on the SAME ARN. It failed silently and late. `generate_live_view_url` only signs locally and calls no API, so app-api minted a URL happily; the denial appeared only when the browser opened the socket and the service closed it, surfacing as DCV auth code 10, "Failed to communicate with server" — which reads as a service fault, not a missing permission. Proven by running the real SDK against a real session outside the browser: `dcv.authenticate` SUCCEEDS with admin credentials (returns sessionId and authToken, socket closes 1000) and fails from app-api. Same SDK, same service, same session — only the signing principal differs. `*` is as narrow as this action can be expressed. The statement is split so it carries that one action and nothing else, and the resource-typed actions stay scoped to the browser ARN. app-api still cannot start, stop or drive a browser. Mutation-checked: re-scoping the connect grant fails the new assertion. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… slot Two problems with where the recap landed, both from it being rendered inside the turn rather than in the tail. It appeared on every finished turn. The number describes what just happened, and a column of durations down the length of a conversation turns a punctuation mark into a metrics readout — every earlier turn's recap is a fact nobody asked for, competing with the answer it sits under. It is now derived once, for the last turn only. It also overlapped the live line it was meant to replace. The old guard was `streamingMessageId`, which clears at message_stop; the loader lives on `isChatLoading`, which clears later at stream close. That gap is exactly the window where both were on screen, so the finished state read as a second, different thing appearing beneath the running one rather than as the running one settling. The recap is now the loading line's `@else`, in the loader's own slot with the same padding, gated on the same signal — one slot, never two, and nothing moves when they swap. Verified across two turns in one session: 1 loader / 0 recaps while running, 0 loaders / 1 recap when done, and the first turn's mark gone the moment the second turn starts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…replaces-live-line fix(ui): show the turn recap only on the latest turn, in the loader's slot
…y served The SPA was synced to S3 with no Cache-Control metadata on any object, so CloudFront served index.html with no Cache-Control at all. With no explicit freshness directive a browser falls back to HEURISTIC caching — roughly 10% of the document's age when it was cached — which for a shell that had been sitting in the bucket for days means hours of reuse with no revalidation. Hashed bundle names do not save you here. The stale shell names the OLD hashes, those objects still exist, and everything loads cleanly. The failure mode is not an error: it is a deploy that appears to have silently not happened, which is exactly how it was found — the orb work in #1188 was live at the origin while the browser kept rendering the previous build. The CloudFront invalidation the deploy already runs does not help, because the stale copy is in the browser, not at the edge. The sync is now two passes: - Hashed bundles (main-ZMUT4FS2.js, chunk-4DOJ3VBG.js, styles-*.css) get `public, max-age=31536000, immutable`. Safe because the name is a function of the content. - Everything else, index.html above all, gets `no-cache` — cache it, but revalidate. That costs a 304 on a 32KB file. The pattern matches an 8-character hash rather than a `.js` suffix, on purpose: `public/` is copied verbatim and unhashed, and it holds `audio/pcm-capture.worklet.js`, which an extension rule would have pinned for a year. Verified against all 66 unhashed files in `public/` plus the five bundle names live on dev: 5 pinned, 0 leaked. Pass order is load-bearing and commented as such — the second pass sweeps the whole tree so --delete keeps pruning stale bundles, and relies on sync skipping the files the first pass just uploaded so they keep the immutable header. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ache-control fix(frontend): send Cache-Control on deploy so a new build is actually served
…es (#1196) The viewer logged an unhandled rejection on every connect: Uncaught (in promise) {code: 2, message: Display channel is not available.} `requestDisplayLayout` was called from `connect().then()`, but the display channel is not up at that point. It returns a PROMISE, so the try/catch around it never saw the failure — it escaped as an unhandled rejection, which is also why the comment claiming "not fatal" was never actually exercised. Moved into the `firstFrame` observer, which by definition fires once the display channel is live, and the returned promise is now caught so a declined layout stays a console note rather than an unhandled rejection. The stream remains usable at whatever size the server chose. Mutation-checked: calling it from `.then()` again fails the new assertion. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#1198) PR-2 made this measurable. With the preamble's eight DynamoDB reads collapsed into one, `preamble.session_state` still measured 17-21ms on dev while doing no IO at all — six helpers each constructing `boto3.resource("dynamodb")` before reaching their snapshot short-circuit, at the ~1.5ms per construction measured at the start of this work. After PR-2b took the quota read, that residual is ~20% of the 95ms the warm preamble now costs. `apis/shared/aws_clients.py` caches resources and clients per (service, region) and exposes `get_dynamodb_table()`. All 29 construction sites in `sessions/metadata.py` are converted; 23 now-dead `import boto3` lines go with them. Seven of those sites used single quotes and the first pass missed them. Worth naming: a partial conversion would have left the residual half-present and read on the dashboard as "the cache underperforms" rather than "the cache was not applied" — much harder to notice than a clean miss. THE MOTO TRAP `mock_aws()` is entered per test. A client cached under one test's mock keeps pointing at a backend torn down when that test ends, so the next test talks to a dead backend — or to a live AWS endpoint. The failure is order-dependent, the same shape as the static-memo leak this repo already paid for across SPA spec files. So the cache is explicitly resettable and all three `aws` fixtures reset it on entry AND exit. An entry-only reset would still let the last test in a file leak into a file that never uses the fixture, which is why the teardown half has its own test. THE LAMBDA IMAGE CLOSURE A new `apis/shared/` module breaks the scheduled-runs image, which COPYs an explicit import closure rather than the whole package. Caught by `test_lambda_image_copies_its_full_import_closure`, which also names the fix: a COPY in `Dockerfile.scheduled-runs` AND a `MANIFESTS` entry in `build-one.sh` so the content-hash tag notices future edits. Both are here. Without the second half the image would build correctly once and then go stale silently. Also records PR-2b's dev validation: preamble 163ms -> 95ms, whole pre-stream window 349ms -> 263ms, 79% below baseline. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Opening the sign-in viewer issued TWO `POST .../browser/live-view` about 30ms apart, measured on dev. `openViewer()` set `minting` but never checked it, so a second activation of the button re-entered it. Each mint is a live SigV4-signed credential for the browser session, which is reason enough on its own. It also produced a second `dcv.authenticate` that failed with code 10 while the first was still connecting, leaving an error in the console of an otherwise healthy stream and sending me looking for a server fault that was not there.⚠️ Test-design note, because the first version of the second test was worthless: it clicked "Open sign-in" again and asserted the mint count had not changed — but that button no longer exists once the viewer is open, so `again?.click()` was a no-op and it passed with the guard REMOVED. It now asserts the affordance is gone, which is a true statement about the UI, and the guard itself is covered by the double-activation test. Mutation-checked: removing the guard fails that one. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The SDK closes the auth WebSocket as soon as it is finished with it, and
reports that close through the `error` callback EVEN THOUGH authentication
succeeded. Measured on dev, with a single mint so nothing else is in play:
23:32:58.148 [connection] WARN ... <- connect() ALREADY running
WebSocket .../auth failed: Close received after close
dcv.authenticate failed {code: 10}
23:32:58.456 [connectionmanager] Added connection 1
23:32:58.467 [dcv::display:1] ...
23:32:59.177 [dcv::input:1] ...
`connect` is running BEFORE the error arrives, so `success` fired first and
the stream establishes normally afterwards. In Node the same close is a
clean code 1000; Chrome surfaces it as a socket error.
`#status` is an absolutely-positioned overlay covering the whole display, so
the error handler was painting "Could not start the session. It may have
ended." across a working stream — and clearing `starting`, so a later post
could open a second one.
Ignore an auth error once a session exists; a real pre-session failure still
reports, and a disconnect resets the flag so a later attempt can fail loudly
again.
This also retires the `code: 10` I chased for several rounds as a service
fault. It was never a failure at all.
Mutation-checked: removing the guard fails the new assertion.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… disproved (#1201) * perf(observability): decompose the agent build into sub-stages Same move that opened the preamble, applied to what is now the largest number left in the prelude. After PR-2/PR-2b/PR-3 the warm preamble is 95ms, down from 455ms. `agent_build` is ~2950ms on a cold turn against 1-47ms warm, and nothing says which part of it that is. The spec's standing hypothesis was the serial MCP `tools/list` pre-flight, and that hypothesis is UNVERIFIED in two ways worth naming: - dev loads exactly ONE external MCP server (`google_tasks`, over two days of logs), so parallelising it there would save nothing and would be unmeasurable in the only environment we can validate in; and - `AgentFactory.create_agent` receives `tools` and `session_manager` already built, so the expensive work is upstream of it in the agent's own constructor — where `SessionFactory.create_session_manager` restores conversation history from AgentCore Memory, which has never been timed. Also worth recording: the fix the spec proposed — `asyncio.gather` over the servers — would have been a NO-OP. `MCPClient.load_tools()` is `async def` with no await in its hot path: `start()` blocks on `_init_future.result()` and `_list_all_tools_sync()` blocks on `_invoke_on_background_thread(...).result()`. Every coroutine blocks the loop before it yields. That is the same shape as the starved-timer bug agent-state-feedback PR-3 already paid for, and measuring first is what caught it. Seven sub-stages — prompt, registry, session_mgr, tools, hooks, plugins, finalize — plus `agent_build.rest` for the remainder. `groups.agent_build` sums them, so the pre-split series stays comparable, reusing the grouping PR-1 built. Eight EMF metrics derive automatically from the mark names, and the dashboard gains a p90 breakdown row (p90 not p50: a warm build is ~0 across the board, so p50 would be a row of flat lines and the cold builds worth fixing live in the upper percentiles). A CONTEXTVAR, AND WHY THAT IS NOT INCONSISTENT WITH PR-2 PR-2 threaded its session snapshot explicitly and rejected an implicit memo. This does the opposite on purpose. There, implicit state going wrong meant stale session data in production — the bug CLAUDE.md names and this repo has shipped twice. Here it means a timing number is missing or mis-attributed: nothing the user sees, nothing persisted, nothing reaching the model. Against that, the explicit route costs a kwarg threaded through a type registry and three agent classes that do not share constructor signatures. Not propagated across threads, deliberately noted: contextvars do not cross into a ThreadPoolExecutor and the MCP load path crosses one, so a future caller marking from inside it would silently record nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(spec): record PR-3's validation, and correct the per-read claim it disproves PR-3 took the warm preamble from 95ms to 22-37ms and the whole pre-stream window to 128-189ms — cumulatively 95% and ~75% below where this spec opened. It also disproves something this spec asserted in bold. After PR-1 it claimed a GSI query from an AgentCore Runtime container costs ~53ms, 4-5x in-region. It does not. `preamble.ownership` is exactly one query and nothing else; with a cached client it measures 6-7ms. The 53ms was ~47ms of boto3 resource construction on throttled container CPU plus ~6ms of entirely normal DynamoDB. The spec anchored on a ~1.5ms-per-construction figure measured on a laptop, which is ~47ms on the container — thirty times off — and then found a story that fit it. So PR-2 was the right fix for the wrong stated reason: it removed eight client constructions, not eight slow network calls. Recorded rather than quietly overwritten, because the reasoning error is more reusable than the result: a laptop number is not a measurement of production, and a hypothesis that keeps fitting the data can still be wrong about mechanism while right about remedy. The sub-stage marks are what made it visible — a single `preamble` number would have shown the same improvement and hidden the bad explanation. Also names the new largest warm stage: `tools` at 98-102ms, never decomposed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…1202) * fix(browser): block Canvas by all its hostnames, not just the vendor one The blocklist was DEFEATED on dev in about ten seconds. During a takeover a human typed `canvas.boisestate.edu` and the Canvas login page loaded normally. The list contained only `boisestatecanvas.instructure.com`. Chromium's URLBlocklist matches on HOST, not on the service behind it. Both names resolve into Instructure's 99.86.101.0/24 and serve the same LMS, so naming one of them stopped nothing that mattered — a student would use the address they already know. The enforcement mechanism was never the problem: the agent is refused with ERR_BLOCKED_BY_ADMINISTRATOR, and that holds under human control too. The DATA was wrong. instructure.com covers the host AND its subdomains, so the boisestatecanvas/.test/.beta instances are included canvas.boisestate.edu is not a subdomain of anything blocked and has to be listed on its own The general lesson, recorded in both the config and the test: when adding a site, enumerate its aliases FIRST — vanity CNAMEs, regional hosts, the mobile hostname. A list naming only the obvious host is a control that looks real and is not, and it fails silently because the blocked name still blocks. Mutation-checked: restoring the single-host list fails both guards. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(browser): don't block the sign-in page; block where it leads Correcting my own overreach. `canvas.boisestate.edu` is Boise State's sign-in / discovery page, not the Canvas app, so seeing it load during a takeover was not the bypass I called it. Blocking it would have been actively wrong: reaching a login page is exactly what an accessibility or VPAT review needs to do, which is the use case this feature exists to serve. I would have broken the primary purpose to defend against a page where no work is submitted. What survives from that alarm is a real improvement. The seed was the single host `boisestatecanvas.instructure.com`; it is now `instructure.com`, which covers that host AND its subdomains, so the `.test.` and `.beta.` Canvas instances are no longer unlisted side doors. Still unverified, and the thing that actually decides whether this control holds: signing in from the discovery page and confirming the destination is refused. The doorway being open is fine; the room must not be. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…en (#1203) Two problems with the inline frame, both reported from a real takeover: the remote browser is letterboxed into a message-column-width box with no room to work, and nothing tells the user they can actually drive it — it reads as a screenshot. Adds "Take over full screen", which expands the viewer over the whole conversation pane, and a plain statement while expanded: "You're driving this browser — click and type in it as you normally would." Escape or "Exit full screen" returns.⚠️ The iframe is ONE element across both states, and the tests enforce it. Re-creating it on expand would reload it, tearing down the DCV stream and losing whatever the user had typed — a password, most likely. That is also why this is NOT a CDK Dialog as the Angular guide otherwise requires: a dialog or a CDK portal re-parents the content, and browsers reload an iframe when it moves in the DOM. Escape, the close control and the ARIA roles are wired by hand instead, which is the trade this one surface has to make. The expanded layer is role="dialog" + aria-modal only while expanded, so the inline state never reads as a modal to assistive tech. Sizing stays driven by the session viewport inline; expanded it fills the pane and letterboxes inside, so the remote display is never scaled to a shape it is not rendering. An effect leaves full screen if the deadline passes, so a lapsed window cannot strand a dead black rectangle over the conversation.⚠️ Test-design note: the lapse test drives `lapsed` through the INPUT, not the clock. The component's 1s ticker is a real `setInterval` created in the constructor, so `vi.useFakeTimers()` installed afterwards never drives it — a timer-based version of this test passed for the wrong reason. Mutation-checked: splitting the frame into per-state iframes fails two guards. Suite 3580 pass; `ng build --configuration=production` clean, which is the check that catches what tsc alone does not here. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#1201 shipped the sub-stage marks but no spec section, so the epic's headline finding existed only in a chat transcript. This writes it down. `agent_build` on a cold turn is 3286ms, and it decomposes as: tools 2039ms 62% <- the largest number in the whole turn session_mgr 830ms 25% <- AgentCore Memory restore, never timed before finalize 194ms prompt 160ms registry 50ms plugins 13ms hooks 0ms `agent_build.tools` is four times what the preamble cost at its worst and larger than everything this epic has fixed so far, combined. It points at the stage this spec always suspected — the MCP pre-flight lives in there — while disproving the fix it proposed. Dev loads exactly ONE external MCP server and the stage still costs 2039ms, so parallelising across servers cannot be the answer. `_build_filtered_tools()` also does registry filtering, gateway integration and local tool assembly, and how the 2039ms splits between them is NOT yet known; given the mechanism error recorded under PR-3, that is not a gap to fill by reasoning. Also records that the proposed `asyncio.gather` would have been a no-op — `MCPClient.load_tools()` has no await in its hot path, so every coroutine blocks the loop before it yields — and re-scopes the remaining work: PR-5 split agent_build.tools (MCP / gateway / local). Measure first; the pattern has been applied twice and overturned an assumption both times. then agent_build.session_mgr at 830ms. drop asyncio.to_thread for the DynamoDB calls. PR-2 removed seven of the eight blocking reads and PR-3 showed the cost was CPU, not IO wait, so its rationale is gone. Kept in the doc as declined-with-reasons rather than deleted. Status header now carries the epic's outcome: warm preamble 455ms -> 22-37ms (-95%), warm pre-stream window 591-656ms -> 128-189ms (-75%), cold agent_build fully attributed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e-epic-close docs(spec): close the agent state feedback epic
…cy-build-findings docs(spec): record what the agent-build decomposition found
…environment The blocklist is the control that stops a human in a browser takeover reaching a system the agent must not act inside. It was hardcoded to `instructure.com` in config.ts, and `platform.yml` never forwarded CDK_BROWSER_URL_BLOCKLIST — so through the deploy pipeline the list was effectively unconfigurable, and every fork of this stack inherited one institution's LMS policy with no breadcrumb. The hosts that matter are a per-deployment policy call, not a property of this stack, so the default is now empty and the value is supplied per environment, following the `domainName` / `corsOrigins` convention. The operational guidance stays in the comment: Chromium matches on HOST, prefer the registrable domain so `.test.`/`.beta.` instances are covered, and don't block a vendor's sign-in page — reaching a login page is what a VPAT review exists to do. New `parseListEnv` treats empty as unset and falls through to context. An unset GitHub Actions variable arrives as '', and the previous `!== undefined` check would have read that as a deliberate "block nothing" the moment the workflow started forwarding it — silently overriding a configured default. Because the list now comes entirely from outside the repo, a deploy that ships an empty one has to say so: the synth log prints the list, or warns that RBAC on browse_web / request_user_login is the only remaining control. The CDK_BROWSER_URL_BLOCKLIST variable is set to `instructure.com` on both the production and development environments, so neither loses the control. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`seed_bootstrap_data.py` is the only writer of the `TOOL#<id>` rows that RBAC grants and the tool picker read, and nothing syncs it with the Python TOOL_CATALOG. `list_spreadsheets` and `analyze_spreadsheet` were catalogued but never seeded, so a freshly bootstrapped deployment got no rows and the tools could not be granted to any role — silently, because an absent row reads as "not in the catalog" rather than as an error. The live environments only have them because the rows were created by hand. Values mirror the live deployments: category `data`, enabledByDefault true, isPublic true. Default-on is safe against the injected-tool cost trap because SPREADSHEET_TOOL_IDS is in KEY_DESCRIBED_INJECTED_TOOL_IDS — the factory closes over (session_id, user_id, assistant_id), all cache-key elements, so it does not force an agent-cache bypass. Seeding skips existing rows, so this cannot touch a deployed catalog. test_seed_matches_tool_catalog pins the two lists together. It is one- directional: seeded-without-catalog is legitimate for context-bound tools, and `document_read` is excluded by name because it is gated on the session having an attachment rather than on enabled_tools, so it must never get a row. Run against the v1.22.0 seeder it fails with exactly the three ids that were missing, so it is not vacuous. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…utral-defaults fix: stop shipping one institution's config as the fork default
The agent stops working in silence, and stops paying to re-read what it already knows. - Live turn narration (agent_status + tool_group_summary): which tool is running, how long it took, and a model-written one-line summary of each finished batch — drained concurrently so a status line lands while its tool is running, not after it. Nothing it produces reaches the prompt. - Document context offload: a digest replaces an attachment's bytes and document_read pulls back the pages the model asks for. 109.1K -> ~15K prefix on a 60-page canary. The instrument was hiding the win — documentTokens understated PDFs ~14x because Bedrock dual-encodes each page as an image on top of the text layer. - Compaction overhaul in five parts: model-relative thresholds, an 8k summary cap, deferred apply when the re-write is free, tool-result offload at intake, and a selective 1h static-prefix TTL. - Browser sign-in handover: the agent pauses and hands the user a live, interactive browser to sign in, then continues in the authenticated session. The automation stream is DISABLED at the service while the user holds it. Off by default behind its own RBAC tool id. - Response feedback: content-free thumbs joined to cost rows, six reason buckets, retry-with-correction, implicit copy/continue signals, and fleet attribution by config arm. - .docx / .pptx / .csv / .xlsx previews in a docked pane, Agent Templates, and ~500ms off the pre-stream window (the session row was being read eight times per turn). - 25 test cases across 6 files were making real authenticated AWS calls, hidden by fail-open handling. An off-box socket guard blocks them and the suite runs in half the time. A CDK deploy is required. No GSI operation against an existing table and no data backfill. Operators must set CDK_BROWSER_URL_BLOCKLIST, which no longer defaults to one institution's hostname. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| ) | ||
| logger.info( | ||
| "Fleet feedback: %d thumbs over %dd, %d sessions joined (%d omitted)", | ||
| report["totals"]["thumbs"], days, len(records_by_session), omitted, |
| logger.error("eval sampling batch failed", exc_info=True) | ||
|
|
||
| background.add_task(task) | ||
| logger.info("Admin queued an eval sampling batch (limit=%d)", limit) |
|
|
||
| def _format_number(value: float) -> str: | ||
| """Format a float without exposing binary-float noise.""" | ||
| if value != value or value in (float("inf"), float("-inf")): |
| payload["groups"] = groups | ||
| if extra: | ||
| payload.update(extra) | ||
| logger.info("turn_prelude %s", json.dumps(payload, default=str)) |
| from __future__ import annotations | ||
|
|
||
| import logging | ||
| from typing import Any, Dict, Optional |
| #: Bedrock document formats the tool can read. Tabular and presentation files | ||
| #: have their own tools; images have no page/text structure to retrieve. | ||
| DOCUMENT_CLASS_FORMATS = frozenset({"pdf", "docx", "txt", "html", "md"}) | ||
| _TEXT_FORMATS = frozenset({"txt", "html", "md"}) |
| DOCUMENT_CLASS_FORMATS = frozenset({"pdf", "docx", "txt", "html", "md"}) | ||
| _TEXT_FORMATS = frozenset({"txt", "html", "md"}) | ||
|
|
||
| _DOCX_MIME = "application/vnd.openxmlformats-officedocument.wordprocessingml.document" |
|
|
||
| async def _owned_session_sk(session_id: str, user_id: str, table) -> str: | ||
| """The session row's SK, or raise :class:`SessionNotOwned`.""" | ||
| from .metadata import _get_session_by_gsi |
| prose about the conversation can never land beside the cost row.""" | ||
| if _carries_explanation(evaluation): | ||
| raise ValueError("evaluation summaries must not carry the judge's explanation") | ||
| from .metadata import _convert_floats_to_decimal |
| from apis.shared.feature_flags import response_feedback_enabled | ||
|
|
||
| if response_feedback_enabled(): | ||
| from .feedback import query_session_feedback |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 1.23.0 —
develop→main.Release gates
agent-templatestable with zero GSIs, which is aCreateTableand exempt. No existing table takes a GSI operation. Verified by reading thegsi-inventory.jsondiff, not only the script.backfill_*.pyadded in this range. Nothing to announce.[PASS]—VERSION1.22.0 → 1.23.0, all manifests and lockfiles regenerated (includingtui/).What is in it
Eight epics. The theme is the agent showing its work and paying less for it.
agent_statusphases and Nova Microtool_group_summarylines, drained concurrently so a status line reaches the client while its tool runs. A three-tool browse turn previously narrated nothing for 4.5s. Nothing it produces reaches the prompt.document_readpulls back the pages the model asks for. 109.1K → ~15K prefix on a 60-page canary. Six validation findings closed before this release, including a ReDoS in pattern mode and a kill switch that only governed one of three paths.DISABLEDat the service while the user holds the browser..docx,.pptx,.csv,.xlsxin a docked pane.Operator actions
CDK_BROWSER_URL_BLOCKLISTper environment. It previously hardcodedinstructure.comand the pipeline never forwarded it. It now ships empty. Already set onproductionanddevelopment; forks and any other environment must set it or the RBAC grant is the only control. The synth log prints the list or warns when empty.CDK_MANAGED_KB_MIGRATION_ENABLEDbefore deploying — it is read at deploy time, so an environment where it was set since the last platform deploy gets the Upgrade card on this release. Nothing migrates on its own (enrolment is a user-initiated POST), but the card appears.request_user_loginto the roles that should have browser sign-in handover. It shipsenabledByDefault: false.DOCUMENT_OFFLOAD_ROLLOUT_PERCENTstays at 100 — decided for this release. Lower it in an environment where you want a concurrent control arm.Deploy order unchanged:
platform.yml→backend.yml→frontend-deploy.yml.After merge
Backmerge
main→developwith a merge commit, not a squash.🤖 Generated with Claude Code