Skip to content

Release/1.23.0 - #1206

Merged
philmerrell merged 2768 commits into
mainfrom
release/1.23.0
Sep 20, 2026
Merged

philmerrell merged 2768 commits into
mainfrom
release/1.23.0

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Release 1.23.0 — develop → main.

Release gates

Gate Result
DynamoDB GSI update limit ✅ PASS — 28 tables compared. The only inventory change is a new agent-templates table with zero GSIs, which is a CreateTable and exempt. No existing table takes a GSI operation. Verified by reading the gsi-inventory.json diff, not only the script.
Pending backfills ✅ PASS — no backfill_*.py added in this range. Nothing to announce.
Version sync ✅ [PASS] — VERSION 1.22.0 → 1.23.0, all manifests and lockfiles regenerated (including tui/).

What is in it

Eight epics. The theme is the agent showing its work and paying less for it.

  • Live turn narration — agent_status phases and Nova Micro tool_group_summary lines, drained concurrently so a status line reaches the client while its tool runs. A three-tool browse turn previously narrated nothing for 4.5s. Nothing it produces reaches the prompt.
  • Document context offload — attachments were 31% of production spend. A digest replaces the bytes; document_read pulls back the pages the model asks for. 109.1K → ~15K prefix on a 60-page canary. Six validation findings closed before this release, including a ReDoS in pattern mode and a kill switch that only governed one of three paths.
  • Compaction overhaul — five parts: model-relative thresholds, 8k summary cap, deferred apply, tool-result offload at intake, selective 1h static-prefix TTL.
  • Browser sign-in handover — the release's one new user-facing capability. Off by default behind its own RBAC tool id; the automation stream is DISABLED at the service while the user holds the browser.
  • Response feedback — thumbs joined to cost rows, retry-with-correction, implicit signals, fleet attribution by config arm. Eval sampling is opt-in at both layers because it sends real conversations to an AWS-managed judge.
  • File previews — .docx, .pptx, .csv, .xlsx in a docked pane.
  • Agent Templates, and a ~500ms pre-stream latency win (the session row was being read eight times per turn).
  • Test integrity — 25 cases across 6 files were making real authenticated AWS calls, hidden by fail-open handling. Blocked, burned down, suite halved.

Operator actions

  1. Set CDK_BROWSER_URL_BLOCKLIST per environment. It previously hardcoded instructure.com and the pipeline never forwarded it. It now ships empty. Already set on production and development; forks and any other environment must set it or the RBAC grant is the only control. The synth log prints the list or warns when empty.
  2. Confirm CDK_MANAGED_KB_MIGRATION_ENABLED before deploying — it is read at deploy time, so an environment where it was set since the last platform deploy gets the Upgrade card on this release. Nothing migrates on its own (enrolment is a user-initiated POST), but the card appears.
  3. Grant request_user_login to the roles that should have browser sign-in handover. It ships enabledByDefault: false.
  4. DOCUMENT_OFFLOAD_ROLLOUT_PERCENT stays at 100 — decided for this release. Lower it in an environment where you want a concurrent control arm.

Deploy order unchanged: platform.yml → backend.yml → frontend-deploy.yml.

After merge

Backmerge main → develop with a merge commit, not a squash.

🤖 Generated with Claude Code

philmerrell and others added 30 commits September 14, 2026 10:48
Renders the `user_question_required` event PR-1 emits: an inline picker with
single- and multi-select, a pager, an "Other" free-text field and Skip, which
resumes the paused turn in place.

Scope change from the original plan. PR-2 was meant to ship single-question
single-select and leave the carousel to PR-3, but measuring PR-1 against real
models showed Haiku 4.5 and Sonnet 4.6 both routinely ask three or four
questions in one call. A single-question renderer would have been broken for
the common case, not a smaller version of it — so the pager and multi-select
are here. PR-3 keeps reload rehydration.

Two defects found by looking at it in the running app, both invisible to unit
tests and to typecheck:

* Every dark-mode rule in the component was dead. Angular's emulated
  encapsulation stamps its `[_ngcontent-…]` attribute onto each compound
  selector *inside* `:where()`, so `:where(.dark, .dark *)` compiled to
  `.dark[_ngcontent-…]` — and <html class="dark"> carries no such attribute.
  That shipped a near-white selected row under near-white text. Switched to
  `:host-context(.dark)`; measured back at rgba(255,255,255,0.04).
* Utility classes set directly on an `<ng-icon>` element do not apply — even
  `text-gray-700` and `size-5` were ignored — leaving the header glyph on the
  inherited near-white at 1.10 contrast in light mode. The colour now lives on
  the wrapper and the icon inherits currentColor: 9.63.

Verified end to end against real Bedrock through the local stack: the model
paused the turn, the picker took a single-select, a two-value multi-select and
a free-text answer, and all three reached the model on resume.

Backend carries one fix from the same session: headers were sliced at 12 chars
mid-word, rendering "DASHBOARD PU" for "Dashboard Purpose". Now trimmed on a
word boundary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…wnload-links

fix: make generated-document download links durable
…r-question-prompt-ui

feat: clarifying-questions picker in the chat transcript (PR-2)
The `user_question_required` event fires once and never re-streams, so a
refresh mid-prompt orphaned the turn: picker gone, agent still paused waiting
for an answer. `GET /messages` already replayed the breadcrumb PR-1 persists;
this reads it back into the picker.

Also fixes a pre-existing bug in the same area, found while building this and
affecting OAuth and tool-approval prompts too. A breadcrumb outlives its
`pausedTurn` snapshot when the user abandons a paused turn by simply typing
something else — nothing cleared it, because the two `remove_pending_interrupts`
call sites are resume cleanup and the explicit dismiss endpoint. Once
rehydration exists that resurrects a prompt the user can no longer answer: the
resume route 400s on an interrupt id the rebuilt agent never saw, so Submit
becomes a guaranteed error. `clear_pending_interrupts` now runs beside
`clear_paused_turn` at the head of every non-resume turn.

It is a separate function rather than an addition to `clear_paused_turn` on
purpose. That one clears only the snapshot — `test_paused_turn_independent_of_
pending_interrupts` pins it deliberately — and it also runs on the
resume-success and expired-snapshot paths, where the narrower cleanup is
already correct. The supersede policy belongs at the call site that owns it,
next to `clear_interrupted_turn` and `clear_truncated_turn`.

Question-shape validation is extracted to `validateUserQuestions`, shared by
the SSE path and the reload path. They arrive by different transports — a
parsed SSE frame and a JSON string out of DynamoDB — but must agree on what is
renderable; two copies would drift, and the failure mode is a prompt that
renders on one path and vanishes on the other.

Verified against the running stack, both halves: refreshed mid-prompt and the
picker came back with its questions, pager and Other field intact, then
answering it resumed the turn through the snapshot-rebuild path with the
selection reaching the model. Separately, abandoned a prompt by typing
something else, refreshed, and got no stale picker.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…r-question-rehydration

feat: rehydrate the clarifying-questions picker after a refresh (PR-3)
…1104)

Arming CDK_MANAGED_KB_NEW_DEFAULT in dev on 2026-09-14 made the handoff's flag table
factually wrong, which is the one kind of staleness that actively misleads. Corrected,
plus the closeout the specs were owed.

- Flag table: dev NEW_DEFAULT is now `true`, with the point that matters — until this
  date born-managed had been merged, deployed and NEVER ONCE EXECUTED in any
  environment. Every test was a unit test against fakes, and a fake cannot tell you
  that AWS accepts CreateKnowledgeBase or that a work key the trigger writes is one the
  dispatcher's index actually returns. Also notes that dev now runs rungs 2 AND 3 at
  once, so an agent can reach managed by either path — needed when reading logs.
- Records that setting the GitHub variable alone does nothing: it is threaded through
  platform.yml into CDK, so it takes effect only on the next Platform Stack deploy.
  Flipping the variable without deploying looks exactly like a broken feature.
- "Do these first" reordered around reality: observing born-managed run in dev is now
  step 1, and it largely SATISFIES task 15.1, whose live create → ingest → retrieve is
  precisely born-managed's happy path. Record the result against 15.1 rather than
  running a separate probe. Also flags the 15-minute first-document latency as the
  thing most likely to be misread as a hang.
- NEW: a "known open defects, none blocking" table, so three findings stop living in
  chat. The ingestion deleting-guard gap, with the correction that RETRIEVAL is not the
  hole (the status filter serves only `complete` and fails closed — the bug is that
  ingestion promotes the row back to a state the filter accepts) and that it is not
  reachable from a single-tab device upload but plainly is from an import or a crawl.
  Whether a large upload can be cancelled at all. And the chunking-strategy finding:
  managed KBs accept DEFAULT/FIXED_SIZE/NONE, semantic is refused, hierarchical is not
  offered, it is immutable after data-source creation, and we send nothing — which is
  why one image's description splits across chunks.
- Status lines closed out on all three specs. born-managed: BUILT AND MERGED, armed in
  dev, dark in prod. kb-chunk-inspector requirements + design: BUILT, live in dev, task
  6 outstanding, with the Requirement 2 amendment (managed-only) pointed at from the
  design doc so a reader of either lands on it.

No archive directory was created: this repo keeps specs in place and uses tasks.md
checkboxes as state, and inventing a convention for one feature would be worse than the
mild untidiness. The superseded new-default-wiring.md already carries its banner.

Documentation only. No code, no behaviour change.
The picker was first modelled on ToolApprovalPromptComponent, and inherited a
visual language built for a different job. That component is a narrow pill with
a hard 2px accent bar and hairline dividers, which is right for a two-button
yes/no that should stay out of the way. This one is a form the reader has to
think about, and at max-w-xl inside a `justify-start` wrapper it read as
cramped and boxy.

- Full width of the assistant column: the `flex justify-start` wrapper and
  `max-w-xl` are both gone.
- Soft ground instead of chrome: rounded-2xl card, no border — a faint tint
  and a soft ring carry the edge. The left accent bar is dropped.
- Options are discrete rounded rows with air between them rather than an
  edge-to-edge divided list, with roomier hit areas and larger type
  (text-sm/6 labels, text-xs/5 descriptions).
- Rounder controls throughout; "Other" now shares the option row's shape so it
  reads as one more choice rather than a stray input.
- The eyebrow loses its icon. It was pure decoration, and a little glyph on an
  AI prompt is the first thing that makes it look generated.

Measured in the browser, light and dark: eyebrow 10.23, option label 17.75,
description 7.56, pager/skip 7.30, selected-row label 14.31. Selection reads
from the tinted surface and the filled marker together, never colour alone.

`background: white` tripped the surface-literal guard; switched to
var(--color-white), which is what that guard exists to enforce.

Behaviour is untouched and re-verified end to end: single-select, a two-value
multi-select and a free-text answer all reached the model on resume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things were stacking a second outline around a selected row.

A focus-outline fallback written for the "Other" row — a <label> whose
focusable element is the input inside it — was scoped to every `.option`. On a
plain option button `:focus-within` is also true after a MOUSE click, so
clicking drew an outline outside the row. Now scoped to `.option--other`.

And selection itself was adding a coloured ring on top of the row's existing
hairline. Selection is now a fill change: every row carries exactly one
hairline whatever its state, and a warm tint plus the filled marker carry the
state together, so it is still never colour alone.

Keyboard focus is unaffected and still draws a visible ring — the stroke that
was removed is the redundant one, not the accessible one.

Measured on the new tinted surface: label 15.63 light / 14.28 dark,
description 6.66 / 10.13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on-and-nightly-allowlist-guard

fix(security): sanitize the remaining log-injection sinks, guard the nightly allowlist
…r-question-ui-polish

style: full-width, softer clarifying-questions picker
…ost-drilldown-ui

feat(admin-costs): conversations on the user page + session profile, diagnoses and trajectory
The tool worked end to end from PR-2, but the model rarely reached for it —
4/24 on deliberately ambiguous requests against a production-shaped tool set.
A short clause in the system prompt takes that to 24/24, and leaves clear
requests at 0/18 so it does not turn direct questions into interrogations.

The clause is appended only when `ask_user_question` is in the turn's effective
tool list, so a user without the tool never carries an instruction to call it.
The check is on the POST-FILTER list rather than the request's `enabled_tools`:
the two diverge — `ToolFilter` drops a catalog id the registry does not know,
as `canvas_faculty` does in dev today — and keying on the request would
advertise a tool absent from `toolConfig`. It also picks up the
`ASK_USER_QUESTION_ENABLED` kill switch for free.

It is applied to the prompt handed to the agent, never to `self.system_prompt`,
which is snapshotted for resume and hashed into the agent cache key. The clause
derives from `enabled_tools`, which that key already covers via `tools_hash`,
so mutating the field would move resume onto a different cache slot for no
benefit. `PrefixFingerprintHook` reads the prompt off the built agent, so
`systemPromptHash` still reflects what was really sent.

Three things the measurements ruled out, recorded so they are not retried:
rewording the tool description (44-56%, within a baseline band of 17-44%),
changing the tool's position in the list (25-38%), and removing the prompt's
"Cost Awareness" clause (38%). Only the system-prompt clause escapes the noise.
The text is therefore load-bearing in this position, and ships byte-identical
to what was measured, with a test pinning that.

Catalog seed flips to enabledByDefault: True.

Cost: ~63 tokens, constant per configuration, inside the cacheable prefix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two user requests are blocked on the same gap: a Library agent that
evaluates the accessibility of subscription databases it renews annually,
and a VPAT agent that must test vendor demo tenants. Neither target is
reachable by a logged-out browser, and there is no path today for a user
to authenticate a browsing session.

Specs the human-in-the-loop takeover flow: the agent calls take_control()
(UpdateBrowserStream -> automation DISABLED, a service-side mutex, not a
convention), raises an interrupt, the user signs in themselves in an
embedded AWS DCV live view, and the turn resumes authenticated. Credentials
never enter the prompt, the conversation, or AgentCore Memory.

Nearly all of the capability already exists and is unused: take_control /
release_control ship in the pinned bedrock-agentcore 1.21.0, and the
Runtime role is already granted UpdateBrowserStream and
ConnectBrowserLiveViewStream. Also covers browser profiles (so a login is
once-per-vendor, not once-per-conversation), a second VPC-mode browser
resource with session recording, and axe-core as an extension.

Two findings worth flagging independently of whether this ships:

- The existing `live_view` action is broken in the way PR #1101 diagnosed.
  generate_live_view_url signs with SigV4QueryAuth, so the signature is in
  the query string, and browse_tool.py:288 returns it as text in a tool
  result -- the model will re-emit it truncated at the `?`. Max expiry is
  300s, which is also far too short for a human login.

- RBAC granularity is exactly one tool_id, so takeover must be its own
  registered tool rather than a browse_web action (D1). As an action it
  would ship to everyone who can browse, which is backwards: students
  should browse and should not be able to drive a browser inside our AWS
  account. As separate tools, an ungranted user pays zero prefix tokens
  for them.

Measured prefix cost: browse_web 493 tokens today, request_user_login
~206, accessibility_scan ~191. A full human login round trip costs about
one short tool result. The real spend for these agents is screenshots and
accumulated page text, which is why D9 requires a 40-target sweep to be 40
sessions rather than one turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…r-question-trigger-guidance

feat: make the clarifying-questions tool actually fire (PR-4)
Fold KB storage usage into the documents-list response so the agent's
knowledge-base card can show how much of the byte cap is in use.

Backend:
- Add KbUsage model (engine, storedBytes, reservedBytes, cap, elevated)
  to DocumentsListResponse.
- _resolve_kb_usage reads the KB_Record once: managed KBs report their
  bytes and the binding effective_cap (min of owner tier and per-KB
  ceiling); legacy S3-Vectors KBs are uncapped (cap=null). Best-effort so
  a record-read failure never breaks the documents list.

Frontend:
- Usage bar on the KB card: 'X of Y used' for managed KBs, green/yellow/red
  at <75/75-90/>=90% of the cap; legacy KBs show 'X stored', always green,
  no denominator.

Tests: backend route tests for managed/legacy/failure paths; frontend
component specs for the thresholds and the uncapped legacy case.
feat(kb): storage usage bar with byte cap on the KB card
…gent (#1109)

Born-managed treated the absence of a KB_Record as 'brand-new agent' and
provisioned a managed KB on the first upload. But legacy KBs are not
first-class: they share one S3-Vectors index and never write a KB_Record,
so an established legacy agent looks identical to a new one. Its NEXT
upload was therefore mistaken for a first upload, flipping retrieval to an
empty managed KB and stranding the existing corpus on the legacy index.

Guard the record-is-None branch on an existing-documents check
(assistant_has_documents: cheap COUNT, Limit=1, no ownership check). Only
provision when the agent has zero documents; fail toward legacy on any
probe error. Adds mutation-guard + helper unit tests (31 pass).
Byte tracking is scoped to managed KBs (Req 12.11), so a legacy KB reports
zeroed counters and cap=null. The usage bar therefore rendered '0 B stored'
beside real documents on every Classic agent, which reads as a bug.

Gate showUsageBar on the managed engine (kbUsage.engine === 'managed')
instead of merely 'kbUsage is present'. Managed KBs are unaffected; legacy
KBs show only per-document sizes. Adds a spec asserting the bar (and its
progressbar element) is absent for a legacy KB in edit mode.
Run the backend suite with pytest-xdist -n auto on the PR gate so the ~3k tests fan across all runner cores instead of running single-threaded. Also drop -v from backend/pytest.ini (it produced thousands of PASSED lines with no diagnostic value). The nightly coverage run (scripts/backend/test.sh) is left serial on purpose.
Make ts-jest transpile-only (isolatedModules) so jest workers stop re-type-checking the whole project each — the dominant cost of the infra suite (the long pole once backend went parallel in #1111). Add a single 'tsc --noEmit' (npm run build) step to the infra CI job so type safety is preserved on PRs (previously only ts-jest enforced it; tsc ran only in teardown.yml). Workers stay at 2 (the --maxWorkers bump regressed 2.5x in #1112). No const enum in the tree, so isolatedModules is safe. Verified: tsc clean; 178 tests green transpile-only.
…icated-web-assessment-spec

docs(spec): authenticated web assessment via browser takeover
… one turn

Mentioning an Agent ran that turn as the Agent and silently reverted the next
one. The thread still looked like the Agent's while its tools, skills and model
were gone, and nothing surfaced the change — not the UI, and not the model,
which cannot know its own toolset shrank. Asked to use a tool it had used a
moment earlier it got `Unknown tool: create_rubric`, and told the user to toggle
that tool in the picker: a confident wrong diagnosis sending them to fix a
setting that was already correct.

D11 chose the per-turn reading deliberately, so this reverses a decision rather
than repairing an oversight. What changed is the evidence. Measured on prod
sessions-metadata (`turnAgentId` on the `C#` rows): of **247 mentions, 247
started the conversation**. Zero were mid-thread consults; zero mentioned a
second Agent inside a thread bound to a first. (Dev: 60 of 61.) The borrow was
paying an invisible failure mode for a case that has never occurred.

A mention now *means* "talk to this Agent", with two outcomes and no third:

  * empty thread   -> the Agent binds the conversation, exactly like launching
                      it from its card. Safe because there is no history for the
                      binding to misrepresent.
  * has messages   -> the message opens a NEW conversation with that Agent, and
                      the SPA says so. The Agent cannot be bound to history
                      written under other instructions, and must not be borrowed.

Persisting `preferences.assistant_id` is necessary but NOT sufficient, which is
the part that is easy to miss: every turn's Agent is resolved from the request,
the SPA's only carrier is the `assistantId` query param, and the self-heal
effect that refills it from preferences runs on session *load*. So the SPA now
sets the param and stops sending `agent_mention` at all. The backend still
honours that flag for clients that predate this change, and `binds_conversation`
gains `thread_is_empty` so a stale tab mentioning into a fresh thread lands
where a current one does. Its thread lookup runs only for mention turns, so no
bound-Agent turn pays a query it cannot act on.

Also fixed, because it is the same invisible loss on the path users are *told*
to use: "Continue" after a max_tokens truncation skipped the whole assistant
block, so a properly launched Agent finished its reply with none of its tools,
skills, model or instructions (spec-acknowledged as a known edge). The SPA was
already resending `rag_assistant_id` there — `continueTruncatedTurn`'s own
comment says "so the backend rebuilds the same model/tools/assistant agent" —
and only the `not is_continuation` guard discarded it. The block now runs for a
continuation, with binding validation and persistence skipped (it binds nothing
new) and RAG skipped (the turn carries an empty message, so a KB search would
spend a query on "" and augment nothing).

A resume still skips the block and always did keep its tools: it rebuilds from
`PausedTurnSnapshot`, replaying the original turn's exact enabled_tools /
system_prompt / enabled_skills to reconstruct the same prompt-cache key.
Re-resolving there would risk a different effective set and orphan the paused
agent. Worth stating because resume rows carry no `turnAgentId`, so a census of
"turns with no Agent" reads them as losses and overcounts badly.

Two known costs retire with the borrow: the ~$0.12-per-mention prefix re-write
(a bound conversation swaps once and stays instead of swapping back), and the
history fork, where the mention agent and the plain agent were two cached
instances that never saw each other's turns.

The rule lives in one testable place at each end — `mention-routing.ts` on the
client, `agent_binding_policy.py` on the server — following the existing
`system_prompt_resolver` precedent: the rule is a handful of lines, the code
around it is a thousand, and a rule no test can reach is a rule that drifts.

Kept as one commit: the mention and continuation halves edit the same guards on
the same block, and splitting them would mean a first commit that knowingly
leaves the block wrong.

Backend 8585 passed / 3 skipped; frontend 3024 passed across 250 files; tsc
clean. Not exercised in a browser — this worktree's code is not what the local
stack serves, and that stack is down.

Spec: docs/specs/agent-marketplace.md (D11 + Phase 7 notes)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swaps heroPaperAirplaneSolid for heroArrowUpSolid on the chat composer's
submit button. Button geometry, colors, the stop-icon branch while
streaming, and aria-labels are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…poser-up-arrow-icon

feat(chat): use up arrow instead of paper airplane for send button
…ention-binds-conversation

fix(agents): an @-mention binds the conversation instead of borrowing one turn
The greeting heading sits in a 616px text column (720px column, 1rem side
padding, 56px logo, 1rem gap) at text-4xl/tight. Measured against that
column, 19 of 51 greetings wrapped to a second line — and because
AnimatedTextComponent types the greeting out a character at a time, the
wrap happens in full view and pushes the composer down mid-animation.

Rewrite the 20 offenders shorter, keeping the voice. Two of them are stock
DEFAULT_GREETING_TEMPLATES entries that brand.defaults.golden.spec.ts
pinned verbatim: "How can I help you today, {name}?" wrapped for any first
name of 8 characters or more, so the pin was preserving a bug. Update the
pin and say why. The worked example config and the rebranding README get
the same budget.

Add greeting-line-length.spec.ts to hold the line. jsdom has no font
metrics, so it sums per-character advance widths captured from the real
InterVariable woff2 in the app's own <h1>; summed advances track browser
layout to within ±7px across 455 name/greeting combinations, which the 8px
tolerance covers. Verified in a browser that all 57 committed greetings
hold one line with a 12-character first name substituted for {name}.

Narrow viewports are deliberately out of scope: below ~720px the column is
the viewport, and no greeting worth writing fits a phone on one line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…g-one-line

fix(branding): keep every greeting on one line in the chat empty state
…evelop-1.22.0

# Conflicts:
#	backend/src/apis/app_api/documents/routes.py
…into-develop-1.22.0

Backmerge: main into develop (1.22.0)
The admin skill form capped `description` at 500 characters, rejecting
valid Anthropic-authored skills on import. The Agent Skills spec allows
1,024 — so the cap was roughly half the spec and blocked real skills
(`math-olympiad` at 704, `hook-development` at 525), including two of
this repo's own `.claude/skills/`.

The field was also enforced at three different numbers for one column:
500 on the admin path, 2,000 on the user-authored path, and no cap at
all on the `Skill` record itself. Replace all of them with a single
spec-derived constant.

A bound is still warranted — `description` is the Level-1 catalog line
injected into the cacheable system prompt for every enabled skill, on
every turn — so this sets the right number rather than removing the cap.

- Add `SKILL_DESCRIPTION_MAX_LENGTH = 1024` to `apis/shared/skills/models`
- Apply it to SkillCreate/Update (was 500) and CreateMy/UpdateMySkill
  (was 2,000)
- Mirror it client-side in `shared/skills/skill-field-limits.ts`, following
  the `skill-resource-types.ts` precedent, and use it in both skill forms
- Interpolate the constant into the admin form's error text so the message
  cannot drift from the validator again

Note: lowering the user-authored cap from 2,000 to 1,024 affects only
writes. The `Skill` model has no read-side cap, so an existing longer
record still loads; it would fail only on a PUT that resends the
description.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
philmerrell and others added 25 commits September 19, 2026 14:39
…ew (#1189)

#1186 removed a real duplicate-connect bug but the viewer still never
streamed: one `dcv.authenticate`, still failing with code 10 behind a socket
that closed immediately.

Measured against real AgentCore, one probe at a time:

  correct single query  -> socket OPEN, held
  DOUBLED query         -> HTTP 403

That is what we were sending. The SDK APPENDS `httpExtraSearchParams` to
whatever URL it is given, and we handed `authenticate` the full signed URL
with its query still attached AND the extras callback — so every SigV4
parameter went twice and the service refused it.

`connect()` has always stripped the query for exactly this reason; its
comment even says "the transport appends to whatever URL it is given".
`authenticate()` was simply missed, and the failure mode (a socket that
opens then closes, surfacing as "Failed to communicate with server") reads
like a service problem rather than a malformed request.

Also ruled out by probe, so the next person does not re-walk them: base-path
SigV4 signing IS valid for the `/auth` sub-path (signing `/auth` itself
403s), the `dcv` subprotocol is requested correctly by the SDK, an `Origin`
header changes nothing, and `take_control` is not implicated.

Mutation-checked: restoring the signed URL fails the new assertion.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ovable (#1184)

* perf(inference-api): decompose the preamble stage into five sub-marks

`agent-state-feedback.md` PR-3 measured the pre-stream window into four
stages, fixed the one that dominated a cold turn (`agent_build`), and left
`preamble` closed. On a warm turn `preamble` is the largest stage in that
table — 460-494ms of a 641-678ms handler total — and it is a single mark
spanning ~500 lines, so nothing says which of its stages owns the time.

Split it into `preamble.ownership` / `.skills` / `.files` / `.session_state`
/ `.quota`, and teach `TurnPrelude.emit` to sum dotted stages back into a
`groups` field. Without that sum, decomposing the stage would discard the
only measurement anyone has of the window — the four-turn dev baseline is
stated in terms of the single coarse number, and nothing else in the log
reproduces it. `groups.preamble` now does.

Measurement only: no behavior changes, and the code reading suggests but
does not prove where the time goes. The accompanying spec records the
hypothesis (eight serialized reads of the same session row, five of them in
`session_state`) *and* what would falsify it, so the next PR is decided by
the log rather than by the reading.

Nothing reaches the model, the conversation, or the cacheable prefix — four
extra `perf_counter()` calls and one log field per turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(observability): make a preamble fix provable, not just attributable

PR-1 makes the preamble attributable. It does not make a fix provable, and
those need different instruments: attribution needs one turn, proof needs a
population.

Three parts.

**The field that would have lied to us.** `time_to_first_token` is the
obvious instrument and the wrong one: it measures from the coordinator's
`stream_start_time`, which since agent-state-feedback PR-3 runs after the
preamble AND after the agent build. It is blind to the app-api hop, the whole
preamble and the whole build — so a successful PR-2 moves it by exactly zero.
Documented in the spec so nobody validates with it. Same trap this repo
already paid for once with `latency.endToEndLatency` vs `turnDurationMs`.

**Percentiles instead of grepped log lines.** `TurnPrelude.emit` now also
writes one EMF record per turn into `AgentCoreStack/TurnLatency` through the
existing helper — no new infrastructure, no new IAM. Metric names are derived
from stage names rather than listed, so a new mark cannot silently go
unmeasured. Dimension-less like every other EMF caller here; `isResume` and
`deferredBuild` ride as log properties and Logs Insights does the slicing a
dimension would have done badly.

**Widgets on the EXISTING AgentCore dashboard, not a fourth one.** A
dedicated construct was built and then removed:
`observability-platform-dashboard.test.ts` pins the stack at three dashboards
because CloudWatch charges $3/month beyond that. The ceiling is a deliberate
cost decision with a test guarding it, and it is not worth breaking as a side
effect of adding widgets. Folding in is the better design anyway — that
dashboard already graphs AWS's own data-plane `Latency`, so the gap between it
and `PreludeTotalMs` is an independent read on the routing overhead no
server-side stage can see.

Two of the widgets exist to catch our own errors rather than the platform's: a
split by turn shape (a shift toward resumes would look exactly like a latency
win) and an unaccounted-time panel (which says when the decomposition is
incomplete).

No alarms: a threshold needs a baseline and there is none yet. Inventing one
is the guessing the spec exists to prevent.

Kill switch `TURN_LATENCY_METRICS_ENABLED` (default on) — unlike PR-1's marks,
which stay ungated, an EMF line costs log ingestion and custom-metric charges.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(observability): give turn latency its own dashboard, and buy the fourth

Reverses the placement decision in the previous commit, not its substance.

PR-1b put the pre-stream stage widgets on the AgentCore Runtime dashboard to
stay inside CloudWatch's free three. That version worked and was still worse:
this board is read *while shipping a latency change*, side by side with
load-test output, and burying the stage breakdown under runtime-health widgets
answering an unrelated question made it harder to use for its one job.

So the fourth dashboard is bought deliberately at $3/month, against a 450-900ms
wait on every turn. The pinned count in
`observability-platform-dashboard.test.ts` moves to four and now reads "three
free + one bought", with the argument recorded next to it — the point of
pinning the number was never that three is sacred, it was that crossing the
tier should be a decision somebody made rather than a side effect of adding
widgets. The next one should have to argue the same way.

The AgentCore Runtime board stays the companion: it graphs AWS's own
data-plane `Latency`, so the gap between that and `PreludeTotalMs` is the
routing overhead no server-side stage can see. Both headers point at each
other and the platform dashboard links to all three drill-downs.

`inference-agentcore-construct.ts` is byte-identical to its pre-PR-1b state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ket (#1190)

* fix(browser): sign the live-view stream socket, not just the auth socket

`connect` reads `httpExtraSearchParams` ONLY from `observers`. We passed it
at the top level, where it is silently ignored — so the stream socket opened
with no query string at all, unsigned, and the service refused it.

Measured by running the shipped 1.13.2 bundle against a stubbed WebSocket:

  top-level  -> wss://.../live-view/ws                  (UNSIGNED)
  observers  -> wss://.../live-view/ws?X-Amz-Algorithm= (signed)

That is exactly what the browser console showed: repeated failures to
`/live-view/ws` carrying no parameters. `authenticate` is the opposite — it
reads the callback from the top level — which is what made the top-level
form look right. The two entry points genuinely differ.

`firstFrame`/`disconnect` move into `observers` with it, because the SDK
honours observers or callbacks and not both.

This is also what AWS's own BrowserLiveView component does; its config was
read from the published package rather than guessed.

Mutation-checked: restoring the top-level form fails the new assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: remove a scratch probe file committed by mistake

`.connprobe.mjs` was a throwaway jsdom harness used to capture the DCV
SDK's WebSocket URL. Its cleanup `rm` sat in the same command line as a
node process that never exited, so `git add -A` swept it into the branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Measured on dev after #1184 shipped the instrumentation
(docs/specs/turn-latency-preamble.md): a warm turn spends ~455ms in the
preamble, and ~445ms of that is DynamoDB reading ONE item eight times.

The per-read cost is the finding. `preamble.ownership` is exactly one GSI
query and nothing else, and it measures 53-55ms — dead stable across turns.
`preamble.files` on a turn with no attachments is also exactly one read, also
53ms. So a GSI query from inside an AgentCore Runtime container costs ~53ms,
not the ~12ms an in-region figure suggests, and the spec's earlier estimate
that this work would "recover roughly a quarter of the stage" was low by 4x.

`load_session_meta()` performs one `SessionLookupIndex` query and returns a
`SessionMetaSnapshot` carrying both answers the preamble needs: this user's
META row, and whether the session belongs to someone else. Both were always in
that one response — `_get_session_by_gsi` discarded the half the ownership
probe needed, which is why the probe was a second query of the same index for
the same key.

The route reads once and threads the snapshot into `pop_pending_attachments`,
`ensure_session_metadata_exists` and the four `clear_*` helpers. Every one
takes `snapshot=None` by default, so non-preamble callers are untouched.

Two decisions worth naming:

**Explicit, not a per-request memo.** Memoising inside `_get_session_by_gsi`
would be a three-line diff and the wrong shape. CLAUDE.md's "never cache
session state" rule has been paid for twice (#741, #751); a snapshot a caller
opts into cannot leak into a caller that needs a fresh read, and a memo can.

**The snapshot object, not the row dict.** `row is None` is a real answer —
"no META row yet" — not "nothing prefetched". Taking the row alone would make
a brand-new session indistinguishable from an absent prefetch, so the first
turn of every conversation would silently fall back to re-reading. Pinned by
test.

Safety: the helpers were already best-effort and fail-open; the GSI is
eventually consistent so re-reading never bought consistency; the single-flight
lease excludes a concurrent turn on the same session; the SK is static so the
conditional writes' key cannot drift. The atomic parts stay atomic — the
`ReturnValues=UPDATED_OLD` writes are untouched, and the snapshot only replaces
the gate read.

Expected: preamble ~455ms -> ~125ms. To be validated on dev, not assumed.
The quota session-notice read (62ms) is deliberately left for PR-2b: it goes
through `get_session_metadata`, which lazily backfills cost aggregates, and
skipping that would quietly stop session notices on legacy rows.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…coder (#1192)

The live view died with "Display channel is not available" and a black
frame. The browser said why:

  Connecting to 'data:application/octet-stream;base64,AGFzbQ...' violates
  the following Content Security Policy directive: "connect-src 'self'
  https://bedrock-agentcore.us-west-2.amazonaws.com ..."

`AGFzbQ` is the WebAssembly magic number. The DCV Web Client SDK fetches
its decoder from an inline `data:` URI, and `connect-src` was the ONLY
directive here that did not allow `data:`/`blob:` — script-, style-, img-,
font- and media-src all already do.

This does not widen exfiltration. `data:` and `blob:` name bytes inside the
document, not a network peer, so nothing can be sent anywhere through them;
the remote origins a page may reach are still exactly `connectDomains`.

The existing tests pinned the old directive string and failed on this
change, which is what they are for — updated, keeping each one's intent.
The malformed-input case asserted "nothing follows 'self'" as a proxy for
"no injected domains"; it now pins the directive exactly, which tests the
same property without depending on the constants being empty.

Mutation-checked: reverting the directive fails 10 of 34.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…read (#1193)

PR-2b of docs/specs/turn-latency-preamble.md, and the eighth and last read of
the session META row in the preamble — a measured 62ms.

`check_quota` / `_resolve_session_notice` take
`session_total_cost: Optional[float] = None`. The inference-api route supplies
it from the snapshot PR-2 introduced, but ONLY when that row actually carries
`totalCost`.

`None` means "not known", never "zero", and that distinction is the whole
safety of this path. `get_session_metadata` lazily backfills the cost
aggregates for legacy rows written before write-time aggregation; treating an
absent attribute as 0.0 would skip the backfill and silence the session notice
on exactly the long-lived conversations it exists to catch — a latency fix
quietly becoming a feature regression.

A genuinely free session reports 0.0, which is falsy, so the check is
`is None` rather than truthiness: a truthy test would send that session back
to the read this PR removes and the bug would be invisible, because the answer
would still be correct — just expensive.

app-api's converse route passes nothing and keeps its read unchanged.

Expected: preamble.quota 62ms -> ~0, warm preamble ~163ms -> ~100ms. To be
validated on dev.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#1194)

This is why the sign-in viewer never streamed. AWS's own service reference
lists the resource types for each action:

  GetBrowserSession            -> ['browser', 'browser-custom']
  UpdateBrowserStream          -> ['browser', 'browser-custom']
  ConnectBrowserLiveViewStream -> []          <- none

An action with NO resource types never matches a resource-scoped statement,
so granting all three on the browser ARN made the third an implicit deny.
`iam simulate-principal-policy` confirms it: allowed for the first two and
implicitDeny for the third, from the SAME statement on the SAME ARN.

It failed silently and late. `generate_live_view_url` only signs locally and
calls no API, so app-api minted a URL happily; the denial appeared only when
the browser opened the socket and the service closed it, surfacing as DCV
auth code 10, "Failed to communicate with server" — which reads as a service
fault, not a missing permission.

Proven by running the real SDK against a real session outside the browser:
`dcv.authenticate` SUCCEEDS with admin credentials (returns sessionId and
authToken, socket closes 1000) and fails from app-api. Same SDK, same
service, same session — only the signing principal differs.

`*` is as narrow as this action can be expressed. The statement is split so
it carries that one action and nothing else, and the resource-typed actions
stay scoped to the browser ARN. app-api still cannot start, stop or drive a
browser.

Mutation-checked: re-scoping the connect grant fails the new assertion.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… slot

Two problems with where the recap landed, both from it being rendered
inside the turn rather than in the tail.

It appeared on every finished turn. The number describes what just
happened, and a column of durations down the length of a conversation
turns a punctuation mark into a metrics readout — every earlier turn's
recap is a fact nobody asked for, competing with the answer it sits
under. It is now derived once, for the last turn only.

It also overlapped the live line it was meant to replace. The old guard
was `streamingMessageId`, which clears at message_stop; the loader lives
on `isChatLoading`, which clears later at stream close. That gap is
exactly the window where both were on screen, so the finished state read
as a second, different thing appearing beneath the running one rather
than as the running one settling. The recap is now the loading line's
`@else`, in the loader's own slot with the same padding, gated on the
same signal — one slot, never two, and nothing moves when they swap.

Verified across two turns in one session: 1 loader / 0 recaps while
running, 0 loaders / 1 recap when done, and the first turn's mark gone
the moment the second turn starts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…replaces-live-line

fix(ui): show the turn recap only on the latest turn, in the loader's slot
…y served

The SPA was synced to S3 with no Cache-Control metadata on any object,
so CloudFront served index.html with no Cache-Control at all. With no
explicit freshness directive a browser falls back to HEURISTIC caching —
roughly 10% of the document's age when it was cached — which for a shell
that had been sitting in the bucket for days means hours of reuse with no
revalidation.

Hashed bundle names do not save you here. The stale shell names the OLD
hashes, those objects still exist, and everything loads cleanly. The
failure mode is not an error: it is a deploy that appears to have silently
not happened, which is exactly how it was found — the orb work in #1188
was live at the origin while the browser kept rendering the previous
build. The CloudFront invalidation the deploy already runs does not help,
because the stale copy is in the browser, not at the edge.

The sync is now two passes:

- Hashed bundles (main-ZMUT4FS2.js, chunk-4DOJ3VBG.js, styles-*.css) get
  `public, max-age=31536000, immutable`. Safe because the name is a
  function of the content.
- Everything else, index.html above all, gets `no-cache` — cache it, but
  revalidate. That costs a 304 on a 32KB file.

The pattern matches an 8-character hash rather than a `.js` suffix, on
purpose: `public/` is copied verbatim and unhashed, and it holds
`audio/pcm-capture.worklet.js`, which an extension rule would have pinned
for a year. Verified against all 66 unhashed files in `public/` plus the
five bundle names live on dev: 5 pinned, 0 leaked.

Pass order is load-bearing and commented as such — the second pass sweeps
the whole tree so --delete keeps pruning stale bundles, and relies on sync
skipping the files the first pass just uploaded so they keep the immutable
header.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ache-control

fix(frontend): send Cache-Control on deploy so a new build is actually served
…es (#1196)

The viewer logged an unhandled rejection on every connect:

  Uncaught (in promise) {code: 2, message: Display channel is not available.}

`requestDisplayLayout` was called from `connect().then()`, but the display
channel is not up at that point. It returns a PROMISE, so the try/catch
around it never saw the failure — it escaped as an unhandled rejection,
which is also why the comment claiming "not fatal" was never actually
exercised.

Moved into the `firstFrame` observer, which by definition fires once the
display channel is live, and the returned promise is now caught so a
declined layout stays a console note rather than an unhandled rejection.
The stream remains usable at whatever size the server chose.

Mutation-checked: calling it from `.then()` again fails the new assertion.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#1198)

PR-2 made this measurable. With the preamble's eight DynamoDB reads collapsed
into one, `preamble.session_state` still measured 17-21ms on dev while doing no
IO at all — six helpers each constructing `boto3.resource("dynamodb")` before
reaching their snapshot short-circuit, at the ~1.5ms per construction measured
at the start of this work. After PR-2b took the quota read, that residual is
~20% of the 95ms the warm preamble now costs.

`apis/shared/aws_clients.py` caches resources and clients per (service, region)
and exposes `get_dynamodb_table()`. All 29 construction sites in
`sessions/metadata.py` are converted; 23 now-dead `import boto3` lines go with
them.

Seven of those sites used single quotes and the first pass missed them. Worth
naming: a partial conversion would have left the residual half-present and read
on the dashboard as "the cache underperforms" rather than "the cache was not
applied" — much harder to notice than a clean miss.

THE MOTO TRAP

`mock_aws()` is entered per test. A client cached under one test's mock keeps
pointing at a backend torn down when that test ends, so the next test talks to
a dead backend — or to a live AWS endpoint. The failure is order-dependent,
the same shape as the static-memo leak this repo already paid for across SPA
spec files.

So the cache is explicitly resettable and all three `aws` fixtures reset it on
entry AND exit. An entry-only reset would still let the last test in a file
leak into a file that never uses the fixture, which is why the teardown half
has its own test.

THE LAMBDA IMAGE CLOSURE

A new `apis/shared/` module breaks the scheduled-runs image, which COPYs an
explicit import closure rather than the whole package. Caught by
`test_lambda_image_copies_its_full_import_closure`, which also names the fix:
a COPY in `Dockerfile.scheduled-runs` AND a `MANIFESTS` entry in
`build-one.sh` so the content-hash tag notices future edits. Both are here.
Without the second half the image would build correctly once and then go stale
silently.

Also records PR-2b's dev validation: preamble 163ms -> 95ms, whole pre-stream
window 349ms -> 263ms, 79% below baseline.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Opening the sign-in viewer issued TWO `POST .../browser/live-view` about
30ms apart, measured on dev. `openViewer()` set `minting` but never checked
it, so a second activation of the button re-entered it.

Each mint is a live SigV4-signed credential for the browser session, which
is reason enough on its own. It also produced a second `dcv.authenticate`
that failed with code 10 while the first was still connecting, leaving an
error in the console of an otherwise healthy stream and sending me looking
for a server fault that was not there.

⚠️ Test-design note, because the first version of the second test was
worthless: it clicked "Open sign-in" again and asserted the mint count had
not changed — but that button no longer exists once the viewer is open, so
`again?.click()` was a no-op and it passed with the guard REMOVED. It now
asserts the affordance is gone, which is a true statement about the UI, and
the guard itself is covered by the double-activation test. Mutation-checked:
removing the guard fails that one.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The SDK closes the auth WebSocket as soon as it is finished with it, and
reports that close through the `error` callback EVEN THOUGH authentication
succeeded. Measured on dev, with a single mint so nothing else is in play:

  23:32:58.148  [connection] WARN ...          <- connect() ALREADY running
                WebSocket .../auth failed: Close received after close
                dcv.authenticate failed {code: 10}
  23:32:58.456  [connectionmanager] Added connection 1
  23:32:58.467  [dcv::display:1] ...
  23:32:59.177  [dcv::input:1] ...

`connect` is running BEFORE the error arrives, so `success` fired first and
the stream establishes normally afterwards. In Node the same close is a
clean code 1000; Chrome surfaces it as a socket error.

`#status` is an absolutely-positioned overlay covering the whole display, so
the error handler was painting "Could not start the session. It may have
ended." across a working stream — and clearing `starting`, so a later post
could open a second one.

Ignore an auth error once a session exists; a real pre-session failure still
reports, and a disconnect resets the flag so a later attempt can fail loudly
again.

This also retires the `code: 10` I chased for several rounds as a service
fault. It was never a failure at all.

Mutation-checked: removing the guard fails the new assertion.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… disproved (#1201)

* perf(observability): decompose the agent build into sub-stages

Same move that opened the preamble, applied to what is now the largest number
left in the prelude.

After PR-2/PR-2b/PR-3 the warm preamble is 95ms, down from 455ms. `agent_build`
is ~2950ms on a cold turn against 1-47ms warm, and nothing says which part of
it that is. The spec's standing hypothesis was the serial MCP `tools/list`
pre-flight, and that hypothesis is UNVERIFIED in two ways worth naming:

  - dev loads exactly ONE external MCP server (`google_tasks`, over two days
    of logs), so parallelising it there would save nothing and would be
    unmeasurable in the only environment we can validate in; and
  - `AgentFactory.create_agent` receives `tools` and `session_manager` already
    built, so the expensive work is upstream of it in the agent's own
    constructor — where `SessionFactory.create_session_manager` restores
    conversation history from AgentCore Memory, which has never been timed.

Also worth recording: the fix the spec proposed — `asyncio.gather` over the
servers — would have been a NO-OP. `MCPClient.load_tools()` is `async def`
with no await in its hot path: `start()` blocks on `_init_future.result()` and
`_list_all_tools_sync()` blocks on `_invoke_on_background_thread(...).result()`.
Every coroutine blocks the loop before it yields. That is the same shape as the
starved-timer bug agent-state-feedback PR-3 already paid for, and measuring
first is what caught it.

Seven sub-stages — prompt, registry, session_mgr, tools, hooks, plugins,
finalize — plus `agent_build.rest` for the remainder. `groups.agent_build` sums
them, so the pre-split series stays comparable, reusing the grouping PR-1 built.
Eight EMF metrics derive automatically from the mark names, and the dashboard
gains a p90 breakdown row (p90 not p50: a warm build is ~0 across the board, so
p50 would be a row of flat lines and the cold builds worth fixing live in the
upper percentiles).

A CONTEXTVAR, AND WHY THAT IS NOT INCONSISTENT WITH PR-2

PR-2 threaded its session snapshot explicitly and rejected an implicit memo.
This does the opposite on purpose. There, implicit state going wrong meant
stale session data in production — the bug CLAUDE.md names and this repo has
shipped twice. Here it means a timing number is missing or mis-attributed:
nothing the user sees, nothing persisted, nothing reaching the model. Against
that, the explicit route costs a kwarg threaded through a type registry and
three agent classes that do not share constructor signatures.

Not propagated across threads, deliberately noted: contextvars do not cross
into a ThreadPoolExecutor and the MCP load path crosses one, so a future caller
marking from inside it would silently record nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(spec): record PR-3's validation, and correct the per-read claim it disproves

PR-3 took the warm preamble from 95ms to 22-37ms and the whole pre-stream
window to 128-189ms — cumulatively 95% and ~75% below where this spec opened.

It also disproves something this spec asserted in bold. After PR-1 it claimed a
GSI query from an AgentCore Runtime container costs ~53ms, 4-5x in-region. It
does not. `preamble.ownership` is exactly one query and nothing else; with a
cached client it measures 6-7ms. The 53ms was ~47ms of boto3 resource
construction on throttled container CPU plus ~6ms of entirely normal DynamoDB.

The spec anchored on a ~1.5ms-per-construction figure measured on a laptop,
which is ~47ms on the container — thirty times off — and then found a story
that fit it. So PR-2 was the right fix for the wrong stated reason: it removed
eight client constructions, not eight slow network calls.

Recorded rather than quietly overwritten, because the reasoning error is more
reusable than the result: a laptop number is not a measurement of production,
and a hypothesis that keeps fitting the data can still be wrong about mechanism
while right about remedy. The sub-stage marks are what made it visible — a
single `preamble` number would have shown the same improvement and hidden the
bad explanation.

Also names the new largest warm stage: `tools` at 98-102ms, never decomposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…1202)

* fix(browser): block Canvas by all its hostnames, not just the vendor one

The blocklist was DEFEATED on dev in about ten seconds. During a takeover a
human typed `canvas.boisestate.edu` and the Canvas login page loaded
normally. The list contained only `boisestatecanvas.instructure.com`.

Chromium's URLBlocklist matches on HOST, not on the service behind it. Both
names resolve into Instructure's 99.86.101.0/24 and serve the same LMS, so
naming one of them stopped nothing that mattered — a student would use the
address they already know.

The enforcement mechanism was never the problem: the agent is refused with
ERR_BLOCKED_BY_ADMINISTRATOR, and that holds under human control too. The
DATA was wrong.

  instructure.com        covers the host AND its subdomains, so the
                         boisestatecanvas/.test/.beta instances are included
  canvas.boisestate.edu  is not a subdomain of anything blocked and has to
                         be listed on its own

The general lesson, recorded in both the config and the test: when adding a
site, enumerate its aliases FIRST — vanity CNAMEs, regional hosts, the
mobile hostname. A list naming only the obvious host is a control that looks
real and is not, and it fails silently because the blocked name still blocks.

Mutation-checked: restoring the single-host list fails both guards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(browser): don't block the sign-in page; block where it leads

Correcting my own overreach. `canvas.boisestate.edu` is Boise State's
sign-in / discovery page, not the Canvas app, so seeing it load during a
takeover was not the bypass I called it.

Blocking it would have been actively wrong: reaching a login page is exactly
what an accessibility or VPAT review needs to do, which is the use case this
feature exists to serve. I would have broken the primary purpose to defend
against a page where no work is submitted.

What survives from that alarm is a real improvement. The seed was the single
host `boisestatecanvas.instructure.com`; it is now `instructure.com`, which
covers that host AND its subdomains, so the `.test.` and `.beta.` Canvas
instances are no longer unlisted side doors.

Still unverified, and the thing that actually decides whether this control
holds: signing in from the discovery page and confirming the destination is
refused. The doorway being open is fine; the room must not be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…en (#1203)

Two problems with the inline frame, both reported from a real takeover: the
remote browser is letterboxed into a message-column-width box with no room
to work, and nothing tells the user they can actually drive it — it reads as
a screenshot.

Adds "Take over full screen", which expands the viewer over the whole
conversation pane, and a plain statement while expanded: "You're driving
this browser — click and type in it as you normally would." Escape or "Exit
full screen" returns.

⚠️ The iframe is ONE element across both states, and the tests enforce it.
Re-creating it on expand would reload it, tearing down the DCV stream and
losing whatever the user had typed — a password, most likely. That is also
why this is NOT a CDK Dialog as the Angular guide otherwise requires: a
dialog or a CDK portal re-parents the content, and browsers reload an iframe
when it moves in the DOM. Escape, the close control and the ARIA roles are
wired by hand instead, which is the trade this one surface has to make.

The expanded layer is role="dialog" + aria-modal only while expanded, so the
inline state never reads as a modal to assistive tech. Sizing stays driven
by the session viewport inline; expanded it fills the pane and letterboxes
inside, so the remote display is never scaled to a shape it is not
rendering. An effect leaves full screen if the deadline passes, so a lapsed
window cannot strand a dead black rectangle over the conversation.

⚠️ Test-design note: the lapse test drives `lapsed` through the INPUT, not
the clock. The component's 1s ticker is a real `setInterval` created in the
constructor, so `vi.useFakeTimers()` installed afterwards never drives it —
a timer-based version of this test passed for the wrong reason.

Mutation-checked: splitting the frame into per-state iframes fails two
guards. Suite 3580 pass; `ng build --configuration=production` clean, which
is the check that catches what tsc alone does not here.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#1201 shipped the sub-stage marks but no spec section, so the epic's headline
finding existed only in a chat transcript. This writes it down.

`agent_build` on a cold turn is 3286ms, and it decomposes as:

  tools        2039ms   62%   <- the largest number in the whole turn
  session_mgr   830ms   25%   <- AgentCore Memory restore, never timed before
  finalize      194ms
  prompt        160ms
  registry       50ms
  plugins        13ms
  hooks           0ms

`agent_build.tools` is four times what the preamble cost at its worst and
larger than everything this epic has fixed so far, combined.

It points at the stage this spec always suspected — the MCP pre-flight lives
in there — while disproving the fix it proposed. Dev loads exactly ONE external
MCP server and the stage still costs 2039ms, so parallelising across servers
cannot be the answer. `_build_filtered_tools()` also does registry filtering,
gateway integration and local tool assembly, and how the 2039ms splits between
them is NOT yet known; given the mechanism error recorded under PR-3, that is
not a gap to fill by reasoning.

Also records that the proposed `asyncio.gather` would have been a no-op —
`MCPClient.load_tools()` has no await in its hot path, so every coroutine
blocks the loop before it yields — and re-scopes the remaining work:

  PR-5  split agent_build.tools (MCP / gateway / local). Measure first; the
        pattern has been applied twice and overturned an assumption both times.
  then  agent_build.session_mgr at 830ms.
  drop  asyncio.to_thread for the DynamoDB calls. PR-2 removed seven of the
        eight blocking reads and PR-3 showed the cost was CPU, not IO wait, so
        its rationale is gone. Kept in the doc as declined-with-reasons rather
        than deleted.

Status header now carries the epic's outcome: warm preamble 455ms -> 22-37ms
(-95%), warm pre-stream window 591-656ms -> 128-189ms (-75%), cold agent_build
fully attributed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e-epic-close

docs(spec): close the agent state feedback epic
…cy-build-findings

docs(spec): record what the agent-build decomposition found
…environment

The blocklist is the control that stops a human in a browser takeover reaching
a system the agent must not act inside. It was hardcoded to `instructure.com`
in config.ts, and `platform.yml` never forwarded CDK_BROWSER_URL_BLOCKLIST — so
through the deploy pipeline the list was effectively unconfigurable, and every
fork of this stack inherited one institution's LMS policy with no breadcrumb.

The hosts that matter are a per-deployment policy call, not a property of this
stack, so the default is now empty and the value is supplied per environment,
following the `domainName` / `corsOrigins` convention. The operational guidance
stays in the comment: Chromium matches on HOST, prefer the registrable domain
so `.test.`/`.beta.` instances are covered, and don't block a vendor's sign-in
page — reaching a login page is what a VPAT review exists to do.

New `parseListEnv` treats empty as unset and falls through to context. An unset
GitHub Actions variable arrives as '', and the previous `!== undefined` check
would have read that as a deliberate "block nothing" the moment the workflow
started forwarding it — silently overriding a configured default.

Because the list now comes entirely from outside the repo, a deploy that ships
an empty one has to say so: the synth log prints the list, or warns that RBAC
on browse_web / request_user_login is the only remaining control.

The CDK_BROWSER_URL_BLOCKLIST variable is set to `instructure.com` on both the
production and development environments, so neither loses the control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`seed_bootstrap_data.py` is the only writer of the `TOOL#<id>` rows that RBAC
grants and the tool picker read, and nothing syncs it with the Python
TOOL_CATALOG. `list_spreadsheets` and `analyze_spreadsheet` were catalogued but
never seeded, so a freshly bootstrapped deployment got no rows and the tools
could not be granted to any role — silently, because an absent row reads as
"not in the catalog" rather than as an error. The live environments only have
them because the rows were created by hand.

Values mirror the live deployments: category `data`, enabledByDefault true,
isPublic true. Default-on is safe against the injected-tool cost trap because
SPREADSHEET_TOOL_IDS is in KEY_DESCRIBED_INJECTED_TOOL_IDS — the factory closes
over (session_id, user_id, assistant_id), all cache-key elements, so it does
not force an agent-cache bypass. Seeding skips existing rows, so this cannot
touch a deployed catalog.

test_seed_matches_tool_catalog pins the two lists together. It is one-
directional: seeded-without-catalog is legitimate for context-bound tools, and
`document_read` is excluded by name because it is gated on the session having
an attachment rather than on enabled_tools, so it must never get a row. Run
against the v1.22.0 seeder it fails with exactly the three ids that were
missing, so it is not vacuous.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…utral-defaults

fix: stop shipping one institution's config as the fork default
The agent stops working in silence, and stops paying to re-read what it
already knows.

- Live turn narration (agent_status + tool_group_summary): which tool is
  running, how long it took, and a model-written one-line summary of each
  finished batch — drained concurrently so a status line lands while its
  tool is running, not after it. Nothing it produces reaches the prompt.
- Document context offload: a digest replaces an attachment's bytes and
  document_read pulls back the pages the model asks for. 109.1K -> ~15K
  prefix on a 60-page canary. The instrument was hiding the win —
  documentTokens understated PDFs ~14x because Bedrock dual-encodes each
  page as an image on top of the text layer.
- Compaction overhaul in five parts: model-relative thresholds, an 8k
  summary cap, deferred apply when the re-write is free, tool-result
  offload at intake, and a selective 1h static-prefix TTL.
- Browser sign-in handover: the agent pauses and hands the user a live,
  interactive browser to sign in, then continues in the authenticated
  session. The automation stream is DISABLED at the service while the
  user holds it. Off by default behind its own RBAC tool id.
- Response feedback: content-free thumbs joined to cost rows, six reason
  buckets, retry-with-correction, implicit copy/continue signals, and
  fleet attribution by config arm.
- .docx / .pptx / .csv / .xlsx previews in a docked pane, Agent Templates,
  and ~500ms off the pre-stream window (the session row was being read
  eight times per turn).
- 25 test cases across 6 files were making real authenticated AWS calls,
  hidden by fail-open handling. An off-box socket guard blocks them and
  the suite runs in half the time.

A CDK deploy is required. No GSI operation against an existing table and
no data backfill. Operators must set CDK_BROWSER_URL_BLOCKLIST, which no
longer defaults to one institution's hostname.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@philmerrell
philmerrell requested a review from a team September 20, 2026 20:31
)
logger.info(
"Fleet feedback: %d thumbs over %dd, %d sessions joined (%d omitted)",
report["totals"]["thumbs"], days, len(records_by_session), omitted,
logger.error("eval sampling batch failed", exc_info=True)

background.add_task(task)
logger.info("Admin queued an eval sampling batch (limit=%d)", limit)

def _format_number(value: float) -> str:
"""Format a float without exposing binary-float noise."""
if value != value or value in (float("inf"), float("-inf")):
payload["groups"] = groups
if extra:
payload.update(extra)
logger.info("turn_prelude %s", json.dumps(payload, default=str))
from __future__ import annotations

import logging
from typing import Any, Dict, Optional
#: Bedrock document formats the tool can read. Tabular and presentation files
#: have their own tools; images have no page/text structure to retrieve.
DOCUMENT_CLASS_FORMATS = frozenset({"pdf", "docx", "txt", "html", "md"})
_TEXT_FORMATS = frozenset({"txt", "html", "md"})
DOCUMENT_CLASS_FORMATS = frozenset({"pdf", "docx", "txt", "html", "md"})
_TEXT_FORMATS = frozenset({"txt", "html", "md"})

_DOCX_MIME = "application/vnd.openxmlformats-officedocument.wordprocessingml.document"

async def _owned_session_sk(session_id: str, user_id: str, table) -> str:
"""The session row's SK, or raise :class:`SessionNotOwned`."""
from .metadata import _get_session_by_gsi
prose about the conversation can never land beside the cost row."""
if _carries_explanation(evaluation):
raise ValueError("evaluation summaries must not carry the judge's explanation")
from .metadata import _convert_floats_to_decimal
from apis.shared.feature_flags import response_feedback_enabled

if response_feedback_enabled():
from .feedback import query_session_feedback
@philmerrell
philmerrell merged commit 32b7ba0 into main Sep 20, 2026
15 checks passed
@philmerrell
philmerrell deleted the release/1.23.0 branch September 20, 2026 20:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants