perf(ci): parallelize backend pytest with xdist - #1111
Merged
Merged
Conversation
Run the backend suite with pytest-xdist -n auto on the PR gate so the ~3k tests fan across all runner cores instead of running single-threaded. Also drop -v from backend/pytest.ini (it produced thousands of PASSED lines with no diagnostic value). The nightly coverage run (scripts/backend/test.sh) is left serial on purpose.
This was referenced Sep 14, 2026
DerrickF
added a commit
that referenced
this pull request
Sep 15, 2026
Make ts-jest transpile-only (isolatedModules) so jest workers stop re-type-checking the whole project each — the dominant cost of the infra suite (the long pole once backend went parallel in #1111). Add a single 'tsc --noEmit' (npm run build) step to the infra CI job so type safety is preserved on PRs (previously only ts-jest enforced it; tsc ran only in teardown.yml). Workers stay at 2 (the --maxWorkers bump regressed 2.5x in #1112). No const enum in the tree, so isolatedModules is safe. Verified: tsc clean; 178 tests green transpile-only.
philmerrell
added a commit
that referenced
this pull request
Sep 20, 2026
* feat: clarifying-questions picker in the chat transcript (PR-2)
Renders the `user_question_required` event PR-1 emits: an inline picker with
single- and multi-select, a pager, an "Other" free-text field and Skip, which
resumes the paused turn in place.
Scope change from the original plan. PR-2 was meant to ship single-question
single-select and leave the carousel to PR-3, but measuring PR-1 against real
models showed Haiku 4.5 and Sonnet 4.6 both routinely ask three or four
questions in one call. A single-question renderer would have been broken for
the common case, not a smaller version of it — so the pager and multi-select
are here. PR-3 keeps reload rehydration.
Two defects found by looking at it in the running app, both invisible to unit
tests and to typecheck:
* Every dark-mode rule in the component was dead. Angular's emulated
encapsulation stamps its `[_ngcontent-…]` attribute onto each compound
selector *inside* `:where()`, so `:where(.dark, .dark *)` compiled to
`.dark[_ngcontent-…]` — and <html class="dark"> carries no such attribute.
That shipped a near-white selected row under near-white text. Switched to
`:host-context(.dark)`; measured back at rgba(255,255,255,0.04).
* Utility classes set directly on an `<ng-icon>` element do not apply — even
`text-gray-700` and `size-5` were ignored — leaving the header glyph on the
inherited near-white at 1.10 contrast in light mode. The colour now lives on
the wrapper and the icon inherits currentColor: 9.63.
Verified end to end against real Bedrock through the local stack: the model
paused the turn, the picker took a single-select, a two-value multi-select and
a free-text answer, and all three reached the model on resume.
Backend carries one fix from the same session: headers were sliced at 12 chars
mid-word, rendering "DASHBOARD PU" for "Dashboard Purpose". Now trimmed on a
word boundary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat: rehydrate the clarifying-questions picker after a refresh (PR-3)
The `user_question_required` event fires once and never re-streams, so a
refresh mid-prompt orphaned the turn: picker gone, agent still paused waiting
for an answer. `GET /messages` already replayed the breadcrumb PR-1 persists;
this reads it back into the picker.
Also fixes a pre-existing bug in the same area, found while building this and
affecting OAuth and tool-approval prompts too. A breadcrumb outlives its
`pausedTurn` snapshot when the user abandons a paused turn by simply typing
something else — nothing cleared it, because the two `remove_pending_interrupts`
call sites are resume cleanup and the explicit dismiss endpoint. Once
rehydration exists that resurrects a prompt the user can no longer answer: the
resume route 400s on an interrupt id the rebuilt agent never saw, so Submit
becomes a guaranteed error. `clear_pending_interrupts` now runs beside
`clear_paused_turn` at the head of every non-resume turn.
It is a separate function rather than an addition to `clear_paused_turn` on
purpose. That one clears only the snapshot — `test_paused_turn_independent_of_
pending_interrupts` pins it deliberately — and it also runs on the
resume-success and expired-snapshot paths, where the narrower cleanup is
already correct. The supersede policy belongs at the call site that owns it,
next to `clear_interrupted_turn` and `clear_truncated_turn`.
Question-shape validation is extracted to `validateUserQuestions`, shared by
the SSE path and the reload path. They arrive by different transports — a
parsed SSE frame and a JSON string out of DynamoDB — but must agree on what is
renderable; two copies would drift, and the failure mode is a prompt that
renders on one path and vanishes on the other.
Verified against the running stack, both halves: refreshed mid-prompt and the
picker came back with its questions, pager and Other field intact, then
answering it resumed the turn through the snapshot-rebuild path with the
selection reaching the model. Separately, abandoned a prompt by typing
something else, refreshed, and got no stale picker.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(kb): close out the KB specs and record NEW_DEFAULT armed in dev (#1104)
Arming CDK_MANAGED_KB_NEW_DEFAULT in dev on 2026-09-14 made the handoff's flag table
factually wrong, which is the one kind of staleness that actively misleads. Corrected,
plus the closeout the specs were owed.
- Flag table: dev NEW_DEFAULT is now `true`, with the point that matters — until this
date born-managed had been merged, deployed and NEVER ONCE EXECUTED in any
environment. Every test was a unit test against fakes, and a fake cannot tell you
that AWS accepts CreateKnowledgeBase or that a work key the trigger writes is one the
dispatcher's index actually returns. Also notes that dev now runs rungs 2 AND 3 at
once, so an agent can reach managed by either path — needed when reading logs.
- Records that setting the GitHub variable alone does nothing: it is threaded through
platform.yml into CDK, so it takes effect only on the next Platform Stack deploy.
Flipping the variable without deploying looks exactly like a broken feature.
- "Do these first" reordered around reality: observing born-managed run in dev is now
step 1, and it largely SATISFIES task 15.1, whose live create → ingest → retrieve is
precisely born-managed's happy path. Record the result against 15.1 rather than
running a separate probe. Also flags the 15-minute first-document latency as the
thing most likely to be misread as a hang.
- NEW: a "known open defects, none blocking" table, so three findings stop living in
chat. The ingestion deleting-guard gap, with the correction that RETRIEVAL is not the
hole (the status filter serves only `complete` and fails closed — the bug is that
ingestion promotes the row back to a state the filter accepts) and that it is not
reachable from a single-tab device upload but plainly is from an import or a crawl.
Whether a large upload can be cancelled at all. And the chunking-strategy finding:
managed KBs accept DEFAULT/FIXED_SIZE/NONE, semantic is refused, hierarchical is not
offered, it is immutable after data-source creation, and we send nothing — which is
why one image's description splits across chunks.
- Status lines closed out on all three specs. born-managed: BUILT AND MERGED, armed in
dev, dark in prod. kb-chunk-inspector requirements + design: BUILT, live in dev, task
6 outstanding, with the Requirement 2 amendment (managed-only) pointed at from the
design doc so a reader of either lands on it.
No archive directory was created: this repo keeps specs in place and uses tasks.md
checkboxes as state, and inventing a convention for one feature would be worse than the
mild untidiness. The superseded new-default-wiring.md already carries its banner.
Documentation only. No code, no behaviour change.
* style: full-width, softer clarifying-questions picker
The picker was first modelled on ToolApprovalPromptComponent, and inherited a
visual language built for a different job. That component is a narrow pill with
a hard 2px accent bar and hairline dividers, which is right for a two-button
yes/no that should stay out of the way. This one is a form the reader has to
think about, and at max-w-xl inside a `justify-start` wrapper it read as
cramped and boxy.
- Full width of the assistant column: the `flex justify-start` wrapper and
`max-w-xl` are both gone.
- Soft ground instead of chrome: rounded-2xl card, no border — a faint tint
and a soft ring carry the edge. The left accent bar is dropped.
- Options are discrete rounded rows with air between them rather than an
edge-to-edge divided list, with roomier hit areas and larger type
(text-sm/6 labels, text-xs/5 descriptions).
- Rounder controls throughout; "Other" now shares the option row's shape so it
reads as one more choice rather than a stray input.
- The eyebrow loses its icon. It was pure decoration, and a little glyph on an
AI prompt is the first thing that makes it look generated.
Measured in the browser, light and dark: eyebrow 10.23, option label 17.75,
description 7.56, pager/skip 7.30, selected-row label 14.31. Selection reads
from the tinted surface and the filled marker together, never colour alone.
`background: white` tripped the surface-literal guard; switched to
var(--color-white), which is what that guard exists to enforce.
Behaviour is untouched and re-verified end to end: single-select, a two-value
multi-select and a free-text answer all reached the model on resume.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* style: single stroke on a selected option
Two things were stacking a second outline around a selected row.
A focus-outline fallback written for the "Other" row — a <label> whose
focusable element is the input inside it — was scoped to every `.option`. On a
plain option button `:focus-within` is also true after a MOUSE click, so
clicking drew an outline outside the row. Now scoped to `.option--other`.
And selection itself was adding a coloured ring on top of the row's existing
hairline. Selection is now a fill change: every row carries exactly one
hairline whatever its state, and a warm tint plus the filled marker carry the
state together, so it is still never colour alone.
Keyboard focus is unaffected and still draws a visible ring — the stroke that
was removed is the redundant one, not the accessible one.
Measured on the new tinted surface: label 15.63 light / 14.28 dark,
description 6.66 / 10.13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat: make the clarifying-questions tool actually fire (PR-4)
The tool worked end to end from PR-2, but the model rarely reached for it —
4/24 on deliberately ambiguous requests against a production-shaped tool set.
A short clause in the system prompt takes that to 24/24, and leaves clear
requests at 0/18 so it does not turn direct questions into interrogations.
The clause is appended only when `ask_user_question` is in the turn's effective
tool list, so a user without the tool never carries an instruction to call it.
The check is on the POST-FILTER list rather than the request's `enabled_tools`:
the two diverge — `ToolFilter` drops a catalog id the registry does not know,
as `canvas_faculty` does in dev today — and keying on the request would
advertise a tool absent from `toolConfig`. It also picks up the
`ASK_USER_QUESTION_ENABLED` kill switch for free.
It is applied to the prompt handed to the agent, never to `self.system_prompt`,
which is snapshotted for resume and hashed into the agent cache key. The clause
derives from `enabled_tools`, which that key already covers via `tools_hash`,
so mutating the field would move resume onto a different cache slot for no
benefit. `PrefixFingerprintHook` reads the prompt off the built agent, so
`systemPromptHash` still reflects what was really sent.
Three things the measurements ruled out, recorded so they are not retried:
rewording the tool description (44-56%, within a baseline band of 17-44%),
changing the tool's position in the list (25-38%), and removing the prompt's
"Cost Awareness" clause (38%). Only the system-prompt clause escapes the noise.
The text is therefore load-bearing in this position, and ships byte-identical
to what was measured, with a test pinning that.
Catalog seed flips to enabledByDefault: True.
Cost: ~63 tokens, constant per configuration, inside the cacheable prefix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(spec): authenticated web assessment via browser takeover
Two user requests are blocked on the same gap: a Library agent that
evaluates the accessibility of subscription databases it renews annually,
and a VPAT agent that must test vendor demo tenants. Neither target is
reachable by a logged-out browser, and there is no path today for a user
to authenticate a browsing session.
Specs the human-in-the-loop takeover flow: the agent calls take_control()
(UpdateBrowserStream -> automation DISABLED, a service-side mutex, not a
convention), raises an interrupt, the user signs in themselves in an
embedded AWS DCV live view, and the turn resumes authenticated. Credentials
never enter the prompt, the conversation, or AgentCore Memory.
Nearly all of the capability already exists and is unused: take_control /
release_control ship in the pinned bedrock-agentcore 1.21.0, and the
Runtime role is already granted UpdateBrowserStream and
ConnectBrowserLiveViewStream. Also covers browser profiles (so a login is
once-per-vendor, not once-per-conversation), a second VPC-mode browser
resource with session recording, and axe-core as an extension.
Two findings worth flagging independently of whether this ships:
- The existing `live_view` action is broken in the way PR #1101 diagnosed.
generate_live_view_url signs with SigV4QueryAuth, so the signature is in
the query string, and browse_tool.py:288 returns it as text in a tool
result -- the model will re-emit it truncated at the `?`. Max expiry is
300s, which is also far too short for a human login.
- RBAC granularity is exactly one tool_id, so takeover must be its own
registered tool rather than a browse_web action (D1). As an action it
would ship to everyone who can browse, which is backwards: students
should browse and should not be able to drive a browser inside our AWS
account. As separate tools, an ungranted user pays zero prefix tokens
for them.
Measured prefix cost: browse_web 493 tokens today, request_user_login
~206, accessibility_scan ~191. A full human login round trip costs about
one short tool result. The real spend for these agents is screenshots and
accumulated page text, which is why D9 requires a 40-target sweep to be 40
sessions rather than one turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(kb): show storage usage bar with byte cap on the KB card
Fold KB storage usage into the documents-list response so the agent's
knowledge-base card can show how much of the byte cap is in use.
Backend:
- Add KbUsage model (engine, storedBytes, reservedBytes, cap, elevated)
to DocumentsListResponse.
- _resolve_kb_usage reads the KB_Record once: managed KBs report their
bytes and the binding effective_cap (min of owner tier and per-KB
ceiling); legacy S3-Vectors KBs are uncapped (cap=null). Best-effort so
a record-read failure never breaks the documents list.
Frontend:
- Usage bar on the KB card: 'X of Y used' for managed KBs, green/yellow/red
at <75/75-90/>=90% of the cap; legacy KBs show 'X stored', always green,
no denominator.
Tests: backend route tests for managed/legacy/failure paths; frontend
component specs for the thresholds and the uncapped legacy case.
* fix(kb): born-managed must not provision over an established legacy agent (#1109)
Born-managed treated the absence of a KB_Record as 'brand-new agent' and
provisioned a managed KB on the first upload. But legacy KBs are not
first-class: they share one S3-Vectors index and never write a KB_Record,
so an established legacy agent looks identical to a new one. Its NEXT
upload was therefore mistaken for a first upload, flipping retrieval to an
empty managed KB and stranding the existing corpus on the legacy index.
Guard the record-is-None branch on an existing-documents check
(assistant_has_documents: cheap COUNT, Limit=1, no ownership check). Only
provision when the agent has zero documents; fail toward legacy on any
probe error. Adds mutation-guard + helper unit tests (31 pass).
* fix(kb): hide the storage usage bar for legacy (Classic) KBs (#1110)
Byte tracking is scoped to managed KBs (Req 12.11), so a legacy KB reports
zeroed counters and cap=null. The usage bar therefore rendered '0 B stored'
beside real documents on every Classic agent, which reads as a bug.
Gate showUsageBar on the managed engine (kbUsage.engine === 'managed')
instead of merely 'kbUsage is present'. Managed KBs are unaffected; legacy
KBs show only per-document sizes. Adds a spec asserting the bar (and its
progressbar element) is absent for a legacy KB in edit mode.
* perf(ci): parallelize backend pytest with xdist (#1111)
Run the backend suite with pytest-xdist -n auto on the PR gate so the ~3k tests fan across all runner cores instead of running single-threaded. Also drop -v from backend/pytest.ini (it produced thousands of PASSED lines with no diagnostic value). The nightly coverage run (scripts/backend/test.sh) is left serial on purpose.
* perf(ci): infra jest transpile-only + shared tsc type-check (#1113)
Make ts-jest transpile-only (isolatedModules) so jest workers stop re-type-checking the whole project each — the dominant cost of the infra suite (the long pole once backend went parallel in #1111). Add a single 'tsc --noEmit' (npm run build) step to the infra CI job so type safety is preserved on PRs (previously only ts-jest enforced it; tsc ran only in teardown.yml). Workers stay at 2 (the --maxWorkers bump regressed 2.5x in #1112). No const enum in the tree, so isolatedModules is safe. Verified: tsc clean; 178 tests green transpile-only.
* fix(agents): an @-mention binds the conversation instead of borrowing one turn
Mentioning an Agent ran that turn as the Agent and silently reverted the next
one. The thread still looked like the Agent's while its tools, skills and model
were gone, and nothing surfaced the change — not the UI, and not the model,
which cannot know its own toolset shrank. Asked to use a tool it had used a
moment earlier it got `Unknown tool: create_rubric`, and told the user to toggle
that tool in the picker: a confident wrong diagnosis sending them to fix a
setting that was already correct.
D11 chose the per-turn reading deliberately, so this reverses a decision rather
than repairing an oversight. What changed is the evidence. Measured on prod
sessions-metadata (`turnAgentId` on the `C#` rows): of **247 mentions, 247
started the conversation**. Zero were mid-thread consults; zero mentioned a
second Agent inside a thread bound to a first. (Dev: 60 of 61.) The borrow was
paying an invisible failure mode for a case that has never occurred.
A mention now *means* "talk to this Agent", with two outcomes and no third:
* empty thread -> the Agent binds the conversation, exactly like launching
it from its card. Safe because there is no history for the
binding to misrepresent.
* has messages -> the message opens a NEW conversation with that Agent, and
the SPA says so. The Agent cannot be bound to history
written under other instructions, and must not be borrowed.
Persisting `preferences.assistant_id` is necessary but NOT sufficient, which is
the part that is easy to miss: every turn's Agent is resolved from the request,
the SPA's only carrier is the `assistantId` query param, and the self-heal
effect that refills it from preferences runs on session *load*. So the SPA now
sets the param and stops sending `agent_mention` at all. The backend still
honours that flag for clients that predate this change, and `binds_conversation`
gains `thread_is_empty` so a stale tab mentioning into a fresh thread lands
where a current one does. Its thread lookup runs only for mention turns, so no
bound-Agent turn pays a query it cannot act on.
Also fixed, because it is the same invisible loss on the path users are *told*
to use: "Continue" after a max_tokens truncation skipped the whole assistant
block, so a properly launched Agent finished its reply with none of its tools,
skills, model or instructions (spec-acknowledged as a known edge). The SPA was
already resending `rag_assistant_id` there — `continueTruncatedTurn`'s own
comment says "so the backend rebuilds the same model/tools/assistant agent" —
and only the `not is_continuation` guard discarded it. The block now runs for a
continuation, with binding validation and persistence skipped (it binds nothing
new) and RAG skipped (the turn carries an empty message, so a KB search would
spend a query on "" and augment nothing).
A resume still skips the block and always did keep its tools: it rebuilds from
`PausedTurnSnapshot`, replaying the original turn's exact enabled_tools /
system_prompt / enabled_skills to reconstruct the same prompt-cache key.
Re-resolving there would risk a different effective set and orphan the paused
agent. Worth stating because resume rows carry no `turnAgentId`, so a census of
"turns with no Agent" reads them as losses and overcounts badly.
Two known costs retire with the borrow: the ~$0.12-per-mention prefix re-write
(a bound conversation swaps once and stays instead of swapping back), and the
history fork, where the mention agent and the plain agent were two cached
instances that never saw each other's turns.
The rule lives in one testable place at each end — `mention-routing.ts` on the
client, `agent_binding_policy.py` on the server — following the existing
`system_prompt_resolver` precedent: the rule is a handful of lines, the code
around it is a thousand, and a rule no test can reach is a rule that drifts.
Kept as one commit: the mention and continuation halves edit the same guards on
the same block, and splitting them would mean a first commit that knowingly
leaves the block wrong.
Backend 8585 passed / 3 skipped; frontend 3024 passed across 250 files; tsc
clean. Not exercised in a browser — this worktree's code is not what the local
stack serves, and that stack is down.
Spec: docs/specs/agent-marketplace.md (D11 + Phase 7 notes)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(chat): use up arrow instead of paper airplane for send button
Swaps heroPaperAirplaneSolid for heroArrowUpSolid on the chat composer's
submit button. Button geometry, colors, the stop-icon branch while
streaming, and aria-labels are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(branding): keep every greeting on one line in the chat empty state
The greeting heading sits in a 616px text column (720px column, 1rem side
padding, 56px logo, 1rem gap) at text-4xl/tight. Measured against that
column, 19 of 51 greetings wrapped to a second line — and because
AnimatedTextComponent types the greeting out a character at a time, the
wrap happens in full view and pushes the composer down mid-animation.
Rewrite the 20 offenders shorter, keeping the voice. Two of them are stock
DEFAULT_GREETING_TEMPLATES entries that brand.defaults.golden.spec.ts
pinned verbatim: "How can I help you today, {name}?" wrapped for any first
name of 8 characters or more, so the pin was preserving a bug. Update the
pin and say why. The worked example config and the rebranding README get
the same budget.
Add greeting-line-length.spec.ts to hold the line. jsdom has no font
metrics, so it sums per-character advance widths captured from the real
InterVariable woff2 in the app's own <h1>; summed advances track browser
layout to within ±7px across 455 name/greeting combinations, which the 8px
tolerance covers. Verified in a browser that all 57 committed greetings
hold one line with a 12-character first name substituted for {name}.
Narrow viewports are deliberately out of scope: below ~720px the column is
the viewport, and no greeting worth writing fits a phone on one line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(skills): align description cap with the 1024-char Agent Skills spec
The admin skill form capped `description` at 500 characters, rejecting
valid Anthropic-authored skills on import. The Agent Skills spec allows
1,024 — so the cap was roughly half the spec and blocked real skills
(`math-olympiad` at 704, `hook-development` at 525), including two of
this repo's own `.claude/skills/`.
The field was also enforced at three different numbers for one column:
500 on the admin path, 2,000 on the user-authored path, and no cap at
all on the `Skill` record itself. Replace all of them with a single
spec-derived constant.
A bound is still warranted — `description` is the Level-1 catalog line
injected into the cacheable system prompt for every enabled skill, on
every turn — so this sets the right number rather than removing the cap.
- Add `SKILL_DESCRIPTION_MAX_LENGTH = 1024` to `apis/shared/skills/models`
- Apply it to SkillCreate/Update (was 500) and CreateMy/UpdateMySkill
(was 2,000)
- Mirror it client-side in `shared/skills/skill-field-limits.ts`, following
the `skill-resource-types.ts` precedent, and use it in both skill forms
- Interpolate the constant into the admin form's error text so the message
cannot drift from the validator again
Note: lowering the user-authored cap from 2,000 to 1,024 affects only
writes. The `Skill` model has no read-side cap, so an existing longer
record still loads; it would fail only on a PUT that resends the
description.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(prompts): point tool-enablement guidance at Customize → Tools
The composer settings drawer was deleted by the Customize epic
(docs/specs/customize-surface.md: "Does the settings icon survive? No."),
but the prompt and UI copy kept directing users to "the gear icon next to
the message input".
Model- and user-facing:
- system_prompt_builder.py HANDLING MISSING TOOLS — intro, step 2, and the
worked example, which is a verbatim reply template the model reproduced
word-for-word.
- inference_api/chat/routes.py — the spreadsheet and PowerPoint diversion
notes, which are appended to the user's own message and render in-thread.
- chat-input.component.ts — the "Enable Spreadsheet Analysis" toast.
Also drops "for this conversation" from the example reply: tool toggles at
/customize/tools are user-level preferences, not conversation-scoped.
Swept the rest of the tree for the same staleness: an admin placeholder
claiming the prompt description is shown in the conversation settings panel
(it is shown nowhere today) and five docstrings naming model-settings, the
drawer, or "the My Skills page".
Note: the code prompt is only the fallback — seed_bootstrap_data.py seeds no
prompt text, so a deployed prompt edited through admin CRUD keeps its old
copy. Deployed environments need checking separately.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(session): extract DockedPaneService from ArtifactStateService
The layout reserves exactly one right-side gutter (`--artifact-pane-width`
plus the `artifact-pane-open` class, read in `app.ts` and seven places in
`chat-container.component.html`), so two docked panes open at once would
overlap and the chat column would be sized for whichever service answered
first. With a second pane arriving, the rail's width and the side-nav
choreography stop being artifact concerns.
`DockedPaneService` now owns the rail: its width, the collapse/restore of
the side nav, and which feature holds it. Mutual exclusion is enforced
there rather than trusted to callers — `claim()` hands the rail over and
implicitly evicts the current holder.
Evicted owners are not called back. Each gates its public open-ref on
`owner()` instead, so eviction is a read rather than a notification. That
keeps the dependency one-way and rules out the write-loop an effect-based
handoff would invite.
`ArtifactStateService`'s public API is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): preview .docx in a docked pane
Adds a Preview button to the inline download card for `.docx` files,
opening a right-docked pane that renders the document in the browser.
Rendering is client-side via `docx-preview`, which walks the OOXML and
reproduces the document's own page layout, styles, tables and numbering.
The reference implementation we looked at frames
`view.officeapps.live.com` instead, which renders server-side at Microsoft
and therefore needs the document reachable by an unauthenticated URL — not
something to do with university documents. Nothing leaves the browser here.
No backend change was needed: `/files/{id}/preview-url` is already
owner-scoped and READY-gated and reports the stored MIME type, and the
user-files bucket already allows CORS GET from the SPA origin. The S3 leg
deliberately uses plain `fetch` with `credentials: 'omit'` — S3 answers a
CORS GET without `Access-Control-Allow-Credentials`, and routing it through
HttpClient would hand S3's own 403s to the global error interceptor, which
treats an auth failure as a reason to bounce the user to login. A presigned
URL going stale is a retry, not a logout.
Notes on the renderer:
- Loaded with a dynamic `import()`. It and its jszip dependency are dead
weight in the initial bundle for the majority of sessions that never open
a Word document (~284 kB, now its own lazy chunk), and it touches
`document` at module scope, so a static import would run during SSR.
- Each render targets detached containers that are swapped in on success.
`renderAsync` captures the elements it is handed and appends to them
asynchronously, so pointing two overlapping renders at the live host
interleaves their output, and a sequence guard cannot undo DOM the
library wrote itself. Fresh containers also stop the injected stylesheet
accumulating a copy per render.
- Pages are scaled to fit with CSS `zoom` rather than `transform: scale()`,
because zoom participates in layout: the flow collapses to the scaled
height instead of reserving the unscaled height and leaving dead space
under every page. Natural page width is read from the inline `width:
612pt` the library stamps, since measuring it back from a zoomed layout
would feed the fit calculation its own output.
- `applyTableConditionalClasses()` re-tags rows and cells after render.
docx-preview puts Word's `tblLook` flags on the `<table>` while its own
emitted CSS targets rows and cells, so its rules for bold headers, first
columns and row banding can never match and tables render flat. This
restores only the hooks its per-style rules already target — it invents
no formatting, and a style with no rule for a band changes nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(word-tool): document page breaks and header bolding
python-docx does no layout, so a generated document contains no page
boundaries at all unless the code adds them, and it never writes the
`lastRenderedPageBreak` hints Word stamps on save. A "three page report"
therefore came out as one continuous run of text — wrong in Word itself,
not just in a preview. The docstring never mentioned `add_page_break()`,
so the model never emitted one.
Header emphasis has the same shape. Setting `table.style` alone leaves the
header row to the style's *conditional* formatting, which not every viewer
applies; direct run formatting always renders. The bolding loop is folded
into the table example rather than added as a separate note, so it travels
with the code the model copies.
This docstring is part of the cacheable `toolConfig` prefix, so it costs a
one-time cache re-write per session on the next turn after deploy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(excel-tool): document freeze panes, number formats and formula caching
Audited 17 sheets across four real agent-generated workbooks pulled from
the dev files store. Three gaps, all traceable to the docstring:
freeze_panes: set on 0 of 17 sheets. Every generated workbook scrolls its
header row off the top. The docstring never mentions it.
number_format: 16 cells across all 17 sheets carry anything but 'General',
and two of the four workbooks have none at all. Currency renders as
1234567.891 and rates as 0.1834. Also never mentioned.
Formulas are written with no cached result -- openpyxl does not evaluate
them. Excel fills them in on open, so a human never sees the problem, but
until then the cell is empty to every other reader: all four formulas in
the corpus read back as None under `data_only=True`. That blinds
read_excel_spreadsheet to totals the agent itself just wrote, and would
show a blank cell in the in-app preview pane.
Deliberately NOT added: column widths, fills, alignment and borders. The
audit shows the model already emits all four heavily (461 fill / 1150
alignment / 1057 border cells) with no example to copy, so documenting
them would be prefix cost for no behavior change. The three above are the
ones it never does unprompted.
Guidance is folded into the example the model copies rather than added as
separate notes, following the word-tool fix in 6498ad77.
This docstring is part of the cacheable `toolConfig` prefix, so it costs a
one-time cache re-write per session on the next turn after deploy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(pptx-tool): document speaker notes on the create tool
read_powerpoint_presentation extracts speaker notes, but
create_powerpoint_presentation never says how to write them -- the
toolset could read a capability it could not produce. Across four real
decks from the dev files store (27 slides), not one slide has notes.
The design guidance already pushes for sparse slides; without notes that
detail is simply lost rather than moved to where a presenter would say
it.
Same cacheable-prefix cost note as the Excel and Word docstring fixes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(deps): add pptx-preview 1.0.7 with echarts stubbed out
pptx-preview reads OOXML and reproduces each slide's own geometry,
including the theme inherited from slide masters and layouts — which is
what preserves branding on decks built from an uploaded template. Tested
against every .pptx in the dev files store: three real agent-generated
decks render correctly, including a 4.3 MB branded template deck with 15
images in 122 ms.
It reaches for ECharts in exactly one place, rendering a *native* OOXML
chart part, through a static namespace import that cannot be tree-shaken.
Our decks never contain one — create_powerpoint_presentation embeds
charts as matplotlib PNGs via add_picture — and the survey found zero
chart parts across all four decks. Shipping ECharts anyway would be
337 kB gzipped, five sixths of the viewer's weight, and every release
below 6.1.0 carries GHSA-fgmj-fm8m-jvvx.
So `echarts` resolves to a local stub instead: the lazy chunk is 48 kB
gzipped rather than 357 kB, and `npm audit` returns to its exact
pre-existing baseline of 11 findings — pptx-preview contributes none. The
cost is that an uploaded deck containing a native chart throws, and the
viewer reports it as unreadable. shims/echarts-stub/README.md records the
trade and why a top-level dependency is the mechanism that works
(tsconfig `paths` does not reach a dependency's own imports, and a
`file:` spec nested under `overrides` resolves relative to wherever npm
places the package).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): preview .pptx in the docked pane
The pane was already format-agnostic apart from the renderer, so this is
a viewer component plus a widened gate: previewKindFor() replaces the
docx-only predicate and drives both the viewer switch and the header's
subtitle, which was hardcoded to "Word document".
The MIME check now cross-references the extension rather than comparing
against a single constant. The extension picks the viewer before any
request is made, so a file named .pptx that the server records as a
.docx has to fail in the service rather than reach a renderer that
cannot read it.
Two things differ from the docx viewer, both forced by the library:
pptx-preview takes a fixed pixel size at init() and has no responsive
mode, so the deck is rendered once at 960px and CSS-zoomed to fit. A rail
drag re-fits without re-parsing. Same zoom-not-transform reasoning as the
docx viewer.
It also writes background, height and overflow as *inline* styles on its
wrapper — a black backdrop and a nested scroller — so those three
overrides carry !important. Without them the deck sits in a black
letterbox inside a second scrollbar, which is what it did until this was
caught in the browser.
The zero-slide guard is the one piece of behaviour with no counterpart in
the docx viewer. pptx-preview resolves *successfully* when it cannot make
sense of a presentation's theme or layout parts, returning no slides —
reproduced against PR873-Verification-Deck.pptx in the dev files store,
whose theme is stripped to 1.8 KB. Unchecked that paints an empty pane
with no error and no retry, which reads as the app being broken rather
than the file being unreadable.
.xlsx is deliberately still excluded, and file-preview.model.ts records
why: the npm build of SheetJS is frozen at a 2022 release with unfixed
advisories, and the only maintained grid renderer is built on ExcelJS,
which throws outright on a workbook containing a native chart — which is
what create_excel_spreadsheet's own documented example produces.
Verified in the browser against all four real decks: fidelity, fit at two
rail widths with no horizontal overflow, resize re-fit, the failure path,
and dark mode.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): preview an uploaded .docx or .pptx, not just a generated one
The docked pane's only entry point was the Preview button on the
generated-file download card, so a deck the platform created previewed and
an identical deck the user uploaded did not. The uploaded one fell through
to the presigned-URL branch, which hands the browser an OOXML file it
cannot render — "open" silently became "download", and the agent's own
fallback was to read the file and describe it in prose.
Attachment cards now route the formats the pane can render to the pane,
keyed off `isPreviewableFilename` — the same gate the download card uses,
so the two surfaces cannot disagree about what is previewable. Markdown
keeps its modal and everything else still opens in a tab.
The button's accessible name follows the action rather than stating "Open"
for something that previews.
Verified against a real uploaded deck in dev: 10 slides, branding intact.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(prompts): point "show me this file" at the viewer, and stop escaping $ into files
Two clauses, both prompted by watching a real session.
Asked to preview an uploaded deck, the agent read it, and when pushed for a
*visual* preview it called create_powerpoint_presentation and rebuilt the
deck from scratch — a lossy copy of a file already sitting in the session,
purely to obtain a Preview button, which only generated files had. It was
not being obtuse: the docked pane has exactly two callers, both user clicks,
and no SSE event can open it, so the model had no way to show anything. Its
only verbs were read and create.
The pane now also opens on uploaded files, so the prompt just has to say the
viewer exists and belongs to the user. It draws the line at "show me" (point
at the button) versus "tell me about" (read the file), because reading is
right for summarize/check/edit and wrong for looking — a text dump is not
what was asked for, and it puts the whole document in the cacheable history
where it is paid for on every later turn. That session had spent 9,053
characters of tool results on it.
The second clause fixes a systematic defect this work surfaced. The KaTeX
guidance said "other uses of $ should be use the HTML entity $" with no
scope, so the model applied it inside generated files: a deck came out
holding the literal string "$100K", which is wrong in PowerPoint too,
not just in our preview. I first wrote this off as noise because none of the
eight real files in dev contained the entity — but none of them contain a
dollar sign at all, so the corpus never had the chance to show it. Every
file that did carry currency reproduced it. The rule is now scoped to chat
markdown and explicitly excluded from files, code and tool arguments.
Verified against the running stack with inference-api pointed at this branch:
"Show me the preview-pane-test deck" now answers "Click the Preview button"
with zero tool calls, and a deck asked for with dollar amounts comes out
holding $100K, not $100K.
Caveat worth knowing: this steers new conversations. In the session that had
already read-then-created four times, Haiku followed its own precedent and
rebuilt the deck anyway.
Base prompt grows 6,244 -> 7,695 chars (~362 tokens), so it costs a one-time
cache re-write per session on the next turn after deploy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(prompts): trim the preview clause to the part that prevents waste
The gap that mattered — previewing an UPLOADED deck — is fixed in the UI
(9ab491e2); it needs no prompt at all. What the prompt still has to prevent
is the model reading a file, or re-creating it, just to show it to someone.
Cut the explanation of what the viewer is, the uploaded-vs-generated aside
and the .xlsx exception, and kept the rule. 362 -> 169 tokens of cacheable
prefix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(files): make the attachment card say it previews
The card is the whole click target — there is no separate button — and the
only hint that clicking did anything was an "open in new tab" glyph that
faded in on hover. So at rest it advertised nothing, on a touch screen it
advertised nothing ever, and the glyph it did show described the wrong
action once .docx/.pptx started opening a side pane instead of a tab.
Previewable types now carry a persistent eye + "PREVIEW" in the header
strip, the same pairing the generated file's download card uses so the two
read as one feature. Other types keep the quieter hover hint.
Two contrast fixes found by measuring rather than eyeballing: the label
started at gray-500, which is 4.55:1 on the pptx header tint and scrapes
past AA by a hair, now gray-600 at 7.12:1; and the hover state used
primary-accessible, which is BSU navy and lands at 1.41:1 in dark mode --
effectively invisible. Hover is now a neutral emphasis that measures
14.92:1 there.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(files): colour the file-type icon, and match PREVIEW to it
The chip icon rendered black in both themes while the label beside it,
carrying the identical class, came out in the file type's accent. Cause:
a text-<colour> utility placed directly on an <ng-icon> never paints it —
the component's own host rule outranks a plain class. It does inherit, so
the accent now sits on the wrapper and reaches icon and label together.
⚠️ This is not confined to this card. Sampling the live DOM, none of the
ng-icons carrying a colour class were getting it: "text-gray-400" rendered
black, "text-gray-500" rendered white. They look right only where the
inherited colour happens to be right, and there are ~449 such call sites.
Out of scope here, but it is a real platform-wide bug.
PREVIEW now takes the same accent so the header strip reads as one unit,
at the same weight and opacity as the type chip rather than a dimmed
version — measured on the live DOM, the accent is already 3.37:1 on its
own light-mode tint, so anything held further back would be worse than a
label that is itself under AA.
⚠️ Also pre-existing and worth its own fix: that 3.37:1 is the type chip's
own contrast in light mode, below the 4.5:1 AA floor for text this size,
across every file type. Dark mode measures 8.60:1 and is fine. The likely
fix is a -700 step for the light-mode accent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(a11y): raise file-type chip text to -700 so it meets AA in light mode
The chip accents were set one step too light for the tints they sit on.
Measured off the rendered DOM rather than the palette, on the chips'
own light-mode grounds:
family -600 -700
pdf 4.12 5.49
doc 4.82 6.28
sheet 3.08 4.72
markdown 5.16 6.58
image 5.78 7.24
code 3.37 4.92
presentation 3.37 4.92
Four of seven were under the 4.5:1 floor for text this size, including
both orange families used by the PPTX and HTML chips. Every family clears
it at -700, and the colour still reads as the file type. Dark mode was
never affected (8.60:1 on the pptx chip) and is untouched.
Applied to the five surfaces where the token colours chip *text*: the
attachment card, the generated-file download badge, the file card, the
artifact library and the shared artifact card. file-browser.page.ts keeps
-600 — there the token tints an icon, which answers to the 3:1 non-text
bar, and per the previous commit a colour class on an <ng-icon> does not
apply at all.
One caveat: the artifact-library and file-card chips sit on a -100 ground
rather than -50, where sheet lands at 4.4957:1 — short of 4.5 by 0.004,
inside sRGB rounding, and reported as 4.5 by tools that round. Every other
family there clears 5.0. Worth a look if that chip ever gets restyled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): open the preview pane on a file the turn just created
Artifacts pop their panel from onArtifact; a generated .docx or .pptx left
the user to find the Preview button themselves. Now the pane opens on it,
which is the whole point of having a viewer.
The hook hangs off tool_result rather than an SSE event of its own,
because the office tools do not have one — the download card is a
file_download inline visual inside the tool result. That placement also
gives the live-vs-hydrated distinction for free: tool_result only ever
arrives mid-stream, so reopening an old conversation replays the card
without seizing the rail, matching seedFromHydration on the artifact side.
Verified both ways in the browser.
It runs before the tool_use block lookup, not after, so an unmatched
result still surfaces its file — the block bookkeeping failing does not
mean the file is absent. Viewed-session only, the same guard as
onArtifact: a conversation streaming in the background must never take
the rail from the thread being read.
Formats the pane cannot render (.xlsx) are skipped rather than opening a
pane that could only show an error, and a legacy card carrying no
upload_id is ignored.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* perf(frontend): lazy-load mermaid out of the eager scripts bundle
mermaid.min.js was listed in angular.json's `scripts` array, which Angular
emits as a plain <script> tag on index.html — outside the module graph, so
lazy routes, @defer and module-graph bundle analyzers all miss it
structurally. At 3.57 MB raw / ~1 MB gzip it was 92% of the scripts bundle
and 71% of the whole 5.02 MB initial bundle, downloaded by every visitor on
every cold load, for a feature that only renders when a message contains a
mermaid fence. The initial budget sat at 5 MB and quietly absorbed it until
the app drifted 23 kB past.
ngx-markdown resolves mermaid off the global scope and throws when it is
absent, so dropping the script tag alone breaks every markdown render.
installLazyMermaid() publishes a stand-in carrying only the two methods
ngx-markdown calls: initialize() records the config, and run() — reached
only once the rendered DOM already contains a .mermaid element —
dynamically imports the real library, replays the config, and swaps itself
out for the real module so later renders go straight through.
scripts bundle 3.89 MB raw / 1.07 MB gzip -> 323 kB raw / 95 kB gzip
initial total 5.02 MB raw / 1.02 MB xfer -> 1.45 MB raw / 326 kB xfer
mermaid now lands in 68 lazy chunks split per diagram type, so a flowchart
no longer drags in gantt/sankey/architecture. Its CommonJS deps (dayjs,
cytoscape layouts, sanitize-url) are allowlisted so the lazy chunks build
without warnings.
Retarget the initial budget to 1.6 MB warning / 2 MB error, close to the
real footprint, so the next multi-megabyte dependency warns on the PR that
adds it. Angular budgets have no transfer-size mode, so raw bytes stay the
lever. KaTeX and Prism are knowingly left eager: ngx-markdown calls both
synchronously in the render pass that inserts the HTML, so deferring them
would flash raw math and unhighlighted code on nearly every message.
Verified in a real browser against the dev server — the stand-in passes
ngx-markdown's guard, the import resolves, the real API replaces the
stand-in on the global, and a flowchart renders as SVG.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(a11y): move success-state text to the AA-passing ramp step
`text-state-success-600` fails WCAG AA for normal-size text in light mode
(3.22:1 on white, 3.08:1 on gray-50). The green ramp is far lighter than
red at the same step, so the step that clears 4.5:1 is not the same for
every state role. Move text call sites to `-700` (4.95:1) and keep the
dark step at `-400` (9.98:1) — `-700` is a light-mode step only, failing
at 3.59:1 in dark.
19 of the 43 `text-state-success-600` sites move. The other 24 stay at
-600 deliberately: aria-hidden glyphs, copy-button check icons, and the
role-form checkbox fill are non-text content governed by the 3:1 rule
(WCAG 1.4.11), which -600 already clears on white.
`user-detail.page.ts` had no dark variant at all — -600 happened to pass
dark at 5.51:1, so bare -700 would have regressed it. It gains an
explicit `dark:text-state-success-400`.
Two further failures found by measurement rather than by the -600 grep:
- A -600 glyph on its own -100 tint misses even the 3:1 floor (success
2.93, warning 2.87). The composer's voice button is the only place
that stacks them, so all three branches of `voiceButtonClass` move to
-700 (success 4.50, warning 4.52, danger 5.27). Danger moves with
them although 3.91 already passed: the three branches are one visual
system, and a thin margin over a 3:1 floor is not worth preserving
for the sake of leaving one branch untouched.
- A -600 icon on a gray-100 HOVER fill is 2.92. The two icon buttons
that pair `hover:text-state-success-600` with `hover:bg-gray-100`
move to -700 (4.49). Icon buttons that only recolor on hover, over
the page background, stay at -600.
All values measured in-browser with transitions and animations disabled
and oklch() resolved to sRGB through a canvas, with the converter
sanity-checked against black/white (21.00) and #595959/white (7.00).
The full table is recorded in the state.css header comment.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(compaction): model-relative thresholds, floor-seeking cut, hysteresis (PR-1)
Compaction fired at a fixed 100k tokens regardless of the model's window,
chose its cut by turn count, re-cut on every over-threshold turn, and
compared a slice-relative index against the persisted absolute checkpoint.
Underneath it, Strands' default 40-message SlidingWindowConversationManager
(AgentFactory passed none) slid the front of agent.messages every turn past
40 messages: a prefix re-write per turn and the D3 coordinate mismatch.
Spec: docs/specs/compaction-model-relative-thresholds.md (rides with this PR).
- CompactionPolicy.resolve(): ceiling = min(0.5 x window, 100k), floor =
0.25 x ceiling, hard ceiling = min(0.7 x window, 1.5 x ceiling), from the
catalog's maxInputTokens (already looked up per turn for the badge).
The window scales the ceiling DOWN, never up: the 2026-09-15 replay of 20
heavy Sonnet 5 sessions priced 200k/50k 43% above 100k/25k on the input
side, because cold re-writes after a >5 min pause scale with the context
at the pause. Unknown window -> the fixed threshold; kill switch
AGENTCORE_MEMORY_COMPACTION_MODEL_RELATIVE_ENABLED=false -> legacy exactly.
- choose_checkpoint(): oldest tool-pair-safe cut whose retained estimate is
at or under the floor (min protected_turns kept); per-message estimates
calibrated to the context-breakdown `messages` partition.
- Hysteresis: CompactionState.armed. A cut disarms; a turn under the
ceiling re-arms; only the hard ceiling forces a cut while disarmed
(logged compaction_forced = the previous cut did not take). The spiral's
56 consecutive cuts become 1 cut + 55 no-ops.
- One coordinate system: _live_offset (absolute index of agent.messages[0],
set by the restore slice); persisted checkpoint = offset + relative cut.
- AgentFactory.build_conversation_manager(): explicit
SlidingWindowConversationManager(window_size=2000, should_truncate_results=True)
(AGENTCORE_CONVERSATION_WINDOW_MESSAGES; 40 restores the SDK default).
Kept rather than Null because its reduce_context is the only
ContextWindowOverflow recovery in the stack, independent of window size.
Consequence: conversations between 40 messages and the ceiling now go to
the model whole (0.1x reads instead of 1.25x re-writes).
- `compaction` SSE payload + CompactionResult carry the policy fields
(additive; SPA validator ignores extras, TS interface gains optionals);
the persisted compaction map records `armed` and a `policy` snapshot.
No change to WHEN history bytes change: the slice still applies at restore.
Bounding the summary is spiral-spec PR-2 (next); in-place apply on warm
agents with paid-when-free scheduling is PR-3.
Tests: policy table, kill switch, estimator, floor-seeking cut, disarm /
forced / re-arm, spiral shape cuts exactly once, live-offset coordinates,
conversation window default + override; existing byte-stability suite
unchanged. Backend suite: 8622 passed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(agent): drop the hour from the system-prompt date line to keep the prompt-cache prefix day-stable
get_current_date_pacific() rendered "YYYY-MM-DD (Weekday) HH:00 TZ" and
SystemPromptBuilder puts it at the tail of the system prompt, which is the
head of the Bedrock prompt-cache prefix. Every Pacific hour boundary
therefore flipped systemPromptHash and re-wrote the whole cached prefix
for every active session. The 2026-09-15 prod cost audit measured 23
distinct systemPromptHash values in one 85-call session and attributed
~2.6% of September cache-write spend to this.
The line now renders date, weekday and timezone only, so it is byte-stable
for a whole Pacific day. Nothing in the prompt, skills, local tools or the
BFF reads the hour (grepped for the helper, the %H:00 format and
"current time"-style language), so no flow needed the hour moved into the
user turn. Function name and signature are unchanged; the docstring and a
comment at the render site record why the hour must stay out of the prefix.
Tests: test_timezone.py gets the new regex plus frozen-clock tests that the
string is identical across all hours of a day, flips at Pacific midnight,
and uses the Pacific (not UTC) date; test_system_prompt_builder.py's mocked
return values lose the hour. test_bedrock_cache_points.py is unchanged and
still passes; the full tests/agents/main_agent suite is green.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(admin-costs): list a user's deleted conversations in the cost drill-down
A soft delete flips `status` and drops the sidebar GSI keys, but the
session's `C#` cost rows and its share of the user's period total survive
it. The per-user conversation list dropped those tombstones, so the audit
could not account for the spend it was built to explain: one prod user
showed a single $3.77 conversation against $20.32 of September cost.
Fleet-wide, 183 of 5,440 September-active sessions (3.4%) are deleted and
carry $58 of $817.
The storage reader gains `include_deleted` (default off, so callers that
list what the user can still open are unchanged); the admin service asks
for tombstones, normalises legacy `deleted` rows to `status="deleted"`,
and reports `deletedSessionCount` / `deletedSessionCost` on the response.
The SPA badges deleted rows and appends "N deleted ($X still counted in
the total)" to the summary line. Profile and anatomy still resolve for a
deleted session because `SessionLookupIndex` keys are kept on the
tombstone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(compaction): "forced" means a cut that ran while disarmed
An armed cut whose input happened to exceed the hard ceiling was being
tagged forced=True. The flag is the spiral signal ("the previous cut did not
take"), so it must only fire for a cut that ran while disarmed because the
hard ceiling was reached. No behavior change to when cuts happen; the
persisted policy snapshot and the SSE field now say what the spec says.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(spa): reset DocxViewerComponent's memoized import between docx specs
CI's "Test frontend (coverage build)" failed twice on this PR, which
touches no frontend file, in file-preview-panel.component.spec.ts:
"keeps download reachable when the document cannot be rendered" received
a fully rendered document instead of the failure message.
DocxViewerComponent memoizes its dynamic import('docx-preview') on a
private static field, and @angular/build runs vitest with isolate: false,
so every spec file in a worker shares one module registry and one
DocxViewerComponent class. When docx-viewer.component.spec.ts runs first
in the worker it pins its own mocked module into that field; the panel
spec's renderAsync.mockRejectedValue() then programs a vi.fn the
component never calls. The received text ("QuarterRev Q1100 Q2120") is
the viewer spec's last table mock, byte for byte.
Both specs now null the memo in beforeEach. Reproduced and verified with
a single-worker, alphabetically ordered runner config (viewer spec before
panel spec), then the full coverage build: 257 files / 3212 tests pass.
Develop's exact frontend tree was never exercised by this job: CI runs
only on PRs, the docx specs landed in 76e022a3, and later PRs that passed
all add or change spec files, which reshuffles worker assignment.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(cost-diagnostics): per-call context ledger — prefix split, window trims, compaction events
Three content-free facts the 2026-09-15 prod cost audit had to reconstruct
by scanning and guessing, now stored on each model call's `C#` row:
- `prefixTokens: {system, tools}` — the agent's stable static prefix, from
the split the context-attribution hook already computes. The fleet
question "how big is the prefix and how much of it is tool schemas" (19%
of September spend; 46% for 16+-tool users) becomes a query.
- `windowRemovedMessages` — the conversation manager's cumulative
`removed_message_count` at the call. A rise between consecutive rows is a
trim (the sliding window or a compaction slice), which re-writes the
cached prefix. Pure-window sessions and compaction spirals took an hour
of fingerprint reading to tell apart; the anatomy now marks the call.
- `compactionEvents: [{kind, checkpoint, summaryTokens, ...}]` — the
decisions the session manager took since the previous call, each with
the summary's token size at that moment: `applied` (restore-time slice
ran), `checkpoint` (a new checkpoint was cut), and `forced` /
`floor_unreachable` reserved for the scheduling policy (#1125), which
calls `TurnBasedSessionManager.record_compaction_event(kind, **ints)`.
This is what will show a summary cap shrinking summaries and a
scheduling rule stopping re-writes without another table scan.
`ContextLedgerHook` mirrors the tool census: per-turn, per-model-call, read
(never drained) at turn end, off with `COST_DIAGNOSTICS_ENABLED=false` so an
absent field reads "not tracked", never 0. Session rows gain
`compactionAppliedCount` / `compactionForcedCount` /
`compactionFloorUnreachableCount` via atomic ADD; `checkpoint` stays on
`compactionCount`, which `_save_compaction_state(record_event=True)`
already bumps, so one event never feeds two counters.
The admin API projects the fields (content policy widened; numbers only),
derives `windowTrimmed` per row so readers need not diff rows, and the
profile reports the latest prefix split, trim calls, event counts by kind
and the last summary size, with `dataCoverage` flags for each. The anatomy
page adds a static-prefix tile (system · tools), a window-trims tile, the
compaction kinds with the latest summary size on the compactions tile, and
per-row `trim −N` / event badges with the numbers in the expanded row.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(compaction): hand forced / floor-unreachable decisions to the compaction ledger
The cost-diagnostics per-call ledger (record_compaction_event →
ContextLedgerHook → `compactionEvents` on the next C# cost row) is landing
on develop separately. Record the two PR-1 decisions the anatomy needs to
show — a cut that ran while disarmed ("forced") and a cut whose protected
tail alone exceeds the floor ("floor_unreachable") — through a small
helper that resolves the recorder by attribute, so it is a no-op on a build
without the ledger and activates when it merges. Int fields only, per the
ledger contract.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(compaction): bound the summary at 8k tokens and emit per-cut metrics (PR-2)
The compaction summary was an unbounded join of AgentCore LTM
ConversationSummary records (165k chars / ~40k tokens in the #833 incident):
a summary that is 40% of the threshold guarantees compaction can never get
back under it. Spiral-spec PR-2; thresholds spec §3.6 / §7.1.
- compaction_summary.bound_summary(): hold the persisted summary at
COMPACTION_SUMMARY_TOKEN_BUDGET (8,000 tokens, chars/4 — the same estimate
the admin SUMMARY_OVER_BUDGET diagnosis uses). Within budget → unchanged.
Over budget → one Nova Micro converse call (side-channel, never touches
agent.messages) with a prompt that keeps standing instructions, decisions,
current state of the work, open items and exact identifiers, and drops
narration and superseded drafts. Model failure, a ceiling-hit generation
or an overshoot → newest-first truncation (keep the newest records that
fit; if none fit, the tail of the newest). Runs once at checkpoint advance
— the turn that already pays a prefix re-write — and the result is
persisted verbatim, so the byte-stability contract is unchanged.
- Kill switch AGENTCORE_MEMORY_COMPACTION_SUMMARY_MODEL_ENABLED=false skips
the model and truncates; AGENTCORE_MEMORY_COMPACTION_SUMMARY_MODEL_ID
selects the model.
- Provenance on the persisted compaction.policy map: summarySource
(ltm|fallback), summaryOutcome, summaryTokensBefore/After,
summaryTokenBudget.
- One content-free EMF record per cut in AgentCoreStack/Compaction:
CompactionCut, CompactionForced, CompactionInputTokens,
CompactionRetainedTokens, CompactionSummaryTokens,
CompactionSummaryOverBudget, with policySource/window/ceiling/floor/
summaryOutcome as queryable properties. Silenced by
PROMPT_CACHE_OBSERVABILITY_ENABLED=false with the rest of the layer.
- forced flag narrowed to "ran while disarmed" (same hunk as the PR-1 fix).
Tests: newest-first truncation, within-budget passthrough, model
compression, model failure / ceiling / overshoot fallbacks, kill switch,
env loading, oversized LTM join bounded and persisted through
update_after_turn, EMF record shape and kill switch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(compaction): park cuts post-turn, apply in place when the re-write is free (PR-3)
Compaction decided WHAT to cut (PR-1) and bounded the summary (PR-2), but a
cut still landed at the next restore regardless of whether the prompt-cache
prefix was warm, and never landed at all on a warm (cached) agent. Under
Bedrock caching a cut costs one re-write of what survives, so the cheapest
turn to pay it on is one that was going to re-write anyway.
Thresholds spec §3.5.
- update_after_turn PARKS the cut: pendingCheckpoint / pendingSummary /
pendingHardCeiling / pendingSince on CompactionState. `checkpoint` stays
the APPLIED value (what _apply_compaction slices at on restore). A second
over-ceiling turn while a cut is parked is a no-op, never a deeper cut.
- apply_pending_compaction(agent, prefix_key) runs at the head of every
turn (stream coordinator, right after the turn lease is stamped), on
cached and freshly restored agents alike. It promotes the pending cut and
slices agent.messages IN PLACE (slice assignment, never rebinding — the
#741 alias) when, in order: cache_expired (more than cache_ttl_seconds
since the previous turn), prefix_changed (model|agent key differs from
the persisted lastPrefixKey), or hard_ceiling (previous input reached
the hard ceiling the cut was computed under). Otherwise it waits.
- The in-place result is byte-identical to what _apply_compaction derives
from stored history under the promoted state (pinned by
test_live_apply_matches_a_cold_restore_of_the_same_state), so a cold
restore after a live apply reads the same prefix.
- _adopt_session_conversation copies _live_offset when it points a new
agent at the live list — the list's coordinate system travels with it.
- Each application persists the reason, cacheGapSeconds and pendingSince
on compaction.policy, logs rewrite_scheduled vs rewrite_forced, and emits
CompactionApplied / CompactionAppliedForced / CompactionCacheGapSeconds;
the cut record gains Compacti…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Runs the backend pytest suite (~3k tests across 469 files) with
pytest-xdist -n autoon the PR gate (tests.yml), so it fans across all runner cores instead of running single-threaded — the backend job is the PR long pole, and the same suite re-runs as the deploy gate, so both get faster.Changes
backend/pyproject.toml: addpytest-xdist==3.6.1to dev deps (+ regenerateduv.lock).github/workflows/tests.yml: backend job nowuv run pytest tests/ -n autobackend/pytest.ini: drop-vfrom addopts (thousands of PASSED lines, no diagnostic value)Deliberately NOT changed
scripts/backend/test.sh) stays serial — xdist + coverage combining is a separate change with its own edge cases.Verified
created: 2/2 workers); targetedtests/costs(218) andtests/security(167) green under-n auto. Full suite runs on CI.