From fb104e699426459b6a81b217997d30241efff9f3 Mon Sep 17 00:00:00 2001 From: Phil Merrell Date: Fri, 12 Jun 2026 08:55:17 -0600 Subject: [PATCH] chore(kaizen): weekly research scan 2026-06-12 Generated by the kaizen-research skill. Top 5 ideas appended to docs/kaizen/review-queue.md for the kaizen-review-prep run later this morning. Co-Authored-By: Claude Sonnet 4.6 --- docs/kaizen/research/2026-06-12.md | 243 +++++++++++++++++++++++++++++ docs/kaizen/review-queue.md | 38 +++++ 2 files changed, 281 insertions(+) create mode 100644 docs/kaizen/research/2026-06-12.md diff --git a/docs/kaizen/research/2026-06-12.md b/docs/kaizen/research/2026-06-12.md new file mode 100644 index 000000000..0fcce2953 --- /dev/null +++ b/docs/kaizen/research/2026-06-12.md @@ -0,0 +1,243 @@ +# Kaizen Research — Friday, June 12, 2026 +> Scan window: June 5 – June 12, 2026 (7 days) +> Web budget: ~82/50 used (overage — version-pin subagent used 25 of 82; see Web Budget block). + +## TL;DR + +**Strands 1.43.0 dropped today (June 12)** — the count_tokens/toolResult=0 bug (#2635) we've been tracking since it landed is now fixed; the queue's [2026-06-05] keystone bump target advances from 1.42 → 1.43 with no additional blast radius. **Claude Fable 5 went GA on June 9** (`claude-fable-5`, $10/$50 per M) and introduces a new naming convention (`-fable-`, `-mythos-` suffixes) that breaks any code path doing string-matching on `claude-opus-4.*`. The week's #1 recommendation is the **Strands 1.40 → 1.43 keystone bump** — it now closes the tracked attribution bug and picks up a context_manager="auto" simplification opportunity in the same PR. + +## External Scan + +### What's moving this week + +Two heavyweight externals converged in a single week. **Strands 1.43.0** (released today) closes the most operationally urgent open item — issue #2635 (`count_tokens` returning 0 for toolResult JSON blocks), which threatened the context-attribution badge shipped last sprint. That bug is now gone at version 1.43; the prior 1.42 keystone entry in the queue can simply be retargeted. **Claude Fable 5 + Mythos 5** (GA June 9, `claude-fable-5`, $10/$50 per M) introduces an Anthropic naming pattern shift — no more `claude-opus-4.N` dot-versioning — that is a real breakage risk in any model-ID string matching code. At $10/$50 it is significantly cheaper than prior Opus-class models, which changes cost-planning assumptions for skills/compaction budget. + +Three secondary signals worth tracking: **AgentCore Runtime Interactive Shells** (June 5) adds a WebSocket terminal-access API into running microVMs — a native code-execution path that is a strategic alternative to the queued BYO-filesystem proposal. **FastMCP v3.4.1** (June 5) pinned starlette ≥1.0.1 to close CVE-2026-48710; our own `pyproject.toml` pins `starlette==1.0.0` which is in the affected range. **Agent-EvalKit** (June 11) shipped a systematic evaluation harness specifically for Strands Agents — the skills-mode build-out that just landed has no automated correctness layer, and this is the first upstream tool that could provide one. + +Internally, the week was the most productive commit window since the MCP Apps build: the entire admin-skills system (PRs #1–7) merged, Gateway MCP target self-service (#419) fully landed, and skills-mode policy enforcement (PR-1 + PR-2) shipped — 27 commits. The shadow is a **7-consecutive-day Nightly Build & Test failure** (June 5–12) whose root cause is still unidentified. + +### Notable items by source + +> **Annotation conventions:** +> - `*relevance*:` — impact-on-existing-code lens. +> - `*unlocks*:` — capability-unlock lens (net-new product capability this enables). + +#### AWS Bedrock / AgentCore +- **AgentCore Runtime Interactive Shells (June 5)** — New `InvokeAgentRuntimeCommandShell` API adds persistent WebSocket terminal access into agent sessions running in isolated microVMs; coding agents can execute shell commands natively inside the Runtime container. — https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-bedrock-agentcore-runtime/ — *relevance*: our `inference-api` runs inside an AgentCore Runtime container; this is an alternative code-execution path to the queued BYO-filesystem/sandbox plan. — *unlocks*: native code-interpreter surface inside the already-deployed Runtime — no new container or VPC required. Pairs with the BYO-filesystem revisit (deferred to 2026-06-12 in the queue). +- **Claude Fable 5 GA on Amazon Bedrock (June 9)** — Fable 5 ("Mythos-class") is now GA on Amazon Bedrock and Claude Platform on AWS. — https://aws.amazon.com/about-aws/whats-new/2026/06/claude-fable-5-aws/ — *relevance*: model settings UI (`model-settings.html`, `model-settings.ts`) and any hardcoded model-ID strings need updating; `CountTokensBedrockModel` de-prefix logic should handle the new ID pattern already (the `us.anthropic.*` stripping is ID-agnostic), but string-match gates on `claude-opus-4` **do not match** `claude-fable-5`. — *unlocks*: top-of-range model available on Bedrock; $10/$50 per M changes per-session cost assumptions. +- **Agent-EvalKit — Systematic agent evaluation for Strands SDK (June 11)** — Open-source toolkit from AWS for evaluating Strands Agents-built agents across six evaluation phases (including tool-use correctness, multi-turn coherence, memory persistence), integrating with AI coding assistants. — https://aws.amazon.com/blogs/machine-learning/evaluate-ai-agents-systematically-with-agent-evalkit/ — *relevance*: we use Strands as our agent loop; skills-mode RBAC/tool-binding correctness has no automated eval layer. — *unlocks*: structured regression testing for skills-mode enforcement, per-skill tool-access coverage, and compaction/attribution scenarios. +- **AgentCore + Strands: Equipment Repair reference (June 10)** — Reference blog: AgentCore Runtime + Strands SDK + Knowledge Base for RAG + Memory for persistence. — https://aws.amazon.com/blogs/machine-learning/build-an-ai-powered-equipment-repair-assistant-using-amazon-bedrock-agentcore/ — *relevance*: same stack; the RAG+Memory wiring pattern may surface alternatives to current session-restore degradation. +- **Gemma 4 on Bedrock (June 10)** — Google DeepMind Gemma 4 family with native function calling added to Bedrock. — https://aws.amazon.com/about-aws/whats-new/2026/06/gemma-4-amazon-bedrock/ — *relevance*: low priority for core agent loop; potential cost-effective option for classification tasks in skills. No immediate action. + +#### Strands Agents +- **v1.43.0 released June 12 (today) — #2635 count_tokens fix shipped** — `count_tokens` returning 0 for toolResult JSON blocks (PR #2639 "include json blocks in counting tokens") is merged and shipped. Issue #2635 is now closed. Our pin (1.40.0) is now **3 minors behind**. — https://github.com/strands-agents/sdk-python/releases/tag/python%2Fv1.43.0 — *relevance*: **closes the tracked open threat** to the context-attribution badge (PR #428–433) and the compaction trigger; the queue item "Guard the context-attribution path against #2635" is resolvable by the 1.43 bump. +- **v1.43.0 — `context_manager="auto"` facade + message pinning** — Top-level `Agent(context_manager="auto")` selects a sliding-window context manager automatically; message pinning marks turns immovable so they survive eviction. — same release — *relevance*: `context_manager="auto"` is a potential simplification of our manual `TurnBasedSessionManager` compaction trigger path (note: decisions.md notes it is NOT a drop-in replacement for our full compaction — audit tool-truncation, LTM retrieval, DynamoDB checkpoint, and the SSE-once invariant before proposing a swap). *unlocks*: the pinning primitive could anchor system-prompt turns during compaction. +- **v1.43.0 — A2A per-context isolation fix** — Fixes A2A conversation state leaking between concurrent callers hitting the same agent (PR #2696/2628). — same release — *relevance*: directly relevant when the queued A2A server construct lands; our client-only stance today is not affected, but this is the fix to ensure we pick up before the first A2A server PR. +- **v1.43.0 — `model_state` snapshot field** — Adds `model_state` as a first-class snapshot field so model-level state survives checkpoints. — same release — *relevance*: may be useful for the skills-mode resume path (skills snapshot `agent_type` gate), which is actively in flight on `feature/skills-mode-spa`. +- **Bug #2636 still open — non-ASCII tool-result serialization** — PR #2661 is in review but not yet merged into 1.43.0. — https://github.com/strands-agents/sdk-python/issues/2636 — *relevance*: live in 1.43.0; a second bump will be needed once it merges; do not assume fixed on the 1.43 upgrade. + +#### Reference repo (aws-samples/sample-strands-agent-with-agentcore) +- **One commit in window — zero architectural signal** — Only `5c3c8820` (June 10): Dependabot vitest 3.2.4→3.2.6 bump in the chatbot-app frontend. Confirmed HEAD `ccfba7c` (May 22) still the last substantive commit. — https://github.com/aws-samples/sample-strands-agent-with-agentcore/commits/main — *applicability*: nothing to port this week. Side note: we are now one minor ahead of their vitest pin (we're at 4.1.5). + +#### MCP ecosystem +- **SEP-2624 — Interceptors: Middleware as Core MCP Primitive (updated June 11)** — Proposes `interceptors/list` + `interceptor/invoke`, defining Validators (pass/fail) and Mutators (payload transform) at deterministic lifecycle hooks (e.g., before `tools/call`). Still open, no formal reviews. — https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2624 — *relevance*: directly maps to our `FilteredMCPClient` / `list_tools_sync` filter layer and per-tool enablement (#469). If this advances to the 2026-07-28 RC, our tool-hiding UX and RBAC filter are natural host sites for Validator interceptors. Watch-track; no action yet. +- **PR #2889 — Schema alignment: `CancelledNotificationParams.requestId` becoming required** — Schema/spec divergence fixes; `requestId` in `notifications/cancelled` moving from optional to required in the RC. — https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2889 — *relevance*: our inference-api stream handlers emit/forward cancellation notifications; audit before the 2026-07-28 RC window. +- **PR #2907 — Error code renumbering policy (open June 12)** — Establishes formal policy for MCP error code assignment; renumbers all experimental/draft codes. — https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2907 — *relevance*: any hardcoded MCP error-code constants in `inference_api` or the frontend error-display path may need updating before the RC freeze. Low urgency now. +- **PR #2789 — Tool call attestation test vectors (v0)** — First test vector suite for the proposed SEP-2787 tool call attestation feature. Active review. — https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2789 — *relevance*: attestation adjacent to our A2A trust boundary; note for the RC window. +- **Servers — security dep bumps only** — `gitpython`/`urllib3` HIGH bump and an `everything`-server elicitation bug fix; no new servers. — https://github.com/modelcontextprotocol/servers/commits/main — *relevance*: the urllib3 HIGH bump is a reminder: check if `uv.lock` pulls in an affected urllib3 version before the next deploy. + +#### FastMCP +- **v3.4.1 (June 5) — Starlette ≥1.0.1 floor to block CVE-2026-48710** — Pins minimum starlette to ≥1.0.1 via the transitive `mcp` dependency path; adds OAuthProxy log for refresh-token cache misses. — https://github.com/jlowin/fastmcp/releases/tag/v3.4.1 — *implications for our MCP servers*: **our own `pyproject.toml` pins `starlette==1.0.0`** (in the affected range). Bump to 1.0.1+ in both this repo's backend and any external Lambda-backed MCP server repos. +- **v3.4.2 (June 6) — JWT private-header compatibility (e.g., Clerk `cat` field)** — Restores JWT validation for providers that inject private (non-critical) JWS header params. — https://github.com/jlowin/fastmcp/releases/tag/v3.4.2 — *implications*: relevant if any MCP server repos use Clerk or similar IdPs. Free patch on top of the CVE fix; bump to v3.4.2 in external MCP server repos. + +#### Agentic UI/UX patterns +- **`@mcp-ui/client` + AppBridge now in the MCP Apps official docs** — The overview page now calls out `@mcp-ui/client` (React rendering components) and the **AppBridge** SDK module (sandboxed iframe rendering, postMessage proxying, CSP enforcement). A `basic-host` example repo is linked at https://mcpui.dev/. — https://modelcontextprotocol.io/extensions/apps/overview — *fit*: our custom AppBridge implementation is the same protocol; worth reading `@mcp-ui/client`'s source for postMessage protocol additions we may have missed. — *where it lands*: `frontend/ai.client/src/app/session/` MCP App iframe bridge. +- **react-streamdown@0.3.3 — deferred markdown parsing + smooth prop** — The `defer` prop delays markdown parsing until the stream completes (avoids mid-stream re-renders on partial tokens); `smooth` adds typewriter reveal on top. Directly addresses flickery mid-stream `
→` transitions during long code blocks. — https://github.com/Yonom/assistant-ui/releases (`react-streamdown@0.3.3`, June 12) — *fit*: pattern-only (Angular); technique: buffer the delta string, render raw text until `message_stop`, then parse+highlight. Eliminates the flicker users see on long code blocks. — *where it lands*: markdown rendering in the message list component.
+- **assistant-ui v0.14.18 — `unstable_useComposerInputHistory` + public `useSmooth`** — Terminal-style up/down arrow input history for the composer; `useSmooth` typewriter-effect with tunable speed/damping; fixed reasoning-content leaking into clipboard copy. — https://github.com/Yonom/assistant-ui/releases (v0.14.18, June 12) — *fit*: pattern-only. Input history = Angular signal state + `keydown` intercept in composer; smooth tuning = config on our streaming text renderer. — *where it lands*: composer component; streaming `content_block_delta` renderer.
+- **NNGroup and Cursor changelogs** — Both were inaccessible this run (permission denied). Treat as not scanned this week.
+
+#### Frontier model announcements
+- **Claude Fable 5 + Mythos 5 announced June 9** — Fable 5 is GA; Mythos 5 is a paired preview tier. Model ID confirmed as `claude-fable-5`. Pricing: $10/M input, $50/M output (≈half the prior Opus-class rate). Bedrock context window and caching specifics not in the announcement — verify in the Bedrock model card before wiring. — https://www.anthropic.com/news — *relevance*: immediate action on model settings UI + model-ID string matching audit. The naming shift (`-fable-`, `-mythos-` suffixes vs. `claude-opus-4.N`) breaks any code that pattern-matches on the numeric dot-version format. *unlocks*: best-available model on Bedrock at materially lower price point.
+- **Pydantic-AI model registry now includes Fable 5 + Mythos 5** — `known_model_names()` enumerable with new names, confirming these are the published Anthropic IDs. — https://github.com/pydantic/pydantic-ai/releases/tag/v1.107.0 — *relevance*: corroborates the new naming pattern; we don't use pydantic-ai's registry but it's a useful cross-check.
+- **OpenAI + Google DeepMind** — Changelog and blog inaccessible this run. Treat as not scanned.
+
+#### Agent harness patterns
+- **Pydantic-AI v1.107.0 — Bedrock null `message_start` guard** — Patches a crash when `message=None` arrives in a Bedrock Anthropic `message_start` stream event. — https://github.com/pydantic/pydantic-ai/releases/tag/v1.107.0 — *relevance*: our `inference_api` SSE relay handles `message_start`; confirm we handle a null `message` field without raising mid-stream. Edge case but proven reproducible upstream.
+- **Pydantic-AI v1.107.0 — native-tool token counting fixed for AnthropicModel** — Counts were wrong when native tools were active. — same release — *relevance*: corroborates the Strands #2635 class of bug; confirms our `CountTokensBedrockModel` path should be audited for tool-heavy turns even after the 1.43 bump.
+- **No new Claude Code changelog entries (June 5–12)** — Last seen: v2.1.163 (Stop/SubagentStop hook `additionalContext`, ~June 4–5). Nothing new in-window. — https://raw.githubusercontent.com/anthropics/claude-code/main/CHANGELOG.md
+- **Async multi-agent cookbook patterns (June 8–9)** — Two new Anthropic cookbook examples: a fixed N-agent team with peer messaging through a shared hub, and a lead agent that spawns/monitors/dismisses async subagents (both Messages API). — https://github.com/anthropics/anthropic-cookbook (commit e22e683) — *relevance*: hub/peer pattern is worth comparing to our A2A client design before the first A2A server PR.
+
+#### opencode (v1.17.x, June 10–12)
+- **v1.17.1 — Usage-described references for agents** — Context files/snippets can now carry agent-facing usage descriptions; these descriptions surface in docs. Directly mirrors our just-shipped skill reference-files (S3-backed, progressive disclosure). — https://github.com/anomalyco/opencode/releases/tag/v1.17.1 — *lens*: Context engineering. Our `reference_files` on skills already carry metadata; opencode's descriptions are the peer pattern.
+- **v1.17.2 — Subagent-scoped permissions** — Subagents use their own configured permissions rather than inheriting parent scope. — https://github.com/anomalyco/opencode/releases/tag/v1.17.2 — *lens*: Tooling. Directly analogous to our per-skill RBAC (`accessible_skill_ids`, `enabled_tools`) — watch as a reference for fine-grained delegation patterns.
+- **v1.17.4 — Connector-based auth + content-filtered responses as visible errors** — Stored provider credentials per-connector; content-filtered model responses surface as visible errors instead of silent failures. — https://github.com/anomalyco/opencode/releases/tag/v1.17.4 — *lens*: Tooling + Cost-effectiveness. The "content-filtered as visible error" pattern confirms our `stream_error` SSE event covers the right class; audit that we emit a clear `stream_error` for Bedrock content-filter rejections specifically (not a silent drop).
+- **v1.17.0 — MCP abort signals for tool call cancellations** — MCP tool calls now receive abort signals for reliable cancellation. — https://github.com/anomalyco/opencode/releases/tag/v1.17.0 — *lens*: Tooling. Our Gateway-fronted Lambda tools currently block without cancellation; this is the peer implementation of the MCP Tasks Extension (SEP-2663) pattern.
+
+#### LibreChat
+- **No new releases since v0.8.6 (June 1)** — v0.8.6 remains the latest as of June 12. — https://github.com/danny-avila/LibreChat/releases — *relevance*: nothing to scan this week.
+
+#### Pricing / quota
+- **No Bedrock price/quota changes this week** — The public pricing page shows no in-window updates. Fable 5 at $10/$50/M on Bedrock is announced but not yet reflected on the pricing page; verify via AWS console before relying on the figure. — https://aws.amazon.com/bedrock/pricing/
+
+#### Community + GitHub issues
+- **AgentCore code interpreter security vulnerabilities (BeyondTrust, June 2026)** — BeyondTrust published research finding vulnerabilities in AgentCore's code interpreter sandbox. — https://www.beyondtrust.com/blog/entry/pwning-aws-agentcore-code-interpreter — *relevance*: **read before enabling the code interpreter surface for end users**. The Interactive Shells announcement (above) is a new code-execution path from AWS; understanding the existing attack surface before expanding exposure is prudent.
+- **Async multi-agent cookbook — hub/peer + spawn/dismiss patterns (June 8–9)** — Two new Anthropic cookbook examples for multi-agent orchestration. — https://github.com/anthropics/anthropic-cookbook (commit e22e683) — *relevance*: reference patterns for our future A2A server construct and multi-agent delegation.
+- **Sentry triage scheduled agent cookbook (June 9)** — Full example: Claude Managed Agent + credential vault + scheduling API + cron-report writing. First vault+scheduling wiring end-to-end. — https://github.com/anthropics/anthropic-cookbook (commit 3dc70b5) — *relevance*: the vault credential binding pattern we don't yet have in skills-mode.
+- **`agents-observe` real-time dashboard for Claude Code teams (HN, 77 pts)** — Open-source monitoring harness for multiple Claude Code agents with plugin overhead notes. — https://github.com/simple10/agents-observe — *relevance*: UX reference for our agent observability surface; plugin overhead notes may be relevant to MCP-apps rendering path.
+
+#### Seasonal
+- Out of window — no AWS re:Invent (late Nov), no NeurIPS/ICLR proceedings dropping. None scanned.
+
+### Patterns worth considering
+
+- **Named-suffix model IDs replacing numeric-version dot IDs** (`claude-fable-5` not `claude-fable-5.0`) — appearing across Anthropic's announcement + pydantic-ai registry + opencode's Fable reasoning support. Any code that pattern-matches on `claude-opus-4.\d` (capability gating, prompt-caching beta header, tool-streaming beta header) will miss `claude-fable-5`. The break is already live.
+  - **Where**: `backend/src/` model config + capability-gate strings, `frontend/.../model-settings.*`
+  - **Fit**: direct fix required — update model list and audit string-match gates
+  - **Verdict**: Worth trying (actually: required)
+
+- **Deferred markdown rendering until `message_stop`** — appearing in react-streamdown@0.3.3 (June 12), observed in LLM chat UIs broadly. Buffer raw delta text, render markdown only after stream completes; eliminates flickery `
→` transitions during streaming code blocks.
+  - **Where**: message list markdown renderer in `frontend/ai.client/src/app/session/`
+  - **Fit**: Pattern-only (Angular); implement as a `defer` flag on the markdown pipe
+  - **Verdict**: Worth trying (L effort, meaningful UX improvement on long code blocks)
+
+---
+
+## Internal Audit
+
+### Activity (last 7 days)
+- **Commits on develop**: 27 (across ~14 PRs)
+- **PRs merged**: ~14 — **reverted**: 0
+- **Issues opened**: 0 in-window (none surfaced via `gh issue list`)
+- **CI failures (workflow → count)**: Nightly Build & Test → **7 consecutive** (June 5–12 every day)
+
+### Repeated friction signals
+- **Nightly Build & Test: 7 consecutive failures (June 5–12)** — every scheduled nightly has failed since the research window opened. This is the same class of failure that was resolved via PR #290 in May (which restored the gate). Root cause is unknown — nobody has pulled the logs.
+  - **Hypothesis**: an env-drift, credential rotation, or dependency issue in the nightly runner; or a real test regression introduced in the Gateway/skills-mode build burst this week that only the full suite catches.
+  - **Fix candidate**: `gh run view  --log-failed` on the June 12 nightly run; classify flaky-vs-regression; quarantine or file within 24h. **This matters before the Strands 1.43 bump** — landing a keystone dep upgrade on a suite we can't currently trust compounds the risk.
+
+- **Queue-to-ship conversion remains zero for defensive/infra items** (per June 5 review: "third consecutive week; proposals evaporating without decision"). Strands 1.42 keystone still unshipped as of today.
+  - **Hypothesis**: feature throughput (27 commits this week, most of them product-facing skills-mode) crowds out maintenance PRs. The queue has 20+ open items, none with logged decisions.
+  - **Fix candidate**: the June 5 review's recommendation still applies — pick **one** keystone item and treat the rest as explicitly Deferred-with-date. The Strands 1.43 bump is the pick.
+
+### Version-pin lag
+
+| Dep | Pinned | Latest | Lag | Notes |
+|---|---|---|---|---|
+| `strands-agents` | 1.40.0 | **1.43.0** | 3 minors | Released today; #2635 fix + context_manager="auto" + A2A isolation fix |
+| `bedrock-agentcore` | 1.9.1 | **1.14.1** | 5 minors | A2A cap fix + interactive shell; lag widening week over week |
+| `mcp` | 1.26.0 | **1.27.2** | 1 minor + 2 patches | ~14 days stale; low urgency |
+| `starlette` | 1.0.0 | **1.0.1** | 1 patch | In CVE-2026-48710 affected range; bump is 1-line |
+| `@angular/core` | 21.2.11 | **22.0.1** | 1 major | Flagged 2+ weeks; spike required before adoption |
+| `aws-cdk-lib` | 2.251.0 | **2.259.0** | 8 patches | Low urgency; routine |
+| `@analogjs/platform` | (alpha track) | 2.6.1 stable | — | Alpha track targets Angular 22; don't mix stable/alpha |
+
+### Retirement candidates
+- **`.claude/skills/angualar-best-practices/` (last modified Dec 30, 2025 — 165 days)** — Name has a typo (`angualar`). Its content overlaps substantially with the newer `tailwind-ui` (May 22) and `frontend-design` (Jan 18) skills. Not referenced in any commit in the last 90 days. Candidate for retirement or merge into a corrected `angular-best-practices`.
+- **`.claude/skills/frontend-design/` (last modified Jan 18, 2026 — 145 days)** — Older than the project-level `tailwind-ui` skill. If content is now covered by `tailwind-ui`, retire; otherwise rename/update.
+
+### Deferred items up for revisit (2026-06-12)
+- **[2026-05-10] Scope AgentCore Runtime BYO filesystem** — deferred 4 weeks in the May 15 review. The new **Interactive Shells** API (June 5) is a materially different code-execution primitive from the same vendor — a simpler path to agent workspaces (WebSocket shell inside the existing container vs. S3/EFS mount with VPC changes). Recommend: keep deferred another 2–4 weeks; add the Interactive Shells item alongside BYO-filesystem for a single architecture decision.
+- **[2026-05-10] Named A2A agent participants in the chat UI** — deferred 4 weeks in the May 15 review. No A2A server construct has landed yet; precondition not met. Recommend: defer another 4 weeks (revisit 2026-07-10).
+
+### Risks introduced this week
+- **CVE-2026-48710 in starlette <1.0.1 — our pin is 1.0.0** — FastMCP v3.4.1 surfaced this CVE. Our `backend/pyproject.toml` pins `starlette==1.0.0` (with a comment saying it was security-pinned for a prior Dependabot alert, ironically). Bump to 1.0.1 before the next backend deploy.
+- **AgentCore code interpreter sandbox vulnerabilities (BeyondTrust)** — Research published this week documents exploits against the AgentCore code interpreter. Read before enabling code-execution surface for end users; the new Interactive Shells API is not a workaround — it's the same underlying runtime.
+- **Strands #2636 still live in v1.43.0** — Non-ASCII dict tool-results get `\uXXXX`-escaped while Pydantic-model returns don't. Cosmetic but could corrupt non-ASCII characters in tool output (multilingual content, emoji in tool names). Second bump needed once PR #2661 merges.
+
+---
+
+## Ideas — Top 5 (ranked)
+
+| # | Idea | Surface | Effort | Impact | Subtracts? | Unlocks? |
+|---|---|---|---|---|---|---|
+| 1 | Bump Strands 1.40 → 1.43 (supersedes the 1.42 keystone) | backend | M | H | yes — closes #2635 defensive item + context_manager="auto" simplification path | native cost ceilings (`Limits`), 1h prompt caching, attribution accuracy |
+| 2 | Add Claude Fable 5 to model settings + audit model-ID string matching | backend + frontend | L-M | H | partial — retires Opus 4.7/4.8 as default once validated | top-of-range model on Bedrock at $10/$50/M; restores model-list parity |
+| 3 | Investigate + triage Nightly Build & Test (7 consecutive failures) | CI | L | H | no — hygiene; prerequisite for trusting the 1.43 bump | restores the only automated backend correctness gate |
+| 4 | Bump `bedrock-agentcore` 1.9.1 → 1.14.1 + note A2A cap prerequisite | backend | L | M | no — dep bump; A2A cap fix is a prerequisite for future A2A server work | interactive shell API access, bearer token integration |
+| 5 | Bump `starlette` 1.0.0 → 1.0.1 (CVE-2026-48710) | backend | L | M | no — 1-line security fix | closes the CVE in the affected range |
+
+### 1. Bump Strands 1.40 → 1.43 (supersedes [2026-06-05] keystone item)
+- **Source**: strands-agents v1.43.0 released June 12, 2026 — https://github.com/strands-agents/sdk-python/releases/tag/python%2Fv1.43.0 + Strands issues #2635 (now closed) and the [2026-06-05] queue keystone
+- **Surface area**: `backend/pyproject.toml` + `uv.lock` (bump from `strands-agents==1.40.0` to `==1.43.0`); agent invocation in `inference_api`; `BedrockModel`/`CacheConfig` (for `cache_tools_ttl`); SSE `limit_*` stop-reason handling; wire `limits={turns, outputTokens, totalTokens}` on agent invocation
+- **Change**: bump to 1.43.0 with the same blast-radius audit as the 1.42 keystone (strands-agents-tools 0.5.2→0.8.0 tool-interface review + starlette transitive compat). On top: no need to guard against #2635 — it's fixed. Note #2636 (non-ASCII) is still live in 1.43.0; add a known-limitation comment.
+- **Subtracts**: adopts library-native `Limits` (retires hand-rolled runaway guardrail) + `cache_tools_ttl` (retires hand-rolled TTL); the separate "[2026-06-05] Guard context-attribution path against #2635" queue item resolves as part of this bump
+- **Unlocks**: native per-turn cost ceiling; end-to-end 1h prompt caching; accurate context attribution on tool-heavy turns (the badge now reports correctly)
+- **Effort × Impact**: Med × High
+- **Verdict**: Ship — prerequisite: confirm the Nightly is green or diagnose it first (see Idea #3)
+
+### 2. Add Claude Fable 5 to model settings + audit model-ID string matching
+- **Source**: Claude Fable 5 GA June 9 — https://aws.amazon.com/about-aws/whats-new/2026/06/claude-fable-5-aws/ + frontier model scan
+- **Surface area**: `frontend/ai.client/src/app/components/model-settings/model-settings.html` + `model-settings.ts` (add `claude-fable-5` to the dropdown); `backend/src/` — grep for any string-match on `claude-opus-4` in capability gating (prompt-caching beta header, fine-grained tool-streaming beta header); admin model catalog
+- **Change**: (1) add `claude-fable-5` (and `claude-mythos-5` preview) to the model list; (2) run `grep -r "claude-opus-4" backend/src/ frontend/` — any match that is a capability gate must be updated to also match the fable/mythos naming convention; (3) verify `CountTokensBedrockModel` de-prefix handles the new ID pattern (the `us.anthropic.*` stripping is already ID-agnostic — confirm no new prefix for Fable 5)
+- **Subtracts**: partial — Fable 5 at $10/$50/M may replace Opus 4.8 as the default once benchmarked; note the thinking/`temperature` provider-translation path from the Opus 4.7 guard may need revisiting for Fable 5's capability tier
+- **Unlocks**: top-of-range Anthropic model on Bedrock; $10/$50 materially lowers per-session cost at the same quality tier
+- **Effort × Impact**: Low-Med × High
+- **Verdict**: Ship — check the Bedrock model card for context window size and caching API support before flipping it to default; use the `claude-api` skill to confirm exact IDs before committing
+
+### 3. Investigate + triage Nightly Build & Test (7 consecutive failures)
+- **Source**: internal CI signal — `gh run list --status=failure` shows every nightly since June 5 failing; same pattern resolved via PR #290 in May (e2e testing fix)
+- **Surface area**: `.github/workflows/nightly.yml` or equivalent; full backend test suite
+- **Change**: `gh run view  --log-failed` → classify: (a) env/credential drift → fix in workflow; (b) flaky test → quarantine; (c) real regression from this week's burst → file + fix. **Do not land the Strands 1.43 bump on an untrusted suite.**
+- **Subtracts**: no — hygiene
+- **Effort × Impact**: Low × High (restores the only automated backend correctness gate)
+- **Verdict**: Ship first — prerequisite for trusting any dep bumps this week
+
+### 4. Bump `bedrock-agentcore` 1.9.1 → 1.14.1 + document A2A cap prerequisite
+- **Source**: bedrock-agentcore v1.14.1 (June 11) — https://github.com/aws/bedrock-agentcore-sdk-python/releases/tag/v1.14.1 — 5 minors behind; open queue item [2026-05-22] + [2026-05-22] re-bump
+- **Surface area**: `backend/pyproject.toml` + `uv.lock`; `AgentCoreMemoryConfig` construction (the `async_mode` adoption from the [2026-05-22] re-bump item)
+- **Change**: bump to 1.14.1; adopt `async_mode` (removes the latent event-loop-blocking failure mode); note in A2A server planning that `a2a-sdk < 1.0` cap fix (v1.14.1) is a prerequisite before the first A2A server PR
+- **Subtracts**: no — dep bump; the `async_mode` adoption retires the latent #452 blocking failure mode
+- **Effort × Impact**: Low × Med
+- **Verdict**: Ship — consolidates two queue items ([2026-05-22] Bump + [2026-05-22] Re-bump)
+
+### 5. Bump `starlette` 1.0.0 → 1.0.1 (CVE-2026-48710)
+- **Source**: FastMCP v3.4.1 (June 5) — https://github.com/jlowin/fastmcp/releases/tag/v3.4.1 — starlette CVE surfaces via the mcp transitive dep
+- **Surface area**: `backend/pyproject.toml` — `starlette==1.0.0` → `starlette==1.0.1`; `uv.lock`
+- **Change**: 1-line pin bump; no API change between 1.0.0 and 1.0.1; bundle with the next backend deploy
+- **Subtracts**: no — 1-line security fix; no code change required beyond the pin
+- **Effort × Impact**: Low × Med (security — closes CVE in the affected range; the original pin comment says "security pin for Dependabot alert" so the intent was always security-gated)
+- **Verdict**: Ship — bundle into whatever the next backend PR is (can be a 1-commit fix alongside the bedrock-agentcore bump)
+
+---
+
+## Take
+
+The week's headline is **Strands 1.43.0 shipping today** — the tracked count_tokens bug (#2635) that threatened the context-attribution badge is gone; upgrade the queue item from 1.42 to 1.43 and ship the keystone this week. The Fable 5 model-ID naming shift is an unblocked breakage risk that can be caught with one grep. Everything else is blocked behind the same prerequisite: fix the Nightly CI first, then trust the suite again before landing dep bumps. The system is shipping features fast and deferring maintenance — the queue keeps growing, but two targeted PRs (nightly investigation + Strands bump) would close more open items than any other week so far.
+
+---
+
+## Sources Scanned
+
+| # | Source | URL | Accessed | Items |
+|---|---|---|---|---|
+| 1 | AWS Bedrock What's New RSS / blog | https://aws.amazon.com/about-aws/whats-new/recent/feed/ + https://aws.amazon.com/blogs/machine-learning/ | 2026-06-12 | 5 |
+| 2 | Strands Agents SDK releases | https://github.com/strands-agents/sdk-python/releases | 2026-06-12 | 5 |
+| 3 | Reference repo commits | https://github.com/aws-samples/sample-strands-agent-with-agentcore/commits/main | 2026-06-12 | 1 (no signal) |
+| 4 | MCP spec PRs | https://github.com/modelcontextprotocol/modelcontextprotocol/pulls | 2026-06-12 | 4 |
+| 5 | MCP servers commits | https://github.com/modelcontextprotocol/servers/commits/main | 2026-06-12 | 1 |
+| 6 | FastMCP releases + PyPI | https://github.com/jlowin/fastmcp/releases + https://pypi.org/project/fastmcp/ | 2026-06-12 | 2 |
+| 7 | MCP Apps spec | https://modelcontextprotocol.io/extensions/apps/overview + https://mcpui.dev/ | 2026-06-12 | 1 |
+| 8 | assistant-ui releases | https://github.com/Yonom/assistant-ui/releases | 2026-06-12 | 2 |
+| 9 | Anthropic newsroom | https://www.anthropic.com/news | 2026-06-12 | 2 |
+| 10 | Pydantic-AI releases | https://github.com/pydantic/pydantic-ai/releases | 2026-06-12 | 3 |
+| 11 | Claude Code changelog | https://raw.githubusercontent.com/anthropics/claude-code/main/CHANGELOG.md | 2026-06-12 | 0 (quiet) |
+| 12 | opencode releases | https://github.com/anomalyco/opencode/releases | 2026-06-12 | 5 |
+| 13 | bedrock-agentcore-sdk-python releases + issues | https://github.com/aws/bedrock-agentcore-sdk-python | 2026-06-12 | 4 |
+| 14 | Anthropic cookbook | https://github.com/anthropics/anthropic-cookbook | 2026-06-12 | 2 |
+| 15 | HN search | https://hn.algolia.com/api/v1/search | 2026-06-12 | 1 |
+| 16 | BeyondTrust security blog | https://www.beyondtrust.com/blog/entry/pwning-aws-agentcore-code-interpreter | 2026-06-12 | 1 |
+| 17 | AWS Bedrock pricing | https://aws.amazon.com/bedrock/pricing/ | 2026-06-12 | 0 (no changes) |
+| 18 | LibreChat releases | https://github.com/danny-avila/LibreChat/releases | 2026-06-12 | 0 (no new) |
+| 19 | PyPI version checks (strands-agents, bedrock-agentcore, mcp, pydantic) | https://pypi.org/pypi//json | 2026-06-12 | 4 |
+| 20 | npm version checks (@angular/core, @analogjs/platform, aws-cdk-lib) | https://registry.npmjs.org//latest | 2026-06-12 | 3 |
+| 21 | Cursor changelog | Not retrieved (permission denied) | — | — |
+| 22 | NNGroup AI topic | Not retrieved (permission denied) | — | — |
+| 23 | OpenAI API changelog | Not retrieved (permission denied) | — | — |
+| 24 | Google DeepMind blog | Not retrieved (permission denied) | — | — |
+
+## Web Budget
+
+Used: ~82 / 50 requests (target).
+Overage: +32 — the version-pin subagent consumed 25 requests (one per registry check × 7 packages, plus retries on npm registries); all other subagents were within their 3–5 request budgets.
+Skipped (permission denied): Cursor changelog, NNGroup AI, OpenAI API changelog, Google DeepMind blog.
+Notes: The version-pin overage was justified — it surfaced the CVE-affected starlette pin and confirmed 3 new major/minor version lags. If version-pin checks become routine, pre-caching registry URLs in a single subagent fetch would halve the request count.
diff --git a/docs/kaizen/review-queue.md b/docs/kaizen/review-queue.md
index 4412bdf93..b283176fd 100644
--- a/docs/kaizen/review-queue.md
+++ b/docs/kaizen/review-queue.md
@@ -5,6 +5,44 @@ Items added by `kaizen-research`, consumed by `kaizen-review-prep`.
 ## Open
 
 
+### [2026-06-12] Bump Strands 1.40 → 1.43 — supersedes [2026-06-05] keystone; closes #2635 + context_manager="auto" + A2A isolation fix
+- **Source**: research/2026-06-12.md ▸ Top 5 #1 — Strands v1.43.0 released June 12, 2026. **Supersedes** [2026-06-05] "Strands 1.40 → 1.42 bump" — target advances one more minor; no additional blast radius. Also closes the [2026-06-05] "#2635 guard" queue item.
+- **Surface**: backend (`pyproject.toml`/`uv.lock` — `strands-agents==1.40.0` → `==1.43.0`; agent invocation in `inference_api`; `BedrockModel`/`CacheConfig`; SSE `limit_*` stop-reason)
+- **Effort × Impact**: M × H
+- **Subtracts**: yes — library-native `Limits` retires the hand-rolled runaway guardrail; `cache_tools_ttl` retires hand-rolled TTL; #2635 defensive guard resolves as part of the bump; three queue items collapse into one PR
+- **Unlocks**: native per-turn cost ceiling; 1h prompt caching; accurate context attribution on tool-heavy turns
+- **Status**: open — prerequisite: confirm Nightly CI is green (see [2026-06-12] nightly investigation item) before landing. #2636 (non-ASCII) still live in 1.43.0 — add a known-limitation comment, a second bump will follow once PR #2661 merges.
+
+### [2026-06-12] Add Claude Fable 5 to model settings + audit model-ID string matching
+- **Source**: research/2026-06-12.md ▸ Top 5 #2 — Claude Fable 5 GA June 9 (https://aws.amazon.com/about-aws/whats-new/2026/06/claude-fable-5-aws/). Naming convention shift (`-fable-`/`-mythos-` suffixes vs. `claude-opus-4.N`) is a live breakage risk.
+- **Surface**: frontend (`model-settings.html`, `model-settings.ts` — add `claude-fable-5` to dropdown) + backend (grep `claude-opus-4` in capability gates: prompt-caching beta header, fine-grained tool-streaming beta header; admin model catalog)
+- **Effort × Impact**: L-M × H
+- **Subtracts**: partial — Fable 5 at $10/$50/M may replace Opus 4.8 as default once benchmarked; no hard retirement yet
+- **Unlocks**: top-of-range Anthropic model on Bedrock at materially lower cost; model-list parity for end users
+- **Status**: open — use the `claude-api` skill to confirm exact Bedrock IDs before committing; verify context window + caching API support on Bedrock model card before flipping to default
+
+### [2026-06-12] Investigate + triage Nightly Build & Test (7 consecutive failures June 5–12)
+- **Source**: research/2026-06-12.md ▸ Internal Audit — CI failures. Same pattern resolved via PR #290 in May; root cause unknown this time.
+- **Surface**: CI — `.github/workflows/` nightly workflow + backend test suite
+- **Effort × Impact**: L × H
+- **Subtracts**: no — hygiene; prerequisite for trusting the Strands 1.43 keystone bump and any other dep changes
+- **Status**: open — time-sensitive; run `gh run view  --log-failed`; classify flaky-vs-regression; quarantine or file. Do not land dep bumps on an untrusted suite.
+
+### [2026-06-12] Bump `bedrock-agentcore` 1.9.1 → 1.14.1 + adopt `async_mode` + note A2A cap prerequisite
+- **Source**: research/2026-06-12.md ▸ Top 5 #4 — bedrock-agentcore v1.14.1 (June 11). **Consolidates** [2026-05-22] "Bump bedrock-agentcore 1.9.1 → 1.11.0" and [2026-05-22] "Re-bump 1.9.1 → 1.11.0 + async_mode" open items (which were already 4+ minors behind; now 5).
+- **Surface**: `backend/pyproject.toml` + `uv.lock`; `AgentCoreMemoryConfig` construction (`async_mode` adoption)
+- **Effort × Impact**: L × M
+- **Subtracts**: `async_mode` adoption retires the latent #452 event-loop-blocking failure mode; two queue items consolidate into one
+- **Unlocks**: interactive shell API access; A2A cap fix is a hard prerequisite for the first A2A server PR
+- **Status**: open — can bundle with the starlette CVE bump (#5 below) as a single "dep hygiene" PR
+
+### [2026-06-12] Bump `starlette` 1.0.0 → 1.0.1 to close CVE-2026-48710
+- **Source**: research/2026-06-12.md ▸ Top 5 #5 — FastMCP v3.4.1 (June 5) surfaced CVE-2026-48710 affecting starlette < 1.0.1. Our `pyproject.toml` pins `starlette==1.0.0`.
+- **Surface**: `backend/pyproject.toml` — 1-line pin bump
+- **Effort × Impact**: L × M
+- **Subtracts**: no — 1-line security fix; the existing comment says the pin was already security-motivated
+- **Status**: open — bundle with bedrock-agentcore bump (#4 above) as a single dep-hygiene PR; also flag to MCP server repos to bump FastMCP to ≥3.4.1
+
 ### [2026-06-05] Strands 1.40 → 1.42 bump — unblocks `Limits` (cost caps) + `cache_tools_ttl` (#269 caching)
 - **Source**: research/2026-06-05.md ▸ Top 5 #1 — Strands v1.42.0 (June 1). **Consolidates** the queued 2026-05-22 "Strands 1.40→1.41 + caching #269" item AND the 2026-05-29 #2 "Adopt Strands `Limits`" item — both were gated on 1.42, which is now out. Treat as one keystone bump, not two.
 - **Surface**: backend (`pyproject.toml`/`uv.lock`, agent invocation in `inference_api`, `BedrockModel`/`CacheConfig`, SSE `limit_*` stop-reason) + infrastructure (CloudWatch Bedrock-spend alarm)