Skip to content

kaizen bundle: /ping reap fix + A2A guard + research URL cleanup - #338

Merged
philmerrell merged 4 commits into
developfrom
kaizen/bundle-2-4-7-9
May 18, 2026
Merged

philmerrell merged 4 commits into
developfrom
kaizen/bundle-2-4-7-9

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Kaizen 2026-05-15 review bundle — proposals #2 + #4 + #7 + #9. Proposal #1 (bedrock-agentcore 1.9.1) already shipped in #337.

Summary

  • move RAG infrastructure into separate stack + fix general permissions issues #2 fix(inference) — /ping now emits integer time_of_last_update + corrected Healthy status casing. Without the field AgentCore's idle reaper kills the microVM at idleRuntimeSessionTimeout mid-stream regardless of status (bedrock-agentcore-sdk-python#471). We have no async-task busy tracking (deferred async_mode work) so we can't report HealthyBusy; a fresh per-ping timestamp is the documented mitigation — it disables ping-based idle reaping for this runtime (the explicit, accepted trade).
  • consolidate multi-environment strategy between CDK stacks and refactor codebase to be environment agnostic #4 docs — A2A is client-only today (no server AgentCard). Added a forward-looking guard in CLAUDE.MD: the first A2A server construct MUST advertise capabilities with streaming=True, else A2A clients hang ~40 min (ref-repo 50c9112).
  • Develop #9 chore(kaizen) — replaced/dropped dead source URLs in the kaizen-research skill; fixed aws/amazon-bedrock-agentcore-* → correct aws/bedrock-agentcore-* slugs (the review flagged the toolkit; the sdk-python line had the same typo).

Not in this PR

Test plan

  • cd backend && uv run python -m pytest tests/ -v (no test asserts the old /ping shape; app_api/health is a separate endpoint)
  • Smoke GET /ping in dev → {"status":"Healthy","time_of_last_update":<int>,"version":...}
  • Long tool-turn in dev does not get reaped mid-stream
  • Skim rendered CLAUDE.MD Multi-Protocol section + kaizen-research/SKILL.md source list

🤖 Generated with Claude Code

philmerrell and others added 4 commits May 18, 2026 14:46
…ping

AgentCore's idle reaper requires an integer `time_of_last_update` field
alongside `status`; when absent, the platform reaps the microVM at
idleRuntimeSessionTimeout even mid-stream regardless of reported status
(bedrock-agentcore-sdk-python#471). We have no async-task busy tracking
(deferred async_mode work), so we cannot report HealthyBusy — returning a
fresh timestamp on every ping is the documented mitigation against silent
mid-generation reaps. Also corrects status casing to match PingStatus.

Kaizen 2026-05-15 review ▸ Proposal #2.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
A2A is currently client-only. When the first A2A server construct lands
(Strands agent.to_a2a(), A2AServer, or a hand-built AgentCard), its
advertised capabilities must include streaming=True — otherwise the A2A
SDK client silently falls back to non-streaming, never receives a
`completed` event, and hangs ~40 minutes (ref-repo
sample-strands-agent-with-agentcore commit 50c9112). Forward-looking
guard since no construction site exists yet.

Kaizen 2026-05-15 review ▸ Proposal #4.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…en-research

- Drop the bedrock/whats-new/ 404 (redundant with the AWS What's New RSS
  feed already listed; annotated that feed to filter for Bedrock/AgentCore)
- Replace the docs.claude.com claude-code release-notes 301->404 with the
  anthropics/claude-code CHANGELOG.md
- Drop anthropics/courses (quiet since Nov 2025)
- Fix aws/amazon-bedrock-agentcore-{sdk-python,starter-toolkit} -> the
  correct aws/bedrock-agentcore-* slugs (review flagged the toolkit; the
  sdk-python line had the identical typo)

Kaizen 2026-05-15 review ▸ Proposal #9.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…f_last_update)

The prior test pinned status to the non-conforming lowercase "healthy".
Assert the real AgentCore contract instead: status is a valid PingStatus
value and time_of_last_update is an int — which now also guards the #471
fix against regression.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit bc16420 into develop May 18, 2026
11 checks passed
@philmerrell
philmerrell deleted the kaizen/bundle-2-4-7-9 branch May 18, 2026 21:01
philmerrell added a commit that referenced this pull request May 18, 2026
#341)

Move 8 proposals Open -> Resolved (shipped via #337/#338/#339/#340 or
declined): bedrock-agentcore 1.9.1, /ping reap fix, A2A guard, dead-URL
cleanup, #266/#267 premise-correction, Reddit decline, Strands 1.40 bump,
tool-renderer registry (MCP Apps PR #0).

Create docs/kaizen/decisions.md with the Reddit decline, the #266/#267
premise-correction, and a scope note that Strands built-in proactive
compression is NOT a drop-in for our TurnBasedSessionManager (tool
truncation + LTM retrieval + DynamoDB checkpoint state) so future
research does not re-propose the bare subtraction.

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
philmerrell added a commit that referenced this pull request Sep 11, 2026
…1051)

`ActiveSessionCount` on the `AgentCore.Runtime` dimension has been graphed
on the observability dashboard since #910, but nothing read it. The sibling
`agentcore-code-interpreter-active-sessions` alarm already watches the same
metric on the CodeInterpreter dimension; this mirrors it onto Runtime.

Runtime bills memory for a session's whole lifetime rather than for compute
time, and AWS still exposes no API to list or force-terminate an active
runtime session (aws/bedrock-agentcore-starter-toolkit#498 reports one
runaway session burning $72.67 in 58 minutes). The DynamoDB lease plus
`cancelRequestedFor` are the only kill switch we have, so session
accumulation is the leading indicator — exactly what the #338 `/ping`
reaper bug did, undetected, for three months at 73% of the platform bill.

Threshold is a tunable (`CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD`)
defaulting to 200, plumbed through platform.yml and load-env.sh like every
other observability threshold. 200 is deliberately not a fraction of the
5,000-session account quota — quota exhaustion is already owned by
`agentcore-throttles`, and #1016 is the standing lesson about comparing an
account-wide roll-up against a number that does not denominate it.

Routed through `AlarmFactory` per #910, so it reaches `{prefix}-alarms` as a
consequence of being created.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
philmerrell added a commit that referenced this pull request Sep 11, 2026
…s by duration (#1058)

#1051 shipped `agentcore-runtime-active-sessions` at threshold 200 over 3
evaluation periods (15 minutes). Validated against 7 days of real prod
`ActiveSessionCount` (2026-09-04..11) it would have fired three times in three
nights — every one of them a planned load test, peaking at 241, 608 and 1404.

No threshold fixes that. Load tests still trip it at 500, and it only goes quiet
around 1500 — which is *above* the ~99 sustained by the #338 `/ping` reaper
regression this alarm exists to catch. Tuning the number either pages on every
load test or never fires on the regression.

Duration does separate them. The regression sustained ~99 for three months; the
load tests ran 15-30 minutes, and the longest continuous run above 75 in that
week was 45 minutes. Simulated against the same data:

| threshold | window | firings |
|-----------|--------|---------|
| 200       | 15 min | 3 (every load test) |
| 75        | 30 min | 2 |
| 75        | 60 min | 0 |

So: threshold 200 -> 75, evaluationPeriods 3 -> 12 (one hour sustained, all 12
datapoints breaching). That clears every observed load test with 15 minutes of
margin while a sustained 99 alarms one hour after onset. Ordinary prod traffic
outside the bursts was max 8-16 (mean 2.4-6.8) and dev peaks at 5, so 75 still
clears a normal day by a wide margin and sits at 1.5% of the 5,000-session
account quota that `agentcore-throttles` owns.

The steering runbook said to respond by raising the threshold; that advice is
now inverted, since raising it is precisely what cannot work. The alarm
description says the window has already elapsed, so a burst alone should not
have reached the responder.

Threshold stays tunable via CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD.
The alarm is not yet deployed to prod (it ships with the next release), so this
lands before it can page anyone.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant