kaizen bundle: /ping reap fix + A2A guard + research URL cleanup - #338
Merged
Merged
Conversation
…ping AgentCore's idle reaper requires an integer `time_of_last_update` field alongside `status`; when absent, the platform reaps the microVM at idleRuntimeSessionTimeout even mid-stream regardless of reported status (bedrock-agentcore-sdk-python#471). We have no async-task busy tracking (deferred async_mode work), so we cannot report HealthyBusy — returning a fresh timestamp on every ping is the documented mitigation against silent mid-generation reaps. Also corrects status casing to match PingStatus. Kaizen 2026-05-15 review ▸ Proposal #2. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
A2A is currently client-only. When the first A2A server construct lands (Strands agent.to_a2a(), A2AServer, or a hand-built AgentCard), its advertised capabilities must include streaming=True — otherwise the A2A SDK client silently falls back to non-streaming, never receives a `completed` event, and hangs ~40 minutes (ref-repo sample-strands-agent-with-agentcore commit 50c9112). Forward-looking guard since no construction site exists yet. Kaizen 2026-05-15 review ▸ Proposal #4. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…en-research
- Drop the bedrock/whats-new/ 404 (redundant with the AWS What's New RSS
feed already listed; annotated that feed to filter for Bedrock/AgentCore)
- Replace the docs.claude.com claude-code release-notes 301->404 with the
anthropics/claude-code CHANGELOG.md
- Drop anthropics/courses (quiet since Nov 2025)
- Fix aws/amazon-bedrock-agentcore-{sdk-python,starter-toolkit} -> the
correct aws/bedrock-agentcore-* slugs (review flagged the toolkit; the
sdk-python line had the identical typo)
Kaizen 2026-05-15 review ▸ Proposal #9.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…f_last_update) The prior test pinned status to the non-conforming lowercase "healthy". Assert the real AgentCore contract instead: status is a valid PingStatus value and time_of_last_update is an int — which now also guards the #471 fix against regression. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
philmerrell
added a commit
that referenced
this pull request
May 18, 2026
#341) Move 8 proposals Open -> Resolved (shipped via #337/#338/#339/#340 or declined): bedrock-agentcore 1.9.1, /ping reap fix, A2A guard, dead-URL cleanup, #266/#267 premise-correction, Reddit decline, Strands 1.40 bump, tool-renderer registry (MCP Apps PR #0). Create docs/kaizen/decisions.md with the Reddit decline, the #266/#267 premise-correction, and a scope note that Strands built-in proactive compression is NOT a drop-in for our TurnBasedSessionManager (tool truncation + LTM retrieval + DynamoDB checkpoint state) so future research does not re-propose the bare subtraction. Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
philmerrell
added a commit
that referenced
this pull request
Sep 11, 2026
…1051) `ActiveSessionCount` on the `AgentCore.Runtime` dimension has been graphed on the observability dashboard since #910, but nothing read it. The sibling `agentcore-code-interpreter-active-sessions` alarm already watches the same metric on the CodeInterpreter dimension; this mirrors it onto Runtime. Runtime bills memory for a session's whole lifetime rather than for compute time, and AWS still exposes no API to list or force-terminate an active runtime session (aws/bedrock-agentcore-starter-toolkit#498 reports one runaway session burning $72.67 in 58 minutes). The DynamoDB lease plus `cancelRequestedFor` are the only kill switch we have, so session accumulation is the leading indicator — exactly what the #338 `/ping` reaper bug did, undetected, for three months at 73% of the platform bill. Threshold is a tunable (`CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD`) defaulting to 200, plumbed through platform.yml and load-env.sh like every other observability threshold. 200 is deliberately not a fraction of the 5,000-session account quota — quota exhaustion is already owned by `agentcore-throttles`, and #1016 is the standing lesson about comparing an account-wide roll-up against a number that does not denominate it. Routed through `AlarmFactory` per #910, so it reaches `{prefix}-alarms` as a consequence of being created. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
philmerrell
added a commit
that referenced
this pull request
Sep 11, 2026
…s by duration (#1058) #1051 shipped `agentcore-runtime-active-sessions` at threshold 200 over 3 evaluation periods (15 minutes). Validated against 7 days of real prod `ActiveSessionCount` (2026-09-04..11) it would have fired three times in three nights — every one of them a planned load test, peaking at 241, 608 and 1404. No threshold fixes that. Load tests still trip it at 500, and it only goes quiet around 1500 — which is *above* the ~99 sustained by the #338 `/ping` reaper regression this alarm exists to catch. Tuning the number either pages on every load test or never fires on the regression. Duration does separate them. The regression sustained ~99 for three months; the load tests ran 15-30 minutes, and the longest continuous run above 75 in that week was 45 minutes. Simulated against the same data: | threshold | window | firings | |-----------|--------|---------| | 200 | 15 min | 3 (every load test) | | 75 | 30 min | 2 | | 75 | 60 min | 0 | So: threshold 200 -> 75, evaluationPeriods 3 -> 12 (one hour sustained, all 12 datapoints breaching). That clears every observed load test with 15 minutes of margin while a sustained 99 alarms one hour after onset. Ordinary prod traffic outside the bursts was max 8-16 (mean 2.4-6.8) and dev peaks at 5, so 75 still clears a normal day by a wide margin and sits at 1.5% of the 5,000-session account quota that `agentcore-throttles` owns. The steering runbook said to respond by raising the threshold; that advice is now inverted, since raising it is precisely what cannot work. The alarm description says the window has already elapsed, so a burst alone should not have reached the responder. Threshold stays tunable via CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD. The alarm is not yet deployed to prod (it ships with the next release), so this lands before it can page anyone. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Kaizen 2026-05-15 review bundle — proposals #2 + #4 + #7 + #9. Proposal #1 (
bedrock-agentcore1.9.1) already shipped in #337.Summary
fix(inference)—/pingnow emits integertime_of_last_update+ correctedHealthystatus casing. Without the field AgentCore's idle reaper kills the microVM atidleRuntimeSessionTimeoutmid-stream regardless of status (bedrock-agentcore-sdk-python#471). We have no async-task busy tracking (deferredasync_modework) so we can't reportHealthyBusy; a fresh per-ping timestamp is the documented mitigation — it disables ping-based idle reaping for this runtime (the explicit, accepted trade).docs— A2A is client-only today (no serverAgentCard). Added a forward-looking guard inCLAUDE.MD: the first A2A server construct MUST advertisecapabilitieswithstreaming=True, else A2A clients hang ~40 min (ref-repo50c9112).chore(kaizen)— replaced/dropped dead source URLs in thekaizen-researchskill; fixedaws/amazon-bedrock-agentcore-*→ correctaws/bedrock-agentcore-*slugs (the review flagged the toolkit; the sdk-python line had the same typo).Not in this PR
Test plan
cd backend && uv run python -m pytest tests/ -v(no test asserts the old/pingshape;app_api/healthis a separate endpoint)GET /pingin dev →{"status":"Healthy","time_of_last_update":<int>,"version":...}CLAUDE.MDMulti-Protocol section +kaizen-research/SKILL.mdsource list🤖 Generated with Claude Code