Repository navigation
fix(observability): alarm on Bedrock TPM quota per model, not account-wide - #1016
Merged
Merged
Conversation
…-wide
`bedrock-tpm-quota-usage` compared `EstimatedTPMQuotaUsage` against 80 as
though the metric were a percentage of quota. It is an absolute token count,
so 80 was crossed by roughly one sentence of model output. In production the
alarm was above threshold for 195 of 197 five-minute datapoints over 24
hours, clearing only when a period had no data at all - it was reporting
whether anyone was using the product. It produced 205 state transitions in
six days and was the source of very nearly all alarm traffic.
Quotas are per model AND per inference profile, so the dimension-less
account-wide roll-up had no single denominator to be a percentage of.
Summing a 40,000,000-quota profile with a 200,000-quota one yields a number
comparable to nothing, and it hides the model closest to its own ceiling:
Claude Sonnet 4 runs at ~45% of a 200,000 quota, while Sonnet 5 carries 94%
of the traffic at ~1% of 40,000,000.
Replaced with one alarm per configured model, thresholded at
`bedrockTpmQuotaPercent` (default 75) of that model's own quota.
Quotas are operator-supplied and default to EMPTY. They are per-account and
adjustable - this account has three increase requests on record, one still
open - so no value is shippable, and an empty map creates no alarm rather
than a confidently wrong one. `bedrock-invocation-throttles` remains the
backstop and needs no quota configured.
The quota map is the first non-scalar observability tunable, which exposed a
transport hazard: `deploy.sh` runs `eval npx cdk synth ${CDK_CONTEXT_PARAMS}`
and eval removes quote characters, so a JSON value passed the way every other
`--context` value is passed arrives as `{model:40000000}` - unparseable, and
it would have fallen back to the empty default creating no alarms with no
error. The parser therefore also accepts a quote-free `modelId=quota,...`
form, `load-env.sh` single-quotes the value, and it fails loudly on a value
containing a single quote instead of degrading silently.
The alarm description leads with "confirm the configured quota is still the
live quota", because nothing checks that automatically; the runbook table in
.kiro/steering/observability.md says the same.
Verified: full suite 782 passing, 8 failing - byte-identical to the failure
set on untouched origin/develop (shell executable-bit checks on Windows, a
committed GSI snapshot, a kb-migration digest pin). No new failures, none
disappeared, +7 tests. tsc clean; bash -n clean; eval quoting verified
empirically against a control.
philmerrell
added a commit
that referenced
this pull request
Sep 11, 2026
…1051) `ActiveSessionCount` on the `AgentCore.Runtime` dimension has been graphed on the observability dashboard since #910, but nothing read it. The sibling `agentcore-code-interpreter-active-sessions` alarm already watches the same metric on the CodeInterpreter dimension; this mirrors it onto Runtime. Runtime bills memory for a session's whole lifetime rather than for compute time, and AWS still exposes no API to list or force-terminate an active runtime session (aws/bedrock-agentcore-starter-toolkit#498 reports one runaway session burning $72.67 in 58 minutes). The DynamoDB lease plus `cancelRequestedFor` are the only kill switch we have, so session accumulation is the leading indicator — exactly what the #338 `/ping` reaper bug did, undetected, for three months at 73% of the platform bill. Threshold is a tunable (`CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD`) defaulting to 200, plumbed through platform.yml and load-env.sh like every other observability threshold. 200 is deliberately not a fraction of the 5,000-session account quota — quota exhaustion is already owned by `agentcore-throttles`, and #1016 is the standing lesson about comparing an account-wide roll-up against a number that does not denominate it. Routed through `AlarmFactory` per #910, so it reaches `{prefix}-alarms` as a consequence of being created. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
bedrock-tpm-quota-usagecomparedEstimatedTPMQuotaUsageagainst80as though themetric were a percentage of quota. It is an absolute token count, so 80 is crossed by
roughly one sentence of model output.
Measured in prod before this change:
It never cleared because anything recovered — it cleared only when a period had no data
at all, which
treatMissingData: notBreachingscores as healthy. In effect it reportedwhether anyone was using the product.
Why per-model
TPM quotas are per model and per inference profile, so the dimension-less account-wide
roll-up has no single denominator to be a percentage of. It also hid the model actually
closest to its ceiling:
global.anthropic.claude-sonnet-5us.anthropic.claude-sonnet-4-20250514-v1:0global.anthropic.claude-haiku-4-5-...amazon.titan-embed-text-v2:0Sonnet 5 carries 94% of the traffic at 1% of quota; Sonnet 4 is barely used but sits at 45%
of a quota 200× smaller. Summing them produces a number comparable to nothing.
Change
One alarm per configured model, thresholded at
bedrockTpmQuotaPercent(default 75) ofthat model's own quota.
Statistic: Maximumis kept deliberately — the quota is per minute,the period is five, so Maximum reads the peak minute; averaging would dilute a real spike.
Quotas are operator-supplied and default to empty. They are per-account and adjustable —
this account has three increase requests on record, one still
CASE_OPENED— so no value isshippable. An empty map creates no alarm rather than a confidently wrong one, and
bedrock-invocation-throttlesremains the backstop with no quota needed.There is no automated drift check, by decision. Instead the alarm description leads with
"confirm the configured quota is still the live quota", so the reminder arrives with the
alert; the runbook table says the same.
Transport hazard found while wiring this up
This is the first non-scalar observability tunable, which exposed a silent failure mode.
deploy.shrunseval npx cdk synth ${CDK_CONTEXT_PARAMS}, and eval removes quotecharacters, so JSON passed the way every other
--contextvalue is passed arrives mangled:Unparseable → falls back to the empty default → no alarms, no error. Verified
empirically against a control (a scalar still arrives as
90).Mitigations: the parser also accepts a quote-free
modelId=quota,...form,load-env.shsingle-quotes the value, and it fails loudly on a value containing a single quote rather
than degrading quietly. Regression test asserts the mangled shape cannot half-parse into a
NaNthreshold.Required before this deploys
The default is empty, so deploying without these deletes the old alarm and creates zero
replacements:
Use the
k=vform, not JSON — it cannot be corrupted by eval. Sonnet 5's 40,000,000 has anincrease case still open from 2026-09-04; re-read it before trusting that entry.
Verification
origin/develop(shell executable-bit checks on Windows, a committed GSI snapshot, akb-migration digest pin). No new failures, none disappeared, +7 tests.
tsc --noEmitclean;bash -n scripts/common/load-env.shclean.cdk synthofboisestateai-v2-PlatformStackproduces exactly 2 quota alarms withthresholds 30,000,000 and 150,000, correct
ModelIddimensions,Maximum, 3×300s,notBreaching, and 1AlarmAction+ 1OKActioneach. No dimension-lessEstimatedTPMQuotaUsagealarm remains.Note for reviewers
A local
cdk deployfrom a checkout without the deployment config is destructive — itplans 11 deletions including the ALB HTTPS listener, the Cognito pre-token-generation Lambda
and the Announcements DynamoDB table, because
cdk.context.jsonships placeholder values andsynth.shexits 0 with only a soft warning. Not addressed here, butload-env.shhard-failson a malformed boolean and could do the same for
certificateArn/domainName.