Skip to content

fix(observability): alarm on Bedrock TPM quota per model, not account-wide - #1016

Merged
ramenNoodles1998 merged 1 commit into
developfrom
fix/bedrock-tpm-per-model-quota
Sep 9, 2026
Merged

ramenNoodles1998 merged 1 commit into
developfrom
fix/bedrock-tpm-per-model-quota

Conversation

@ramenNoodles1998

Copy link
Copy Markdown
Contributor

Problem

bedrock-tpm-quota-usage compared EstimatedTPMQuotaUsage against 80 as though the
metric were a percentage of quota. It is an absolute token count, so 80 is crossed by
roughly one sentence of model output.

Measured in prod before this change:

5-minute datapoints over threshold, last 24h 195 of 197
State transitions, 09-03 → 09-08 205 (28–41/day)
Share of all alarm notifications very nearly all of them
Datapoint range 2 → 445,831

It never cleared because anything recovered — it cleared only when a period had no data
at all
, which treatMissingData: notBreaching scores as healthy. In effect it reported
whether anyone was using the product.

Why per-model

TPM quotas are per model and per inference profile, so the dimension-less account-wide
roll-up has no single denominator to be a percentage of. It also hid the model actually
closest to its ceiling:

Model Recent peak Applied quota Utilisation
global.anthropic.claude-sonnet-5 445,831 40,000,000 1.1%
us.anthropic.claude-sonnet-4-20250514-v1:0 89,503 200,000 45%
global.anthropic.claude-haiku-4-5-... 51,755 5,000,000 1.0%
amazon.titan-embed-text-v2:0 6,438 300,000 2.1%

Sonnet 5 carries 94% of the traffic at 1% of quota; Sonnet 4 is barely used but sits at 45%
of a quota 200× smaller. Summing them produces a number comparable to nothing.

Change

One alarm per configured model, thresholded at bedrockTpmQuotaPercent (default 75) of
that model's own quota. Statistic: Maximum is kept deliberately — the quota is per minute,
the period is five, so Maximum reads the peak minute; averaging would dilute a real spike.

Quotas are operator-supplied and default to empty. They are per-account and adjustable —
this account has three increase requests on record, one still CASE_OPENED — so no value is
shippable. An empty map creates no alarm rather than a confidently wrong one, and
bedrock-invocation-throttles remains the backstop with no quota needed.

There is no automated drift check, by decision. Instead the alarm description leads with
"confirm the configured quota is still the live quota", so the reminder arrives with the
alert; the runbook table says the same.

Transport hazard found while wiring this up

This is the first non-scalar observability tunable, which exposed a silent failure mode.
deploy.sh runs eval npx cdk synth ${CDK_CONTEXT_PARAMS}, and eval removes quote
characters, so JSON passed the way every other --context value is passed arrives mangled:

value: {"global.anthropic.claude-sonnet-5":40000000}
after eval: observability.bedrockTpmQuotas={global.anthropic.claude-sonnet-5:40000000}

Unparseable → falls back to the empty default → no alarms, no error. Verified
empirically against a control (a scalar still arrives as 90).

Mitigations: the parser also accepts a quote-free modelId=quota,... form, load-env.sh
single-quotes the value, and it fails loudly on a value containing a single quote rather
than degrading quietly. Regression test asserts the mangled shape cannot half-parse into a
NaN threshold.

Required before this deploys

The default is empty, so deploying without these deletes the old alarm and creates zero
replacements:

CDK_OBSERVABILITY_BEDROCK_TPM_QUOTA_PERCENT = 75
CDK_OBSERVABILITY_BEDROCK_TPM_QUOTAS = global.anthropic.claude-sonnet-5=40000000,us.anthropic.claude-sonnet-4-20250514-v1:0=200000,global.anthropic.claude-haiku-4-5-20251001-v1:0=5000000,us.anthropic.claude-haiku-4-5-20251001-v1:0=5000000,us.amazon.nova-micro-v1:0=8000000,amazon.titan-embed-text-v2:0=300000,amazon.titan-embed-text-v1=300000

Use the k=v form, not JSON — it cannot be corrupted by eval. Sonnet 5's 40,000,000 has an
increase case still open from 2026-09-04; re-read it before trusting that entry.

Verification

  • Full suite 782 passing / 8 failing — byte-identical failure set to untouched
    origin/develop (shell executable-bit checks on Windows, a committed GSI snapshot, a
    kb-migration digest pin). No new failures, none disappeared, +7 tests.
  • tsc --noEmit clean; bash -n scripts/common/load-env.sh clean.
  • Real cdk synth of boisestateai-v2-PlatformStack produces exactly 2 quota alarms with
    thresholds 30,000,000 and 150,000, correct ModelId dimensions, Maximum, 3×300s,
    notBreaching, and 1 AlarmAction + 1 OKAction each. No dimension-less
    EstimatedTPMQuotaUsage alarm remains.

Note for reviewers

A local cdk deploy from a checkout without the deployment config is destructive — it
plans 11 deletions including the ALB HTTPS listener, the Cognito pre-token-generation Lambda
and the Announcements DynamoDB table, because cdk.context.json ships placeholder values and
synth.sh exits 0 with only a soft warning. Not addressed here, but load-env.sh hard-fails
on a malformed boolean and could do the same for certificateArn / domainName.

…-wide

`bedrock-tpm-quota-usage` compared `EstimatedTPMQuotaUsage` against 80 as
though the metric were a percentage of quota. It is an absolute token count,
so 80 was crossed by roughly one sentence of model output. In production the
alarm was above threshold for 195 of 197 five-minute datapoints over 24
hours, clearing only when a period had no data at all - it was reporting
whether anyone was using the product. It produced 205 state transitions in
six days and was the source of very nearly all alarm traffic.

Quotas are per model AND per inference profile, so the dimension-less
account-wide roll-up had no single denominator to be a percentage of.
Summing a 40,000,000-quota profile with a 200,000-quota one yields a number
comparable to nothing, and it hides the model closest to its own ceiling:
Claude Sonnet 4 runs at ~45% of a 200,000 quota, while Sonnet 5 carries 94%
of the traffic at ~1% of 40,000,000.

Replaced with one alarm per configured model, thresholded at
`bedrockTpmQuotaPercent` (default 75) of that model's own quota.

Quotas are operator-supplied and default to EMPTY. They are per-account and
adjustable - this account has three increase requests on record, one still
open - so no value is shippable, and an empty map creates no alarm rather
than a confidently wrong one. `bedrock-invocation-throttles` remains the
backstop and needs no quota configured.

The quota map is the first non-scalar observability tunable, which exposed a
transport hazard: `deploy.sh` runs `eval npx cdk synth ${CDK_CONTEXT_PARAMS}`
and eval removes quote characters, so a JSON value passed the way every other
`--context` value is passed arrives as `{model:40000000}` - unparseable, and
it would have fallen back to the empty default creating no alarms with no
error. The parser therefore also accepts a quote-free `modelId=quota,...`
form, `load-env.sh` single-quotes the value, and it fails loudly on a value
containing a single quote instead of degrading silently.

The alarm description leads with "confirm the configured quota is still the
live quota", because nothing checks that automatically; the runbook table in
.kiro/steering/observability.md says the same.

Verified: full suite 782 passing, 8 failing - byte-identical to the failure
set on untouched origin/develop (shell executable-bit checks on Windows, a
committed GSI snapshot, a kb-migration digest pin). No new failures, none
disappeared, +7 tests. tsc clean; bash -n clean; eval quoting verified
empirically against a control.
@ramenNoodles1998
ramenNoodles1998 merged commit f5b9fcd into develop Sep 9, 2026
4 checks passed
philmerrell added a commit that referenced this pull request Sep 11, 2026
…1051)

`ActiveSessionCount` on the `AgentCore.Runtime` dimension has been graphed
on the observability dashboard since #910, but nothing read it. The sibling
`agentcore-code-interpreter-active-sessions` alarm already watches the same
metric on the CodeInterpreter dimension; this mirrors it onto Runtime.

Runtime bills memory for a session's whole lifetime rather than for compute
time, and AWS still exposes no API to list or force-terminate an active
runtime session (aws/bedrock-agentcore-starter-toolkit#498 reports one
runaway session burning $72.67 in 58 minutes). The DynamoDB lease plus
`cancelRequestedFor` are the only kill switch we have, so session
accumulation is the leading indicator — exactly what the #338 `/ping`
reaper bug did, undetected, for three months at 73% of the platform bill.

Threshold is a tunable (`CDK_OBSERVABILITY_AGENTCORE_ACTIVE_SESSION_THRESHOLD`)
defaulting to 200, plumbed through platform.yml and load-env.sh like every
other observability threshold. 200 is deliberately not a fraction of the
5,000-session account quota — quota exhaustion is already owned by
`agentcore-throttles`, and #1016 is the standing lesson about comparing an
account-wide roll-up against a number that does not denominate it.

Routed through `AlarmFactory` per #910, so it reaches `{prefix}-alarms` as a
consequence of being created.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants