You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Setting the platform default model to claude-opus-5-5 (selectable since #2987) broke every execution that does not pin a model, on a production instance. Two defects combined to make a single settings change look like a subscription outage:
The agent base image cannot pick up a new Claude Code.docker/base-image/Dockerfile installs @anthropic-ai/claude-code@latest. Docker caches that layer by its text, so a rebuild reuses whatever version was cached. The instance's trinity-agent-base:latest was rebuilt on 2026-09-24 and still carries 2.1.278, although 2.1.280 was published on 2026-09-22 and 2.1.281 on 2026-09-23. claude-opus-5-5 needs ≥ 2.1.280.
A model rejection is classified as an auth failure. The run exits 1 with [claude-code:unrecognized_model], the agent server answers HTTP 503, and classify_switch_failure (services/execution_classification.py) maps 503 to failure_kind="auth". SUB-003 then rotates the agent through every subscription and finally falls back to the platform API key (bug(subscriptions): a Workspace message on a rate-limited subscription fails instead of switching and completing — close the SUB-003 gaps on the turn path #2638). The Workspace user is told "The agent hit its usage limit, so it was moved onto the platform API key. Send that again and it should go through." (client_portal/service.py). Resending fails the same way.
Evidence
One agent, 2026-09-24 (UTC):
Time
Event
16:08
admin sets platform_default_model = claude-opus-5-5 via Settings (PUT /api/settings/platform_default_model)
16:12
reminder run fails unrecognized_model → SUB-003 switch #1 (failure_kind: auth)
Agent trinity-pm responded: HTTP 503 (1940ms)
[SUB-003] Auto-switching agent ... after an authentication failure
Agent ... post-switch retry responded: HTTP 503 (1995ms)
Failed to execute task: Execution failed with no output (exit code 1): [claude-code:unrecognized_model] {"model":"claude-opus-5-5","query_source":"sdk"}
Auth failure detected on ...
Reproduced outside Trinity, claude -p --model claude-opus-5-5 "Reply with exactly: OK":
Claude Code
Result
2.1.278
API Error: 400 Claude Code 2.1.278 does not support this model; version 2.1.280 or newer is required.
2.1.280 / 2.1.281 / 2.1.282
OK
The API's 400 names the fix, but that sentence never reaches the backend; the stored error holds only the [claude-code:unrecognized_model] marker.
Proposed fix
1. Pin Claude Code in the base image.ARG CLAUDE_CODE_VERSION=2.1.281 and install @anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}. Bumping the value invalidates the layer, so a rebuild really upgrades, and the bump can ship in the same PR as any model-catalog addition that needs it.
2. Treat a model rejection as its own failure, not auth.
Agent server: map unrecognized_model (and the API's does not support this model 400) to a distinct error_code (e.g. MODEL_UNSUPPORTED), carrying the API's message text, instead of a 503.
Backend: classify_switch_failure must return None for it, so SUB-003 does not switch subscriptions or fall back to the API key.
Workspace / chat: show the API's sentence (e.g. "Claude Code 2.1.278 does not support this model; version 2.1.280 or newer is required") instead of the usage-limit copy.
With both, a bad model setting fails once with the correct reason, and no subscription state changes.
Acceptance criteria
docker/base-image/Dockerfile installs a pinned Claude Code version; changing it rebuilds that layer.
A run rejected with unrecognized_model produces no subscription_auto_switch and no subscription_api_key_fallback.
The user-facing error for that run names the model / version problem, not a usage limit.
Regression test: a 503 carrying unrecognized_model classifies as non-auth.
Why P0
A single admin setting change silently rotated an agent through all its subscriptions and onto the platform API key, and every Workspace user got a wrong "usage limit, send again" instruction. The fallback does not undo itself when the setting is fixed.
Summary
Setting the platform default model to
claude-opus-5-5(selectable since #2987) broke every execution that does not pin a model, on a production instance. Two defects combined to make a single settings change look like a subscription outage:docker/base-image/Dockerfileinstalls@anthropic-ai/claude-code@latest. Docker caches that layer by its text, so a rebuild reuses whatever version was cached. The instance'strinity-agent-base:latestwas rebuilt on 2026-09-24 and still carries 2.1.278, although 2.1.280 was published on 2026-09-22 and 2.1.281 on 2026-09-23.claude-opus-5-5needs ≥ 2.1.280.[claude-code:unrecognized_model], the agent server answers HTTP 503, andclassify_switch_failure(services/execution_classification.py) maps 503 tofailure_kind="auth". SUB-003 then rotates the agent through every subscription and finally falls back to the platform API key (bug(subscriptions): a Workspace message on a rate-limited subscription fails instead of switching and completing — close the SUB-003 gaps on the turn path #2638). The Workspace user is told "The agent hit its usage limit, so it was moved onto the platform API key. Send that again and it should go through." (client_portal/service.py). Resending fails the same way.Evidence
One agent, 2026-09-24 (UTC):
platform_default_model = claude-opus-5-5via Settings (PUT /api/settings/platform_default_model)unrecognized_model→ SUB-003 switch #1 (failure_kind: auth)subscription_api_key_fallbackBackend log for one turn:
Reproduced outside Trinity,
claude -p --model claude-opus-5-5 "Reply with exactly: OK":API Error: 400 Claude Code 2.1.278 does not support this model; version 2.1.280 or newer is required.OKThe API's 400 names the fix, but that sentence never reaches the backend; the stored error holds only the
[claude-code:unrecognized_model]marker.Proposed fix
1. Pin Claude Code in the base image.
ARG CLAUDE_CODE_VERSION=2.1.281and install@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}. Bumping the value invalidates the layer, so a rebuild really upgrades, and the bump can ship in the same PR as any model-catalog addition that needs it.2. Treat a model rejection as its own failure, not auth.
unrecognized_model(and the API'sdoes not support this model400) to a distincterror_code(e.g.MODEL_UNSUPPORTED), carrying the API's message text, instead of a 503.classify_switch_failuremust returnNonefor it, so SUB-003 does not switch subscriptions or fall back to the API key.With both, a bad model setting fails once with the correct reason, and no subscription state changes.
Acceptance criteria
docker/base-image/Dockerfileinstalls a pinned Claude Code version; changing it rebuilds that layer.unrecognized_modelproduces nosubscription_auto_switchand nosubscription_api_key_fallback.unrecognized_modelclassifies as non-auth.Why P0
A single admin setting change silently rotated an agent through all its subscriptions and onto the platform API key, and every Workspace user got a wrong "usage limit, send again" instruction. The fallback does not undo itself when the setting is fixed.