Repository navigation
Cache breakpoint placed on a thinking block permanently 400s long Anthropic sessions (thinking.cache_control: Extra inputs are not permitted) #51141
Description
Activity
I'd like to work on this.
Confirmed the mechanism in the tree: the auto cache policy places its breakpoint on the last content part of the chosen message when that message has no text part.
markMessageAtfalls back tocontent.length - 1(packages/llm/src/cache-policy.ts, mirrored aspackages/ai/src/cache-policy.tson v2), so for an assistant turn whose trailing block isreasoningthe marker is stamped onto a block Anthropic does not acceptcache_controlon. The chosen index is stable, so the session stays poisoned exactly as reported.Plan:
- make the picker walk past
thinking/redacted_thinkingcandidates to the nearest legally cacheable block (text, tool_use, tool_result, image, document) on both auto paths - add regression coverage in
packages/llm/test/cache-policy.test.tsand a wire-level assertion next to the existing cases inpackages/opencode/test/provider/transform.test.ts, in the shape of the Bedrock fix fix(provider): filter unreplayable Bedrock reasoning before caching #45769
#31687is the same root cause on the Bedrock surface, so the fix should not reintroduce it there.Will target
devunless you would rather see it onv2.- make the picker walk past
Reporter here — I read
packages/llm/src/cache-policy.tsondevand can confirm your diagnosis is correct:markMessageAtmarks the lasttextpart, and when the message has no text part it falls back tocontent.length - 1regardless of type. An assistant message whose trailing part isreasoninggets the marker stamped on a block the API rejectscache_controlon, and since the chosen index is stable the session stays poisoned. Mirror confirmed present onv2(packages/ai/src/cache-policy.ts) as well.One edge case your plan should cover explicitly: the production failure path was
messages.660.content.0.thinking.cache_control— block 0 with thelength - 1fallback means the message content was a single thinking block. In that shape there is no legally cacheable block within the message to walk to, so a within-message walk-past is not sufficient. The picker needs to either move the breakpoint to an adjacent message whose last markable block is legal, or omit that breakpoint entirely (fewer breakpoints beats a guaranteed 400). For reference, microsoft/amplifier-module-provider-anthropic#99 fixed the identical failure by making its breakpoint index picker walk past whole messages whose candidate block isthinking/redacted_thinking.Suggested regression shapes:
[reasoning, text]— marks the text part (legal; already works viafindLastIndex; worth pinning so the fix doesn't regress it)[tool_use, reasoning](trailing reasoning, no text) — fallback marks the reasoning part today (fails); fixed picker should walk back within the message and marktool_use[reasoning]only — the production shape (fails today); no markable block exists in the message, so it must not be marked at all (skip to adjacent message or drop the breakpoint)
On
devvsv2: the bug is live in both mirrors, so from the reporter's side dev-first (matching the #45769 Bedrock precedent) with a v2 port sounds right — maintainers' call. Thanks for picking this up.Checked on V2: thinking and redacted-thinking blocks are sent without cache markers, including when the cache picker selects them. The invalid request described here therefore isn’t produced by the tested V2 path.
The picker still needs improvement: selecting thinking and then discarding its marker misses a caching opportunity. We’ll address that separately as a cache optimization.
Closing for V2. Please upgrade and reopen if you still encounter this error, including your V2 version and reproduction details.
See here: https://opencode.ai/v2/docs
Description
In a long session using the built-in
anthropicprovider with extendedthinking, every request eventually fails with:
Anthropic permits
cache_controlon text/tool_use/tool_result/image/documentblocks but NOT on
thinkingorredacted_thinking. OpenCode's rollingconversation cache breakpoint lands on an assistant message whose first/last
content block is a thinking block once the history is long enough. Because the
chosen breakpoint index is stable, the session is then permanently poisoned:
every subsequent turn rebuilds the same invalid request and 400s. The only
user-visible recovery is abandoning the session.
Short sessions never hit it, which makes it look intermittent.
Steps to reproduce
anthropicprovider (API key auth), extended-thinking model.breakpoint's stable message is an assistant turn whose block 0 is
thinking.Expected
The breakpoint picker should walk past messages whose candidate block is a
thinking/redacted_thinkingblock and stamp the nearest legally cacheableblock instead (same approach as microsoft/amplifier-module-provider-anthropic
PR #99, which fixed the identical failure shape).
Environment
anthropicprovider,apiKeyauth("Cache point cannot be inserted after reasoning block").