Skip to content

Cache breakpoint placed on a thinking block permanently 400s long Anthropic sessions (thinking.cache_control: Extra inputs are not permitted) #51141

Description

@anguslab-gh

Description

In a long session using the built-in anthropic provider with extended
thinking, every request eventually fails with:

Anthropic API error: 400 {"type":"error","error":{"type":"invalid_request_error",
"message":"messages.660.content.0.thinking.cache_control: Extra inputs are not
permitted"},"request_id":"req_011CfLwdUDbS6AEZanP2zGVJ"}

Anthropic permits cache_control on text/tool_use/tool_result/image/document
blocks but NOT on thinking or redacted_thinking. OpenCode's rolling
conversation cache breakpoint lands on an assistant message whose first/last
content block is a thinking block once the history is long enough. Because the
chosen breakpoint index is stable, the session is then permanently poisoned:
every subsequent turn rebuilds the same invalid request and 400s. The only
user-visible recovery is abandoning the session.

Short sessions never hit it, which makes it look intermittent.

Steps to reproduce

  1. Built-in anthropic provider (API key auth), extended-thinking model.
  2. Run a session long enough (~hundreds of messages) that the cache
    breakpoint's stable message is an assistant turn whose block 0 is
    thinking.
  3. Every turn from then on fails with the 400 above; retrying does not help.

Expected

The breakpoint picker should walk past messages whose candidate block is a
thinking/redacted_thinking block and stamp the nearest legally cacheable
block instead (same approach as microsoft/amplifier-module-provider-anthropic
PR #99, which fixed the identical failure shape).

Environment

Activity

  1. argszero commented on Sep 29, 2026

    @argszero
    Contributor

    I'd like to work on this.

    Confirmed the mechanism in the tree: the auto cache policy places its breakpoint on the last content part of the chosen message when that message has no text part. markMessageAt falls back to content.length - 1 (packages/llm/src/cache-policy.ts, mirrored as packages/ai/src/cache-policy.ts on v2), so for an assistant turn whose trailing block is reasoning the marker is stamped onto a block Anthropic does not accept cache_control on. The chosen index is stable, so the session stays poisoned exactly as reported.

    Plan:

    • make the picker walk past thinking / redacted_thinking candidates to the nearest legally cacheable block (text, tool_use, tool_result, image, document) on both auto paths
    • add regression coverage in packages/llm/test/cache-policy.test.ts and a wire-level assertion next to the existing cases in packages/opencode/test/provider/transform.test.ts, in the shape of the Bedrock fix fix(provider): filter unreplayable Bedrock reasoning before caching #45769

    #31687 is the same root cause on the Bedrock surface, so the fix should not reintroduce it there.

    Will target dev unless you would rather see it on v2.

  2. anguslab-gh commented on Sep 29, 2026

    @anguslab-gh
    Author

    Reporter here — I read packages/llm/src/cache-policy.ts on dev and can confirm your diagnosis is correct: markMessageAt marks the last text part, and when the message has no text part it falls back to content.length - 1 regardless of type. An assistant message whose trailing part is reasoning gets the marker stamped on a block the API rejects cache_control on, and since the chosen index is stable the session stays poisoned. Mirror confirmed present on v2 (packages/ai/src/cache-policy.ts) as well.

    One edge case your plan should cover explicitly: the production failure path was messages.660.content.0.thinking.cache_control — block 0 with the length - 1 fallback means the message content was a single thinking block. In that shape there is no legally cacheable block within the message to walk to, so a within-message walk-past is not sufficient. The picker needs to either move the breakpoint to an adjacent message whose last markable block is legal, or omit that breakpoint entirely (fewer breakpoints beats a guaranteed 400). For reference, microsoft/amplifier-module-provider-anthropic#99 fixed the identical failure by making its breakpoint index picker walk past whole messages whose candidate block is thinking/redacted_thinking.

    Suggested regression shapes:

    1. [reasoning, text] — marks the text part (legal; already works via findLastIndex; worth pinning so the fix doesn't regress it)
    2. [tool_use, reasoning] (trailing reasoning, no text) — fallback marks the reasoning part today (fails); fixed picker should walk back within the message and mark tool_use
    3. [reasoning] only — the production shape (fails today); no markable block exists in the message, so it must not be marked at all (skip to adjacent message or drop the breakpoint)

    On dev vs v2: the bug is live in both mirrors, so from the reporter's side dev-first (matching the #45769 Bedrock precedent) with a v2 port sounds right — maintainers' call. Thanks for picking this up.

  3. rekram1-node commented on Sep 30, 2026

    @rekram1-node
    Collaborator

    Checked on V2: thinking and redacted-thinking blocks are sent without cache markers, including when the cache picker selects them. The invalid request described here therefore isn’t produced by the tested V2 path.

    The picker still needs improvement: selecting thinking and then discarding its marker misses a caching opportunity. We’ll address that separately as a cache optimization.

    Closing for V2. Please upgrade and reopen if you still encounter this error, including your V2 version and reproduction details.

    See here: https://opencode.ai/v2/docs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions