Skip to content

fix(llm): preserve replay state and agent output budgets - #390

Draft
WH-2099 wants to merge 2 commits into
mainfrom
fix/preserve-opaque-invocations
Draft

WH-2099 wants to merge 2 commits into
mainfrom
fix/preserve-opaque-invocations

Conversation

@WH-2099

@WH-2099 WH-2099 commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Blocking ordinary and structured-output LLM invocations dropped provider replay snapshots and discarded content arrays between text chunks.
Preserve the last non-null opaque_body, retain mixed text and content blocks in order, and treat assistant state or tool calls as nonempty.
Each call starts with fresh state, including after an earlier call captured a snapshot.
Both invocation APIs reuse the same chunk merger while retaining structured output, tool calls, and usage.

Agent strategies now reject an exhausted context before invocation instead of reducing the configured output budget below a thinking budget.
The check uses the configured parameter, its template alias, or its schema default and preserves valid small output budgets; it does not mutate model parameters.
The Function Calling and ReAct caller update is included in official-plugins #3919.

Closes #389.
Related: langgenius/graphon#260 and langgenius/dify-official-plugins#3916.

Validation: just check, just test (303 passed, 1 skipped), and just build passed.
The skipped marketplace integration test requires the Dify CLI.
New regression tests use the real Session, request serialization, response parsing, and both stream modes of ordinary and structured-output invocation, with only the HTTP transport replaced.
They cover mixed content order, per-message and per-block JSON values and types, trailing empty chunks, and consecutive calls.
No provider API or deployed Dify environment was used.

Compatibility: the existing optional field and package version remain unchanged; no release version is invented by this change.
Agent contexts that cannot fit the configured or default output budget now fail explicitly instead of silently lowering the limit.
Token counts still use the existing local estimate.
Pure text still returns a string; responses containing content arrays now retain their blocks instead of dropping them.
Provider snapshots are internal replay state and are not appended to visible text.

Pull Request Checklist

Thank you for your contribution! Before submitting your PR, please make sure you have completed the following checks:

Compatibility Check

  • I have checked whether this change affects the backward compatibility of the plugin declared in README.md
  • I have checked whether this change affects the forward compatibility of the plugin declared in README.md
  • If this change introduces a breaking change, I have discussed it with the project maintainer and specified the release version in the README.md — not applicable; this restores existing message fields.
  • I have described the compatibility impact and the corresponding version number in the PR description — no release bump; current package metadata remains 0.10.2.
  • I have checked whether the plugin version is updated in the README.md — no version update is required for this fix.

Available Checks

  • just build has passed
  • Relevant documentation has been updated (if necessary) — no public configuration or schema changed.

@WH-2099 WH-2099 changed the title fix(llm): preserve opaque payloads in blocking invocation fix(llm): preserve replay state and agent output budgets Sep 24, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(llm): preserve opaque snapshots and content blocks in blocking invocation

1 participant