Skip to content

fix(server): soft-accept OpenAI response_format json_object/json_schema - #604

Open
berdachuk wants to merge 3 commits into
FlashML-org:mainfrom
berdachuk:soft-accept-response-format
Open

berdachuk wants to merge 3 commits into
FlashML-org:mainfrom
berdachuk:soft-accept-response-format

Conversation

@berdachuk

Copy link
Copy Markdown

Summary

  • Soft-accept OpenAI response_format types json_object and json_schema instead of returning HTTP 400 no constrained decoding.
  • Prepend a best-effort system hint so the model is steered toward JSON; this is not token-level constrained/guided decoding.
  • Keep unknown response_format.type values rejected, with a clearer error message.

OpenAI-compatible clients (structured outputs / JSON mode) currently break against FreeToken even when the model can emit valid JSON from a prompt. This unblocks that wire format without claiming full guided decoding.

Test plan

  • POST /v1/chat/completions without response_format still returns 200 (no regression).
  • POST /v1/chat/completions with "response_format": {"type": "json_object"} returns 200 and JSON content.
  • Same with "type": "json_schema" (+ simple object schema) returns 200 with top-level fields matching the schema (best-effort).
  • Unknown response_format.type still returns 400 with the updated error text.
  • Optional: stream path with json_object still completes.

Verified in production against FreeToken 0.1.3 / Qwen3.6-35B-A3B (chat completions via an OpenAI gateway): baseline, json_object, and json_schema all returned HTTP 200.

Siarhei Berdachuk added 3 commits October 5, 2026 19:11
Rejecting response_format with "no constrained decoding" breaks OpenAI-compatible
clients that rely on json_object/json_schema even when the model can emit valid
JSON from a prompt. Accept those wire types and prepend a best-effort system hint;
schema is not token-level enforced. Unknown response_format types still return 400.
After soft-accepting json_object/json_schema, the old "no constrained decoding"
message was misleading for remaining unknown types.
Soft-accepting json_object/json_schema is enough when clients already
request JSON in the prompt. The injected schema wrapper was redundant
and could distort the output shape.
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 6, 2026
FlashML-org#604 rejected (accepts JSON response formats without enforcing them); FlashML-org#605 deferred
(pre-sm70 only); drafts FlashML-org#614, FlashML-org#615, FlashML-org#616 deferred with what next already covers.

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant