feat: ask_user_question — structured clarifying questions via mid-turn interrupt (PR-1) - #1100
Merged
Merged
Conversation
Lets the agent pause a turn to ask the user 1-4 structured multiple-choice questions, then continue in place with their answer as the tool's result. Backend only: the tool, its Strands interrupt, the `user_question_required` SSE event and the reload breadcrumb. The SPA renderer is PR-2, so the catalog seed ships `enabledByDefault: False` — a paused turn with no picker is a worse failure than a guessed assumption. The pause *is* the tool here, so it raises its own interrupt through `ToolContext` rather than from a `BeforeToolCall` hook the way OAuth consent and per-tool approval do. That flavor was unproven on the resume path, so `test_user_question_interrupt_integration.py` drives the real Strands event loop and pins it: the turn pauses, `pending_tool_execution` is captured, the interrupt id is tool-scoped, and resuming feeds the answer back into the same tool call without re-invoking it. Two asymmetric contracts, deliberately: - questions come from the model and are validated strictly, at the tool boundary, while it still has a turn left to fix them — emitting an unrenderable prompt would pause the turn with nothing on screen; - answers come from the client and are parsed leniently — that payload is the user's only route out of a paused turn, so a shape mismatch must never strand them. Cost: the tool spec is a constant in the cacheable `toolConfig` prefix (static registry tool, not an `extra_tools` injection, so it also avoids the injected-tool agent-cache bypass). Questions ride the SSE channel; only the one-line formatted answer block re-enters the conversation. Gated by `ASK_USER_QUESTION_ENABLED`, default on with a kill switch. Spec: docs/specs/ask-user-question.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both found by validating against real Bedrock (Haiku 4.5 + Sonnet 4.6 on
dev-ai), not by the unit tests.
Haiku supplied its own "Other" option on roughly half of sampled turns despite
the tool description saying not to. That is worse than a duplicate chip: the
picker's own Other carries a free-text field, so selecting a model-supplied
lookalike records the bare string "Other" and teaches the model nothing.
Stripped in `normalize_questions` rather than trusted away — a description is
guidance, not a guarantee. A question that falls below the two-option minimum
as a result is dropped: the model padded a binary choice with an escape hatch
the interface already provides.
Haiku also sometimes opened a second round of questions after receiving the
first answers ("Great! To finalize the details:"). That is legal — each round
is a fresh interrupt id and resumes correctly — but it reads as an
interrogation, so the docstring now tells the model to ask once and get on
with the work.
Live validation now passes 42/42 across both models: asks on an ambiguous
request, does not over-trigger on unambiguous ones, emits one well-formed
`user_question_required` frame, resumes into the same tool call, and the skip
path lets the model proceed instead of stalling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Sep 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Backend for a Claude-style follow-up questions UI: the agent pauses a turn to ask the user 1–4 structured multiple-choice questions, then continues in place with their answer as the tool's result.
PR 1 of 4 — backend only. No SPA renderer yet, so the catalog seed ships
enabledByDefault: False. A paused turn with no picker is a worse failure than a guessed assumption.Spec:
docs/specs/ask-user-question.mdWhy this shape
The mechanism already existed. Per-tool approval (
MCPExternalApprovalHook) is this same thing with one question and two fixed options — a Strands interrupt pauses the turn, an SSE event drives an inline prompt, and the decision POSTs back as aninterrupt_responsesentry. Generalizing the payload was cheaper than a second pause mechanism, and it inherits the paused-turn snapshot, the resume guard and the reload breadcrumbs for free.An MCP App (
ui_resource) was considered and rejected: App frames render alongside a running turn, they don't pause it.The one real risk, and how it was closed
Every interrupt shipped so far is raised from a
BeforeToolCallEventhook. This one is raised by the tool itself viaToolContext(which implements_Interruptible), because here the pause is the tool. The resume path —PausedTurnSnapshot→ rebuilt agent → restored_interrupt_state→pending_tool_executionreplayed — was built and proven against the hook flavor only.test_user_question_interrupt_integration.pydrives the real Strands event loop (realAgent, real@tool, scripted model only) and pins the tool-scoped flavor: the turn pauses,pending_tool_executionis captured, the id isv1:tool_call:{toolUseId}:…, resuming feeds the answer back into the same tool call, and the tool is not re-invoked.Two asymmetric contracts, deliberately
Cost
toolConfigprefix — registered once at startup, never injected per-turn. Do not make its presence conditional on conversation state.extra_toolsinjection, so it captures no identity and doesn't touch the injected-tool agent-cache bypass.Validation
Unit: 82 new tests; full backend suite green (8,540), ruff clean.
Live against real Bedrock on dev-ai (Haiku 4.5 + Sonnet 4.6), 42/42: asks on an ambiguous request, does not over-trigger on unambiguous ones, emits one well-formed
user_question_requiredframe with a tool-scoped id, resumes into the same tool call, and the skip path lets the model proceed rather than stall.Two defects the unit tests could not have found, fixed in c60b91c:
"Other"option on ~half of sampled turns despite the description. Worse than a duplicate — the picker's Other carries a free-text field, a model-supplied one records the bare string. Now stripped server-side.Regression: a normal chat turn through the local stack streams and completes unchanged (
POST /chat/stream → 200), with the new extractor in thedonepath.Follow-ups
UserQuestionService+ stream-parser wiring + single-question single-select prompt; generalizeresumeFromToolApprovalto carry an object. Ships usable.enabledByDefault, description tuning, RBAC grants.🤖 Generated with Claude Code