Skip to content

feat: ask_user_question — structured clarifying questions via mid-turn interrupt (PR-1) - #1100

Merged
philmerrell merged 2 commits into
developfrom
feature/ask-user-question-interrupt
Sep 14, 2026
Merged

philmerrell merged 2 commits into
developfrom
feature/ask-user-question-interrupt

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Backend for a Claude-style follow-up questions UI: the agent pauses a turn to ask the user 1–4 structured multiple-choice questions, then continues in place with their answer as the tool's result.

PR 1 of 4 — backend only. No SPA renderer yet, so the catalog seed ships enabledByDefault: False. A paused turn with no picker is a worse failure than a guessed assumption.

Spec: docs/specs/ask-user-question.md

Why this shape

The mechanism already existed. Per-tool approval (MCPExternalApprovalHook) is this same thing with one question and two fixed options — a Strands interrupt pauses the turn, an SSE event drives an inline prompt, and the decision POSTs back as an interrupt_responses entry. Generalizing the payload was cheaper than a second pause mechanism, and it inherits the paused-turn snapshot, the resume guard and the reload breadcrumbs for free.

An MCP App (ui_resource) was considered and rejected: App frames render alongside a running turn, they don't pause it.

The one real risk, and how it was closed

Every interrupt shipped so far is raised from a BeforeToolCallEvent hook. This one is raised by the tool itself via ToolContext (which implements _Interruptible), because here the pause is the tool. The resume path — PausedTurnSnapshot → rebuilt agent → restored _interrupt_state → pending_tool_execution replayed — was built and proven against the hook flavor only.

test_user_question_interrupt_integration.py drives the real Strands event loop (real Agent, real @tool, scripted model only) and pins the tool-scoped flavor: the turn pauses, pending_tool_execution is captured, the id is v1:tool_call:{toolUseId}:…, resuming feeds the answer back into the same tool call, and the tool is not re-invoked.

Two asymmetric contracts, deliberately

  • Questions come from the model → validated strictly, at the tool boundary, while it still has a turn left to fix them. Emitting an unrenderable prompt would pause the turn with nothing on screen.
  • Answers come from the client → parsed leniently (keyed, positional, bare string, junk). That payload is the user's only route out of a paused turn, so a shape mismatch must never strand them.

Cost

  • The tool spec is a constant in the cacheable toolConfig prefix — registered once at startup, never injected per-turn. Do not make its presence conditional on conversation state.
  • It is a static registry tool, not an extra_tools injection, so it captures no identity and doesn't touch the injected-tool agent-cache bypass.
  • Questions ride the SSE channel; only the one-line formatted answer block re-enters the conversation.

Validation

Unit: 82 new tests; full backend suite green (8,540), ruff clean.

Live against real Bedrock on dev-ai (Haiku 4.5 + Sonnet 4.6), 42/42: asks on an ambiguous request, does not over-trigger on unambiguous ones, emits one well-formed user_question_required frame with a tool-scoped id, resumes into the same tool call, and the skip path lets the model proceed rather than stall.

Two defects the unit tests could not have found, fixed in c60b91c:

  • Haiku supplied its own "Other" option on ~half of sampled turns despite the description. Worse than a duplicate — the picker's Other carries a free-text field, a model-supplied one records the bare string. Now stripped server-side.
  • Haiku sometimes opened a second round of questions after the first answers. Legal (fresh interrupt id, resumes fine) but reads as an interrogation; the docstring now says ask once.

Regression: a normal chat turn through the local stack streams and completes unchanged (POST /chat/stream → 200), with the new extractor in the done path.

Follow-ups

PR Scope
2 UserQuestionService + stream-parser wiring + single-question single-select prompt; generalize resumeFromToolApproval to carry an object. Ships usable.
3 Multi-select, multi-question carousel, "Other", "Skip", reload rehydration.
4 Flip enabledByDefault, description tuning, RBAC grants.

🤖 Generated with Claude Code

philmerrell and others added 2 commits September 14, 2026 09:16
Lets the agent pause a turn to ask the user 1-4 structured multiple-choice
questions, then continue in place with their answer as the tool's result.

Backend only: the tool, its Strands interrupt, the `user_question_required`
SSE event and the reload breadcrumb. The SPA renderer is PR-2, so the catalog
seed ships `enabledByDefault: False` — a paused turn with no picker is a worse
failure than a guessed assumption.

The pause *is* the tool here, so it raises its own interrupt through
`ToolContext` rather than from a `BeforeToolCall` hook the way OAuth consent
and per-tool approval do. That flavor was unproven on the resume path, so
`test_user_question_interrupt_integration.py` drives the real Strands event
loop and pins it: the turn pauses, `pending_tool_execution` is captured, the
interrupt id is tool-scoped, and resuming feeds the answer back into the same
tool call without re-invoking it.

Two asymmetric contracts, deliberately:
- questions come from the model and are validated strictly, at the tool
  boundary, while it still has a turn left to fix them — emitting an
  unrenderable prompt would pause the turn with nothing on screen;
- answers come from the client and are parsed leniently — that payload is the
  user's only route out of a paused turn, so a shape mismatch must never
  strand them.

Cost: the tool spec is a constant in the cacheable `toolConfig` prefix (static
registry tool, not an `extra_tools` injection, so it also avoids the
injected-tool agent-cache bypass). Questions ride the SSE channel; only the
one-line formatted answer block re-enters the conversation.

Gated by `ASK_USER_QUESTION_ENABLED`, default on with a kill switch.

Spec: docs/specs/ask-user-question.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both found by validating against real Bedrock (Haiku 4.5 + Sonnet 4.6 on
dev-ai), not by the unit tests.

Haiku supplied its own "Other" option on roughly half of sampled turns despite
the tool description saying not to. That is worse than a duplicate chip: the
picker's own Other carries a free-text field, so selecting a model-supplied
lookalike records the bare string "Other" and teaches the model nothing.
Stripped in `normalize_questions` rather than trusted away — a description is
guidance, not a guarantee. A question that falls below the two-option minimum
as a result is dropped: the model padded a binary choice with an escape hatch
the interface already provides.

Haiku also sometimes opened a second round of questions after receiving the
first answers ("Great! To finalize the details:"). That is legal — each round
is a fresh interrupt id and resumes correctly — but it reads as an
interrogation, so the docstring now tells the model to ask once and get on
with the work.

Live validation now passes 42/42 across both models: asks on an ambiguous
request, does not over-trigger on unambiguous ones, emits one well-formed
`user_question_required` frame, resumes into the same tool call, and the skip
path lets the model proceed instead of stalling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit 2e6108f into develop Sep 14, 2026
6 checks passed
@philmerrell
philmerrell deleted the feature/ask-user-question-interrupt branch September 14, 2026 16:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant