docs(spec): authenticated web assessment via browser takeover - #1107
Merged
Merged
Conversation
Two user requests are blocked on the same gap: a Library agent that evaluates the accessibility of subscription databases it renews annually, and a VPAT agent that must test vendor demo tenants. Neither target is reachable by a logged-out browser, and there is no path today for a user to authenticate a browsing session. Specs the human-in-the-loop takeover flow: the agent calls take_control() (UpdateBrowserStream -> automation DISABLED, a service-side mutex, not a convention), raises an interrupt, the user signs in themselves in an embedded AWS DCV live view, and the turn resumes authenticated. Credentials never enter the prompt, the conversation, or AgentCore Memory. Nearly all of the capability already exists and is unused: take_control / release_control ship in the pinned bedrock-agentcore 1.21.0, and the Runtime role is already granted UpdateBrowserStream and ConnectBrowserLiveViewStream. Also covers browser profiles (so a login is once-per-vendor, not once-per-conversation), a second VPC-mode browser resource with session recording, and axe-core as an extension. Two findings worth flagging independently of whether this ships: - The existing `live_view` action is broken in the way PR #1101 diagnosed. generate_live_view_url signs with SigV4QueryAuth, so the signature is in the query string, and browse_tool.py:288 returns it as text in a tool result -- the model will re-emit it truncated at the `?`. Max expiry is 300s, which is also far too short for a human login. - RBAC granularity is exactly one tool_id, so takeover must be its own registered tool rather than a browse_web action (D1). As an action it would ship to everyone who can browse, which is backwards: students should browse and should not be able to drive a browser inside our AWS account. As separate tools, an ungranted user pays zero prefix tokens for them. Measured prefix cost: browse_web 493 tokens today, request_user_login ~206, accessibility_scan ~191. A full human login round trip costs about one short tool result. The real spend for these agents is screenshots and accumulated page text, which is why D9 requires a 40-target sweep to be 40 sessions rather than one turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Two user requests are blocked on the same missing capability:
browse_webcan only read what a logged-out, AWS-hosted Chromium can reach. There is no path today for a user to authenticate a browsing session — and the obvious workaround (let the user paste credentials so the agent types them) is the one option that must never ship: the secret would land in the prompt, in AgentCore Memory, and in the cacheable prefix, re-read every turn for the life of the conversation.Spec only — no code. Sized for four PRs.
What it specs
Human-in-the-loop takeover. The agent calls
take_control()(UpdateBrowserStream→ automation streamDISABLED— a service-side mutex, not a convention in our code), raises an interrupt, the user signs in themselves in an embedded AWS DCV live view of that exact session, and the turn resumes authenticated. Credentials never enter the conversation.Plus browser profiles (so a login is once-per-vendor, not once-per-conversation — this is what makes an annual sweep viable), a second VPC-mode browser resource with session recording for VPAT evidence, and axe-core as a browser extension.
Nearly all of the capability already exists and is unused.
take_control/release_controlship in the pinnedbedrock-agentcore==1.21.0, and the Runtime execution role is already grantedUpdateBrowserStreamandConnectBrowserLiveViewStream.Two findings worth acting on regardless of whether this ships
1. The existing
live_viewaction is broken in the way PR #1101 diagnosed.generate_live_view_urlsigns withSigV4QueryAuth, so the signature lives in the query string, andbrowse_tool.py:288returns that URL as text in a tool result — the model will re-emit it truncated at the?. Same failure as the.docxdownload link. Max expiry is also 300 seconds, far too short for a human to find a password and log in.2. RBAC granularity is exactly one
tool_id. An earlier draft made takeover an action onbrowse_web; that is not grantable, so it would ship to everyone who can browse. Backwards — students should browse, and should not be able to drive a browser inside our AWS account. D1 now makes takeover andaccessibility_scanseparate registered tools with their own catalog entries, granted independently likeask_user_question.Cost
Measured against real serialized tool specs:
browse_web(today)request_user_login(new)accessibility_scan(new)Separate tools means an ungranted user pays zero — and
live_viewleavesbrowse_web's enum, so the marginal cost for a student is slightly negative. A complete human login round trip costs about one short tool result: the live-view URL is never model-visible, the DCV stream is video the model never sees, and the full axe report goes to the workspace.The real spend for these agents is screenshots (vision tokens) and accumulated page text at 2k tokens per
extract_text. Hence D9: a 40-target sweep must be 40 sessions, not one turn — otherwise compaction fires repeatedly and makes a nearly-free feature look expensive.Open questions for review
networkConfigurationcreate-only onCfnBrowserCustom? Affects PR 4's shape.connect-srcallowances the MCP-sandbox policy doesn't grant.bedrock-agentcore(npm) is a new dependency and needs explicit approval before install.🤖 Generated with Claude Code