Skip to content

docs(spec): authenticated web assessment via browser takeover - #1107

Merged
philmerrell merged 1 commit into
developfrom
feature/authenticated-web-assessment-spec
Sep 15, 2026
Merged

philmerrell merged 1 commit into
developfrom
feature/authenticated-web-assessment-spec

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Why

Two user requests are blocked on the same missing capability:

"in terms of sharing an agent with the Library for them to take over the accessibility evaluation of all the databases they annually renew, they'd need the agent to be able to scan non-public web apps."

"I'm also creating an agent that evaluates VPATs, which may require testing against demo versions of apps that aren't public."

browse_web can only read what a logged-out, AWS-hosted Chromium can reach. There is no path today for a user to authenticate a browsing session — and the obvious workaround (let the user paste credentials so the agent types them) is the one option that must never ship: the secret would land in the prompt, in AgentCore Memory, and in the cacheable prefix, re-read every turn for the life of the conversation.

Spec only — no code. Sized for four PRs.

What it specs

Human-in-the-loop takeover. The agent calls take_control() (UpdateBrowserStream → automation stream DISABLED — a service-side mutex, not a convention in our code), raises an interrupt, the user signs in themselves in an embedded AWS DCV live view of that exact session, and the turn resumes authenticated. Credentials never enter the conversation.

Plus browser profiles (so a login is once-per-vendor, not once-per-conversation — this is what makes an annual sweep viable), a second VPC-mode browser resource with session recording for VPAT evidence, and axe-core as a browser extension.

Nearly all of the capability already exists and is unused. take_control / release_control ship in the pinned bedrock-agentcore==1.21.0, and the Runtime execution role is already granted UpdateBrowserStream and ConnectBrowserLiveViewStream.

Two findings worth acting on regardless of whether this ships

1. The existing live_view action is broken in the way PR #1101 diagnosed. generate_live_view_url signs with SigV4QueryAuth, so the signature lives in the query string, and browse_tool.py:288 returns that URL as text in a tool result — the model will re-emit it truncated at the ?. Same failure as the .docx download link. Max expiry is also 300 seconds, far too short for a human to find a password and log in.

2. RBAC granularity is exactly one tool_id. An earlier draft made takeover an action on browse_web; that is not grantable, so it would ship to everyone who can browse. Backwards — students should browse, and should not be able to drive a browser inside our AWS account. D1 now makes takeover and accessibility_scan separate registered tools with their own catalog entries, granted independently like ask_user_question.

Cost

Measured against real serialized tool specs:

Tool Prefix tokens
browse_web (today) 493
request_user_login (new) ~206
accessibility_scan (new) ~191

Separate tools means an ungranted user pays zero — and live_view leaves browse_web's enum, so the marginal cost for a student is slightly negative. A complete human login round trip costs about one short tool result: the live-view URL is never model-visible, the DCV stream is video the model never sees, and the full axe report goes to the workspace.

The real spend for these agents is screenshots (vision tokens) and accumulated page text at 2k tokens per extract_text. Hence D9: a 40-target sweep must be 40 sessions, not one turn — otherwise compaction fires repeatedly and makes a nearly-free feature look expensive.

Open questions for review

  1. Is networkConfiguration create-only on CfnBrowserCustom? Affects PR 4's shape.
  2. Do profiles survive a browser resource replacement? If they're scoped to the browser id, a PR-4 deploy could invalidate every saved login — verify before shipping PR 3.
  3. Does the DCV viewer work in a cross-origin iframe under our CSP? It opens a WebSocket and may need connect-src allowances the MCP-sandbox policy doesn't grant.
  4. bedrock-agentcore (npm) is a new dependency and needs explicit approval before install.
  5. Session recording captures the user typing into a password field. DCV masks nothing — needs disclosure in the UI before handoff, and the same bucket treatment as a credential store.

🤖 Generated with Claude Code

Two user requests are blocked on the same gap: a Library agent that
evaluates the accessibility of subscription databases it renews annually,
and a VPAT agent that must test vendor demo tenants. Neither target is
reachable by a logged-out browser, and there is no path today for a user
to authenticate a browsing session.

Specs the human-in-the-loop takeover flow: the agent calls take_control()
(UpdateBrowserStream -> automation DISABLED, a service-side mutex, not a
convention), raises an interrupt, the user signs in themselves in an
embedded AWS DCV live view, and the turn resumes authenticated. Credentials
never enter the prompt, the conversation, or AgentCore Memory.

Nearly all of the capability already exists and is unused: take_control /
release_control ship in the pinned bedrock-agentcore 1.21.0, and the
Runtime role is already granted UpdateBrowserStream and
ConnectBrowserLiveViewStream. Also covers browser profiles (so a login is
once-per-vendor, not once-per-conversation), a second VPC-mode browser
resource with session recording, and axe-core as an extension.

Two findings worth flagging independently of whether this ships:

- The existing `live_view` action is broken in the way PR #1101 diagnosed.
  generate_live_view_url signs with SigV4QueryAuth, so the signature is in
  the query string, and browse_tool.py:288 returns it as text in a tool
  result -- the model will re-emit it truncated at the `?`. Max expiry is
  300s, which is also far too short for a human login.

- RBAC granularity is exactly one tool_id, so takeover must be its own
  registered tool rather than a browse_web action (D1). As an action it
  would ship to everyone who can browse, which is backwards: students
  should browse and should not be able to drive a browser inside our AWS
  account. As separate tools, an ungranted user pays zero prefix tokens
  for them.

Measured prefix cost: browse_web 493 tokens today, request_user_login
~206, accessibility_scan ~191. A full human login round trip costs about
one short tool result. The real spend for these agents is screenshots and
accumulated page text, which is why D9 requires a 40-target sweep to be 40
sessions rather than one turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit 51fdbc7 into develop Sep 15, 2026
6 checks passed
@philmerrell
philmerrell deleted the feature/authenticated-web-assessment-spec branch September 15, 2026 01:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant