fix(agents): an @-mention binds the conversation instead of borrowing one turn - #1115
Merged
Merged
Conversation
… one turn
Mentioning an Agent ran that turn as the Agent and silently reverted the next
one. The thread still looked like the Agent's while its tools, skills and model
were gone, and nothing surfaced the change — not the UI, and not the model,
which cannot know its own toolset shrank. Asked to use a tool it had used a
moment earlier it got `Unknown tool: create_rubric`, and told the user to toggle
that tool in the picker: a confident wrong diagnosis sending them to fix a
setting that was already correct.
D11 chose the per-turn reading deliberately, so this reverses a decision rather
than repairing an oversight. What changed is the evidence. Measured on prod
sessions-metadata (`turnAgentId` on the `C#` rows): of **247 mentions, 247
started the conversation**. Zero were mid-thread consults; zero mentioned a
second Agent inside a thread bound to a first. (Dev: 60 of 61.) The borrow was
paying an invisible failure mode for a case that has never occurred.
A mention now *means* "talk to this Agent", with two outcomes and no third:
* empty thread -> the Agent binds the conversation, exactly like launching
it from its card. Safe because there is no history for the
binding to misrepresent.
* has messages -> the message opens a NEW conversation with that Agent, and
the SPA says so. The Agent cannot be bound to history
written under other instructions, and must not be borrowed.
Persisting `preferences.assistant_id` is necessary but NOT sufficient, which is
the part that is easy to miss: every turn's Agent is resolved from the request,
the SPA's only carrier is the `assistantId` query param, and the self-heal
effect that refills it from preferences runs on session *load*. So the SPA now
sets the param and stops sending `agent_mention` at all. The backend still
honours that flag for clients that predate this change, and `binds_conversation`
gains `thread_is_empty` so a stale tab mentioning into a fresh thread lands
where a current one does. Its thread lookup runs only for mention turns, so no
bound-Agent turn pays a query it cannot act on.
Also fixed, because it is the same invisible loss on the path users are *told*
to use: "Continue" after a max_tokens truncation skipped the whole assistant
block, so a properly launched Agent finished its reply with none of its tools,
skills, model or instructions (spec-acknowledged as a known edge). The SPA was
already resending `rag_assistant_id` there — `continueTruncatedTurn`'s own
comment says "so the backend rebuilds the same model/tools/assistant agent" —
and only the `not is_continuation` guard discarded it. The block now runs for a
continuation, with binding validation and persistence skipped (it binds nothing
new) and RAG skipped (the turn carries an empty message, so a KB search would
spend a query on "" and augment nothing).
A resume still skips the block and always did keep its tools: it rebuilds from
`PausedTurnSnapshot`, replaying the original turn's exact enabled_tools /
system_prompt / enabled_skills to reconstruct the same prompt-cache key.
Re-resolving there would risk a different effective set and orphan the paused
agent. Worth stating because resume rows carry no `turnAgentId`, so a census of
"turns with no Agent" reads them as losses and overcounts badly.
Two known costs retire with the borrow: the ~$0.12-per-mention prefix re-write
(a bound conversation swaps once and stays instead of swapping back), and the
history fork, where the mention agent and the plain agent were two cached
instances that never saw each other's turns.
The rule lives in one testable place at each end — `mention-routing.ts` on the
client, `agent_binding_policy.py` on the server — following the existing
`system_prompt_resolver` precedent: the rule is a handful of lines, the code
around it is a thousand, and a rule no test can reach is a rule that drifts.
Kept as one commit: the mention and continuation halves edit the same guards on
the same block, and splitting them would mean a first commit that knowingly
leaves the block wrong.
Backend 8585 passed / 3 skipped; frontend 3024 passed across 250 files; tsc
clean. Not exercised in a browser — this worktree's code is not what the local
stack serves, and that stack is down.
Spec: docs/specs/agent-marketplace.md (D11 + Phase 7 notes)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 task done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
Mentioning an Agent ran that turn as the Agent and silently reverted the next one. The thread still looked like the Agent's while its tools, skills and model were gone, and nothing surfaced the change — not the UI, and not the model, which cannot know its own toolset shrank.
Found from a real dev session: the Rubric Builder drafted a rubric, the user said "publish it", and
create_rubriccame backUnknown tool: create_rubric. The model then told the user to toggle that tool off and on in the picker — a confident wrong diagnosis sending them to fix a setting that was already correct. The catalog entry, the Agent's binding and the live MCP server were all fine; the turn simply had no Agent.The tell is content-free, on the session's
C#rows:turnAgentIdast-9149ef191614None(withagentSwitched=Trueon the first)Why reverse D11
D11 chose the per-turn borrow deliberately, so this reverses a decision rather than repairing an oversight. What changed is the evidence — measured on prod
sessions-metadata:The borrow was paying an invisible failure mode for a case that has never occurred.
What changes
A mention now means "talk to this Agent", with two outcomes and no third:
Persisting
preferences.assistant_idis necessary but not sufficient, which is the easy part to miss: every turn's Agent is resolved from the request, the SPA's only carrier is theassistantIdquery param, and the self-heal effect that refills it from preferences runs on session load. So the SPA sets the param and stops sendingagent_mentionentirely.The backend still honours that flag for clients predating this change, and
binds_conversationgainsthread_is_emptyso a stale tab mentioning into a fresh thread lands where a current one does. Its thread lookup runs only for mention turns, so no bound-Agent turn pays a query it cannot act on.Also fixed: "Continue" after a truncation
Same invisible loss, on the path users are told to use.
continue_truncatedskipped the whole assistant block, so a properly launched Agent finished its reply with none of its tools, skills, model or instructions — spec-acknowledged as a known edge. The SPA was already resendingrag_assistant_idthere (continueTruncatedTurn's comment says "so the backend rebuilds the same model/tools/assistant agent"); only that guard discarded it. The block now runs for a continuation, with binding validation and persistence skipped (it binds nothing new) and RAG skipped (empty message → a KB search would spend a query on"").PausedTurnSnapshot, replaying the original turn's exactenabled_tools/system_prompt/enabled_skillsto reconstruct the same prompt-cache key. Worth stating because resume rows carry noturnAgentId, so a census of "turns with no Agent" reads them as losses and overcounts badly (I overcounted 4,127 prod calls that way before noticing the field only exists since 2026-08).Cost
Two known costs retire with the borrow: the ~$0.12-per-mention prefix re-write (a bound conversation swaps once and stays instead of swapping back), and the history fork, where the mention agent and the plain agent were two cached instances that never saw each other's turns.
Testing
tsc --noEmitcleanmention-routing.tsandagent_binding_policy.py— following thesystem_prompt_resolverprecedent, and the two mirror each other@-mention an Agent as the first message, then send an ordinary follow-up and confirm it still reaches the Agent's tools — that is the exact sequence that failed.Spec:
docs/specs/agent-marketplace.md(D11 + Phase 7 notes)🤖 Generated with Claude Code