Skip to content

feat(queen): runners take tasks and hand the work back for review - #521

Merged
dmitrii-f-t27 merged 14 commits into
fix/queen-worker-provider-and-prompt-sizefrom
claude/runner-claim-complete
Oct 4, 2026
Merged

dmitrii-f-t27 merged 14 commits into
fix/queen-worker-provider-and-prompt-sizefrom
claude/runner-claim-complete

Conversation

@dmitrii-f-t27

Copy link
Copy Markdown
Collaborator

Phase 2 of runner lanes. It is stacked on #518, which registers runners, so review and merge that one first. GitHub will retarget this PR when #518 merges.

A runner created in the app.t27.ai cabinet can now be given a task. It runs the task on its owner's machine with its owner's key and hands back a pushed branch. The existing review judges that branch by the same rules as any container bee's work. The provider key never reaches this server.

How a task flows

  1. Offer. dispatchBee asks offerToRunner first, before any key, memory or disk is measured.

    • An idle runner gets a dispatch row on its lane (100000000 + id). The row is started = true and stores the brief, the system prompt and the start commit.
    • Nothing is started locally: no worktree, no /chat, and no provider or model is recorded. Because the row has no provider or model, the daily spend cap never prices the lender's spend.
  2. Which issues are offered. Only issues a runner can actually continue:

    • no branch yet;
    • a branch with nothing beyond the base;
    • the runner's own previous push, after a send-back.

    A branch with commits that exist only in the container, or a dirty worktree, stays with the container.

  3. Claim. POST /queen/runner/claim hands over the task. Every runner call renews the lease. Claiming again returns the same task, so a runner that restarts picks up where it left off.

  4. Complete. POST /queen/runner/complete with {conversationId, remoteUrl, branch, headSha, said}. The server:

    • fetches that one branch into a scratch ref, only from https://github.com/..., with protocol.allow=never and no credential helper;
    • checks the commit is the one named and descends from the offered start;
    • moves queen-<issue> and gives it a clean worktree, as a container bee would leave;
    • writes the agent's answer as the say transcript the review reads;
    • only then closes the row and fires the refill signal.

    A runner can also send {gaveUp: true} to release the issue immediately.

  5. Reap. Runner rows are skipped by the boot reaper and the stall reaper. Instead they are released by lease:

    • an offer unclaimed for 10 minutes;
    • a lease not renewed for 15 minutes;
    • any task older than 6 hours.

Effect on the round

  • A runner's in-flight row goes on the board as queued. queend already treats that state as a live claim that holds its issue and files, but it does not count it against canStartAnother. So runners add capacity rather than consuming container slots.
  • When queend refuses because the container is full, a runner-only pass re-asks it with local bees shown as queued. Each answer goes to an idle runner only. Boundaries, claims and the spend cap still apply.
  • Runner lanes are excluded from provider-key accounting:
    • the in-flight takenKeys and keyCursor;
    • runningKeys, used for the reviewer lane;
    • the public research telemetry.
  • QUEEN_RUNNERS=off turns offers off, and the swarm runs as it did before. RUNNER_PROTOCOL is now 2.

The runner script

trios/agent-server/tools/queen-runner/queen-runner.mjs plus a README. It needs Node 20+ and git, with no packages to install. It:

  • heartbeats and claims;
  • checks out the start commit in a worktree on queen-<issue>;
  • runs the volunteer's own agent command with the prompt on stdin. The default is claude -p --permission-mode acceptEdits;
  • commits whatever the agent left, pushes to the volunteer's public fork, and completes;
  • hands the task back on Ctrl-C, on an agent that failed and changed nothing, or when the swarm cannot take the branch.

Tests

What Result
bun run typecheck clean
bun run test:api 1359 pass, 59 skip, 0 fail
bun run test:pglive (local PostgreSQL 16) 5 pass, 0 fail
bun run test:root 68 pass, 0 fail
biome check on changed files no errors; the complexity warnings are the same kind the existing runRound and reviewFinishedDispatches already carry

New tests:

  • tests/api/queen-runner-work.test.ts. Uses real git in a temp directory. It covers:
    • which issues a runner may take;
    • the offer row;
    • that an offered task cuts no worktree and measures no container;
    • the runner-only path;
    • body validation: file://, ssh, ext::, other hosts, userinfo, --upload-pack, .. and abbreviated SHAs are all refused;
    • bringing a branch home, wrong SHA, unrelated history, and a checked-out tree;
    • completion ordering, 404 for someone else's task, give-up;
    • the routes.
  • tests/api/queen-round-runners.test.ts. Drives runRound against a stand-in policy that applies queend's capacity and claim rules. The real queend is not built in CI, so queen-round.test.ts skips there. Mutation-checked: removing the queued mapping fails one test, and treating runner dispatches as running fails another.
  • tests/pglive/queen-runner-work-live.test.ts. The whole life of a task on a migrated scratch database: register, idle, offer, claim, renew, complete, then a review-ready row with the verdict in its transcript. It also covers reaping.
  • tests/api/queen-runner-cli.test.ts. Runs the script end to end against a stand-in swarm with local upstream and fork repos, under Bun. It was also run locally under Node 22.

Two existing tests were adjusted:

  • queen-runners.test.ts now expects protocol 2.
  • The container memory-refusal test in queen-dispatch.test.ts passes offer: null. That test asserts that no queen_dispatch statement is sent, and the runner lookup is a read of that table.

Security notes

  • What a runner can trigger in the container. Only a fetch of one branch from github.com over https, pinned and timed out. The criteria runner already executes only an allowlisted read-only command set taken from the issue, as the bee's uid, in a scratch worktree. A runner's commit therefore adds no execution surface that a bee's commit doesn't already have.
  • Download size. Any signed-in person can mint a runner, and the container will fetch that runner's branch. The scope is one claimed task per runner, with a 300 s timeout and the existing volume guard. There is no explicit size limit yet. That is a follow-up if abuse appears.
  • The project URL. The URL a runner is given has credentials stripped from TRIOS_REPO_URL.

Not done

🤖 Generated with Claude Code

https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2


Generated by Claude Code

claude added 10 commits October 1, 2026 14:07
A signed-in person (app.t27.ai session, verified by asking its issuer's
whoami for telegram_id) can mint, list and revoke runner tokens at
/queen/me/runners. A runner process on the lender's own machine presents
that token at /queen/runner/heartbeat and learns its lane. The provider
key never reaches the Queen: no route accepts, stores or returns one.

- queen_runner table: only the token's SHA-256 and last four characters
  are kept; at most 5 live runners per person.
- A runner's lane is 100000000 + id, a key_index block no operator pool
  reaches, so dispatch rows on it are credited by the existing
  leaderboard arithmetic.
- Leaderboard gathers runner lanes by person (telegram_id), never by
  name, so a runner cannot merge into an operator's row, and never links
  a runner to an unverified GitHub login.
- CORS for /queen/me/* is exactly https://app.t27.ai, bearer only, no
  credentials; both new mounts are on the route-guard allowlist with an
  own-bearer reason, and the audit pins are re-measured.

Taking tasks and handing work back (claim/complete, remote review,
runner CLI) is the next stage; the heartbeat says so instead of
offering work.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…dules symlink

Every `Tests / *` job on #517 and #518 was cancelled at the 20-minute
budget while still inside `bun ci`, before any test ran.

Two things combined:
- trios/agent-server/apps/server/node_modules was committed as a symlink
  to a local macOS path. `.gitignore` said `node_modules/`, which only
  matches directories, so the symlink slipped through.
- setup-bun ran without a version. The `packageManager: bun@1.3.6` pin
  lives in trios/agent-server/package.json, not at the repo root, so CI
  got the latest release (1.4.2), which hangs on that dangling symlink.
  1.3.x installs past it.

Untrack the symlink, make the ignore rule match files too, and read the
Bun version from the workspace package.json in test.yml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com

With installs no longer hanging, server-tools ran for the first time
since 2026-09-23 and failed one test: get_page_content read
https://example.com 57 ms after opening it and found no "Example
Domain". The test is about extracting text, so it now writes that text
into about:blank with evaluate_script, as get_page_links already does.

Locally (BrowserOS AppImage, headless, --no-sandbox): the old test
fails the same way; the new one passes 3/3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… process

server-tools still exited 1 after every test in observation.test.ts
passed: before navigation-newtab-guard.test.ts the helper ran
`lsof -ti :<cdp port>`, which also lists clients still connected to the
port. One of them was the bun test process itself (its CDP socket to the
previous file's browser), so the SIGTERM ended the whole run and no
junit report was written ("workflow > server-tools setup").

Use `lsof -ti tcp:<port> -sTCP:LISTEN` and drop process.pid.

Locally, input.test.ts + navigation-newtab-guard.test.ts in one process:
before, exit 143 right after "Terminating process(es) <own pid>, ...";
after, 18 pass / 0 fail. The whole test:tools group now runs to the end
(242 pass; the 2 local failures load https://example.com, which this
sandbox's browser cannot reach and CI can).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com

With the run no longer killing itself, server-tools finished in CI with
243 pass / 1 fail: `wait_for finds text on page` waited its full 10 s
for "Example Domain" on https://example.com and never saw it - the same
page get_page_content could not read either.

The page now adds that text itself 500 ms after load, so the test still
proves wait_for waits, with nothing outside the runner involved.

Locally: 2/2 wait_for tests pass on repeat; the whole test:tools group
is 243 pass, the one local failure being take_screenshot (a 60 s hang
in this sandbox only - it passes in CI).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… too

server-tools on 3d57649 ran clean except one test that had passed on
both earlier runs: `search_dom > finds multiple elements with CSS class
selector` (123 ms, fewer than 3 matches). It searches once, straight
after new_page - the race this file already names and fixes with
searchUntil for two sibling tests. Use the same helper here.

Locally: search_dom 13/13, three runs in a row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…udget

`the salvage commit > never splits a rename across the path cap` runs
real git over 205 files and salvageWorktree. It takes ~2 s for the whole
file locally and passed on the two CI runs before, then hit bun's 5 s
default once on a loaded runner (job 110500921083) with nothing in the
change touching salvage. A git-heavy fixture test should not share the
budget of a pure unit test.

Locally: queen-salvage-guards.test.ts 13 pass / 0 fail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Phase 2 of the runner lanes. A runner registered in the cabinet can now be
given work, run it on its owner's machine with its owner's key, and hand
back a pushed branch that the existing review judges like any other bee.

- dispatchBee offers the chosen issue to an idle runner before the
  container is measured (offerToRunner). The dispatch row goes on the
  runner's lane with the brief, the system prompt and the start commit
  stored on it; no worktree, no /chat, no key, no provider/model.
- Only issues a runner can continue are offered: no branch yet, a branch
  with nothing beyond the base, or the runner's own previous push. A
  branch with container-only commits or a dirty worktree stays local.
- The round shows a runner's task as `queued`: a live claim that holds its
  issue and files but no container worker slot. When queend says the
  container is full, a runner-only pass re-asks with local bees shown as
  queued and hands further tasks to idle runners only.
- /queen/runner/claim hands over the task; every call renews the lease.
  /queen/runner/complete fetches the named branch from an https GitHub
  remote into a scratch ref, checks the commit and that it descends from
  the start, moves queen-<issue>, writes the agent's answer as the
  transcript the review reads, then closes the row.
- Runners are reaped by lease (unclaimed offer 10 min, silent lease
  15 min, cap 6 h), not by the container's boot or stall reapers.
- QUEEN_RUNNERS=off keeps the swarm local. RUNNER_PROTOCOL is now 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
A zero-dependency Node script (tools/queen-runner/queen-runner.mjs) that
takes the Queen's tasks on the volunteer's own machine:

- heartbeats, claims the task its lane was offered, checks the project
  out at the commit the task starts from (or its own previous push after
  a send-back), in a worktree on queen-<issue>;
- runs the volunteer's own agent command in the project directory with
  the prompt on stdin (default: claude -p --permission-mode acceptEdits),
  commits whatever the agent left, pushes to the volunteer's public fork
  and completes with the exact commit and the agent's answer;
- hands the task back on Ctrl-C, on an agent that fails and changes
  nothing, or when the swarm cannot take the branch.

The provider key stays wherever the volunteer's agent already keeps it.
An end-to-end test drives the script against a stand-in swarm with local
upstream and fork repositories.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Without one, a send-back the container takes instead of a runner cut a
fresh tree with `worktree add -B queen-<issue> ... base`, which reset the
branch and started the retry beside the runner's work rather than from
it. A runner's finished issue now looks like a container bee's: branch
and clean tree, which prepareWorktree reuses. Not fatal if it fails; the
review reads the branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

✅ Tests passed — 2596/2656

Suite Passed Failed Skipped
✅ agent 87/87 0 0
✅ build 9/9 0 0
✅ cdp-protocol 5/5 0 0
✅ eval 93/93 0 0
✅ server-agent 280/280 0 0
✅ server-api 1443/1502 0 59
✅ server-browser 6/6 0 0
✅ server-integration 10/11 0 1
✅ server-lib 279/279 0 0
✅ server-pglive 27/27 0 0
✅ server-root 68/68 0 0
✅ server-skills 31/31 0 0
✅ server-tools 244/244 0 0
✅ shared 14/14 0 0

View workflow run

Copy link
Copy Markdown
Collaborator Author

The cla check is red. The CLA Assistant action exits before it reads any signature because the PERSONAL_ACCESS_TOKEN secret is empty (Please add a personal access token as an environment variable...). This PR doesn't cause it, and it fails the same way on #518. No code fix exists: a maintainer needs to set that secret. Every test suite on this head is green (2482 passed, 0 failed, including server-pglive). I can't re-run checks here.


Generated by Claude Code

claude added 2 commits October 2, 2026 03:24
Conflicts, both additive:
- server.ts: keep the runner mounts and add /queen/contributor-keys.
- queen-leaderboard.ts: rank() takes the operator map merged with
  contributorOwnerNames(), and the runner owners beside it.

Also ports the route-guard fix from #520: #522 left
/queen/contributor-keys out of the audit, which turns four route-guard
tests red on the base. It is allowlisted with its own capability guard
and the pins are re-measured (48 mounts; 25 /queen: 8/8/9).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
One conflict, in dispatchBee: the runner offer still comes first, and
the container path then picks its key with #522's contributor runtime
(resolveWorkerProvider(..., runtime)). Contributor keys use negative
indices and runner lanes sit above 100000000, so the runner filters
(NOT_A_RUNNER, isRunnerLane) leave contributor keys with the container.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
claude added 2 commits October 3, 2026 20:43
Retargets the runner cabinet at the branch production builds from.
Conflicts, all additive:
- server.ts: runner mounts beside /queen/public-credits and
  /queen/public-earnings.
- pg-migrate.ts: queen_runner beside queen_tri_earnings.
- route-guard-audit.mjs: the runner allowlist entries beside the
  production branch's /queen/contributor-keys entry (same reason text).
- route-guard.test.ts: the production branch's pins plus the two runner
  mounts: 51 mounts; 28 /queen mounts as 10 public-read, 9 wrapper,
  9 allowlisted; eleven unguarded without the allowlist.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
The production branch now has its own runners (#525): operator replicas
that claim rows the Queen queued (queued_at / claimed_by) straight from
the database, with the operator's keys. Volunteer runners stay separate:
their offers never set queued_at, and operator rows never sit on a
volunteer lane.

Conflicts:
- queen-dispatch.ts reapers: a row is the container's to reap only when
  neither kind of runner holds it (queued_at IS NULL AND NOT_A_RUNNER).
  The stall reaper keeps the production branch's silent-runner clause
  for claimed rows and still skips volunteer lanes, which their lease
  reaper releases.
- queen-tick.ts imports, pg-migrate.ts columns: both kept.

dispatchBee still asks a volunteer runner first, then takes the
production path (a local bee, or a queued row for an operator replica).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
@dmitrii-f-t27
dmitrii-f-t27 changed the base branch from claude/gallant-bardeen-h2hgaa to fix/queen-worker-provider-and-prompt-size October 3, 2026 20:46

Copy link
Copy Markdown
Collaborator Author

I've retargeted this PR to fix/queen-worker-provider-and-prompt-size, the production branch, and merged it in (78eb418). This PR still contains #518, so please merge #518 first. Merging this one alone would also bring in #518.

The production branch has gained its own runners (#525). These are the operator's replicas. They claim rows the Queen queued (queued_at / claimed_by) directly from the database and use the operator's keys. The runners in this PR are volunteer machines. They speak only HTTP with a runner token and use their owner's key. The two don't collide:

  • Volunteer offers never set queued_at, so operator replicas never claim them.
  • Operator rows never sit on a volunteer lane above 100000000.
  • dispatchBee asks an idle volunteer runner first. Otherwise it takes the production path: a local bee, or a queued row for an operator replica.

Reapers. The production side had changed the reapers too; I combined both changes. A row is the container's to reap only when neither kind of runner holds it (queued_at IS NULL AND NOT_A_RUNNER). The stall reaper keeps the production branch's silent-runner clause for claimed rows and still skips volunteer lanes; those lanes are released by their own lease reaper.

Local results: test:api 1443 passed, 0 failed; test:pglive 27 passed (including the volunteer-runner live test); test:root 68 passed.

Once this merges, trios/agent-server/tools/queen-runner/queen-runner.mjs exists on the production branch. gHashTag/trinity#1230 will point its download link there.


Generated by Claude Code

@dmitrii-f-t27
dmitrii-f-t27 merged commit 879799a into fix/queen-worker-provider-and-prompt-size Oct 4, 2026
16 of 17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants