Repository navigation
feat(queen): runners take tasks and hand the work back for review - #521
dmitrii-f-t27 merged 14 commits into
Conversation
A signed-in person (app.t27.ai session, verified by asking its issuer's whoami for telegram_id) can mint, list and revoke runner tokens at /queen/me/runners. A runner process on the lender's own machine presents that token at /queen/runner/heartbeat and learns its lane. The provider key never reaches the Queen: no route accepts, stores or returns one. - queen_runner table: only the token's SHA-256 and last four characters are kept; at most 5 live runners per person. - A runner's lane is 100000000 + id, a key_index block no operator pool reaches, so dispatch rows on it are credited by the existing leaderboard arithmetic. - Leaderboard gathers runner lanes by person (telegram_id), never by name, so a runner cannot merge into an operator's row, and never links a runner to an unverified GitHub login. - CORS for /queen/me/* is exactly https://app.t27.ai, bearer only, no credentials; both new mounts are on the route-guard allowlist with an own-bearer reason, and the audit pins are re-measured. Taking tasks and handing work back (claim/complete, remote review, runner CLI) is the next stage; the heartbeat says so instead of offering work. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…dules symlink Every `Tests / *` job on #517 and #518 was cancelled at the 20-minute budget while still inside `bun ci`, before any test ran. Two things combined: - trios/agent-server/apps/server/node_modules was committed as a symlink to a local macOS path. `.gitignore` said `node_modules/`, which only matches directories, so the symlink slipped through. - setup-bun ran without a version. The `packageManager: bun@1.3.6` pin lives in trios/agent-server/package.json, not at the repo root, so CI got the latest release (1.4.2), which hangs on that dangling symlink. 1.3.x installs past it. Untrack the symlink, make the ignore rule match files too, and read the Bun version from the workspace package.json in test.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com With installs no longer hanging, server-tools ran for the first time since 2026-09-23 and failed one test: get_page_content read https://example.com 57 ms after opening it and found no "Example Domain". The test is about extracting text, so it now writes that text into about:blank with evaluate_script, as get_page_links already does. Locally (BrowserOS AppImage, headless, --no-sandbox): the old test fails the same way; the new one passes 3/3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… process
server-tools still exited 1 after every test in observation.test.ts
passed: before navigation-newtab-guard.test.ts the helper ran
`lsof -ti :<cdp port>`, which also lists clients still connected to the
port. One of them was the bun test process itself (its CDP socket to the
previous file's browser), so the SIGTERM ended the whole run and no
junit report was written ("workflow > server-tools setup").
Use `lsof -ti tcp:<port> -sTCP:LISTEN` and drop process.pid.
Locally, input.test.ts + navigation-newtab-guard.test.ts in one process:
before, exit 143 right after "Terminating process(es) <own pid>, ...";
after, 18 pass / 0 fail. The whole test:tools group now runs to the end
(242 pass; the 2 local failures load https://example.com, which this
sandbox's browser cannot reach and CI can).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…example.com With the run no longer killing itself, server-tools finished in CI with 243 pass / 1 fail: `wait_for finds text on page` waited its full 10 s for "Example Domain" on https://example.com and never saw it - the same page get_page_content could not read either. The page now adds that text itself 500 ms after load, so the test still proves wait_for waits, with nothing outside the runner involved. Locally: 2/2 wait_for tests pass on repeat; the whole test:tools group is 243 pass, the one local failure being take_screenshot (a 60 s hang in this sandbox only - it passes in CI). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
… too server-tools on 3d57649 ran clean except one test that had passed on both earlier runs: `search_dom > finds multiple elements with CSS class selector` (123 ms, fewer than 3 matches). It searches once, straight after new_page - the race this file already names and fixes with searchUntil for two sibling tests. Use the same helper here. Locally: search_dom 13/13, three runs in a row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
…udget `the salvage commit > never splits a rename across the path cap` runs real git over 205 files and salvageWorktree. It takes ~2 s for the whole file locally and passed on the two CI runs before, then hit bun's 5 s default once on a loaded runner (job 110500921083) with nothing in the change touching salvage. A git-heavy fixture test should not share the budget of a pure unit test. Locally: queen-salvage-guards.test.ts 13 pass / 0 fail. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Phase 2 of the runner lanes. A runner registered in the cabinet can now be given work, run it on its owner's machine with its owner's key, and hand back a pushed branch that the existing review judges like any other bee. - dispatchBee offers the chosen issue to an idle runner before the container is measured (offerToRunner). The dispatch row goes on the runner's lane with the brief, the system prompt and the start commit stored on it; no worktree, no /chat, no key, no provider/model. - Only issues a runner can continue are offered: no branch yet, a branch with nothing beyond the base, or the runner's own previous push. A branch with container-only commits or a dirty worktree stays local. - The round shows a runner's task as `queued`: a live claim that holds its issue and files but no container worker slot. When queend says the container is full, a runner-only pass re-asks with local bees shown as queued and hands further tasks to idle runners only. - /queen/runner/claim hands over the task; every call renews the lease. /queen/runner/complete fetches the named branch from an https GitHub remote into a scratch ref, checks the commit and that it descends from the start, moves queen-<issue>, writes the agent's answer as the transcript the review reads, then closes the row. - Runners are reaped by lease (unclaimed offer 10 min, silent lease 15 min, cap 6 h), not by the container's boot or stall reapers. - QUEEN_RUNNERS=off keeps the swarm local. RUNNER_PROTOCOL is now 2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
A zero-dependency Node script (tools/queen-runner/queen-runner.mjs) that takes the Queen's tasks on the volunteer's own machine: - heartbeats, claims the task its lane was offered, checks the project out at the commit the task starts from (or its own previous push after a send-back), in a worktree on queen-<issue>; - runs the volunteer's own agent command in the project directory with the prompt on stdin (default: claude -p --permission-mode acceptEdits), commits whatever the agent left, pushes to the volunteer's public fork and completes with the exact commit and the agent's answer; - hands the task back on Ctrl-C, on an agent that fails and changes nothing, or when the swarm cannot take the branch. The provider key stays wherever the volunteer's agent already keeps it. An end-to-end test drives the script against a stand-in swarm with local upstream and fork repositories. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Without one, a send-back the container takes instead of a runner cut a fresh tree with `worktree add -B queen-<issue> ... base`, which reset the branch and started the retry beside the runner's work rather than from it. A runner's finished issue now looks like a container bee's: branch and clean tree, which prepareWorktree reuses. Not fatal if it fails; the review reads the branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
✅ Tests passed — 2596/2656
|
|
The Generated by Claude Code |
Conflicts, both additive: - server.ts: keep the runner mounts and add /queen/contributor-keys. - queen-leaderboard.ts: rank() takes the operator map merged with contributorOwnerNames(), and the runner owners beside it. Also ports the route-guard fix from #520: #522 left /queen/contributor-keys out of the audit, which turns four route-guard tests red on the base. It is allowlisted with its own capability guard and the pins are re-measured (48 mounts; 25 /queen: 8/8/9). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
One conflict, in dispatchBee: the runner offer still comes first, and the container path then picks its key with #522's contributor runtime (resolveWorkerProvider(..., runtime)). Contributor keys use negative indices and runner lanes sit above 100000000, so the runner filters (NOT_A_RUNNER, isRunnerLane) leave contributor keys with the container. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Retargets the runner cabinet at the branch production builds from. Conflicts, all additive: - server.ts: runner mounts beside /queen/public-credits and /queen/public-earnings. - pg-migrate.ts: queen_runner beside queen_tri_earnings. - route-guard-audit.mjs: the runner allowlist entries beside the production branch's /queen/contributor-keys entry (same reason text). - route-guard.test.ts: the production branch's pins plus the two runner mounts: 51 mounts; 28 /queen mounts as 10 public-read, 9 wrapper, 9 allowlisted; eleven unguarded without the allowlist. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
The production branch now has its own runners (#525): operator replicas that claim rows the Queen queued (queued_at / claimed_by) straight from the database, with the operator's keys. Volunteer runners stay separate: their offers never set queued_at, and operator rows never sit on a volunteer lane. Conflicts: - queen-dispatch.ts reapers: a row is the container's to reap only when neither kind of runner holds it (queued_at IS NULL AND NOT_A_RUNNER). The stall reaper keeps the production branch's silent-runner clause for claimed rows and still skips volunteer lanes, which their lease reaper releases. - queen-tick.ts imports, pg-migrate.ts columns: both kept. dispatchBee still asks a volunteer runner first, then takes the production path (a local bee, or a queued row for an operator replica). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
|
I've retargeted this PR to The production branch has gained its own runners (#525). These are the operator's replicas. They claim rows the Queen queued (
Reapers. The production side had changed the reapers too; I combined both changes. A row is the container's to reap only when neither kind of runner holds it ( Local results: Once this merges, Generated by Claude Code |
879799a
into
fix/queen-worker-provider-and-prompt-size
Phase 2 of runner lanes. It is stacked on #518, which registers runners, so review and merge that one first. GitHub will retarget this PR when #518 merges.
A runner created in the app.t27.ai cabinet can now be given a task. It runs the task on its owner's machine with its owner's key and hands back a pushed branch. The existing review judges that branch by the same rules as any container bee's work. The provider key never reaches this server.
How a task flows
Offer.
dispatchBeeasksofferToRunnerfirst, before any key, memory or disk is measured.100000000 + id). The row isstarted = trueand stores the brief, the system prompt and the start commit./chat, and no provider or model is recorded. Because the row has no provider or model, the daily spend cap never prices the lender's spend.Which issues are offered. Only issues a runner can actually continue:
A branch with commits that exist only in the container, or a dirty worktree, stays with the container.
Claim.
POST /queen/runner/claimhands over the task. Every runner call renews the lease. Claiming again returns the same task, so a runner that restarts picks up where it left off.Complete.
POST /queen/runner/completewith{conversationId, remoteUrl, branch, headSha, said}. The server:https://github.com/..., withprotocol.allow=neverand no credential helper;queen-<issue>and gives it a clean worktree, as a container bee would leave;saytranscript the review reads;A runner can also send
{gaveUp: true}to release the issue immediately.Reap. Runner rows are skipped by the boot reaper and the stall reaper. Instead they are released by lease:
Effect on the round
queued. queend already treats that state as a live claim that holds its issue and files, but it does not count it againstcanStartAnother. So runners add capacity rather than consuming container slots.queued. Each answer goes to an idle runner only. Boundaries, claims and the spend cap still apply.takenKeysandkeyCursor;runningKeys, used for the reviewer lane;QUEEN_RUNNERS=offturns offers off, and the swarm runs as it did before.RUNNER_PROTOCOLis now 2.The runner script
trios/agent-server/tools/queen-runner/queen-runner.mjsplus a README. It needs Node 20+ and git, with no packages to install. It:queen-<issue>;claude -p --permission-mode acceptEdits;Tests
bun run typecheckbun run test:apibun run test:pglive(local PostgreSQL 16)bun run test:rootbiome checkon changed filesrunRoundandreviewFinishedDispatchesalready carryNew tests:
tests/api/queen-runner-work.test.ts. Uses real git in a temp directory. It covers:file://, ssh,ext::, other hosts, userinfo,--upload-pack,..and abbreviated SHAs are all refused;tests/api/queen-round-runners.test.ts. DrivesrunRoundagainst a stand-in policy that applies queend's capacity and claim rules. The realqueendis not built in CI, soqueen-round.test.tsskips there. Mutation-checked: removing thequeuedmapping fails one test, and treating runner dispatches as running fails another.tests/pglive/queen-runner-work-live.test.ts. The whole life of a task on a migrated scratch database: register, idle, offer, claim, renew, complete, then a review-ready row with the verdict in its transcript. It also covers reaping.tests/api/queen-runner-cli.test.ts. Runs the script end to end against a stand-in swarm with local upstream and fork repos, under Bun. It was also run locally under Node 22.Two existing tests were adjusted:
queen-runners.test.tsnow expects protocol 2.queen-dispatch.test.tspassesoffer: null. That test asserts that noqueen_dispatchstatement is sent, and the runner lookup is a read of that table.Security notes
TRIOS_REPO_URL.Not done
claude -prun. The CLI test uses a stand-in agent.claude/runner-claim-complete, stacked on feat(queen): MY RUNNERS cabinet on the leaderboard tab trinity#1205.🤖 Generated with Claude Code
https://claude.ai/code/session_01SJ8KjRoGNBoHBoDR92fAo2
Generated by Claude Code