Skip to content

Pin the model on every claude launch, not just the first - #351

Draft
dkrattiger wants to merge 1 commit into
mainfrom
panopticon/deterministic-task-model
Draft

dkrattiger wants to merge 1 commit into
mainfrom
panopticon/deterministic-task-model

Conversation

@dkrattiger

Copy link
Copy Markdown
Contributor

Summary

container/agent.py::_claude_argv only passed --model on a fresh session. Since the per-task
config volume persists across container respawn/recreate (ADR 0011 §5), a resumed session
(--continue) is the common case after the first launch — and resumes got no --model flag at
all, falling through to the claude CLI's own account default. A task's recorded
starting_model (e.g. opus) would silently diverge from what actually ran (observed running
Fable/Sonnet instead).

  • _claude_argv now passes --model <starting_model> on every launch — first run, --continue
    resumes, and any later fresh session — whenever the task has one set. starting_model=None
    remains the explicit escape hatch (claude's own default), on every launch.
  • Added PUT /tasks/{id}/starting-model (+ TaskService.set_starting_model) so a task's model
    can be changed mid-task, following the existing single-field-setter convention
    (slug/url/token-estimate/…). There was no PATCH for tasks to reuse (only for repos).
  • Fixed _update_task in the SQLAlchemy store, which never wrote starting_model back on save
    (it was write-once from creation) — without this fix the new endpoint would silently no-op.
  • Updated the stale "first run/launch only" docstrings in core/models.py, core/workflow.py,
    sessionservice/local_runner.py, sessionservice/spawner.py, and docs/tasks.md.

See the plan artifact (plan.md) on this task for the full analysis.

_claude_argv only passed --model on a fresh session; every --continue
resume (which is most relaunches, since the per-task config volume
persists across container respawn) fell through to the CLI's own
account default, so a task's recorded starting_model quietly diverged
from what actually ran.

Also add PUT /tasks/{id}/starting-model so the model can be changed
mid-task, and fix _update_task in the SQLAlchemy store, which never
wrote starting_model back on save — without it the new endpoint would
silently no-op.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
dkrattiger added a commit that referenced this pull request Sep 25, 2026
`default_model` was the bare alias "opus", which resolves to the *latest* Opus
at launch time. That silently re-points every new task's agent whenever a new
Opus ships — the same class of hazard as the unpinned claude install in the base
image, and it's how tasks ended up on Opus 5 without anyone choosing it.

Pin an explicit id so a bump is a deliberate, reviewable edit. 4.8 is also the
more conservative model on scope: Opus 5 expands task scope and writes longer
deliverables, which lands directly on whoever reviews the PR.

Only affects a task's first launch — on resume claude keeps whatever model the
conversation is already on, so existing tasks stay where they are. Moving one of
those means `/model claude-opus-4-8` in its pane, until #351 lands.

The test now asserts the value is a pinned id and explicitly not an alias, so
re-introducing one fails loudly rather than drifting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dkrattiger added a commit that referenced this pull request Sep 30, 2026
`reviewDecision == "REVIEW_REQUIRED"` is not evidence of an external wait. On a
branch with a protection rule it means only that *a* review is required and none
has been given — which, with nobody requested, makes that review ours.

Every open task PR here reads exactly that way:

    #351    REVIEW_REQUIRED  author=dkrattiger  reviewRequests=[]
    #10990  REVIEW_REQUIRED  author=dkrattiger  reviewRequests=[]
    #10983  REVIEW_REQUIRED  author=dkrattiger  reviewRequests=[]
    #10980  REVIEW_REQUIRED  author=dkrattiger  reviewRequests=[]

So the first cut would have labelled all four "ext review" and dimmed them —
hiding precisely the work most in need of attention, which is worse than no
marker at all. Verified against the live PRs: all four now derive to None.

Ask the narrower question instead: is a review pending from someone *other than
us*? Reads `reviewRequests` and compares against the authenticated login
(`gh api user`, resolved once and cached for the daemon's life). A requested team
counts as external — we are never a team.

Where it can't tell, it errs toward "ours". An unresolvable identity, or a request
naming only us, both fall through to no marker. The asymmetry is deliberate: a
wrong "external" hides work, while a missing one costs a glance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant