Skip to content

CI: raise timeout-minutes on the three pytest jobs from 30 to 45 #713

Description

@JarryShaw

.github/workflows/unit-tests.yml:96 sets timeout-minutes: 30 on the integration job. Measured, that is
below the job's worst-case runtime
, so integration checks are being killed mid-suite on PRs that have nothing
wrong with them.

The measurement

Durations of integration jobs that succeeded, sampled across branches and Python versions from recent
Unit Tests runs:

duration job
29 min Integration Python 3.10
28 min Integration Python 3.15
28 min Integration Python 3.12
27 min Integration Python 3.14
26 min Integration Python 3.14
26 min Integration Python 3.10
21 min Integration Python 3.15
21 min Integration Python 3.14
20 min Integration Python 3.13

The slowest passing run is 29 minutes against a 30-minute ceiling — one minute of margin.

Two jobs already killed

PR job id started killed duration step reached
#691 107202753543 14:12:13Z 14:42:32Z 30 min 19 s Run full test suite, steps 1-5 all green
#695 107202791532 14:31:48Z 15:02:05Z 30 min 17 s Run full test suite, steps 1-5 all green

Both are Integration Python 3.12, but that is coincidence rather than a property of 3.12 — the table above
shows 3.10, 3.14 and 3.15 all landing at 26-29 minutes. Whichever job draws a slow runner tips over.

Why this blocks merges rather than just being noisy

Integration Python 3.10 through 3.14 are five of the fifteen required status checks in ruleset 23497679.
A timed-out job is not re-run automatically, so any PR that loses one cannot merge until someone re-runs it
by hand — and re-running into a busy queue can hit the same wall.

A timeout is reported as cancelled, not failure. That makes it easy to miss: a conclusion-based failure
tally returns zero while statusCheckRollup.state reads FAILURE, and the usual explanation for a cancelled
job — a force-push superseding the run — does not apply. Anything automating over these checks needs to treat
cancelled as "needs attention" rather than folding it in with superseded runs.

Suggested fix

Raise timeout-minutes on the integration job from 30 to 45, roughly 1.55x the observed maximum. That
keeps a real runaway bounded while leaving headroom for a slow runner. 45 is a suggestion, not a measurement —
the right number is a judgement about CI budget.

Worth considering alongside it, though out of scope here: the integration job runs the full suite per Python
version with no parallelism inside the job, which is why it sits near half an hour. Sharding it or using
pytest-xdist would cut wall-clock more durably than raising the ceiling, at the cost of a more complex
workflow.

Related sites with the same setting, unchanged by this and listed for completeness: unit-tests.yml:54,
unit-tests.yml:207, lint.yml:72 (all 30), and unit-tests.yml:179 (5).

Activity

  1. added
    bugIssues reporting a defect (set by the bug report template; a default, not an assessment)
    ciPull requests that change CI or workflow configuration (ci: subject prefix)
    on Sep 23, 2026
  2. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Correction: this is not confined to the integration job — the unit job is hitting the same wall

    The issue as filed named only integration (.github/workflows/unit-tests.yml:96). A third timeout has now
    landed on a unit job, so the scope is wider than stated and the title has been updated.

    PR #711, job 107203030981, Python 3.15: started 14:39:37Z, killed 15:09:53Z — 30 min 16 s —
    inside step 5 Run unit tests, with steps 1-4 green. That is timeout-minutes: 30 at
    .github/workflows/unit-tests.yml:54, the unit job's own ceiling, not the integration one.

    Durations of unit jobs that succeeded, same sampling method as the table above:

    duration job
    27 min Python 3.10
    24 min Python 3.12
    23 min Python 3.14
    23 min Python 3.13
    23 min Python 3.12
    23 min Python 3.11
    23 min Python 3.10
    22 min Python 3.15

    Worst passing case 27 minutes against a 30-minute ceiling — three minutes of margin. Tighter than it looks,
    because the spread between 22 and 27 minutes is pure runner variance on identical work.

    Running tally of killed jobs

    PR job id job ceiling started killed duration required context?
    #691 107202753543 Integration Python 3.12 :96 14:12:13Z 14:42:32Z 30 min 19 s yes
    #695 107202791532 Integration Python 3.12 :96 14:31:48Z 15:02:05Z 30 min 17 s yes
    #711 107203030981 Python 3.15 :54 14:39:37Z 15:09:53Z 30 min 16 s no — 3.15 runs but cannot block

    Three in about an hour, across two different jobs and two different Python versions. #711 is not blocked by its
    one, since Python 3.15 is not among ruleset 23497679's fifteen required contexts, but #691 and #695 both are
    blocked.

    What this changes about the fix

    Raising only the integration ceiling would leave the unit job one slow runner away from the same outcome. Both
    ceilings need the same treatment — unit-tests.yml:54 and unit-tests.yml:96, currently 30 each. The
    suggestion stands at 45 for both, with the same caveat that the number is a CI-budget judgement rather than a
    measurement.

    unit-tests.yml:207 is also 30 and lint.yml:72 is 30; neither has produced a timeout yet, so they are
    listed for awareness rather than as part of the fix. unit-tests.yml:179 is 5 and is a different kind of step.

    The deeper point is unchanged and worth weighing against simply raising the numbers: both jobs run the whole
    suite serially per Python version, which is why they sit within a few minutes of half an hour. Sharding or
    pytest-xdist would move the distribution rather than move the ceiling.

  3. changed the title [-]CI: integration jobs are killed by timeout-minutes: 30 when the suite's worst case is 29 min[/-] [+]CI: unit and integration jobs are both killed by timeout-minutes: 30, with only 3 min of margin[/+] on Sep 23, 2026
  4. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Update: re-running two of the killed jobs under a drained queue — both passed

    The issue as written reads as though the suite cannot complete inside 30 minutes. It can. Two of the
    timed-out jobs were re-run when the Actions queue had fully drained, and both finished:

    PR job started outcome
    #691 Integration Python 3.12 16:11:46Z SUCCESS
    #695 Integration Python 3.12 16:11:16Z SUCCESS

    Conditions at the time: 0 runs queued, 7 in progress, the emptiest the repository had been all day. Those
    same two jobs had each been killed at 30 minutes earlier, when 20+ runs were competing for runners.

    What that changes about the diagnosis

    The six timeouts were contention pushing 26-29 minute jobs past a 30-minute wall, not a suite that
    overruns the ceiling on its own. The duration table above still stands — the worst passing integration job is
    29 minutes and the worst passing unit job is 27 — so the margin is genuinely 1-3 minutes. But the failure mode
    is now precise:

    • Under a quiet queue the suite fits, with a few minutes to spare.
    • Under a busy queue the same job takes several minutes longer and dies.

    So raising timeout-minutes is insurance against contention rather than a correction of an impossible
    limit. That is a weaker claim than the issue originally made, and the recommendation is unchanged — 45 on both
    .github/workflows/unit-tests.yml:54 and :96 — but the justification should be stated honestly: it stops CI
    from being a lottery on queue depth, not from being arithmetically impossible.

    A second, cheaper mitigation discovered by accident

    A rebase clears a timeout casualty for free. When main moved and all open PRs were rebased onto it, every
    dead cancelled check vanished, because the new head gets a fresh run of every job. #703's Integration 3.13,
    #711's Python 3.15 and #712's Integration 3.10 were all resolved this way without a single re-run.

    Which means the manual gh run rerun --job <id> path only matters for a PR that is terminal and already
    up-to-date with main
    — nothing queued, nothing running, one required check dead. Any PR that will need a
    rebase anyway gets its casualties cleared as a side effect.

    What is still true and still worth fixing

    Five of the six timeouts hit required contexts (Integration Python 3.10-3.14 are five of the fifteen in
    ruleset 23497679), and a timed-out job is reported as cancelled rather than failure, so a
    conclusion-based failure tally reads zero while statusCheckRollup.state reads FAILURE. Anything automating
    over these checks still needs to treat cancelled as "needs attention" rather than folding it in with runs
    superseded by a force-push.

  5. changed the title [-]CI: unit and integration jobs are both killed by timeout-minutes: 30, with only 3 min of margin[/-] [+]CI: raise timeout-minutes on the three pytest jobs from 30 to 45[/+] on Sep 23, 2026
  6. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Rescoped: this issue is now the ceiling raise alone

    The structural work — cutting job count and duplicated execution — has moved to #715, along with the
    measurements behind it. This issue keeps the one change that can land immediately and close on its own.

    What to change

    .github/workflows/unit-tests.yml, timeout-minutes: 30 → 45 on the three jobs that run pytest:

    line job name as a check context measured worst passing run
    :54 test Python 3.1x 27.6 min
    :96 integration Integration Python 3.1x 29.4 min
    :207 gate Gate (full suite, Python 3.14) ~27.2 min

    Leave unit-tests.yml:179 (changelog, 5 min) and lint.yml:72 (30 min) alone — lint completes in 4.0
    min and changelog in 6-10 s, so both have ample margin and neither has ever been killed.

    Why raise it at all, given #715 is the real fix

    Because the margin is 36 seconds. The worst passing integration job measured 29.4 min against a 30-minute
    wall, and six jobs have already been killed today — five of them on required contexts
    (Integration Python 3.10-3.14 are five of the fifteen in ruleset 23497679), each needing a manual re-run
    before its PR can merge.

    Raising the ceiling does not make CI faster and is not a substitute for #715. It stops required checks dying
    while the structural work lands. Two re-runs under a drained queue both passed, which confirms the suite fits in
    30 minutes when runners are free — the kills happen when they are not.

    The counter-argument, stated honestly

    A longer ceiling lets a contended job hold a runner slot longer, which can deepen the backlog for everything
    behind it. That is a real cost and the reason this should be a temporary measure: once #715's items land, the
    job count falls, contention falls, and this can be reverted to 30
    — at which point the margin is no longer
    36 seconds because the jobs themselves are shorter.

    45 is chosen as ~1.55× the observed maximum: enough headroom that contention alone cannot kill a job, tight
    enough that a genuine runaway still gets bounded rather than burning a full hour.

    Note for whoever automates over these checks

    A timeout-minutes kill is reported as conclusion: "cancelled", not "timed_out" — a scan of the 90 most
    recent completed runs found zero jobs with timed_out. That makes a timeout indistinguishable from a
    concurrency-group cancellation by conclusion alone, and cancel-in-progress is enabled on all eight workflows
    (#394), so both happen routinely.

    The tell is duration. In run 35873084988, Integration Python 3.10 ran 30.3 min and was killed by the
    ceiling, while three sibling jobs in the same run were cancelled after 17-21 min in lockstep — the latter
    being the concurrency group correctly superseding a stale run. Filter on duration ≈ the ceiling to isolate real
    timeouts; a raw cancelled count conflates a bug with a feature.

  7. added a commit that references this issue on Sep 23, 2026
    b60eaea
  8. added this to the 1.5 milestone on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugIssues reporting a defect (set by the bug report template; a default, not an assessment)ciPull requests that change CI or workflow configuration (ci: subject prefix)

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions