Repository navigation
CI: raise timeout-minutes on the three pytest jobs from 30 to 45 #713
Description
Activity
- addedbugIssues reporting a defect (set by the bug report template; a default, not an assessment)Issues reporting a defect (set by the bug report template; a default, not an assessment)ciPull requests that change CI or workflow configuration (ci: subject prefix)Pull requests that change CI or workflow configuration (ci: subject prefix)
on Sep 23, 2026 Correction: this is not confined to the integration job — the unit job is hitting the same wall
The issue as filed named only
integration(.github/workflows/unit-tests.yml:96). A third timeout has now
landed on a unit job, so the scope is wider than stated and the title has been updated.PR #711, job
107203030981,Python 3.15: started14:39:37Z, killed15:09:53Z— 30 min 16 s —
inside step 5Run unit tests, with steps 1-4 green. That istimeout-minutes: 30at
.github/workflows/unit-tests.yml:54, theunitjob's own ceiling, not the integration one.Durations of unit jobs that succeeded, same sampling method as the table above:
duration job 27 min Python 3.1024 min Python 3.1223 min Python 3.1423 min Python 3.1323 min Python 3.1223 min Python 3.1123 min Python 3.1022 min Python 3.15Worst passing case 27 minutes against a 30-minute ceiling — three minutes of margin. Tighter than it looks,
because the spread between 22 and 27 minutes is pure runner variance on identical work.Running tally of killed jobs
PR job id job ceiling started killed duration required context? #691 107202753543Integration Python 3.12:9614:12:13Z14:42:32Z30 min 19 s yes #695 107202791532Integration Python 3.12:9614:31:48Z15:02:05Z30 min 17 s yes #711 107203030981Python 3.15:5414:39:37Z15:09:53Z30 min 16 s no — 3.15 runs but cannot block Three in about an hour, across two different jobs and two different Python versions. #711 is not blocked by its
one, sincePython 3.15is not among ruleset23497679's fifteen required contexts, but #691 and #695 both are
blocked.What this changes about the fix
Raising only the
integrationceiling would leave the unit job one slow runner away from the same outcome. Both
ceilings need the same treatment —unit-tests.yml:54andunit-tests.yml:96, currently30each. The
suggestion stands at 45 for both, with the same caveat that the number is a CI-budget judgement rather than a
measurement.unit-tests.yml:207is also30andlint.yml:72is30; neither has produced a timeout yet, so they are
listed for awareness rather than as part of the fix.unit-tests.yml:179is5and is a different kind of step.The deeper point is unchanged and worth weighing against simply raising the numbers: both jobs run the whole
suite serially per Python version, which is why they sit within a few minutes of half an hour. Sharding or
pytest-xdistwould move the distribution rather than move the ceiling.- changed the title
[-]CI: integration jobs are killed by timeout-minutes: 30 when the suite's worst case is 29 min[/-][+]CI: unit and integration jobs are both killed by timeout-minutes: 30, with only 3 min of margin[/+]on Sep 23, 2026 Update: re-running two of the killed jobs under a drained queue — both passed
The issue as written reads as though the suite cannot complete inside 30 minutes. It can. Two of the
timed-out jobs were re-run when the Actions queue had fully drained, and both finished:PR job started outcome #691 Integration Python 3.1216:11:46ZSUCCESS #695 Integration Python 3.1216:11:16ZSUCCESS Conditions at the time: 0 runs queued, 7 in progress, the emptiest the repository had been all day. Those
same two jobs had each been killed at 30 minutes earlier, when 20+ runs were competing for runners.What that changes about the diagnosis
The six timeouts were contention pushing 26-29 minute jobs past a 30-minute wall, not a suite that
overruns the ceiling on its own. The duration table above still stands — the worst passing integration job is
29 minutes and the worst passing unit job is 27 — so the margin is genuinely 1-3 minutes. But the failure mode
is now precise:- Under a quiet queue the suite fits, with a few minutes to spare.
- Under a busy queue the same job takes several minutes longer and dies.
So raising
timeout-minutesis insurance against contention rather than a correction of an impossible
limit. That is a weaker claim than the issue originally made, and the recommendation is unchanged —45on both
.github/workflows/unit-tests.yml:54and:96— but the justification should be stated honestly: it stops CI
from being a lottery on queue depth, not from being arithmetically impossible.A second, cheaper mitigation discovered by accident
A rebase clears a timeout casualty for free. When
mainmoved and all open PRs were rebased onto it, every
deadcancelledcheck vanished, because the new head gets a fresh run of every job. #703'sIntegration 3.13,
#711'sPython 3.15and #712'sIntegration 3.10were all resolved this way without a single re-run.Which means the manual
gh run rerun --job <id>path only matters for a PR that is terminal and already
up-to-date withmain— nothing queued, nothing running, one required check dead. Any PR that will need a
rebase anyway gets its casualties cleared as a side effect.What is still true and still worth fixing
Five of the six timeouts hit required contexts (
Integration Python 3.10-3.14are five of the fifteen in
ruleset23497679), and a timed-out job is reported ascancelledrather thanfailure, so a
conclusion-based failure tally reads zero whilestatusCheckRollup.statereadsFAILURE. Anything automating
over these checks still needs to treatcancelledas "needs attention" rather than folding it in with runs
superseded by a force-push.- changed the title
[-]CI: unit and integration jobs are both killed by timeout-minutes: 30, with only 3 min of margin[/-][+]CI: raise timeout-minutes on the three pytest jobs from 30 to 45[/+]on Sep 23, 2026 Rescoped: this issue is now the ceiling raise alone
The structural work — cutting job count and duplicated execution — has moved to #715, along with the
measurements behind it. This issue keeps the one change that can land immediately and close on its own.What to change
.github/workflows/unit-tests.yml,timeout-minutes: 30→45on the three jobs that run pytest:line job name as a check context measured worst passing run :54testPython 3.1x27.6 min :96integrationIntegration Python 3.1x29.4 min :207gateGate (full suite, Python 3.14)~27.2 min Leave
unit-tests.yml:179(changelog, 5 min) andlint.yml:72(30 min) alone —lintcompletes in 4.0
min andchangelogin 6-10 s, so both have ample margin and neither has ever been killed.Why raise it at all, given #715 is the real fix
Because the margin is 36 seconds. The worst passing
integrationjob measured 29.4 min against a 30-minute
wall, and six jobs have already been killed today — five of them on required contexts
(Integration Python 3.10-3.14are five of the fifteen in ruleset23497679), each needing a manual re-run
before its PR can merge.Raising the ceiling does not make CI faster and is not a substitute for #715. It stops required checks dying
while the structural work lands. Two re-runs under a drained queue both passed, which confirms the suite fits in
30 minutes when runners are free — the kills happen when they are not.The counter-argument, stated honestly
A longer ceiling lets a contended job hold a runner slot longer, which can deepen the backlog for everything
behind it. That is a real cost and the reason this should be a temporary measure: once #715's items land, the
job count falls, contention falls, and this can be reverted to 30 — at which point the margin is no longer
36 seconds because the jobs themselves are shorter.45is chosen as ~1.55× the observed maximum: enough headroom that contention alone cannot kill a job, tight
enough that a genuine runaway still gets bounded rather than burning a full hour.Note for whoever automates over these checks
A
timeout-minuteskill is reported asconclusion: "cancelled", not"timed_out"— a scan of the 90 most
recent completed runs found zero jobs withtimed_out. That makes a timeout indistinguishable from a
concurrency-group cancellation by conclusion alone, andcancel-in-progressis enabled on all eight workflows
(#394), so both happen routinely.The tell is duration. In run
35873084988,Integration Python 3.10ran 30.3 min and was killed by the
ceiling, while three sibling jobs in the same run were cancelled after 17-21 min in lockstep — the latter
being the concurrency group correctly superseding a stale run. Filter on duration ≈ the ceiling to isolate real
timeouts; a rawcancelledcount conflates a bug with a feature.- added a commit that references this issue
on Sep 23, 2026
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
.github/workflows/unit-tests.yml:96setstimeout-minutes: 30on theintegrationjob. Measured, that isbelow the job's worst-case runtime, so integration checks are being killed mid-suite on PRs that have nothing
wrong with them.
The measurement
Durations of integration jobs that succeeded, sampled across branches and Python versions from recent
Unit Testsruns:Integration Python 3.10Integration Python 3.15Integration Python 3.12Integration Python 3.14Integration Python 3.14Integration Python 3.10Integration Python 3.15Integration Python 3.14Integration Python 3.13The slowest passing run is 29 minutes against a 30-minute ceiling — one minute of margin.
Two jobs already killed
10720275354314:12:13Z14:42:32ZRun full test suite, steps 1-5 all green10720279153214:31:48Z15:02:05ZRun full test suite, steps 1-5 all greenBoth are
Integration Python 3.12, but that is coincidence rather than a property of 3.12 — the table aboveshows 3.10, 3.14 and 3.15 all landing at 26-29 minutes. Whichever job draws a slow runner tips over.
Why this blocks merges rather than just being noisy
Integration Python 3.10through3.14are five of the fifteen required status checks in ruleset23497679.A timed-out job is not re-run automatically, so any PR that loses one cannot merge until someone re-runs it
by hand — and re-running into a busy queue can hit the same wall.
A timeout is reported as
cancelled, notfailure. That makes it easy to miss: a conclusion-based failuretally returns zero while
statusCheckRollup.statereadsFAILURE, and the usual explanation for acancelledjob — a force-push superseding the run — does not apply. Anything automating over these checks needs to treat
cancelledas "needs attention" rather than folding it in with superseded runs.Suggested fix
Raise
timeout-minuteson theintegrationjob from30to 45, roughly 1.55x the observed maximum. Thatkeeps a real runaway bounded while leaving headroom for a slow runner.
45is a suggestion, not a measurement —the right number is a judgement about CI budget.
Worth considering alongside it, though out of scope here: the integration job runs the full suite per Python
version with no parallelism inside the job, which is why it sits near half an hour. Sharding it or using
pytest-xdistwould cut wall-clock more durably than raising the ceiling, at the cost of a more complexworkflow.
Related sites with the same setting, unchanged by this and listed for completeness:
unit-tests.yml:54,unit-tests.yml:207,lint.yml:72(all30), andunit-tests.yml:179(5).