Repository navigation
CI: cut job count and duplicated work — contention, not job duration, is what makes CI slow #715
Description
Activity
- addedenhancementIssues requesting a new capability (set by the feature request template)Issues requesting a new capability (set by the feature request template)perfPull requests that improve performance (perf: subject prefix)Pull requests that improve performance (perf: subject prefix)ciPull requests that change CI or workflow configuration (ci: subject prefix)Pull requests that change CI or workflow configuration (ci: subject prefix)
on Sep 23, 2026 Measurement landed — and it moves the biggest lever off the CI config entirely
A per-directory profile of the suite is in. The dominant cost is not the workflow structure. It is that
pcapkitis cold-reimported once per test method.Where the 24.5 minutes actually goes
Measured single-process, one
pytest --durations -qper directory, on an otherwise-quiet 16-core box. The
passed+skipped total (1739) matchespytest --collect-only -qexactly, so nothing was double-counted:directory wall tests subtests % tests/protocols/742.70s 729 1829 50.5% tests/foundation/291.38s 250/11 skip 394 19.8% tests/corekit/145.22s 179 413 9.9% tests/vendor/64.33s 55 54 4.4% tests/toolkit/63.94s 49/4 skip 51 4.3% tests/integration/55.76s 90/2 skip 115 3.8% tests/utilities/26.51s 104 100 1.8% tests/interface/19.81s 22 9 1.3% tests/const/18.06s 40 798 1.2% tests/dumpkit/17.93s 15 68 1.2% tests/root13.47s 57 35 0.9% tests/project/9.65s 126 479 0.7% tests/cli/1.05s 6 0 0.07% total 1469.8s ≈ 24.5 min 1739 4345 protocols+foundation+corekit= 1179.3s, 80.2% of the suite.The new top lever:
purge_modulesinsetUprather thansetUpClasstests/project/averages 0.077s per test.tests/corekit/averages 0.811s per test. The difference is that
corekitpurgespcapkitfromsys.modulesinsetUp, so every test method pays a cold reimport.Verified independently by grepping the tree: 144
purge_modulescall sites sit insetUp, against only
10 insetUpClass/setUpModule, spread across 108 files. A coldimport pcapkitmeasures 654ms, of
whichpcapkit/const/reg/apptype.pyalone is 160ms (24.4%) — aStrEnumwith ~8,949 members modelling the
IANA port registry, built eagerly at class-body evaluation.Moving those purges to class scope would pay the tax once per class instead of once per method: roughly 10-13
minutes off a 24.5-minute suite (40-55%).This is the highest-value change available and also the riskiest.
tests/conftest.pydocuments that the
per-test isolation exists because of two real regressions (#660, #688) where a leakedsys.modulesbinding
broke unrelated tests, and — line 128 — that the suite currently passes partly because unrelated directories
"sort between the polluter and its victims". That is an accidental order-dependency. Any change here needs to
prove no sibling test in a class depends on a purge another sibling would otherwise leave behind.Item 1 is confirmed and now quantified — it is bigger than first stated
The delta between the full suite and the unit job's subset is 164.9s ≈ 2.75 min, 11.2% of total:
tests/integration/55.76s, plus the*_runtime.py/*_regression.pyfiles at 108.30s inprotocols(led by
test_pcapng_regression.py18.77s,test_registry_runtime.py16.95s,test_http_runtime.py13.01s) and 0.81s
infoundation.So partitioning
unit-tests.yml:156takes eachIntegration Python 3.1xjob from ~24.5 min to ~2.75 min —
~21.75 min saved per job, ~130 min per PR push across the six. That is the largest CI-config win by a wide
margin.The risk is the inverse of what this issue first stated. It is not that unit-scoped tests depend on fixtures
and would break — thetestjob generates no fixtures at all (verified:unit-tests.ymllines 43-90 contain
only an install step and the pytest step, nomake_samples.py), so they already pass without them. The real
exposure is that 39 files callsample_path()and 4 use anexcept FileNotFoundErrorskip idiom, so some
tests may be skipping in the unit job and running for real only in the integration job. Partitioning would mean
those never run for real anywhere. Before landing item 1, diff the skip counts between the unit invocation and
the full invocation; any test that runs in one and skips in the other must stay on the integration side.Two hypotheses refuted, worth recording so they are not re-chased
Collection overhead is negligible —
pytest --collect-only -qover all 1739 tests takes 2.75s.Capture I/O is not the driver. The largest capture,
http.pcap(171 KB, 1117 frames), parses once in
1.22s CPU and is referenced in only 8 files. Aggregate parse/IO cost is on the order of 1-2% of the suite.
make_samples.pyruns in 2.84s.Subtests are numerically concentrated but not a wall-clock driver. Per-test-method cost is uniform regardless
of subtest density —test_ipv6_extension_unit.pyis 1.18s/phase andtest_mh_unit.py1.34s/phase, the same as
everywhere else, because each method already pays the fixed reimport tax and the subtest iterations inside are
cheap. Sampling or caching subtests would save little.A correction to #575's framing
#575 attributes 16.7% of extraction self-time to
aenum.extend_enum. At the function level that is refuted
for the test suite:extend_enumis 245 calls, 0.010s self — under 0.1%. The cost is theEnumMeta
class-construction machinery it depends on: allaenumfunctions together are 6.615s self (23.8%) and 22.339s
cumulative (80.5%) of a profiledtests/const/run, withbuiltins.__build_class__at 21.280s cumulative
(76.7%) and 1.53Mabc.__subclasscheck__calls from aenum's descriptor emulation. So the finding stands in
substance and is wrong in attribution — optimisingextend_enumitself would achieve nothing; shrinking eager
enum construction would.The top-20 files, for whoever picks up the work
Within
protocols(exhaustive--durations=0):test_ipv6_extension_unit.py72.22s,test_mh_unit.py67.21s,
test_pcapng_unit.py65.49s,test_hip_unit.py38.25s,test_link_unit.py35.97s,
test_protocol_code_registration_unit.py34.98s,test_http_unit.py34.66s,test_esp_unit.py34.60s,
test_tcp_udp_unit.py33.15s,test_ipv4_unit.py29.93s,test_protocol_base_unit.py25.70s,
test_schema_unit.py19.35s,test_pcapng_regression.py18.77s,test_registry_runtime.py16.95s,
test_transport_unit.py16.36s,test_dispatch_bindings_unit.py15.48s,test_ipv6_unit.py13.49s,
test_http_runtime.py13.01s,test_tcp_runtime.py10.13s,test_dispatch_registry_unit.py9.59s.Those 20 files are 605.29s = 41.2% of the entire suite, and all of them call
purge_modules(['pcapkit'])in
setUp.Revised ordering
purge_modules→ class scope (new). 10-13 min, ~40-55%. Highest value, highest risk, needs the
order-dependency work above. Source change, not CI.- Partition the integration job (item 1). ~130 min/PR push. Needs the skip-count diff first.
- Stop the gate re-running on main pushes (item 2). ~108 min/main push, no signal lost, no measurement
needed — can land now. - Trim 3.15 from the PR matrix (item 3). 3 jobs/PR push, safest.
pytest-xdist(item 4). Still assessing — and note the profiler independently flagged the same
order-dependency atconftest.py:128as the reason parallelism may resurface the twelve tests/project/ tests fail withTypeError: type 'ProtocolBase' is not subscriptabledepending on what ran first: the fake-module helpers purge on entry and never restore on exit #660/tests/cli/test_main.py leaves hand-written sys.modules stand-ins behind: 10 errors when paired with test_public_api #688 class of bug.
One caveat on provenance: the tail of the
tests/protocols/measurement overlapped with a concurrent
pytest -nexperiment from another investigation on the same host, pushing load to ~15 for the final 1-2
minutes. Every measured invocation was single-process at ~99% of one core, so contention effects are minor, but
that one number may be slightly inflated.Two things about item 2 found while raising the timeout PR — one is a prerequisite, one is a correction
A test parses this workflow and will not notice a positive selection
tests/test_tier_guard.py:92-109,WorkflowAgreementTests.test_ignore_flags_match_the_fixture_tier_constants,
reads.github/workflows/unit-tests.ymland asserts its flags equal two constants intests/_tiers.py:globs = set(re.findall(r"--ignore-glob='\*([^']+)'", text)) directories = set(re.findall(r'--ignore=tests/(\S+)', text)) self.assertEqual(globs, set(_tiers.FIXTURE_TIER_SUFFIXES)) self.assertEqual(directories, set(_tiers.FIXTURE_TIER_DIRS))
with
FIXTURE_TIER_SUFFIXES = ('_runtime.py', '_regression.py')and
FIXTURE_TIER_DIRS = frozenset({'integration'})attests/_tiers.py:140and:143. Its docstring names the
failure it guards: "somebody adds a fourth fixture-dependent naming convention to the workflow and the guard
goes on classifying those modules as unit-tier, flagging their perfectly legal capture reads."tests/_tiers.py:134also holdsUNIT_TIER_SELECTION, thetestjob's command quoted verbatim, with the
comment thatis_unit_tier()"is the executable statement of the same rule, and the two have to agree."Consequence for item 2. Partitioning the integration job means giving it a positive selection
(tests/integration/plus the two globs) instead of the unit job's negative one. Both regexes above only match
--ignore/--ignore-glob, so a positive selection is invisible to them — the workflow and_tierscould drift
apart in exactly the way this test exists to prevent, while it stays green. Whoever lands item 2 must extend this
test to cover the integration job's selection as well, and keepUNIT_TIER_SELECTIONaccurate. That is the test
change item 2 needs; it is not a change with "nothing to assert".Correction: these constants do not shortcut the skip-count prerequisite
An earlier reading suggested
FIXTURE_TIER_SUFFIXES/FIXTURE_TIER_DIRSmight already enumerate which tests
depend on fixtures, making the skip-count diff unnecessary. That is wrong, and the mechanism is the opposite.Those two constants define the fixture tier — modules that are fixture-dependent in their entirety and are
therefore excluded from the unit job. They say nothing about unit-tier modules that read a capture
conditionally. And such modules exist by design:tests/test_tier_guard.py:198reads "The skip idiom opts a
call out, because it is tier-safe already", andtests/_tiers.py:149names
SKIP_IDIOM_EXAMPLE = 'tests/toolkit/test_dpkt_unit.py'as the worked example, quoted in the guard's own failure
message at:552.So the guard's rule is that a unit-tier module may read a capture provided it handles absence at the call
site. Which means the population of unit-tier tests that skip when fixtures are absent and run for real only in
the integration job is real and non-empty — 39 files callsample_path()and 4 use an
except FileNotFoundErrorskip.The skip-count diff remains a prerequisite for item 2, and it is the cheap way to size it: compare skip counts
between the unit invocation (no fixtures generated — verified, thetestjob has nomake_samples.pystep) and
the full invocation aftermake_samples.py. Every test that runs in one and skips in the other must stay on the
integration side of the partition, or it stops being exercised anywhere.Item 4 measured:
pytest-xdistis worth adopting, with one prerequisite-n autowith the default--dist loadis safe today and roughly quadruples throughput.--dist loadfileand--dist loadscopeare not safe — they surface a pre-existingsys.modulesleak, now filed as #720.Directory Tests Serial -n 4load(×2 runs)Speedup tests/corekit179 (413 subtests) 138.53 s 39.32 / 38.51 s ~3.6× tests/protocols729 (1829 subtests) 719.06 s 178.43 / 189.57 s ~3.9× tests/foundation250 passed / 11 skipped (394 subtests) 277.04 s 70.73 / 73.69 s ~3.9× Pass/skip/warning/subtest counts were byte-identical to serial in every
loadrun.pytest-xdistis absent frompyproject.tomltoday (0 mentions).Prerequisite: #720. Until it lands, pin
--dist loadexplicitly rather than relying on the default, because switching toloadfilelater would reintroduce the failure.Not measured:
tests/vendor,tests/interface,tests/cli,tests/utilities,tests/integration,tests/dumpkit,tests/project. A single large indivisible subtest sweep in one of those would cap the speedup — underloadfilespecifically,tests/constalready drops to ~1.2× becausetest_const_enum_builtin_parity.pyconcentrates 500+ subtests in one file.A measurement trap worth recording
completed_at − started_atis not execution time. A job cancelled while still queued reportsstarted_at == created_atwithsteps: [], which manufactures a 78-minute "run" out of a 5-minute job. Filter on non-emptystepsbefore measuring anything from the Actions API.Two more timeout casualties since this issue was filed
Both at the 30-minute ceiling, both on required contexts, both blocking a merge-ready PR:
- docs(tests): stop stating how many captures are committed (#700) #703
Python 3.10— job107282865176, 30.2 min, inner stepRun unit testscancelled - fix(docs): qualify the ambiguous Type refs in the foundation rst tree #714
Integration Python 3.13— job107300119290, 30.28 min, inner stepRun full test suitecancelled
A
timeout-minuteskill reportsconclusion: cancelled, nevertimed_out, so it is indistinguishable from a force-push cancellation by conclusion alone. The tell is duration ≈ the ceiling with a cancelled inner step.These are the case for merging #713's ceiling raise (PR #716, currently merge-ready) before re-running either job — a re-run inherits the same 30-minute ceiling and can die the same way.
- docs(tests): stop stating how many captures are committed (#700) #703
- addedwipWork in flight - a covering PR is open or an agent is actively on itWork in flight - a covering PR is open or an agent is actively on it
on Sep 24, 2026 All four items have shipped. Audit against
origin/mainat55513f69e:item landed as evidence on origin/main1 — partition the integration job to its own files #727 Integration Python ${{ matrix.python-version }}job atunit-tests.yml:922 — stop the gate re-running on every push to main#731 gate chained off Unit Tests 3 — move 3.15 off the PR-blocking matrix #725 # 3.15 deliberately excludedat:58-59, repeated at:1094 — adopt pytest-xdist#725 python -m pytest -q -n auto --dist loadat:95and:193Plus #716 raised the pytest job timeouts 30 → 45 min.
Item 4's prerequisite was honoured:
--dist loadonly, becauseloadfileandloadscopesurface thesys.modulesleak — filed as #720 and fixed since. #725 also added theflockmutex that closed the xdist race in the tier guard.Removing the
wiplabel: nothing is in flight, so it was asserting something false.Recommend closing. Leaving it open in case there is further CI scope you have not written down — if there is, it wants listing here or in a fresh issue, since the four items above no longer describe any remaining work. #729 (dpkt never installed on any CI path) came out of the same area but is tracked separately and has a worker on it.
- addedneeds: decisionWaiting on the maintainer to decide — not blocked by other workWaiting on the maintainer to decide — not blocked by other workand removedwipWork in flight - a covering PR is open or an agent is actively on itWork in flight - a covering PR is open or an agent is actively on it
on Sep 24, 2026 Okay, let's close it with a closure note then.
Closure note
All four items shipped, verified against
origin/mainat9b2d927c2:item landed as evidence 1 — partition the integration job #727 Integration Python ${{ matrix.python-version }}job atunit-tests.yml:922 — stop the gate re-running on push #731 gate chained off Unit Tests 3 — 3.15 off the PR-blocking matrix #725 # 3.15 deliberately excludedat:58-59,:1094 — adopt pytest-xdist#725 pytest -q -n auto --dist loadat:95,:193Plus #716 raised the pytest job timeouts 30 → 45 min.
Item 4's prerequisite was honoured:
--dist loadonly, becauseloadfileandloadscopesurface thesys.modulesleak filed as #720 and since fixed. #725 also added theflockmutex that closed the xdist race in the tier guard — measured 18/32 failures with the lock off, 0/60 with it on.Not carried forward from this issue, tracked elsewhere: #729 (dpkt never installed on any CI path, 28 methods) came out of the same area and merged as #737. #738 records the remaining 106 dark methods across seven flags.
Closing per the maintainer's decision.
- removedneeds: decisionWaiting on the maintainer to decide — not blocked by other workWaiting on the maintainer to decide — not blocked by other work
on Sep 24, 2026
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Splitting the structural half of #713 out, so the one-line ceiling raise there can land and close on its own
instead of waiting on measurement.
The finding that reframed this
Raising
timeout-minutesdoes not make CI faster — it stops jobs dying. The actual cost is runnercontention, measured across five real runs:
3589019030835867536092358675318183586752618535867523310~285 minutes of real execution every time; wall-clock to green swings 6× purely on how contended the runner
pool was, with individual jobs staggered 82-161 minutes apart waiting for a runner. The fix is therefore to
remove jobs and duplicated work, not to give each job a longer leash.
Per-job execution for reference: unit 14.4-27.6 min, integration 17.4-29.4 min, gate ~27.2 min, all six
Compat Pythonjobs 1.4 min combined (they runcompileallplus an import smoke test, no pytest),changelog~6-10 s,lint4.0 min, CodeQL 2.4 min. A PR push is 22 jobs, 12 of which run pytest; a pushto
mainis 32 jobs. The repository is public, so Linux minutes are free — the cost is queue depth andwall-clock, not billing.
Already shipped — listed so nobody proposes them again
gate-only(e87ec884e, ci: run the unit tier once per event, not once per caller #392):unit-tests.yml:10defines the input, jobs gate on it at:51,:93,:205, and all four callers passgate-only: true(deploy-pages.yml:39,cron-vendor.yml:29,cron-conda.yml:31,create-release.yml:87). Reusable-workflow callers already run only the singlegatejob, not the 12-job matrix.concurrency:withcancel-in-progress(eabf89417, ci: stop dropping a commit's validation when the next one lands #394): present on all eight workflows.unit-tests.yml:38-40keys ongithub.refforpull_requestwithcancel-in-progress: true.Item 1 — partition the integration job to its own files
unit-tests.yml:156runs a barepython -m pytest -q, i.e. the whole suite, while thetestjob at:84-89runs
pytest -q --ignore=tests/integration --ignore-glob='*_runtime.py' --ignore-glob='*_regression.py'. So eachof the six
Integration Python 3.1xjobs re-runs every one of the ~118 files the matching unit job just ran,adding only ~30 files of its own.
Change
:156to run onlytests/integration/plus the two globs the unit job excludes.Expected saving: large but not yet quantified — it depends on whether integration's ~30 unique files are
cheap or expensive relative to the ~118 shared ones. A per-directory profile is the input needed before landing
this, and that measurement is in progress.
Risk to check first: whether any unit-scoped test implicitly depends on fixtures
examples/generators/make_samples.pygenerates only inside the integration and gate jobs. If so, those testspass today only because integration re-runs them after fixture generation, and partitioning would break them
silently.
Item 2 — stop the gate re-running on every push to
mainThe
gatejob (Gate (full suite, Python 3.14),:203) is byte-for-byte identical work toIntegration Python 3.14from the same commit: same interpreter, same barepytest -q, same tree. On a push tomainit fires up to four times, once pergate-onlycaller, on top of the matrix that already ran.Replace the
push-triggeredgate-only: truecalls indeploy-pages.yml,cron-vendor.ymlandcron-conda.ymlwith aworkflow_runtrigger onUnit Testscompletion, gated onconclusion == 'success'andchecked out at
head_sha.create-release.yml:6-9already uses exactly this pattern, chaining offVendor Update's completion, so there is a working template in-tree.Saving: ~108 min of duplicate execution per main push; main-push job count 32 → ~20-23.
Signal lost: none.
Item 3 — move Python 3.15 off the PR-blocking matrix
Ruleset
23497679requires 15 contexts —Python,Integration PythonandCompat Pythonfor 3.10-3.14only. 3.15 is verified absent from all fifteen, so it cannot block a merge, yet it costs 3 jobs per PR
push (22 → 19) and its legs are among the longest-running.
Remove the 3.15 legs from
unit-tests.yml'stestandintegrationmatrices and frompython-compatibility.yml, and add a nightly or weekly schedule instead.Signal lost: none for merge-gating. Only cost is detection latency — a 3.15 regression surfaces on the next
scheduled run rather than in the PR.
Item 4 —
pytest-xdist, undecidedpytest-xdistis absent frompyproject.tomland all three pytest invocations are barepytest -q, so there isno parallelism anywhere.
-n autowould dwarf items 1-3 if it works.The reason it may not: this suite mutates global state that parallel workers would race on — the registries
under
pcapkit/foundation/registry/**,aenum.extend_enumadding members to shared enum classes at runtime(measured at 16.7% of extraction self-time in #575), the
sys.modulesswapping that #674/#687/#688 alreadyhad to fix once, and
__warningregistry__ordering that warning-count assertions depend on. Subtest-heavysweeps are also indivisible units — a file reporting thousands of subtests lands in one worker and caps the
achievable speedup.
A feasibility assessment is in progress. Do not add
-nuntil it reports; a suite that passes serially andfails under
-nis worse than a slow one.Ordering
Item 1 is the largest win for the PR queue specifically and should go first once its measurement lands. Item 2 is
the largest absolute win and can land immediately — it needs no measurement and loses no signal. Item 3 is the
cheapest and safest. Item 4 is gated.
Each item is independent and should be its own PR; they touch different files and carry different risk.