Skip to content

CI: cut job count and duplicated work — contention, not job duration, is what makes CI slow #715

Description

@JarryShaw

Splitting the structural half of #713 out, so the one-line ceiling raise there can land and close on its own
instead of waiting on measurement.

The finding that reframed this

Raising timeout-minutes does not make CI faster — it stops jobs dying. The actual cost is runner
contention
, measured across five real runs:

run condition sum(execution) wall-clock to done
35890190308 main push, quiet queue 272 min 31 min
35867536092 PR, busy queue 292 min 178 min
35867531818 PR, busy queue 290 min 161 min
35867526185 PR, busy queue 284 min 103 min
35867523310 PR, busy queue 283 min 115 min

~285 minutes of real execution every time; wall-clock to green swings 6× purely on how contended the runner
pool was
, with individual jobs staggered 82-161 minutes apart waiting for a runner. The fix is therefore to
remove jobs and duplicated work, not to give each job a longer leash.

Per-job execution for reference: unit 14.4-27.6 min, integration 17.4-29.4 min, gate ~27.2 min, all six
Compat Python jobs 1.4 min combined (they run compileall plus an import smoke test, no pytest),
changelog ~6-10 s, lint 4.0 min, CodeQL 2.4 min. A PR push is 22 jobs, 12 of which run pytest; a push
to main is 32 jobs. The repository is public, so Linux minutes are free — the cost is queue depth and
wall-clock, not billing.

Already shipped — listed so nobody proposes them again

Item 1 — partition the integration job to its own files

unit-tests.yml:156 runs a bare python -m pytest -q, i.e. the whole suite, while the test job at :84-89
runs pytest -q --ignore=tests/integration --ignore-glob='*_runtime.py' --ignore-glob='*_regression.py'. So each
of the six Integration Python 3.1x jobs re-runs every one of the ~118 files the matching unit job just ran,
adding only ~30 files of its own.

Change :156 to run only tests/integration/ plus the two globs the unit job excludes.

Expected saving: large but not yet quantified — it depends on whether integration's ~30 unique files are
cheap or expensive relative to the ~118 shared ones. A per-directory profile is the input needed before landing
this, and that measurement is in progress.

Risk to check first: whether any unit-scoped test implicitly depends on fixtures
examples/generators/make_samples.py generates only inside the integration and gate jobs. If so, those tests
pass today only because integration re-runs them after fixture generation, and partitioning would break them
silently.

Item 2 — stop the gate re-running on every push to main

The gate job (Gate (full suite, Python 3.14), :203) is byte-for-byte identical work to
Integration Python 3.14 from the same commit: same interpreter, same bare pytest -q, same tree. On a push to
main it fires up to four times, once per gate-only caller, on top of the matrix that already ran.

Replace the push-triggered gate-only: true calls in deploy-pages.yml, cron-vendor.yml and
cron-conda.yml with a workflow_run trigger on Unit Tests completion, gated on conclusion == 'success' and
checked out at head_sha. create-release.yml:6-9 already uses exactly this pattern, chaining off
Vendor Update's completion, so there is a working template in-tree.

Saving: ~108 min of duplicate execution per main push; main-push job count 32 → ~20-23.
Signal lost: none.

Item 3 — move Python 3.15 off the PR-blocking matrix

Ruleset 23497679 requires 15 contexts — Python, Integration Python and Compat Python for 3.10-3.14
only
. 3.15 is verified absent from all fifteen, so it cannot block a merge, yet it costs 3 jobs per PR
push
(22 → 19) and its legs are among the longest-running.

Remove the 3.15 legs from unit-tests.yml's test and integration matrices and from
python-compatibility.yml, and add a nightly or weekly schedule instead.

Signal lost: none for merge-gating. Only cost is detection latency — a 3.15 regression surfaces on the next
scheduled run rather than in the PR.

Item 4 — pytest-xdist, undecided

pytest-xdist is absent from pyproject.toml and all three pytest invocations are bare pytest -q, so there is
no parallelism anywhere. -n auto would dwarf items 1-3 if it works.

The reason it may not: this suite mutates global state that parallel workers would race on — the registries
under pcapkit/foundation/registry/**, aenum.extend_enum adding members to shared enum classes at runtime
(measured at 16.7% of extraction self-time in #575), the sys.modules swapping that #674/#687/#688 already
had to fix once, and __warningregistry__ ordering that warning-count assertions depend on. Subtest-heavy
sweeps are also indivisible units — a file reporting thousands of subtests lands in one worker and caps the
achievable speedup.

A feasibility assessment is in progress. Do not add -n until it reports; a suite that passes serially and
fails under -n is worse than a slow one.

Ordering

Item 1 is the largest win for the PR queue specifically and should go first once its measurement lands. Item 2 is
the largest absolute win and can land immediately — it needs no measurement and loses no signal. Item 3 is the
cheapest and safest. Item 4 is gated.

Each item is independent and should be its own PR; they touch different files and carry different risk.

Activity

  1. added
    enhancementIssues requesting a new capability (set by the feature request template)
    perfPull requests that improve performance (perf: subject prefix)
    ciPull requests that change CI or workflow configuration (ci: subject prefix)
    on Sep 23, 2026
  2. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Measurement landed — and it moves the biggest lever off the CI config entirely

    A per-directory profile of the suite is in. The dominant cost is not the workflow structure. It is that
    pcapkit is cold-reimported once per test method.

    Where the 24.5 minutes actually goes

    Measured single-process, one pytest --durations -q per directory, on an otherwise-quiet 16-core box. The
    passed+skipped total (1739) matches pytest --collect-only -q exactly, so nothing was double-counted:

    directory wall tests subtests %
    tests/protocols/ 742.70s 729 1829 50.5%
    tests/foundation/ 291.38s 250/11 skip 394 19.8%
    tests/corekit/ 145.22s 179 413 9.9%
    tests/vendor/ 64.33s 55 54 4.4%
    tests/toolkit/ 63.94s 49/4 skip 51 4.3%
    tests/integration/ 55.76s 90/2 skip 115 3.8%
    tests/utilities/ 26.51s 104 100 1.8%
    tests/interface/ 19.81s 22 9 1.3%
    tests/const/ 18.06s 40 798 1.2%
    tests/dumpkit/ 17.93s 15 68 1.2%
    tests/ root 13.47s 57 35 0.9%
    tests/project/ 9.65s 126 479 0.7%
    tests/cli/ 1.05s 6 0 0.07%
    total 1469.8s ≈ 24.5 min 1739 4345

    protocols + foundation + corekit = 1179.3s, 80.2% of the suite.

    The new top lever: purge_modules in setUp rather than setUpClass

    tests/project/ averages 0.077s per test. tests/corekit/ averages 0.811s per test. The difference is that
    corekit purges pcapkit from sys.modules in setUp, so every test method pays a cold reimport.

    Verified independently by grepping the tree: 144 purge_modules call sites sit in setUp, against only
    10 in setUpClass/setUpModule, spread across 108 files. A cold import pcapkit measures 654ms, of
    which pcapkit/const/reg/apptype.py alone is 160ms (24.4%) — a StrEnum with ~8,949 members modelling the
    IANA port registry, built eagerly at class-body evaluation.

    Moving those purges to class scope would pay the tax once per class instead of once per method: roughly 10-13
    minutes off a 24.5-minute suite (40-55%)
    .

    This is the highest-value change available and also the riskiest. tests/conftest.py documents that the
    per-test isolation exists because of two real regressions (#660, #688) where a leaked sys.modules binding
    broke unrelated tests, and — line 128 — that the suite currently passes partly because unrelated directories
    "sort between the polluter and its victims". That is an accidental order-dependency. Any change here needs to
    prove no sibling test in a class depends on a purge another sibling would otherwise leave behind.

    Item 1 is confirmed and now quantified — it is bigger than first stated

    The delta between the full suite and the unit job's subset is 164.9s ≈ 2.75 min, 11.2% of total:
    tests/integration/ 55.76s, plus the *_runtime.py/*_regression.py files at 108.30s in protocols (led by
    test_pcapng_regression.py 18.77s, test_registry_runtime.py 16.95s, test_http_runtime.py 13.01s) and 0.81s
    in foundation.

    So partitioning unit-tests.yml:156 takes each Integration Python 3.1x job from ~24.5 min to ~2.75 min —
    ~21.75 min saved per job, ~130 min per PR push across the six. That is the largest CI-config win by a wide
    margin.

    The risk is the inverse of what this issue first stated. It is not that unit-scoped tests depend on fixtures
    and would break — the test job generates no fixtures at all (verified: unit-tests.yml lines 43-90 contain
    only an install step and the pytest step, no make_samples.py), so they already pass without them. The real
    exposure is that 39 files call sample_path() and 4 use an except FileNotFoundError skip idiom, so some
    tests may be skipping in the unit job and running for real only in the integration job. Partitioning would mean
    those never run for real anywhere. Before landing item 1, diff the skip counts between the unit invocation and
    the full invocation; any test that runs in one and skips in the other must stay on the integration side.

    Two hypotheses refuted, worth recording so they are not re-chased

    Collection overhead is negligible — pytest --collect-only -q over all 1739 tests takes 2.75s.

    Capture I/O is not the driver. The largest capture, http.pcap (171 KB, 1117 frames), parses once in
    1.22s CPU and is referenced in only 8 files. Aggregate parse/IO cost is on the order of 1-2% of the suite.
    make_samples.py runs in 2.84s.

    Subtests are numerically concentrated but not a wall-clock driver. Per-test-method cost is uniform regardless
    of subtest density — test_ipv6_extension_unit.py is 1.18s/phase and test_mh_unit.py 1.34s/phase, the same as
    everywhere else, because each method already pays the fixed reimport tax and the subtest iterations inside are
    cheap. Sampling or caching subtests would save little.

    A correction to #575's framing

    #575 attributes 16.7% of extraction self-time to aenum.extend_enum. At the function level that is refuted
    for the test suite: extend_enum is 245 calls, 0.010s self — under 0.1%. The cost is the EnumMeta
    class-construction machinery it depends on: all aenum functions together are 6.615s self (23.8%) and 22.339s
    cumulative (80.5%)
    of a profiled tests/const/ run, with builtins.__build_class__ at 21.280s cumulative
    (76.7%)
    and 1.53M abc.__subclasscheck__ calls from aenum's descriptor emulation. So the finding stands in
    substance and is wrong in attribution — optimising extend_enum itself would achieve nothing; shrinking eager
    enum construction would.

    The top-20 files, for whoever picks up the work

    Within protocols (exhaustive --durations=0): test_ipv6_extension_unit.py 72.22s, test_mh_unit.py 67.21s,
    test_pcapng_unit.py 65.49s, test_hip_unit.py 38.25s, test_link_unit.py 35.97s,
    test_protocol_code_registration_unit.py 34.98s, test_http_unit.py 34.66s, test_esp_unit.py 34.60s,
    test_tcp_udp_unit.py 33.15s, test_ipv4_unit.py 29.93s, test_protocol_base_unit.py 25.70s,
    test_schema_unit.py 19.35s, test_pcapng_regression.py 18.77s, test_registry_runtime.py 16.95s,
    test_transport_unit.py 16.36s, test_dispatch_bindings_unit.py 15.48s, test_ipv6_unit.py 13.49s,
    test_http_runtime.py 13.01s, test_tcp_runtime.py 10.13s, test_dispatch_registry_unit.py 9.59s.

    Those 20 files are 605.29s = 41.2% of the entire suite, and all of them call purge_modules(['pcapkit']) in
    setUp.

    Revised ordering

    1. purge_modules → class scope (new). 10-13 min, ~40-55%. Highest value, highest risk, needs the
      order-dependency work above. Source change, not CI.
    2. Partition the integration job (item 1). ~130 min/PR push. Needs the skip-count diff first.
    3. Stop the gate re-running on main pushes (item 2). ~108 min/main push, no signal lost, no measurement
      needed — can land now.
    4. Trim 3.15 from the PR matrix (item 3). 3 jobs/PR push, safest.
    5. pytest-xdist (item 4). Still assessing — and note the profiler independently flagged the same
      order-dependency at conftest.py:128 as the reason parallelism may resurface the twelve tests/project/ tests fail with TypeError: type 'ProtocolBase' is not subscriptable depending on what ran first: the fake-module helpers purge on entry and never restore on exit #660/tests/cli/test_main.py leaves hand-written sys.modules stand-ins behind: 10 errors when paired with test_public_api #688 class of bug.

    One caveat on provenance: the tail of the tests/protocols/ measurement overlapped with a concurrent
    pytest -n experiment from another investigation on the same host, pushing load to ~15 for the final 1-2
    minutes. Every measured invocation was single-process at ~99% of one core, so contention effects are minor, but
    that one number may be slightly inflated.

  3. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Two things about item 2 found while raising the timeout PR — one is a prerequisite, one is a correction

    A test parses this workflow and will not notice a positive selection

    tests/test_tier_guard.py:92-109, WorkflowAgreementTests.test_ignore_flags_match_the_fixture_tier_constants,
    reads .github/workflows/unit-tests.yml and asserts its flags equal two constants in tests/_tiers.py:

    globs = set(re.findall(r"--ignore-glob='\*([^']+)'", text))
    directories = set(re.findall(r'--ignore=tests/(\S+)', text))
    self.assertEqual(globs, set(_tiers.FIXTURE_TIER_SUFFIXES))
    self.assertEqual(directories, set(_tiers.FIXTURE_TIER_DIRS))

    with FIXTURE_TIER_SUFFIXES = ('_runtime.py', '_regression.py') and
    FIXTURE_TIER_DIRS = frozenset({'integration'}) at tests/_tiers.py:140 and :143. Its docstring names the
    failure it guards: "somebody adds a fourth fixture-dependent naming convention to the workflow and the guard
    goes on classifying those modules as unit-tier, flagging their perfectly legal capture reads."

    tests/_tiers.py:134 also holds UNIT_TIER_SELECTION, the test job's command quoted verbatim, with the
    comment that is_unit_tier() "is the executable statement of the same rule, and the two have to agree."

    Consequence for item 2. Partitioning the integration job means giving it a positive selection
    (tests/integration/ plus the two globs) instead of the unit job's negative one. Both regexes above only match
    --ignore/--ignore-glob, so a positive selection is invisible to them
    — the workflow and _tiers could drift
    apart in exactly the way this test exists to prevent, while it stays green. Whoever lands item 2 must extend this
    test to cover the integration job's selection as well, and keep UNIT_TIER_SELECTION accurate. That is the test
    change item 2 needs; it is not a change with "nothing to assert".

    Correction: these constants do not shortcut the skip-count prerequisite

    An earlier reading suggested FIXTURE_TIER_SUFFIXES/FIXTURE_TIER_DIRS might already enumerate which tests
    depend on fixtures, making the skip-count diff unnecessary. That is wrong, and the mechanism is the opposite.

    Those two constants define the fixture tier — modules that are fixture-dependent in their entirety and are
    therefore excluded from the unit job. They say nothing about unit-tier modules that read a capture
    conditionally. And such modules exist by design: tests/test_tier_guard.py:198 reads "The skip idiom opts a
    call out, because it is tier-safe already"
    , and tests/_tiers.py:149 names
    SKIP_IDIOM_EXAMPLE = 'tests/toolkit/test_dpkt_unit.py' as the worked example, quoted in the guard's own failure
    message at :552.

    So the guard's rule is that a unit-tier module may read a capture provided it handles absence at the call
    site. Which means the population of unit-tier tests that skip when fixtures are absent and run for real only in
    the integration job
    is real and non-empty — 39 files call sample_path() and 4 use an
    except FileNotFoundError skip.

    The skip-count diff remains a prerequisite for item 2, and it is the cheap way to size it: compare skip counts
    between the unit invocation (no fixtures generated — verified, the test job has no make_samples.py step) and
    the full invocation after make_samples.py. Every test that runs in one and skips in the other must stay on the
    integration side of the partition, or it stops being exercised anywhere.

  4. JarryShaw commented on Sep 23, 2026

    @JarryShaw
    OwnerAuthor

    Item 4 measured: pytest-xdist is worth adopting, with one prerequisite

    -n auto with the default --dist load is safe today and roughly quadruples throughput. --dist loadfile and --dist loadscope are not safe — they surface a pre-existing sys.modules leak, now filed as #720.

    Directory Tests Serial -n 4 load (×2 runs) Speedup
    tests/corekit 179 (413 subtests) 138.53 s 39.32 / 38.51 s ~3.6×
    tests/protocols 729 (1829 subtests) 719.06 s 178.43 / 189.57 s ~3.9×
    tests/foundation 250 passed / 11 skipped (394 subtests) 277.04 s 70.73 / 73.69 s ~3.9×

    Pass/skip/warning/subtest counts were byte-identical to serial in every load run. pytest-xdist is absent from pyproject.toml today (0 mentions).

    Prerequisite: #720. Until it lands, pin --dist load explicitly rather than relying on the default, because switching to loadfile later would reintroduce the failure.

    Not measured: tests/vendor, tests/interface, tests/cli, tests/utilities, tests/integration, tests/dumpkit, tests/project. A single large indivisible subtest sweep in one of those would cap the speedup — under loadfile specifically, tests/const already drops to ~1.2× because test_const_enum_builtin_parity.py concentrates 500+ subtests in one file.

    A measurement trap worth recording

    completed_at − started_at is not execution time. A job cancelled while still queued reports started_at == created_at with steps: [], which manufactures a 78-minute "run" out of a 5-minute job. Filter on non-empty steps before measuring anything from the Actions API.

    Two more timeout casualties since this issue was filed

    Both at the 30-minute ceiling, both on required contexts, both blocking a merge-ready PR:

    A timeout-minutes kill reports conclusion: cancelled, never timed_out, so it is indistinguishable from a force-push cancellation by conclusion alone. The tell is duration ≈ the ceiling with a cancelled inner step.

    These are the case for merging #713's ceiling raise (PR #716, currently merge-ready) before re-running either job — a re-run inherits the same 30-minute ceiling and can die the same way.

  5. added
    wipWork in flight - a covering PR is open or an agent is actively on it
    on Sep 24, 2026
  6. JarryShaw commented on Sep 24, 2026

    @JarryShaw
    OwnerAuthor

    All four items have shipped. Audit against origin/main at 55513f69e:

    item landed as evidence on origin/main
    1 — partition the integration job to its own files #727 Integration Python ${{ matrix.python-version }} job at unit-tests.yml:92
    2 — stop the gate re-running on every push to main #731 gate chained off Unit Tests
    3 — move 3.15 off the PR-blocking matrix #725 # 3.15 deliberately excluded at :58-59, repeated at :109
    4 — adopt pytest-xdist #725 python -m pytest -q -n auto --dist load at :95 and :193

    Plus #716 raised the pytest job timeouts 30 → 45 min.

    Item 4's prerequisite was honoured: --dist load only, because loadfile and loadscope surface the sys.modules leak — filed as #720 and fixed since. #725 also added the flock mutex that closed the xdist race in the tier guard.

    Removing the wip label: nothing is in flight, so it was asserting something false.

    Recommend closing. Leaving it open in case there is further CI scope you have not written down — if there is, it wants listing here or in a fresh issue, since the four items above no longer describe any remaining work. #729 (dpkt never installed on any CI path) came out of the same area but is tracked separately and has a worker on it.

  7. added
    needs: decisionWaiting on the maintainer to decide — not blocked by other work
    and removed
    wipWork in flight - a covering PR is open or an agent is actively on it
    on Sep 24, 2026
  8. JarryShaw commented on Sep 24, 2026

    @JarryShaw
    OwnerAuthor

    Okay, let's close it with a closure note then.

  9. JarryShaw commented on Sep 24, 2026

    @JarryShaw
    OwnerAuthor

    Closure note

    All four items shipped, verified against origin/main at 9b2d927c2:

    item landed as evidence
    1 — partition the integration job #727 Integration Python ${{ matrix.python-version }} job at unit-tests.yml:92
    2 — stop the gate re-running on push #731 gate chained off Unit Tests
    3 — 3.15 off the PR-blocking matrix #725 # 3.15 deliberately excluded at :58-59, :109
    4 — adopt pytest-xdist #725 pytest -q -n auto --dist load at :95, :193

    Plus #716 raised the pytest job timeouts 30 → 45 min.

    Item 4's prerequisite was honoured: --dist load only, because loadfile and loadscope surface the sys.modules leak filed as #720 and since fixed. #725 also added the flock mutex that closed the xdist race in the tier guard — measured 18/32 failures with the lock off, 0/60 with it on.

    Not carried forward from this issue, tracked elsewhere: #729 (dpkt never installed on any CI path, 28 methods) came out of the same area and merged as #737. #738 records the remaining 106 dark methods across seven flags.

    Closing per the maintainer's decision.

  10. removed
    needs: decisionWaiting on the maintainer to decide — not blocked by other work
    on Sep 24, 2026
  11. added this to the 1.5 milestone on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ciPull requests that change CI or workflow configuration (ci: subject prefix)enhancementIssues requesting a new capability (set by the feature request template)perfPull requests that improve performance (perf: subject prefix)

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions