fix(tests): three reds that keep CI wrong — signal-guard race, process-wide caplog, stale release allowlist - #2935
Conversation
…a process-wide caplog, and a stale release allowlist Three failure classes have been redding `backend-unit-test` on dev and the regression diff on unrelated PRs (#2920, #2924, #2927, the 09-18 train): 1. `test_subprocess_pgroup` — `signal_guard.guarded_killpg` walked the group's members, then read each member's cgroup; the harness parent exiting in between (that IS the scenario under test) made `_cgroup_of` answer None, which read as "outside the session cgroup", and a kill of a group that was entirely ours a millisecond earlier was refused. Membership is now decided per LIVE member: a pid that vanished (or is a zombie) is not a member. A live pid whose cgroup cannot be read stays foreign — fail closed, unchanged. Two tests pin both halves; reverting the guard turns the race test red. 2. `test_2789…test_retry_budget_is_logged_with_its_cause` asserted over EVERY caplog record in the process, so a background task left by an earlier test (order-dependent under a random seed) logging an unrelated ERROR read as `['ERROR', 'WARNING'] == ['WARNING']`. It now asserts over the module's own logger. 3. `test_2814_workflow_trigger_parity` — every ACCEPTED_UNTIL_RELEASE entry is declared on `main` since the v0.9.5 cut and the guard has said "prune these" on every run since. Pruned, as the guard was designed to demand. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Full unit suite on this branch under The 15 are the same on both seeds and pre-existing on |
…ived `mine()` was introduced to stop a background task's unrelated record from reading as this test's own, and applied at four sites. Six reads of the process-wide `caplog.records` survived in the same function, after the last `mine()` call — including `:523`, the exact shape the fix was written for, and two `assert not caplog.records` that any stray record from any logger reddens. Verified by negative control: with an unrelated ERROR emitted inside the third phase's `caplog.at_level` window, the pre-fix assertions fail at `assert "30s already spent" in caplog.records[0].message`; with `mine()` they pass. `caplog.records` now appears once, in `mine()` itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…cstring `_is_gone` returned True for every OSError, so a LIVE pid whose `/proc/<pid>/status` cannot be read (EACCES under a `hidepid=` mount, a malformed line) was classified as gone, dropped from `live`, and the `killpg` proceeded. The docstring two lines above states the opposite contract: "a LIVE pid whose cgroup we cannot read is treated as foreign (fail closed)". Only the vanished-pid case is `gone` — FileNotFoundError. Every other OSError/IndexError now keeps the pid in `live` so the cgroup check can refuse the kill. This guard replaces os.kill/os.killpg for the whole unit suite and exists because a mis-fire SIGKILLed a developer's desktop twice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
merge-train — two mechanical fixes pushed to this branch (
|
vybe
left a comment
There was a problem hiding this comment.
merge-train: batch validated on train/20260921-1644 (#2939) — full suite green, regression diff shows 0 failures across all three seeds and records this PR as fixing dev's standing red (test_2814). Two mechanical fixes pushed to this branch during assembly and announced above.
Three failure classes have been redding
backend-unit-testondevand theregression diffon unrelated PRs (#2920, #2924, #2927, the 09-18 train). None is a product defect; all three are test-side and each is fixed at its cause, not by retrying.test_subprocess_pgroup(order/timing flake) —tests/signal_guard.py::guarded_killpgwalked the group's members and then read each member's cgroup. The harness parent exiting in between — the scenario those tests exist to exercise — makes_cgroup_ofanswerNone, which read as "group contains a process outside the session cgroup", refusing a kill of a group that was entirely ours a millisecond earlier. Membership is now decided per live member (_is_gone:/proc/<pid>absent or stateZ). A live pid whose cgroup cannot be read stays foreign — fail closed, unchanged. Two tests intest_2845_signal_guard.pypin both halves; with the guard fix reverted the race test goes red (negative-controlled).test_2789…test_retry_budget_is_logged_with_its_causeasserted over every caplog record in the process; a background task left by an earlier test (seed-dependent) logging an unrelated ERROR read as['ERROR','WARNING'] == ['WARNING']. It now asserts overservices.task_execution_service's own records.test_2814_workflow_trigger_parity— everyACCEPTED_UNTIL_RELEASEentry is declared onmainsince v0.9.5 and the guard has said "prune these" on every run since. Pruned, as it was designed to demand — this is the standing red ondev's ownbackend-unit-test.Verification: the affected suites 77 pass locally; the full unit suite under
pytest-randomlyat CI seeds 12345 and 67890 (results appended below when they land).🤖 Generated with Claude Code