Skip to content

test: fix two racy tests - #151

Merged
jonhoo merged 1 commit into
mainfrom
fix-racy-tests
Aug 8, 2026
Merged

jonhoo merged 1 commit into
mainfrom
fix-racy-tests

Conversation

@jonhoo

@jonhoo jonhoo commented Aug 8, 2026

Copy link
Copy Markdown
Collaborator

record_nodrop and recorder_drop_staged deadlock permanently in CI. When it happens the test binary never exits, so the job runs until GitHub's timeout with no failure output. It shows up most often under cargo llvm-cov, but nothing about coverage instrumentation is special here; it just makes the whole suite slower, and the suite also runs mt_record_static/mt_record_dynamic, which keep 32 CPU-bound threads on a 2-4 vCPU runner. The reader thread is then far more likely to be preempted in the one window where this race is lost.

refresh() reads recorders under truth, bumps phase, and then blocks in recv() until it has received that many histograms. That recv() cannot fail, since the Sender lives in the Arc<Shared> the SyncHistogram itself holds, so the wait is unbounded. A live Recorder only sends when a subsequent write observes the new phase, or when it is dropped.

Both tests had the writer perform a fixed number of writes and then park on a Barrier that the reader only reaches after refresh() returns. If the writer got through all of its writes before the reader's phase.fetch_add, it never observed the phase shift, parked on the barrier, and could not write again -- while the reader blocked forever on a histogram nobody would send, and so never reached its own barrier.wait(). Neither side can make progress. Confirmed by forcing exactly that interleaving under gdb: breaking just before the recorders read is enough to hang both tests.

So the writers now record until the reader tells them its phase shift is through, which is the pattern clone_idle_recorder already uses. The handoff is explicit instead of depending on who gets scheduled first.

The cost is that record_nodrop loses its exact-count assertions: the number of writes is no longer fixed, so it can only check that the phase shift carried at least one sample across and that everything it carried was TEST_VALUE_LEVEL. What the test actually guards -- that refresh() is released by a write while the recorder is still alive -- is unchanged, and is now guaranteed rather than hoped for. recorder_drop_staged keeps its exact assertions, since the writer returns its own count.

`record_nodrop` and `recorder_drop_staged` deadlock permanently in CI.
When it happens the test binary never exits, so the job runs until
GitHub's timeout with no failure output. It shows up most often under
`cargo llvm-cov`, but nothing about coverage instrumentation is special
here; it just makes the whole suite slower, and the suite also runs
`mt_record_static`/`mt_record_dynamic`, which keep 32 CPU-bound threads
on a 2-4 vCPU runner. The reader thread is then far more likely to be
preempted in the one window where this race is lost.

`refresh()` reads `recorders` under `truth`, bumps `phase`, and then
blocks in `recv()` until it has received that many histograms. That
`recv()` cannot fail, since the `Sender` lives in the `Arc<Shared>` the
`SyncHistogram` itself holds, so the wait is unbounded. A live
`Recorder` only sends when a *subsequent* write observes the new phase,
or when it is dropped.

Both tests had the writer perform a fixed number of writes and then park
on a `Barrier` that the reader only reaches after `refresh()` returns.
If the writer got through all of its writes before the reader's
`phase.fetch_add`, it never observed the phase shift, parked on the
barrier, and could not write again -- while the reader blocked forever
on a histogram nobody would send, and so never reached its own
`barrier.wait()`. Neither side can make progress. Confirmed by forcing
exactly that interleaving under gdb: breaking just before the
`recorders` read is enough to hang both tests.

So the writers now record until the reader tells them its phase shift is
through, which is the pattern `clone_idle_recorder` already uses. The
handoff is explicit instead of depending on who gets scheduled first.

The cost is that `record_nodrop` loses its exact-count assertions: the
number of writes is no longer fixed, so it can only check that the phase
shift carried at least one sample across and that everything it carried
was `TEST_VALUE_LEVEL`. What the test actually guards -- that
`refresh()` is released by a write while the recorder is still alive --
is unchanged, and is now guaranteed rather than hoped for.
`recorder_drop_staged` keeps its exact assertions, since the writer
returns its own count.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jonhoo
jonhoo merged commit ea926c4 into main Aug 8, 2026
18 checks passed
@jonhoo
jonhoo deleted the fix-racy-tests branch August 8, 2026 19:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant