Skip to content

Docs: record why the host-log claim budget is not the knob it looks like - #2108

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:hostlog-queue-investigation
Sep 3, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:hostlog-queue-investigation

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

kProducerClaimAttempts = 1024 is the most knob-shaped constant in the host-log
queue #2029 introduced, so the first response to a run reporting drops is to reach
for it. It is the wrong lever, and establishing that took a measurement rather
than an argument — so this records the measurement instead of leaving the next
person to redo it.

Docs-only. No code changes.

What it says

Verdict: do not change it. A producer wins its MPSC slot on the first
attempt
57–78% of the time, and the worst count across every workload shape
tried is 86 — twelve times under the bound.

threads (paced) won on 1st try max attempts to win
4 78% 4
16 57% 14
64 65% 37

Unpaced 64 threads: max 86.

A bound of 16 — the specific alternative considered — would have turned 787
successful writes (0.06%)
at 64 threads into claim_exhausted drops, to save a
worst-case CPU burn of ~1.6 µs instead of ~100 µs that never occurs.

Every loss lands in queue_full, at 4, 16 and 64 threads. That is structural,
not incidental
: a full queue exits on difference < 0 before spending any
attempt, so queue_full and claim_exhausted are nearly mutually exclusive by
construction. The constraint on loss is kQueueCapacity — the writer's drain
rate.

And the conclusion that was wrong first

The entry keeps the failed reading, because it is the one a reader is likely to
repeat. A saturating benchmark reports claim_exhausted == 0 and looks like
an answer — but saturation routes producers through the queue-full early exit and
barely visits the claim loop, so it says nothing about the attempt distribution
below the bound. That distribution is what decides whether the loop holds the
calling thread and whether a smaller bound would fit. Pacing the producers is
what puts the claim path under test.

Two questions were being conflated: does the budget cause drops (the counter
answers it) and how many attempts does a producer spend (the counter cannot).
Only the per-cause breakdown #2029 added makes either answerable, since a single
total cannot separate a queue that is too small from a budget that is too tight
from a destination that is broken.

Two properties recorded, not fixed

Neither is required by #1792 item 6 — which asked for "never block; drop and
count" and gets both — and neither is covered by a test:

  • Head-of-line blocking. pop() is strictly in-order, so a producer
    preempted between claiming position P and publishing sequence = P+1 parks
    the writer on P; records at P+1 and beyond cannot drain even though they
    are published, the queue fills behind the gap, and other producers begin
    dropping. The writer sleeping rather than spinning is tested
    (WriterSleepsWhenAdjacentProducersPublishOutOfOrder); the queue not backing
    up is not — different properties.
  • Fairness. The losses under contention fall on the slowest producers, which
    is the opposite of the useful bias for diagnostics: the thread that is stuck is
    often the one worth observing.

Caveats stated in the entry

The 45–94% drop rates it quotes come from deliberately pathological unpaced
workloads and are not representative; what they establish is the queue's role
— it absorbs bursts, not sustained overload, which is what its own comment claims
and now has a number behind. Everything is one machine (320-core aarch64 at load
~80) and two workload shapes, so the defensible reading is "1024 has ~12× headroom
over the worst observed and 16 does not", not "1024 is optimal". A value covering
everything observed with margin would be ~128–256; that is stated, along with why
shrinking to it still is not worth doing.

Testing

Docs-only, so no build or suite is affected. markdownlint-cli2,
check_english_only and check_retired_names clean; both relative links
(../logging.md, ../dfx/host-trace.md) resolve.

Indexed in docs/investigations/README.md per discipline.md §4 — an unlinked
entry is invisible, and the index is the only discovery surface.

Follows #2029 (item 6 of #1792).

`kProducerClaimAttempts = 1024` is the most knob-shaped constant in the host-log
queue, so a run reporting drops invites reaching for it first. It is the wrong
lever, and the reason takes a measurement to establish rather than an argument.

A producer wins its MPSC slot on the first attempt 57–78% of the time, and the
worst count across every workload shape tried is 86 — twelve times under the
bound. A bound of 16, the alternative considered, would have turned 787
successful writes at 64 threads into drops to save a worst case of ~100 µs that
never occurs. Every loss lands in `queue_full` instead, and that is structural:
a full queue exits before spending any attempt, so the two causes are nearly
mutually exclusive.

The entry also records the first conclusion, which was wrong. A saturating
benchmark reports `claim_exhausted == 0` and looks like an answer, but
saturation routes producers through the queue-full early exit and barely visits
the claim loop, so it says nothing about the attempt distribution below the
bound. Pacing the producers is what puts that path under test — and only the
per-cause breakdown hw-native-sys#2029 added makes either reading possible, since a single
total cannot separate a queue that is too small from a budget that is too tight
from a destination that is broken.

Two properties of the implementation go in alongside, neither required by
hw-native-sys#1792 item 6 and neither covered by a test: a producer preempted between
claiming a position and publishing it parks the writer on that position, so
later published records cannot drain and other producers begin dropping; and
the losses under contention fall on the slowest producers, which is the
opposite of the useful bias for diagnostics.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 3, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 50 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Team

Run ID: 41cd33b3-bd59-47ec-9042-b6e3c29dcba0

📥 Commits

Reviewing files that changed from the base of the PR and between d42d465 and 322117d.

📒 Files selected for processing (2)
  • docs/investigations/2026-09-host-log-queue-claim-budget.md
  • docs/investigations/README.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ChaoWao
ChaoWao merged commit baca34e into hw-native-sys:main Sep 3, 2026
16 checks passed
@ChaoWao
ChaoWao deleted the hostlog-queue-investigation branch September 3, 2026 08:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant