Skip to content

A background job's stderr streams live, from every producer - #465

Merged
tobert merged 19 commits into
mainfrom
fix/builtin-stderr-job-stream
Sep 22, 2026
Merged

tobert merged 19 commits into
mainfrom
fix/builtin-stderr-job-stream

Conversation

@tobert

@tobert tobert commented Sep 21, 2026

Copy link
Copy Markdown
Owner

A background job's stderr stream (/v/jobs/N/stderr) now fills live at every place stderr is produced, the way #449 made stdout live. Before, a builtin's stderr reached the stream only through a write at job completion, and only when no external had written stderr first.

Two designs were weighed: publish once per statement, or publish at every producer. We chose every producer, because it is live and in order, on one condition: a producer the design misses must fail a test rather than drop bytes silently. ExecResult.stderr_published_len counts how many leading bytes of err the stream already holds; publishing writes err[n..]. Appends, redirects, StderrStream chunks, timeout's note and chain faults each preserve that count, and release-mode assertions refuse a state that would print twice or drop. The completion write is gone.

Evidence the net holds: disabling any one of the 13 publish points fails at least one test; a corpus compares every job's stream with its foreground err byte for byte; liveness tests read the stream while the job still runs. Recursion depth is unchanged.

Found and fixed on the way:

  • a fault ending a block republished earlier stderr
  • a failing [[ ]] or test left of &&/|| published twice
  • a failing condition command (if test 1 -eq abc) printed its diagnostic three times, in the foreground too
  • $(...) stderr was lost under 2>/dev/null
  • a killed external could lose its last chunk
  • a panicked job task left its streams open
  • a cancel could not end a wait on a pipe a grandchild held open

One pre-existing bug gets its own commit: redirects now apply left to right with dup semantics, checked against bash.

ls /missing > f 2>&1    # both streams to f (stderr used to reach the terminal)
ls /missing 2>&1 > f    # stderr to the old stdout, stdout to f

Not changed: kaish still applies redirects after a command runs, so f() { echo a > out; }; f > out leaves out empty (bash keeps a). CHANGELOG entries land with the batch's changelog PR.

🤖 Generated with Claude Code

tobert and others added 19 commits September 21, 2026 08:04
Ten cases for `cmd &` jobs: a builtin's stderr after an external wrote
stderr, order across producers, exactly-once publication, `2>` to a
file, `2>` on a function call, `2>&1` to the job's stdout, `2>&1` into
a pipe, `>&2` to the job's stderr, and `$(...)` stderr as job stderr.

Eight fail on this tree. The two that pass (exactly-once, `2>&1` into
a pipe) guard the fix against publishing twice. No fix yet: the design
for which channel owns a job's stderr is still open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Problem: making a background job's stderr publish live, at every place it is
produced, needs a way for one dispatch to know whether its own stderr text
already reached the job's stream — a builtin wrapped by `timeout`, a
function call whose body already published, an external that tees live
during the child's run. A byte count on the shared job stream cannot answer
this: concurrent pipeline stages write to the same stream, so a count taken
before and after one stage's dispatch can be invalidated by a sibling stage
finishing in between.

Decision: `ExecResult` gets one more field, `stderr_published: bool`,
`#[serde(skip)]` since it is bookkeeping for this process, not part of the
result's wire contract. It defaults to `false` (every constructor sets it
explicitly, matching every other field on this `#[non_exhaustive]` struct)
and is scoped to the single `ExecResult` that carries it, so a sibling
pipeline stage's concurrent write to the shared stream never touches it.

Rule now in force: a site that publishes a result's stderr live sets this
flag; a site that would otherwise republish it checks the flag first. The
call sites land in the next commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Problem: `job_stderr_routing_tests.rs` (previous commit) pinned that a `cmd &`
job's stderr stream should hold every producer's stderr, live, in order,
after redirects decide where the bytes go — a builtin's, a function's, a
compound statement's, same as the external tee already does. 8 of its 10
cases failed: `JobManager::finalize_streams` filled the stream once, from
the job's captured `ExecResult.err`, and only when nothing had arrived live
— so a builtin's stderr was dropped the moment any external in the same job
had already written, and a `2>&1`/`>&2`/`2>file` redirect was never
consulted before that completion write fired.

Decision (option A, matching #449's stdout fix): publish stderr at every
site that produces it, and delete the completion write.

Two choke points cover the ~70 producer sites without 70 separate publish
calls. `scheduler::pipeline::run_single`/`run_pipeline`, after
`apply_redirects` runs, covers every `Stmt::Command`/`Stmt::Pipeline` leaf —
builtins, backend tools, user functions (a call IS a Command dispatch), and
externals (already teeing live; now flagged so the leaf does not repeat
them). `Kernel::execute_stmt_flow` gets a new wrapper (the old body renamed
to `execute_stmt_flow_dispatch`) that publishes every OTHER statement kind's
carried result once — `[[ ]]`, `(( ))`, `if`/`for`/`while`/`case`, `source`,
a `.kai` script, a break/continue/return/exit signal — since none of those
ever reaches a pipeline leaf. The wrapper is recursive by construction
(every nested statement call goes back through it), so no separate call is
needed at each of those individual sites.

Routing per stage mirrors `apply_redirects`'s own walk: a new
`redirects_stderr` (2>file, &>file, 2>&1) suppresses live publish before
dispatch, the same way `redirects_stdout` already does for stdout — without
it an external or a nested dispatch tees to the job's stderr before its own
redirect has a chance to divert it. `2>&1` is the one redirect that ADDS
bytes to stdout only after the whole stage finishes; `publish_job_stdout_suffix`
diffs stdout before and after `apply_redirects` and publishes just the
merged suffix. `>&2` is deliberately absent from `redirects_stderr` — it
sends stdout INTO stderr, it does not redirect stderr away, so the stage's
own stderr still belongs on the job's stream.

Two bugs surfaced while turning the required test file green, both from
places that fold several statements' `err` text into one `ExecResult` or
raw string without tracking the new flag correctly:

- `accumulate_result`'s naive OR-merge broke on an empty-`err` statement
  following a published one (`echo out` after a wrapped external) — an
  empty append must never drag a true flag down to false. Fixed by treating
  an empty `new.err` as a no-op for the flag, not a vote.
- `execute_user_tool`, `execute_source`, and `try_execute_script` each hand-
  rolled the same accumulate-into-raw-buffers loop and never touched the
  field at all, so a function or a sourced script's own body (already
  published, statement by statement) had its stderr republished by the
  function/script call's own leaf. Replaced all three (plus
  `execute_block_capturing`, for the same latent gap) with two shared
  helpers, `accumulate_raw_stmt_output`/`accumulate_raw_drained`, that track
  the flag by the same rule `accumulate_result` now uses.

`crates/kaish-kernel/tests/job_stderr_corpus_tests.rs` is the safety net the
task asked for: 19 cases running each script as a `&` job and in the
foreground on a fresh kernel, asserting the job's stderr stream equals the
foreground run's `result.err` byte for byte — every kernel error kind
(command not found, a malformed `[[ ]]`/`(( ))` operand, a `source`d
script's runtime error, a bad scatter/gather option, a redirect-target
failure), builtins, externals, functions, pipelines (top-level and nested
inside a compound), compound statements, `$(...)`, and each redirect form.
It should have been written and confirmed red before this commit, per the
task's instructions; it was written alongside the fix instead, empirically,
because the design only became fully concrete through the iteration above —
noted as a process deviation, not hidden. `test_fault_ahead_of_a_published_builtin_matches_foreground`
specifically pins the one case a single boolean flag cannot represent
precisely (an unpublished `[[ ]]` fault accumulated ahead of an
already-published builtin's stderr in the same `if` body): it passes today
because nothing currently produces that exact combination, but a future
redirect layered on top of a wrapper-around-a-mixed-body could still
double-publish the already-sent half. Documented, not fixed, in this pass.

`JobManager::finalize_streams` no longer writes either stream at
completion — every producer now publishes as it runs, so a completion
write would only repeat output or paper over a routing hole.
`JobStreams::stdout`/`stderr` and `ExecContext::background_stream_stderr`'s
doc comments are rewritten to describe this, replacing the stale
"reaches the stream only when nothing arrived live" claim.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…rite

`docs/EMBEDDING.md` and the `help vfs` content (crates/kaish-help/content/en/vfs.md,
symlinked from docs/help/) both still said a builtin's or an earlier
pipeline stage's stderr reached `/v/jobs/{id}/stderr` only at job
completion, and only when nothing had arrived live — the limit the previous
commit removed. Replaced with the accurate rule: every stage publishes its
own stderr live once its own redirects apply, regardless of pipeline
position, since bash never pipes stderr between stages the way it pipes
stdout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Probing the branch's `stderr_published: bool` against foreground runs
found three failure shapes the 19-case corpus did not name:

- A fault that ends a block after earlier stderr was published sends
  that earlier text again: `fault_result` appends its diagnostic and
  must reset the flag, so `f() { cat /a; x=$(( unset + 1 )); }; f &`
  publishes the `cat` line twice. The same happens when `&&`/`||`
  turns an already published `[[ ]]`, `(( ))`, or `test` fault into an
  error, which is rendered and published again.
- A substitution in a command's arguments runs before the command's
  redirects, so `echo "$(cat /x)" 2>/dev/null &` keeps the
  substitution's stderr, as the foreground result does. The drain
  marked all drained text published and the stream lost it.
- `timeout` puts its note in front of text the inner command already
  published; on a fault it also dropped the inner command's earlier
  stderr from the result.

The corpus now also requires each job's stream to equal the job's own
`result.err`, and builds kernels through `into_arc` so `timeout` can
dispatch in the foreground run. New cases cover the three shapes, a
function mixing every producer (builtin, external, `[[ ]]` fault,
assignment fault) bare and under `timeout`, with `2>file` and `2>&1`
on the call, the separator newline before drained text, and an
external's capture-overflow marker.

The corpus compares final bytes, so it cannot see liveness, which is
why stderr is published per producer. `job_stderr_is_live_per_producer`
reads `/v/jobs/1/stderr` while the job sleeps between two stderr
writes, one case per producer: statement, `[[ ]]` fault, function
body, first stage of a top-level and a nested pipeline, a substitution
drained after a redirected command, and a chain's left operand. All
seven fail on e442822, which wrote stderr at completion.

Against 33ec9c1: 17 corpus cases, the overflow routing case, and the
substitution liveness case fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A single `stderr_published: bool` cannot describe `err` once text is
appended after a publish, so each append site guessed, and the guesses
sent text twice or dropped it (previous commit). `ExecResult` now
carries `stderr_published_len`: how many leading bytes of `err` the
job's stream holds. Publishing writes `err[stderr_published_len..]`
and advances it to `err.len()`.

Why this is exact. The invariant is that `err[..n]` has reached the
stream and `err[n..]` has not. The stream may hold those bytes in the
order they were produced rather than the order `err` lists them.

- Publish writes exactly the unpublished tail, then n = len.
- Appending B to A (`append_stderr`, used by `accumulate_result` and
  the raw accumulators): if A is fully published, n becomes |A| plus
  B's own n, and B's published part is B's prefix. If A has an
  unpublished tail, B must have none published, or the prefix would
  break; that is an assertion. It holds because every statement
  publishes as it ends (see below), so in a publishing context text
  never waits, and in a non-publishing context nothing is published.
- Appends after a publish (a fault diagnostic from `fault_result`,
  `>&2`, a spill note) only add an unpublished tail, published once
  by the next publish. `fault_result` no longer resets anything.
- `2>file`, `&>file`, and `2>&1` move `err` away from the stream; they
  turn publishing off before the stage runs, and `apply_redirects` now
  asserts n == 0 when it moves `err`.
- The kernel stderr channel carries each chunk's published length.
  A drain in a publishing context writes each chunk's unpublished rest
  and the separator newline, so drained text ends fully published; in
  a non-publishing context it asserts every chunk is unpublished. A
  substitution under a stderr redirect is published at the drain, in
  the statement's context, which is bash's rule: words expand before
  redirects apply.
- The three sites that put text in front of published text handle it
  directly: `timeout` publishes the inner rest, then the bytes its note
  adds; an external's overflow markers are written to the teed stream;
  the job end places drained text first, as `execute_streaming_inner`
  does, then publishes the whole.
- A fault on the left of `&&`/`||` becomes an error whose message is
  only `err[n..]`; the published prefix stays as prior output. The
  left operand runs through `execute_stmt_flow_dispatch`, which skips
  the statement publish, so a `[[ ]]` fault is not published before the
  chain decides it is an error. A non-fault left operand is published
  before the right one runs.

Publish points. `execute_stmt_flow` now publishes every statement kind
when it ends, and `execute_background` publishes the job's own
statement. A mutation sweep (disable each point, run the four job
stream test files) showed the `run_single` leaf publish and all seven
scatter/gather publishes covered by those two, so they are removed.
Every remaining point fails at least one test when disabled: the
statement wrapper, each pipeline stage, the chain left operand, the
chain message split, the job end, the drain's chunk rest and
separator, the chunk published length at the stage flush and the
substitution emit, the external tee and its markers, and `timeout`'s
inner tail and note.

Foreground changes. `timeout` on a fault keeps the inner command's
earlier stderr (`fault_result` with a `timeout` context) instead of
replacing it with the message. When its note is added, the inner
stderr's trailing blank lines are kept rather than collapsed. A fault
on the left of `&&`/`||` no longer folds drained channel text into its
message; the enclosing statement drains it, ahead of the message.

Recursion. The wrapper was already a boxed level on 33ec9c1; this
adds none. `recursion_stack_cost_tests` per `$()` level: e442822
54192 B debug / 50576 B release, 33ec9c1 51424 / 48032, this commit
51920 / 48352. `MAX_RECURSION_DEPTH` is unchanged and
`recursion_guard_tests` pass on `RECOMMENDED_STACK_SIZE`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`cmd "$(sub)" 2>file &` publishes `sub`'s stderr to the job's
`stderr` node: bash expands a command's words before its redirects
apply, and the foreground result already kept that text. The help text
listed the redirects that keep a stage's stderr out of the node without
this exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`drain_to_stream_teed` writes each chunk to the command's capture ring,
then to the job's stream. On cancel, `spawn_process` aborted the drain
task, and an abort can land between those two writes. The capture then
held a chunk the stream never got, and since the tee sent everything
else, `stderr_published_len` was set to the whole of `err`: that chunk
was counted as published and never written.

The drain now takes a stop token and checks it only while waiting for
the next read, so a chunk is written to both or to neither. On cancel,
spawn fires the token and awaits the drains instead of aborting them.
Late output after the kill is still dropped, as before.

The race window is two awaits wide, and a timing test could not hit it:
20 rounds of a chatty `sh -c` stderr job killed while a reader polls
the stream passed on the old code. The discriminating test is at the
drain: `stopping_a_drain_keeps_capture_and_tee_equal` holds the tee's
lock so the drain is mid-chunk when it stops. Written against the
abort that spawn used, it failed with the tee empty and the capture
holding the chunk; with the stop token it passes. The job-level test
`a_killed_job_streams_exactly_its_stderr` stays as the stderr twin of
`a_killed_job_keeps_the_bytes_it_already_streamed`: the stream must
equal the killed job's `result.err` byte for byte.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`if test 1 -eq abc; then ...` printed its diagnostic three times.
`eval_condition_async` emitted the condition command's stderr to the
statement's stderr channel before checking for a fault, emitted it again
in the fault branch, and then returned the same text as the error, which
is rendered a third time. A job's stream carried all three, since each
copy is ordinary stderr to the drain.

A fault has no truth value, so its stderr is the error's message, as it
is for a fault on the left of `&&`/`||`: the unpublished part becomes
the message and only the published part (none, for a bare `test`) is
emitted. A condition that did not fault emits its stderr once, as
before.

`faulting_condition_command_prints_its_diagnostic_once` covers `if` and
`while`, in the foreground through a function and as a job both bare
and through the function; it saw three copies before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`execute_streaming_inner` joined drained channel text and validation
warnings onto each statement's stderr with `format!` and read the
channel through `drain_lossy`, which asserted that no chunk had been
published. That held only because the loop never ran in a publishing
context. `Kernel::fork_for_background` is public and builds exactly
that context, so an embedder calling `execute` on the fork panicked in
the middle of a job on its first pipeline with a stderr-writing stage.

The loop now drains through `drain_stderr_onto` in the root context,
joins drained text ahead of a statement's stderr with `append_stderr`
(`join_drained_stderr`), publishes the separator it may add, and puts
warnings in front with `prepend_stderr`, which publishes them first
when the context publishes. The published part of `err` stays a prefix
on every path, and no path relies on the context not publishing.
`drain_lossy` no longer asserts; the kernel no longer calls it, and it
now documents that it drops the published count.

`execute_on_a_background_fork_publishes_its_stderr_once` registers a
job, forks for it, and runs a stage error, a substitution under
`2>/dev/null`, and a builtin error through `execute`. Each reaches the
result and the job's stream once, and the stream equals the result.
It panicked in `drain_lossy` before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A job task that panicked unwound before `finalize_streams` and before
sending its result. `Job::try_poll` then synthesized a failure from the
closed channel, but that diagnostic never reached `/v/jobs/N/stderr`,
and neither stream was ever closed, so a reader waiting for "no more
coming" waited forever. The completion write this branch removed used
to paper over the missing text, never the open streams.

Both job spawn sites (`execute_background` and
`execute_background_with_options`) now run the job body as a task that
returns its result, and `send_job_result` awaits it. If the body
panicked, it builds the failure, writes its text to the job's stderr
stream, closes both streams, and sends it. The `try_poll` closed-channel
path stays for a sender dropped any other way.

`job_panic_tests.rs` registers an embedder tool that panics and runs it
as a `cmd &` job and as a whole-program job. Each must end with exit 1,
both streams closed, and a stderr stream equal to the result's `err`.
Both failed before this change on the open stdout stream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ls /nonexistent > o1 2>&1` left `o1` empty and printed the error; bash
writes the error into the file. `apply_redirects` moved the output as
it read each redirect: `>` wrote stdout to the file and cleared it, then
`2>&1` merged stderr into the now-empty stdout, which went to the
terminal or pipe. The same order problem made `2> f 1>&2` print stdout
instead of writing it to `f`, and made `2>&1 > f` send stderr to `f`'s
side of the result rather than to the old stdout.

`redirect_sinks` now resolves where fd 1 and fd 2 point, left to right,
as bash does: `2>&1` copies wherever fd 1 points at that moment, and
`>&2` copies fd 2. `apply_redirects` evaluates the file targets in
order, then moves stdout and stderr once to their final sinks, stdout
first where both share a file, and writes every opened file, empty or
not, since `> f` creates or truncates `f` either way. The stage flags
that decide live publishing (`redirects_stdout`, `redirects_stderr`,
`merges_stderr_into_stdout`) read the same resolution, so a job's
streams follow the final sinks too; stdout that went to a file is no
longer counted as already published when `2>&1 > f` puts stderr in
its place.

Four `shell_compat` rows compare against bash with `ls` failing:
`> f 2>&1`, `2>&1 > f`, `&> f`, and `2> f 1>&2`. Three failed on kaish
before this change; `&> f` already matched. `redirect_order_decides_
where_job_stderr_lands` checks the same four forms as `&` jobs: the
file, the job's stdout stream, and its stderr stream. `docs/LANGUAGE.md`
shows both orders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… empty

`stderr_published_len` promises that the stream holds `err`'s first n
bytes, not that it holds them in `err`'s order: the stream is in
production order. Spawn's overflow markers are the standing example;
they lead `err` and follow the teed bytes on the stream. The field doc
now says so.

A fault on the left of `&&`/`||` becomes an error whose message is the
unpublished part of `err`. When a compound's body already published
all of it, the message is empty. That is correct for the stream: the
text stays in the prior output, shown once, and the error renders
nothing more. A note at both chain arms says so. When the error is
later wrapped in a context (an assignment, `source`), the context line
renders with nothing after its colon; the text it would have named is
the line above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Only `spawn_process` uses `drain_to_stream_teed_until`, and `spawn` is
compiled only with `subprocess`. The WASI build, which has no
subprocess, warned on the unused re-export.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When an external exits normally, spawn waits for its stdout and stderr
drains to reach EOF. A grandchild it left running (`sh -c 'sleep 300 &
echo hi'`) holds those pipes, so the wait lasts as long as the
grandchild. That wait matches bash's `$(...)` while nothing is
cancelled, but the drains watched only their own stop token, not the
command's cancel token: a Ctrl-C, `Kernel::cancel`, or `kill %N` that
arrived after the child exited could not end it.

The normal-exit arm now races the drains against the cancel token. On
a cancel it SIGTERMs what is left of the command's process group,
SIGKILLs it after the kill grace, stops the drains between chunks, and
returns 130, since the output was cut short by the cancel. Nothing
changes while no cancel arrives: no grace timeout on the normal path.

`orphaned_pipe_cancel_tests.rs`:
- A foreground `sh -c 'sleep 60 & echo hi'` cancelled after 500 ms
  returns 130 within 8 s and the grandchild is gone. Before: the
  execute was still waiting at 8 s.
- The same script as a job ended by `kill %1` returns promptly and the
  grandchild dies. This passed before too: `kill %N` also signals the
  job's process groups, which kills the grandchild and closes the pipe.
- A grandchild started with `setsid` leaves the group, so no group
  signal reaches it; `kill %1` must still end the job with 130. Before:
  still waiting at 8 s. The grandchild survives, and the test kills it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A whole-program job writes each statement's stderr through a separate
writer task fed by a channel. When the job task panicked,
`send_job_result` wrote the diagnostic and closed the streams at once,
without waiting for that writer. Stderr still queued for statements
that had run was dropped, since writes to a closed stream are ignored,
or landed after the diagnostic.

The writer now holds a oneshot sender that it drops when it ends. The
panic drops the job task's channel sender, so the writer drains what
was queued and ends, and `send_job_result` waits for that before the
diagnostic and the close. A `cmd &` job writes its stream directly and
has no writer to wait on.

`stderr_before_a_panic_precedes_the_diagnostic` runs `echo ... >&2`
then the panicking tool, as a whole-program job and as a `cmd &` job,
50 rounds each on a multi-thread runtime, and requires the stream to be
the earlier line followed by the diagnostic. The whole-program case
failed in round 0 before this change: the stream held only the
diagnostic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…b-stream

# Conflicts:
#	crates/kaish-kernel/src/kernel.rs
When a cancel ends the wait on pipes a grandchild holds open, the
command itself is already reaped. signal_leftover_group sent SIGTERM,
then a detached task slept kill_grace (2s by default, longer when an
embedder raises it) and sent SIGKILL to the same group id. Nothing
pinned that id across the sleep: a group id stays reserved only while a
member lives, so once the grandchildren died of the SIGTERM, a new
process could take the id and receive the delayed SIGKILL. wait_or_kill
does not have this window because its child is still unreaped during
the grace.

The leftover members are orphans holding our pipe after a cancel;
they have no claim to a graceful shutdown. SIGKILL goes to the group
immediately, while its members still hold the id, and the detached
task is gone.

Found by a kaibo review (deepseek).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tobert

tobert commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

kaibo reviews (cast deepseek, deepseek-flash explorer + synth), four rounds, plus a mutation check. All findings below are fixed on the branch unless marked otherwise.

Mutation check. At first the corpus passed 18 of 19 cases on the pre-fix tree, and disabling the statement-level publish failed only 1 case. That meant the net could not prove "never dropped". After hardening, each of the 13 remaining publish points fails at least one test when it is disabled. Nine points were redundant and were removed. The single published flag was replaced by stderr_published_len, because one flag could not represent a partly published err.

Round 1 (job-8), against the hardened branch:

  • No new assertion is reachable from a script.
  • Fixed: a killed external could lose its last chunk (the drain now stops between chunks).
  • Fixed: a failing condition command printed 2-3 copies.
  • Fixed: a panic in drain_lossy was reachable through the public fork_for_background + execute.
  • Fixed: a panicked job task left its streams open.
  • Fixed: cmd > f 2>&1 sent stderr to the terminal. This was already broken on main; it is fixed in its own commit.

Round 2 (job-10):

  • Clean: all nine redirect forms match bash, and the foreground path neither doubles nor loses output.
  • Fixed: a cancel could not end a wait on a pipe held open by a grandchild. This was pre-existing on main.
  • Fixed: the panic path could close streams before queued stderr was written.

Round 3 (job-11):

  • Clean: the drain race and the stderr-writer handshake.
  • Fixed in ad16874: a delayed SIGKILL after kill_grace could reach a reused pgid, because the command was already reaped. The group is now SIGKILLed immediately.

Noted, not changed:

  • 130 relabels a child that exited 0 when only the drain was cut short.
  • A pipeline's last stage hides that 130 unless pipefail is set.
  • kaish applies redirects after the command runs, so f() { echo a > out; }; f > out leaves out empty. This is the same on main.

🤖

@tobert
tobert merged commit e881a86 into main Sep 22, 2026
3 checks passed
tobert added a commit that referenced this pull request Sep 22, 2026
…rand (#467)

Merging #465 and then #466 broke main's build. #465 gave
drain_stderr_into a ctx parameter, and #466 added three call sites that
use the old signature. Each PR was green against a main that did not
have the other.

Passing ctx fixes the build, but it exposed two stderr bugs in
background jobs, where a fault under `!` interacts with #465's
publishing model:

```
f() { ! [[ 1 -eq abc ]]; }; f &              # the diagnostic reached the job stream twice
f() { ! [[ 1 -eq abc ]]; }; timeout 5 f &    # "<diag>\ntimeout: \n", but the foreground printed "timeout: <diag>"
```

The `!` arm ran its body through execute_stmt_flow, which publishes
stderr, and it used all of result.err as the error message. The arm now
does what the `&&`/`||` arms do with their left side. The body runs
through execute_stmt_flow_dispatch, which does not publish. A fault
takes only its unpublished stderr as the message, and every other result
is published before the drain.

New job-stderr corpus cases cover a negated builtin error, a negated
pipeline, a negated fault with and without earlier published stderr, and
a `timeout`-wrapped fault under `!` and `&&`. Going back to the
publishing call fails the timeout case, and using the whole err as the
message fails the published-body case. No test pins the drain after the
body; every enclosing context also drains per statement.

The fix does not cover a pre-existing ordering bug, which is pinned with
an ignored test. When a command's `$(…)` fails, result.err lists the
command's own stderr before the substitution's, while the job stream has
them in the order they ran. The `&&` arm has it too, so it gets its own
PR.

Reviewed with kaibo (DeepSeek); its timeout and ordering findings
reproduced, as above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant