Skip to content

Reset modules stuck in a transient state when their worker dies - #145

Merged
jgruberf5 merged 1 commit into
stagingfrom
fix/6-stuck-transient-module-status
Aug 18, 2026
Merged

jgruberf5 merged 1 commit into
stagingfrom
fix/6-stuck-transient-module-status

Conversation

@jgruberf5

Copy link
Copy Markdown
Collaborator

Description

reset_stale_tasks flipped the Task row to failed for every stale task, but only recovered the owning module for DESTROY tasks. The code said so outright:

# Only destroy tasks need the module/chain recovery below; deploy tasks
# keep their existing reset behaviour (Task row flipped, nothing else).

So an apply/plan/init whose worker died between the route setting applying and the task's error handler left the module transient forever — the UI showed a perpetually-applying module with no error and no Retry button (Retry only renders for *_failed states), recoverable only by a manual UPDATE project_modules SET status=..., exactly as the issue reports.

Deploy-side modules are now driven to the matching terminal state: initializing → init_failed, planning → plan_failed, applying → apply_failed.

Fixes / Implements: #6

Why this shape

The issue offers three options. This is option 2 (a background sweep), because the infrastructure already exists and already runs — execution_janitor computes the live Celery task set at boot and periodically, and is where the identical DESTROY recovery already lives. So "the worker is dead" means exactly what it already means elsewhere in the codebase, and no new schedule, table, or config is introduced.

Option 1 (a finally in each task function) does not actually cover the reported cause: a worker killed by SIGKILL/OOM never runs any Python handler, so a finally block never executes. It would help the narrower "unexpected exception type" case, and remains worth doing, but it cannot be the backstop on its own.

Two deliberate exclusions:

  • DESTROYING is not in the mapping. That path re-drives the destroy chain rather than doing a plain reset, because the next module in the reverse DAG depends on the transition. Adding it here would bypass _trigger_next_destroy_module.
  • A real deployment_error is preserved. The janitor's note only fills an empty field, so a genuine cause recorded before the worker died is never overwritten with a generic message.

Architectural Decision Record (ADR)

  • N/A (Bug fix or minor docs tweak).

Type of Change

  • Bug fix (non-breaking change fixing an issue)

Verification & Testing

Six cases added to backend/tests/component/test_execution_janitor.py (14 passed in that file):

  • parametrized over all three transient states → correct terminal state and the "Worker no longer alive" note;
  • a live task leaves the module alone — the important negative case, since resetting a running deploy out from under itself would be far worse than the bug;
  • a module already in a terminal state is not rewritten;
  • an existing deployment_error wins over the janitor's note.

Regression diff on the janitor / stale / lifecycle selection in bnk-forge-test:latest:

baseline (staging):   138 passed, 0 failed
this branch:          144 passed, 0 failed

Environment validation needed

Yes — the trigger is a worker death, which I can only simulate. My tests pass a dead celery_task_id rather than actually killing a worker. Worth reproducing the original scenario on a real deployment:

  1. start an apply, then docker kill the celery worker mid-run (SIGKILL, so no Python handler runs);
  2. wait for the periodic janitor pass — the module should move to apply_failed with the note, and the Retry button should appear;
  3. confirm a healthy long-running apply is never reset while it is still executing. This is the regression that would matter most, and it depends on get_live_task_ids() being accurate under real broker conditions — something no unit test here covers.

Checklist

  • My code follows the project's code style and formatting guidelines.
  • I have updated documentation where necessary — the _TRANSIENT_TO_FAILED mapping carries a comment explaining why DESTROYING is excluded, since that is the non-obvious part.
  • N/A — no backend route/schema changes.

reset_stale_tasks flipped the Task row to failed for every stale task but
only recovered the owning module for DESTROY tasks -- "deploy tasks keep
their existing reset behaviour (Task row flipped, nothing else)", as the
comment put it. So an apply/plan/init whose worker died between the route
setting `applying` and the task's error handler left the module transient
forever: the UI showed a perpetually-applying module with no error and no
Retry button (Retry only renders for *_failed states), recoverable only by
a manual UPDATE against the database.

Deploy-side modules are now driven to the matching terminal state --
initializing -> init_failed, planning -> plan_failed, applying ->
apply_failed -- reusing the janitor that already computes the live Celery
task set, so "dead worker" means exactly what it already means elsewhere.
This is option 2 from the issue; it needs no new schedule, since the
janitor already runs at boot and periodically.

DESTROYING is deliberately absent from the mapping: that path re-drives the
destroy chain rather than doing a plain reset, because the next module in
the reverse DAG depends on the transition.

A real deployment_error is preserved if one was recorded -- the janitor's
note only fills an empty field, so the actual cause is never overwritten by
a generic message. Modules already in a terminal state are left alone, and
a module whose task is still live is never reset out from under itself.

Fixes #6

Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4

@mwiget mwiget left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review Summary

Goal: Fixes issue #6 where killed Celery workers left modules stuck indefinitely in transient states (initializing, planning, applying).

Key Technical Observations:

  • Janitor Cleanup: Extends execution_janitor to transition modules with dead Celery workers from transient states to corresponding terminal failure states (init_failed, plan_failed, apply_failed), restoring the UI Retry button.
  • Safety Exclusions: Excludes DESTROYING state to prevent skipping DAG reverse destroy chain triggers. Preserves existing non-empty deployment_error messages.
  • Test Coverage: Added component tests in test_execution_janitor.py covering transient-to-failed mappings, live worker non-interference, and error message preservation.

Verdict: LGTM! Approved.

@jgruberf5

Copy link
Copy Markdown
Collaborator Author

Self-review

No defects found. Three things I checked that were not obviously safe:

Checked — the module change is always committed

reset_stale_executions only commits when tasks_reset or destroy_reset["stale_stack_ids"] is non-empty:

if tasks_reset or destroy_reset["stale_stack_ids"]:
    db.commit()

A module reset outside that condition would be silently discarded, and the bug would look unfixed intermittently. It cannot happen: a module only enters deploy_modules_to_reset inside the same loop iteration that appends to reset_ids, so tasks_reset is non-empty whenever any module was reset. The commit condition did not need widening.

Checked — DESTROYING really is excluded, not just undocumented

_TRANSIENT_TO_FAILED deliberately omits destroying, and the .get(...) is None → continue guard is what enforces it. This matters for a case that is easy to miss: a module sitting in destroying whose stale task is not task_type="destroy" lands in deploy_modules_to_reset, and the guard is the only thing stopping it from being reset out from under _trigger_next_destroy_module — which the destroy chain depends on. Verified the mapping keys are plain strings matching the column values ({'initializing': 'init_failed', 'planning': 'plan_failed', 'applying': 'apply_failed'}).

Checked — neighbouring comment still accurate

parallel_execution_service.py:513 documents a residual race with "The janitor does NOT recover it: reset_stale_tasks only touches NON-TERMINAL tasks, and this one is terminal." Still true — this PR does not change which tasks are selected, only what happens to their modules.

Deliberate consistency choice, flagged for the reviewer

The reset writes module.status directly rather than going through transition_module_status, so no module_state_transitions audit row is recorded. That matches the destroy recovery immediately below it (execution_janitor.py:163 does the same), so this PR is internally consistent — but it does mean a janitor-driven reset is invisible in the audit log.

Given #101 is specifically about that audit log being uninformative, routing janitor resets through the state machine with a reason looks like the better end state. I did not do it here because it changes the transition contract (lock/fence expectations) for a path that runs without a task, and I would rather not fold that into a bug fix. Worth a follow-up if you agree.

Unchanged

The environment-validation request stands, and point 3 is the one that matters: a healthy long-running apply must never be reset while still executing. That depends on get_live_task_ids() being accurate under real broker conditions, which no test here covers.

Regression diff: 144 passed vs 138 baseline, no new failures.

@jgruberf5
jgruberf5 merged commit c8e9934 into staging Aug 18, 2026
25 checks passed
@jgruberf5
jgruberf5 deleted the fix/6-stuck-transient-module-status branch August 18, 2026 15:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants