Never delete an interrupted-apply module without attempting a destroy - #151
Conversation
…stroy Stack destroy silently orphaned cloud resources on one specific path. BUG-008 recovery handles modules stuck in a transitional status with no live task; for `applying` with NO terminal task on record -- the worker died mid-run -- it reset the module to not_initialized. That status routed it into modules_to_delete, which deleted the row with no destroy attempt at all. But `tofu apply` creates IAM roles, VPCs and the like well before it finishes. A module that reached `applying` may already own resources, and Forge holds the only record of them. Deleting the row made them invisible: they kept existing, and the next deploy hit EntityAlreadyExists (#30). Two changes: - Recovery lands an interrupted apply in apply_failed (and an interrupted destroy in destroy_failed), so the module is queued for a real destroy. Destroy is idempotent if nothing was created; the row-delete is not recoverable if something was. initializing/planning still recover to not_initialized -- nothing is applied during those phases, so the fast delete path stays. PLANNING + a FAILED task now recovers to plan_failed rather than collapsing into init_failed. - The delete-vs-destroy classification now uses NO_INFRA_STATUSES -- the same vocabulary the project-DELETE guard adopted in #129 -- instead of a hand-rolled list, and fails CLOSED: anything not positively known to have no infrastructure gets a destroy attempt rather than a row delete. This file already imported NO_INFRA_STATUSES for _execute_stack_destroy; the two sites had drifted apart in the dangerous direction. The tests fail against the unpatched service and pass with it. Fixes #30 Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4
mwiget
left a comment
There was a problem hiding this comment.
Code Review Summary
Goal: Fixes issue #30 by ensuring modules whose apply was interrupted by worker death are transitioned to apply_failed and queued for a tofu destroy attempt rather than being silently deleted from the DB.
Key Technical Observations:
- Resource Leak Fix: Prevents orphaned cloud infrastructure (e.g., AWS IAM roles, VPCs) from surviving stack deletions when a worker dies mid-apply.
- Fail-Closed Classification: Stack destruction uses
NO_INFRA_STATUSESto fail closed; any status not explicitly known to have zero infrastructure receives a destroy task. - Fast-Path Preservation: Modules in
initializingorplanningstill skip destroy and are deleted directly since no cloud resources exist. - Test Coverage: 3 component tests in
test_stack_deployment_service.pyasserting non-orphaning of interrupted applies and destroys.
Verdict: LGTM! Approved.
Self-reviewNo defects found. Two things surfaced that the PR body didn't say, one of them a genuine second fix hiding in the change. 🟢 Found — this also closes a second orphan path I hadn't noticedEnumerated how every The interesting one is bare ( 🟢 Verified — composes correctly with #145's janitorSince #145 merged, the janitor may reset an interrupted apply to Verified non-vacuous, and targeted at the right branchThe orphan test creates zero Task rows (checked: 0 occurrences of The one thing I could not demonstrateThe PR's argument rests on "destroy is idempotent if nothing was created" — i.e. Checked and found correct
Regression sweep unchanged: 530 passed, 2 pre-existing Scope, restatedThis fixes a real, reproducible orphan path — two of them, it turns out — that produce exactly #30's symptom. It is still not proof that this was the reporter's cause; the PR body's diagnostic stands (if |
Description
Stack destroy silently orphaned cloud resources on one specific path. BUG-008 recovery handles modules stuck in a transitional status with no live task; for
applyingwith no terminal task on record — the worker died mid-run — it reset the module tonot_initialized. That status routed it intomodules_to_delete, which deleted the row with no destroy attempt at all.But
tofu applycreates IAM roles, VPCs and the like well before it finishes. A module that reachedapplyingmay already own resources, and Forge holds the only record of them. Deleting the row made them invisible: they kept existing, and the next deploy hitEntityAlreadyExists.Fixes / Implements: #30
Honesty about scope
#30 is one line long, and I could not reproduce its stated symptom ("stack-delete tears down K8s but not AWS-provider modules") — there is no provider conditional anywhere in destroy selection. What I found instead is a concrete, testable path that produces exactly the outcome the issue describes: orphaned AWS/IAM resources +
EntityAlreadyExistson redeploy. It's the candidate I flagged on the issue as "the one to look at first."I'm confident this path is real and worth fixing on its own merits. I'm not certain it's the only cause behind the original report — if the reporter's case was a
tofu destroythat partially failed but was still marked destroyed, this PR doesn't touch that. Worth keeping in mind before treating #30 as fully closed.The two changes
Recovery lands an interrupted apply in
apply_failed(and an interrupted destroy indestroy_failed), so the module is queued for a real destroy. Destroy is idempotent if nothing was created; the row-delete is unrecoverable if something was.initializing/planningstill recover tonot_initialized— nothing is applied during those phases, so the fast delete path stays. AlsoPLANNING+FAILEDtask now recovers toplan_failedrather than collapsing intoinit_failed.The delete-vs-destroy classification uses
NO_INFRA_STATUSES— the same vocabulary the project-DELETE guard adopted in Refuse project deletion while modules still own cloud resources #129 — instead of a hand-rolled list, and fails closed: anything not positively known to have no infrastructure gets a destroy attempt rather than a row delete. (Self-review found this closes a second orphan path too: the genericfailedstatus fell into the oldelseand was deleted untried.) This file already importedNO_INFRA_STATUSESfor_execute_stack_destroy; the two sites had drifted apart in the dangerous direction.Architectural Decision Record (ADR)
Type of Change
Behaviour change to be aware of: a stack destroy that previously "succeeded instantly" on an interrupted-apply module (by deleting it) will now actually run a destroy task for it. If that destroy fails — e.g. the workspace is gone — the stack destroy reports the failure instead of quietly succeeding. That's the fix working, but it will look like a new failure to anyone who had this happen before.
Verification & Testing
Three tests added to
backend/tests/component/test_stack_deployment_service.py(95 passed in file):test_destroy_does_not_orphan_an_interrupted_apply—applying, no Task row: recovered toapply_failed, queued for destroy, row still present;destroy_failed, queued;initializing/planninginterrupted → still deleted directly, no destroy dispatched (the fast path is preserved).Verified non-vacuous: both orphan tests fail against the unpatched service and pass with it; the init/plan control passes both ways. I also wrote and then deleted a fourth test that only asserted "row is gone" — that was true before the fix too, so it proved nothing.
Regression sweep on
stack/destroy/project_service/janitor: 530 passed, 2 failed — both pre-existingtest_cli_smoke.pyfailures from the known baseline.ruffclean.Environment validation needed
Yes — the interesting case needs a real cloud account. One thing to confirm first, because the whole argument rests on it and there is no
tofubinary in any local image to check:tofu destroyagainst a workspace with no state must exit 0 ("No objects need to be destroyed"). That is documented behaviour and CP-008 already relies on it, but if it were ever false, an interrupted-apply module would make the stack destroy fail where it previously succeeded by deleting the row. Then, against AWS:docker killthe celery worker oncetofuhas begun creating resources, then destroy the stack — confirm the module goes toapply_failed, a destroy task runs, and the IAM resources are actually removed;EntityAlreadyExists(this is the reported symptom);not_initialized/plannedmodules still destroys instantly with no destroy tasks.If step 2 still throws
EntityAlreadyExistsafter this, that would tell us the reporter's case is a different path (likely partialtofu destroyfailure), and #30 should be reopened with that detail.Checklist
applyingmay own resources, since "no terminal task → not_initialized" reads as reasonable until you know that.