Problem
Some resources cannot be deleted until something else is done to them first, and that something is not expressible as a single statement in the resource's .iql file. The canonical case is an S3 bucket: DeleteBucket fails with 409 BucketNotEmpty (surfaced by stackql as no response body for operation = DeleteBucket) while the bucket holds objects, and emptying a bucket needs an object listing loop (and a version listing loop for versioned buckets). A Databricks workspace root bucket is never empty at teardown time because the control plane writes DBFS and log data into it.
Today the only options are:
skip_on_delete: true on the bucket, which retains it (fine when the data should be kept, wrong when it should not),
--on-failure ignore, which finishes the teardown but leaves the bucket behind, or
- an out-of-band
aws s3 rm --recursive before running teardown.
Other resources with the same shape: ECR repositories with images, Cognito user pools with a deletion protection flag, RDS instances with deletion protection, Key Vaults with soft delete, GitHub repositories with branch protection rules, or any resource where a dependent that is not managed by the stack has to be cleared first.
Proposal
A predelete hook that teardown runs for a resource after the exists check reports the resource present and before its delete statement. Two forms:
- A
/*+ predelete */ anchor in the resource's .iql file for cases that are a single statement (for example flipping a deletion protection property with an UPDATE, or a DELETE on a dependent resource that the provider can address in one call). Rendered with the same context as delete, including this.* fields captured by the exists check.
- A
predelete script in the manifest for cases that need a loop or an external tool, using the same contract as type: script resources (run with sh -c, exports optional):
- name: aws/s3/root_bucket
file: aws/s3/bucket.iql
predelete:
run: aws s3 rm s3://{{ root_bucket_name }} --recursive
props:
- name: bucket_name
value: "{{ root_bucket_name }}"
If both are present the anchor runs first, then the script.
Behaviour to pin down
- Ordering. exists check, then
predelete, then delete, then callback:delete, then the post-delete check (per postdelete_retries / postdelete_retry_delay). When the delete is re-issued (the delete anchor's retries), predelete should run again before each attempt, since the reason for the retry may be that more objects appeared.
- Failure. A failing
predelete is a failed delete: abort under --on-failure error, log and report the resource as not confirmed deleted under --on-failure ignore. Fatal errors abort in both modes.
- Dry run. Render and log the anchor / script without executing, like
delete.
- Not run when the resource does not exist, when the resource has
skip_on_delete: true, or when the delete is skipped because a query references an export that could not be collected.
- Not run by
build or test.
- Idempotency is the author's responsibility, as with
script resources; document that a predelete may run more than once per teardown.
- The anchor form should support the usual options (
retries, retry_delay).
Related
Problem
Some resources cannot be deleted until something else is done to them first, and that something is not expressible as a single statement in the resource's
.iqlfile. The canonical case is an S3 bucket:DeleteBucketfails with409 BucketNotEmpty(surfaced by stackql asno response body for operation = DeleteBucket) while the bucket holds objects, and emptying a bucket needs an object listing loop (and a version listing loop for versioned buckets). A Databricks workspace root bucket is never empty at teardown time because the control plane writes DBFS and log data into it.Today the only options are:
skip_on_delete: trueon the bucket, which retains it (fine when the data should be kept, wrong when it should not),--on-failure ignore, which finishes the teardown but leaves the bucket behind, oraws s3 rm --recursivebefore runningteardown.Other resources with the same shape: ECR repositories with images, Cognito user pools with a deletion protection flag, RDS instances with deletion protection, Key Vaults with soft delete, GitHub repositories with branch protection rules, or any resource where a dependent that is not managed by the stack has to be cleared first.
Proposal
A
predeletehook thatteardownruns for a resource after the exists check reports the resource present and before itsdeletestatement. Two forms:/*+ predelete */anchor in the resource's.iqlfile for cases that are a single statement (for example flipping a deletion protection property with anUPDATE, or aDELETEon a dependent resource that the provider can address in one call). Rendered with the same context asdelete, includingthis.*fields captured by the exists check.predeletescript in the manifest for cases that need a loop or an external tool, using the same contract astype: scriptresources (run withsh -c, exports optional):If both are present the anchor runs first, then the script.
Behaviour to pin down
predelete, thendelete, thencallback:delete, then the post-delete check (perpostdelete_retries/postdelete_retry_delay). When the delete is re-issued (thedeleteanchor'sretries),predeleteshould run again before each attempt, since the reason for the retry may be that more objects appeared.predeleteis a failed delete: abort under--on-failure error, log and report the resource as not confirmed deleted under--on-failure ignore. Fatal errors abort in both modes.delete.skip_on_delete: true, or when the delete is skipped because a query references an export that could not be collected.buildortest.scriptresources; document that apredeletemay run more than once per teardown.retries,retry_delay).Related
skip_on_deleteproperty forqueryresources #56 / skip_on_delete, teardown hardening, --on-failure ignore, integration tests (2.2.0) #57:skip_on_delete,--on-failure ignoreand the end-of-run summary of unconfirmed deletes, which are the current workarounds.tests/live_stacks/aws_ssm_onfailurelive stack already has an S3DeleteBucketfailure case; apredeletethat empties a bucket could be exercised there, or in a small dedicated S3 stack (an empty standard bucket is free).