Skip to content

fix(ateapi): tear a crashed actor's workload down before it is deleted - #256

Closed
teemow wants to merge 2 commits into
giantswarmfrom
fix/237-delete-actor-teardown-2
Closed

teemow wants to merge 2 commits into
giantswarmfrom
fix/237-delete-actor-teardown-2

Conversation

@teemow

@teemow teemow commented Oct 10, 2026

Copy link
Copy Markdown
Member

Problem

Seen on v1.7.0-rc.1 (gVisor, a kind cluster with the NFS CSI driver) with an ActorTemplate whose snapshot_config is on_pause: FULL, on_commit: DATA and no durable_dir volume, only two existing_volumes from one read-write-many NFS PersistentVolume: every SuspendActor failed on the node (no durable-dir volumes found for DATA snapshot), the actor crashed, and DeleteActor then removed the record while the sandbox kept running on the worker with its four nfs4 mounts under /var/lib/ate/actors/<uid>/volumes/ and its worker capacity; the PersistentVolume stayed Released after its claim was deleted, the NFS driver's DeleteVolume failing on directory not empty. Fixes #237.

Two defects: a crash released the worker before anything stopped the sandbox, so the delete found no worker to reach; and a template whose commit scope is DATA without a durable-dir volume was accepted though no checkpoint of its actors can succeed.

Proposed solution

One carried commit, upstream-shaped:

  • crashActor terminates the sandbox on atelet, detaches the actor's volumes and only then releases the worker. When a step fails the actor is still CRASHED but keeps its worker assignment, so a later DeleteActor or RevertActor retries the teardown against the same worker.
  • DeleteActor refuses with FailedPrecondition while the sandbox cannot be stopped or a volume cannot be released, keeping the assignment for the retry: a record never goes while its workload may still be placed.
  • CreateActorTemplate refuses a DATA on_pause or on_commit scope unless a container mounts a durable-dir volume, InvalidArgument naming actor_template.snapshot_config.on_pause / on_commit (the node snapshots the mounted directory, so a declared but unmounted volume is not enough).
  • The existing-volumes e2e suite gains DeleteAfterFailedSuspend: the actor's checkpoint directory on the node is made a file so the checkpoint fails before touching the sandbox; the crash leaves no directory, mount or process of the actor on the node, the worker's allocation drops, the crashed record deletes and the volume's PersistentVolume deletes with its claim.

Upstream check (agent-substrate/substrate): issue agent-substrate#1936 reports the DATA-scope-without-durable-dir acceptance and is being fixed by PR agent-substrate#1972 (open, a newer preferred_fidelity API); PR agent-substrate#1953 (open) moves the crash's teardown ahead of the CRASHED transition and keeps the assignment when it fails, which is the shape adapted here. agent-substrate#2355, agent-substrate#1665 and agent-substrate#641 are related; agent-substrate#2038 is a closed duplicate of agent-substrate#1665. The patch falls away at the re-pin onto a release that carries agent-substrate#1953 and agent-substrate#1972.

Acceptance criteria

  • After a failed SuspendActor, DeleteActor leaves no sandbox and no volume mount of the actor on the node, or is refused
  • A PersistentVolume an actor mounted as an existing volume is deletable once the actor is deleted
  • CreateActorTemplate with a DATA commit scope and no durable-dir volume is refused at create
  • e2e: the existing-volumes suite covers delete after a failed suspend
  • Unit tests: TestCrashActor_TerminatesAndDetachesBeforeRelease, TestCrashActorReleaseFailureKeepsWorkerAssignment, TestDeleteActor_TeardownFailureRefusesWithFailedPrecondition, TestDeleteActor_CrashedWithPlacedWorkload, TestValidateCreateActorTemplateRequest_DataScopeNeedsMountedDurableDir
  • The suite green on a kind cluster built from this branch, and again on the release candidate

This pull request was written by an agent.

A failed checkpoint crashed the actor and released its worker at once,
with the sandbox still running on it: DeleteActor then found no worker
to reach, removed the record, and the sandbox, its existing-volume
mounts and the worker's capacity leaked until the worker pod restarted.
A template whose on_pause or on_commit scope was DATA without any
durable-dir volume was accepted, though no checkpoint of its actors
could succeed on the node.

crashActor now terminates the sandbox on atelet, detaches the actor's
volumes and only then releases the worker; when a step fails the actor
is still CRASHED but keeps its worker assignment, so a later
DeleteActor or RevertActor retries the teardown against the same
worker. DeleteActor refuses with FailedPrecondition while the sandbox
cannot be stopped or a volume cannot be released, and keeps the
assignment for the retry: a record never goes while its workload may
still be placed. CreateActorTemplate refuses a DATA on_pause or
on_commit scope unless a container mounts a durable-dir volume, with
InvalidArgument naming the field: a DATA snapshot captures those
volumes and nothing else, so without one every pause or suspend of
the template's actors fails on the node and crashes them.

The existing-volumes e2e suite covers the delete after a failed
suspend: the checkpoint is made to fail on the node, the crash leaves
no directory, mount or process of the actor there, the worker's
capacity is freed, the crashed record deletes and the volume's
PersistentVolume deletes with its claim.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
@teemow

teemow commented Oct 10, 2026

Copy link
Copy Markdown
Member Author

Replaced by #258: the same carried commit with the e2e subtest reworked to pass on both sandbox classes (an immutable file at the checkpoint-state path, which micro-VM's checkpoint cannot clear and gVisor's cannot create; the test now also proves the refusal branch). This comment was written by an agent.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: DeleteActor after a failed suspend leaks the sandbox and its mounts

1 participant