Repository navigation
Conversation
A failed checkpoint crashed the actor and released its worker at once, with the sandbox still running on it: DeleteActor then found no worker to reach, removed the record, and the sandbox, its existing-volume mounts and the worker's capacity leaked until the worker pod restarted. A template whose on_pause or on_commit scope was DATA without any durable-dir volume was accepted, though no checkpoint of its actors could succeed on the node. crashActor now terminates the sandbox on atelet, detaches the actor's volumes and only then releases the worker; when a step fails the actor is still CRASHED but keeps its worker assignment, so a later DeleteActor or RevertActor retries the teardown against the same worker. DeleteActor refuses with FailedPrecondition while the sandbox cannot be stopped or a volume cannot be released, and keeps the assignment for the retry: a record never goes while its workload may still be placed. CreateActorTemplate refuses a DATA on_pause or on_commit scope unless a container mounts a durable-dir volume, with InvalidArgument naming the field: a DATA snapshot captures those volumes and nothing else, so without one every pause or suspend of the template's actors fails on the node and crashes them. The existing-volumes e2e suite covers the delete after a failed suspend: the checkpoint is made to fail on the node, the crash leaves no directory, mount or process of the actor there, the worker's capacity is freed, the crashed record deletes and the volume's PersistentVolume deletes with its claim. Signed-off-by: Timo Derstappen <teemow@gmail.com>
6 tasks
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Member
Author
|
Replaced by #258: the same carried commit with the e2e subtest reworked to pass on both sandbox classes (an immutable file at the checkpoint-state path, which micro-VM's checkpoint cannot clear and gVisor's cannot create; the test now also proves the refusal branch). This comment was written by an agent. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Seen on v1.7.0-rc.1 (gVisor, a kind cluster with the NFS CSI driver) with an ActorTemplate whose
snapshot_configison_pause: FULL,on_commit: DATAand nodurable_dirvolume, only twoexisting_volumes from one read-write-many NFS PersistentVolume: everySuspendActorfailed on the node (no durable-dir volumes found for DATA snapshot), the actor crashed, andDeleteActorthen removed the record while the sandbox kept running on the worker with its four nfs4 mounts under/var/lib/ate/actors/<uid>/volumes/and its worker capacity; the PersistentVolume stayedReleasedafter its claim was deleted, the NFS driver's DeleteVolume failing ondirectory not empty. Fixes #237.Two defects: a crash released the worker before anything stopped the sandbox, so the delete found no worker to reach; and a template whose commit scope is DATA without a durable-dir volume was accepted though no checkpoint of its actors can succeed.
Proposed solution
One carried commit, upstream-shaped:
crashActorterminates the sandbox on atelet, detaches the actor's volumes and only then releases the worker. When a step fails the actor is stillCRASHEDbut keeps its worker assignment, so a laterDeleteActororRevertActorretries the teardown against the same worker.DeleteActorrefuses withFailedPreconditionwhile the sandbox cannot be stopped or a volume cannot be released, keeping the assignment for the retry: a record never goes while its workload may still be placed.CreateActorTemplaterefuses a DATAon_pauseoron_commitscope unless a container mounts a durable-dir volume,InvalidArgumentnamingactor_template.snapshot_config.on_pause/on_commit(the node snapshots the mounted directory, so a declared but unmounted volume is not enough).DeleteAfterFailedSuspend: the actor's checkpoint directory on the node is made a file so the checkpoint fails before touching the sandbox; the crash leaves no directory, mount or process of the actor on the node, the worker's allocation drops, the crashed record deletes and the volume's PersistentVolume deletes with its claim.Upstream check (agent-substrate/substrate): issue agent-substrate#1936 reports the DATA-scope-without-durable-dir acceptance and is being fixed by PR agent-substrate#1972 (open, a newer
preferred_fidelityAPI); PR agent-substrate#1953 (open) moves the crash's teardown ahead of theCRASHEDtransition and keeps the assignment when it fails, which is the shape adapted here. agent-substrate#2355, agent-substrate#1665 and agent-substrate#641 are related; agent-substrate#2038 is a closed duplicate of agent-substrate#1665. The patch falls away at the re-pin onto a release that carries agent-substrate#1953 and agent-substrate#1972.Acceptance criteria
SuspendActor,DeleteActorleaves no sandbox and no volume mount of the actor on the node, or is refusedCreateActorTemplatewith a DATA commit scope and no durable-dir volume is refused at createTestCrashActor_TerminatesAndDetachesBeforeRelease,TestCrashActorReleaseFailureKeepsWorkerAssignment,TestDeleteActor_TeardownFailureRefusesWithFailedPrecondition,TestDeleteActor_CrashedWithPlacedWorkload,TestValidateCreateActorTemplateRequest_DataScopeNeedsMountedDurableDirThis pull request was written by an agent.