Summary
With ARGO_POD_STATUS_CAPTURE_FINALIZER=true, an aged, successfully completed Pod can lose workflows.argoproj.io/status before its result is persisted to the Workflow. If Pod cleanup runs while Workflow reconciliation is delayed, the Pod disappears and subsequent reconciliation marks the node Error: pod deleted.
This was reproduced using a single public BusyBox container and an isolated stock controller. No workflow compiler, DAG, GPU, Kueue labels, artifacts, application callbacks, retries, or workflow PodGC policy are needed.
Version and test scope
- Runtime tested: workflow-controller and argoexec v4.1.3, controller git commit
5fdad0fe6f6d740c55d8289f26c912ab711ca407.
- Kubernetes: EKS 1.36, worker kubelet 1.36.2.
- Native Pod status-capture finalizer enabled before creating the Pod.
- One isolated, namespaced controller instance, with its own instance ID; two Workflow workers normally.
- Deterministic fault injection temporarily sets
--workflow-workers=0, leaving the independent Pod cleanup workers active. This is not a recommended operating configuration or a claim about failure frequency under normal load.
I have not run the reproduction with the mutable :latest image or v4.1.4. The controlled experiment was pinned to the installed v4.1.3 release to vary one factor at a time and avoid an unrelated controller upgrade. I separately inspected the tagged v4.1.4 source and current main: commonPodEvent has the same timed-removal logic. That is source evidence, not a runtime claim for those builds.
Minimal workflow
Use a dedicated test namespace, an isolated controller managing only that namespace, and a workflow-runner ServiceAccount with the standard WorkflowTaskResult create/patch permissions. The label below must match that controller's configured instance ID.
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
name: status-capture-repro
namespace: status-capture-test
labels:
workflows.argoproj.io/controller-instanceid: status-capture-test
spec:
serviceAccountName: workflow-runner
entrypoint: main
templates:
- name: main
container:
image: busybox:1.37
command: [sh, -c]
args: ['sleep 30; echo capture-ok']
resources:
requests:
cpu: 10m
memory: 16Mi
limits:
memory: 32Mi
Names are normalized here. The executed test additionally pinned this CPU Pod to an existing node by hostname to avoid provisioning or node disruption during the experiment; that environment-specific selector is omitted above.
Reproduction sequence
Only use a disposable, isolated controller for these steps.
- Start that controller with
ARGO_POD_STATUS_CAPTURE_FINALIZER=true and --workflow-workers=2. Submit the workflow and wait until the persisted Workflow node phase is Running.
- Scale only the isolated controller to zero before the sleep finishes. Do not stop the executor Pod.
- Wait until the Pod is
Succeeded, verify both main and wait terminated with exit code 0, and record the original Pod UID. The Workflow still records Running.
- Wait until at least 190 seconds have elapsed since the latest Pod condition
lastTransitionTime.
- Request deletion of that completed Pod. Verify the same UID remains with a deletion timestamp and
workflows.argoproj.io/status while the controller is stopped. Do not remove any finalizer manually.
- Restart the isolated controller with
--workflow-workers=0. Pod cleanup remains active. Observe that the native finalizer is removed and the Pod disappears while the persisted Workflow node is still Running.
- Stop the isolated controller, restore
--workflow-workers=2, and start it again.
- Observe Workflow/node phase
Error and message pod deleted.
No Workflow status field was patched to manufacture this outcome. A control with normal Workflow workers after the same aging/deletion sequence succeeded, which is consistent with reconciliation winning the race in that run.
Expected and actual behavior
Expected: while the owning Workflow exists and is active, retain a terminal Pod's status-capture finalizer until its result is persisted, even if reconciliation is delayed. Cancellation, Workflow deletion and orphan cleanup still need to work.
Actual: elapsed time independently authorizes removal before Workflow persistence; a successful execution is subsequently reported as an error.
Controller evidence
The following is an excerpt from the isolated controller, with resource names omitted:
16:31:11.132 Current Worker Numbers ... workflowWorkers=0 ... podCleanup=4
16:31:11.257 Removing finalizers during a delete ... pod.Finalizers=[workflows.argoproj.io/status]
16:31:11.257 queuing pod delay ... delay=-1m24.257168224s action=removeFinalizer
16:31:11.350 cleaning up pod ... action=removeFinalizer
16:31:11.363 delete pod event
At this point the Pod was absent and the Workflow node remained Running. After restoring Workflow workers, its recorded terminal state was:
phase: Error
message: pod deleted
taskResultSynced: true
Executor completion evidence
The test did not retain the wait container's stdout; the retained Kubernetes container statuses show:
main:
exitCode: 0
reason: Completed
startedAt: '2026-09-19T16:27:14Z'
finishedAt: '2026-09-19T16:27:44Z'
wait:
exitCode: 0
reason: Completed
startedAt: '2026-09-19T16:27:14Z'
finishedAt: '2026-09-19T16:27:44Z'
Latest Pod condition transition: 2026-09-19T16:27:47Z. Deletion was requested at approximately 16:31:04Z. Both containers had completed before deletion.
Observed images:
docker.io/library/busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0
quay.io/argoproj/argoexec@sha256:6aa5cf3452b12c8498eccc62455e460e94e6d092931932aa4dbda6bb1512036c
Source analysis
commonPodEvent in v4.1.3 queues removeFinalizer when a Pod has a deletion timestamp. The deadline is based on the latest Pod condition transition plus the greater of PodGC delay and two minutes. For an aged Pod the remaining delay is nonpositive, so cleanup becomes immediately eligible, independently of Workflow persistence.
By contrast, persistUpdates invokes normal Pod cleanup after updating the Workflow, and queuePodsForCleanup checks that node state is fulfilled with task results synchronized.
Is timed removal during external deletion intended to override status-capture protection for a terminal Pod with an active owner, or should this path retain protection until reconciliation records the result?
Related issues checked
I did not find an open report specifically reproducing timed finalizer removal with the flag enabled and delayed Workflow workers. This report concerns that remaining path; it does not claim to establish a new regression version or a measured production failure rate.
Summary
With
ARGO_POD_STATUS_CAPTURE_FINALIZER=true, an aged, successfully completed Pod can loseworkflows.argoproj.io/statusbefore its result is persisted to the Workflow. If Pod cleanup runs while Workflow reconciliation is delayed, the Pod disappears and subsequent reconciliation marks the nodeError: pod deleted.This was reproduced using a single public BusyBox container and an isolated stock controller. No workflow compiler, DAG, GPU, Kueue labels, artifacts, application callbacks, retries, or workflow PodGC policy are needed.
Version and test scope
5fdad0fe6f6d740c55d8289f26c912ab711ca407.--workflow-workers=0, leaving the independent Pod cleanup workers active. This is not a recommended operating configuration or a claim about failure frequency under normal load.I have not run the reproduction with the mutable
:latestimage or v4.1.4. The controlled experiment was pinned to the installed v4.1.3 release to vary one factor at a time and avoid an unrelated controller upgrade. I separately inspected the tagged v4.1.4 source and current main:commonPodEventhas the same timed-removal logic. That is source evidence, not a runtime claim for those builds.Minimal workflow
Use a dedicated test namespace, an isolated controller managing only that namespace, and a
workflow-runnerServiceAccount with the standard WorkflowTaskResult create/patch permissions. The label below must match that controller's configured instance ID.Names are normalized here. The executed test additionally pinned this CPU Pod to an existing node by hostname to avoid provisioning or node disruption during the experiment; that environment-specific selector is omitted above.
Reproduction sequence
Only use a disposable, isolated controller for these steps.
ARGO_POD_STATUS_CAPTURE_FINALIZER=trueand--workflow-workers=2. Submit the workflow and wait until the persisted Workflow node phase isRunning.Succeeded, verify bothmainandwaitterminated with exit code 0, and record the original Pod UID. The Workflow still recordsRunning.lastTransitionTime.workflows.argoproj.io/statuswhile the controller is stopped. Do not remove any finalizer manually.--workflow-workers=0. Pod cleanup remains active. Observe that the native finalizer is removed and the Pod disappears while the persisted Workflow node is stillRunning.--workflow-workers=2, and start it again.Errorand messagepod deleted.No Workflow status field was patched to manufacture this outcome. A control with normal Workflow workers after the same aging/deletion sequence succeeded, which is consistent with reconciliation winning the race in that run.
Expected and actual behavior
Expected: while the owning Workflow exists and is active, retain a terminal Pod's status-capture finalizer until its result is persisted, even if reconciliation is delayed. Cancellation, Workflow deletion and orphan cleanup still need to work.
Actual: elapsed time independently authorizes removal before Workflow persistence; a successful execution is subsequently reported as an error.
Controller evidence
The following is an excerpt from the isolated controller, with resource names omitted:
At this point the Pod was absent and the Workflow node remained
Running. After restoring Workflow workers, its recorded terminal state was:Executor completion evidence
The test did not retain the wait container's stdout; the retained Kubernetes container statuses show:
Latest Pod condition transition:
2026-09-19T16:27:47Z. Deletion was requested at approximately16:31:04Z. Both containers had completed before deletion.Observed images:
docker.io/library/busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0quay.io/argoproj/argoexec@sha256:6aa5cf3452b12c8498eccc62455e460e94e6d092931932aa4dbda6bb1512036cSource analysis
commonPodEventin v4.1.3 queuesremoveFinalizerwhen a Pod has a deletion timestamp. The deadline is based on the latest Pod condition transition plus the greater of PodGC delay and two minutes. For an aged Pod the remaining delay is nonpositive, so cleanup becomes immediately eligible, independently of Workflow persistence.By contrast,
persistUpdatesinvokes normal Pod cleanup after updating the Workflow, andqueuePodsForCleanupchecks that node state is fulfilled with task results synchronized.Is timed removal during external deletion intended to override status-capture protection for a terminal Pod with an active owner, or should this path retain protection until reconciliation records the result?
Related issues checked
pod deleted#8783 introduced the status-capture finalizer direction and is closed.Runningwhen Karpenterpod deleted#13152 is closed as a duplicate of earlier Pod-loss work.I did not find an open report specifically reproducing timed finalizer removal with the flag enabled and delayed Workflow workers. This report concerns that remaining path; it does not claim to establish a new regression version or a measured production failure rate.