Skip to content

Status-capture finalizer can expire before successful Pod results are persisted during delayed reconciliation #17024

Description

@jalvarezz13

Summary

With ARGO_POD_STATUS_CAPTURE_FINALIZER=true, an aged, successfully completed Pod can lose workflows.argoproj.io/status before its result is persisted to the Workflow. If Pod cleanup runs while Workflow reconciliation is delayed, the Pod disappears and subsequent reconciliation marks the node Error: pod deleted.

This was reproduced using a single public BusyBox container and an isolated stock controller. No workflow compiler, DAG, GPU, Kueue labels, artifacts, application callbacks, retries, or workflow PodGC policy are needed.

Version and test scope

  • Runtime tested: workflow-controller and argoexec v4.1.3, controller git commit 5fdad0fe6f6d740c55d8289f26c912ab711ca407.
  • Kubernetes: EKS 1.36, worker kubelet 1.36.2.
  • Native Pod status-capture finalizer enabled before creating the Pod.
  • One isolated, namespaced controller instance, with its own instance ID; two Workflow workers normally.
  • Deterministic fault injection temporarily sets --workflow-workers=0, leaving the independent Pod cleanup workers active. This is not a recommended operating configuration or a claim about failure frequency under normal load.

I have not run the reproduction with the mutable :latest image or v4.1.4. The controlled experiment was pinned to the installed v4.1.3 release to vary one factor at a time and avoid an unrelated controller upgrade. I separately inspected the tagged v4.1.4 source and current main: commonPodEvent has the same timed-removal logic. That is source evidence, not a runtime claim for those builds.

Minimal workflow

Use a dedicated test namespace, an isolated controller managing only that namespace, and a workflow-runner ServiceAccount with the standard WorkflowTaskResult create/patch permissions. The label below must match that controller's configured instance ID.

apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  name: status-capture-repro
  namespace: status-capture-test
  labels:
    workflows.argoproj.io/controller-instanceid: status-capture-test
spec:
  serviceAccountName: workflow-runner
  entrypoint: main
  templates:
    - name: main
      container:
        image: busybox:1.37
        command: [sh, -c]
        args: ['sleep 30; echo capture-ok']
        resources:
          requests:
            cpu: 10m
            memory: 16Mi
          limits:
            memory: 32Mi

Names are normalized here. The executed test additionally pinned this CPU Pod to an existing node by hostname to avoid provisioning or node disruption during the experiment; that environment-specific selector is omitted above.

Reproduction sequence

Only use a disposable, isolated controller for these steps.

  1. Start that controller with ARGO_POD_STATUS_CAPTURE_FINALIZER=true and --workflow-workers=2. Submit the workflow and wait until the persisted Workflow node phase is Running.
  2. Scale only the isolated controller to zero before the sleep finishes. Do not stop the executor Pod.
  3. Wait until the Pod is Succeeded, verify both main and wait terminated with exit code 0, and record the original Pod UID. The Workflow still records Running.
  4. Wait until at least 190 seconds have elapsed since the latest Pod condition lastTransitionTime.
  5. Request deletion of that completed Pod. Verify the same UID remains with a deletion timestamp and workflows.argoproj.io/status while the controller is stopped. Do not remove any finalizer manually.
  6. Restart the isolated controller with --workflow-workers=0. Pod cleanup remains active. Observe that the native finalizer is removed and the Pod disappears while the persisted Workflow node is still Running.
  7. Stop the isolated controller, restore --workflow-workers=2, and start it again.
  8. Observe Workflow/node phase Error and message pod deleted.

No Workflow status field was patched to manufacture this outcome. A control with normal Workflow workers after the same aging/deletion sequence succeeded, which is consistent with reconciliation winning the race in that run.

Expected and actual behavior

Expected: while the owning Workflow exists and is active, retain a terminal Pod's status-capture finalizer until its result is persisted, even if reconciliation is delayed. Cancellation, Workflow deletion and orphan cleanup still need to work.

Actual: elapsed time independently authorizes removal before Workflow persistence; a successful execution is subsequently reported as an error.

Controller evidence

The following is an excerpt from the isolated controller, with resource names omitted:

16:31:11.132 Current Worker Numbers ... workflowWorkers=0 ... podCleanup=4
16:31:11.257 Removing finalizers during a delete ... pod.Finalizers=[workflows.argoproj.io/status]
16:31:11.257 queuing pod delay ... delay=-1m24.257168224s action=removeFinalizer
16:31:11.350 cleaning up pod ... action=removeFinalizer
16:31:11.363 delete pod event

At this point the Pod was absent and the Workflow node remained Running. After restoring Workflow workers, its recorded terminal state was:

phase: Error
message: pod deleted
taskResultSynced: true

Executor completion evidence

The test did not retain the wait container's stdout; the retained Kubernetes container statuses show:

main:
  exitCode: 0
  reason: Completed
  startedAt: '2026-09-19T16:27:14Z'
  finishedAt: '2026-09-19T16:27:44Z'
wait:
  exitCode: 0
  reason: Completed
  startedAt: '2026-09-19T16:27:14Z'
  finishedAt: '2026-09-19T16:27:44Z'

Latest Pod condition transition: 2026-09-19T16:27:47Z. Deletion was requested at approximately 16:31:04Z. Both containers had completed before deletion.

Observed images:

  • docker.io/library/busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0
  • quay.io/argoproj/argoexec@sha256:6aa5cf3452b12c8498eccc62455e460e94e6d092931932aa4dbda6bb1512036c

Source analysis

commonPodEvent in v4.1.3 queues removeFinalizer when a Pod has a deletion timestamp. The deadline is based on the latest Pod condition transition plus the greater of PodGC delay and two minutes. For an aged Pod the remaining delay is nonpositive, so cleanup becomes immediately eligible, independently of Workflow persistence.

By contrast, persistUpdates invokes normal Pod cleanup after updating the Workflow, and queuePodsForCleanup checks that node state is fulfilled with task results synchronized.

Is timed removal during external deletion intended to override status-capture protection for a terminal Pod with an active owner, or should this path retain protection until reconciliation records the result?

Related issues checked

I did not find an open report specifically reproducing timed finalizer removal with the flag enabled and delayed Workflow workers. This report concerns that remaining path; it does not claim to establish a new regression version or a measured production failure rate.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions