fix(kubernetes): serialize lifecycle cleanup with sandbox restart - #4078
matthewgrossman wants to merge 1 commit into
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Label |
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
078d7e5 to
9db51c0
Compare
|
/ok-to-test 9db51c0 |
elezar
left a comment
There was a problem hiding this comment.
Reviewed at 9db51c0. The per-sandbox gate and refresh under the gate address the reproduced lifecycle/reconciliation race within one Kubernetes driver instance. No blocking findings; all 274 Kubernetes driver unit tests passed locally, and the passing Kubernetes E2E run covers this commit.
We have started investigating the broader compute-driver lifecycle contract: observations and cleanup from an earlier sandbox operation must not stop, delete, or invalidate a runtime established by a later operation. That investigation includes local reproduction and remediation of the Podman restart/watch race, with the aim of addressing Podman failures; attribution to historical CI failures remains unproven. Cross-process Kubernetes/HA coordination also remains a separate follow-up. These broader gaps do not block this scoped fix.
Summary
A Kubernetes reconciliation pass can retain a stopped or stopping Sandbox LIST snapshot while restart creates a replacement supervisor Pod at the same name. Cleanup can then delete the replacement before a later CR update detects the conflict. Serialize lifecycle mutations and reconciliation per sandbox within one driver instance, and refresh the CR under the gate before deciding whether to clean up.
This is a proposed quick fix for a reproduced race. The historical sandbox-stop-start CI failure lacked Kubernetes Pod diagnostics, so its exact cause remains unproven.
Related Issue
Refs #3954. No new issue was created, as explicitly requested; this PR does not close the broader investigation. This cleanup guard complements the bootstrap fence publication work in #4067 and diagnostics in #4072.
Changes
The gate is process-local: separate gateway or external driver processes can still race. This draft does not change generic Kubernetes 404 error mapping or replace generation/UID fencing with distributed exclusion.
Testing
Rebased without conflicts onto
76cfd0e31d5e1633db7ccd86ad9023ef7a2461b2(includes #4076). Current candidate:9db51c01936c6eeafdc59c74612e9839e1623e1d.git range-diffconfirms this PR's patch is unchanged.cargo clippy --offline -p openshell-driver-kubernetes --all-targets -- -D warningsandgit diff --checkpass.PR head, workflow heads, binary artifact sources, preparation checkouts, and image tags were verified against the candidate SHA. Carried-forward GitHub results with regenerated job IDs are excluded from repeat counts using actual execution timestamps and logs. No failures were observed in this campaign.
Optional Kubernetes HA and credential-driver suites were disabled in the full run and are not counted as passing coverage. Passing K3s logs show actual conformance archive execution and its success assertion; the pinned archive includes the unignored lifecycle scenario, but passing per-command output is suppressed.
The original commit passed the full configured pre-commit hook. Local cluster E2E previously stopped before scenarios because Docker was unavailable; that preflight attempt is not qualification evidence for this candidate.
Checklist