Conversation
The Kubernetes e2e wrapper decided whether the target cluster was OpenShift with a single `kubectl api-resources --api-group=route.openshift.io` call piped into `grep -q .`. Three properties made failure silent: kubectl's exit status was discarded because it was the left side of a pipe, its stderr went to /dev/null, and the probe ran exactly once. `api-resources` is a discovery call that fans out to every aggregated APIService, so it routinely returns partial results or fails during a transient unavailability, throttling, or auth blip -- none of which mean "this is not OpenShift." When that happened, OPENSHIFT_DETECTED kept its default of 0 and the run proceeded down the vanilla-Kubernetes path. That flag gates the SCC values overlay, the privileged SCC grant for openshell-sandbox, and using a passthrough Route instead of a port-forward as the gateway transport. The misdetection surfaced far from its cause as pod admission failures and `sandbox connect` timeouts, reading as product bugs rather than a harness misconfiguration. Replace the probe with `kubectl get --raw /apis/route.openshift.io/v1`, which asks about one API group and does not depend on full aggregated discovery. Observe kubectl's exit status directly and keep its stderr. Treat only a NotFound answer as conclusive evidence that the cluster is not OpenShift; retry anything else with exponential backoff and, if no conclusive answer arrives, exit non-zero naming the probe and the underlying error instead of continuing with OPENSHIFT_DETECTED=0. Log the outcome in both directions so "not OpenShift" is distinguishable from "the probe never ran", and add OPENSHELL_E2E_OPENSHIFT so a contributor who has diagnosed the problem can force the answer. The logic lives in e2e/support/gateway-common.sh so it can be unit tested with a fake kubectl on PATH, following the existing test-e2e-image-overrides.sh precedent. Fixes NVIDIA#4088 Signed-off-by: Russell Bryant <rbryant@redhat.com>
russellb
requested review from
a team,
derekwaynecarr,
mrunalp and
sjenning
as code owners
October 1, 2026 23:36
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Kubernetes e2e wrapper decided whether the target cluster was OpenShift with a probe that could fail silently, leaving
OPENSHIFT_DETECTED=0and reconfiguring the entire run down the vanilla-Kubernetes path. Detection now distinguishes a conclusive "not OpenShift" answer from a discovery failure, retries the latter, and aborts rather than guessing.Related Issue
Closes #4088
Changes
The old probe was:
Three properties made failure silent:
kubectl's exit status was discarded because it was the left side of a pipe, its stderr went to/dev/null, and the probe ran exactly once.api-resourcesis a discovery call that fans out to every aggregatedAPIService, so it returns partial results or fails outright during transient unavailability, throttling, or an auth blip — none of which mean the cluster is not OpenShift.kubectl get --raw /apis/route.openshift.io/v1instead. It asks about one API group and does not depend on full aggregated discovery.kubectl's exit status directly rather than masking it behind a pipeline, and keep its stderr so a failure can be reported.OPENSHIFT_DETECTED=0.OPENSHELL_E2E_OPENSHIFT=1|0to force the answer and skip the probe, following the existingOPENSHELL_E2E_*convention. An unrecognized value is a hard error rather than a silent default.The logic lives in
e2e/support/gateway-common.shso it is sourceable and unit-testable with a fakekubectlonPATH, following the existingtest-e2e-image-overrides.shprecedent. Theoc-is-required check is preserved, moved into the branch that follows detection.Why this matters
OPENSHIFT_DETECTEDgates seven behaviors, including the SCC values overlay, the privileged SCC grant for theopenshell-sandboxservice account, and using a passthrough Route instead of a port-forward as the gateway transport. Misdetection did not produce a detection error — it produced security contexts that SCC admission rejects andsandbox connectattempting SSH over a port-forward, which the script's own header comment explains can never complete. The failures surfaced far from the cause and read as product bugs.Testing
mise run pre-commitpasses.shellcheck -xis clean on all three changed shell files, andbash -n e2e/with-kube-gateway.shis clean.New
tasks/scripts/test-e2e-openshift-detection.sh, registered astest:e2e-openshift-detectionand added to the[test]depends list. It puts a fakekubectlahead ofPATHwith an invocation counter, so it can assert how many times the probe ran:1, one probe call0, exactly one call (no pointless retry), and the not-OpenShift line is logged1, three calls0, stderr names the probe path and the underlying errorOPENSHELL_E2E_OPENSHIFT=1against a broken kubectl →1, zero probe callsOPENSHELL_E2E_OPENSHIFT=falseagainst an OpenShift kubectl →0, zero probe callsOPENSHELL_E2E_OPENSHIFT=maybe→ non-zero exit with an explanatory messageThe suite was mutation-tested to confirm it has teeth: reintroducing the original bug (treating every failure as conclusive absence) and separately removing the retry loop each make case 3 fail.
Verified against a real OpenShift cluster
The one judgement call here is that conclusive-absence is recognized from kubectl's error text, since
get --rawreturns exit 1 for every failure class. That was checked against a live OpenShift 4.x cluster (RHCOS 9.6) by calling the helper directly:route.openshift.io/v11APIResourceList0Real kubectl emits
Error from server (NotFound): the server could not find the requested resourcefor an absent group, which satisfies both matchers. Note the third row: on persistent failure the result is empty rather than0, so even if that wording changed in a future kubectl, the failure mode is a loud abort and never the silent-vanilla regression this issue is about.Not verified: a full
with-kube-gateway.shrun end to end on OpenShift. The detection function itself was exercised against the live cluster as above.Checklist