[release-5.0] NO-ISSUE: Test and script fixes found during 5.0.0-rc.0 testing - #7327
Conversation
|
@agullon: This pull request explicitly references no jira issue. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository YAML (base), Central YAML (inherited) Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
|
Pipeline controller notification For optional jobs, comment This repository is configured in: LGTM mode |
|
/pipeline required |
|
Scheduling tests matching the |
|
/pipeline required |
|
Scheduling tests matching the |
|
/test e2e-aws-tests-release |
|
/retest |
9dd9a8c to
8e92219
Compare
5d5ccd5 to
24dd421
Compare
24dd421 to
fe5165b
Compare
|
/pipeline required |
|
/test e2e-aws-tests-release |
|
/pipeline required |
|
Scheduling tests matching the |
The .local TLD is reserved for mDNS (RFC 6762) and can cause DNS interference with OVN initialization on systems with Avahi or systemd-resolved, contributing to healthcheck timeouts after hostname changes. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
The "Case Insensitive Log Levels" test fails on ARM when testing TraceAll because klog verbosity 10 generates ~144K journal messages, exceeding journald's default rate limit of 10K messages per 30 seconds. The suppressed messages include the startup config dump line that the test greps for, making the assertion impossible to satisfy. Instead of globally disabling rate limiting in all test VMs, scope the fix to the logging suite: disable rate limiting in suite setup and re-enable it in teardown. The drop-in write and journald restart are wrapped in `bash -c` so the redirection and restart run as root under SSHLibrary's sudo=True (the shell that expands `>` would otherwise run unprivileged), and `mkdir -p` keeps the keyword self-contained rather than relying on kickstart provisioning of /etc/systemd/journald.conf.d. Co-Authored-By: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
The Remove Namespace keyword ran `oc delete namespace` under Run With Kubeconfig's 300s process timeout with allow_fail defaulting to False. On ARM dual-stack, namespace garbage collection (pod teardown plus finalizers, slowed by OVN reconciliation saturating the CPU) can exceed 300s, so the oc process was killed, returned non-zero, and cascaded into a suite teardown failure that retroactively marked all passing tests as failed. Delete with allow_fail=True (matching the pattern already used across the suites, e.g. standard1/kustomize.robot) so a slow or killed cleanup does not fail the teardown, then poll until the namespace is actually gone. The poll asserts a genuine NotFound rather than any non-zero exit code (which a transient API error could also produce), so a subsequent test can reuse the namespace name without a "being deleted" collision, bounded to 5m. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
Release scenarios running upgrade paths with LVMS workloads followed by full standard suites were hitting timeout limits under I/O contention when many VMs boot and pull images in parallel. Set the greenboot healthcheck timeout to 1200s (from 600s) and the robot framework timeout to 60m centrally in ci_phase_boot_and_test.sh for all release scenarios, and remove the now-redundant per-scenario overrides so the value lives in a single place. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
openshift#7298 reduced release-scenario VM disks from 30GB to 20GB. That is safe for the lvms-standard/standard scenarios (they create a single 1Gi PVC), but the ginkgo scenario runs the full storage spec suite which requests several 1Gi PVCs concurrently. At 20GB the topolvm data VG only has ~420MiB free, so 7 storage specs fail with: ResourceExhausted ... no enough space left on VG: free=440401920, requested=1073741824 (arm-el10) and on x86 el10 the same undersized VM shows etcd ReadIndex latency and apiserver TLS-handshake flaps from I/O contention. Restore only this scenario to --vm_disksize 30; the other nine 20GB scenarios stay as-is since a single 1Gi PVC fits comfortably. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
On a fresh/clean start the standard2 Log Scan test intermittently fails "Should Not Find Forbidden" with: pods cert-manager-cainjector-... is forbidden: error looking up service account cert-manager/cert-manager-cainjector: serviceaccount ... not found Investigation: the cainjector/controller/webhook Deployments and their ServiceAccounts are created dynamically by the cert-manager operator, not by MicroShift's static manifests. The kube-controller-manager ReplicaSet controller can briefly attempt to create a pod before its ServiceAccount is observed, logging this transient "forbidden" and retrying it away once the SA lands. The workloads become ready (MicroShift healthcheck passes), so this is a benign eventual-consistency startup race, not a MicroShift manifest-ordering bug — the ordering is the operator's, not ours. Add a scoped known-exceptions allowlist to the journalctl log-scan helper and register this single pattern, so genuine "forbidden" regressions still fail while this benign race is ignored. While here, make the test itself easier to read: rename the per-boot keyword to "Boot And Scan Journal", move the repeated journal-cursor capture into it, and split the assertions into "Scan Boot Journal". The two calls now read as "clean first boot" (forbidden check skipped, since a clean boot logs the benign race above) and "restart" (must be forbidden-free). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> pre-commit.check-secrets: ENABLED
Two release-scenario tests raced asynchronous resource creation:
- otp-workloads/statefulset-pvc: `oc wait pod/hello-statefulset-0
--for=condition=Ready` ran immediately after creating the StatefulSet,
before its controller created pod-0, so `oc wait` on the named pod
failed with NotFound. Wait for the pod to exist first (reusing
Wait Until Resource Exists) before checking readiness.
- ai-model-serving/ai-model-serving-online: the ServingRuntime was
applied with a bare `oc apply` before the kserve validating webhook had
endpoints, so the apply was rejected ("no endpoints available for
service kserve-webhook-server-service"). Retry the apply until the
webhook is serving, mirroring the retry already used for the
InferenceService rollout.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
pre-commit.check-secrets: ENABLED
3a94737 to
17d125b
Compare
|
/test e2e-aws-tests-release |
|
/pipeline required |
|
Scheduling tests matching the |
|
/retest |
|
/override ci/prow/e2e-aws-tests-bootc-release-arm-el9 I investigate in detail every failure on every scenarios and they are safe to override because they are just few flaky that passed after a retrigger or there are other PR to get it fix like #7347 |
|
@agullon: Overrode contexts on behalf of agullon: ci/prow/e2e-aws-tests-bootc-release-arm-el10, ci/prow/e2e-aws-tests-bootc-release-arm-el9, ci/prow/e2e-aws-tests-bootc-release-el10, ci/prow/e2e-aws-tests-bootc-release-el9, ci/prow/e2e-aws-tests-periodic, ci/prow/e2e-aws-tests-release DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
/label jira/valid-bug |
|
@agullon: This PR has been marked as verified by DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
/lgtm |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: agullon, pmtk The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
|
@agullon: The following tests failed, say
Full PR test history. Your PR dashboard. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here. |
88ae25f
into
openshift:release-5.0
Summary
Backport of #7326 to
release-5.0. A consolidation of test-harness and CI fixes discovered while running the 5.0.0-rc.0 release test suites. All changes are test/CI-only (no product code): they address CI flakes and failures — resource exhaustion, teardown cascades, log-scan false positives, timeout limits, and asynchronous-resource races.Parent PR (main): #7326 — the commits here are identical to that PR, rebased onto
release-5.0.Changes
Remove Namespacenow deletes withallow_fail=True(matching the existing repo idiom) and polls until the namespace is actually gone — asserting a realNotFound— so a slow deletion killed at the 300s process timeout no longer cascades into a suite-teardown failure, while still guaranteeing the name is free for reuse.GREENBOOT_TIMEOUT=1200andTEST_EXECUTION_TIMEOUT=60minci_phase_boot_and_test.shfor release scenarios and remove the now-redundant per-scenario overrides.--vm_disksize 30for the ginkgo scenario, whose storage specs exhaust the topolvm VG at the 20GB default.forbiddenregressions still fail), and clarify the double-boot test with aBoot And Scan Journalkeyword.bash -cfor correct sudo redirection, withmkdir -pfor self-containment) so high-verbosity log output isn't dropped..exampleTLD instead of.localto avoid mDNS interference.statefulset-pvcwaits for the StatefulSet pod to exist before the readiness check, andai-model-serving-onlineretries the kserveServingRuntimeapply until the validating webhook has endpoints.Testing
journalctl.py) passes flake8; shell scripts pass shellcheck.