Skip to content

Flaky test: KubernetesClusterDockerTest (k3s pod never becomes Ready) #20432

Description

@FrankChen021

This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.

Status: Open: no fix PR
Subject: KubernetesClusterDockerTest class setup (embedded-tests, docker-tests)
Failures: 2 · First seen: 2026-09-10 · Last seen: 2026-09-18

Root cause

K3sClusterResource.waitUntilPodIsReady waits 300 s for each pod's Ready=True condition. On timeout it throws an ISE that contains only the pod name. The pods have no readiness probe and run with DRUID_XMX=128m, so a container restart or a node NotReady blip on a busy runner leaves the pod unready, and the cause is not visible in the log.

Suggested fix

On timeout, include the pod status (conditions, restart count, last termination reason), its events and the log tail in the exception, and dump pod logs to druid-container-logs. Add a /status/health readinessProbe with an initial delay to manifests/druid-service.yaml.

Occurrences

Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.

Date Commit Job Failure log Detail Reported in
2026-09-10 61ed0a3 (#20272) docker-tests job 102731457503 router pod never became Ready #20312
2026-09-18 3cbd00d (#20363) docker-tests job 105449382571 coordinator pod not Ready within 300 s #20385

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions