Skip to content

[CI] Make transient no-CN startup probes distinguishable in standalone BVT #450

Description

@XuPeng-SH

Summary

The standalone multi-CN BVT startup probe reaches the proxy before CN heartbeat registration / the proxy cluster view has a routable CN. During that short startup window, the probe logs repeated ERROR 20101: no available CN server, then succeeds once CNs become visible. The timing is consistent with service-registration / snapshot convergence; confirm the exact boundary before changing product behavior.

This message is not the cause of the cited red BVT checks. The BVT suites continue to completion; their terminal failures are separate SQL/result assertions. This task is for clearer readiness signaling and faster, more accurate CI triage.

Evidence

MatrixOne consumes this workflow via matrixorigin/CI/.github/workflows/e2e-standalone-parallel.yaml@main.

Desired outcome

  • Make the BVT startup readiness state explicit and bounded, with a concise time-to-routable-CN result.
  • Distinguish expected pre-readiness connection attempts from actual test failures in the job summary, while preserving useful details in diagnostics.
  • If no CN becomes routable before the existing readiness deadline, fail with evidence that helps distinguish CN registration, HAKeeper/cluster refresh, and routing-filter causes.
  • Do not add fixed sleeps, unbounded retries, or suppress failures after the suite starts. Do not change proxy error semantics just to hide the startup window.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions