Skip to content

orchestratord: add reconciliation metrics and Kubernetes events - #39426

Merged
Alphadelta14 merged 1 commit into
MaterializeInc:mainfrom
Alphadelta14:heather/CLO-188-orchestrator-metrics-events4
Oct 1, 2026
Merged

Alphadelta14 merged 1 commit into
MaterializeInc:mainfrom
Alphadelta14:heather/CLO-188-orchestrator-metrics-events4

Conversation

@Alphadelta14

@Alphadelta14 Alphadelta14 commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Motivation

Part of CLO-188. This reimplements #39030 on top of k8s-controller 0.13.0, which now provides reconciliation metrics and Kubernetes events itself (MaterializeInc/k8s-controller#52). The metric names, steps and lifecycle events are the same; orchestratord no longer carries its own reconciliation wrapper and event publisher.

Description

  • Dependency: bumps k8s-controller to 0.13.0 with its prometheus feature.
  • Metrics: registers k8s-controller's Prometheus metrics under the orchestratord prefix and reports every controller's passes to them:
    • orchestratord_reconciliations_total{controller, phase, outcome}
    • orchestratord_reconciliation_duration_seconds{controller, phase}
    • orchestratord_reconciliation_steps_total{controller, step, outcome}
    • orchestratord_reconciliation_step_duration_seconds{controller, step}
  • Steps: the Materialize, Balancer and Console reconcilers mark named steps through the TraceMetadata each pass receives, with the same step names as CLO-188 Add events and more metrics to orchestratord (attempt 3) #39030.
  • Failure events: each controller gets its own EventRecorder, so the event's reporting controller identifies it: orchestratord.materialize.cloud/materialize, /balancer or /console, with the pod name as the instance. A failed pass publishes a ReconcileFailed warning event on the resource, with the error and its causes as the note.
  • Lifecycle events: the Materialize controller publishes each UpToDate transition as an event once the status is written. The condition's reason and message become the event's. It's a warning when the environment is not up to date, except while waiting for approval.
  • RBAC: the operator's ClusterRole gains create and patch on events.k8s.io events.
  • Metrics catalog: mz-metrics-catalog only saw metric! invocations, so it now also recognizes PrometheusMetrics::new("<namespace>"). It documents those metrics by constructing them and reading back their descriptions, as it already does for tokio's runtime metrics. It now depends on k8s-controller and k8s-openapi, which makes the catalog lint slower to build.

Differences from #39030 that reviewers may notice:

  • The failure event reason is k8s-controller's ReconcileFailed, not ReconciliationFailed.
  • The phase label is phase rather than event_type.
  • Each controller reports as itself instead of all sharing orchestratord.materialize.cloud.
  • An identical lifecycle event repeated within 10 minutes increments the existing event's count instead of creating a new event. The messages include the generation number, so this only applies when a transition really repeats.

Verification

  • mz-metrics-catalog: a test that PrometheusMetrics constructions are documented with the right names and labels.
  • mz-orchestratord: a test of which lifecycle transitions are published as warnings.
  • Helm: the ClusterRole test asserts the events permission.
  • test/orchestratord:
    • A new failure-events workflow, with a nightly step. It drives an environment through two different failures, a missing backend secret and then an invalid license key, and checks that the ReconcileFailed event's note follows the current cause.
    • manually-promote now checks that each rollout phase was published as an event.
    • Deployments expected to fail (post_run_check with expect_fail) now also check for a ReconcileFailed event.

Release notes

This release will publish Kubernetes events on Materialize, Balancer and Console resources when the operator fails to reconcile them, and on each Materialize rollout phase, so kubectl describe shows why a resource isn't progressing. The operator's ClusterRole now includes create and patch on events.k8s.io events.

🤖 Generated with Claude Code

Adopt k8s-controller 0.13.0, which provides reconciliation metrics and
events, and use it in place of orchestratord's own wrapper:

- report every controller's passes and named steps to k8s-controller's
  Prometheus metrics, under the orchestratord prefix
- give each controller its own event recorder, so a failed pass is
  published as a ReconcileFailed warning event whose reporting
  controller names the controller that failed
- publish each UpToDate transition of a Materialize as an event
- grant the operator create and patch on events.k8s.io events
- teach mz-metrics-catalog to document metrics constructed with
  PrometheusMetrics::new, which it could not see as metric! invocations

Part of CLO-188.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Alphadelta14
Alphadelta14 merged commit 28cc791 into MaterializeInc:main Oct 1, 2026
88 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants