Skip to content

spec: what Plan B's spike found about Debezium metrics and sink lag - #442

Merged
turbolytics merged 1 commit into
mainfrom
spec/connect-plan-b-spike
Oct 5, 2026
Merged

turbolytics merged 1 commit into
mainfrom
spec/connect-plan-b-spike

Conversation

@turbolytics

Copy link
Copy Markdown
Owner

A spike against Kafka Connect 3.9 and Debezium 3.0.8 answered the questions Plan B (Debezium event lag and snapshot progress, sink lag) depends on. This PR adds the findings to the spec.

  • Stale Debezium metrics. A task restarted mid-snapshot kept the old run's Debezium MBeans registered for 80 seconds. Debezium then gave up registering the new task's, and logged "metrics will not be available". The reporter treats an MBean registered before the task's last start as stale. A Debezium connector without current snapshot metrics sends backfill.state: unknown.
  • Matching Debezium MBeans to tasks. The reporter matches server to the connector's topic.prefix, and task= where SQL Server and MongoDB add it. SQL Server's per-database MBeans collapse as shards.
  • Sink lag per task. The broker lists each sink task's partitions under its client.id. A failed sink task leaves the consumer group, so it sends no lag.
  • Broker credentials. The extension receives the worker's connection and security settings, so the admin client can connect as the worker does.

If stale MBeans were read as current, a restarted task would report the previous run's snapshot progress and lag as its own.

A spike against Kafka Connect 3.9 and Debezium 3.0.8 answered the
questions Plan B depends on:

- Restarted mid-snapshot, a task's old Debezium MBeans stayed registered
  80 s; Debezium gave up registering the new ones and the task ran without
  snapshot metrics. MBeans registered before the task's last start are
  stale, and a Debezium connector without current snapshot metrics sends
  backfill.state unknown.
- Debezium MBeans match a task by server=topic.prefix and, for SQL Server
  and MongoDB, task=; per-database MBeans collapse as shards.
- The broker reports each sink task's partitions by client.id; a failed
  sink task leaves the group and sends no lag.
- The extension receives the worker's connection and security settings, so
  the admin client connects as the worker does.

If stale MBeans were read as current, a restarted task would report the
previous run's snapshot progress and lag as its own.
@turbolytics
turbolytics merged commit 511335b into main Oct 5, 2026
7 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 5, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant