Skip to content

feat(health): page when the database snapshot stops happening - #140

Merged
Fl0p merged 1 commit into
mainfrom
flo-1016-probe-snapshot-health
Oct 6, 2026
Merged

Fl0p merged 1 commit into
mainfrom
flo-1016-probe-snapshot-health

Conversation

@Fl0p

@Fl0p Fl0p commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

The gap

docs/operations/duckdb-recovery.md now says the recovery point comes from snapshots. The mechanism works; nothing watches it. snapshotHealth exists only on /api/v1/health, while the scheduled probe reads /healthz, whose body carries ok, spans, last_ingest_at, newest_span_age_seconds and nothing else. The snapshot worker could die or start failing and the first notice would be a human opening the endpoint by hand.

What changed

scripts/probe-healthz.sh follows a green /healthz with a second question to /api/v1/health:

snapshot object Verdict
status: "error" red, exit 6, last_error quoted into the line
status: "ok", last_run_at older than SNAPSHOT_STALE_AFTER_SECONDS (default 43200 = 2 × the shipped 6h interval) red, exit 6
status: "ok", last_run_at absent or unparseable red, exit 6
status: "unknown" not asserted, said on the green line
no snapshot field, or /api/v1/health unreadable not asserted, said on the green line
  • unknown is the ambiguous half — it means both "the worker has not run yet" and "snapshots are disabled", and disabled is the shipped default everywhere but production. Reddening on it would lie on every local instance, so it is silent by default. SNAPSHOT_CHECK=require is the caller asserting that snapshots are expected on this instance and turns it — plus a missing field and an unreadable endpoint — red. SNAPSHOT_CHECK=off skips the question entirely.
  • No ~/ops edit needed. The /api/v1/health address is derived from the /healthz URL the wrapper already passes (same entry point, other path), so the merge reaches the timer by itself. API_HEALTH_URL overrides the derivation.
  • Exit 6, not 5. ~/ops/cotel-healthz.sh already spends 5 on "the LAN half is green and the public ingest half is not", and passes a probe's own code through otherwise.
  • /healthz is untouched, in fields and in the meaning of ok. It is the container liveness contract (cotel --healthcheck, the Docker HEALTHCHECK, scripts/wait-for-healthy.sh), and a failed backup must not mark a working container unhealthy or fail a deploy.
  • The edge half is unaffected — it runs probe-edge-ingest.sh and never reads /healthz. Through Cloudflare Access the default auto degrades quietly; the docs say not to set require on a vantage point behind Access.

Verification

Green on live production (LAN):

$ HEALTHZ_URL=http://robmini.local:8080/healthz bash scripts/probe-healthz.sh
probe-healthz: OK — HTTP 200 ingest age 2s (threshold 21600s) url=http://robmini.local:8080/healthz | snapshot last run 2026-10-06T17:39:54Z (age 909s, threshold 43200s) dir=/snapshots/2026-10-06T17-39-54Z
rc=0

Classification tests — bash scripts/probe-healthz_test.sh → passed=27 failed=0 (16 new; snapshot error, overdue, missing/unparseable timestamp, unknown under both modes, absent field, unreadable endpoint, off, config validation, and the URL derivation, which every /snap-*/healthz case exercises because the probe is given only the /healthz URL).

/healthz contract — go test ./internal/dashboard/ -run TestHealthz green (in golang:1.24, no Go toolchain on this host). New TestHealthzIgnoresTheSnapshotWorker writes snapshot_last_status=error to settings and pins the status code, ok, and the exact field set.

Docs — npm run build in docs/ completes clean.

The alert path, not just the check. Driven through the real host wrapper and the real pager on the drill marker, with a mock serving a green /healthz and an /api/v1/health reporting status: "error":

$ COTEL_HEALTHZ_REPO=<this branch> COTEL_HEALTHZ_REF=flo-1016-probe-snapshot-health \
  PC_ORIGIN_ID=cotel-health-probe-drill COTEL_HEALTHZ_URL=http://127.0.0.1:PORT/broken/healthz \
  ~/ops/cotel-healthz.sh --half local
[…] RED (streak 1/2, not paging yet) — probe-healthz: FAILED — no current database snapshot: snapshot worker reported error last_run_at=… last_error=export failed: IO Error: No space left on device …
[…] RED (streak 2), paged — … | page-cotel-health: opened FLO-1017 …

A real alert issue was created and assigned, and the assignment wake was claimed by a live run — delivery, not just detection. The green tick that closes it is run after the assignment run finishes, with the recovery wake checked in diagnostics/wakes for a request id of its own.

Docs

  • docs/operations/duckdb-snapshots.md — new Who is watching it section: what is red, why the threshold is two intervals, how to check by hand.
  • docs/operations/health-probe.md — the snapshot question, the verdict table, the two vantage points it must not be enabled for, exit 6, and the probe env table.
  • README.md / docs/index.md — the three new probe env vars in the env tables, marked as read by scripts/probe-healthz.sh rather than the binary; /healthz section states why worker health is deliberately absent from it.
  • CHANGELOG.md — entry under Unreleased → Added.

Known limitation (needs a ~/ops edit, not made here)

The alert's title still reads cotel prod /healthz is red on a snapshot failure: PC_ALERT_SUBJECT is chosen by the host wrapper per half, not per verdict, and the wrapper is out of scope for this repo. The probe's verdict line in the body names the snapshot failure, and --probe-only reprints it. Say the word and I will write it up for the board.

Summary by CodeRabbit

  • New Features
    • The scheduled health probe now checks snapshot-worker status and freshness after confirming service readiness. Failed, invalid, or older-than-12-hour snapshots trigger an alert.
    • Configure snapshot checks to run automatically, require a verifiable snapshot, or be turned off. The freshness window and health endpoint can also be customized.
  • Documentation
    • Clarified that /healthz reports container readiness, while snapshot and retention health are available through /api/v1/health. Snapshot issues affect alerting, not /healthz liveness.

The snapshot worker's report lives only on /api/v1/health, and the
scheduled probe reads /healthz, so a backup could die or start failing
and nothing would say so until a human opened the endpoint by hand.

A green /healthz is now followed by a second question to /api/v1/health,
whose address is derived from the /healthz URL the probe is already
given - so the host wrapper on the Pi needs no edit and the check
reaches the timer with the merge. Red (exit 6, not 5, which the wrapper
already spends on "LAN green, public ingest not") on status "error", on
a last_run_at older than SNAPSHOT_STALE_AFTER_SECONDS, and on a
status "ok" whose timestamp is missing or unparseable.

status "unknown" stays silent: it means both "the worker has not run
yet" and "snapshots are disabled", and disabled is the shipped default
everywhere but production, so reddening on it would lie on every local
instance. SNAPSHOT_CHECK=require is the caller asserting snapshots are
expected here and turns that - plus a missing snapshot field and an
unreadable /api/v1/health - red too; off skips the question.

/healthz is unchanged in fields and in the meaning of ok, with a test
pinning both against a worker reporting a failed cycle: that endpoint is
the container liveness contract, and a failed backup must not mark a
working container unhealthy or fail a deploy.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Review in Change Stack →

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 5d208657-fefb-4251-b541-a2cc3f7d55b0
📥 Commits

Reviewing files that changed from the base of the PR and between 7058390 and 04afcd6.

📒 Files selected for processing (8)
  • CHANGELOG.md
  • README.md
  • docs/index.md
  • docs/operations/duckdb-snapshots.md
  • docs/operations/health-probe.md
  • internal/dashboard/healthz_test.go
  • scripts/probe-healthz.sh
  • scripts/probe-healthz_test.sh
 __________________________________
< Bazinga! You missed a semicolon. >
 ----------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Fl0p
Fl0p merged commit 3d06113 into main Oct 6, 2026
7 of 8 checks passed
@Fl0p
Fl0p deleted the flo-1016-probe-snapshot-health branch October 6, 2026 18:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant