Skip to content

feat(storage): timed Parquet snapshots of the live database, and a restore path - #121

Merged
Fl0p merged 6 commits into
mainfrom
flo-974-prod-db-snapshot
Oct 6, 2026
Merged

Fl0p merged 6 commits into
mainfrom
flo-974-prod-db-snapshot

Conversation

@Fl0p

@Fl0p Fl0p commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Production had no backup of the live database. The only recovery point was a September copy that stops being data around 2026-10-27, when its newest span falls out of the 30-day raw retention window.

This ships the mechanism the accepted ADR picked, on the same branch as the ADR.

What changed

  • ADR-0023 (docs/decisions/0023-production-database-snapshots.md) — the decision plus the measurements it was chosen on. Renumbered from 0019 → 0020 → 0023 as the decision index moved twice during review; 0023 is the next free slot on main.
  • RunSnapshotWorker (internal/storage/snapshot.go) — EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION ZSTD) into one dated directory per snapshot under COTEL_SNAPSHOT_DIR, keeping the COTEL_SNAPSHOT_KEEP newest. Runs on the same single connection as ingest and the dashboard, so it cannot race the WAL checkpoint and never writes the live file. snapshot.json is written last and by nothing else, so its presence is the only "complete" signal; it carries per-table row counts read back out of the Parquet files. A cycle only runs when the newest complete snapshot is older than the interval, so a restart or crash loop cannot churn through the retained window. Pruning keeps the newest N complete, deletes incomplete ones, runs only after a successful export, and ignores directories whose name is not a snapshot instant.
  • cotel --db-import <dir> — restores into COTEL_DB_PATH and verifies every table against the manifest. Needs only the image already on the host: no DuckDB CLI with its version matched by hand. Refuses a populated target and a snapshot with no manifest.
  • /api/v1/health gains a snapshot object; a failed export degrades the top-level status, like a failed retention roll-up.
  • compose — second volume cotel-snapshots at /snapshots, name overridable via COTEL_SNAPSHOT_VOLUME, and the three new env vars with defaults (/snapshots, 6h, 56).
  • docs — new docs/operations/duckdb-snapshots.md (check it is working, restore into a probe volume, promote, what it does not cover, and that a snapshot carries the plaintext ingest tokens); env rows in README.md and docs/index.md; snapshot row in the API reference; CHANGELOG; recovery page points here and drops its "there is no backup" section. The recovery page's ARM CLI asset name is also corrected — DuckDB publishes …-aarch64.zip up to v1.2.x and …-arm64.zip from v1.3.0 (verified against four releases' asset lists).

Verification

Measured on robmini against a probe copy of the live 152.6 MB database (65 184 spans), through the production image's binary with access_mode=read_only: the export holds the connection for 0.15-0.3 s and writes 4.2 MiB; IMPORT DATABASE restores in under a second with every row count, the schema version and all four indexes matching the source. Full table in the ADR.

On this branch:

  • go vet ./... clean; go test ./internal/storage/ ./internal/api/ ./cmd/cotel/ green (10 new snapshot tests, incl. a round-trip restore and a test pinning that load.sql carries absolute paths).
  • npm run build in docs/ green (VitePress dead-link check).
  • End-to-end against the built binary with COTEL_SNAPSHOT_INTERVAL=10s COTEL_SNAPSHOT_KEEP=2: three cycles ran, the oldest was pruned at each one, /healthz stayed ok throughout, /api/v1/health reported the snapshot block, --db-import restored the newest snapshot and the same binary read the result back, and a second import into the now-populated target was refused.

Post-merge (merge = deploy to robmini) still to be shown on the issue: a snapshot on robmini created by the mechanism, prod /healthz before/after, and one restore rehearsal on a probe volume.

Summary by CodeRabbit

  • New Features
    • Added automatic, periodic database snapshots with configurable storage location, schedule, and retention.
    • Added snapshot restore with manifest and table-count verification.
    • Added snapshot health details to the health endpoint; snapshot failures now mark overall health as degraded.
  • Documentation
    • Added guidance for configuring, restoring, and managing snapshots, including recovery and storage considerations.
    • Updated database recovery instructions for ARM DuckDB CLI asset names.

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 2df2417f-e672-419b-8abe-738fa722536b
📥 Commits

Reviewing files that changed from the base of the PR and between 1bd7f98 and 6d29809.

📒 Files selected for processing (15)
  • CHANGELOG.md
  • README.md
  • cmd/cotel/main.go
  • docker-compose.yml
  • docs/.vitepress/config.js
  • docs/decisions/0023-production-database-snapshots.md
  • docs/decisions/index.md
  • docs/index.md
  • docs/operations/api-reference.md
  • docs/operations/duckdb-recovery.md
  • docs/operations/duckdb-snapshots.md
  • internal/api/handler.go
  • internal/api/handler_test.go
  • internal/storage/snapshot.go
  • internal/storage/snapshot_test.go
 ___________________________________________________
< My threat model includes gremlins after midnight. >
 ---------------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Fl0p

Fl0p commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Heads up on the number, not the content: 0019 is taken on main as of #126 and #127 - the alert-recovery record moved there from a duplicate 0017, and this branch was cut before that. Renumber this one to 0020 (file, heading and the index row) before merge, or it lands as the third ADR-0019 this morning.

@Fl0p

Fl0p commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

ADR number collision - renumber before merge. origin/main already carries docs/decisions/0019-ci-never-mutates-an-issue.md (landed in #126 / #127 while this branch was open). This PR adds docs/decisions/0019-production-database-snapshots.md.

Git will not warn you: the two filenames differ, so the only conflict is in index.md - resolve that by hand and main ends up with two ADR-0019s. Rebase on current origin/main, rename this one to 0020-production-database-snapshots.md, update index.md and any internal links.

Check with ls docs/decisions/ | cut -c1-4 | uniq -d - must print nothing.

-- Daedalus (CTO)

@Fl0p
Fl0p force-pushed the flo-974-prod-db-snapshot branch from 4d5dec2 to bddf2b3 Compare October 5, 2026 08:40
@Fl0p Fl0p changed the title docs(decisions): ADR-0019 - timed Parquet snapshots of the live database docs(decisions): ADR-0020 - timed Parquet snapshots of the live database Oct 5, 2026
Wayland and others added 3 commits October 6, 2026 19:09
Production has no recovery point: the 2026-10-04 recovery left three volumes
behind, two of which hold the same damaged file and the third is production
itself. The surviving September copy stops being data around 2026-10-27, when
its newest span falls out of the raw retention window.

ADR-0023 picks the mechanism and records the numbers it was picked on, measured
on robmini against a probe copy of the live 152.6 MB database: a full
EXPORT DATABASE to Parquet+ZSTD costs 0.15-0.3 s of connection time and 4.2 MiB,
and IMPORT DATABASE restores it in under a second with every row count, the
schema version and all four indexes matching the source.

Numbered 0023 because the decision index moved twice while the proposal was in
review; the earlier drafts claimed 0019 and 0020, both now taken on main.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…re from it

The snapshot has to come from inside cotel: DuckDB has one writer and the live
process holds the file lock, so no external process can open the database even
read-only. RunSnapshotWorker runs EXPORT DATABASE (FORMAT PARQUET, COMPRESSION
ZSTD) into one dated directory per snapshot under COTEL_SNAPSHOT_DIR, keeping the
COTEL_SNAPSHOT_KEEP newest. It runs on the same single connection as ingest and
the dashboard, which is why it cannot race the WAL checkpoint - the two are
serialised by construction rather than by a lock - and why the decisive cost is
how long the statement holds the connection, not how large its output is.

snapshot.json is written last and by nothing else, so its presence is the only
"complete" signal. It carries the row count per table read back out of the
Parquet files, not out of the live tables, so it describes what the snapshot
holds rather than what the database held a moment later. The export writes
straight into its final directory because DuckDB bakes absolute paths into
load.sql, which a rename would break; a test pins that.

A cycle only runs when the newest complete snapshot is older than the interval,
so a restart - or a crash loop - cannot spend the retained window on snapshots
minutes apart. Pruning keeps the newest N complete, deletes incomplete ones,
runs only after a successful export so the last one standing survives, and
ignores directories whose name is not a snapshot instant.

cotel --db-import <dir> restores into COTEL_DB_PATH and verifies every table
against the manifest, so a restore needs only the image already on the host
instead of a DuckDB CLI with its version matched by hand. It refuses a populated
target and a snapshot with no manifest.

The worker's outcome lands in settings and on /api/v1/health as a snapshot
object; a failed export degrades the top-level status, the way a failed
retention roll-up does.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A new operations page covers how to tell whether snapshots are happening, how to
restore one into a probe volume and promote it, and the three things snapshots
do not cover: host loss, docker volume prune on a stopped deploy, and anything
newer than the last snapshot. It also states plainly that a snapshot carries the
plaintext ingest tokens, because users.token is plaintext by design and a
snapshot without it would not restore to a working instance.

The recovery page gains a pointer to it, loses its "there is no backup" section,
and gets the ARM CLI asset name corrected: DuckDB publishes
duckdb_cli-linux-aarch64.zip up to v1.2.x and duckdb_cli-linux-arm64.zip from
v1.3.0, verified against the release assets of v1.1.3, v1.2.2, v1.3.0 and
v1.5.6.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Fl0p
Fl0p force-pushed the flo-974-prod-db-snapshot branch from bddf2b3 to aeb1b38 Compare October 6, 2026 17:21
@Fl0p Fl0p changed the title docs(decisions): ADR-0020 - timed Parquet snapshots of the live database feat(storage): timed Parquet snapshots of the live database, and a restore path Oct 6, 2026
@Fl0p
Fl0p marked this pull request as ready for review October 6, 2026 17:22
Wayland and others added 3 commits October 6, 2026 19:23
…not positive

A zero or negative COTEL_SNAPSHOT_INTERVAL made every cycle due the moment the
last one finished, which is a busy loop on the one connection ingest uses.

Also note in the operations page that a failed import leaves a half-populated
file the refusal check then rejects on retry: discard the probe volume rather
than trying to clean it.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A snapshot directory dated in the future - a snapshots volume carried over from
a host whose clock ran fast - would otherwise defer the next snapshot by the
whole skew rather than by the interval.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Punctuation only, confined to the 22 lines this branch introduced across the
docs, the compose comment and the sidebar label. Surrounding text is untouched.

Co-Authored-By: Wayland <wayland@agents.flopbut.local>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Fl0p

Fl0p commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@Fl0p
Fl0p merged commit 7058390 into main Oct 6, 2026
7 checks passed
@Fl0p
Fl0p deleted the flo-974-prod-db-snapshot branch October 6, 2026 17:39
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Pull request is closed.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant