Skip to content

docs(ops): add retention-gap recovery runbook - #181

Open
Hustler490 wants to merge 1 commit into
wraith-protocol:developfrom
Hustler490:docs/retention-gap-recovery
Open

Hustler490 wants to merge 1 commit into
wraith-protocol:developfrom
Hustler490:docs/retention-gap-recovery

Conversation

@Hustler490

Copy link
Copy Markdown

Closes #159

What

The deployment guide documents Soroban RPC retention windows and the mitigation options, but stops there. It never tells an operator what to actually do when a scanner is offline past the boundary. This adds a dedicated recovery runbook.

Changes

  • guides/ops/retention-gap-recovery.mdx (new) - step-by-step procedure covering:
    • Detection - compare lastProcessedLedger to the network tip and classify the gap against the ~17,280-ledger window; confirm via retention_window_exceeded / RetentionExceededError.
    • Checkpoint handling - stop the scanner, capture lastProcessedLedger before and after, never advance the checkpoint past unpersisted events.
    • Archive replay - full-history node with fallbackPolicy: "archive_node".
    • Dual-provider recovery - fail over to an independent provider for the missed range.
    • Horizon fallback - best-effort reconstruction when no archive/secondary covers the range.
    • Safe rescan + deduplication - stable event identity (contract_id, ledger, tx_hash, event_index) with a unique constraint and ON CONFLICT DO NOTHING upsert.
    • Alert thresholds - approaching-boundary (medium) and past-boundary (critical) Prometheus rules.
    • Operator evidence checklist and a post-recovery verification list.
  • guides/ops/monitoring-and-on-call.mdx - links the runbook from Playbook 3 (Indexer Stall) and Cross-References.
  • docs.json - registers the page in the Operations navigation group.

Validation

  • node localization-qa.cjs ? 0 issues (new page recognised as a canonical page).
  • node scripts/check-nav-coverage.mjs ? 72 pages checked, 72 nav entries verified.
  • docs.json parses as valid JSON.

@truthixify

Copy link
Copy Markdown
Contributor

The replacement removes the fake CLI commands, but several examples still use APIs that do not exist. fetchAnnouncementsStream is a Stellar entry-point export and takes (chain, options); it has no contractIds, retention, or fallbackPolicy. Horizon has no getTransactions helper, and announcements do not expose event_index. Please align the runbook with the current SDK and existing stellar_network_ledger_height metric.

The deployment guide explains Soroban RPC retention windows but stops at mitigation options, leaving operators without a step-by-step recovery procedure for a scanner that fell behind the retention boundary. Add a runbook covering detection, checkpoint handling, archive replay, dual-provider and Horizon recovery, safe rescan with deduplication, retention-aware alert thresholds, and the operator evidence to collect. Link it from the monitoring guide and register it in the Operations navigation.

Closes wraith-protocol#159
@Hustler490
Hustler490 force-pushed the docs/retention-gap-recovery branch from 34b215a to d9f412d Compare October 2, 2026 15:39

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Wave 9] Add a retention-gap recovery runbook for operators

2 participants