diff --git a/docs.json b/docs.json
index dd51c93..24db23b 100644
--- a/docs.json
+++ b/docs.json
@@ -181,7 +181,8 @@
"guides/stellar/subscriptions-with-wraith-names",
"guides/wraith-names-stellar",
"guides/ops/self-hosted-deployment",
- "guides/ops/monitoring-and-on-call"
+ "guides/ops/monitoring-and-on-call",
+ "guides/ops/retention-gap-recovery"
]
}
]
diff --git a/guides/ops/monitoring-and-on-call.mdx b/guides/ops/monitoring-and-on-call.mdx
index 73f8fda..ea49c2b 100644
--- a/guides/ops/monitoring-and-on-call.mdx
+++ b/guides/ops/monitoring-and-on-call.mdx
@@ -403,6 +403,13 @@ Each playbook follows a standard structure: **Symptoms → Diagnosis → Immedia
**Severity**: High ([Auditor Guide Sev 2](/reference/auditor-guide#severity-matrix))
**Response SLA**: 1 hour to acknowledge, 4 hours to mitigation
+
+ If the indexer's `lastProcessedLedger` gap exceeds the Soroban RPC retention
+ window (~17,280 ledgers / ~24 hours), this is no longer a stall — it is a
+ retention gap. Follow the
+ [Retention‑Gap Recovery Runbook](/guides/ops/retention-gap-recovery) instead.
+
+
#### Symptoms
- `IndexerBacklogBuildup` or `AnnouncementLagHigh` alert firing
- Users report delayed stealth payment notifications
@@ -675,6 +682,7 @@ Run different scenarios each quarter:
## Cross-References
- [Self-Hosted Deployment](/guides/ops/self-hosted-deployment) — prerequisite deployment guide
+- [Retention‑Gap Recovery Runbook](/guides/ops/retention-gap-recovery) — recover a scanner that fell behind the RPC retention window
- [Auditor Guide Severity Matrix](/reference/auditor-guide#severity-matrix) — severity definitions and response SLAs
- [Multisig Authority Rotation](/guides/stellar/multisig-authority-rotation) — key rotation procedures
- [Stellar Cryptography](/architecture/stellar-cryptography) — contract architecture for debugging
diff --git a/guides/ops/retention-gap-recovery.mdx b/guides/ops/retention-gap-recovery.mdx
new file mode 100644
index 0000000..90e61a8
--- /dev/null
+++ b/guides/ops/retention-gap-recovery.mdx
@@ -0,0 +1,314 @@
+---
+title: 'Retention‑Gap Recovery Runbook'
+sidebarTitle: 'Retention‑Gap Recovery'
+description: 'Step-by-step operator procedure for recovering a scanner that fell behind the Soroban RPC retention window'
+---
+
+This runbook is the recovery procedure for the failure mode described in
+[Soroban RPC Retention Windows](/guides/stellar-mainnet-deployment#soroban-rpc-retention-windows):
+a scanner (indexer or watcher) was offline long enough that its
+`lastProcessedLedger` checkpoint fell **behind the tip of the RPC node's
+retention window**, so the node no longer serves the events the scanner still
+needs.
+
+
+ This is a **data-recovery** procedure. Every step is written to avoid
+ advancing the checkpoint past events that have not actually been persisted,
+ and to make the recovery safe to re-run. Do not skip the
+ [evidence collection](#operator-evidence-to-collect) step — it is what makes
+ the incident reviewable.
+
+
+## When to use this runbook
+
+Use it when any of the following is true:
+
+- An alert fired because `lastProcessedLedger` lagged past the retention
+ boundary (see [alert thresholds](#alert-thresholds)).
+- Logs contain `retention_window_exceeded` (see
+ [Error Codes](/reference/error-codes#soroban-rpc--indexer-errors)).
+- `fetchAnnouncementsStream` / `getEvents` throws `RetentionExceededError`
+ because the requested `fromLedger` was pruned.
+- An operator reports missing announcements that are older than ~24 hours.
+
+If the scanner is merely lagging but still **inside** the window, use
+[Playbook 3: Indexer Stall](/guides/ops/monitoring-and-on-call#playbook-3-indexer-stall)
+instead — this runbook is only for crossing the retention boundary.
+
+## 1. Detect the gap
+
+Detection is a comparison between the scanner's checkpoint and the network tip.
+
+1. Read the current network tip ledger:
+
+ ```bash
+ curl -s https://soroban-mainnet.stellar.org \
+ -X POST -H 'Content-Type: application/json' \
+ -d '{"jsonrpc":"2.0","id":1,"method":"getLatestLedger"}'
+ ```
+
+2. Read the scanner's checkpoint (`lastProcessedLedger`) from its datastore.
+3. Compute the gap: `gap = latest_ledger - lastProcessedLedger`.
+4. Compare the gap against the provider's retention window
+ (`~17,280` ledgers ≈ 24 hours for public nodes, per
+ [Soroban RPC Retention Windows](/guides/stellar-mainnet-deployment#retention-limits)).
+
+| Gap (ledgers) | Meaning | Action |
+|---|---|---|
+| < 8,640 (~12h) | Inside the window | Continue normal operation; [Playbook 3](/guides/ops/monitoring-and-on-call#playbook-3-indexer-stall) applies |
+| 8,640 – 17,280 | Approaching the boundary | Treat as a **warning**; verify failover is healthy |
+| > 17,280 (~24h) | **Past the boundary** | Continue with this runbook from step 2 |
+
+Also confirm the gap in the error signal, not just the arithmetic:
+
+```bash
+grep -E "retention_window_exceeded|RetentionExceededError" /var/log/wraith-indexer.log | tail -20
+```
+
+
+ The storage of `lastProcessedLedger` is implementation-specific (a table, a
+ KV entry, or a file). Locate your scanner's checkpoint store before starting —
+ you will read and write it several times below.
+
+
+## 2. Stop the scanner and freeze the checkpoint
+
+1. Stop the scanner so it cannot advance the checkpoint mid-recovery:
+
+ ```bash
+ systemctl stop wraith-indexer # and/or wraith-watcher
+ ```
+
+2. Record the checkpoint value **before** changing anything:
+
+ ```bash
+ # Example: checkpoint persisted in Postgres
+ psql "$DATABASE_URL" -c \
+ "SELECT value FROM indexer_state WHERE key = 'lastProcessedLedger';" \
+ | tee checkpoint-before.txt
+ ```
+
+3. Record the network tip as the upper bound of the recovery range:
+
+ ```bash
+ echo "recovery_range = (lastProcessedLedger, latest_ledger]"
+ ```
+
+
+ Never move the checkpoint forward to `latest_ledger` to "clear" the error. Doing
+ so permanently discards the announcements in the gap and is unrecoverable.
+
+
+## 3. Recover the missed range
+
+Pick exactly one of the following, in order of preference. All of them replay
+the range `(lastProcessedLedger, latest_ledger]`; they differ only in where the
+history comes from.
+
+### Option A — Archive replay (most reliable)
+
+Replay the pruned range from a node that retains full history.
+
+```typescript
+import { fetchAnnouncementsStream } from '@wraith-protocol/sdk';
+
+const stream = fetchAnnouncementsStream({
+ // Archive node with full history enabled.
+ rpcUrl: process.env.STELLAR_ARCHIVE_RPC_URL,
+ contractIds: [ANNOUNCER_CONTRACT_ID],
+ fromLedger: lastProcessedLedger, // replay exactly the missed range
+ retention: {
+ fallbackPolicy: 'archive_node', // fail loudly if even the archive lacks it
+ },
+});
+
+for await (const announcement of stream) {
+ await persistAnnouncement(announcement); // idempotent — see step 4
+ if (announcement.ledger > latestLedger) break;
+}
+```
+
+### Option B — Dual-provider recovery
+
+If your primary provider pruned the range, fail over to the independent
+secondary provider and replay from there. This is why dual-provider setups are
+recommended in
+[Mitigating Retention Gaps](/guides/stellar-mainnet-deployment#mitigating-retention-gaps).
+
+```bash
+# Confirm the fallback provider still serves the missed range.
+curl -s "$STELLAR_RPC_URL_FALLBACK" \
+ -X POST -H 'Content-Type: application/json' \
+ -d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"getEvents\",
+ \"params\":{\"startLedger\":$lastProcessedLedger,\"filters\":[]}}"
+```
+
+If the fallback returns events for `startLedger`, replay the range against it,
+then restore the primary once it is healthy. If **both** providers have pruned
+the range, use Option A (archive) or Option C (Horizon).
+
+### Option C — Horizon fallback (best-effort)
+
+Horizon retains transaction history longer than the RPC event window and can
+reconstruct announcement events from transaction results, at the cost of extra
+parsing. Use this only when no archive or secondary provider covers the range.
+
+```typescript
+// Reconstruct announcements from Horizon transactions in the gap.
+// See /guides/stellar-troubleshooting for the retention-window error context.
+const txs = await horizon.getTransactions({ from: gapStartTime, to: gapEndTime });
+const announcements = txs
+ .flatMap(reconstructAnnouncementsFromTransaction) // announce events in tx meta
+ .filter(isWraithAnnouncement);
+```
+
+
+ Horizon reconstruction is lossy if it cannot fully decode contract event
+ metadata. Record how many events were recovered this way in the operator
+ evidence (step 6) so the coverage gap is documented.
+
+
+## 4. Safe rescan and deduplication
+
+A rescan overlaps data the scanner may have already persisted, so replay
+**must** be idempotent. Never assume the gap is clean of already-processed
+events.
+
+### Deduplicate on a stable event identity
+
+Derive a deterministic key from the event's on-chain identity — not from
+insertion time — and enforce it with a unique constraint.
+
+```sql
+-- Stable identity for a Wraith announcement.
+ALTER TABLE announcements
+ ADD CONSTRAINT announcements_event_key
+ UNIQUE (contract_id, ledger, tx_hash, event_index);
+```
+
+Then persist with an upsert so a replay can be run repeatedly:
+
+```sql
+INSERT INTO announcements (contract_id, ledger, tx_hash, event_index, payload)
+VALUES ($1, $2, $3, $4, $5)
+ON CONFLICT (contract_id, ledger, tx_hash, event_index) DO NOTHING;
+```
+
+### Rescan procedure
+
+
+
+ Run one of the recovery options above for the exact range
+ `(lastProcessedLedger, latest_ledger]`.
+
+
+ Compare the number of announcements persisted for the range against the
+ number the source returns:
+
+ ```sql
+ SELECT count(*) FROM announcements
+ WHERE ledger > $lastProcessedLedger AND ledger <= $latestLedger;
+ ```
+
+
+ Only once every event in the range is committed (or confirmed absent),
+ advance `lastProcessedLedger` to `latest_ledger`.
+
+
+ ```bash
+ systemctl start wraith-indexer # and/or wraith-watcher
+ ```
+
+
+ Watch `wraith_announcement_lag_seconds` and `wraith_indexer_sync_height`
+ until both return to their healthy baselines.
+
+
+
+
+ Keep the unique constraint in place after recovery. It is the mechanism that
+ makes any future rescan safe, not a one-time fix.
+
+
+## 5. Alert thresholds
+
+Add retention-aware alerts on top of the existing
+[Indexer Metrics](/guides/ops/monitoring-and-on-call#indexer-metrics). The
+existing alerts catch a stalled indexer; these catch a scanner that is about to
+cross the point of no return.
+
+```yaml
+groups:
+ - name: wraith_retention_alerts
+ interval: 1m
+ rules:
+ - alert: ScannerApproachingRetentionBoundary
+ expr: (wraith_latest_ledger - wraith_indexer_sync_height) > 8640
+ for: 10m
+ labels:
+ severity: medium
+ component: indexer
+ annotations:
+ summary: "Scanner sync gap > 8,640 ledgers (~12h)"
+ description: "The scanner is approaching the ~17,280-ledger retention boundary. Verify failover and prepare retention-gap recovery."
+ playbook: "https://docs.usewraith.xyz/guides/ops/retention-gap-recovery"
+
+ - alert: ScannerPastRetentionBoundary
+ expr: (wraith_latest_ledger - wraith_indexer_sync_height) > 17280
+ for: 1m
+ labels:
+ severity: critical
+ component: indexer
+ annotations:
+ summary: "Scanner gap exceeds the RPC retention window (~24h)"
+ description: "Events in the gap may be pruned. Run the retention-gap recovery runbook immediately."
+ playbook: "https://docs.usewraith.xyz/guides/ops/retention-gap-recovery"
+```
+
+| Threshold | Expression | Severity | Why |
+|---|---|---|---|
+| Approaching boundary | gap > 8,640 ledgers for 10m | Medium | Half the window; time remains to fail over |
+| Past boundary | gap > 17,280 ledgers for 1m | Critical | Range may be pruned; recovery is now required |
+| Retention error observed | `increase(wraith_retention_window_errors_total[5m]) > 0` | Critical | The boundary has already been crossed |
+
+If you do not yet export a ledger-gap gauge, derive it from
+`wraith_indexer_sync_height` plus the network tip, or from the
+`wraith_announcement_lag_seconds` metric already described in the monitoring
+guide.
+
+## 6. Operator evidence to collect
+
+Capture all of the following before closing the incident. Missing evidence is
+the most common reason a retention-gap postmortem cannot determine coverage.
+
+- [ ] `checkpoint-before.txt` — the `lastProcessedLedger` value at detent time.
+- [ ] Network tip ledger at detent time and the computed gap (ledgers and hours).
+- [ ] The exact recovery range `(lastProcessedLedger, latest_ledger]`.
+- [ ] Which recovery option was used (archive / dual-provider / Horizon) and why.
+- [ ] The `retention_window_exceeded` / `RetentionExceededError` log lines.
+- [ ] Announcement count the source returned for the range, and the count persisted after dedup.
+- [ ] Any events that could **not** be recovered, with the affected ledger range.
+- [ ] Before/after values for `wraith_announcement_lag_seconds` and `wraith_indexer_sync_height`.
+- [ ] The checkpoint value after recovery.
+
+## 7. Verify recovery
+
+1. `wraith_indexer_sync_height` lag is back under the healthy baseline
+ (< 10 ledgers).
+2. `wraith_announcement_lag_seconds` is back under 30s.
+3. The reconciled announcement count for the recovery range matches the source
+ (minus any documented unrecoverable events).
+4. `wraith_scan_miss_rate` is back under 0.1%.
+5. A spot-check of known stealth payments from the gap appears in the datastore.
+
+Once all five pass, remove the incident hold and note the outcome in the
+post-incident review.
+
+## Cross-References
+
+- [Soroban RPC Retention Windows](/guides/stellar-mainnet-deployment#soroban-rpc-retention-windows) — the retention limits this runbook recovers from
+- [Monitoring, Alerting, and On-Call](/guides/ops/monitoring-and-on-call) — metrics, alerts, and the other incident playbooks
+- [Indexer Stall (Playbook 3)](/guides/ops/monitoring-and-on-call#playbook-3-indexer-stall) — use instead when the scanner is still inside the window
+- [Error Codes](/reference/error-codes#soroban-rpc--indexer-errors) — `retention_window_exceeded`
+- [fetchAnnouncementsStream](/api-reference/fetch-announcements-stream) — `RetentionConfig`, `fallbackPolicy`, and checkpoint resumption
+- [Stellar Troubleshooting](/guides/stellar-troubleshooting) — general retention-window error handling