Skip to content

feat(health): add dependency-aware readiness and recovery diagnostics - #140

Open
woahwhattheheck wants to merge 1 commit into
RemitFlow:mainfrom
woahwhattheheck:latch/remitflow-134-readiness
Open

woahwhattheheck wants to merge 1 commit into
RemitFlow:mainfrom
woahwhattheheck:latch/remitflow-134-readiness

Conversation

@woahwhattheheck

Copy link
Copy Markdown

Summary

Closes #134.

Separates process liveness from traffic readiness and adds bounded, redacted dependency probes so a process can no longer report healthy while its store, payment provider, or FX dependency is down.

What changed

  • GET /api/health/live stays process-only and dependency-free so orchestrators do not restart instances during dependency outages.
  • GET /api/health/ready probes store, payments (Stellar), and fx under a per-check timeout (HEALTH_CHECK_TIMEOUT_MS, default 1000ms).
  • Failures return HTTP 503 with stable redacted reason codes (STORE_UNAVAILABLE, PAYMENTS_TIMEOUT, etc.). Raw messages, stacks, and connection material never appear in the payload.
  • Probes are re-evaluated on every request — clearing a dependency outage restores readiness without a process restart.
  • Light ping() hooks on rateService and stellarService; shared logic lives in dependencyHealthService.

Design tradeoffs

  • Bounded races, not cancellation: hung probes are raced against a timer. Node cannot abort an in-flight plain Promise, but readiness never waits longer than the budget, so the route cannot hang.
  • Redacted codes over messages: monitors get stable enums; operators keep logs for detail. Prefer this over leaking driver/SQL/auth strings on a public health path.
  • No sticky failure cache: recovering without restart means every ready check hits live probes. Cost is three cheap local checks (in-memory store + mock Stellar/FX); acceptable for this stack.
  • Compatibility: /api/health and /api/health/live payloads are unchanged. /api/health/ready checks values move from plain "ok" strings to { status, latencyMs, reason? } objects — callers that only checked top-level status continue to work; callers that deep-equalled checks.store === 'ok' need the new shape (covered in smoke).

Acceptance criteria mapping

Criterion Evidence
Liveness remains responsive during outages test/dependencyHealth.test.js — live stays 200 while ready is 503
Readiness accurately changes and recovers failure + clearForcedStates recovery tests; no restart
Dependency checks cannot hang timeout tests assert *_TIMEOUT and elapsed ≪ hang delay
Redacted reason codes throw-with-secret tests assert secrets absent from JSON
Status codes 200 when ready, 503 when any dependency fails
Regression for original failure mode payment/store/FX unavailable → not_ready

Test evidence

npm test
# 269 pass / 0 fail (includes new dependencyHealth suite + updated smoke)

Out of scope

  • No broad rewrite of transfer/user flows
  • No real external DB/Stellar clients (demo remains in-memory / mock)

Separate process liveness from traffic readiness. Ready probes now check
store, payments (Stellar), and FX under a per-dependency timeout, return
redacted reason codes on failure, and re-evaluate every request so recovery
does not require a restart. Liveness stays dependency-free so outages do not
flap orchestrator restarts.

Closes RemitFlow#134
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(backend): add dependency-aware readiness and recovery diagnostics

1 participant