Bug Description
The health check endpoint at /health (lines 209-235 in src/server.ts) calls Stellar Horizon and Soroban RPC as part of the health check. If either Stellar service is down, the health check returns 503 even though the API itself is healthy.
Location
src/server.ts lines 209-235
The Problem
The health check performs 4 checks:
- Database (SELECT 1)
- Redis (PING)
- Stellar Horizon (root call)
- Soroban RPC (health check)
If any of these fail, the entire health check returns 503 "degraded". This means:
- A Stellar network outage makes the API appear "down" to load balancers
- Kubernetes/Docker health checks may restart healthy API instances
- Users cannot access non-Stellar features (course browsing, quiz taking) during Stellar outages
Recommended Fix
Separate the health check into:
- Liveness (
/health/live) - Is the process alive? (always 200)
- Readiness (
/health/ready) - Can it serve requests? (DB + Redis only)
- Full Health (
/health) - Detailed status including external dependencies
The readiness check should only verify core infrastructure (DB + Redis), not external blockchain services. Stellar-dependent endpoints can handle their own degradation.
Current State
Note: /health/live already exists and returns { status: "ok" }. However, /health/ready has the same 4-check structure as /health, making it equally fragile.
Acceptance Criteria
/health/live - Returns 200 if process is alive (already done)
/health/ready - Returns 200 if DB and Redis are healthy (remove Stellar checks)
/health - Keep as detailed status endpoint for monitoring dashboards
- Update documentation to explain the difference between readiness and liveness
Severity
medium - Can cause unnecessary container restarts and load balancer failures during Stellar outages.
Bug Description
The health check endpoint at
/health(lines 209-235 insrc/server.ts) calls Stellar Horizon and Soroban RPC as part of the health check. If either Stellar service is down, the health check returns 503 even though the API itself is healthy.Location
src/server.tslines 209-235The Problem
The health check performs 4 checks:
If any of these fail, the entire health check returns 503 "degraded". This means:
Recommended Fix
Separate the health check into:
/health/live) - Is the process alive? (always 200)/health/ready) - Can it serve requests? (DB + Redis only)/health) - Detailed status including external dependenciesThe readiness check should only verify core infrastructure (DB + Redis), not external blockchain services. Stellar-dependent endpoints can handle their own degradation.
Current State
Note:
/health/livealready exists and returns{ status: "ok" }. However,/health/readyhas the same 4-check structure as/health, making it equally fragile.Acceptance Criteria
/health/live- Returns 200 if process is alive (already done)/health/ready- Returns 200 if DB and Redis are healthy (remove Stellar checks)/health- Keep as detailed status endpoint for monitoring dashboardsSeverity
medium - Can cause unnecessary container restarts and load balancer failures during Stellar outages.