Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
166 changes: 166 additions & 0 deletions backend/docs/EXCHANGE_RATE_ORACLE_CACHE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
# Exchange Rate Oracle Cache

Caching and concurrency control for path-payment exchange-rate quotes
(`GET /api/path-payment-quote/:id`).

Covers issues **#1445** (distributed concurrency control and locking) and
**#1446** (integration and stress test suite).

---

## 1. Layout

| File | Role |
|---|---|
| `src/lib/exchange-rate-cache.js` | In-process LRU + TTL cache with single-flight `getOrLoad()` |
| `src/lib/exchange-rate-coordinator.js` | Cross-instance coordination: Redis lock + shared quote store |
| `src/services/exchangeRateService.js` | `getExchangeRateQuote()` composes both layers around the Horizon query |
| `src/lib/path-payment-metrics.js` | Prometheus series (existing cache metrics + concurrency metrics) |

```
getExchangeRateQuote(key)
│
├─ L1: ExchangeRateCache.getOrLoad(key) per process
│ fresh hit ─────────────────────────────▶ return (cached: true)
│ load in flight ────────────────────────▶ join it (cached: true)
│ otherwise run ONE loader ↓
│
└─ L2: ExchangeRateCoordinator.load(key) across instances (if Redis)
shared quote present ──────────────────▶ return (cached: true)
SET exrate:lock:<key> NX PX ── won ─────▶ query Horizon, publish, release
└ lost ────▶ poll shared store / retry lock
until waitTimeoutMs → direct query
```

## 2. The problems this solves (#1445)

| Problem | Before | After |
|---|---|---|
| Thundering herd, one process | N concurrent misses → N Horizon calls | 1 call; the rest join it |
| Thundering herd, many instances | 1 call per instance | 1 call total (lock leader); peers reuse the shared quote |
| Invalidation race | a load started before `invalidateExchangeRateQuote()` wrote the old quote back afterwards | the detached load serves its own callers but never writes to the cache |
| Hung Horizon call | every waiter hung with it | loads time out (`EXCHANGE_RATE_LOAD_TIMEOUT_MS`, 504) and the slot is freed |
| Invalidation across instances | only the local process forgot the quote | the shared quote is deleted too |

## 3. In-process single-flight (`ExchangeRateCache.getOrLoad`)

- The first miss for a key registers an in-flight entry, then calls the
loader. Later callers await the same promise.
- The loader is invoked on a later microtask, so a synchronously throwing
loader still reaches the cleanup code.
- Errors reach every waiter and are **not** cached. The next call retries.
- `delete(key)` / `clear()` flag the in-flight entry as invalidated and
detach it. The next caller starts a fresh load, and the old load cannot
write back. Memory is bounded by the number of active loads, with no
per-key history.
- A stale-but-tolerable entry is refreshed through the same single-flight
path.

## 4. Distributed coordination (`ExchangeRateCoordinator`)

Enabled by `createApp()` whenever the Redis client is connected
(`configureExchangeRateCoordination`). Without Redis, only the in-process
layer runs.

**Lock.** `SET exrate:lock:<key> <uuid> PX <lockTtlMs> NX`. Release uses a
Lua compare-and-delete, so a holder whose lease expired can never delete the
next owner's lock. No fencing token is needed: the protected work (a
read-only Horizon query followed by an idempotent cache write) is safe to
duplicate in the rare lease-expiry case. That case is logged with a hint to
raise `EXCHANGE_RATE_LOCK_TTL_MS`.

**Shared store.** `exrate:quote:<key>` holds `{ v: 1, insertedAt, data }` with
`PX = sharedTtlMs`, so any present entry is fresh. Keys are SHA-256 hashes, so
client input never reaches Redis key names.

**Follower loop.** Each iteration reads the shared store (hit → done), then
tries the lock (won → become leader, double-check the store, load). If the
lock is still held, the follower sleeps `pollIntervalMs` and repeats. If the
leader fails, or crashes and its lease expires, a follower wins the lock on a
later iteration. After `waitTimeoutMs`, the follower queries Horizon directly.

**Failure policy: fail open.** Any Redis error, or a closed client, falls
back to a direct Horizon query and is counted in
`exchange_rate_coordination_fallbacks_total{reason="redis_error"}`. Loader
errors are never mistaken for Redis errors, so they are never retried by the
coordinator.

**Validation of shared data.** Entries read from Redis must parse, carry
`v: 1` and pass `isValidQuote` (string amounts matching
`^\d{1,19}(\.\d{1,7})?$` and an array `path`). Anything else counts as a miss
(`result="invalid"`) and is overwritten by the next leader. A corrupted or
tampered entry can never reach a payer as `send_max`.

### Configuration

| Variable | Default | Meaning |
|---|---|---|
| `EXCHANGE_RATE_CACHE_TTL_MS` | `30000` | L1 freshness and shared-quote lifetime |
| `EXCHANGE_RATE_LOCK_TTL_MS` | `5000` | Lock lease; keep above a typical Horizon call |
| `EXCHANGE_RATE_LOCK_WAIT_MS` | `2000` | Max follower wait before querying directly |
| `EXCHANGE_RATE_LOCK_POLL_MS` | `50` | Follower poll interval |
| `EXCHANGE_RATE_LOAD_TIMEOUT_MS` | `15000` | Upper bound on one load; waiters get 504 |

## 5. Metrics

Added to the path-payment registry, which `/metrics` already serves:

| Metric | Type | Labels |
|---|---|---|
| `exchange_rate_cache_coalesced_requests_total` | counter | `cache` |
| `exchange_rate_cache_inflight_loads` | gauge | `cache` |
| `exchange_rate_cache_load_timeouts_total` | counter | `cache` |
| `exchange_rate_cache_stale_writes_prevented_total` | counter | `cache` |
| `exchange_rate_lock_acquisitions_total` | counter | `result` (acquired/contended/error) |
| `exchange_rate_lock_wait_seconds` | histogram | `outcome` (shared_hit/acquired/timeout/error) |
| `exchange_rate_shared_cache_lookups_total` | counter | `result` (hit/miss/invalid/error) |
| `exchange_rate_coordination_fallbacks_total` | counter | `reason` (wait_timeout/redis_error) |

`path_payment_quote_cache_size` is now updated on every write and delete,
not only on `prune()`.

## 6. Security notes

- Redis is treated as trusted infrastructure, but its content is still
validated before use (see above).
- Lock tokens are random UUIDs, and release is owner-checked atomically.
- Fail-open is deliberate: the quote is public DEX data and coordination is
only an optimization. An attacker who can break Redis gains nothing beyond
the pre-#1445 behavior of one Horizon query per request.
- The existing per-IP rate limit on the quote endpoint still applies.
Coalescing reduces the Horizon load an attacker can cause by bursting
identical requests.
- Unchanged from before: the cache key covers asset pair, amount and issuers,
not `slippage` or `source_account`. The route always uses the default
slippage.

## 7. Tests (#1446)

| Suite | Scope |
|---|---|
| `src/lib/exchange-rate-cache.test.js` | LRU/TTL plus single-flight, invalidation races, sync-throwing loaders, timeouts, gauges |
| `src/lib/exchange-rate-coordinator.test.js` | lock ownership and expiry, leader/follower, leader failure and crash takeover, wait-timeout fallback, fail-open, poisoned shared entries |
| `src/services/exchangeRateService.test.js` | service wiring, burst coalescing, `NoPathFoundError` fan-out, shared reuse, cross-instance invalidation, Redis outage |
| `tests/integration/exchange-rate-cache.test.js` | real HTTP stack on `GET /api/path-payment-quote/:id`: 40-request burst → 1 Horizon call, distinct pairs, 404/502 fan-out without caching, 504 on hung Horizon, invalidation mid-flight, `/metrics`, Redis coordination, poisoned entry, Redis down |
| `load-tests/exchange-rate-cache-stress.test.js` | 5k concurrent → 1 call; 20k over 250 keys → 250 calls; LRU churn; invalidation storm; failure isolation; 8 simulated instances sharing Redis → 1 call per key; cold-instance reuse; lock-holder crash; Redis outage mid-burst; stuck lock; leak check |

`tests/helpers/fake-redis.js` is an in-memory Redis (SET NX/PX, GET, DEL,
compare-and-delete EVAL, TTL expiry, injectable latency and failures). Many
coordinators can share one instance to simulate a scaled deployment
deterministically, without a live Redis.

HTTP bursts wait until every request has joined the in-flight load, using
`exchange_rate_cache_coalesced_requests_total`, before releasing the stubbed
Horizon response. supertest opens a separate server per request, so arrival
order is otherwise not guaranteed.

The suites were mutation-checked against `exchange-rate-cache.js`. Disabling
single-flight fails 16 tests. Dropping the invalidation guard fails 4.

```
npx vitest run src/lib/exchange-rate-cache.test.js \
src/lib/exchange-rate-coordinator.test.js \
src/services/exchangeRateService.test.js \
tests/integration/exchange-rate-cache.test.js
npm run test:load -- exchange-rate-cache-stress
```
197 changes: 197 additions & 0 deletions backend/docs/PAYMENT_SESSION_VALIDATOR.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,197 @@
# Payment Session Validator

Validation stage that runs before a payment session is persisted, used by both
`POST /api/create-payment` / `POST /api/sessions` (`src/routes/payments.js`)
and `paymentService.createPaymentSession` (`src/services/paymentService.js`).

Covers issues **#1447** (payload sanitization and strict validation) and
**#1448** (Prometheus alert metrics and health telemetry). Builds on the shared
rules from #1087 (see `PAYMENT_PROCESSOR.md`).

---

## 1. Layout

| File | Role |
|---|---|
| `src/lib/payment-session-rules.js` | Pure rule functions: no I/O, logging or metrics |
| `src/lib/payment-session-validator.js` | `validatePaymentSession()`: runs the rules in order, records metrics, logs, tracks health |
| `src/lib/payment-session-validator-metrics.js` | Separate Prometheus registry, merged into `/metrics` |
| `docs/alerts/payment-session-validator.rules.yml` | Prometheus alerting rules |

Request pipeline for the HTTP routes:

```
zod schema (paymentSessionZodSchema) type coercion, required fields, memo rules
sanitizeMetadataMiddleware metadata XSS/NoSQL scrubbing, forbidden keys → 400
validatePaymentSession({ source: "http" })
sanitization → payload → issuer → limits → allowlist first failure → 400
insert into payments uses the *sanitized* payload
```

In the service layer, the validator now runs **before** the on-chain issuer
lookup (`AssetIssuerErrorRecovery.verifyIssuerOnChain`). An invalid request
never costs a Horizon round-trip.

## 2. Rules

| Rule | Function | Rejection reasons |
|---|---|---|
| `sanitization` | `sanitizeSessionPayload` | `malformed_payload`, `forbidden_key`, `field_too_long` |
| `payload` | `validateSessionAsset`, `validateSessionAmount` | `invalid_asset`, `invalid_amount` |
| `issuer` | `resolveAndValidateIssuer` | `missing_issuer`, `invalid_issuer` |
| `limits` | `validatePerAssetLimits` | `below_min`, `above_max` (response includes `min`/`max` + `delta`) |
| `allowlist` | `validateAllowedIssuers` | `issuer_not_allowed` |

### Sanitization (#1447)

- The body must be a plain JSON object. Arrays, primitives and class
instances are rejected.
- `__proto__`, `constructor` and `prototype` keys are rejected at any depth
(walk capped at 8 levels). `JSON.parse` creates `__proto__` as an own
property, so the walk sees it.
- For the string fields `asset`, `asset_issuer`, `recipient`, `description`,
`message`, `memo`, `memo_type`, `webhook_url` and `client_id`, each value is:
- NFC-normalized
- stripped of C0/C1 control characters, bidi overrides/isolates
(U+202A–U+202E, U+2066–U+2069, U+200E/F) and zero-width characters
- trimmed
- length-checked **after** stripping (limits in `SESSION_FIELD_MAX_LENGTHS`)
- An optional field that sanitizes to `""` is dropped. A required field
(`asset`, `recipient`) stays `""` and fails later with an explicit error.
- The input is never mutated. Callers persist `validation.payload`.

### Strict validation (#1447)

- `asset` must match `^[A-Z0-9]{1,12}$` after upper-casing.
- `amount` must be a finite number > 0, ≤ `922337203685.4775807` (int64
stroops) and have at most 7 decimal places. Floating-point artifacts such as
`0.1 + 0.2` are rejected.

### Hardening of existing rules

- Limit lookup uses `Object.hasOwn`, so asset codes like `CONSTRUCTOR` or
`__proto__` can never resolve to inherited properties.
- Limit bounds are coerced with `Number()`, so legacy numeric-string configs
such as `"10"` keep working. Unusable bounds (`"abc"`, objects) are
**ignored** and reported as config anomalies instead of silently comparing
as `NaN`.
- Allowlist entries that are not strings are ignored rather than coerced.

### Related fixes

- `src/lib/request-schemas.js` called `isValidStellarPublicKey` without
importing it. Every non-XLM session with an issuer hit a `ReferenceError`
inside schema validation.
- `src/lib/sanitize-metadata.js` copied keys into a fresh object with
`sanitized[key] = …`. A `__proto__` key therefore replaced that object's
prototype. Metadata containing forbidden keys is now rejected with 400.

## 3. Metrics (#1448)

Every label value comes from a fixed server-side set. None comes from request
data, so clients cannot inflate cardinality. Unknown `source` values collapse
to `unknown`.

| Metric | Type | Labels |
|---|---|---|
| `payment_session_validator_evaluations_total` | counter | `source` (http/service/unknown), `outcome` (accepted/rejected/error) |
| `payment_session_validator_rejections_total` | counter | `source`, `rule`, `reason` |
| `payment_session_validator_duration_seconds` | histogram | `source`, `outcome` |
| `payment_session_validator_sanitized_fields_total` | counter | `field` |
| `payment_session_validator_suspicious_payloads_total` | counter | `signal` (forbidden_key/malformed_payload/bidi_control/oversized_field) |
| `payment_session_validator_config_anomalies_total` | counter | `kind` (invalid_entry/invalid_min/invalid_max/min_greater_than_max) |
| `payment_session_validator_health_state` | gauge | 0 healthy, 1 degraded, 2 unhealthy |
| `payment_session_validator_rejection_ratio` | gauge | rolling window |
| `payment_session_validator_error_ratio` | gauge | rolling window |
| `payment_session_validator_last_evaluation_timestamp_seconds` | gauge | none |

The rolling-window gauges are recomputed at scrape time, so they decay back to
healthy when traffic stops.

## 4. Health telemetry (#1448)

`GET /health/payment-session-validator` (public, no merchant data):

```json
{
"status": "healthy",
"reasons": [],
"window_ms": 300000,
"total": 42, "accepted": 40, "rejected": 2, "errors": 0, "suspicious": 0,
"rejection_ratio": 0.0476, "error_ratio": 0,
"last_evaluation_at": "2026-09-26T12:00:00.000Z",
"thresholds": { "min_samples": 20, "error_ratio": 0.05, "rejection_ratio": 0.5, "suspicious": 10 }
}
```

| Status | Condition | HTTP |
|---|---|---|
| `unhealthy` | ≥ `min_samples` evaluations and error ratio ≥ threshold | 503 |
| `degraded` | ≥ `min_samples` and rejection ratio ≥ threshold, **or** suspicious count ≥ threshold | 200 |
| `healthy` | otherwise (including no traffic) | 200 |

`GET /health` also reports `services.payment_session_validator`. The value is
informational only and does not change `ok` or the status code.

The window uses 5-second buckets in a fixed ring, so memory stays constant
under any load. Tuning via environment:

| Variable | Default |
|---|---|
| `PAYMENT_SESSION_VALIDATOR_HEALTH_WINDOW_MS` | `300000` |
| `PAYMENT_SESSION_VALIDATOR_HEALTH_MIN_SAMPLES` | `20` |
| `PAYMENT_SESSION_VALIDATOR_ERROR_RATIO_THRESHOLD` | `0.05` |
| `PAYMENT_SESSION_VALIDATOR_REJECTION_RATIO_THRESHOLD` | `0.5` |
| `PAYMENT_SESSION_VALIDATOR_SUSPICIOUS_THRESHOLD` | `10` |

## 5. Alerts

`docs/alerts/payment-session-validator.rules.yml` defines:

| Alert | Severity | Fires when |
|---|---|---|
| `PaymentSessionValidatorInternalErrors` | critical | any `outcome="error"` in 5m |
| `PaymentSessionValidatorUnhealthy` | critical | `health_state >= 2` for 5m |
| `PaymentSessionValidatorHighRejectionRatio` | warning | > 50% rejected over 10m (≥ 20 samples) |
| `PaymentSessionValidatorSuspiciousPayloadSpike` | warning | > 10 suspicious payloads in 5m |
| `PaymentSessionValidatorPrototypePollutionAttempt` | warning | any `forbidden_key` in 15m |
| `PaymentSessionValidatorMerchantConfigAnomaly` | info | any config anomaly in 30m |
| `PaymentSessionValidatorSlow` | warning | p99 > 25ms for 10m |

A unit test checks that every metric referenced in the rules file exists in
the registry.

## 6. Logging

- Rejection: `info` with `{ merchantId, source, rule, reason }`
- Suspicious payload: `warn` with `{ merchantId, source, signals }`
- Config anomaly: `warn` with `{ merchantId, source, anomalies }`
- Internal error: `error` with `{ err, merchantId, source }`, then rethrown

The raw payload is never logged.

## 7. Security notes

- **Fail-closed on validator errors:** unexpected exceptions are recorded and
rethrown, so the route returns 500 and no session is created.
- **Config anomalies fail open, but only per bound:** a malformed bound is
ignored, and the other bound and all other rules still apply. The choice
favors availability for merchants with bad configs. It is visible through
the `MerchantConfigAnomaly` alert.
- **Trojan Source / display spoofing:** bidi and zero-width characters are
stripped before the text reaches the database or the hosted checkout page.
- **DoS:** limits and allowlist checks now run before the Horizon issuer
lookup, and the forbidden-key walk is depth-capped.
- **Behavior changes clients may notice:** amounts with more than 7 decimals,
amounts above the Stellar maximum, `description` > 1000 chars and
`client_id` > 128 chars are now rejected with 400.

## 8. Tests

```
npx vitest run src/lib/payment-session-rules.test.js \
src/lib/payment-session-validator.test.js \
src/lib/sanitize-metadata.test.js \
tests/integration/payment-session-validator.test.js
```
Loading
Loading