diff --git a/docs/code-style.md b/docs/code-style.md
new file mode 100644
index 0000000..755e400
--- /dev/null
+++ b/docs/code-style.md
@@ -0,0 +1,227 @@
+# Code Style
+
+This page covers:
+
+- the formatters and linters in force for each language
+- how to run them locally
+- which ones CI blocks on
+- the conventions no tool enforces, which reviewers are expected to check
+
+There are no pre-commit hooks (no husky, lint-staged or pre-commit). Run
+the commands below before pushing.
+
+## Summary
+
+| Language | Tool | Config | Command | CI gate |
+| ----------------- | ------------------------------- | ------------------------------------------------------------------------------------------------ | -------------------------------------------------------- | -------------------------------------------------------------------- |
+| TS, TSX, MD, JSON | Prettier `3.9.1` (pinned) | [`.prettierrc.json`](../.prettierrc.json), [`.prettierignore`](../.prettierignore) | `pnpm format:check` / `pnpm format` (repo root) | **Partial.** Only `apps/backend/src/**/*.ts` (Backend CI). |
+| TS (backend) | ESLint 9 + `typescript-eslint` | [`apps/backend/eslint.config.js`](../apps/backend/eslint.config.js) | `pnpm --filter backend lint` | **Errors only** (Backend CI, PR Check) |
+| TS/TSX (web) | ESLint 9 + `eslint-config-next` | [`apps/web/eslint.config.mjs`](../apps/web/eslint.config.mjs) | `pnpm --filter web lint` | **Errors only** (Frontend CI, PR Check) |
+| TS type check | `tsc` (`strict: true`) | `apps/*/tsconfig.json` | `pnpm --filter backend build`; `pnpm --filter web build` | **Web only.** `next build` type-checks; Backend CI never runs `tsc`. |
+| Python | ruff `0.15.20` (lint + format) | [`apps/ai_agent/pyproject.toml`](../apps/ai_agent/pyproject.toml) `[tool.ruff]` | `uv run ruff check .` / `uv run ruff format --check .` | **Yes** (AI Agent CI) |
+| Python | mypy `2.1.0` | `pyproject.toml` `[tool.mypy]` | `uv run mypy main.py` | **Yes**, `main.py` only (AI Agent CI) |
+| Rust | rustfmt (default style) | none; `rustfmt` component in [`contracts/rust-toolchain.toml`](../contracts/rust-toolchain.toml) | `cargo fmt --check` / `cargo fmt` (in `contracts/`) | **No** |
+| Rust | clippy, `-D warnings` | [`contracts/Cargo.toml`](../contracts/Cargo.toml) `[workspace.lints]` + CI flags | see [Rust](#rust-rustfmt--clippy) | **Yes**, zero warnings (Contracts CI) |
+
+Python versions are pinned by `apps/ai_agent/uv.lock`. Prettier is pinned
+exactly in both the root and backend `package.json`. The Rust toolchain is
+**not** pinned: it tracks `stable`, both locally and in CI.
+
+## Prettier
+
+- **Pinned** to `3.9.1` (exact, no caret) in the root `package.json` and in
+ `apps/backend/package.json`. A Prettier upgrade reformats code, so it gets a
+ PR of its own and never rides along with other changes.
+- **One shared config** at the repo root, `.prettierrc.json`, which every
+ package inherits: `semi`, `singleQuote`, `trailingComma: all`,
+ `printWidth: 100`, `tabWidth: 2`. Packages must not add their own Prettier
+ config.
+- **Commands:**
+ - `pnpm format:check` and `pnpm format` at the root cover
+ `**/*.{ts,tsx,md,json}`.
+ - `pnpm --filter backend format:check` covers only `apps/backend/src/**/*.ts`.
+ This is the only Prettier check CI runs.
+- **CI gap:** web `.tsx`, Markdown and JSON are not format-checked in CI.
+ Run the root `pnpm format:check` before opening a PR that touches them.
+
+### `.prettierignore`: generated output is never formatted
+
+**Prettier formats source that people edit. It never formats files a tool
+generates.** Reformatting generated files produces noisy diffs that hide real
+changes, and the next build or test run rewrites them anyway. Each ignore entry
+follows from that rule:
+
+| Entry | Why it is ignored |
+| --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
+| `node_modules/`, `dist/`, `.next/`, `.turbo/`, `coverage/`, `target/` | Dependency and build output |
+| `apps/backend/drizzle/meta/` | drizzle-kit migration snapshots and `_journal.json`, rewritten by `db:generate` |
+| `**/*.d.ts` | TypeScript declaration output (for example `drizzle.config.d.ts`) |
+| `apps/web/next-env.d.ts` | Next.js ambient types, regenerated by `next dev` and `next build` |
+| `contracts/**/test_snapshots/` | Soroban test snapshots, regenerated by `cargo test` (also gitignored in `contracts/.gitignore`) |
+| `contracts/target/`, `apps/ai_agent/.venv/` | Rust and Python build output |
+
+When you add a tool that writes files into the tree, add its output to
+`.prettierignore` in the same PR, under the "Generated artifacts" heading and
+with a comment naming the tool. The generated SQL files in
+`apps/backend/drizzle/*.sql` are not listed because the root glob does not
+match `.sql`. Do not hand-edit them either.
+
+## ESLint (TypeScript / TSX)
+
+**Backend** ([`eslint.config.js`](../apps/backend/eslint.config.js)) uses the
+flat config: `@eslint/js` recommended plus `typescript-eslint` recommended,
+with these overrides:
+
+- `@typescript-eslint/no-unused-vars`: **error**. Parameters prefixed with `_`
+ are exempt, so name deliberately unused arguments `_req`, `_next` and so on.
+- `@typescript-eslint/no-explicit-any`: **warn**.
+- Node globals (`process`, `Buffer`, `NodeJS`, …) are declared explicitly,
+ because `no-undef` cannot see `@types/node`.
+
+**Web** ([`eslint.config.mjs`](../apps/web/eslint.config.mjs)) uses
+`eslint-config-next` (`core-web-vitals` + `typescript`) with the default
+Next.js ignores.
+
+**Commands:** `pnpm --filter backend lint` and `pnpm --filter web lint` (add
+`lint:fix` to auto-fix). `pnpm lint` at the root runs both through Turborepo.
+
+**CI:** ESLint exits non-zero only on **errors**, so errors block the build and
+warnings do not. The linters run in Backend CI, Frontend CI, and in `pr.yml`
+(`pnpm run lint`) on every PR.
+
+## Python: ruff + mypy
+
+Config lives in [`apps/ai_agent/pyproject.toml`](../apps/ai_agent/pyproject.toml):
+
+- **ruff:** `target-version = "py312"`, `line-length = 100` (the same as
+ Prettier's `printWidth`), and lint rules `E`, `F`, `I` (import sorting) and
+ `W`. `ruff format` is the formatter; there is no black or isort.
+- **mypy:** `python_version = "3.12"`, `warn_return_any`,
+ `warn_unused_configs`, `ignore_missing_imports`.
+
+Run from `apps/ai_agent/`:
+
+```bash
+uv sync --group dev
+uv run ruff check . # lint (add --fix for safe fixes)
+uv run ruff format --check . # format check (drop --check to apply)
+uv run mypy main.py # type check
+```
+
+**CI:** all three block the build (AI Agent CI). mypy only checks `main.py`,
+so `tests/` is not type-checked.
+
+## Rust: rustfmt + clippy
+
+Run from `contracts/`:
+
+```bash
+cargo fmt --check # format check (drop --check to apply)
+cargo clippy --workspace --target wasm32-unknown-unknown -- \
+ -D warnings -A dead_code -A clippy::too-many-arguments # exactly what CI runs
+```
+
+- **rustfmt** uses the default style (there is no `rustfmt.toml`). It is **not
+ checked in CI**. Run `cargo fmt` before pushing.
+- **clippy** runs with **`-D warnings`**, so any warning fails Contracts CI.
+ Lint against the `wasm32-unknown-unknown` target, because that is how the
+ contracts are built and warnings can differ from a host build.
+- **The gate is looser than `-D warnings` suggests.** `[workspace.lints]` in
+ `contracts/Cargo.toml` sets `dead_code`, `unused_variables`, `unused_mut`,
+ `unused_imports` and `clippy::too_many_arguments` to `allow`. The CI command
+ allows `dead_code` and `too_many_arguments` again. None of those ever fail
+ the build.
+- **The toolchain is `stable` and unpinned** (`rust-toolchain.toml` and
+ `dtolnay/rust-toolchain@stable`). A new Rust release can add clippy lints
+ that fail CI without any code change. Fix those in a dedicated PR.
+- `cargo audit` also blocks the build in Contracts CI.
+
+## Warning counts: non-blocking, but must not grow
+
+Measured on 2026-09-29 against code identical to `main` at `23eaf29`.
+clippy could not be run on the measuring machine; its zero comes from the CI
+gate.
+
+| Check | Count | Blocking? |
+| --------------------------------------- | ----------------------------------------------------------------- | -------------- |
+| Backend ESLint | **58 warnings**, all `@typescript-eslint/no-explicit-any` | No |
+| Web ESLint | **4 warnings**, all `@typescript-eslint/no-unused-vars` | No |
+| Backend Prettier (`src/**/*.ts`) | 0 | Yes |
+| Root Prettier (`**/*.{ts,tsx,md,json}`) | 1 file (`apps/ai_agent/docs/concepts-rag-search-architecture.md`) | No (not in CI) |
+| ruff check / ruff format | 0 / 0 | Yes |
+| rustfmt | 0 diffs | No (not in CI) |
+| clippy | 0 (enforced by `-D warnings`) | Yes |
+
+These warnings are tolerated, not accepted:
+
+- **A PR must not increase any count.** A new `any` or unused variable gets
+ fixed in the PR that introduces it. Reviewers compare the lint output
+ against this table.
+- **Touching a file is a good time to fix its warnings**, but do it in a
+ separate commit so the behaviour change stays reviewable.
+- When a count reaches zero, lock it in: pass `--max-warnings 0` to that
+ package's `lint` script, or raise the rule to `error`. Then update this
+ table.
+
+## Conventions lint cannot enforce
+
+### Comments explain _why_ and cite the issue behind them
+
+This is the house style throughout `apps/backend`: 56 of the 83 non-test
+source files cite an issue in their comments. Every comment worth writing
+does two things:
+
+1. **Explains why, not what.** The code already says what it does. The
+ comment records the constraint, threat or failure that makes it look this
+ way, the thing a reader could not work out from the code.
+2. **References the issue number that motivated the code**, so the reader can
+ find the full discussion.
+
+It takes one of two forms: a leading `#N —` or a trailing `(#N)`.
+
+```ts
+// #330 — dev/test only: serves the fs-backed object store so presigned URLs
+// issued locally are real, working URLs. Deliberately outside requireAuth
+// (a presigned URL carries its own HMAC + expiry, exactly like S3), and never
+// mounted in production, where the real object store answers these requests.
+```
+
+```ts
+// Refuse to boot a non-dev gateway that would accept plaintext transport (#374).
+assertTransportSecurityConfig();
+```
+
+Module-level JSDoc headers follow the same pattern: a title with the issue
+number, then the design rationale. See `lib/transportSecurity.ts` (#374) and
+`config/rateLimits.ts` (#375). Rust follows it too, for example the `///` on
+`token_transfer::upgrade` (#44).
+
+Things reviewers push back on:
+
+- A comment that restates the code, such as `// increment counter`.
+- A bare issue number with no explanation. The issue tracker may not outlive
+ the code, so the comment must make sense on its own.
+- An explanation with no issue number, when an issue exists.
+- A comment that no longer matches the code. Update the comment in the same
+ PR that changes the behaviour.
+
+### Other conventions
+
+- **Section banners.** Long files are split into sections with box-drawing
+ banners, for example `// ── Socket events ───` in TypeScript and
+ `# ── Helpers ───` in Python. Use the same style in new long files.
+- **Environment reads take a `source` parameter** that defaults to
+ `process.env`, parse with a safe fallback, and never throw on a bad value.
+ Examples are `getRateLimitRule` and `isTlsEnforced`. This keeps them
+ testable without mutating the global environment. Anything that must crash
+ on start belongs in `config.ts`. See
+ [environment-variables.md](environment-variables.md).
+- **Never log message content.** Log IDs, counts and durations only. The
+ pino redaction in `lib/logger.ts` is a backstop, not permission to log
+ payloads.
+- **The backend uses NodeNext import specifiers.** Relative imports end in
+ `.js` (`'./lib/redis.js'`) even though the source file is `.ts`. `tsc`
+ enforces this, but Backend CI does not run `tsc`, so run
+ `pnpm --filter backend build` locally.
+- **Commit messages** follow Conventional Commits (`feat:`, `fix:`, `docs:`,
+ optionally scoped, like `fix(major): …`), as the README shows.
diff --git a/docs/data-retention.md b/docs/data-retention.md
new file mode 100644
index 0000000..b341682
--- /dev/null
+++ b/docs/data-retention.md
@@ -0,0 +1,306 @@
+# Data Retention
+
+For every kind of data the system stores, this page gives:
+
+- what is kept, and whether the server can read it
+- how long it is kept
+- what deletes it
+- what a user-initiated erasure can and cannot reach
+
+It describes the code as it is, not as intended. Where the two differ, the gap is
+called out; see [Known gaps](#known-gaps) for the full list.
+
+Stores covered: Postgres (backend), Redis (backend), the S3-compatible object
+store, backend process memory, and Weaviate (AI agent). Retention windows that
+can be tuned are listed with their environment variable; defaults come from
+[environment-variables.md](environment-variables.md).
+
+## Content versus metadata
+
+The server stores two very different kinds of data. See
+[threat-model.md](threat-model.md) for the full visibility analysis.
+
+- **Content is end-to-end encrypted and the server cannot read it.** This
+ covers message ciphertext (per-device envelopes, and MLS group ciphertext on
+ `messages`), file blobs in the object store, and opaque MLS commit and
+ Welcome material. It is encrypted on the client. The keys never reach the
+ server: file keys travel inside envelopes, and private keys and ratchet state
+ stay on the client
+ ([threat-model → What the server cannot see](threat-model.md#what-the-server-cannot-see)).
+ Deleting content is still worth doing, because it limits what an attacker who
+ later gets a client key could decrypt.
+- **Metadata is visible to the server.** This covers who is in which
+ conversation, who sent what to whom and when, sizes, content types, device
+ names and platforms, presence, IP addresses and user agents in audit logs,
+ push endpoints, and wallet addresses. **Most metadata outlives the content it
+ describes.** Deleting a message removes its ciphertext, but the row recording
+ that the message existed stays
+ ([threat-model → Residual metadata risk](threat-model.md#residual-metadata-risk)).
+- **Exceptions, where the server can read content:**
+ - **AI assistant prompts and replies.** Prompts are sent to the backend as
+ plaintext and forwarded to the AI agent and OpenAI. Replies are stored in
+ `messages.ciphertext` as **plaintext**, because the server writes them.
+ - **System messages** (`contentType = 'system'`), which carry a plaintext
+ `systemPayload`.
+ - **Weaviate `Message` objects** in the AI agent, which hold plaintext
+ `content` (see [AI agent](#ai-agent-weaviate)).
+
+## Retention table
+
+Legend: **C** is content (encrypted), **M** is metadata (visible), **P** is
+plaintext content that the server can read.
+
+| Data | Store | Kind | Retention | Removed by |
+| ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
+| **Messages and envelopes** | | | | |
+| Message row: sender, conversation, timestamps, `contentType`, `fileId`, `mlsEpoch` | Postgres `messages` | M | **Indefinite**, even after the message is deleted | Only a cascade when the conversation row is deleted |
+| Message ciphertext (MLS group messages) | Postgres `messages.ciphertext` | C | Until the sender deletes the message | Sender deletes the message (sets the column to `NULL`) |
+| Per-device envelope ciphertext | Postgres `message_envelopes` | C | **7 days after delivery**, or **30 days after creation** if never delivered | Envelope GC. Sender deletion removes them immediately. |
+| AI assistant replies | Postgres `messages.ciphertext` + envelopes | **P** | Row: **indefinite**. Envelopes: as above. | Nothing (see [cannot reach](#what-erasure-cannot-reach)) |
+| System messages | Postgres `messages.systemPayload` | M | Indefinite | Conversation cascade only |
+| **Files** | | | | |
+| File blob (encrypted) | Object store | C | While any live message references it | File GC hard-deletes it once every referencing message is deleted and the grace period has passed |
+| File row: uploader, conversation, size, MIME type, sha256, storage key | Postgres `files` | M | **Indefinite**. The row is kept with `hardDeletedAt` set after the blob is gone. | Conversation cascade only |
+| Unconfirmed upload (`status = 'pending'`) | Object store + `files` | C + M | **24 hours** | File GC deletes both the blob and the row |
+| **Devices and keys** | | | | |
+| Device row: identity public key, name, platform, `lastSeenAt`, capabilities | Postgres `devices` | M | **Indefinite**. Revocation sets `revokedAt`; after 180 days the GC sets `staleFlaggedAt`. | Nothing. Rows are never hard-deleted. |
+| One-time prekeys | Postgres `device_prekeys` | M (public keys) | Consumed: **30 days** after creation. Unconsumed: **90 days** after creation. | Device GC. Device revocation deletes them immediately. |
+| Signed prekey | Postgres `device_prekeys` | M (public key) | Until replaced by the next upload | Replaced on upload. Device revocation deletes it. |
+| MLS KeyPackages | Postgres `mls_key_packages` | M (public) | Same windows as one-time prekeys | Device GC. **Not** removed by revocation. |
+| Identity-key change log | Postgres `device_key_history` | M | **Indefinite, by design.** It exists to detect silent key swaps. | Nothing |
+| **Conversations and groups** | | | | |
+| Conversation and membership | Postgres `conversations`, `conversation_members` | M | Membership: until the user leaves. Group: until its last member leaves. **DMs: indefinite.** | Leaving a group. The last member leaving deletes the conversation and cascades. |
+| MLS group state, commits, Welcomes; group control log | Postgres `mls_*`, `group_control_events` | C (opaque payloads) + M | Indefinite | Conversation cascade only |
+| **Accounts** | | | | |
+| User profile and privacy settings | Postgres `users` | M | **Indefinite** | Nothing. There is no account-deletion path. |
+| Wallet addresses | Postgres `wallets` | M | Indefinite | Nothing |
+| **Audit** | | | | |
+| Security audit events: actor, subject, target, IP, user agent, metadata | Postgres `audit_logs` | M | **Indefinite. Append-only by design.** | Only a deliberate, privileged manual prune (see [below](#audit-logs)) |
+| **Payments and governance** | | | | |
+| Token transfers (mirrors on-chain `transfer` events) | Postgres `token_transfers` | M | Indefinite | Conversation cascade (the chain copy is permanent) |
+| Treasury proposals and votes | Postgres `treasury_proposals`, `proposal_votes` | M | Indefinite | Nothing (the chain copy is permanent) |
+| On-chain transactions, balances, proposals, votes | Stellar ledger | M (public) | **Permanent** | Nothing, ever |
+| **Push** | | | | |
+| Push subscription: endpoint, `p256dh`, `auth` | Postgres `push_subscriptions` | M | Until unsubscribed or the push service reports it dead | `DELETE /push/subscriptions`, or automatic pruning on HTTP 410/404. **Not** removed by device revocation. |
+| **Redis (ephemeral)** | | | | |
+| Presence: `presence:user:*`, `presence:sockets:*`, `presence:device_sockets*`, `presence:socket:*` | Redis | M | Per-device and socket keys: **90 s TTL**, refreshed by heartbeats. The `presence:user:{userId}` hash has no TTL; its entry is removed when the device goes offline, and the key is deleted with the last device. | TTL expiry, disconnect or offline handling |
+| Resume stream (`resume:events:{userId}`): typing, presence and other ephemeral events for reconnect replay | Redis stream | M | **300 s** after the last event, capped at about 500 entries | TTL + `XADD MAXLEN ~ 500` |
+| Replay-protection markers (`replay:{deviceId}:{eventId}`) | Redis | M | **300 s** (`REPLAY_PROTECTION_TTL_SECONDS`) | TTL |
+| Rate-limit counters (`rl:*`), abuse counters (`abuse:*`) | Redis | M | Bucket window; `abuse:*` 1 h | TTL |
+| Prekeys-low latch | Redis | M | 30 days | TTL. Released on device revocation. |
+| Conversation list cache | Redis | M | 30 s | TTL, invalidated on writes |
+| **Process memory** | | | | |
+| Auth and device-link nonces | Backend memory | M | 5 min / 2 min, single use | Consumed on read, or lost on restart |
+| **AI agent** | | | | |
+| Indexed messages and embeddings | Weaviate `Message` | **P** | **Indefinite** | Nothing. There is no delete endpoint. |
+| Prompts and replies | OpenAI | **P** | Governed by OpenAI's API data policy | Outside this system |
+
+Out of scope, owned by the operator: application logs (stdout/pino), database
+and object-store backups, and metrics. Their retention is whatever the
+deployment configures. **Backups keep erased data until they expire.**
+
+## Garbage-collection jobs
+
+Three in-process jobs enforce the windows above. Every backend instance starts
+them from `apps/backend/src/index.ts` with `setInterval(...).unref()`.
+
+- **The first pass runs one interval after boot, not at boot.**
+- On a multi-instance deployment, every instance runs every job. The passes
+ are plain idempotent `DELETE … WHERE` / `UPDATE … WHERE` statements, so
+ concurrent runs are safe.
+- Failures are logged (`[device-gc]`, `[envelope-gc]`, `[file-cleanup]`) and
+ retried on the next tick.
+
+| Job | Source | Schedule (default) | What each pass does | Windows |
+| ---------------- | ------------------------- | -------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
+| **Device GC** | `services/deviceGc.ts` | Every **1 h** (`DEVICE_GC_INTERVAL_MS`) | 1. Deletes one-time prekeys past their window.
2. Deletes MLS KeyPackages past the same window.
3. Sets `staleFlaggedAt` on long-revoked devices (flag only, no delete). | Consumed 30 d (`PREKEY_CONSUMED_RETENTION_DAYS`)
Unconsumed 90 d (`PREKEY_UNCONSUMED_MAX_AGE_DAYS`)
Stale 180 d (`DEVICE_STALE_AFTER_DAYS`) |
+| **Envelope GC** | `services/envelopeGc.ts` | Every **30 min** (`ENVELOPE_GC_INTERVAL_MS`) | Deletes envelopes that were delivered longer ago than the delivered window, **or** created longer ago than the max-age window. | Delivered 7 d (`ENVELOPE_DELIVERED_RETENTION_DAYS`)
Max age 30 d (`ENVELOPE_MAX_AGE_DAYS`) |
+| **File cleanup** | `services/fileCleanup.ts` | Every **5 min** (`FILE_GC_INTERVAL_MS`) | 1. Hard-deletes soft-deleted blobs that have no live message reference, then sets `hardDeletedAt`.
2. Deletes pending uploads (blob and row) past their TTL.
3. Re-enables push subscriptions whose 5-minute back-off has expired. | Grace 0 (`FILE_HARD_DELETE_GRACE_MS`)
Pending 24 h (`PENDING_UPLOAD_TTL_MS`) |
+
+Details that affect the windows:
+
+- **Prekey and KeyPackage age is measured from `createdAt`**, not from
+ consumption. `device_prekeys` has no `consumedAt` column, so a key consumed
+ on day 29 is deleted on day 30.
+- **The envelope max-age applies to undelivered envelopes.** A device that
+ stays offline for more than 30 days loses those messages permanently. Keep
+ `ENVELOPE_MAX_AGE_DAYS` at or above the longest offline period you support.
+- **The `/sync` window is separate from GC.** `/sync` returns envelopes from
+ the last 7 days (`ENVELOPE_TTL_SECONDS`). That controls what is served, not
+ what is kept.
+- **File hard-delete is idempotent.** `hardDeletedAt` is set only after the
+ object-store delete succeeds, so a crash between the two steps is retried.
+ Before each delete, the job re-checks that no live message references the
+ file.
+- **Push dead-endpoint pruning is not a GC job.** It happens inline when a
+ send returns 410/404. A transient failure sets `disabledAt` for 5 minutes,
+ and the file-cleanup tick re-enables the subscription.
+
+## Per-category notes
+
+### Messages and envelopes
+
+A message is stored as one `messages` row plus one `message_envelopes` row per
+recipient device. For MLS groups, the single group ciphertext lives on the
+`messages` row instead. When the **sender** deletes a message (`DELETE
+/messages/:id`, or the `delete_message` socket event):
+
+1. `messages.deletedAt` is set and `messages.ciphertext` is set to `NULL`.
+2. Every envelope for the message is deleted.
+3. The attached file, if any, is soft-deleted, but only once no other live
+ message references it.
+4. `message_deleted` is broadcast so clients drop their copy.
+
+The `messages` row itself stays: sender, conversation, timestamps, content
+type, file link and MLS epoch. Edits are separate `messages` rows linked by
+`editsMessageId`. Deleting the original does not delete its edits, so each
+edit must be deleted on its own.
+
+### Files
+
+The lifecycle is `pending → ready → deleted`:
+
+- **Unconfirmed uploads** are deleted after 24 hours.
+- **Soft delete** (`deletedAt`) happens when the last live message referencing
+ the file is deleted.
+- **Hard delete** removes the blob from the object store on the next
+ file-cleanup pass after the grace period (default: immediately, so within
+ about 5 minutes). The `files` row survives with `hardDeletedAt` set, keeping
+ its size, MIME type, sha256 and uploader. The file key never touches the
+ server: it lives inside the envelope ciphertext.
+
+### Devices and prekeys
+
+Revoking a device (`DELETE /devices/:id`, or log-out-everywhere):
+
+- sets `revokedAt`
+- deletes all of the device's `device_prekeys`
+- releases the prekeys-low latch
+- disconnects the device's sockets on every instance
+
+It does **not** delete the device row, its MLS KeyPackages, its push
+subscriptions, its undelivered envelopes, or its key history. Those fall
+under their own windows in the table above. The device row is kept on
+purpose, to preserve audit history. The GC only flags it as stale after 180
+days, and there is no hard delete.
+
+### Audit logs
+
+`audit_logs` records security events: device linked or revoked,
+log-out-everywhere, key-bundle drains, failed auth, file access denied, group
+membership changes. It is **append-only by design**, and the actor and subject
+columns are **deliberately not foreign keys**
+([security/audit-logging.md](security/audit-logging.md#append-only)):
+
+- A cascade would delete a user's history along with the account it
+ incriminates.
+- `ON DELETE SET NULL` would issue an `UPDATE` that the append-only rule
+ refuses.
+
+As a result, **no user action and no GC job removes audit rows.** Removing
+them is a deliberate, privileged operation: drop the trigger, prune, then
+recreate the trigger. Rows hold identifiers, IP address, user agent and
+sanitised metadata, and never message content.
+
+> **Gap:** the trigger that enforces append-only (`audit_logs_no_mutation`) is
+> described in `security/audit-logging.md` and `db/schema.ts`, but it is **not
+> present in any migration**. `apps/backend/drizzle/` contains only
+> `0000_lean_scrambler.sql`, and the `0001_audit_logs.sql` the audit doc cites
+> does not exist. Until it is added, audit rows are append-only by application
+> convention only, and the database account can update or delete them.
+
+### Presence and resume streams (Redis)
+
+Everything in Redis is short-lived metadata with a TTL:
+
+- **Presence keys** expire 90 seconds after the last heartbeat and are deleted
+ on disconnect.
+- **The resume stream** holds ephemeral events (typing, presence and similar,
+ not message content) for replay after a reconnect. It expires 300 seconds
+ after its last write and is trimmed to about 500 entries.
+
+The durable part of presence is `devices.lastSeenAt` in Postgres, which is kept
+as long as the device row: indefinitely. If Redis is flushed, all of this is
+lost with no user-visible effect beyond a presence blip.
+
+### Push subscriptions
+
+A subscription stores the browser's push endpoint and its encryption keys
+(`p256dh`, `auth`). It is removed when:
+
+- the device calls `DELETE /push/subscriptions`, or
+- the push service returns 410 Gone or 404, in which case it is pruned on the
+ next send.
+
+It is **not** removed when the device is revoked, and the FK cascade never
+fires because device rows are never deleted. The push provider (the browser
+vendor) keeps its own delivery metadata, outside this system.
+
+### AI agent (Weaviate)
+
+`POST /index/message` in `apps/ai_agent` stores the **plaintext** message
+`content` plus its embedding in Weaviate. `GET /search` reads it back. There is
+no delete endpoint and no retention job. No code in this repo calls
+`/index/message` today. **Wiring it up would store server-readable plaintext
+indefinitely**, which contradicts the end-to-end encryption model in
+[threat-model.md](threat-model.md). Any caller needs an explicit opt-in, a
+retention window, and deletion that follows message deletion.
+
+## User-initiated erasure
+
+There is **no account-deletion or "erase my data" endpoint.** Users can only
+remove data object by object:
+
+| User action | What it removes | What it leaves |
+| ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| Delete own message | Its ciphertext, all its envelopes, and the file blob if no other live message uses it (within about 5 minutes) | The message row (metadata), edit rows, the `files` row, and copies already on recipients' devices |
+| Revoke a device / log out everywhere | The device's prekeys and live sockets | The device row, key history, MLS KeyPackages (until GC), push subscriptions, undelivered envelopes (until GC), and an audit row recording the revocation |
+| Leave a group | The membership row. If they are the last member, the whole conversation cascades (messages, envelopes, files rows, MLS state, control log). | The user's past messages in a group that still has members. A `member_left` control event and an audit row. **For the last member: file blobs** (see [Known gaps](#known-gaps)). |
+| Unsubscribe from push | That push subscription | Nothing else |
+| Change privacy settings | Future presence and read-receipt exposure to _other users_ | What the server already holds. The server sees presence regardless. |
+
+## What erasure cannot reach
+
+These records survive every action a user can take. Some survive by design,
+some because of gaps in the code:
+
+1. **Audit logs, by design.** They are append-only and not FK-linked, so that
+ deleting an account cannot erase the history of what was done with it. Only
+ an operator can prune them, deliberately.
+2. **On-chain data, permanently.** Token transfers, treasury deposits and
+ withdrawals, proposals and votes live on the Stellar ledger. They are
+ public, and no one can delete them: not the user, not the operator. The
+ Postgres mirrors (`token_transfers`, `treasury_proposals`,
+ `proposal_votes`) can be removed, but the chain copy cannot. Wallet
+ addresses tie those records to the user.
+3. **Account and device records.** There is no deletion path for `users`,
+ `wallets` or `devices`. `device_key_history` is permanent by design.
+4. **Message metadata.** Deleted messages leave rows that record sender,
+ conversation and time. DM conversations cannot be left, so their messages,
+ `files` rows and membership are kept indefinitely.
+5. **AI assistant replies.** They are stored as server-readable plaintext
+ under the assistant's user ID. The delete route only lets the _sender_
+ delete, so no user can remove them.
+6. **Copies outside the server:** recipients' devices, the push provider,
+ OpenAI (assistant prompts), and operator backups and logs.
+
+## Known gaps
+
+These are the places where the code falls short of what this page, or other
+docs, would lead a reader to expect:
+
+- **No account deletion.** Building one has to decide:
+ - what cascades (`users` cascades widely through FKs)
+ - what is tombstoned instead (audit rows must stay; see above)
+ - how to handle DMs, where the other party's copy is legitimately theirs
+- **Audit append-only is not enforced in the database.** The trigger is
+ missing from migrations (see [Audit logs](#audit-logs)).
+- **Orphaned file blobs on conversation deletion.** When the last member
+ leaves a group, `files` rows are cascade-deleted by the database, so the
+ file-cleanup job never sees them and **their encrypted blobs stay in the
+ object store forever**.
+- **Revocation leaves push subscriptions and KeyPackages behind.** Push
+ subscriptions for a revoked device are never removed unless the push service
+ reports them dead.
+- **Weaviate stores plaintext with no retention.** See
+ [AI agent](#ai-agent-weaviate).
diff --git a/docs/environment-variables.md b/docs/environment-variables.md
new file mode 100644
index 0000000..26b0ecf
--- /dev/null
+++ b/docs/environment-variables.md
@@ -0,0 +1,195 @@
+# Environment Variables
+
+Every environment variable read by the four apps in this repo: `apps/backend`,
+`apps/web`, `apps/ai_agent` and `apps/tests`. Sources:
+
+- [`.env.example`](../.env.example)
+- [`apps/backend/src/config.ts`](../apps/backend/src/config.ts) (boot-time schema)
+- [`apps/backend/src/config/rateLimits.ts`](../apps/backend/src/config/rateLimits.ts)
+- every direct `process.env` / `os.environ` read under `apps/`
+
+`apps/tests` reads no environment variables. The backend test suite sets its
+own values in `apps/backend/src/__tests__/setup.ts`.
+
+## How to read the table
+
+**Loaded** says when the value is read, and so how a bad value shows up:
+
+| Loaded | Meaning | If the value is missing or malformed |
+| ------------ | ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ |
+| `boot` | Validated by `loadEnv()` in `config.ts`, called at the top of `apps/backend/src/index.ts`. | **Crashes on start:** logs `Missing or invalid environment variables: …` and exits 1. |
+| `boot (TLS)` | Checked by `assertTransportSecurityConfig()` right after `loadEnv()`. | **Crashes on start** for the contradictions described in the row. Other cases only log a warning. |
+| `module` | Read once when the module is first imported. No validation. | **Fails silently.** Usually the default is used, but a few `parseInt` reads become `NaN` (see the row). A restart is needed to pick up a change. |
+| `lazy` | Read on every call or job run. Parsed with a fallback. | **Fails silently.** An invalid value falls back to the default, sometimes with a `console.warn`. |
+| `build` | Inlined into the web client bundle by `next build`. | **Fails silently** in the browser. A rebuild is needed to pick up a change. |
+| `unused` | Declared in `.env.example` or `config.ts` but read by no code. | Nothing. Setting it has no effect. |
+
+`boot` only checks that a value is **present and well-formed**. A
+`DATABASE_URL` that points at the wrong host still passes and only fails on
+the first query.
+
+**Secret** 🔒 marks credentials. See [Secrets](#secrets).
+
+**Public** 🌐 marks values compiled into the browser bundle. See
+[`NEXT_PUBLIC_*` values are public](#next_public_-values-are-public).
+
+## Caveat: the backend loads `.env` too late for `module` reads
+
+`apps/backend/src/index.ts` calls `dotenv.config()` **after** its `import`
+statements. The backend is an ES module (`"type": "module"`), so every imported
+module has already run its top-level code before dotenv loads the file. As a
+result:
+
+- Values supplied **only** through a `.env` file are invisible to every
+ `module` row except those read in `index.ts` itself (`PORT`, `REDIS_URL` for
+ the Socket.IO adapter, `STELLAR_RPC_URL`, `TOKEN_TRANSFER_CONTRACT_ID`,
+ `GROUP_TREASURY_CONTRACT_ID`). The other `module` readers fall back to their
+ defaults.
+- **`JWT_SECRET` is affected.** `lib/jwt.ts` runs first, finds no value, falls
+ back to `'test-secret'` and writes that back into `process.env`. dotenv then
+ refuses to overwrite it, and `loadEnv()` passes because the variable is now
+ set. **Tokens are signed with `'test-secret'`**, whatever the `.env` file
+ says.
+- `DATABASE_URL` (`db/index.ts`), `REDIS_URL` (`lib/redis.ts`) and the
+ `VAPID_*` keys (`services/pushNotification.ts`) are affected the same way.
+
+Values injected by the process environment (Docker `environment:`, Kubernetes,
+systemd, `export …`) are not affected. **In any deployed environment, set
+variables through the real process environment, not a `.env` file**, until
+dotenv is loaded before the other imports (for example with
+`node --import dotenv/config`).
+
+## All variables
+
+| Variable | App | Required | Default | Loaded | What breaks if it is wrong |
+| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------ | ------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| **Auth** | | | | | |
+| `JWT_SECRET` 🔒 | backend | **Yes** | none (`lib/jwt.ts` falls back to `'test-secret'`, see caveat) | `boot` + `module` | A weak or leaked value lets anyone forge session tokens for any user. Changing it invalidates every existing session. |
+| **Core infrastructure** | | | | | |
+| `DATABASE_URL` 🔒 | backend | **Yes** | none (`db/index.ts` falls back to `postgres://user:password@localhost:5432/testdb`) | `boot` + `module` | Boot passes if the value is non-empty. A wrong host or credentials only show up as failed queries at runtime. It is also read by `drizzle.config.ts` for migrations (empty string if unset). 🔒 because it contains the DB password. |
+| `REDIS_URL` 🔒 | backend | **Yes** | none (the Socket.IO adapter falls back to `redis://localhost:6379`) | `boot` + `module` | The Socket.IO Redis adapter logs a warning and **degrades to single-instance mode**: multi-instance rooms and presence break. If the value is missing at import time, `lib/redis.ts` sets `redis = null`, which disables caching, presence reconciliation and device-revocation fan-out. 🔒 when the URL embeds a password. |
+| `PORT` | backend | **Yes** | none (the `3001` fallback in `index.ts` is never reached because `loadEnv()` exits first) | `boot` | Must be a positive integer, or boot fails. A mismatch with the web app's URLs means the web app cannot reach the API. It is also used to build local-disk presigned URLs in dev. |
+| `NODE_ENV` | backend | Required in prod | unset → treated as `development` | `lazy` | `production` switches to the real S3 object store and unmounts `/local-storage`. **If it is unset in production**, uploads go to the container's local disk (`.local-storage/`) with `http://localhost` URLs. If `APP_ENV` is also unset, TLS enforcement, HSTS and the origin check are off as well. `test` silences `morgan` request logging. |
+| `APP_ENV` | backend | Optional | falls back to `NODE_ENV`, then `development` | `lazy` | Selects the transport-security posture. Only `development` and `test` allow plaintext. **Any other value is treated as production.** |
+| **Transport security** (see [`security/tls-and-pinning.md`](security/tls-and-pinning.md)) | | | | | |
+| `ENFORCE_TLS` | backend | Optional | on outside `development`/`test` | `boot (TLS)` + `lazy` | Accepts `true/1/yes/false/0/no`; any other value is ignored. `false` outside dev logs a warning and **accepts plaintext `http`/`ws`**. |
+| `ALLOWED_ORIGINS` | backend | Required in prod | empty | `boot (TLS)` + `lazy` | **Boot fails** if TLS is enforced and the list contains an `http://` origin. If empty outside dev, any `https://` origin can call the API and open a socket (a warning is logged). If it omits the web app's origin, CORS and the socket handshake are refused. |
+| `TRUST_PROXY` | backend | Optional | `1` | `lazy` | Non-integer → `1`. `1` with the gateway exposed directly lets clients forge `X-Forwarded-Proto: https`. `0` behind a load balancer makes every request look plaintext, so TLS enforcement refuses all traffic. |
+| `HSTS_MAX_AGE` | backend | Optional | `31536000` | `lazy` | Invalid → default. `0` removes the HSTS header. |
+| `HSTS_PRELOAD` | backend | Optional | `true` | `lazy` | Adds `preload` to HSTS. Hard to undo once browsers have preloaded the domain. |
+| `TLS_PINNED_HOSTS` | backend | Optional | empty | `lazy` | Hostnames advertised at `GET /security/transport-policy`. Public, not secret. |
+| `TLS_PINNED_SPKI_SHA256` | backend | Optional | empty | `lazy` | Malformed pins are dropped with a warning. Mobile clients enforce pins only when a backup pin is also set. **A wrong pin plus a backup locks every installed client out.** |
+| `TLS_BACKUP_SPKI_SHA256` | backend | Optional (required with the above) | empty | `lazy` | Without it, pinning is advertised as not enforced. |
+| `TLS_PIN_MAX_AGE_SECONDS` | backend | Optional | `5184000` (60 days) | `lazy` | Invalid → default. How long clients cache the pin set. Too long makes rotation slow. |
+| `TLS_PIN_REPORT_URI` | backend | Optional | none | `lazy` | Where clients report pin failures. |
+| **Blockchain** | | | | | |
+| `TOKEN_TRANSFER_CONTRACT_ID` | backend | **Yes** | none | `boot` | Also turns on the Stellar transfer listener, together with `STELLAR_RPC_URL`. **A wrong ID passes boot**; the listener just watches the wrong contract and transfer events never arrive. Must match `NEXT_PUBLIC_TOKEN_TRANSFER_CONTRACT`. |
+| `STELLAR_RPC_URL` | backend | Optional | none | `module` (index.ts) | **Not in `.env.example`** (which lists `RPC_URL` instead). If unset, the transfer and treasury listener is disabled with a single log line. |
+| `GROUP_TREASURY_CONTRACT_ID` | backend | Optional | `'stub'` in `routes/treasury.ts` | `module` + `lazy` | If unset, treasury events are not watched and the treasury route reports contract `stub`. |
+| `RPC_URL` | — | — | — | `unused` | Listed in `.env.example`, but the backend reads `STELLAR_RPC_URL`. Setting only `RPC_URL` leaves the listener off. |
+| `PROPOSALS_CONTRACT_ID` | — | — | — | `unused` | No code reads it. |
+| **Object storage** (used only when `NODE_ENV=production`, but validated in every environment) | | | | | |
+| `OBJECT_STORE_ENDPOINT` | backend | **Yes** | `.env.example`: `http://localhost:9000` | `boot` | A wrong endpoint passes boot. Presigned upload and download URLs then fail when a client uses them. |
+| `OBJECT_STORE_BUCKET` | backend | **Yes** | `.env.example`: `clicked` | `boot` | Same as above: uploads fail at runtime. |
+| `OBJECT_STORE_ACCESS_KEY` 🔒 | backend | **Yes** | `.env.example`: `clicked` (local MinIO only) | `boot` | Presigned URLs are rejected by the store (`SignatureDoesNotMatch`). |
+| `OBJECT_STORE_SECRET_KEY` 🔒 | backend | **Yes** | `.env.example`: `clickedsecret` (local MinIO only) | `boot` | Same as above. A leak gives full read/write access to every stored attachment. |
+| `OBJECT_STORE_REGION` | backend | **Yes** | `.env.example`: `us-east-1` | `boot` | Signature mismatch on region-strict providers. |
+| `OBJECT_STORE_FORCE_PATH_STYLE` | backend | **Yes** | `.env.example`: `true` | `boot` | Must be `true/false/1/0`, or boot fails. The wrong style for the provider produces unreachable URLs (use `true` for MinIO, `false` for AWS S3 and R2). |
+| `S3_ENDPOINT`, `S3_REGION`, `S3_ACCESS_KEY_ID`, `S3_SECRET_ACCESS_KEY` 🔒, `S3_BUCKET`, `S3_FORCE_PATH_STYLE` | — | — | — | `unused` | Declared optional in `config.ts` but never read. Use the `OBJECT_STORE_*` variables instead. |
+| `LOCAL_STORAGE_DIR` | backend | Optional (dev) | `/.local-storage` | `lazy` | Dev and test only: where the local-disk object store writes files. |
+| `STORAGE_ENDPOINT` | backend | Optional (dev) | `http://localhost:$PORT/local-storage` | `lazy` | Dev and test only: the base URL of local presigned URLs. If wrong, uploads from another device or container fail. |
+| **Push notifications (VAPID)** | | | | | |
+| `VAPID_PUBLIC_KEY` | backend | Optional | none | `boot` (optional, unchecked) + `module` + `lazy` | If unset, `GET /push/vapid-public-key` returns `configured: false` and the web app skips push. If it does not pair with the private key, push delivery fails silently. |
+| `VAPID_PRIVATE_KEY` 🔒 | backend | Optional | none | `boot` (optional, unchecked) + `module` | If unset, push is silently disabled. A leak lets anyone send push notifications to your subscribers. |
+| `VAPID_SUBJECT` | backend | Optional | `mailto:admin@clicked.app` | `boot` (optional, unchecked) + `module` | Push services may reject a subject that is not a `mailto:` or `https:` URL. |
+| **Messaging and sync** | | | | | |
+| `REPLAY_PROTECTION_TTL_SECONDS` | backend | Optional | `300` | `lazy` | **Not in `.env.example`.** This is the window that actually controls replay and duplicate `eventId` rejection. Values outside `1`–`86400` fall back to the default. |
+| `IDEMPOTENCY_TTL_SECONDS` | backend | Optional | `.env.example` says `86400` | `boot` | A non-positive or non-integer value **crashes boot**, but the parsed value is **never used**. Replay protection reads `REPLAY_PROTECTION_TTL_SECONDS` instead. |
+| `PRESENCE_OFFLINE_GRACE_MS` | backend | Optional | `5000` | `lazy` | Invalid → default. Too high delays "offline" status. `0` produces offline/online flapping on brief disconnects. |
+| `PREKEY_LOW_THRESHOLD` | backend | Optional | `20` | `module` | Invalid → default. Too low and devices run out of one-time prekeys before they are told to replenish. |
+| `SOCKET_EVENT_MAX_AGE_MS` | backend | Optional | `300000` | `module` | **Not in `.env.example`.** A non-numeric value becomes `NaN` and **every socket event is rejected as stale**. |
+| `SOCKET_EVENT_MAX_FUTURE_SKEW_MS` | backend | Optional | `30000` | `module` | **Not in `.env.example`.** Same `NaN` failure as above. |
+| `ENVELOPE_TTL_SECONDS` | backend | Optional | `604800` (7 days) | `module` | **Not in `.env.example`.** How far back `/sync` returns envelopes. A non-numeric value becomes `NaN` and breaks the sync cutoff. |
+| `SYNC_PAGE_SIZE` | backend | Optional | `50` | `module` | **Not in `.env.example`.** Maximum `/sync` page size. A non-numeric value becomes `NaN`. |
+| `MAX_PAYLOAD_SIZE` | backend | Optional | `16384` bytes | `lazy` | **Not in `.env.example`.** Too low and legitimate socket messages are rejected. |
+| `MAX_ENVELOPE_SIZE` | backend | Optional | `4096` bytes | `lazy` | **Not in `.env.example`.** Too low and legitimate encrypted envelopes are rejected. |
+| `FIRST_CONTACT_HOUR_LIMIT` | backend | Optional | `5` per hour | `module` | **Not in `.env.example`.** A non-numeric value becomes `NaN` and **every first-contact message is blocked**. |
+| `GROUP_INVITE_HOUR_LIMIT` | backend | Optional | `10` per hour | `module` | **Not in `.env.example`.** A non-numeric value becomes `NaN` and **every group invite is blocked**. |
+| `XMTP_ENV` | — | — | — | `unused` | Listed in `.env.example`; no code reads it. |
+| **Socket backpressure** | | | | | |
+| `SOCKET_BUFFER_THRESHOLD` | backend | Optional | `65536` bytes | `lazy` | **Not in `.env.example`.** A socket whose send buffer exceeds this is disconnected. |
+| `SOCKET_SHED_THRESHOLD` | backend | Optional | `32768` bytes | `lazy` | **Not in `.env.example`.** Above this, the socket is marked "shed" and a warning and metric are emitted. Nothing reads the flag yet (`isSocketShed` has no callers), so no traffic is actually dropped. Should stay below `SOCKET_BUFFER_THRESHOLD`. |
+| **Background GC jobs** (none in `.env.example`) | | | | | |
+| `DEVICE_GC_INTERVAL_MS` | backend | Optional | `3600000` (1 h) | `lazy` (at job start) | Invalid → default. |
+| `PREKEY_CONSUMED_RETENTION_DAYS` | backend | Optional | `30` | `lazy` | Invalid → default. How long consumed prekeys are kept for audit. |
+| `PREKEY_UNCONSUMED_MAX_AGE_DAYS` | backend | Optional | `90` | `lazy` | Invalid → default. Too low deletes prekeys that peers still need to start sessions. |
+| `DEVICE_STALE_AFTER_DAYS` | backend | Optional | `180` | `lazy` | Invalid → default. |
+| `ENVELOPE_GC_INTERVAL_MS` | backend | Optional | `1800000` (30 min) | `lazy` (at job start) | Invalid → default. |
+| `ENVELOPE_DELIVERED_RETENTION_DAYS` | backend | Optional | `7` | `lazy` | Invalid → default. |
+| `ENVELOPE_MAX_AGE_DAYS` | backend | Optional | `30` | `lazy` | Invalid → default. Too low and **messages to offline devices are deleted before they are delivered**. |
+| `FILE_GC_INTERVAL_MS` | backend | Optional | `300000` (5 min) | `lazy` | Invalid → default. |
+| `FILE_HARD_DELETE_GRACE_MS` | backend | Optional | `0` | `lazy` | Invalid → default. |
+| `PENDING_UPLOAD_TTL_MS` | backend | Optional | `86400000` (24 h) | `lazy` | Invalid → default. Too low deletes uploads that are still in progress. |
+| **Rate limits** (see [`security/rate-limits.md`](security/rate-limits.md)). Format: `[/]`. A malformed value logs `[rateLimit] ignoring malformed …` and uses the default. All are `lazy`: read on every check, so no restart is needed. | | | | | |
+| `RATE_LIMIT_DISABLED` | backend | Optional | `false` | `lazy` | Only the exact string `true` disables **every** limit. **Never set it in production**, because it re-opens enumeration and resource exhaustion. |
+| `RATE_LIMIT_GLOBAL_IP` | backend | Optional | `600/60` | `lazy` | Per-IP ceiling across every HTTP endpoint. Too low throttles clients behind shared NAT. |
+| `RATE_LIMIT_AUTH_CHALLENGE` | backend | Optional | `10/60` | `lazy` | Wallet challenge nonce issuance. |
+| `RATE_LIMIT_AUTH_VERIFY` | backend | Optional | `5/60` | `lazy` | Signature verification attempts. Too high weakens brute-force protection. |
+| `RATE_LIMIT_DEVICE_LINK_CHALLENGE` | backend | Optional | `10/60` | `lazy` | Device-link challenge issuance. |
+| `RATE_LIMIT_DEVICE_LINK_VERIFY` | backend | Optional | `5/60` | `lazy` | Device-link verification attempts. |
+| `RATE_LIMIT_KEY_BUNDLE` | backend | Optional | `30/60` | `lazy` | X3DH prekey bundle fetches. |
+| `RATE_LIMIT_KEY_BUNDLE_DAILY` | backend | Optional | `200/86400` | `lazy` | Daily bundle quota. Too high allows draining one-time prekeys. |
+| `RATE_LIMIT_UPLOAD_SLOT` | backend | Optional | `20/60` | `lazy` | Presigned upload slot requests. |
+| `RATE_LIMIT_UPLOAD_BYTES_DAILY` | backend | Optional | `2147483648/86400` (2 GiB) | `lazy` | Daily upload volume per user. |
+| `RATE_LIMIT_FILE_DOWNLOAD` | backend | Optional | `120/60` | `lazy` | Presigned download URL issuance. |
+| `RATE_LIMIT_PUSH_SUBSCRIBE` | backend | Optional | `10/60` | `lazy` | Web-push subscription registration. |
+| `RATE_LIMIT_SOCKET_DEFAULT` | backend | Optional | `10/1` | `lazy` | Any socket event without its own bucket. |
+| `RATE_LIMIT_SOCKET_SEND_MESSAGE` | backend | Optional | `30/10` | `lazy` | `send_message` and `send_file_message`. |
+| `RATE_LIMIT_SOCKET_TYPING` | backend | Optional | `20/5` | `lazy` | `typing_start` and `typing_stop`. |
+| `RATE_LIMIT_SOCKET_ASK_ASSISTANT` | backend | Optional | `5/60` | `lazy` | AI assistant invocations, which drive the OpenAI bill. |
+| `SOCKET_RATE_LIMIT_PER_SEC` | backend | Optional | none | `lazy` | Legacy setting. Used only when `RATE_LIMIT_SOCKET_DEFAULT` is unset, and sets that bucket to `/1`. |
+| **Logging** | | | | | |
+| `LOG_LEVEL` | backend | Optional | `info` | `module` | **Not in `.env.example`.** An unknown level makes pino throw at import, which **crashes on start**. |
+| **AI agent** | | | | | |
+| `OPENAI_API_KEY` 🔒 | ai_agent | **Yes** (for AI endpoints) | none | `lazy` (per request) | Listed in `.env.example` under "AI Service", but **only the AI agent reads it**; the backend does not. The agent does not load `.env` files, so the key must be in its process environment. If unset or invalid, the service still starts and `/health` is fine, but `/chat`, `/proposals/summarise` and sub-threshold `/transfers/analyse` return 500, and `/index/message` and `/search` return 503. A leak lets others spend on your OpenAI account. The agent has no other settings: its port (`8000`) and Weaviate (`connect_to_local()`, `localhost:8080`) are hard-coded. |
+| **Web client** (all 🌐 public, `build`) | | | | | |
+| `NEXT_PUBLIC_API_URL` 🌐 | web | Required outside local dev | **`http://localhost:4000`** in `lib/api.ts`, but `http://localhost:3001` in `NewConversationModal.tsx` | `build` | The two fallbacks disagree, and the backend defaults to port 3001, so **leaving it unset breaks most REST calls in local dev**. In production it must be the `https://` API origin. |
+| `NEXT_PUBLIC_SOCKET_URL` 🌐 | web | Required outside local dev | `http://localhost:3001` | `build` | Socket never connects. Precedence is inconsistent: `hooks/useSocket.ts` prefers this variable over `NEXT_PUBLIC_BACKEND_URL`, while `lib/socket.ts` prefers `NEXT_PUBLIC_BACKEND_URL`. Set both to the same value. It must be `https://` when TLS is enforced. |
+| `NEXT_PUBLIC_BACKEND_URL` 🌐 | web | Required outside local dev | `http://localhost:3001` | `build` | See `NEXT_PUBLIC_SOCKET_URL`. |
+| `NEXT_PUBLIC_SOROBAN_RPC_URL` 🌐 | web | Optional | `https://soroban-testnet.stellar.org` | `build` | If unset in production, transfers are built against testnet. |
+| `NEXT_PUBLIC_NETWORK_PASSPHRASE` 🌐 | web | Optional | `Networks.TESTNET` | `build` | If it does not match the RPC's network, Freighter-signed transactions are rejected. |
+| `NEXT_PUBLIC_TOKEN_TRANSFER_CONTRACT` 🌐 | web | **Yes** for transfers | `'REPLACE_WITH_TOKEN_TRANSFER_CONTRACT_ID'` | `build` | If unset, every in-chat transfer fails. Must equal the backend's `TOKEN_TRANSFER_CONTRACT_ID`, or the backend listener never sees the transfer. |
+| `NEXT_PUBLIC_NETWORK` 🌐 | web | Optional | `test` | `build` | If wrong, transfer cards link to the wrong network on the Stellar explorer. |
+| `NEXT_PUBLIC_AUTH_TOKEN` 🌐 | web | **Never set outside local dev** | none | `build` | A fallback session token for when none is stored. Because it is compiled into the bundle, **every visitor receives this token and is signed in as that user**. |
+| `NEXT_PUBLIC_VAPID_PUBLIC_KEY` 🌐 | web | — | `''` | `build` (`next.config.ts`) | Effectively unused: declared in `next.config.ts`, but no code reads it. The web app fetches the key from `GET /push/vapid-public-key` at runtime (#349). |
+
+## Secrets
+
+These values are credentials: 🔒 `JWT_SECRET`, `DATABASE_URL`, `REDIS_URL`
+(when it embeds a password), `OBJECT_STORE_ACCESS_KEY`,
+`OBJECT_STORE_SECRET_KEY`, `VAPID_PRIVATE_KEY`, `OPENAI_API_KEY`, and the
+unused `S3_SECRET_ACCESS_KEY`.
+
+- **Never commit them.** `.env` is ignored by the root `.gitignore`, and
+ `.env*` by `apps/web/.gitignore`. Only `.env.example` is tracked, and it must
+ hold placeholders or local-only values.
+- The `OBJECT_STORE_*` values in `.env.example` (`clicked` / `clickedsecret`)
+ match the local MinIO in `infra/docker-compose.yml`. They are for local use
+ only and must never be reused anywhere reachable.
+- Never give a secret a `NEXT_PUBLIC_` prefix. That publishes it (see below).
+- If a secret is committed or leaked, rotate it; deleting the commit is not
+ enough. Rotating `JWT_SECRET` signs out every user.
+
+## `NEXT_PUBLIC_*` values are public
+
+Next.js replaces every `process.env.NEXT_PUBLIC_*` reference with its literal
+value at `next build`. The value ends up in the JavaScript sent to every
+browser, so anyone can read it with view-source. It follows that:
+
+- Every `NEXT_PUBLIC_*` variable above is **public**. None of them may hold a
+ secret. Most are harmless (URLs, network passphrase, contract ID), but
+ `NEXT_PUBLIC_AUTH_TOKEN` is a live credential and must never be set for a
+ shared or deployed build.
+- Changing one requires a **rebuild**. Restarting `next start` is not enough.
+- Next.js loads `.env*` files from `apps/web/`, not the repo root. The root
+ `.env.example` does not list any web variables.
diff --git a/docs/release-process.md b/docs/release-process.md
new file mode 100644
index 0000000..dfe409e
--- /dev/null
+++ b/docs/release-process.md
@@ -0,0 +1,342 @@
+# Release Process
+
+How a change gets from a contributor's pull request to production, for all
+four deployables:
+
+- the backend service (`apps/backend`)
+- the Next.js web app (`apps/web`)
+- the AI agent (`apps/ai_agent`)
+- the Soroban contracts (`contracts/`)
+
+For checks after a deploy and for incident handling, use the
+[operator runbook](runbook.md). For every setting mentioned here, see
+[environment variables](environment-variables.md).
+
+## What is automated today
+
+Be clear about this before reading further: **the repository contains no
+deploy automation.** CI verifies code; people deploy it.
+
+| Stage | Automated? | Where |
+| --------------------------------------------------- | --------------------------------- | ---------------------------------------------------------------- |
+| Lint, test, build on every PR and push | Yes | `.github/workflows/*-ci.yml`, filtered by app path |
+| Security regression and crypto-dependency CVE audit | Yes, on PRs and on push to `main` | `security-ci.yml` |
+| Closing PRs to `main` that a non-maintainer opened | Yes | `guard-main-branch.yml` |
+| Closing linked issues when a PR merges to `dev` | Yes | `close-linked-issues.yml` |
+| Version bumps, tags, changelog | **No** | Manual (see [Versioning](#versioning)) |
+| Database migrations | **No** | Manual (`pnpm --filter backend db:migrate`) |
+| Deploying the backend, web app or AI agent | **No** | Manual. There are no Dockerfiles or hosting configs in the repo. |
+| Contract deploys | **No** | Manual scripts in `contracts/scripts/`, testnet only |
+
+## Branch flow: contributor PR → `dev` → `main`
+
+```
+contributor fork ──PR──▶ dev ──(maintainer PR)──▶ main ──▶ deploy
+```
+
+1. **A contributor opens a PR against `dev`.** PRs against `main` from anyone
+ other than the maintainer are closed automatically by
+ `guard-main-branch.yml`, with a comment asking the author to retarget `dev`.
+2. **CI must pass.** The workflows that run depend on the paths the PR touches:
+ Backend CI, Frontend CI, AI Agent CI, Contracts CI, Security CI, and the
+ repo-wide lint in `pr.yml`.
+3. **A maintainer reviews and merges into `dev`.** Any `Closes #N` / `Fixes #N`
+ issues in the PR are closed at that point by `close-linked-issues.yml`,
+ because GitHub only auto-closes them on merges to the default branch.
+4. **The maintainer promotes `dev` to `main`.** They open a PR from `dev` into
+ `main`, wait for CI on the combined changes, and merge it. **Only the
+ maintainer merges to `main`.** In this repo, "maintainer" means the repo
+ owner or a collaborator with `admin` or `maintain` permission, which is the
+ same check `guard-main-branch.yml` uses.
+5. **A merge to `main` is the release candidate.** Nothing deploys
+ automatically. The maintainer tags the release and runs the
+ [release checklist](#release-checklist) below.
+
+### Gaps to close before this flow is real
+
+- **There is no `dev` branch on `origin` yet.** At the time of writing, only
+ `main` (plus feature branches) exists, and recent contributor PRs were merged
+ straight into `main`. Create `dev` from `main` and make it the base for new
+ contributor PRs.
+- **The guard only checks who _opened_ a PR, not who _merges_ it.** It closes
+ non-maintainer PRs to `main`, but it does not stop a collaborator with write
+ access from merging a PR that the maintainer opened. To enforce "only the
+ maintainer merges to `main`", add a branch-protection rule (or ruleset) on
+ `main` that restricts who can push and merge to the maintainer, and requires
+ the CI checks to pass.
+
+## Versioning
+
+**Today:** there are no git tags, no changelog, and no release automation. The
+package versions have never been bumped: backend `1.0.0`, web `0.1.0`,
+AI agent `0.1.0`, and the contracts have no versions. The backend reports its
+`package.json` version from `GET /health`, so that value is what operators see.
+
+**Convention from now on:**
+
+- **One tag per promotion to `main`**, in the form `vMAJOR.MINOR.PATCH` (SemVer),
+ on the merge commit of the `dev → main` PR. The whole monorepo shares one
+ version, because the apps are deployed together and depend on each other's
+ API.
+- **How to choose the bump:**
+ - **MAJOR** for any change that breaks a deployed client or peer: a
+ WebSocket event or REST contract change that old web or mobile clients
+ cannot handle, a destructive database migration, or a contract redeploy
+ that changes a contract ID.
+ - **MINOR** for new features and additive migrations.
+ - **PATCH** for fixes only.
+- **Bump `apps/backend/package.json` to the tag version** in the promotion PR,
+ so that `/health` reports what is actually running. Bumping the web and
+ AI agent versions to match is optional.
+- **Write release notes in the GitHub Release for the tag.** List the merged PRs,
+ migrations, new or changed environment variables, and any contract ID changes.
+
+## Release checklist
+
+Deploy in this order. Each step depends on the ones before it:
+
+1. **Contracts**, only if the release changes them. Contract IDs feed the
+ backend's environment and are baked into the web build. See
+ [Soroban contracts](#soroban-contracts).
+2. **Database migrations.**
+3. **Backend.**
+4. **AI agent.** It is independent of the others and can go at any point
+ after step 3.
+5. **Web app.** It goes last because it calls the new backend API and has
+ contract IDs compiled in.
+6. **Verify.** See [Post-deploy verification](#post-deploy-verification).
+
+## Backend service
+
+### Migrations and their order relative to the rollout
+
+Migrations live in `apps/backend/drizzle/` and are applied with
+`drizzle-kit migrate`. **Run them before rolling out the new backend code**, and
+write them so that the _previous_ backend version still works against the
+migrated schema:
+
+- The gateway runs as several instances coordinated through Redis (see
+ [runbook → Scaling gateways](runbook.md#scaling-gateways)). During a rolling
+ deploy, old and new instances serve traffic against the same database at the
+ same time.
+- **Additive changes go in the same release as the code.** Examples: a new
+ table, a new nullable column, a new index. Migrate first, then roll out.
+- **Destructive changes need two releases** (expand, then contract). Examples:
+ dropping or renaming a column, adding `NOT NULL` without a default. Release
+ _N_ stops reading and writing the column. Release _N+1_ ships the migration
+ that removes it, once no running instance uses it.
+- **Never run `db:push` against a shared database.** It diffs the schema and
+ applies changes directly, can drop columns, and records no migration.
+
+```bash
+# From the repo root, with the target database in the environment.
+# drizzle.config.ts reads DATABASE_URL from process.env; if it is unset,
+# the URL is an empty string.
+export DATABASE_URL='postgres://…'
+pnpm --filter backend db:generate # only when authoring: commit the generated SQL in the PR
+pnpm --filter backend db:migrate # at release time: applies pending migrations
+```
+
+Take a database backup or snapshot immediately before `db:migrate`. It is the
+only rollback path for a migration (see below).
+
+### Deploy
+
+```bash
+pnpm install --frozen-lockfile
+pnpm --filter backend build # tsc → apps/backend/dist
+NODE_ENV=production pnpm --filter backend start # node dist/index.js
+```
+
+- **Set configuration through the real process environment, not a `.env`
+ file.** The backend loads `.env` after its modules have read their settings,
+ so values that exist only in `.env` are silently ignored. `JWT_SECRET` then
+ falls back to `'test-secret'`. See the caveat in
+ [environment-variables.md](environment-variables.md#caveat-the-backend-loads-env-too-late-for-module-reads).
+- **`NODE_ENV=production` is required.** Without it, the backend uses the
+ local-disk object store and turns TLS enforcement off (unless `APP_ENV` is
+ set).
+- **Roll instances one at a time with `SIGTERM`.** The shutdown handler keeps
+ presence intact for clients that reconnect to another instance (see
+ [runbook → Scaling gateways](runbook.md#scaling-gateways)). Sticky sessions
+ are not required.
+- If the startup check in `config.ts` fails, the process exits with code 1 and
+ lists the missing or invalid variables. Treat that as a failed deploy and
+ keep the old instances running.
+
+### Rollback
+
+- **Code: yes.** Redeploy the previous tag's build. This is safe **only if**
+ that release's migrations were additive, which the expand/contract rule
+ above guarantees.
+- **Migrations: no automated rollback.** `drizzle-kit` has no down-migration
+ runner, and the repo contains no down scripts.
+ `apps/backend/docs/message-encryption-migration.md` refers to
+ `drizzle/rollback/0003_…down.sql`, but that file does not exist. If a
+ migration must be undone, the options are:
+ 1. Write and ship a forward migration that reverses it (preferred).
+ 2. Restore the pre-migration backup. **This loses every write since the
+ backup**, including messages and prekey consumption.
+- **The Stellar listener's position is lost on restart.** The listener keeps
+ its event cursor in memory only. After a deploy or rollback, check that
+ transfer and treasury events are still arriving (see
+ [Post-deploy verification](#post-deploy-verification)).
+
+## Next.js web app
+
+### Deploy
+
+```bash
+pnpm install --frozen-lockfile
+# NEXT_PUBLIC_* values are compiled into the bundle at build time
+pnpm --filter web build
+pnpm --filter web start # or deploy the build output to your host
+```
+
+- **Every `NEXT_PUBLIC_*` value is fixed at build time.** This includes the
+ API and socket URLs and `NEXT_PUBLIC_TOKEN_TRANSFER_CONTRACT`. Changing one
+ means **rebuilding**, not restarting. Put the values in `apps/web/.env.production`
+ or the host's build environment. They must not come from the repo root.
+- Never set `NEXT_PUBLIC_AUTH_TOKEN` for a deployed build. It would be sent to
+ every visitor.
+- The repo has no hosting configuration. `apps/web/README.md` only contains
+ the default Next.js "Deploy on Vercel" note.
+
+### Rollback
+
+- **Yes, but only as good as your build artifacts.** If the host keeps previous
+ builds (for example, Vercel's instant rollback), promote the previous build.
+- If you have to rebuild an old tag instead, you must rebuild with **the same
+ `NEXT_PUBLIC_*` values** that tag was built with. After a contract redeploy,
+ the old build points at the old contract ID. That is only correct if you
+ are also rolling the contract ID back in the backend.
+- Browsers that already loaded the new bundle keep running it until they
+ reload.
+
+## AI agent
+
+### Deploy
+
+```bash
+cd apps/ai_agent
+uv sync --frozen # installs from uv.lock
+OPENAI_API_KEY=… uv run python main.py # uvicorn on 0.0.0.0:8000 (hard-coded)
+```
+
+- The only setting is `OPENAI_API_KEY`, and it must be in the process
+ environment because the agent does not read `.env` files.
+- The agent connects to Weaviate via `connect_to_local()`, which is hard-coded
+ to `localhost:8080`, so Weaviate must run on the same host or network
+ namespace.
+- `/health` only confirms that the process is up. It does not check OpenAI
+ or Weaviate.
+
+### Rollback
+
+- **Code: yes.** The service is stateless. Redeploy the previous tag.
+- **Indexed data: no.** Message vectors in Weaviate's `Message` collection are
+ not versioned. A release that changes the embedding model
+ (`text-embedding-3-small`) makes existing vectors incompatible, and rolling
+ the code back does not un-index them. Such a release needs a planned
+ re-index and should be treated as MAJOR.
+
+## Soroban contracts
+
+Contract deploys are **a separate path with no rollback**. Treat them as their
+own change, reviewed and scheduled separately from app releases. The
+step-by-step commands are in the
+[contract build, deployment & invocation guide](../contracts/docs/api-deployment-invocation.md).
+This section covers how contract deploys fit into a release.
+
+### How a deploy works
+
+Each `contracts/scripts/deploy_*.sh` script builds the WASM, uploads it,
+**creates a new contract instance**, and calls `initialize`. Run the scripts
+from `contracts/` (see guide §4.4). The `make deploy-contracts` target runs
+them from the repo root, where the relative `cargo build` and WASM paths do not
+resolve. It also skips `proposals`.
+
+- **Every run produces a new contract ID with empty state.** Balances,
+ treasury members, proposals and votes do not move across. Funds held by the
+ old `group_treasury` instance stay in the old instance.
+- **The scripts are hard-coded to `NETWORK="testnet"`.** There is no scripted
+ mainnet path. A mainnet deploy needs its own reviewed procedure.
+- **New IDs must be propagated by hand**, in this order:
+ 1. Backend environment: `TOKEN_TRANSFER_CONTRACT_ID`,
+ `GROUP_TREASURY_CONTRACT_ID`, and `STELLAR_RPC_URL` (not `RPC_URL`). Then
+ restart the backend.
+ 2. Web build environment: `NEXT_PUBLIC_TOKEN_TRANSFER_CONTRACT`. Then
+ **rebuild** the web app.
+
+ If the two sides disagree, the backend listener watches one contract while
+ users transact on another (guide §6).
+
+### Upgrading in place
+
+Only `token_transfer` exposes `upgrade(new_wasm_hash)`. It is admin-only and
+keeps the same contract ID and storage (guide §5.1). `group_treasury` and
+`proposals` have no upgrade entry point. **The only way to change their code is
+to deploy a new instance**, with all the consequences above.
+
+### Rollback: there is none
+
+- Contract transactions are final. Transfers, deposits, withdrawals and votes
+ executed against a bad contract cannot be reverted.
+- For `token_transfer`, you can call `upgrade` with the previous WASM hash.
+ That is a _roll forward to old code_, not a rollback: storage the new code
+ wrote stays as written. If the bad WASM breaks `upgrade` itself or the
+ admin key check, the contract cannot be changed again.
+- For `group_treasury` and `proposals`, "rolling back" means pointing the
+ apps at the old contract ID again. That works only if the old instance is
+ still valid and nobody has moved state to the new one.
+- Because of this, before any contract deploy that users depend on:
+ - run Contracts CI and `cargo test` on the exact commit;
+ - deploy to testnet first and exercise the flows in guide §5;
+ - record the WASM hash, contract ID and deployer account in the release notes.
+
+### Known script defects to fix before the next contract deploy
+
+- `deploy_group_treasury.sh` calls `initialize` without the required
+ `threshold` argument, so the invoke fails. **The instance is then deployed
+ but uninitialised, and anyone can call `initialize` on it and become admin.**
+ Do not publish that contract ID. Fix the script, or initialise the instance
+ immediately by hand.
+- `deploy_proposals.sh` requires `TREASURY_CONTRACT_ID` and `MIN_VOTES`, but
+ `initialize` only takes `admin`, so neither value is applied.
+
+## Rollback summary
+
+| App | Code rollback | State rollback | Notes |
+| --------- | --------------------------------------------------- | ------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------- |
+| Backend | Yes: redeploy the previous tag | **No.** Forward-fix or restore a backup (loses writes since the backup). | Safe only if migrations followed expand/contract. |
+| Web app | Yes: promote the previous build | n/a | Old builds keep old `NEXT_PUBLIC_*` values. Loaded tabs keep the new bundle until reload. |
+| AI agent | Yes: redeploy the previous tag | **No** for Weaviate vectors | An embedding-model change needs a re-index. |
+| Contracts | `token_transfer` only, via `upgrade` to an old hash | **None.** On-chain transactions are final. | `group_treasury` and `proposals` can only be replaced, not changed. |
+
+## Post-deploy verification
+
+Run these after every release. The [operator runbook](runbook.md) has the
+diagnosis and recovery steps for anything that fails.
+
+1. **Backend health.** `GET /health` returns `200` with `"db": "connected"` and
+ the new `version` from each instance. A `503` or `"db": "unreachable"` means
+ Postgres is unreachable. See
+ [runbook → Storage outage (Postgres)](runbook.md#storage-outage-postgres).
+2. **Redis adapter.** The boot logs show `[socket.io] Redis adapter attached`,
+ not `Redis unavailable … single-instance mode`. See
+ [runbook → Redis / message bus down](runbook.md#redis--message-bus-down).
+3. **Object storage.** `/health` does not check it. Do a manual upload and
+ download through the app, as described in
+ [runbook → Rotating VAPID / storage credentials](runbook.md#rotating-vapid--storage-credentials).
+ Also see
+ [runbook → Storage outage (S3-compatible object store)](runbook.md#storage-outage-s3-compatible-object-store).
+4. **Stellar listener.** The boot logs do _not_ show
+ `[stellar-listener] … listener disabled`, and a test transfer shows up in
+ the chat.
+5. **Metrics.** `GET /metrics` is being scraped, and the dashboards in
+ [observability.md](observability.md) show no spike in errors, disconnects
+ or backpressure after the rollout.
+6. **Web app.** Sign in, open a conversation, send a message, and check that
+ the socket connects over `wss://`.
+7. **AI agent.** `GET /health` returns `ok`, and one `/chat` call returns a
+ reply. That call is what proves `OPENAI_API_KEY` and outbound access work.