From 362bcd38fa87f223bb8ba793fe4f38cc0409882f Mon Sep 17 00:00:00 2001 From: josephchimebuka Date: Tue, 29 Sep 2026 00:45:26 +0000 Subject: [PATCH] docs: add backend concept pages and ADR index for #550 #556 #559 #561 Surface-level docs for configuration load order, Stellar listener reconciliation, backpressure tiers as actually wired, and the six already-made architecture decisions, so operators and reviewers have one place to look without duplicating the env-var catalog. Co-authored-by: Cursor --- IMPLEMENTATION_DOCS.md | 31 +++++ apps/backend/docs/concepts-backpressure.md | 100 ++++++++++++++ apps/backend/docs/concepts-configuration.md | 119 +++++++++++++++++ .../backend/docs/concepts-stellar-listener.md | 124 ++++++++++++++++++ docs/adr/0000-template.md | 31 +++++ docs/adr/0001-single-devices-table.md | 44 +++++++ .../0002-message-ordering-by-createdat-id.md | 46 +++++++ .../0003-ciphertext-only-message-storage.md | 46 +++++++ docs/adr/0004-per-device-envelopes.md | 52 ++++++++ .../0005-sealed-box-to-signal-migration.md | 54 ++++++++ docs/adr/0006-mls-for-groups.md | 51 +++++++ docs/adr/README.md | 29 ++++ 12 files changed, 727 insertions(+) create mode 100644 IMPLEMENTATION_DOCS.md create mode 100644 apps/backend/docs/concepts-backpressure.md create mode 100644 apps/backend/docs/concepts-configuration.md create mode 100644 apps/backend/docs/concepts-stellar-listener.md create mode 100644 docs/adr/0000-template.md create mode 100644 docs/adr/0001-single-devices-table.md create mode 100644 docs/adr/0002-message-ordering-by-createdat-id.md create mode 100644 docs/adr/0003-ciphertext-only-message-storage.md create mode 100644 docs/adr/0004-per-device-envelopes.md create mode 100644 docs/adr/0005-sealed-box-to-signal-migration.md create mode 100644 docs/adr/0006-mls-for-groups.md create mode 100644 docs/adr/README.md diff --git a/IMPLEMENTATION_DOCS.md b/IMPLEMENTATION_DOCS.md new file mode 100644 index 0000000..e483a81 --- /dev/null +++ b/IMPLEMENTATION_DOCS.md @@ -0,0 +1,31 @@ +# Implementation docs + +Index of the docs that describe how this tree actually behaves. Concept +and contract pages live next to the code they cover; this file is only +the map. + +## Decisions + +- [docs/adr/README.md](./docs/adr/README.md) — Architecture Decision + Records (devices table, message ordering, ciphertext-only storage, + per-device envelopes, sealed-box → Signal, MLS for groups). + +## Backend concepts + +- [apps/backend/docs/concepts-configuration.md](./apps/backend/docs/concepts-configuration.md) — + boot-validated vs lazy env, `loadEnv()`, object-store singleton. +- [apps/backend/docs/concepts-backpressure.md](./apps/backend/docs/concepts-backpressure.md) — + socket buffer tiers; only disconnect is enforced today. +- [apps/backend/docs/concepts-stellar-listener.md](./apps/backend/docs/concepts-stellar-listener.md) — + Soroban event watch, cursors, idempotency, RPC failure. +- [apps/backend/docs/concepts-gateway-architecture.md](./apps/backend/docs/concepts-gateway-architecture.md) +- [apps/backend/docs/concepts-delivery-fanout.md](./apps/backend/docs/concepts-delivery-fanout.md) +- [apps/backend/docs/concepts-storage-push-jobs.md](./apps/backend/docs/concepts-storage-push-jobs.md) + +## Operations and security + +- [docs/runbook.md](./docs/runbook.md) +- [docs/observability.md](./docs/observability.md) +- [docs/threat-model.md](./docs/threat-model.md) +- [docs/security/rate-limits.md](./docs/security/rate-limits.md) +- [`.env.example`](./.env.example) — environment-variable reference diff --git a/apps/backend/docs/concepts-backpressure.md b/apps/backend/docs/concepts-backpressure.md new file mode 100644 index 0000000..7aeafe5 --- /dev/null +++ b/apps/backend/docs/concepts-backpressure.md @@ -0,0 +1,100 @@ +# Backpressure and slow-consumer handling + +How `services/backpressure.ts` watches per-socket send-buffer occupancy, what +the two thresholds do **today**, and how a client recovers after a hard +disconnect. + +Gateway architecture mentions this module in passing +([concepts-gateway-architecture.md](./concepts-gateway-architecture.md) §8). +That overview currently describes shed as "stop sending". This page is the +accurate account of what is wired. + +## Monitoring + +On connect, `index.ts` calls `registerForBackpressure(socket)`. The socket is +added to an in-process set. A single `setInterval(checkBuffers, 5000)` runs +while any socket is registered and is cleared when the last one +unregisters (disconnect). + +Each tick reads the engine.io transport's `bufferedAmount` (bytes waiting +in the WebSocket send buffer). If that field is missing or throws, occupancy +is treated as `0` — the socket is not shed or disconnected on a read error. + +This is **per socket**, not per user. Two devices of the same account are +independent. The sets live in process memory; they are not shared across +gateway instances. + +## Two tiers + +| Tier | Env | Default | What the code does today | +| --- | --- | --- | --- | +| Shed | `SOCKET_SHED_THRESHOLD` | `32768` bytes | Marks the socket id in `shedSockets`, increments `clicked_backpressure_events_total{action="shed"}`, logs a warning. **Does not change emit behaviour.** | +| Disconnect | `SOCKET_BUFFER_THRESHOLD` | `65536` bytes | Same mark + metric `{action="disconnect"}`, then `socket.disconnect(true)`. **This is the only tier that affects delivery.** | + +Both env values are parsed on every tick (`parseInt`, must be a positive +integer). Invalid or unset values fall back to the defaults above. + +### Which tier changes emit behaviour + +`isSocketShed(socketId)` exists and is exported. **Nothing in the backend +calls it.** No dispatcher, fan-out, or `socket.emit` path consults +`shedSockets` before sending. + +So: + +- **Disconnect** is enforced. Crossing `SOCKET_BUFFER_THRESHOLD` force-closes + the socket. +- **Shed** is telemetry and a flag only. Crossing `SOCKET_SHED_THRESHOLD` + does **not** stop new events being queued. A reader of the gateway + overview must not infer that the server sheds load today — the buffer + will keep growing until the disconnect threshold (or the client catches + up and the flag is cleared on a later tick). + +When occupancy drops back below the shed threshold, the id is removed from +`shedSockets`. That only matters for the metric/flag, not for emits. + +## Hard disconnect — what the client sees + +`socket.disconnect(true)` is a server-initiated close (`close` / disconnect +on the client, reason typically `io server disconnect`). The socket is +removed from rooms, presence, the device-delivery subscriber, and +backpressure monitoring. In-flight emits for that connection are gone. + +The client is not told "you were too slow". There is no dedicated +backpressure event. From the client's point of view this is the same shape +as any other unexpected drop: the connection is dead and it must open a +new one and authenticate again. + +## How resume recovers + +A new connection is a new socket. After auth it re-joins conversation / +user / device rooms and is registered for backpressure from scratch +(`shedSockets` does not survive the old socket id). + +Missed **durable** chat is not in the WebSocket buffer and is not in the +resume stream. The client recovers it with envelope sync +(`GET /sync`, `syncRequired: true` on `resume_complete`). See +[api-messages-sync.md](./api-messages-sync.md) and +[concepts-delivery-fanout.md](./concepts-delivery-fanout.md). + +Missed **ephemeral** events (read/delivery receipts, presence, system +notices) are recovered by the resume/replay path: + +1. Client emits `resume { lastEventId }` with the last Redis stream id it + persisted. +2. The gateway reads `resume:events:${userId}` after that id (exclusive) + and emits `ephemeral_replay` for each entry still within the 300s TTL / + 500-entry cap. +3. It finishes with `resume_complete { lastEventId, syncRequired: true }`. + +Implementation: `services/resumeStream.ts`, handler in +`socket/messaging.ts`. Protocol notes: +[concepts-gateway-architecture.md](./concepts-gateway-architecture.md) §5 +and [api-websocket-events.md](./api-websocket-events.md) (`resume` / +`resume_complete` / `ephemeral_replay`). + +Events that were only sitting in the killed socket's send buffer and were +never recorded to Postgres or the resume stream are not replayed. That is +the loss window a slow consumer accepts when it is hard-disconnected. + +Cross-referenced from [`IMPLEMENTATION_DOCS.md`](../../../IMPLEMENTATION_DOCS.md). diff --git a/apps/backend/docs/concepts-configuration.md b/apps/backend/docs/concepts-configuration.md new file mode 100644 index 0000000..50a51b5 --- /dev/null +++ b/apps/backend/docs/concepts-configuration.md @@ -0,0 +1,119 @@ +# Backend configuration loading and validation + +How the gateway reads environment variables: which ones fail boot, which ones +are read later, and how that interacts with the object-store singleton. + +This page is about **when** a variable is consumed. The catalog of names, +defaults, and meanings lives in the environment-variable reference +([`.env.example`](../../../.env.example)). Rate-limit buckets and override +syntax are in [`docs/security/rate-limits.md`](../../../docs/security/rate-limits.md). +Do not treat the tables below as a second copy of those lists. + +## Boot-validated set (`src/config.ts`) + +`index.ts` calls `loadEnv()` immediately after `dotenv.config()`. `loadEnv()` +parses `process.env` against `EnvSchema` (Zod). On success it returns the +typed object and prints nothing. On failure it logs the offending names and +**exits the process with code 1** — the HTTP server is never bound. + +Required (missing or empty → hard startup failure): + +| Variable | Constraint | +| --- | --- | +| `DATABASE_URL` | non-empty string | +| `REDIS_URL` | non-empty string | +| `JWT_SECRET` | non-empty string | +| `PORT` | positive integer (string is coerced) | +| `TOKEN_TRANSFER_CONTRACT_ID` | non-empty string | +| `OBJECT_STORE_ENDPOINT` | non-empty string | +| `OBJECT_STORE_BUCKET` | non-empty string | +| `OBJECT_STORE_ACCESS_KEY` | non-empty string | +| `OBJECT_STORE_SECRET_KEY` | non-empty string | +| `OBJECT_STORE_REGION` | non-empty string | +| `OBJECT_STORE_FORCE_PATH_STYLE` | `true` / `false` / `1` / `0` | + +Optional in the same schema (validated only when present; absence does not +fail boot): `VAPID_PUBLIC_KEY`, `VAPID_PRIVATE_KEY`, `VAPID_SUBJECT`, the +legacy `S3_*` aliases, and `IDEMPOTENCY_TTL_SECONDS`. + +`assertTransportSecurityConfig()` runs next. That is a separate check +(`ENFORCE_TLS`, `ALLOWED_ORIGINS`) and is not part of `EnvSchema`. + +### Failure output + +If `DATABASE_URL` is missing, stderr looks like this and the process exits 1: + +```text +Missing or invalid environment variables: DATABASE_URL + - DATABASE_URL: DATABASE_URL is required +``` + +An empty environment lists every required name, one `- name: message` line +each. A non-numeric `PORT` is reported the same way (`PORT must be an integer`). + +## Lazy vs eager + +**Eager (boot).** `EnvSchema` variables above. The process does not start +without them. `index.ts` then immediately calls `getObjectStore()`, so the +object-store singleton is constructed on the boot path too — but the +construction itself is still lazy at the *module* level (see below). + +**Lazy (first use, no boot failure).** Everything else is read from +`process.env` at the call site, with a hardcoded default when unset or +malformed: + +- Rate-limit buckets in `src/config/rateLimits.ts` via `getRateLimitRule()`. + Each call re-reads `RATE_LIMIT_` (and the legacy + `SOCKET_RATE_LIMIT_PER_SEC` for `socket_default`). A malformed override is + warned and ignored; the default stands. `RATE_LIMIT_DISABLED=true` is a + kill switch, also read on demand. +- Backpressure thresholds (`SOCKET_BUFFER_THRESHOLD`, `SOCKET_SHED_THRESHOLD`) + — see [concepts-backpressure.md](./concepts-backpressure.md). +- Stellar RPC wiring (`STELLAR_RPC_URL`, `GROUP_TREASURY_CONTRACT_ID`) — + missing values disable the listener rather than crashing; see + [concepts-stellar-listener.md](./concepts-stellar-listener.md). +- Sync/envelope knobs such as `ENVELOPE_TTL_SECONDS` and `SYNC_PAGE_SIZE`. + +Rate-limit rules are deliberately not cached at import time so tests and the +runbook can change a limit without restarting the module graph. + +## `loadEnv()` and the object-store singleton + +`lib/objectStore.ts` keeps a process-wide singleton behind `getObjectStore()`. +The client is **not** built at import time. The first call constructs it: + +- `NODE_ENV === 'production'` → `createObjectStore(loadEnv())` (real S3 / + MinIO / R2 client, credentials from the boot-validated `OBJECT_STORE_*` + fields). +- otherwise → `getLocalObjectStore()` (fs-backed store; no live S3 needed). + +Building the singleton in `objectStore.ts` (instead of exporting a client +from `index.ts`) avoids a circular import: + +```text +index.ts → app.ts → routes → lib/storage.ts → index.ts +``` + +`storage.ts` and `services/fileCleanup.ts` both call `getObjectStore()`, so +there is exactly one client/bucket pair per process. `loadEnv()` is invoked +again inside that first production call; by then boot validation has already +run, so it is a typed re-parse, not a second chance to start with a broken +env. + +`resetObjectStoreForTests()` clears the memo so tests can change env. + +## Production vs development in `lib/storage.ts` + +`generatePresignedPut` / `generatePresignedGet` branch on +`NODE_ENV === 'production'`: + +| Branch | Implementation | +| --- | --- | +| production | `getObjectStore()` — the S3-compatible client from `OBJECT_STORE_*` | +| any other `NODE_ENV` | `getLocalObjectStore()` — local-disk store, URLs served by `routes/localStorage.ts` | + +The same `NODE_ENV` split exists inside `getObjectStore()` itself. The +storage helpers short-circuit to the local store in non-production so +presigned-URL callers never touch the S3 constructor on a laptop or in CI. + +Cross-referenced from [`IMPLEMENTATION_DOCS.md`](../../../IMPLEMENTATION_DOCS.md). diff --git a/apps/backend/docs/concepts-stellar-listener.md b/apps/backend/docs/concepts-stellar-listener.md new file mode 100644 index 0000000..16105bc --- /dev/null +++ b/apps/backend/docs/concepts-stellar-listener.md @@ -0,0 +1,124 @@ +# Stellar chain listener and on-chain reconciliation + +How `services/stellarListener.ts` watches Soroban contract events, writes +`token_transfers` / `treasury_proposals`, tracks its place in the ledger, and +behaves after downtime or an unreachable RPC. + +The listener is started from `index.ts` only when both `STELLAR_RPC_URL` and +`TOKEN_TRANSFER_CONTRACT_ID` are set. Missing either logs + +```text +[stellar-listener] STELLAR_RPC_URL or TOKEN_TRANSFER_CONTRACT_ID unset; listener disabled. +``` + +and leaves the API up. `GROUP_TREASURY_CONTRACT_ID` is optional: when set, a +second fetcher is attached for treasury events. + +`runForever` never rethrows. RPC and DB errors are logged inside the loop so +a dead chain cannot take the HTTP/WebSocket process down. + +## Watched events and database writes + +Two independent pollers share the same loop (default interval 5s). + +### `token_transfer` — topic `transfer` + +`buildRpcFetcher` calls Soroban RPC `getEvents` filtered to +`TOKEN_TRANSFER_CONTRACT_ID` and topic `transfer`. Each event is mapped to: + +| RPC field | Persisted column (`token_transfers`) | +| --- | --- | +| `txHash` | `tx_hash` (unique) | +| `value.to` | `recipient_address` | +| `value.amount` | `amount` (decimal string) | +| `value.memo` | `memo` (hex, optional) | + +`conversation_id` / `sender_id` are filled by decoding `memo` as a message +UUID and looking up `messages`. If that misses, the writer falls back to an +arbitrary existing conversation and user; if the database has neither, the +event is dropped. + +Write path: `INSERT … ON CONFLICT (tx_hash) DO UPDATE SET created_at = now()`. +A replayed ledger page refreshes `created_at` on the same row; it does not +insert a second transfer. + +### `group_treasury` — proposal lifecycle + +`buildTreasuryRpcFetcher` watches `GROUP_TREASURY_CONTRACT_ID` for: + +| Contract event | Row status | +| --- | --- | +| `proposal_created` | `active` | +| `proposal_approved` | `approved` | +| `proposal_rejected` | `rejected` | +| `proposal_executed` | `executed` | +| `proposal_expired` | `expired` | + +Persistence is an **update** on `(contract_id, proposal_id)` (unique index +`treasury_proposals_contract_proposal_idx`), copying `approvals` / +`rejections` when present. If no matching row exists, the event is ignored +(`if (!row) return`) — the listener does not insert treasury proposals. After +a successful update it emits `treasury_proposal_updated` to the linked +conversation Socket.IO room. + +## Cursor / position tracking + +Both pollers keep an **in-process** cursor (`pagingToken` from the last +successfully persisted event). There is no cursor table and nothing is +written to Redis or Postgres. + +- First poll after process start: `cursor` is `null`. `getEvents` is called + with `startLedger` and `cursor` both unset (the fetcher comment: "resume + on cursor only"). +- After each successful persist, the cursor advances to that event's + `pagingToken`. A persist failure leaves the cursor unmoved so the next + poll retries the same page. +- Token-transfer and treasury streams have separate cursors. + +Cursors die with the process. A restart is a cold start: both cursors are +`null` again. + +## Catch-up after downtime + +Two different "down" cases: + +1. **Process was down (restart / deploy).** Cursors reset. The next + `getEvents` page is whatever the RPC returns without a cursor. Events + still inside the RPC's retention window are re-delivered and upserted. + Events that have aged out of that window are not replayed by this + listener — those mirrored rows stay missing until something else writes + them. +2. **Process stayed up but a poll failed.** `consecutiveFailures` + increments. The loop waits `min(1000 * 2^(n-1), 30000)` ms and retries + **from the last good in-memory cursor**, so it continues where it left + off rather than rewinding. + +## Idempotency + +A replayed ledger event must not double-write: + +- Token transfers: unique `tx_hash` + `ON CONFLICT DO UPDATE`. The same + hash can be persisted any number of times; there is still one row. +- Treasury proposals: unique `(contract_id, proposal_id)`. A repeated + event updates status / counts on the existing row. + +Because the cursor only advances after a successful persist, a crash +mid-page can re-offer the same events; the unique keys absorb that. + +## RPC unreachable — drift from on-chain state + +When `getEvents` throws (timeout, 5xx, network partition): + +- The error is logged as `fetch failed; reconnecting after backoff`. +- The API and WebSocket server keep running. +- No new rows are written for the duration of the outage. + +**Yes, on-chain state can drift from the mirrored tables.** The listener is +a best-effort projection, not a consensus participant. During an RPC outage +the chain continues; `token_transfers` and `treasury_proposals` do not. +After reconnect, catch-up depends on RPC retention and on the treasury +update-only rule (a `proposal_*` event for a row that was never inserted +is silently skipped). There is no compensating backfill job in this +service. + +Cross-referenced from [`IMPLEMENTATION_DOCS.md`](../../../IMPLEMENTATION_DOCS.md). diff --git a/docs/adr/0000-template.md b/docs/adr/0000-template.md new file mode 100644 index 0000000..1b7e34a --- /dev/null +++ b/docs/adr/0000-template.md @@ -0,0 +1,31 @@ +# ADR NNNN: Title + +- Status: proposed | accepted | superseded by ADR NNNN | deprecated +- Date: YYYY-MM-DD +- Deciders: (optional) + +## Context + +What problem is being decided, and what constraints matter. Point at the +code, schema comment, or review thread that made the decision necessary. + +## Decision + +The choice that was made, in one short paragraph. Name the current +implementation (table, column, module) so a later reader can verify it. + +## Alternatives rejected + +What was considered and not chosen, and **why**. If the repo does not +record a rationale, say so explicitly rather than inventing one. + +## Consequences + +What becomes easier, what becomes harder, and what follow-up is now +required (or explicitly deferred). + +## Status + +`accepted` means the tree still implements this. `superseded` means a +later ADR replaced it; link that ADR and leave this file in place so the +history stays readable. diff --git a/docs/adr/0001-single-devices-table.md b/docs/adr/0001-single-devices-table.md new file mode 100644 index 0000000..abdee09 --- /dev/null +++ b/docs/adr/0001-single-devices-table.md @@ -0,0 +1,44 @@ +# ADR 0001: One canonical `devices` table + +- Status: accepted +- Date: 2026-09-29 + +## Context + +Auth, prekeys, realtime delivery, push, and MLS all need a stable device +identity. An earlier split stored some of that on a separate +`user_devices` table. Two registries can silently disagree about whether +a device exists, is revoked, or owns a given identity key. + +`apps/backend/src/db/schema.ts` and `routes/devices.ts` record the merge: +`devices` is the single canonical device registry. `messages.senderDeviceId`, +`messageEnvelopes.recipientDeviceId`, and `pushSubscriptions.deviceId` all +FK here. `revokedAt` keeps the row for audit instead of deleting it. + +## Decision + +One `devices` table per physical client. Identity public key, platform, +`lastSeenAt`, `pushEnabled`, `capabilities`, and revocation live on that +row. `GET /user-devices/:id/public-key` remains as a lookup alias; it is +not a second store (`routes/userDevices.ts`). + +## Alternatives rejected + +**A separate `user_devices` table** (auth/prekeys vs realtime/push). +Rejected because the two copies drifted: a device could be live for +messaging and missing for prekeys, or revoked in one table and not the +other. Issue #103 / #107 and the header on `routes/devices.ts` are the +review record of that merge. + +**Deleting revoked rows.** Rejected so fingerprint history, envelope +foreign keys, and "when was this revoked" stay queryable. `revokedAt` plus +the GC `staleFlaggedAt` flag cover retirement without a hard delete. + +## Consequences + +Every device-shaped feature adds columns or child tables that FK +`devices`, not a parallel registry. Linking a new device and revoking an +old one are the same row lifecycle. Clients must treat `/user-devices` as +a compatibility path only. + +See [api-devices.md](../../apps/backend/docs/api-devices.md). diff --git a/docs/adr/0002-message-ordering-by-createdat-id.md b/docs/adr/0002-message-ordering-by-createdat-id.md new file mode 100644 index 0000000..84ee730 --- /dev/null +++ b/docs/adr/0002-message-ordering-by-createdat-id.md @@ -0,0 +1,46 @@ +# ADR 0002: Order messages by `(createdAt, id)`, not `sequenceNumber` + +- Status: accepted +- Date: 2026-09-29 + +## Context + +Offline sync and in-conversation display both need a total order. A +per-conversation monotonic `sequenceNumber` was implemented briefly, then +dropped. The schema comment on `messages` and `routes/sync.ts` state why: +the same counter cannot also be a coherent **cross-conversation** cursor +for a device catching up on every chat it belongs to, and keeping two +counters was judged not worth the complexity. + +`id` (UUID) is only a same-millisecond tiebreaker, not a clock. + +## Decision + +Order and paginate by `(createdAt, id)`. Envelope sync encodes that pair +as the cursor (`createdAt.getTime():id`). `GET /sync` does not emit +`sequenceNumber`. Group *control* events still use an epoch/sequence log +on the conversation row; that is a different problem (see +[group-epoch-sync.md](../group-epoch-sync.md)). + +## Alternatives rejected + +**Per-conversation `sequenceNumber` as the only order key.** Rejected +because a device's offline cursor has to move across many conversations. +A number that restarts per chat cannot be compared globally. Maintaining +a second, cross-conversation counter alongside it was rejected as +duplicated machinery for the same `createdAt` fact the row already has. + +(The schema comment cites `#137` / `#128` for the sync argument; those +issue numbers in comments do not all match their original tickets. The +behaviour is what the code does.) + +## Consequences + +Clients must sort and resume with timestamps + ids. The web app still +types `sequenceNumber` in several places and falls back to `0` +(DRIFT-1 in +[contracts-response-types.md](../../apps/web/docs/contracts-response-types.md)); +that is a known client bug, not a revival of the column. + +See `apps/backend/src/db/schema.ts` (`messages`) and +`apps/backend/src/routes/sync.ts`. diff --git a/docs/adr/0003-ciphertext-only-message-storage.md b/docs/adr/0003-ciphertext-only-message-storage.md new file mode 100644 index 0000000..854f34a --- /dev/null +++ b/docs/adr/0003-ciphertext-only-message-storage.md @@ -0,0 +1,46 @@ +# ADR 0003: Ciphertext-only message storage + +- Status: accepted +- Date: 2026-09-29 + +## Context + +`messages` once had a plaintext `content` column (and a GIN index for +server-side search) beside the E2EE work. The server cannot search +ciphertext it cannot read. The cutover is recorded in +[message-encryption-migration.md](../../apps/backend/docs/message-encryption-migration.md). + +## Decision + +Live `messages` rows store `ciphertext` (and optional per-device +envelopes). There is no plaintext `content` column. Pre-cutover plaintext +was copied to `message_content_archive` and then purged from `messages` +(archive-then-purge). `GET /conversations/:id/search` returns `410 Gone`. +Search is client-side over decrypted history. + +System rows use `systemPayload` JSON, never ciphertext; a check constraint +keeps the two from mixing. + +## Alternatives rejected + +**In-place tombstone** (keep a nulled `content` / `legacyPlaintext` column +forever). Rejected because every future query and migration would carry a +legacy-shaped column for a closed set of old rows, and a future path could +accidentally serve that field as if it were ciphertext. + +**Re-encrypting old plaintext to look like E2EE.** Rejected: archived rows +with no ciphertext surface as `{ unavailable: true }`. The product does +not backdate an "encrypted" claim. + +**Server-side search over ciphertext.** Not implemented; there is no +searchable-encryption scheme in this tree. The 410 is the recorded +refusal. + +## Consequences + +Hot-path serializers only have a ciphertext (or envelope, or unavailable) +branch. The archive table has no API. Legal/compliance access to old +plaintext is out of band. + +See also ADR 0004 (how ciphertext is keyed per device) and ADR 0006 +(group ciphertext via MLS). diff --git a/docs/adr/0004-per-device-envelopes.md b/docs/adr/0004-per-device-envelopes.md new file mode 100644 index 0000000..d677b45 --- /dev/null +++ b/docs/adr/0004-per-device-envelopes.md @@ -0,0 +1,52 @@ +# ADR 0004: Per-device envelopes over a shared ciphertext + +- Status: accepted +- Date: 2026-09-29 + +## Context + +A user has many devices. Each device has its own identity key. A message +the server stores as one blob encrypted to a conversation-wide secret +would let any device (or a leak of that secret) read everything, and +would not match the Phase-1 sealed-box model (one independent box per +recipient device). + +Fan-out helpers in `lib/messageFanout.ts` require one +`message_envelopes` row per recipient device in the same transaction as +the `messages` insert. + +## Decision + +DM / sealed-box / Signal sends carry **per-device envelopes**: +`recipientDeviceId` + `ciphertext` (+ `protocol`). Sibling-device +coverage is checked on send (`fetchSiblingDeviceIds`). The conversation +room may get a ciphertext-free `new_message` for UI; the decryptable +bytes go to `device:${deviceId}`. + +MLS group messages are the documented exception: one group ciphertext +on `messages`, no per-device envelopes (ADR 0006). Mixing `envelopes` +with `mlsEpoch` is rejected in `validateMessagePayload`. + +## Alternatives rejected + +**One shared conversation ciphertext (static group key).** Rejected for +1:1 and multi-device DMs: every device would share a decrypt key, add +device / revoke device would require rotating that key for all history +or leaving old devices able to read new mail, and it collapses to a +server-visible or single-leak plaintext if the key is ever exposed. The +repo does not record a long design thread against a static shared key +beyond the envelope requirement itself; the implemented model is +per-device boxes. + +**Encrypting only to the recipient user, not each of their devices.** +Rejected: a newly linked laptop would not be able to decrypt a message +already stored for the phone. Sibling envelopes exist so every active +device of the sender (and each recipient) can open the message locally +(#188). + +## Consequences + +Send payload size grows with the device set. A missing sibling envelope +fails the send rather than silently dropping a device. Group chats that +need one copy of the bytes use MLS instead of multiplying envelopes +(ADR 0006). diff --git a/docs/adr/0005-sealed-box-to-signal-migration.md b/docs/adr/0005-sealed-box-to-signal-migration.md new file mode 100644 index 0000000..fc34430 --- /dev/null +++ b/docs/adr/0005-sealed-box-to-signal-migration.md @@ -0,0 +1,54 @@ +# ADR 0005: Phase-1 sealed box → Phase-2 Signal, per device pair + +- Status: accepted +- Date: 2026-09-29 + +## Context + +The product shipped on a Phase-1 sealed box: ECDH ephemeral key + HKDF + +AES-256-GCM, one box per recipient device per message. No ratchet, no +forward secrecy. Phase 2 is the Signal Double Ratchet +([signal-integration.md](../signal-integration.md), +[signal-migration.md](../../apps/backend/docs/signal-migration.md)). + +Clients upgrade at different times. A conversation will contain both +protocols until the last device moves. + +## Decision + +- **History is never re-encrypted.** `message_envelopes.protocol` + (`sealed_box` | `signal` | `mls`, default `sealed_box`) records what + actually produced that ciphertext. Old rows keep decrypting on the + Phase-1 path. +- **Negotiation is per device pair**, not per conversation. + `selectProtocol` picks the strongest protocol both devices advertise + in `devices.capabilities`. A Signal-capable pair is not held back + because another member is still on sealed box. +- A sender claiming a weaker protocol than both sides support is + refused (stops a client from pinning a peer on sealed box forever). + +The web Signal adapter (`signalClient.ts`) still throws if invoked; the +migration path is accepted in the data model and not fully activated in +the running client. + +## Alternatives rejected + +**Conversation-wide readiness gate** ("everyone must advertise Signal +before anyone uses it"). Rejected: it delays forward secrecy for every +pair until the slowest member upgrades, which in a large group may be +never. + +**Re-encrypt existing history on cutover.** Rejected: the server has no +plaintext, and rewriting history would break devices that only know +sealed box. The `protocol` column exists so history stays labelled. + +**Inferring the construction from ciphertext bytes or send time.** +Rejected: capabilities change; last month's envelope is not "whatever +this device supports now". + +## Consequences + +Capability advertisement and the per-envelope `protocol` column are both +required. Rollout can be incremental. Pairwise Signal and group MLS +(ADR 0006) are different constructions and must not be mixed on one +message (ADR 0004). diff --git a/docs/adr/0006-mls-for-groups.md b/docs/adr/0006-mls-for-groups.md new file mode 100644 index 0000000..9331c8e --- /dev/null +++ b/docs/adr/0006-mls-for-groups.md @@ -0,0 +1,51 @@ +# ADR 0006: MLS for groups + +- Status: accepted +- Date: 2026-09-29 + +## Context + +Per-device envelopes (ADR 0004) scale linearly with members × devices. +Group file sharing and group text need one ciphertext the epoch's +members can open, without giving the server the key. The chosen protocol +is MLS (RFC 9420) via `@openmls/wasm` on the client; the backend stores +public group state only +([mls-group-membership.md](../../apps/backend/docs/mls-group-membership.md), +[mls-integration.md](../../apps/web/src/lib/mls-integration.md)). + +## Decision + +Group conversations use MLS. The server holds `mls_groups`, +`mls_group_members` (membership as an epoch interval), `mls_commits`, +`mls_welcomes`, and `mls_key_packages`. `messages.mls_epoch` names the +epoch that encrypted `ciphertext`. A newly joined device cannot decrypt +pre-join epochs; those messages render as unavailable, not as errors. + +File keys for group files ride inside that single MLS message so object +storage stores one object, not one copy per device +([mls-group-files.md](../../apps/backend/docs/mls-group-files.md)). + +## Alternatives rejected + +**Per-device envelope fan-out for group content.** Rejected on cost and +churn: "in a 30-person group with 3 devices each that is 90 copies of +the same bytes, and it grows again on every new device." + +**Sender keys / a static group key on the server.** Rejected: the server +would be in the key-distribution path, and removal would not have MLS's +epoch forward secrecy. Not implemented; no extra rationale is recorded +beyond the MLS membership doc. + +**Re-encrypting history to a joining device.** Deferred (would need an +existing member to resend plaintext). Until that exists, "no history for +a new device" is the behaviour. + +**Mixing `envelopes` and `mlsEpoch` on one message.** Rejected in +`validateMessagePayload` so recipients have one rule for which bytes to +decrypt. + +## Consequences + +DMs stay on sealed box / Signal (ADR 0004, ADR 0005). Groups accept +weaker history for new devices in exchange for one ciphertext and +epoch-based membership. The backend remains an untrusted relay. diff --git a/docs/adr/README.md b/docs/adr/README.md new file mode 100644 index 0000000..ef39ca7 --- /dev/null +++ b/docs/adr/README.md @@ -0,0 +1,29 @@ +# Architecture Decision Records + +Short records of decisions this repo has already made and re-litigated in +review. New decisions use [0000-template.md](./0000-template.md) +(context / decision / alternatives rejected / consequences / status). + +Cross-referenced from [`IMPLEMENTATION_DOCS.md`](../../IMPLEMENTATION_DOCS.md). + +## Index + +| ADR | Title | Status | +| --- | --- | --- | +| [0001](./0001-single-devices-table.md) | One canonical `devices` table | accepted | +| [0002](./0002-message-ordering-by-createdat-id.md) | `(createdAt, id)` ordering; no `sequenceNumber` | accepted | +| [0003](./0003-ciphertext-only-message-storage.md) | Ciphertext-only message storage | accepted | +| [0004](./0004-per-device-envelopes.md) | Per-device envelopes over a shared ciphertext | accepted | +| [0005](./0005-sealed-box-to-signal-migration.md) | Phase-1 sealed box → Phase-2 Signal | accepted | +| [0006](./0006-mls-for-groups.md) | MLS for groups | accepted | + +None of these records is superseded. If a later change replaces one, set +that file to `superseded` and point at the new ADR; leave the old file in +this list. + +## Adding a record + +1. Copy `0000-template.md` to `NNNN-short-title.md` (next free number). +2. Fill every section, including **what was rejected and why**. +3. Add a row here with status `accepted` or `superseded`. +4. Link the ADR from the code or concept doc that implements it.