Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion .github/workflows/backup.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ on:
- full
- incremental
- verify
- drill

env:
BACKUP_DIR: /tmp/backups
Expand Down Expand Up @@ -82,6 +83,18 @@ jobs:
export DATABASE_URL="postgresql://postgres:postgres@localhost:5432/agenticpay"
bash scripts/backup.sh verify

- name: Run isolated restore drill
if: ${{ github.event.inputs.type == 'drill' }}
run: |
export DATABASE_URL="postgresql://postgres:postgres@localhost:5432/agenticpay"
createdb -h localhost -U postgres agenticpay_drill
psql "$DATABASE_URL" -v ON_ERROR_STOP=1 -c 'CREATE TABLE IF NOT EXISTS dr_evidence (id integer primary key);'
bash scripts/backup.sh full
export DRILL_DATABASE_URL="postgresql://postgres:postgres@localhost:5432/agenticpay_drill"
bash scripts/disaster-recovery-drill.sh
env:
PGPASSWORD: postgres

- name: Upload backup artifacts
uses: actions/upload-artifact@v4
with:
Expand All @@ -101,4 +114,4 @@ jobs:
}]
}
env:
SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK_URL }}
SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK_URL }}
41 changes: 41 additions & 0 deletions .github/workflows/k6-load-test.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
name: k6 Load Test

on:
workflow_dispatch:
inputs:
base_url:
description: Isolated or staging AgenticPay URL
required: true
type: string
target_rate:
description: Requests per second at the steady stage
required: false
default: '25'
type: string

permissions:
contents: read

jobs:
load-test:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: grafana/setup-k6-action@v1
- name: Run health smoke test
env:
BASE_URL: ${{ inputs.base_url }}
run: k6 run tests/load/k6/smoke.js
- name: Run staged payment load test
env:
BASE_URL: ${{ inputs.base_url }}
TARGET_RATE: ${{ inputs.target_rate }}
run: k6 run --summary-export=k6-summary.json tests/load/k6/payment-flow.js
- name: Upload load-test summary
if: always()
uses: actions/upload-artifact@v4
with:
name: k6-summary
path: k6-summary.json
if-no-files-found: ignore
11 changes: 8 additions & 3 deletions backend/docs/COHORT_ANALYTICS.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Subscription Cohort Retention Analytics

Issue #629. In-memory, event-sourced cohort analytics for on-chain-managed
Issues #629 and #851. In-memory, event-sourced cohort analytics for on-chain-managed
subscriptions (see `backend/src/jobs/subscription.service.ts`). There is no
Prisma-persisted subscription/customer table in this codebase — subscriptions
are managed on-chain — so this service, like `backend/src/services/analytics.ts`
Expand Down Expand Up @@ -195,6 +195,12 @@ List all cohort months present with their size.
}
```

### `GET /api/v1/analytics/cohorts/retention`

Returns every cohort as an aligned retention matrix suitable for a heatmap.
Offsets without enough observed history are `null`, not zero, so future months
are not misclassified as complete churn.

### `GET /api/v1/analytics/cohorts/:cohortMonth/retention`

`cohortMonth` must match `YYYY-MM`.
Expand Down Expand Up @@ -285,8 +291,7 @@ cohortMonth,monthOffset,activeCustomers,retentionPct

## Mounting

This route module is not wired into `src/index.ts` by this change (per task
scope — index.ts is owned by another workstream). To enable it, mount:
The route is mounted in `src/index.ts` at:

```ts
import { cohortAnalyticsRouter } from './routes/cohort-analytics.js';
Expand Down
7 changes: 7 additions & 0 deletions backend/src/routes/cohort-analytics.ts
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,13 @@ cohortAnalyticsRouter.get(
}),
);

cohortAnalyticsRouter.get(
'/retention',
asyncHandler(async (_req, res) => {
res.json({ data: cohortAnalyticsService.getRetentionMatrix() });
}),
);

cohortAnalyticsRouter.get(
'/:cohortMonth/retention',
asyncHandler(async (req, res) => {
Expand Down
16 changes: 16 additions & 0 deletions backend/src/routes/reports.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@ import { Router } from 'express';
import { asyncHandler } from '../middleware/errorHandler.js';
import { customReportService } from '../services/reports/custom-report.js';
import { AppError } from '../middleware/errorHandler.js';
import { exportReportRows, type ReportExportFormat } from '../services/reports/report-export.js';

export const reportsRouter = Router();

Expand Down Expand Up @@ -78,3 +79,18 @@ reportsRouter.get('/:id/data', asyncHandler(async (req, res) => {
const data = await customReportService.generateReportData(req.params.id);
res.json(data);
}));

reportsRouter.get('/:id/export', asyncHandler(async (req, res) => {
const format = req.query.format as ReportExportFormat | undefined;
if (format !== 'csv' && format !== 'json' && format !== 'excel') {
throw new AppError(400, 'format must be csv, json, or excel', 'INVALID_EXPORT_FORMAT');
}

const generated = await customReportService.generateReportData(req.params.id);
const exported = exportReportRows(generated.data as Record<string, unknown>[], format);
const safeName = generated.report.name.replace(/[^a-zA-Z0-9_-]+/g, '-').replace(/^-|-$/g, '') || 'report';

res.setHeader('Content-Type', exported.contentType);
res.setHeader('Content-Disposition', `attachment; filename="${safeName}.${exported.extension}"`);
res.send(exported.content);
}));
10 changes: 10 additions & 0 deletions backend/src/services/__tests__/cohort-analytics.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -148,6 +148,16 @@ describe('CohortAnalyticsService', () => {
expect(comparison.summary.averageRetentionPctByCohort['2025-01']).toBe(100);
expect(comparison.summary.averageRetentionPctByCohort['2025-02']).toBe(75);
});

it('builds an aligned retention matrix without treating unavailable months as churn', () => {
const matrix = service.getRetentionMatrix();

expect(matrix.maxMonthOffset).toBe(1);
expect(matrix.cohorts).toEqual([
{ cohortMonth: '2025-01', cohortSize: 1, retentionByMonth: [100, 100] },
{ cohortMonth: '2025-02', cohortSize: 2, retentionByMonth: [100, 50] },
]);
});
});

describe('exportToCsv', () => {
Expand Down
35 changes: 35 additions & 0 deletions backend/src/services/cohort-analytics.ts
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,10 @@ export interface CohortSummary {
cohortSize: number;
}

export interface RetentionMatrixRow extends CohortSummary {
retentionByMonth: Array<number | null>;
}

export interface RetentionPoint {
monthOffset: number;
activeCustomers: number;
Expand Down Expand Up @@ -127,6 +131,37 @@ export class CohortAnalyticsService {
.sort((a, b) => a.cohortMonth.localeCompare(b.cohortMonth));
}

/**
* Returns an aligned retention matrix for dashboard heatmaps. Missing future
* offsets are represented by null rather than zero so they are not mistaken
* for churn.
*/
getRetentionMatrix(): { maxMonthOffset: number; cohorts: RetentionMatrixRow[] } {
const summaries = this.getCohorts();
const curves = summaries.map((summary) => ({
summary,
curve: this.getRetentionCurve(summary.cohortMonth),
}));
const maxMonthOffset = curves.reduce(
(max, { curve }) => Math.max(max, curve.at(-1)?.monthOffset ?? 0),
0,
);

return {
maxMonthOffset,
cohorts: curves.map(({ summary, curve }) => {
const byOffset = new Map(curve.map((point) => [point.monthOffset, point.retentionPct]));
return {
...summary,
retentionByMonth: Array.from(
{ length: maxMonthOffset + 1 },
(_, offset) => byOffset.get(offset) ?? null,
),
};
}),
};
}

/**
* Retention curve for a cohort: for each month-offset N, the % of the cohort's
* original customers still active in that offset month.
Expand Down
33 changes: 33 additions & 0 deletions backend/src/services/reports/report-export.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
import { describe, expect, it } from 'vitest';
import { exportReportRows } from './report-export.js';

const rows = [
{ merchant: 'Alice, Inc.', payments: 2, note: 'quoted "value"' },
{ merchant: 'Björk & Co', payments: 1, note: '=HYPERLINK("https://example.com")' },
];

describe('exportReportRows', () => {
it('exports valid JSON', () => {
const result = exportReportRows(rows, 'json');
expect(result.extension).toBe('json');
expect(JSON.parse(result.content)).toEqual(rows);
});

it('exports RFC-compatible CSV with an Excel UTF-8 BOM', () => {
const result = exportReportRows(rows, 'csv');
expect(result.contentType).toContain('text/csv');
expect(result.content.startsWith('\uFEFFmerchant,payments,note')).toBe(true);
expect(result.content).toContain('"Alice, Inc."');
expect(result.content).toContain('"quoted ""value"""');
expect(result.content).toContain("'=HYPERLINK");
});

it('exports an Excel-readable SpreadsheetML workbook with escaped cells', () => {
const result = exportReportRows(rows, 'excel');
expect(result.extension).toBe('xls');
expect(result.contentType).toContain('application/vnd.ms-excel');
expect(result.content).toContain('ss:Type="Number">2</Data>');
expect(result.content).toContain('Björk &amp; Co');
expect(result.content).toContain('&apos;=HYPERLINK');
});
});
85 changes: 85 additions & 0 deletions backend/src/services/reports/report-export.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
export type ReportExportFormat = 'csv' | 'json' | 'excel';

export interface ReportExportResult {
content: string;
contentType: string;
extension: 'csv' | 'json' | 'xls';
}

type ReportRow = Record<string, unknown>;

function stringifyCell(value: unknown): string {
if (value === null || value === undefined) return '';
if (value instanceof Date) return value.toISOString();
if (typeof value === 'object') return JSON.stringify(value);
return String(value);
}

function neutralizeSpreadsheetFormula(value: unknown): string {
const text = stringifyCell(value);
return /^[=+\-@]/.test(text) ? `'${text}` : text;
}

function escapeCsv(value: unknown): string {
const text = neutralizeSpreadsheetFormula(value);
return /[",\r\n]/.test(text) ? `"${text.replace(/"/g, '""')}"` : text;
}

function escapeXml(value: unknown, neutralizeFormula = false): string {
return (neutralizeFormula ? neutralizeSpreadsheetFormula(value) : stringifyCell(value))
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;')
.replace(/'/g, '&apos;');
}

function columnsFor(rows: ReportRow[]): string[] {
const columns = new Set<string>();
for (const row of rows) {
for (const key of Object.keys(row)) columns.add(key);
}
return [...columns];
}

/** Serialize report rows without evaluating formulas or emitting unsafe HTML. */
export function exportReportRows(rows: ReportRow[], format: ReportExportFormat): ReportExportResult {
const columns = columnsFor(rows);

if (format === 'json') {
return {
content: JSON.stringify(rows, null, 2),
contentType: 'application/json; charset=utf-8',
extension: 'json',
};
}

if (format === 'csv') {
const lines = [columns.map(escapeCsv).join(',')];
for (const row of rows) lines.push(columns.map((column) => escapeCsv(row[column])).join(','));
return {
// UTF-8 BOM keeps non-ASCII data intact when opened directly in Excel.
content: `\uFEFF${lines.join('\r\n')}`,
contentType: 'text/csv; charset=utf-8',
extension: 'csv',
};
}

const header = columns.map((column) => `<Cell><Data ss:Type="String">${escapeXml(column)}</Data></Cell>`).join('');
const body = rows
.map((row) => {
const cells = columns.map((column) => {
const value = row[column];
const numeric = typeof value === 'number' && Number.isFinite(value);
return `<Cell><Data ss:Type="${numeric ? 'Number' : 'String'}">${escapeXml(value, !numeric)}</Data></Cell>`;
}).join('');
return `<Row>${cells}</Row>`;
})
.join('');

return {
content: `<?xml version="1.0"?><?mso-application progid="Excel.Sheet"?><Workbook xmlns="urn:schemas-microsoft-com:office:spreadsheet" xmlns:ss="urn:schemas-microsoft-com:office:spreadsheet"><Worksheet ss:Name="Report"><Table><Row>${header}</Row>${body}</Table></Worksheet></Workbook>`,
contentType: 'application/vnd.ms-excel; charset=utf-8',
extension: 'xls',
};
}
64 changes: 64 additions & 0 deletions docs/DISASTER_RECOVERY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Disaster recovery runbook

This runbook defines the recovery process for the AgenticPay PostgreSQL data
plane. The objectives are an RPO of one hour and an RTO of four hours. Incident
commanders must record actual recovery times and data loss in the incident log.

## Preparation and ownership

- Primary owner: on-call platform engineer.
- Approver for production restore: incident commander plus database owner.
- Full backups run daily; incremental backups run every six hours. Production
deployments should use WAL archiving to meet the one-hour RPO.
- Backups require SHA-256 verification before upload or restore.
- Quarterly restore drills must target an isolated database and preserve the
workflow artifact/log as evidence.

## Incident procedure

1. Declare the incident, freeze database migrations and payment writes, and
record the suspected corruption time.
2. Confirm whether the primary can be recovered in place. Prefer failover to a
healthy replica for infrastructure-only failures.
3. List and verify restore candidates:

```bash
BACKUP_DIR=/var/backups/agenticpay bash scripts/backup.sh verify
```

4. Select the newest verified full backup before the incident, then either:

```bash
DATABASE_URL="$RECOVERY_DATABASE_URL" bash scripts/backup.sh restore /path/to/full_backup.sql.gz
DATABASE_URL="$RECOVERY_DATABASE_URL" bash scripts/backup.sh pitr '2026-09-25 10:00:00'
```

5. Validate migrations, table counts, a read-only payment query, queue health,
and reconciliation totals before directing traffic to the recovered system.
6. Rotate credentials exposed during response, re-enable writes gradually, and
monitor payment failures, lag, and reconciliation mismatches.
7. Publish the achieved RPO/RTO, missing transactions, and follow-up actions.

## Automated restore drill

The drill refuses to run when source and target URLs are identical:

```bash
DATABASE_URL="$SOURCE_DATABASE_URL" \
DRILL_DATABASE_URL="$ISOLATED_DATABASE_URL" \
BACKUP_DIR=/tmp/agenticpay-drill \
bash scripts/disaster-recovery-drill.sh
```

Use the `Database Backup` workflow's `drill` dispatch option for CI evidence.
The target database is disposable; never use a production connection string as
`DRILL_DATABASE_URL`.

## Recovery verification checklist

- [ ] Backup gzip and SHA-256 validation passed.
- [ ] Restore used an isolated target first.
- [ ] Schema and migration state match the release being recovered.
- [ ] Payment totals reconcile against ledger/provider records.
- [ ] Background queues and webhook delivery resume without duplication.
- [ ] RPO and RTO measurements are attached to the incident.
Loading
Loading