Category: DevOps, CI/CD & Observability
Description: uptime-check.yml pings URLs every 30 minutes. Replace/extend it with a scheduled Playwright synthetic that exercises a real user path on staging (load, feed renders, WebSocket connects, quote fetch) and alerts on failure with deduplicated GitHub issues and an optional chat webhook.
Problem Statement & Context: A 200 OK from the homepage says nothing about whether the swap card can fetch quotes or the live feed connects — the failures users actually experience.
Scope & Acceptance Criteria:
- Scheduled workflow (every 15 min) running a small Playwright suite against configurable URLs: page loads without console errors,
/api/health OK, the feed shows items (or a valid empty state), WebSocket reaches open within N seconds (observed via the UI "Live" indicator), quote request for a canned route returns within budget.
- Failure handling: retry once, then open (or update) a single GitHub issue labelled
incident with logs/trace links, and optionally notify a Slack/Discord webhook secret; auto-close/comment on recovery; avoids alert storms via dedupe.
- Timing metrics appended to a JSON artifact/
gh-pages branch to generate a simple static status/latency page (uptime %, last 30 days) published via GitHub Pages.
- Least-privilege token usage; runs in read-only mode (no signing, no submissions); cost/time limits documented.
- Out of scope: paging/on-call rotation.
Implementation Guidelines (Suggested Execution):
- Key Files/Modules:
.github/workflows/uptime-check.yml, new e2e/synthetic/*.spec.ts, new scripts/status-page/*, README.md (status badge).
- Design/Architecture: Synthetic specs independent from the PR e2e suite; results normalised to a JSON schema; status page is static HTML generated from history.
- Edge Cases/Constraints: GitHub scheduled workflow delays/disable-after-inactivity, staging downtime windows (maintenance flag), secrets absent on forks, flaky third-party dependencies.
- Testing: Unit tests for issue-dedupe/history-aggregation scripts; a demo run (or
workflow_dispatch with a forced failure) shown in the PR.
Definition of "Done": Baseline DoD, plus published status page URL and runbook in docs/.
Resources: .github/workflows/uptime-check.yml, GitHub Actions scheduled workflow docs.
Complexity: High (200 points)
Baseline Definition of Done (applies to every Wave issue)
Each issue's own "Definition of Done" is in addition to this baseline:
- Code, tests and documentation are included in one PR that references the issue.
npm run lint, npm run typecheck, npm test, npm run check:i18n and npm run check:editorconfig pass locally and in CI. (If a command is broken by pre-existing repo debris, see the Repository Health & Build Integrity issues — do not silence the check; note the blocker in the PR.)
- No new
any, no // @ts-ignore / eslint-disable without a justification comment, and no new console.* (use secureLogger from src/lib/secureLogging.ts).
- All new user-facing strings go through the i18n catalog (
src/lib/i18n/messages/en.ts and es.ts).
- New interactive UI is keyboard-operable, has visible focus, correct ARIA semantics, and works in both the dark and light palettes defined in
src/app/globals.css.
- UI changes include before/after screenshots (desktop and ~400px mobile). Behaviour changes include a short screen recording or test output.
- Relevant docs under
docs/ (and README.md if routes/scripts/env vars change) are updated.
- The PR is reviewed and approved by a maintainer listed in
.github/CODEOWNERS.
Category: DevOps, CI/CD & Observability
Description:
uptime-check.ymlpings URLs every 30 minutes. Replace/extend it with a scheduled Playwright synthetic that exercises a real user path on staging (load, feed renders, WebSocket connects, quote fetch) and alerts on failure with deduplicated GitHub issues and an optional chat webhook.Problem Statement & Context: A 200 OK from the homepage says nothing about whether the swap card can fetch quotes or the live feed connects — the failures users actually experience.
Scope & Acceptance Criteria:
/api/healthOK, the feed shows items (or a valid empty state), WebSocket reachesopenwithin N seconds (observed via the UI "Live" indicator), quote request for a canned route returns within budget.incidentwith logs/trace links, and optionally notify a Slack/Discord webhook secret; auto-close/comment on recovery; avoids alert storms via dedupe.gh-pagesbranch to generate a simple static status/latency page (uptime %, last 30 days) published via GitHub Pages.Implementation Guidelines (Suggested Execution):
.github/workflows/uptime-check.yml, newe2e/synthetic/*.spec.ts, newscripts/status-page/*,README.md(status badge).workflow_dispatchwith a forced failure) shown in the PR.Definition of "Done": Baseline DoD, plus published status page URL and runbook in
docs/.Resources:
.github/workflows/uptime-check.yml, GitHub Actions scheduled workflow docs.Complexity: High (200 points)
Baseline Definition of Done (applies to every Wave issue)
Each issue's own "Definition of Done" is in addition to this baseline:
npm run lint,npm run typecheck,npm test,npm run check:i18nandnpm run check:editorconfigpass locally and in CI. (If a command is broken by pre-existing repo debris, see the Repository Health & Build Integrity issues — do not silence the check; note the blocker in the PR.)any, no// @ts-ignore/eslint-disablewithout a justification comment, and no newconsole.*(usesecureLoggerfromsrc/lib/secureLogging.ts).src/lib/i18n/messages/en.tsandes.ts).src/app/globals.css.docs/(andREADME.mdif routes/scripts/env vars change) are updated..github/CODEOWNERS.