Skip to content

Merge dev into main: bring the prod deploy pipeline and 40 commits of dev work to main - #460

Merged
saqibmanan merged 40 commits into
mainfrom
dev
Sep 11, 2026
Merged

saqibmanan merged 40 commits into
mainfrom
dev

Conversation

@saqibmanan

Copy link
Copy Markdown
Contributor

Summary

Notes

  • Merge-tree dry run confirmed clean — zero conflicts.
  • `production` GH environment has a working required-reviewer gate (public repo). Merging this does NOT deploy by itself — the deploy job will sit pending your approval, same as DataSpaceFrontend/DataSpaceBackend.

Saqib and others added 30 commits March 25, 2026 19:52
…d health check

- Fix remote script mixing && and ; — a failed npm install could
  previously fall through to deleting the working build and restarting
  pm2 on an incomplete install. Now fully &&-chained under set -e.
- Back up current .next/public before swapping in the new build; roll
  back and restart pm2 automatically if the post-deploy health check
  fails, instead of leaving a silently broken deploy running.
- Add HEALTH_CHECK_URL-based curl health check after pm2 restart.
- Add concurrency group so overlapping pushes to CD queue instead of
  racing on the same EC2 target directory.
- Add workflow_dispatch for manual redeploys.
- Add timeout-minutes: 15 to bound hung scp/ssh steps.
- Switch npm install to npm ci (both in Actions and on the EC2 side)
  for reproducible, lockfile-exact installs.
Lockfile was out of sync with package.json (missing/mismatched
@emnapi/* entries pulled in via sharp and rolldown's wasm32-wasi
binding) — npm install was silently tolerating this on every prior
install, masking it. npm ci correctly rejects it, which is what
surfaced this. Regenerated via npm install; verified npm ci now
succeeds and all platform-specific optional deps (including
linux-x64, needed by CI and the EC2 server) are still present.
husky is a devDependency, but "prepare" ran it unconditionally —
any --omit=dev install (like the EC2-side deploy install) failed
outright with "husky: not found" once npm ci correctly enforced
lockfile sync and the deploy script's new set -e stopped tolerating
the failure. husky's own documented fix: fall back to a no-op when
the binary isn't present, since prepare only matters for local dev
git hooks anyway.
npm run build already runs "generate:ci && next build" internally,
and postinstall runs generate:ci again on top of that — the explicit
Generate step made it a pointless third invocation, always a no-op
fallback since no gql-tagged documents exist in the codebase anymore.
Build step already inherits NEXT_PUBLIC_USI_GRAPHQL_ENDPOINT from the
job-level env, and next build forces production mode itself, so
nothing else needs to move.
- Updated `opub-ui` dependency version in `package.json` and `package-lock.json` to `^0.5.3`.
- Added new constants for evaluation status and audit type labels to improve code readability.
- Refactored evaluation components to utilize new constants for audit types and evaluation modes.
- Cleaned up unused code and improved the structure of evaluation-related components for better maintainability.
- Added Filter and pagination in evaluation filters
- Imported `AuditType` in `EvaluationConfiguration.tsx` and `NewEvaluationContent.tsx` for improved type safety.
- Updated the logic for parsing and setting the default audit type to utilize the `AUDIT_TYPE` constant.
- Cleaned up unused imports to enhance code clarity and maintainability.
npm audit fix (no --force) resolved all 4 critical CVEs without any
major version bumps:
- next-auth 4.24.13 -> 4.24.15 (fixes a critical Auth.js email
  normalizer homoglyph bypass, plus a high-severity malformed Bearer
  token crash) - patch-level, already within the ^4.24.13 range
- basic-ftp, handlebars, shell-quote - transitive, non-breaking bumps

package.json is unchanged; only lockfile-level resolution moved.
Verified locally: clean npm ci, npm run build, dev server homepage
and /api/auth/signin both 200, and the Keycloak OAuth redirect still
constructs correctly (same client_id/redirect_uri/PKCE shape as
before).

75 -> 57 vulnerabilities remain, all now requiring --force (major
version bumps: Next.js 14->16, @sentry/nextjs 8->10, sharp 0.33->0.35,
ESLint/graphql-codegen tooling) - deliberately not forced here, that's
real framework-migration work with regression risk, not an audit fix.
Retires the dedicated CD branch as the deploy trigger. Once this
change is merged into dev (via the upcoming CD -> dev PR), every
subsequent merge to dev will auto-trigger this same hardened deploy
workflow, matching how DataSpaceFrontend's deploy-Dataspace.yml
already works. workflow_dispatch stays available for manual runs.

Note: pushing this commit to CD itself won't trigger a deploy, since
the new trigger no longer matches the CD branch.
Harden ParakhAI dev frontend deploy workflow + fix critical npm audit findings
Temporarily reverts the trigger back to push:CD for testing this
flow end-to-end (will switch back to push:dev once verified).

Adds three jobs after deploy:
- smoke-tests: calls ParakhAI_test's reusable run-smoke.yml
- rollback-on-smoke-failure: if smoke tests fail, restores the
  pre-deploy .next.bak/public.bak build on the EC2 box, restarts
  pm2, health-checks the rollback itself, then fails the job loudly
  so it's visible as a red check even though the mitigation worked
- finalize-deploy: if smoke tests pass, clears the backup

This required removing the "rm -rf .next.bak public.bak" that used
to run at the end of deploy's own health check - the backup now has
to survive until the smoke-tests job (a separate job/runner) has had
a chance to gate on it. deploy's own internal rollback-on-health-
check-failure is unchanged; the new jobs only fire when deploy itself
already succeeded but the subsequent smoke suite didn't.
…sts']

GitHub Actions expression syntax doesn't support hyphens in dot-property
access - needs.smoke-tests.result parses as needs.smoke MINUS
tests.result, not a property lookup. Both rollback-on-smoke-failure and
finalize-deploy silently evaluated to false and skipped on the first
real run despite smoke-tests genuinely failing. Bracket notation with a
string key is required for hyphenated job ids.
needs['smoke-tests'].result stayed skipped on two consecutive real
runs despite smoke-tests genuinely failing both times - the bracket-
notation fix (previous commit) didn't actually resolve it, so the
hyphen-in-dot-notation theory wasn't the real cause. Switching to
GitHub Actions' built-in failure()/success() functions instead, which
is the same pattern already proven in DataSpaceFrontend's draft
automation branch. rollback-on-smoke-failure also guards on
needs.deploy.result == 'success' (no hyphen, unambiguous) so it only
fires when deploy itself succeeded but smoke-tests specifically
failed - deploy's own internal rollback already handles the case
where deploy itself fails, no need to double-fire.
GitHub Actions only exposes environment-scoped secrets to a job that
explicitly declares that environment - the smoke-tests job (a
workflow_call to ParakhAI_test) never declared environment: development,
so every secrets.X reference in its secrets: block resolved to empty
regardless of what existed in that environment. Every dispatched run so
far showed "14 passed, 43 skipped" - identical before and after adding
TEST_EMAIL_2/TEST_PASSWORD_2/SANDBOX_ORG_SLUG - because TEST_EMAIL_1 was
never actually reaching the reusable workflow either; the earlier "it's
working" read was based on the deploy/finalize pipeline mechanics
firing correctly, not on actually checking the test pass count.
Saqib and others added 10 commits August 4, 2026 10:48
GitHub Actions rejects the combination outright (run 30880188994
failed at 0s with zero jobs created - schema-level rejection, not a
runtime error). A workflow_call job's secrets: block can only resolve
${{ secrets.X }} against repo-level secrets, never environment-scoped
ones - there's no way to grant a reusable-workflow-calling job access
to a specific environment's secrets. The real fix is moving
TEST_EMAIL_1/2, TEST_PASSWORD_1/2, and SANDBOX_ORG_SLUG to repo-level
secrets instead of the development environment.
- Replaced string literals for evaluation methods with the `EVALUATION_MODE` constants for improved type safety and consistency.
- Removed the `EvaluationMethod` type definition as it is no longer needed.
- Updated state initialization and reset logic to utilize the new constants, enhancing code clarity.
Testing on push:CD is done - full pipeline verified end-to-end
(deploy, real smoke tests with correct secret scope, automatic
rollback on failure, automatic finalize on success: 55 passed,
2 skipped). This is the version that should land on dev via PR #448 /
issue #447, so every future merge to dev auto-deploys and runs the
smoke gate for real.
Add smoke-test-gated deploy: real tests, auto-rollback, auto-finalize
chore: update dependencies and refactor evaluation components
- Updated CSS styles for the evaluations list page to improve layout consistency.
- Replaced `flex-direction: row-reverse` with `justify-content: flex-start` for better alignment of header labels and icons.
- Applied `max-width` and `width` properties to ensure elements do not stretch across the cell, enhancing visual clarity.
style: enhance evaluations page layout and alignment
- Removed redundant styles for evaluation mode columns to streamline the CSS.
- Updated header label and icon alignment to enhance visual consistency across the evaluations table.
- Adjusted flex properties to ensure proper grouping of elements within table headers.
style: refine evaluations page CSS for improved layout and alignment
@saqibmanan
saqibmanan merged commit b6c8cbd into main Sep 11, 2026
4 checks passed
saqibmanan added a commit that referenced this pull request Sep 11, 2026
PR #460 (dev -> main) reported zero conflicts and merged clean, but the
actual result was wrong: main kept the stale, pre-rename
deploy-parakh-dev.yml (old dev-only content, no prod logic) instead of
adopting deploy-parakh.yml.

Root cause: PR #459 renamed the file AND changed most of its content in
the same commit. Git's rename detection (~50% similarity threshold)
didn't recognize the two as the same file, so the 3-way merge treated it
as an unrelated delete (of the old path, unchanged on main / gone on dev)
+ add (of the new path, dev-only) -- and for reasons not fully traced,
the add didn't land, leaving only the stale file behind. Confirmed main
had zero trace of deploy-parakh.yml anywhere in its tree.

Fix: explicit rm + checkout from dev's real, already-proven content --
verified byte-identical to what's currently running successfully on dev.

This branch was successfully deployed

1 active deployment
development — 15d6157a Deployed Aug 25, 2026 by Abhishekfm via Finalize Deploy #34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants