From 0376b3d81c54d6e563e00105decede261b1001a2 Mon Sep 17 00:00:00 2001 From: Lily Shen <115414357+lilyshen0722@users.noreply.github.com> Date: Sun, 27 Sep 2026 04:21:05 -0700 Subject: [PATCH] ci(deploy): still probe chat when the rollout gate fails MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The agent-visible bad deploy right now is readiness refusing a pod that is in fact serving: Sentry wrapping the express layer handle made the mount probe answer false, so /ready refused a pod whose /api/pg/messages was mounted, the new pod never became Ready, and `kubectl rollout status --timeout=8m` failed the job. The step that carries the diagnosis does not run in that case. The rollout step exit 1s at line 227 inside itself, so the job stops there and skips `Verify the new backend serves chat` at line 238, whose error text is the one line that distinguishes a mounted route (401) from an unmounted one (404) — "readiness is lying about a working pod" vs "the pod is broken". The operator sees only "backend did not become ready within 8m". Give the rollout step an id and run the probe when that step failed. Not always()/!cancelled(): if an earlier step failed (image push, helm, GKE credentials) the cluster is unreachable or unchanged and the probe would report "Chat unavailable on the new backend" for a reason unrelated to chat. Cannot turn a green job red: the added runs are on a job that already failed, and the success and early-failure paths are unchanged. --- .github/workflows/deploy-dev.yml | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/.github/workflows/deploy-dev.yml b/.github/workflows/deploy-dev.yml index 7a13cf38f..ff7a59b39 100644 --- a/.github/workflows/deploy-dev.yml +++ b/.github/workflows/deploy-dev.yml @@ -204,6 +204,7 @@ jobs: # it. A deployment scaled to zero passes instantly (0 of 0 updated), # which is correct: parked is not stuck. - name: Verify rollout of the deployed workloads + id: rollout run: | set -uo pipefail FAILED="" @@ -236,6 +237,19 @@ jobs: # (TASK-168), ask the newest backend pod from the inside. Unauthenticated, # a mounted /api/pg/messages answers 401 and an unmounted one 404. - name: Verify the new backend serves chat + # Deliberately not `always()` / `!cancelled()`: if an earlier step failed + # (image push, helm upgrade, GKE credentials) the cluster is unreachable + # or unchanged, and this step would then report "Chat unavailable on the + # new backend" for a reason that has nothing to do with chat. The case + # that needs it is the one below — the rollout gate failing on its own. + # + # 2026-09-27: a stuck rollout (readiness refusing a pod that was in fact + # serving 401s) exited 1 inside the rollout step, so this step was + # skipped and the operator got "backend did not become ready within 8m" + # with no 401-vs-404 discriminator attached. That discriminator is what + # separates "readiness is lying about a working pod" from "the pod is + # broken", and it was absent exactly when it was needed. + if: ${{ success() || steps.rollout.outcome == 'failure' }} run: | set -uo pipefail POD=$(kubectl get pod -n "$NAMESPACE" -l app=backend \