Skip to content

fix(ci): stop the ci-health scheduled sweep timing out on a cold cargo build - #4273

Closed
rysweet wants to merge 1 commit into
mainfrom
engineer/steward-ci-github-actions-health-across-all-gov-e06d9e64-1784307875-ec0046
Closed

rysweet wants to merge 1 commit into
mainfrom
engineer/steward-ci-github-actions-health-across-all-gov-e06d9e64-1784307875-ec0046

Conversation

@rysweet

@rysweet rysweet commented Jul 17, 2026

Copy link
Copy Markdown
Owner

Summary

Steward CI/GitHub-Actions health sweep across the governed fleet. The daily
ci-health scheduled sweep (.github/workflows/ci-health.yml) was the one
red default-branch workflow on rysweet/Simard: run
29557011590 was
cancelled at its 20-minute timeout-minutes cap. This PR diagnoses the root
cause and fixes it.

Root cause (diagnosed from the failed run log)

The sweep is a Rust binary, so the job must build simard before it can audit
the fleet. Its Swatinem/rust-cache step set no shared-key, so rust-cache
derived the cache key from the job name (sweep) → v0-rust-sweep-*. Because
the job is save-if: false, nothing ever writes to that namespace, so every
scheduled run was a guaranteed cache miss that rebuilt the entire workspace
from scratch. The failed run's log proves this:

  • ... Restoring cache ... → No cache found.
  • At the 20m kill, the orphaned processes were cargo + rustc + rustc —
    i.e. it was still compiling; the sweep binary never even started.

Fix

Reuse the same warm simard-ci-v2 cargo cache that verify.yml already
populates on main (its main-branch run is the sole writer), staying
read-only (save-if: false) so the scheduled sweep still never poisons the
shared PR/verify cache. This mirrors verify.yml's existing read-only
fallback-build job and turns the cold rebuild into a few-minute incremental one,
well under the 20m budget. No Rust code changes — behaviour is unchanged.

Diff: 2 files, +25/−2 (.github/workflows/ci-health.yml, docs/reference/ci-health-sweep.md).

Dedup

Merge-ready evidence

Criterion 1 — qa-team / gadugi scenarios. The change is CI-workflow config;
it introduces no new CLI/behavioural surface, so the existing governed-fleet
scenario tests/gadugi/ci-health-sweep.yaml remains the coverage and still
passes end-to-end against committed offline fixtures:

$ bash tests/gadugi/ci-health-sweep.sh
ci-health-sweep: PASS   (exit 0)

(gadugi-test validate/run are unavailable in this environment; the scenario's
committed driver script — what the gadugi step executes — was run directly.)

Criterion 2 — Docs. Changed surface is the ci-health.yml workflow (internal
CI infra). docs/reference/ci-health-sweep.md updated: new Build cache bullet
under "Scheduled recurring sweep" documenting the shared-key reuse and why the
prior per-job key caused the timeout.

Criterion 3 — Quality audit. SEEK→VALIDATE→FIX over the diff across 3 cycles,
final cycle clean (see PR comment). No critical/high; no medium
correctness/security findings. Change is config-only, read-only cache reuse,
matching an established in-repo pattern.

Criterion 4 — CI green. actionlint clean on the changed workflow; the
authoritative verify gates run on this PR (see checks). Local evidence:

$ actionlint .github/workflows/ci-health.yml   # clean
$ cargo test --locked --lib ci_health::         # 66 passed; 0 failed; 0 ignored
$ bash tests/gadugi/ci-health-sweep.sh          # ci-health-sweep: PASS

Criterion 6 — Focused diff. Only the ci-health cache config + its reference
doc. No unrelated edits.

Refs #4172

The daily `ci-health` sweep (`.github/workflows/ci-health.yml`) was being
cancelled at its 20-minute `timeout-minutes` cap before the sweep binary
ever ran. Root cause: its `Swatinem/rust-cache` step set no `shared-key`,
so rust-cache derived a per-job key from the job name (`sweep`) — a cache
namespace nothing ever writes to (the job is `save-if: false`). Every
scheduled run was therefore a guaranteed cache MISS that rebuilt the whole
workspace from scratch; the failed run's log shows `No cache found` and
`cargo`/`rustc` still compiling when the job was killed at 20m.

Fix: reuse the same warm `simard-ci-v2` cargo cache that `verify.yml`
populates on `main`, staying read-only (`save-if: false`) so the sweep
still never poisons the shared PR/verify cache. This mirrors verify.yml's
read-only fallback-build job and turns the cold rebuild into a few-minute
incremental one, well under the 20m budget.

No Rust code changes; behaviour is unchanged. Docs updated in
docs/reference/ci-health-sweep.md (Scheduled recurring sweep § Build cache).

Evidence:
- actionlint .github/workflows/ci-health.yml -> clean
- cargo test --locked --lib ci_health:: -> 66 passed; 0 failed
- tests/gadugi/ci-health-sweep.sh -> ci-health-sweep: PASS (exit 0)

Refs #4172

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@rysweet

rysweet commented Jul 17, 2026

Copy link
Copy Markdown
Owner Author

Quality audit — 3× SEEK→VALIDATE→FIX (final cycle clean)

Scope: the full diff (.github/workflows/ci-health.yml, docs/reference/ci-health-sweep.md).

Cycle 1 — correctness of the cache reuse.

  • SEEK: Does shared-key: simard-ci-v2 match the cache verify.yml populates, and does that cache hold the dev-profile artifacts the sweep (cargo run --bin simard) needs?
  • VALIDATE: verify.yml uses shared-key: simard-ci-v2 on every rust-cache step; its main-branch job runs cargo test (dev/test profile) and cargo build --bin simard --locked (dev), so the dev-profile graph is cached. The failed sweep run compiled cleanly (no lbug/link error) and died only on the 20m timeout, so the sweep's default build needs no extra provisioning that the cache restore would omit.
  • FIX: none required.

Cycle 2 — security / cache-poisoning / staleness.

  • SEEK: Could reusing verify's cache poison the shared cache or restore stale artifacts?
  • VALIDATE: save-if: false is retained → the sweep is a read-only consumer and can never write simard-ci-v2 (verify's main run stays the sole writer). rust-cache keys include the Cargo.lock hash + rustc version, and the sweep builds the same lockfile at the same default-branch HEAD verify built; cargo fingerprinting rebuilds anything stale. This is byte-for-byte the same pattern as verify.yml's already-trusted read-only fallback-build job.
  • FIX: none required.

Cycle 3 — accuracy, lint, focus.

  • SEEK: Are the doc/comment claims accurate, is the workflow valid, and is the diff focused?
  • VALIDATE: "sole writer of simard-ci-v2" is correct — verify.yml's writer job is save-if: ${{ github.ref == 'refs/heads/main' }}, all consumers save-if: false. actionlint .github/workflows/ci-health.yml is clean. git diff --stat shows only the 2 intended files (+25/−2); no unrelated edits.
  • FIX: none required. Final cycle clean — zero critical/high; zero medium correctness/security findings.

@rysweet

rysweet commented Jul 17, 2026

Copy link
Copy Markdown
Owner Author

Closing as a duplicate — superseded by #4251 (fix(ci-health): restore warm shared cache + prebuilt lbug so the scheduled sweep builds within its 20-min timeout), commit dd01c9fd, already merged to main.

This PR and #4251 diagnosed the identical root cause and applied the identical primary fix — adding shared-key: simard-ci-v2 to the ci-health sweep's Swatinem/rust-cache step so the daily sweep reuses verify.yml's warm cargo cache instead of cold-building the whole workspace and overrunning its 20-minute timeout-minutes. #4251 additionally provisions the prebuilt lbug static archive for the warm-cache build, so it is the more complete change.

Per the CI-health steward's dedup rule (one PR per distinct failure), I'm not landing a competing edit to the same lines. The fix is on main; I've dispatched a fresh ci-health run on main HEAD (which includes dd01c9fd) to confirm it now completes green within budget.

Refs #4251, #4172

@rysweet rysweet closed this Jul 17, 2026
@rysweet
rysweet deleted the engineer/steward-ci-github-actions-health-across-all-gov-e06d9e64-1784307875-ec0046 branch July 17, 2026 17:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant