Repository navigation
Conversation
The daily `ci-health` sweep (`.github/workflows/ci-health.yml`) was being cancelled at its 20-minute `timeout-minutes` cap before the sweep binary ever ran. Root cause: its `Swatinem/rust-cache` step set no `shared-key`, so rust-cache derived a per-job key from the job name (`sweep`) — a cache namespace nothing ever writes to (the job is `save-if: false`). Every scheduled run was therefore a guaranteed cache MISS that rebuilt the whole workspace from scratch; the failed run's log shows `No cache found` and `cargo`/`rustc` still compiling when the job was killed at 20m. Fix: reuse the same warm `simard-ci-v2` cargo cache that `verify.yml` populates on `main`, staying read-only (`save-if: false`) so the sweep still never poisons the shared PR/verify cache. This mirrors verify.yml's read-only fallback-build job and turns the cold rebuild into a few-minute incremental one, well under the 20m budget. No Rust code changes; behaviour is unchanged. Docs updated in docs/reference/ci-health-sweep.md (Scheduled recurring sweep § Build cache). Evidence: - actionlint .github/workflows/ci-health.yml -> clean - cargo test --locked --lib ci_health:: -> 66 passed; 0 failed - tests/gadugi/ci-health-sweep.sh -> ci-health-sweep: PASS (exit 0) Refs #4172 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Quality audit — 3× SEEK→VALIDATE→FIX (final cycle clean)Scope: the full diff ( Cycle 1 — correctness of the cache reuse.
Cycle 2 — security / cache-poisoning / staleness.
Cycle 3 — accuracy, lint, focus.
|
|
Closing as a duplicate — superseded by #4251 ( This PR and #4251 diagnosed the identical root cause and applied the identical primary fix — adding Per the CI-health steward's dedup rule (one PR per distinct failure), I'm not landing a competing edit to the same lines. The fix is on |
Summary
Steward CI/GitHub-Actions health sweep across the governed fleet. The daily
ci-healthscheduled sweep (.github/workflows/ci-health.yml) was the onered default-branch workflow on
rysweet/Simard: run29557011590 was
cancelled at its 20-minute
timeout-minutescap. This PR diagnoses the rootcause and fixes it.
Root cause (diagnosed from the failed run log)
The sweep is a Rust binary, so the job must build
simardbefore it can auditthe fleet. Its
Swatinem/rust-cachestep set noshared-key, so rust-cachederived the cache key from the job name (
sweep) →v0-rust-sweep-*. Becausethe job is
save-if: false, nothing ever writes to that namespace, so everyscheduled run was a guaranteed cache miss that rebuilt the entire workspace
from scratch. The failed run's log proves this:
... Restoring cache ...→No cache found.cargo+rustc+rustc—i.e. it was still compiling; the sweep binary never even started.
Fix
Reuse the same warm
simard-ci-v2cargo cache thatverify.ymlalreadypopulates on
main(its main-branch run is the sole writer), stayingread-only (
save-if: false) so the scheduled sweep still never poisons theshared PR/verify cache. This mirrors verify.yml's existing read-only
fallback-build job and turns the cold rebuild into a few-minute incremental one,
well under the 20m budget. No Rust code changes — behaviour is unchanged.
Diff: 2 files, +25/−2 (
.github/workflows/ci-health.yml,docs/reference/ci-health-sweep.md).Dedup
Dependabot PRs chore(ci): Bump actions/checkout from 4.3.1 to 7.0.0 #4066 / chore(ci): Bump actions/setup-node from 4.4.0 to 7.0.0 #4065 — deliberately not touched here to avoid a
duplicate fix (one PR per distinct failure).
Merge-ready evidence
Criterion 1 — qa-team / gadugi scenarios. The change is CI-workflow config;
it introduces no new CLI/behavioural surface, so the existing governed-fleet
scenario
tests/gadugi/ci-health-sweep.yamlremains the coverage and stillpasses end-to-end against committed offline fixtures:
(
gadugi-test validate/runare unavailable in this environment; the scenario'scommitted driver script — what the gadugi step executes — was run directly.)
Criterion 2 — Docs. Changed surface is the
ci-health.ymlworkflow (internalCI infra).
docs/reference/ci-health-sweep.mdupdated: new Build cache bulletunder "Scheduled recurring sweep" documenting the shared-key reuse and why the
prior per-job key caused the timeout.
Criterion 3 — Quality audit. SEEK→VALIDATE→FIX over the diff across 3 cycles,
final cycle clean (see PR comment). No critical/high; no medium
correctness/security findings. Change is config-only, read-only cache reuse,
matching an established in-repo pattern.
Criterion 4 — CI green. actionlint clean on the changed workflow; the
authoritative
verifygates run on this PR (see checks). Local evidence:Criterion 6 — Focused diff. Only the ci-health cache config + its reference
doc. No unrelated edits.
Refs #4172