diff --git a/.github/actions/ci-status/action.yml b/.github/actions/ci-status/action.yml index 4407c2d..cec75dd 100644 --- a/.github/actions/ci-status/action.yml +++ b/.github/actions/ci-status/action.yml @@ -12,7 +12,8 @@ inputs: Whitespace-separated job results, one per aggregated lane — the caller builds this from `needs..result`. Empty fails: a gate with nothing to aggregate is a miswired `needs` list, not a pass. Ignored when - `contract-only` is true, where every lane is `skipped` by construction. + `contract-only` is true, where every lane is `skipped` by construction, + and when `record-pending` is true. required: true treat-skipped-as: description: >- @@ -78,24 +79,55 @@ inputs: Reaching the ceiling fails the run; there is no pass-on-timeout path. Two cases run to the ceiling and fail closed: two or more contract-only runs on a SHA with no status and no full run in flight to write one, and a - writer re-run slower than the ceiling. Every red names both remedies: - re-run the full workflow, or, once the status is `success`, re-run this - run. `0` disables the wait and goes straight to a single status read. - The calling job needs `actions: read`: under an explicit `permissions:` - block the default is none, and a 403 warns naming the scope and degrades - to the status read alone. - SIZE THIS FROM THE REPOSITORY'S OWN MEASURED FULL RUN, not from the - calling job's budget. The ceiling must cover the queue wait plus the p95 - wall of the full `ci` run; set the calling job's `timeout-minutes` to at - least that figure plus two minutes, then set this input to - `timeout-minutes * 60 - 60`. Staying 60 seconds below `timeout-minutes` - is a constraint on the result, not the sizing rule: it keeps the job - timeout from preempting the fail-closed error, but a ceiling derived from - an underived budget cannot outlast the run it waits for. The `240` - default suits only a repository whose full run finishes in well under two - minutes; see the composite's README section for the measured per- - repository values in use. + writer re-run slower than the ceiling. The ceiling counts elapsed + wall-clock time from the start of the wait, API calls included, so it + overruns by at most one poll's calls. Every red names the remedy that + fits the status it read (see `rerun-contract-only-siblings`). + `0` reads the status once, never waits, and makes no Actions call. With + `record-pending` and `rerun-contract-only-siblings` it is the + recommended setting, with a calling-job `timeout-minutes` of 3: no run + then waits for another. + A wait above `0` needs `actions: read` on the calling job: under an + explicit `permissions:` block the default is none, and a 403 warns + naming the scope and degrades to the status read alone. + A WAIT ABOVE `0` IS SIZED FROM THE REPOSITORY'S OWN MEASURED FULL RUN, + not from the calling job's budget. The ceiling must cover the queue wait + plus the p95 wall of the full `ci` run; set the calling job's + `timeout-minutes` to at least that figure plus two minutes, then set + this input to `timeout-minutes * 60 - 60`. Staying 60 seconds below + `timeout-minutes` keeps the job timeout from preempting the fail-closed + error. The `240` default suits only a repository whose full run finishes + in well under two minutes; see the composite's README section for the + values in use. default: '240' + rerun-contract-only-siblings: + description: >- + `true` to have a full run that records `success` re-run every failed + contract-only run of this same workflow on the same head SHA, so each + red `ci-status` check run is replaced by one that reads the success. + A contract-only run is recognized by its jobs: in its latest attempt + every job but one was skipped, and that one failed. A full run always + runs more than its gate and this run is excluded by id, so neither is + ever re-run. A re-run attempt keeps its original event and is + contract-only again, so it never re-runs anything itself. A + carry-forward red then says the full run will re-run it instead of + telling the reader to. The calling job needs `actions: write`; a + refusal warns and never changes this run's verdict. Off by default. + default: 'false' + record-pending: + description: >- + `true` to mark `status-context` `pending` on `sha` and stop, with no + aggregation and no carry-forward. Call it from the first job of the + full run: a contract-only run that reads the status while this run is + in flight then finds `pending` and fails, instead of carrying an older + run's `success` forward (a draft run's, for example), and this run's + own gate records the real verdict over it. It writes only on a + same-repository pull request event that is not contract-only, and + notes the skip and passes otherwise: a fork's token is read-only, and + neither a push nor a contract-only run is the full run a contract-only + run waits for. `results` is ignored. The calling job needs + `statuses: write`. Off by default. + default: 'false' token: description: >- Token used to write and read the commit status. The calling job needs @@ -121,6 +153,8 @@ runs: SAME_REPO: ${{ inputs.same-repo }} STATUS_CONTEXT: ${{ inputs.status-context }} CARRY_FORWARD_WAIT_SECONDS: ${{ inputs.carry-forward-wait-seconds }} + RERUN_CONTRACT_ONLY_SIBLINGS: ${{ inputs.rerun-contract-only-siblings }} + RECORD_PENDING: ${{ inputs.record-pending }} GH_TOKEN: ${{ inputs.token }} REPOSITORY: ${{ inputs.repository }} SHA: ${{ inputs.sha }} diff --git a/.github/actions/ci-status/run.sh b/.github/actions/ci-status/run.sh index 70a774d..5eee19c 100755 --- a/.github/actions/ci-status/run.sh +++ b/.github/actions/ci-status/run.sh @@ -8,6 +8,18 @@ # run cannot say which event produced it — a chain of contract-only runs could # otherwise self-certify. # +# With `rerun-contract-only-siblings`, a full run that records `success` then +# re-runs every failed contract-only run of this workflow on the SHA, so its red +# check run is replaced by one that reads the success. See +# `rerun_failed_contract_only_siblings` for how a contract-only run is +# recognized and why the re-run cannot loop. +# +# Pending mode (`record-pending` true): the first job of a full run marks +# `status-context` `pending` and stops. A contract-only run that reads the +# status while that full run is in flight then fails instead of carrying an +# older run's `success` forward, and the full run's gate overwrites the marker +# with its verdict. +# # Carry-forward mode (`contract-only` true): the lanes were gated off by # construction, so aggregation is skipped and the recorded commit status for # `status-context` on the same SHA decides. It is read from the status LIST @@ -58,16 +70,17 @@ # passing on an unsettled status; there is no pass-on-timeout path. Two cases # run to the ceiling: two or more contract-only runs on a SHA with no status and # no full run in flight to write one (they wait on each other and then both fail -# closed), and a writer re-run that outlasts the ceiling. +# closed), and a writer re-run that outlasts the ceiling. The ceiling is elapsed +# wall-clock time since the wait began, API calls included, not the sum of the +# sleeps, so a large ceiling stays under the job budget it was sized against. # # Reading a settled `success` before the wait set empties is a deliberate trade: # it releases the mutual wait above, and it can carry an older `success` forward # while a re-run of the same SHA is in flight to overwrite it. The defenses # against a forged status (context, creator login, Bot type, newest id) hold. # -# Every red names both remedies. The run cannot see a later verdict, so the -# message says to re-run the full workflow, and to re-run this run instead once -# `status-context` on the SHA is `success`. +# Every red names the remedy that fits the state it read; see +# `fail_carry_forward`. # # `same-repo` false is a fork pull request. Its token is read-only on # `pull_request` whatever `permissions:` requests, so it cannot record lane @@ -81,6 +94,8 @@ set -euo pipefail RESULTS="${RESULTS:-}" CONTRACT_ONLY="${CONTRACT_ONLY:-}" SAME_REPO="${SAME_REPO:-}" +RERUN_CONTRACT_ONLY_SIBLINGS="${RERUN_CONTRACT_ONLY_SIBLINGS:-}" +RECORD_PENDING="${RECORD_PENDING:-}" STATUS_CONTEXT="${STATUS_CONTEXT:-ci-lanes}" REPOSITORY="${REPOSITORY:-}" SHA="${SHA:-}" @@ -162,6 +177,14 @@ fi if ! same_repo="$(read_boolean same-repo "$SAME_REPO" true)"; then exit 1 fi +# shellcheck disable=SC2310 # read_boolean reports a bad value through its status; the caller exits on it. +if ! rerun_contract_only_siblings="$(read_boolean rerun-contract-only-siblings "$RERUN_CONTRACT_ONLY_SIBLINGS" false)"; then + exit 1 +fi +# shellcheck disable=SC2310 # read_boolean reports a bad value through its status; the caller exits on it. +if ! record_pending="$(read_boolean record-pending "$RECORD_PENDING" false)"; then + exit 1 +fi # GitHub's documented escaping for workflow-command data, so a value echoed back # in an annotation cannot close it and inject a second command. `%` first, or the @@ -195,6 +218,67 @@ require_pattern sha "$SHA" '^[0-9a-f]{40}$' 'a full 40-character lowercase commi # a no-op, either of which silently changes the branch the caller asked for. require_pattern carry-forward-wait-seconds "$CARRY_FORWARD_WAIT_SECONDS" '^[0-9]+$' 'a non-negative integer number of seconds' +# POST one `status-context` entry on `sha` naming this run, retrying after 1s, +# 2s and 4s. Every write is load-bearing, not best-effort: the carry-forward +# branch reads nothing else, so a silently missing status turns every later +# contract-only run red, or lets one carry an older verdict, with no way to tell +# a refused write from a real result. Returns 1 after printing the error. +write_status() { + local state="$1" description="$2" attempt delay + jq -n \ + --arg state "$state" \ + --arg context "$STATUS_CONTEXT" \ + --arg description "$description" \ + --arg target_url "${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${GITHUB_RUN_ID:-0}" \ + '{state: $state, context: $context, description: $description, target_url: $target_url}' \ + >"$scratch/status-payload.json" + for attempt in 1 2 3 4; do + # shellcheck disable=SC2310 # gh_api handles its own errexit; the retry loop classifies the status. + if gh_api POST "repos/${REPOSITORY}/statuses/${SHA}" --input "$scratch/status-payload.json"; then + return 0 + fi + if [[ "$attempt" -lt 4 ]]; then + delay=$((STATUS_RETRY_BASE_DELAY * (1 << (attempt - 1)))) + echo "::warning::could not record ${STATUS_CONTEXT} on ${SHA} (HTTP ${GH_HTTP_STATUS:-unknown}); retrying in ${delay}s" + sleep "$delay" + fi + done + cat "$gh_stderr" >&2 + echo "::error::could not record ${STATUS_CONTEXT} on ${SHA} (${GH_HTTP_STATUS:-unknown}); the calling job needs statuses: write" + return 1 +} + +# --------------------------------------------------------------------------- +# Pending mode. Only the full run a contract-only run would wait for marks its +# verdict pending: a contract-only run marking it would fail itself and every +# sibling with nothing coming to replace the marker, a fork's token cannot +# write, and a push has no contract-only sibling to protect. Each of those +# passes with a notice, so the step is safe in a job that runs on every event. +# --------------------------------------------------------------------------- +if [[ "$record_pending" == true ]]; then + if [[ "$contract_only" == true ]]; then + echo "::notice::contract-only event: ${STATUS_CONTEXT} is not marked pending; only a full run marks its own verdict pending." + exit 0 + fi + if [[ "$same_repo" != true ]]; then + echo "::notice::fork pull request: ${STATUS_CONTEXT} is not marked pending; a fork never carries a verdict forward." + exit 0 + fi + case "${GITHUB_EVENT_NAME:-}" in + pull_request | pull_request_target) ;; + *) + echo "::notice::${GITHUB_EVENT_NAME:-unknown} event: ${STATUS_CONTEXT} is not marked pending; only a pull request has contract-only runs." + exit 0 + ;; + esac + # shellcheck disable=SC2310 # write_status prints its own error; the caller exits on it. + if ! write_status pending 'Full run in flight; lanes not yet aggregated.'; then + exit 1 + fi + echo "Recorded ${STATUS_CONTEXT}=pending on ${SHA}; a contract-only run fails until this run's gate records its verdict." + exit 0 +fi + # --------------------------------------------------------------------------- # Carry-forward mode. Branched on first, before `same-repo`: the caller's # predicate makes `contract-only` false for every fork event, so the @@ -255,22 +339,44 @@ read_carried_state() { carried_state_read=true } -# Both Actions reads need `actions: read`, which an explicit `permissions:` block -# does not grant by default. Name the scope rather than printing a bare 403, and -# never treat the refusal as permission to pass: the caller falls through to the -# status read it would have done without the wait. -warn_actions_read_failed() { - local endpoint="$1" +# The Actions calls need a scope an explicit `permissions:` block does not grant +# by default: `actions: read` to wait, `actions: write` to re-run siblings. Name +# the scope rather than printing a bare 403, and say what the run does instead. +# Neither caller treats a refusal as permission to pass. +warn_actions_failed() { + local endpoint="$1" scope="$2" consequence="$3" if [[ "$GH_HTTP_STATUS" == 403 ]]; then - echo "::warning::${endpoint} returned HTTP 403; the ci-status job needs 'actions: read' to wait for a sibling run. Reading the recorded status without waiting." + echo "::warning::${endpoint} returned HTTP 403; the ci-status job needs '${scope}'. ${consequence}" else - echo "::warning::could not read ${endpoint} (HTTP ${GH_HTTP_STATUS:-unknown}); reading the recorded status without waiting." + echo "::warning::could not read ${endpoint} (HTTP ${GH_HTTP_STATUS:-unknown}). ${consequence}" fi cat "$gh_stderr" >&2 } -# Seconds actually slept, and the sibling run ids waited on, for the ceiling -# message. Set by wait_for_sibling_runs. +# Sets `workflow_id` to this run's workflow, the key both the wait and the +# sibling re-run list runs under, and validates `GITHUB_RUN_ID`, which both +# exclude from that list. Returns 1 after a warning ending in `consequence`. +workflow_id="" +resolve_workflow_id() { + local scope="$1" consequence="$2" run_id="${GITHUB_RUN_ID:-}" + if [[ ! "$run_id" =~ ^[0-9]+$ ]]; then + echo "::warning::GITHUB_RUN_ID is not a run id; cannot exclude this run from the runs on ${SHA}. ${consequence}" + return 1 + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${run_id}"; then + warn_actions_failed "repos/${REPOSITORY}/actions/runs/${run_id}" "$scope" "$consequence" + return 1 + fi + workflow_id="$(jq -r '.workflow_id // ""' <"$gh_stdout")" + if [[ ! "$workflow_id" =~ ^[0-9]+$ ]]; then + echo "::warning::run ${run_id} reported no workflow_id. ${consequence}" + return 1 + fi +} + +# Wall-clock seconds since the wait began, and the sibling run ids waited on, +# for the ceiling message. Set by wait_for_sibling_runs. carry_forward_waited=0 carry_forward_wait_note="" @@ -295,19 +401,14 @@ set_wait_note() { # that fails closed on its own failure; there is no outcome that passes without # one completed read. wait_for_sibling_runs() { - local run_id="${GITHUB_RUN_ID:-}" workflow_id ids all_ids="" sleep_for remaining - if [[ ! "$run_id" =~ ^[0-9]+$ ]]; then - echo "::warning::GITHUB_RUN_ID is not a run id; cannot exclude this run from its own wait set. Reading the recorded status without waiting." - return 0 - fi - # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. - if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${run_id}"; then - warn_actions_read_failed "repos/${REPOSITORY}/actions/runs/${run_id}" - return 0 - fi - workflow_id="$(jq -r '.workflow_id // ""' <"$gh_stdout")" - if [[ ! "$workflow_id" =~ ^[0-9]+$ ]]; then - echo "::warning::run ${run_id} reported no workflow_id; reading the recorded status without waiting." + local ids all_ids="" sleep_for remaining wait_started now + local no_wait='Reading the recorded status without waiting.' + # `date`, not bash's `SECONDS`: an external clock is one the harness can + # advance from its `sleep` and `gh` shims, so the ceiling arithmetic stays + # real in tests without a test-only knob in this script. + wait_started="$(date +%s)" + # shellcheck disable=SC2310 # resolve_workflow_id warns itself; the caller degrades to one read. + if ! resolve_workflow_id 'actions: read' "$no_wait"; then return 0 fi while :; do @@ -316,7 +417,7 @@ wait_for_sibling_runs() { # is already far past the burst this wait exists for. # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. if ! gh_api GET "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs?head_sha=${SHA}&per_page=100"; then - warn_actions_read_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" + warn_actions_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" 'actions: read' "$no_wait" # Discard any earlier poll's read so the caller takes the single fresh # read the warning above promises. Without this, a listing that fails on # the second or later poll would decide on the state read before the @@ -332,12 +433,14 @@ wait_for_sibling_runs() { # covers a run object with no `id`, keeping a malformed entry out rather # than waiting on it forever. Sorted so the ids read in run order and the # log line is stable from one poll to the next. - ids="$(jq -r --argjson incomplete "$INCOMPLETE_RUN_STATUSES" --argjson self "$run_id" \ + ids="$(jq -r --argjson incomplete "$INCOMPLETE_RUN_STATUSES" --argjson self "$GITHUB_RUN_ID" \ '[ .workflow_runs[]? | select(.status as $s | $incomplete | index($s)) | select((.id // $self) != $self) | .id ] | sort | join(" ")' \ <"$gh_stdout")" # Kept for the settled-failure check below, which runs after the status # read has overwritten gh_stdout. cp -- "$gh_stdout" "$scratch/runs.json" + now="$(date +%s)" + carry_forward_waited=$((now - wait_started)) # The status read comes AFTER the listing above, never before it: a sibling # writes the status and flips to `completed` moments later, and the other # order lets that completion land between the two calls. @@ -400,17 +503,41 @@ wait_for_sibling_runs() { fi echo "Waiting ${sleep_for}s for in-flight run(s) ${ids} on ${SHA} to finish (waited ${carry_forward_waited}s of ${CARRY_FORWARD_WAIT_SECONDS}s)." sleep "$sleep_for" - carry_forward_waited=$((carry_forward_waited + sleep_for)) done set_wait_note "$all_ids" } -# Every carry-forward red. The run cannot see a verdict recorded after it, so -# the message names both remedies: a failed or missing full run needs the full -# workflow again, and once the status is `success` re-running this run is -# enough, where a new commit would re-run every lane. +# Every carry-forward red. The remedy follows the state read. `pending` names +# the full run in flight, whose own gate check run supersedes this one, so +# nothing needs re-running unless that run was cancelled. A `failure` or `error` +# names the run that recorded it. An absent status, or a read that failed, +# cannot tell whether a full run is coming, so it gives both cases. The closing +# sentence says who replaces this red once the lanes pass: the full run itself +# when this workflow re-runs contract-only siblings (the same step runs in both +# modes, so this run's input is the full run's), otherwise whoever re-runs this +# run, which is cheaper than a new commit that re-runs every lane. fail_carry_forward() { - echo "::error::no successful ${STATUS_CONTEXT} status on ${SHA}; re-run the full workflow${carry_forward_wait_note}. Once ${STATUS_CONTEXT} on ${SHA} is success, re-run this run instead." + local writer="" remedy closing + if [[ -n "$carried_writer_run_id" ]]; then + writer="${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${carried_writer_run_id}" + fi + case "$carried_state" in + pending) + remedy="full run ${writer:-on this SHA} is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled" + ;; + failure | error) + remedy="the lanes verdict is ${carried_state}${writer:+ (${writer})}; fix the failing lane, or re-run that run's failed jobs if the failure was transient" + ;; + *) + remedy="if a full run on this SHA is in flight, its ci-status check supersedes this one; otherwise re-run the full workflow" + ;; + esac + if [[ "$rerun_contract_only_siblings" == true ]]; then + closing="A full run that records ${STATUS_CONTEXT}=success on ${SHA} re-runs this run; re-run it yourself only if it stays red after that." + else + closing="Once ${STATUS_CONTEXT} on ${SHA} is success, re-run this run instead." + fi + echo "::error::no successful ${STATUS_CONTEXT} status on ${SHA}; ${remedy}${carry_forward_wait_note}. ${closing}" exit 1 } @@ -491,10 +618,7 @@ if [[ "$lanes_state" == success ]]; then fi # --------------------------------------------------------------------------- -# Record the verdict as a commit status. Load-bearing on a same-repository run, -# not best-effort: the carry-forward branch reads nothing else, so a silently -# missing status turns every later contract-only run red with no way to tell a -# refused write from a genuinely failing lane. +# Record the verdict as a commit status (see `write_status`). # # A fork pull request is the one exception: its token cannot write a status at # all, and nothing will ever read one for it, so the run reports the lanes @@ -508,32 +632,8 @@ if [[ "$same_repo" != true ]]; then exit 0 fi -target_url="${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${GITHUB_RUN_ID:-0}" -jq -n \ - --arg state "$lanes_state" \ - --arg context "$STATUS_CONTEXT" \ - --arg description "$lanes_description" \ - --arg target_url "$target_url" \ - '{state: $state, context: $context, description: $description, target_url: $target_url}' \ - >"$scratch/status-payload.json" - -status_written=false -for attempt in 1 2 3 4; do - # shellcheck disable=SC2310 # gh_api handles its own errexit; the retry loop classifies the status. - if gh_api POST "repos/${REPOSITORY}/statuses/${SHA}" --input "$scratch/status-payload.json"; then - status_written=true - break - fi - if [[ "$attempt" -lt 4 ]]; then - delay=$((STATUS_RETRY_BASE_DELAY * (1 << (attempt - 1)))) - echo "::warning::could not record ${STATUS_CONTEXT} on ${SHA} (HTTP ${GH_HTTP_STATUS:-unknown}); retrying in ${delay}s" - sleep "$delay" - fi -done - -if [[ "$status_written" != true ]]; then - cat "$gh_stderr" >&2 - echo "::error::could not record ${STATUS_CONTEXT} on ${SHA} (${GH_HTTP_STATUS:-unknown}); the ci-status job needs statuses: write" +# shellcheck disable=SC2310 # write_status prints its own error; the caller exits on it. +if ! write_status "$lanes_state" "$lanes_description"; then exit 1 fi @@ -542,3 +642,67 @@ echo "Recorded ${STATUS_CONTEXT}=${lanes_state} on ${SHA}." if [[ "$lanes_state" == failure ]]; then exit 1 fi + +# Re-run every failed contract-only run of this workflow on this SHA, so its red +# check run is replaced by one that reads the success just recorded. Runs only +# after that write, so a re-run cannot read anything older. +# +# A contract-only run is recognized by its jobs, not its event, which the run +# object does not carry: in its latest attempt every job but one was skipped, +# and that one failed. A full run always runs more than its gate, and this run +# is excluded by id, so neither is ever re-run. A re-run attempt keeps its +# original event, so it is contract-only again and never reaches this code: no +# loop. A contract-only run that is red for a contract reason (title, +# `do-not-merge`) just goes red again, at the cost of one short job. +# +# The known gap: a contract-only run still in flight when this lists the runs +# read the status before this run wrote it, finishes red after, and is not +# re-run. Its message therefore ends by telling the reader to re-run it if it +# stays red after the full run's success. +# +# Nothing here changes this run's verdict: the success is already recorded, so +# a refusal warns and moves on. +rerun_failed_contract_only_siblings() { + local skip='Not re-running failed contract-only runs.' candidates id contract_shaped rerun_count=0 + # shellcheck disable=SC2310 # resolve_workflow_id warns itself; a refusal only skips the re-run. + if ! resolve_workflow_id 'actions: write' "$skip"; then + return 0 + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs?head_sha=${SHA}&per_page=100"; then + warn_actions_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" 'actions: write' "$skip" + return 0 + fi + candidates="$(jq -r --argjson self "$GITHUB_RUN_ID" \ + '[ .workflow_runs[]? | select(.status == "completed" and .conclusion == "failure" and (.id // $self) != $self) | .id ] | sort | join(" ")' \ + <"$gh_stdout")" + for id in $candidates; do + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${id}/jobs?filter=latest&per_page=100"; then + warn_actions_failed "repos/${REPOSITORY}/actions/runs/${id}/jobs" 'actions: write' "Not re-running run ${id}." + continue + fi + contract_shaped="$(jq -r '[ .jobs[]? | .conclusion ] | ((map(select(. != "skipped")) == ["failure"]) and (length > 1))' <"$gh_stdout")" + if [[ "$contract_shaped" != true ]]; then + continue + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api POST "repos/${REPOSITORY}/actions/runs/${id}/rerun-failed-jobs"; then + if [[ "$GH_HTTP_STATUS" == 403 ]]; then + echo "::warning::re-running run ${id} returned HTTP 403; rerun-contract-only-siblings needs 'actions: write' on the ci-status job." + else + echo "::warning::could not re-run run ${id} (HTTP ${GH_HTTP_STATUS:-unknown})." + fi + cat "$gh_stderr" >&2 + continue + fi + echo "Re-running failed contract-only run ${id} on ${SHA}." + rerun_count=$((rerun_count + 1)) + done + echo "Re-ran ${rerun_count} failed contract-only run(s) on ${SHA}." +} + +if [[ "$rerun_contract_only_siblings" == true ]]; then + # shellcheck disable=SC2310 # errexit is off inside on purpose: the success is recorded, and a body jq cannot parse must not turn it red. + rerun_failed_contract_only_siblings || true +fi diff --git a/.github/actions/ci-status/run.test.sh b/.github/actions/ci-status/run.test.sh index 39e7fa5..e729d0b 100755 --- a/.github/actions/ci-status/run.test.sh +++ b/.github/actions/ci-status/run.test.sh @@ -22,21 +22,44 @@ mkdir -p "$fixtures" "$calls" "$shim_directory" sha=deadbeefdeadbeefdeadbeefdeadbeefdeadbeef repository=melodic-software/ci-workflows -# A no-op `sleep` first on PATH. The carry-forward wait accounts for elapsed -# time arithmetically against its own 15-second interval, so shimming the sleep -# keeps the ceiling arithmetic real while the suite runs instantly. Unlike a -# test-only poll-interval knob, it adds nothing to the shipped runner. +# The carry-forward red, one remedy per state read; see `fail_carry_forward`. +fail_prefix="::error::no successful ci-lanes status on ${sha}; " +absent_remedy='if a full run on this SHA is in flight, its ci-status check supersedes this one; otherwise re-run the full workflow' +failure_remedy="the lanes verdict is failure; fix the failing lane, or re-run that run's failed jobs if the failure was transient" +writer_failure_remedy="the lanes verdict is failure (https://github.com/${repository}/actions/runs/4000); fix the failing lane, or re-run that run's failed jobs if the failure was transient" +pending_remedy='full run on this SHA is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled' +manual_closing="Once ci-lanes on ${sha} is success, re-run this run instead." +automatic_closing="A full run that records ci-lanes=success on ${sha} re-runs this run; re-run it yourself only if it stays red after that." + +# A `sleep` and a `date` first on PATH that share a fake clock: sleeping +# advances it, and `date +%s` reads it. The wait's ceiling is elapsed wall-clock +# time, so this keeps the ceiling arithmetic real while the suite runs +# instantly, and adds nothing test-only to the shipped runner. Every sleep is +# logged so a case can assert that none happened. +clock_file="$temporary_directory/clock" +sleep_log="$temporary_directory/sleep.log" cat >"$shim_directory/sleep" <<'SLEEP_SHIM' #!/usr/bin/env bash -exit 0 +printf '%s\n' "$1" >>"$SLEEP_LOG" +printf '%s\n' "$(($(cat "$FAKE_CLOCK") + $1))" >"$FAKE_CLOCK" SLEEP_SHIM chmod +x "$shim_directory/sleep" +cat >"$shim_directory/date" <<'DATE_SHIM' +#!/usr/bin/env bash +if [[ "$*" != '+%s' ]]; then + echo "date shim: only +%s is supported, got: $*" >&2 + exit 1 +fi +cat "$FAKE_CLOCK" +DATE_SHIM +chmod +x "$shim_directory/date" # A `gh` shim first on PATH: it serves fixture JSON keyed by method plus API # path, records every call so the harness can assert on writes that did and did # not happen, can fail a keyed call a fixed number of times before succeeding -# (the retry case), and can serve a DIFFERENT body on the Nth call to the same -# key (`..json`), which is how a status appearing mid-wait is fixtured. +# (the retry case), can serve a DIFFERENT body on the Nth call to the same key +# (`..json`), which is how a status appearing mid-wait is fixtured, and +# can make every call to a key take time on the fake clock (`.advance`). cat >"$shim_directory/gh" <<'SHIM' #!/usr/bin/env bash set -uo pipefail @@ -93,6 +116,10 @@ if [[ -f "$count_file" ]]; then fi printf '%s\n' "$call_number" >"$count_file" +if [[ -f "$GH_FIXTURES/${key}.advance" ]]; then + printf '%s\n' "$(($(cat "$FAKE_CLOCK") + $(cat "$GH_FIXTURES/${key}.advance")))" >"$FAKE_CLOCK" +fi + fail_times="$GH_FIXTURES/${key}.fail-times" if [[ -f "$fail_times" ]]; then remaining="$(cat "$fail_times")" @@ -140,6 +167,8 @@ run_case() { shift $(($# > 5 ? 5 : $#)) local actual_status : >"$gh_log" + : >"$sleep_log" + printf '0\n' >"$clock_file" rm -rf -- "$calls" mkdir -p "$calls" set +e @@ -148,6 +177,8 @@ run_case() { GH_LOG="$gh_log" \ GH_FIXTURES="$fixtures" \ GH_CALLS="$calls" \ + FAKE_CLOCK="$clock_file" \ + SLEEP_LOG="$sleep_log" \ GH_TOKEN=fixture-token \ RESULTS="$results" \ TREAT_SKIPPED_AS="$treat_skipped_as" \ @@ -155,11 +186,14 @@ run_case() { SAME_REPO="$same_repo" \ STATUS_CONTEXT=ci-lanes \ CARRY_FORWARD_WAIT_SECONDS=0 \ + RERUN_CONTRACT_ONLY_SIBLINGS=false \ + RECORD_PENDING=false \ REPOSITORY="$repository" \ SHA="$sha" \ STATUS_RETRY_BASE_DELAY=0 \ GITHUB_SERVER_URL=https://github.com \ GITHUB_RUN_ID=4242 \ + GITHUB_EVENT_NAME=pull_request \ "$@" \ bash "$script_directory/run.sh" >"$log_file" 2>&1 actual_status=$? @@ -223,6 +257,30 @@ expect_gh_call_before() { fi } +# expect_status_reads +# The shim counts calls per key, so this pins how many times the status list +# was read, independent of which fixture answered. +expect_status_reads() { + local expected="$1" actual=0 + local count_file="$calls/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.count" + if [[ -f "$count_file" ]]; then + actual="$(cat "$count_file")" + fi + if [[ "$actual" -ne "$expected" ]]; then + echo "FAIL: expected $expected status read(s), got $actual" + cat "$gh_log" + failures=$((failures + 1)) + fi +} + +expect_no_sleep() { + if [[ -s "$sleep_log" ]]; then + echo 'FAIL: expected no sleep, got:' + cat "$sleep_log" + failures=$((failures + 1)) + fi +} + # The shim logs `$*`, which never contains the literal `gh api` — it starts at # `api -X GET …` — so asserting on that string could never fire. Assert the log # is empty instead. @@ -352,7 +410,7 @@ echo 'case: full mode fails the run when the status write is refused after retri printf '%s\n' 99 >"$fixtures/POST_repos_melodic-software_ci-workflows_statuses_${sha}.fail-times" run_case 1 'success success' pass '' expect_log 'All lanes passed or were skipped.' -expect_log "::error::could not record ci-lanes on ${sha} (500); the ci-status job needs statuses: write" +expect_log "::error::could not record ci-lanes on ${sha} (500); the calling job needs statuses: write" rm -f -- "$fixtures/POST_repos_melodic-software_ci-workflows_statuses_${sha}.fail-times" # Without the retry loop the first 500 fails the run. @@ -460,19 +518,22 @@ expect_no_log 'All lanes passed' echo 'case: carry-forward fails on a recorded ci-lanes failure' status_list "[$(bot_status 100 failure)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}. ${manual_closing}" # Without the explicit `== success` test, a pending status would ride through. +# Without the pending branch of the message, the reader is told to re-run a +# full run that is still in flight. echo 'case: carry-forward fails on a pending ci-lanes status' status_list "[$(bot_status 100 pending)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${pending_remedy}. ${manual_closing}" +expect_no_log 're-run the full workflow' # Without the context filter, another context's success would satisfy the gate. echo 'case: carry-forward fails when no entry carries the ci-lanes context' status_list '[{"context":"other-lane","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the context filter, the FIRST entry (a failure under another context) # would decide. @@ -488,7 +549,7 @@ echo 'case: carry-forward fails closed when the status list cannot be read' rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.json" printf '%s\n' 'gh: Not Found (HTTP 404)' >"$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.err" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.err" # --- carry-forward: forged statuses ---------------------------------------- @@ -501,7 +562,7 @@ rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statu echo 'case: a forged success by a user account does not satisfy the carry-forward' status_list "[$(user_status 200 success),$(bot_status 100 failure)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" # Without newest-first selection an older user failure would shadow the bot's # real success. @@ -515,26 +576,35 @@ expect_log "Carried forward: ci-lanes is success on ${sha}" echo 'case: a bot failure newer than a bot success fails' status_list "[$(bot_status 200 failure),$(bot_status 100 success)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" + +# The marker a full run's first job writes under `record-pending`. Without +# newest-id selection the older success is carried forward while that full run +# is still in flight, which is the stale green the marker exists to stop. +echo 'case: a bot pending newer than a bot success fails' +status_list "[$(writer_status 200 pending 2026-09-05T12:00:20Z 4000),$(bot_status 100 success)]" +run_case 1 'skipped skipped' pass true +expect_log "${fail_prefix}full run https://github.com/${repository}/actions/runs/4000 is still in flight" +expect_no_log 'Carried forward' # Without the empty-list guard an absent status would read as an empty state. echo 'case: an empty status list fails the carry-forward' status_list '[]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the creator filter, a context that ONLY a user ever wrote satisfies # the gate — the plant-then-label attack in its simplest form. echo 'case: the ci-lanes context present only from a user account fails' status_list "[$(user_status 100 success)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the Bot type check, an account merely NAMED like the bot passes. echo 'case: a user account impersonating the bot login fails' status_list '[{"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"User"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without `max_by(.id)` the selection depends on the array order the API # happens to return; an oldest-first list would then hand back the stale @@ -542,7 +612,7 @@ expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full wo echo 'case: an oldest-first status list still selects the newest bot entry' status_list '[{"id":10,"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}},{"id":20,"context":"ci-lanes","state":"failure","creator":{"login":"github-actions[bot]","type":"Bot"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" echo 'case: an oldest-first status list still carries a newer bot success forward' status_list '[{"id":10,"context":"ci-lanes","state":"failure","creator":{"login":"github-actions[bot]","type":"Bot"}},{"id":20,"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}}]' @@ -586,7 +656,25 @@ current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${pending_remedy} (waited 30s of 30s on in-flight run(s): 4000)" + +# The ceiling is elapsed wall-clock time, API calls included. Each listing here +# takes 10 seconds, so after one 15-second sleep 35 seconds have passed and the +# 30-second ceiling is reached on the second poll. Counting only the sleeps +# (15, then 30) would poll a third time and report 30 of 30, which is how a +# large ceiling used to outlast the job budget sized for it. +echo 'case: the ceiling counts elapsed wall-clock time, not summed sleeps' +clear_status_fixtures +clear_run_fixtures +status_list "[$(bot_status 100 pending)]" +current_run +workflow_runs "[${earlier_full_run}]" +printf '%s\n' 10 >"$fixtures/${workflow_runs_key}.advance" +run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 +expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 10s of 30s)." +expect_log "${fail_prefix}${pending_remedy} (waited 35s of 30s on in-flight run(s): 4000)" +expect_no_log 'waited 30s of 30s' +rm -f -- "$fixtures/${workflow_runs_key}.advance" # What decides the gate is the verdict the sibling left behind, not whatever was # on the SHA when this run started. Serving nothing on the first read and a @@ -601,7 +689,7 @@ workflow_runs_on_call 1 "[${earlier_full_run}]" workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log 'Waiting 15s for in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${failure_remedy} (waited 15s of 60s on in-flight run(s): 4000)" # A settled failure ends the wait even with a sibling incomplete, when that # sibling is not the writer re-running (this status names no writer at all). @@ -614,7 +702,7 @@ current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" # Without the self-exclusion term this run waits on itself forever; without the # status filter it waits on a run that already finished. @@ -628,7 +716,7 @@ run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_gh_call 'actions/workflows/777/runs' expect_no_log 'Waiting ' expect_no_log 'in-flight run(s):' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the 403 branch the missing scope reads as "no earlier run", which is # the same outcome but unattributable. The run degrades to @@ -641,7 +729,7 @@ current_run printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/${workflow_runs_key}.err" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::repos/${repository}/actions/workflows/777/runs returned HTTP 403; the ci-status job needs 'actions: read'" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' # Without discarding the earlier poll's read, a listing that fails PART WAY @@ -672,16 +760,27 @@ expect_log "::warning::repos/${repository}/actions/runs/4242 returned HTTP 403; expect_no_gh_call 'actions/workflows/' # Without the `> 0` guard, a consumer that disabled the wait still pays two -# Actions API calls and still needs the `actions: read` scope. -echo 'case: carry-forward-wait-seconds 0 disables the wait and makes no Actions API call' +# Actions API calls, still needs the `actions: read` scope, and holds its runner +# while a sibling is in flight. Wait 0 is one status read and nothing else. +echo 'case: carry-forward-wait-seconds 0 reads the status once, never waits, and makes no Actions API call' clear_status_fixtures clear_run_fixtures status_list '[]' current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=0 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_gh_call 'actions/' +expect_status_reads 1 +expect_no_sleep + +echo 'case: carry-forward-wait-seconds 0 carries a recorded success after one read' +clear_status_fixtures +status_list "[$(bot_status 100 success)]" +run_case 0 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=0 +expect_log "Carried forward: ci-lanes is success on ${sha}" +expect_status_reads 1 +expect_no_sleep # Without the runner-side default, an action.yml regression that stopped passing # the input would silently disable the wait rather than fall back to 240. @@ -742,7 +841,7 @@ current_run status_list '[]' workflow_runs "[$(run_entry 3000 completed 2026-09-05T11:59:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' expect_no_log 'waited ' @@ -760,7 +859,7 @@ workflow_runs "[$(run_entry 4100 in_progress "$current_run_created_at"),$(run_en run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log "Waiting 15s for in-flight run(s) 4100 4300 on ${sha} to finish (waited 0s of 30s)." expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4100 4300' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4100 4300)" +expect_log "${fail_prefix}${absent_remedy} (waited 30s of 30s on in-flight run(s): 4100 4300)" # Without `error` in the settled set this waits the full 60s on a verdict that # can never change, then fails anyway. `error` is a terminal commit-status state @@ -773,7 +872,7 @@ status_list "[$(bot_status 100 error)]" workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}the lanes verdict is error; fix the failing lane" # --- carry-forward: a stale failure while its writer re-runs ---------------- # @@ -809,7 +908,7 @@ status_list "[${stale_failure}]" workflow_runs "[${writer_rerun}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4000). Once ci-lanes on ${sha} is success, re-run this run instead." +expect_log "${fail_prefix}${writer_failure_remedy} (waited 30s of 30s on in-flight run(s): 4000). ${manual_closing}" # Without the start-time term this waits on the writer's own first attempt, # which records the status a moment before it completes, and turns a prompt @@ -822,7 +921,7 @@ status_list "[${stale_failure}]" workflow_runs "[$(attempt_entry 4000 in_progress 2026-09-05T11:59:00Z 1 2026-09-05T11:59:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${writer_failure_remedy}" # Without narrowing the wait to the writer, a failed SHA brings back the mutual # wait: each contract-only run holds the failure open for the other until the @@ -836,7 +935,7 @@ status_list "[${stale_failure}]" workflow_runs "[$(attempt_entry 4100 in_progress "$current_run_created_at" 1 "$current_run_created_at"),$(attempt_entry 4300 in_progress 2026-09-05T12:00:00Z 2 2026-09-05T12:01:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow." +expect_log "${fail_prefix}${writer_failure_remedy}. ${manual_closing}" # Without the unsettled-state fail a re-run of the full run that ends without # recording anything would leave an absent status to pass. No status names no @@ -850,20 +949,36 @@ workflow_runs_on_call 1 "[${writer_rerun}]" workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 0s of 60s)." -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)." +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000)." expect_no_log 'Carried forward' -# Without the second remedy the message only ever says to re-run the full -# workflow, and a red that outlived its failure gets a new commit (every lane -# again) where re-running this run is enough. -echo 'case: the failure message names the re-run-this-run remedy for a later success' +# Without the closing remedy a red that outlived its failure gets a new commit +# (every lane again) where re-running this run is enough. Without naming the +# writer the reader has to hunt for the run whose failed jobs to re-run. +echo 'case: the failure message names the writer and the re-run-this-run remedy for a later success' clear_status_fixtures clear_run_fixtures current_run status_list "[${stale_failure}]" workflow_runs '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow. Once ci-lanes on ${sha} is success, re-run this run instead." +expect_log "${fail_prefix}${writer_failure_remedy}. ${manual_closing}" + +# The same step runs in both modes, so this run's rerun-contract-only-siblings +# is the full run's. Without branching the closing on it, the reader is told to +# re-run by hand a run the full run is about to re-run itself. +echo 'case: with rerun-contract-only-siblings the closing says the full run re-runs this run' +clear_status_fixtures +clear_run_fixtures +status_list "[$(writer_status 200 pending 2026-09-05T12:00:20Z 4000)]" +run_case 1 'skipped skipped' pass true true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "${fail_prefix}full run https://github.com/${repository}/actions/runs/4000 is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled. ${automatic_closing}" +expect_no_log "$manual_closing" +expect_no_log 're-run the full workflow' +# A contract-only run never re-runs anything, even with the input on: the +# re-run is a full-mode step, which is what keeps it from looping. +expect_no_gh_call 'rerun-failed-jobs' +expect_no_gh_call 'actions/' # The documented trade, pinned deliberately from the passing side. A recorded # success ends the wait even with a sibling in flight, so a re-run of this SHA @@ -893,7 +1008,7 @@ workflow_runs "[${earlier_full_run}]" printf '%s\n' 'gh: Not Found (HTTP 404)' >"$fixtures/${statuses_key}.err" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::could not read repos/${repository}/commits/${sha}/statuses (HTTP 404)" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_log 'gh: Not Found (HTTP 404)' rm -f -- "$fixtures/${statuses_key}.err" @@ -929,7 +1044,7 @@ workflow_runs_on_call 1 "[${earlier_full_run}]" workflow_runs_on_call 2 "[${earlier_full_run},$(run_entry 4700 in_progress "$current_run_created_at")]" printf '%s\n' 2 >"$fixtures/${statuses_key}.fail-on-call" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000 4700)" +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000 4700)" rm -f -- "$fixtures/${statuses_key}.fail-on-call" # A sibling can finish without ever recording a verdict: cancelled, or failed @@ -945,7 +1060,7 @@ workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 0s of 60s)." expect_log "No run on ${sha} is still in flight after 15s; using the ci-lanes status." -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000)" # A truncated body is what a cut-off response looks like: valid JSON up to the # point the connection dropped. `status_list` writes its argument verbatim, so @@ -961,11 +1076,182 @@ status_list "[$(bot_status 100 success)" workflow_runs '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::could not read repos/${repository}/commits/${sha}/statuses" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' clear_status_fixtures +# --- pending mode ---------------------------------------------------------- + +status_write_key="POST_repos_melodic-software_ci-workflows_statuses_${sha}" + +# Without pending mode, a contract-only run that reads the status while a full +# run is in flight carries an older run's success forward over lanes nobody has +# run yet. The marker names this run, which is how a contract-only red names +# the run to wait for. No aggregation: `results` is empty here and must not fail. +echo 'case: record-pending marks ci-lanes pending on a same-repository pull request and stops' +run_case 0 '' pass false true RECORD_PENDING=true +expect_status_payload '"state": "pending"' +expect_status_payload '"context": "ci-lanes"' +expect_status_payload "\"target_url\": \"https://github.com/${repository}/actions/runs/4242\"" +expect_log "Recorded ci-lanes=pending on ${sha}" +expect_no_log 'results is required' +expect_no_gh_call 'actions/' +expect_no_gh_call "commits/${sha}/statuses" + +echo 'case: record-pending writes on a pull_request_target event too' +run_case 0 '' pass false true RECORD_PENDING=true GITHUB_EVENT_NAME=pull_request_target +expect_status_payload '"state": "pending"' + +# Without the contract-only guard this run would mark the verdict pending, and +# it and every contract-only sibling would go red with no full run coming to +# overwrite the marker. +echo 'case: record-pending on a contract-only event writes nothing and passes' +run_case 0 '' pass true true RECORD_PENDING=true +expect_log '::notice::contract-only event: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the fork guard the read-only token's refused write fails the job. +echo 'case: record-pending on a fork pull request writes nothing and passes' +run_case 0 '' pass false false RECORD_PENDING=true +expect_log '::notice::fork pull request: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the event guard a push leaves a pending marker on a main commit that +# no full pull-request run will ever overwrite. +echo 'case: record-pending on a push writes nothing and passes' +run_case 0 '' pass '' true RECORD_PENDING=true GITHUB_EVENT_NAME=push +expect_log '::notice::push event: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the load-bearing write a refused marker passes silently, and the stale +# green it exists to stop comes back with nothing to say why. +echo 'case: record-pending fails the run when the write is refused after retries' +printf '%s\n' 99 >"$fixtures/${status_write_key}.fail-times" +run_case 1 '' pass false true RECORD_PENDING=true +expect_log "::error::could not record ci-lanes on ${sha} (500); the calling job needs statuses: write" +rm -f -- "$fixtures/${status_write_key}.fail-times" + +echo 'case: an unrecognized record-pending value is rejected' +run_case 1 '' pass false true RECORD_PENDING=yes +expect_log "::error::record-pending must be 'true' or 'false', got: yes" +expect_no_gh_calls_at_all + +# --- full mode: re-running failed contract-only siblings ------------------- + +# completed_run +completed_run() { + printf '{"id":%s,"status":"completed","conclusion":"%s"}' "$1" "$2" +} + +# jobs_for ... +# The latest attempt's jobs of run , one per conclusion. +jobs_for() { + local id="$1" jobs="" conclusion + shift + for conclusion in "$@"; do + jobs="${jobs}${jobs:+,}{\"name\":\"job\",\"conclusion\":\"${conclusion}\"}" + done + printf '{"jobs":[%s]}' "$jobs" >"$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_${id}_jobs.json" +} + +clear_rerun_fixtures() { + rm -f -- "$fixtures"/GET_repos_melodic-software_ci-workflows_actions_runs_*_jobs.* \ + "$fixtures"/POST_repos_melodic-software_ci-workflows_actions_runs_* +} + +# 5000 and 4242 (this run, completed only in this fixture) are contract-only +# shaped: one failed gate, every other job skipped. 5100 is a failed full run, +# 5500 a failed single-job run, and 5200, 5300 and 5400 did not fail. +clear_run_fixtures +clear_rerun_fixtures +current_run +workflow_runs "[$(completed_run 5000 failure),$(completed_run 5100 failure),$(completed_run 5200 success),$(run_entry 5300 in_progress 2026-09-05T12:00:40Z),$(completed_run 5400 cancelled),$(completed_run 5500 failure),$(completed_run 4242 failure)]" +jobs_for 5000 failure skipped skipped +jobs_for 5100 success failure failure +jobs_for 5500 failure +jobs_for 4242 failure skipped skipped + +# Without the re-run the contract-only red stays beside this run's green until +# someone re-runs it by hand. Without the shape test the failed full run 5100 is +# re-run too, every lane again for a verdict already recorded; without the +# single-job guard so is 5500; without the conclusion filter the green, +# in-flight and cancelled runs are fetched; without excluding this run's id it +# re-runs itself. +echo 'case: a full-mode success re-runs only the failed contract-only siblings' +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "Recorded ci-lanes=success on ${sha}." +expect_gh_call "POST repos/${repository}/actions/runs/5000/rerun-failed-jobs" +expect_log "Re-running failed contract-only run 5000 on ${sha}." +expect_log "Re-ran 1 failed contract-only run(s) on ${sha}." +for not_rerun in 5100 5500 4242; do + expect_no_gh_call "actions/runs/${not_rerun}/rerun-failed-jobs" +done +for not_fetched in 5200 5300 5400 4242; do + expect_no_gh_call "actions/runs/${not_fetched}/jobs" +done +# The re-run reads the status, so it must come after the write. +expect_gh_call_before "POST repos/${repository}/statuses/${sha}" 'rerun-failed-jobs' + +# Without the success gate a failing run re-runs siblings that can only read +# its failure and go red again. +echo 'case: a full-mode failure re-runs nothing' +run_case 1 'success failure' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_status_payload '"state": "failure"' +expect_no_gh_call 'actions/' + +# Without the opt-in every consumer would need `actions: write` on upgrade. +echo 'case: rerun-contract-only-siblings off by default makes no Actions call' +run_case 0 'success success' pass false true +expect_no_gh_call 'actions/' + +echo 'case: a fork run re-runs nothing' +run_case 0 'success success' pass false false RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_no_gh_call 'actions/' + +# The verdict is already recorded, so a refused re-run must not turn it red, +# and must not stop the next candidate. +echo 'case: a refused re-run warns naming actions: write and moves to the next sibling' +jobs_for 5600 failure skipped +workflow_runs "[$(completed_run 5000 failure),$(completed_run 5600 failure)]" +printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/POST_repos_melodic-software_ci-workflows_actions_runs_5000_rerun-failed-jobs.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::re-running run 5000 returned HTTP 403; rerun-contract-only-siblings needs 'actions: write' on the ci-status job." +expect_gh_call "POST repos/${repository}/actions/runs/5600/rerun-failed-jobs" +expect_log "Re-ran 1 failed contract-only run(s) on ${sha}." +rm -f -- "$fixtures/POST_repos_melodic-software_ci-workflows_actions_runs_5000_rerun-failed-jobs.err" + +echo 'case: a jobs read failure skips that sibling and still passes' +printf '%s\n' 'gh: Internal Server Error (HTTP 500)' >"$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_5000_jobs.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::could not read repos/${repository}/actions/runs/5000/jobs (HTTP 500). Not re-running run 5000." +expect_no_gh_call 'actions/runs/5000/rerun-failed-jobs' +expect_gh_call "POST repos/${repository}/actions/runs/5600/rerun-failed-jobs" +rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_5000_jobs.err" + +echo 'case: a 403 listing the runs warns naming actions: write and still passes' +printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/${workflow_runs_key}.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::repos/${repository}/actions/workflows/777/runs returned HTTP 403; the ci-status job needs 'actions: write'. Not re-running failed contract-only runs." +expect_no_gh_call 'rerun-failed-jobs' + +# Without errexit off inside the re-run, jq failing on a cut-off body exits the +# script nonzero after the success is recorded, turning a green gate red. +echo 'case: a malformed run list re-runs nothing and still passes' +rm -f -- "$fixtures/${workflow_runs_key}.err" +printf '%s' '{"workflow_runs":[' >"$fixtures/${workflow_runs_key}.json" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "Recorded ci-lanes=success on ${sha}." +expect_no_gh_call 'rerun-failed-jobs' + +echo 'case: an unrecognized rerun-contract-only-siblings value is rejected' +run_case 1 'success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=on +expect_log "::error::rerun-contract-only-siblings must be 'true' or 'false', got: on" +expect_no_gh_calls_at_all + +clear_run_fixtures +clear_rerun_fixtures + # --- input validation ------------------------------------------------------ # # Every one of these values is interpolated into a `gh api` path. @@ -991,19 +1277,29 @@ expect_status_payload '"context": "CI Lanes"' echo 'case: action.yml still wires every input this harness exercises' action_metadata="$script_directory/action.yml" -for input_name in results treat-skipped-as contract-only same-repo status-context carry-forward-wait-seconds token repository sha; do +for input_name in results treat-skipped-as contract-only same-repo status-context carry-forward-wait-seconds rerun-contract-only-siblings record-pending token repository sha; do if ! grep -qE "^ ${input_name}:" "$action_metadata"; then echo "FAIL: action.yml declares no '${input_name}' input" failures=$((failures + 1)) fi done -for environment_name in RESULTS TREAT_SKIPPED_AS CONTRACT_ONLY SAME_REPO STATUS_CONTEXT CARRY_FORWARD_WAIT_SECONDS GH_TOKEN REPOSITORY SHA; do +for environment_name in RESULTS TREAT_SKIPPED_AS CONTRACT_ONLY SAME_REPO STATUS_CONTEXT CARRY_FORWARD_WAIT_SECONDS RERUN_CONTRACT_ONLY_SIBLINGS RECORD_PENDING GH_TOKEN REPOSITORY SHA; do if ! grep -qF " ${environment_name}: " "$action_metadata"; then echo "FAIL: action.yml does not pass '${environment_name}' to run.sh" failures=$((failures + 1)) fi done +# Both A+ inputs are opt-in: a consumer that passes nothing keeps today's +# behavior and needs no new permission. +for input_name in rerun-contract-only-siblings record-pending; do + input_default="$(awk -v key=" ${input_name}:" '$0 == key { found = 1; next } found && /^ default:/ { print; exit }' "$action_metadata")" + if [[ "$input_default" != " default: 'false'" ]]; then + echo "FAIL: action.yml does not default ${input_name} to 'false' (got: ${input_default})" + failures=$((failures + 1)) + fi +done + # The published default is the contract consumers inherit when they pass # nothing, and the design fixes it at 240 seconds (below cursor-plugins' # five-minute ci-status job budget). A drift here is invisible to every case diff --git a/README.md b/README.md index 67aa81a..51f539e 100644 --- a/README.md +++ b/README.md @@ -276,14 +276,25 @@ consumer to audit it. land between the two calls and report both no status and nothing in flight, failing a SHA that does carry a verdict. - **The wait never turns a verdict green.** Reaching the ceiling prints - `::error::no successful status on ; re-run the full workflow`, - extended with how long it waited and on which run ids, and exits 1. There is - no pass-on-timeout path. Every carry-forward red ends with `Once - on is success, re-run this run instead.`: the run cannot see a verdict - recorded after it, and once the lanes pass, re-running the red run replaces - its check run where a new commit would re-run every lane. - - **The calling job needs `actions: read`.** Under an explicit `permissions:` + `::error::no successful status on ; `, extended + with how long it waited and on which run ids, and exits 1. There is no + pass-on-timeout path. The ceiling counts elapsed wall-clock time from the + start of the wait, API calls included, so it overruns by at most the few + seconds one poll's calls take. + + Every carry-forward red names the remedy for the state it read. `pending` + names the full run in flight, whose own `ci-status` check run supersedes the + red one, so nothing needs re-running unless that run was cancelled. A + `failure` or `error` names the run that recorded it: fix the lane, or re-run + that run's failed jobs. An absent status says a full run in flight will + supersede the red one and to re-run the full workflow only when none is. The + closing sentence says who replaces the red once the lanes pass: with + `rerun-contract-only-siblings` on, the full run re-runs it (the same step runs + in both modes, so the contract-only run knows); otherwise `Once on + is success, re-run this run instead.`, which replaces its check run + where a new commit would re-run every lane. + + **A wait above `0` needs `actions: read`.** Under an explicit `permissions:` block the default for every scope is none, so both Actions calls 403 without it. A 403 prints a `::warning::` naming the missing scope and then degrades to a single status read, which is the pre-6b contract: loud, and never a pass on @@ -300,8 +311,73 @@ consumer to audit it. The 15-second poll interval is deliberately not a caller input: the only knob a consumer should have to reason about is the ceiling. - **Size the ceiling from the repository's own measured full run, then derive - `timeout-minutes` from it.** `carry-forward-wait-seconds` must cover the wall + **Recommended: no run waits for another.** Three opt-in settings together + remove the wait, and with it the dependency on how long the full run takes: + + - `carry-forward-wait-seconds: '0'` and `timeout-minutes: 3` on the + `ci-status` job. A contract-only run reads the status once, makes no + Actions call, and finishes in seconds: green on a recorded `success`, red + otherwise. + - `record-pending: 'true'`, called as the first step of the full run's first + job (for example `changes`). It marks `status-context` `pending` on the + head SHA, so a contract-only run on that SHA while the lanes run reads + `pending` and goes red instead of carrying an older `success` forward (a + draft run's, for example). The full run's gate overwrites the marker with + its verdict. It writes only on a same-repository pull request event that is + not contract-only, and passes with a notice otherwise, so it is safe in a + job that runs on every event. `results` is ignored. + - `rerun-contract-only-siblings: 'true'` on the `ci-status` step. After the + full run records `success`, it re-runs, through + `POST /repos/{owner}/{repo}/actions/runs/{run_id}/rerun-failed-jobs`, every + failed run of the same workflow on the SHA whose latest attempt skipped + every job but one failed gate. Full runs and the run itself are never + re-run. A re-run attempt keeps its original event, so it is contract-only + again, reads the new `success`, replaces the red check run, and never + re-runs anything itself. A refusal warns and does not change the verdict. + + A red contract-only check does not hold a merge in the meantime: GitHub's + docs do not say how several same-name check runs on one SHA are judged, and + on claude-code-plugins the merge gate followed the newest one in every case + measured, so the full run's later `ci-status` supersedes the red. The + remaining gap: a contract-only run still in flight when the full run lists + its siblings read the status before it was written, finishes red, and is not + re-run; its message ends by saying to re-run it if it stays red. + + Permissions: the first job needs `statuses: write`; the `ci-status` job needs + `statuses: write` and `actions: write` (which covers the `actions: read` a + wait above `0` needs). + + ```yaml + jobs: + changes: + permissions: + contents: read + statuses: write + steps: + - name: Mark the lanes verdict pending + uses: melodic-software/ci-workflows/.github/actions/ci-status@ # + with: + record-pending: 'true' + # ... change detection + ci-status: + if: always() + timeout-minutes: 3 + permissions: + contents: read + statuses: write + actions: write + steps: + - name: Aggregate lane results + uses: melodic-software/ci-workflows/.github/actions/ci-status@ # + with: + carry-forward-wait-seconds: '0' + rerun-contract-only-siblings: 'true' + results: ${{ needs.changes.result }} ... + ``` + + **To keep a wait instead, size the ceiling from the repository's own measured + full run, then derive `timeout-minutes` from it.** + `carry-forward-wait-seconds` must cover the wall time the contract-only run may have to wait out: the queue wait plus the p95 wall of the full `ci` run on this repository. Set `timeout-minutes` to at least that figure plus two minutes, then set `carry-forward-wait-seconds` to @@ -321,10 +397,11 @@ consumer to audit it. The 60-second margin survives as a **constraint, not a sizing rule**: a ceiling at or above the job budget lets the job timeout preempt the fail-closed error, which reports as a cancelled job rather than the - actionable "re-run the full workflow" message. + actionable message. Because the ceiling counts wall-clock time, the margin + holds at any ceiling size. The `240` default therefore suits only a repository whose full run finishes in - well under two minutes. The values in use across the fleet today: + well under two minutes. The waiting values in use across the fleet today: | repository | `timeout-minutes` | `carry-forward-wait-seconds` | | --- | --- | --- |