From 89b13739d20c1f1747f2f589ba4e3b5d97429a01 Mon Sep 17 00:00:00 2001 From: Kyle Sexton <153232337+kyle-sexton@users.noreply.github.com> Date: Fri, 2 Oct 2026 14:55:31 -0400 Subject: [PATCH] feat(ci-status): add an opt-in mode in which no run waits for another A contract-only run could only wait for the full run's ci-lanes verdict, holding a runner for up to the full run's wall and going red at the ceiling whenever the full run was slower. Two opt-in inputs remove the wait: - record-pending, called from the full run's first job, marks ci-lanes pending on the head SHA and stops. A contract-only run that reads the status while the lanes run then fails instead of carrying an older success forward. It writes only on a same-repository pull request event that is not contract-only, and passes with a notice otherwise. - rerun-contract-only-siblings makes a full run that records success call rerun-failed-jobs on every failed run of the same workflow on the SHA whose latest attempt has one failed job and every other job skipped. Full runs and the run itself are never re-run, and a re-run is contract-only again, so it cannot loop. A refusal only warns. carry-forward-wait-seconds '0' already read the status once with no Actions call; tests now pin that, and the README recommends it with a 3-minute job budget. The wait ceiling now counts elapsed wall-clock time instead of summed sleeps, and each carry-forward red names the remedy for the state it read (pending, failure or error, absent) instead of always saying to re-run the full workflow. Every existing input keeps its default; both new inputs default to false. Co-Authored-By: Claude Opus 5.5 (1M context) --- .github/actions/ci-status/action.yml | 70 +++-- .github/actions/ci-status/run.sh | 294 +++++++++++++++----- .github/actions/ci-status/run.test.sh | 384 +++++++++++++++++++++++--- README.md | 101 ++++++- 4 files changed, 710 insertions(+), 139 deletions(-) diff --git a/.github/actions/ci-status/action.yml b/.github/actions/ci-status/action.yml index 4407c2d..cec75dd 100644 --- a/.github/actions/ci-status/action.yml +++ b/.github/actions/ci-status/action.yml @@ -12,7 +12,8 @@ inputs: Whitespace-separated job results, one per aggregated lane — the caller builds this from `needs..result`. Empty fails: a gate with nothing to aggregate is a miswired `needs` list, not a pass. Ignored when - `contract-only` is true, where every lane is `skipped` by construction. + `contract-only` is true, where every lane is `skipped` by construction, + and when `record-pending` is true. required: true treat-skipped-as: description: >- @@ -78,24 +79,55 @@ inputs: Reaching the ceiling fails the run; there is no pass-on-timeout path. Two cases run to the ceiling and fail closed: two or more contract-only runs on a SHA with no status and no full run in flight to write one, and a - writer re-run slower than the ceiling. Every red names both remedies: - re-run the full workflow, or, once the status is `success`, re-run this - run. `0` disables the wait and goes straight to a single status read. - The calling job needs `actions: read`: under an explicit `permissions:` - block the default is none, and a 403 warns naming the scope and degrades - to the status read alone. - SIZE THIS FROM THE REPOSITORY'S OWN MEASURED FULL RUN, not from the - calling job's budget. The ceiling must cover the queue wait plus the p95 - wall of the full `ci` run; set the calling job's `timeout-minutes` to at - least that figure plus two minutes, then set this input to - `timeout-minutes * 60 - 60`. Staying 60 seconds below `timeout-minutes` - is a constraint on the result, not the sizing rule: it keeps the job - timeout from preempting the fail-closed error, but a ceiling derived from - an underived budget cannot outlast the run it waits for. The `240` - default suits only a repository whose full run finishes in well under two - minutes; see the composite's README section for the measured per- - repository values in use. + writer re-run slower than the ceiling. The ceiling counts elapsed + wall-clock time from the start of the wait, API calls included, so it + overruns by at most one poll's calls. Every red names the remedy that + fits the status it read (see `rerun-contract-only-siblings`). + `0` reads the status once, never waits, and makes no Actions call. With + `record-pending` and `rerun-contract-only-siblings` it is the + recommended setting, with a calling-job `timeout-minutes` of 3: no run + then waits for another. + A wait above `0` needs `actions: read` on the calling job: under an + explicit `permissions:` block the default is none, and a 403 warns + naming the scope and degrades to the status read alone. + A WAIT ABOVE `0` IS SIZED FROM THE REPOSITORY'S OWN MEASURED FULL RUN, + not from the calling job's budget. The ceiling must cover the queue wait + plus the p95 wall of the full `ci` run; set the calling job's + `timeout-minutes` to at least that figure plus two minutes, then set + this input to `timeout-minutes * 60 - 60`. Staying 60 seconds below + `timeout-minutes` keeps the job timeout from preempting the fail-closed + error. The `240` default suits only a repository whose full run finishes + in well under two minutes; see the composite's README section for the + values in use. default: '240' + rerun-contract-only-siblings: + description: >- + `true` to have a full run that records `success` re-run every failed + contract-only run of this same workflow on the same head SHA, so each + red `ci-status` check run is replaced by one that reads the success. + A contract-only run is recognized by its jobs: in its latest attempt + every job but one was skipped, and that one failed. A full run always + runs more than its gate and this run is excluded by id, so neither is + ever re-run. A re-run attempt keeps its original event and is + contract-only again, so it never re-runs anything itself. A + carry-forward red then says the full run will re-run it instead of + telling the reader to. The calling job needs `actions: write`; a + refusal warns and never changes this run's verdict. Off by default. + default: 'false' + record-pending: + description: >- + `true` to mark `status-context` `pending` on `sha` and stop, with no + aggregation and no carry-forward. Call it from the first job of the + full run: a contract-only run that reads the status while this run is + in flight then finds `pending` and fails, instead of carrying an older + run's `success` forward (a draft run's, for example), and this run's + own gate records the real verdict over it. It writes only on a + same-repository pull request event that is not contract-only, and + notes the skip and passes otherwise: a fork's token is read-only, and + neither a push nor a contract-only run is the full run a contract-only + run waits for. `results` is ignored. The calling job needs + `statuses: write`. Off by default. + default: 'false' token: description: >- Token used to write and read the commit status. The calling job needs @@ -121,6 +153,8 @@ runs: SAME_REPO: ${{ inputs.same-repo }} STATUS_CONTEXT: ${{ inputs.status-context }} CARRY_FORWARD_WAIT_SECONDS: ${{ inputs.carry-forward-wait-seconds }} + RERUN_CONTRACT_ONLY_SIBLINGS: ${{ inputs.rerun-contract-only-siblings }} + RECORD_PENDING: ${{ inputs.record-pending }} GH_TOKEN: ${{ inputs.token }} REPOSITORY: ${{ inputs.repository }} SHA: ${{ inputs.sha }} diff --git a/.github/actions/ci-status/run.sh b/.github/actions/ci-status/run.sh index 70a774d..5eee19c 100755 --- a/.github/actions/ci-status/run.sh +++ b/.github/actions/ci-status/run.sh @@ -8,6 +8,18 @@ # run cannot say which event produced it — a chain of contract-only runs could # otherwise self-certify. # +# With `rerun-contract-only-siblings`, a full run that records `success` then +# re-runs every failed contract-only run of this workflow on the SHA, so its red +# check run is replaced by one that reads the success. See +# `rerun_failed_contract_only_siblings` for how a contract-only run is +# recognized and why the re-run cannot loop. +# +# Pending mode (`record-pending` true): the first job of a full run marks +# `status-context` `pending` and stops. A contract-only run that reads the +# status while that full run is in flight then fails instead of carrying an +# older run's `success` forward, and the full run's gate overwrites the marker +# with its verdict. +# # Carry-forward mode (`contract-only` true): the lanes were gated off by # construction, so aggregation is skipped and the recorded commit status for # `status-context` on the same SHA decides. It is read from the status LIST @@ -58,16 +70,17 @@ # passing on an unsettled status; there is no pass-on-timeout path. Two cases # run to the ceiling: two or more contract-only runs on a SHA with no status and # no full run in flight to write one (they wait on each other and then both fail -# closed), and a writer re-run that outlasts the ceiling. +# closed), and a writer re-run that outlasts the ceiling. The ceiling is elapsed +# wall-clock time since the wait began, API calls included, not the sum of the +# sleeps, so a large ceiling stays under the job budget it was sized against. # # Reading a settled `success` before the wait set empties is a deliberate trade: # it releases the mutual wait above, and it can carry an older `success` forward # while a re-run of the same SHA is in flight to overwrite it. The defenses # against a forged status (context, creator login, Bot type, newest id) hold. # -# Every red names both remedies. The run cannot see a later verdict, so the -# message says to re-run the full workflow, and to re-run this run instead once -# `status-context` on the SHA is `success`. +# Every red names the remedy that fits the state it read; see +# `fail_carry_forward`. # # `same-repo` false is a fork pull request. Its token is read-only on # `pull_request` whatever `permissions:` requests, so it cannot record lane @@ -81,6 +94,8 @@ set -euo pipefail RESULTS="${RESULTS:-}" CONTRACT_ONLY="${CONTRACT_ONLY:-}" SAME_REPO="${SAME_REPO:-}" +RERUN_CONTRACT_ONLY_SIBLINGS="${RERUN_CONTRACT_ONLY_SIBLINGS:-}" +RECORD_PENDING="${RECORD_PENDING:-}" STATUS_CONTEXT="${STATUS_CONTEXT:-ci-lanes}" REPOSITORY="${REPOSITORY:-}" SHA="${SHA:-}" @@ -162,6 +177,14 @@ fi if ! same_repo="$(read_boolean same-repo "$SAME_REPO" true)"; then exit 1 fi +# shellcheck disable=SC2310 # read_boolean reports a bad value through its status; the caller exits on it. +if ! rerun_contract_only_siblings="$(read_boolean rerun-contract-only-siblings "$RERUN_CONTRACT_ONLY_SIBLINGS" false)"; then + exit 1 +fi +# shellcheck disable=SC2310 # read_boolean reports a bad value through its status; the caller exits on it. +if ! record_pending="$(read_boolean record-pending "$RECORD_PENDING" false)"; then + exit 1 +fi # GitHub's documented escaping for workflow-command data, so a value echoed back # in an annotation cannot close it and inject a second command. `%` first, or the @@ -195,6 +218,67 @@ require_pattern sha "$SHA" '^[0-9a-f]{40}$' 'a full 40-character lowercase commi # a no-op, either of which silently changes the branch the caller asked for. require_pattern carry-forward-wait-seconds "$CARRY_FORWARD_WAIT_SECONDS" '^[0-9]+$' 'a non-negative integer number of seconds' +# POST one `status-context` entry on `sha` naming this run, retrying after 1s, +# 2s and 4s. Every write is load-bearing, not best-effort: the carry-forward +# branch reads nothing else, so a silently missing status turns every later +# contract-only run red, or lets one carry an older verdict, with no way to tell +# a refused write from a real result. Returns 1 after printing the error. +write_status() { + local state="$1" description="$2" attempt delay + jq -n \ + --arg state "$state" \ + --arg context "$STATUS_CONTEXT" \ + --arg description "$description" \ + --arg target_url "${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${GITHUB_RUN_ID:-0}" \ + '{state: $state, context: $context, description: $description, target_url: $target_url}' \ + >"$scratch/status-payload.json" + for attempt in 1 2 3 4; do + # shellcheck disable=SC2310 # gh_api handles its own errexit; the retry loop classifies the status. + if gh_api POST "repos/${REPOSITORY}/statuses/${SHA}" --input "$scratch/status-payload.json"; then + return 0 + fi + if [[ "$attempt" -lt 4 ]]; then + delay=$((STATUS_RETRY_BASE_DELAY * (1 << (attempt - 1)))) + echo "::warning::could not record ${STATUS_CONTEXT} on ${SHA} (HTTP ${GH_HTTP_STATUS:-unknown}); retrying in ${delay}s" + sleep "$delay" + fi + done + cat "$gh_stderr" >&2 + echo "::error::could not record ${STATUS_CONTEXT} on ${SHA} (${GH_HTTP_STATUS:-unknown}); the calling job needs statuses: write" + return 1 +} + +# --------------------------------------------------------------------------- +# Pending mode. Only the full run a contract-only run would wait for marks its +# verdict pending: a contract-only run marking it would fail itself and every +# sibling with nothing coming to replace the marker, a fork's token cannot +# write, and a push has no contract-only sibling to protect. Each of those +# passes with a notice, so the step is safe in a job that runs on every event. +# --------------------------------------------------------------------------- +if [[ "$record_pending" == true ]]; then + if [[ "$contract_only" == true ]]; then + echo "::notice::contract-only event: ${STATUS_CONTEXT} is not marked pending; only a full run marks its own verdict pending." + exit 0 + fi + if [[ "$same_repo" != true ]]; then + echo "::notice::fork pull request: ${STATUS_CONTEXT} is not marked pending; a fork never carries a verdict forward." + exit 0 + fi + case "${GITHUB_EVENT_NAME:-}" in + pull_request | pull_request_target) ;; + *) + echo "::notice::${GITHUB_EVENT_NAME:-unknown} event: ${STATUS_CONTEXT} is not marked pending; only a pull request has contract-only runs." + exit 0 + ;; + esac + # shellcheck disable=SC2310 # write_status prints its own error; the caller exits on it. + if ! write_status pending 'Full run in flight; lanes not yet aggregated.'; then + exit 1 + fi + echo "Recorded ${STATUS_CONTEXT}=pending on ${SHA}; a contract-only run fails until this run's gate records its verdict." + exit 0 +fi + # --------------------------------------------------------------------------- # Carry-forward mode. Branched on first, before `same-repo`: the caller's # predicate makes `contract-only` false for every fork event, so the @@ -255,22 +339,44 @@ read_carried_state() { carried_state_read=true } -# Both Actions reads need `actions: read`, which an explicit `permissions:` block -# does not grant by default. Name the scope rather than printing a bare 403, and -# never treat the refusal as permission to pass: the caller falls through to the -# status read it would have done without the wait. -warn_actions_read_failed() { - local endpoint="$1" +# The Actions calls need a scope an explicit `permissions:` block does not grant +# by default: `actions: read` to wait, `actions: write` to re-run siblings. Name +# the scope rather than printing a bare 403, and say what the run does instead. +# Neither caller treats a refusal as permission to pass. +warn_actions_failed() { + local endpoint="$1" scope="$2" consequence="$3" if [[ "$GH_HTTP_STATUS" == 403 ]]; then - echo "::warning::${endpoint} returned HTTP 403; the ci-status job needs 'actions: read' to wait for a sibling run. Reading the recorded status without waiting." + echo "::warning::${endpoint} returned HTTP 403; the ci-status job needs '${scope}'. ${consequence}" else - echo "::warning::could not read ${endpoint} (HTTP ${GH_HTTP_STATUS:-unknown}); reading the recorded status without waiting." + echo "::warning::could not read ${endpoint} (HTTP ${GH_HTTP_STATUS:-unknown}). ${consequence}" fi cat "$gh_stderr" >&2 } -# Seconds actually slept, and the sibling run ids waited on, for the ceiling -# message. Set by wait_for_sibling_runs. +# Sets `workflow_id` to this run's workflow, the key both the wait and the +# sibling re-run list runs under, and validates `GITHUB_RUN_ID`, which both +# exclude from that list. Returns 1 after a warning ending in `consequence`. +workflow_id="" +resolve_workflow_id() { + local scope="$1" consequence="$2" run_id="${GITHUB_RUN_ID:-}" + if [[ ! "$run_id" =~ ^[0-9]+$ ]]; then + echo "::warning::GITHUB_RUN_ID is not a run id; cannot exclude this run from the runs on ${SHA}. ${consequence}" + return 1 + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${run_id}"; then + warn_actions_failed "repos/${REPOSITORY}/actions/runs/${run_id}" "$scope" "$consequence" + return 1 + fi + workflow_id="$(jq -r '.workflow_id // ""' <"$gh_stdout")" + if [[ ! "$workflow_id" =~ ^[0-9]+$ ]]; then + echo "::warning::run ${run_id} reported no workflow_id. ${consequence}" + return 1 + fi +} + +# Wall-clock seconds since the wait began, and the sibling run ids waited on, +# for the ceiling message. Set by wait_for_sibling_runs. carry_forward_waited=0 carry_forward_wait_note="" @@ -295,19 +401,14 @@ set_wait_note() { # that fails closed on its own failure; there is no outcome that passes without # one completed read. wait_for_sibling_runs() { - local run_id="${GITHUB_RUN_ID:-}" workflow_id ids all_ids="" sleep_for remaining - if [[ ! "$run_id" =~ ^[0-9]+$ ]]; then - echo "::warning::GITHUB_RUN_ID is not a run id; cannot exclude this run from its own wait set. Reading the recorded status without waiting." - return 0 - fi - # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. - if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${run_id}"; then - warn_actions_read_failed "repos/${REPOSITORY}/actions/runs/${run_id}" - return 0 - fi - workflow_id="$(jq -r '.workflow_id // ""' <"$gh_stdout")" - if [[ ! "$workflow_id" =~ ^[0-9]+$ ]]; then - echo "::warning::run ${run_id} reported no workflow_id; reading the recorded status without waiting." + local ids all_ids="" sleep_for remaining wait_started now + local no_wait='Reading the recorded status without waiting.' + # `date`, not bash's `SECONDS`: an external clock is one the harness can + # advance from its `sleep` and `gh` shims, so the ceiling arithmetic stays + # real in tests without a test-only knob in this script. + wait_started="$(date +%s)" + # shellcheck disable=SC2310 # resolve_workflow_id warns itself; the caller degrades to one read. + if ! resolve_workflow_id 'actions: read' "$no_wait"; then return 0 fi while :; do @@ -316,7 +417,7 @@ wait_for_sibling_runs() { # is already far past the burst this wait exists for. # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. if ! gh_api GET "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs?head_sha=${SHA}&per_page=100"; then - warn_actions_read_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" + warn_actions_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" 'actions: read' "$no_wait" # Discard any earlier poll's read so the caller takes the single fresh # read the warning above promises. Without this, a listing that fails on # the second or later poll would decide on the state read before the @@ -332,12 +433,14 @@ wait_for_sibling_runs() { # covers a run object with no `id`, keeping a malformed entry out rather # than waiting on it forever. Sorted so the ids read in run order and the # log line is stable from one poll to the next. - ids="$(jq -r --argjson incomplete "$INCOMPLETE_RUN_STATUSES" --argjson self "$run_id" \ + ids="$(jq -r --argjson incomplete "$INCOMPLETE_RUN_STATUSES" --argjson self "$GITHUB_RUN_ID" \ '[ .workflow_runs[]? | select(.status as $s | $incomplete | index($s)) | select((.id // $self) != $self) | .id ] | sort | join(" ")' \ <"$gh_stdout")" # Kept for the settled-failure check below, which runs after the status # read has overwritten gh_stdout. cp -- "$gh_stdout" "$scratch/runs.json" + now="$(date +%s)" + carry_forward_waited=$((now - wait_started)) # The status read comes AFTER the listing above, never before it: a sibling # writes the status and flips to `completed` moments later, and the other # order lets that completion land between the two calls. @@ -400,17 +503,41 @@ wait_for_sibling_runs() { fi echo "Waiting ${sleep_for}s for in-flight run(s) ${ids} on ${SHA} to finish (waited ${carry_forward_waited}s of ${CARRY_FORWARD_WAIT_SECONDS}s)." sleep "$sleep_for" - carry_forward_waited=$((carry_forward_waited + sleep_for)) done set_wait_note "$all_ids" } -# Every carry-forward red. The run cannot see a verdict recorded after it, so -# the message names both remedies: a failed or missing full run needs the full -# workflow again, and once the status is `success` re-running this run is -# enough, where a new commit would re-run every lane. +# Every carry-forward red. The remedy follows the state read. `pending` names +# the full run in flight, whose own gate check run supersedes this one, so +# nothing needs re-running unless that run was cancelled. A `failure` or `error` +# names the run that recorded it. An absent status, or a read that failed, +# cannot tell whether a full run is coming, so it gives both cases. The closing +# sentence says who replaces this red once the lanes pass: the full run itself +# when this workflow re-runs contract-only siblings (the same step runs in both +# modes, so this run's input is the full run's), otherwise whoever re-runs this +# run, which is cheaper than a new commit that re-runs every lane. fail_carry_forward() { - echo "::error::no successful ${STATUS_CONTEXT} status on ${SHA}; re-run the full workflow${carry_forward_wait_note}. Once ${STATUS_CONTEXT} on ${SHA} is success, re-run this run instead." + local writer="" remedy closing + if [[ -n "$carried_writer_run_id" ]]; then + writer="${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${carried_writer_run_id}" + fi + case "$carried_state" in + pending) + remedy="full run ${writer:-on this SHA} is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled" + ;; + failure | error) + remedy="the lanes verdict is ${carried_state}${writer:+ (${writer})}; fix the failing lane, or re-run that run's failed jobs if the failure was transient" + ;; + *) + remedy="if a full run on this SHA is in flight, its ci-status check supersedes this one; otherwise re-run the full workflow" + ;; + esac + if [[ "$rerun_contract_only_siblings" == true ]]; then + closing="A full run that records ${STATUS_CONTEXT}=success on ${SHA} re-runs this run; re-run it yourself only if it stays red after that." + else + closing="Once ${STATUS_CONTEXT} on ${SHA} is success, re-run this run instead." + fi + echo "::error::no successful ${STATUS_CONTEXT} status on ${SHA}; ${remedy}${carry_forward_wait_note}. ${closing}" exit 1 } @@ -491,10 +618,7 @@ if [[ "$lanes_state" == success ]]; then fi # --------------------------------------------------------------------------- -# Record the verdict as a commit status. Load-bearing on a same-repository run, -# not best-effort: the carry-forward branch reads nothing else, so a silently -# missing status turns every later contract-only run red with no way to tell a -# refused write from a genuinely failing lane. +# Record the verdict as a commit status (see `write_status`). # # A fork pull request is the one exception: its token cannot write a status at # all, and nothing will ever read one for it, so the run reports the lanes @@ -508,32 +632,8 @@ if [[ "$same_repo" != true ]]; then exit 0 fi -target_url="${GITHUB_SERVER_URL:-https://github.com}/${REPOSITORY}/actions/runs/${GITHUB_RUN_ID:-0}" -jq -n \ - --arg state "$lanes_state" \ - --arg context "$STATUS_CONTEXT" \ - --arg description "$lanes_description" \ - --arg target_url "$target_url" \ - '{state: $state, context: $context, description: $description, target_url: $target_url}' \ - >"$scratch/status-payload.json" - -status_written=false -for attempt in 1 2 3 4; do - # shellcheck disable=SC2310 # gh_api handles its own errexit; the retry loop classifies the status. - if gh_api POST "repos/${REPOSITORY}/statuses/${SHA}" --input "$scratch/status-payload.json"; then - status_written=true - break - fi - if [[ "$attempt" -lt 4 ]]; then - delay=$((STATUS_RETRY_BASE_DELAY * (1 << (attempt - 1)))) - echo "::warning::could not record ${STATUS_CONTEXT} on ${SHA} (HTTP ${GH_HTTP_STATUS:-unknown}); retrying in ${delay}s" - sleep "$delay" - fi -done - -if [[ "$status_written" != true ]]; then - cat "$gh_stderr" >&2 - echo "::error::could not record ${STATUS_CONTEXT} on ${SHA} (${GH_HTTP_STATUS:-unknown}); the ci-status job needs statuses: write" +# shellcheck disable=SC2310 # write_status prints its own error; the caller exits on it. +if ! write_status "$lanes_state" "$lanes_description"; then exit 1 fi @@ -542,3 +642,67 @@ echo "Recorded ${STATUS_CONTEXT}=${lanes_state} on ${SHA}." if [[ "$lanes_state" == failure ]]; then exit 1 fi + +# Re-run every failed contract-only run of this workflow on this SHA, so its red +# check run is replaced by one that reads the success just recorded. Runs only +# after that write, so a re-run cannot read anything older. +# +# A contract-only run is recognized by its jobs, not its event, which the run +# object does not carry: in its latest attempt every job but one was skipped, +# and that one failed. A full run always runs more than its gate, and this run +# is excluded by id, so neither is ever re-run. A re-run attempt keeps its +# original event, so it is contract-only again and never reaches this code: no +# loop. A contract-only run that is red for a contract reason (title, +# `do-not-merge`) just goes red again, at the cost of one short job. +# +# The known gap: a contract-only run still in flight when this lists the runs +# read the status before this run wrote it, finishes red after, and is not +# re-run. Its message therefore ends by telling the reader to re-run it if it +# stays red after the full run's success. +# +# Nothing here changes this run's verdict: the success is already recorded, so +# a refusal warns and moves on. +rerun_failed_contract_only_siblings() { + local skip='Not re-running failed contract-only runs.' candidates id contract_shaped rerun_count=0 + # shellcheck disable=SC2310 # resolve_workflow_id warns itself; a refusal only skips the re-run. + if ! resolve_workflow_id 'actions: write' "$skip"; then + return 0 + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs?head_sha=${SHA}&per_page=100"; then + warn_actions_failed "repos/${REPOSITORY}/actions/workflows/${workflow_id}/runs" 'actions: write' "$skip" + return 0 + fi + candidates="$(jq -r --argjson self "$GITHUB_RUN_ID" \ + '[ .workflow_runs[]? | select(.status == "completed" and .conclusion == "failure" and (.id // $self) != $self) | .id ] | sort | join(" ")' \ + <"$gh_stdout")" + for id in $candidates; do + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api GET "repos/${REPOSITORY}/actions/runs/${id}/jobs?filter=latest&per_page=100"; then + warn_actions_failed "repos/${REPOSITORY}/actions/runs/${id}/jobs" 'actions: write' "Not re-running run ${id}." + continue + fi + contract_shaped="$(jq -r '[ .jobs[]? | .conclusion ] | ((map(select(. != "skipped")) == ["failure"]) and (length > 1))' <"$gh_stdout")" + if [[ "$contract_shaped" != true ]]; then + continue + fi + # shellcheck disable=SC2310 # gh_api handles its own errexit; the caller classifies the status. + if ! gh_api POST "repos/${REPOSITORY}/actions/runs/${id}/rerun-failed-jobs"; then + if [[ "$GH_HTTP_STATUS" == 403 ]]; then + echo "::warning::re-running run ${id} returned HTTP 403; rerun-contract-only-siblings needs 'actions: write' on the ci-status job." + else + echo "::warning::could not re-run run ${id} (HTTP ${GH_HTTP_STATUS:-unknown})." + fi + cat "$gh_stderr" >&2 + continue + fi + echo "Re-running failed contract-only run ${id} on ${SHA}." + rerun_count=$((rerun_count + 1)) + done + echo "Re-ran ${rerun_count} failed contract-only run(s) on ${SHA}." +} + +if [[ "$rerun_contract_only_siblings" == true ]]; then + # shellcheck disable=SC2310 # errexit is off inside on purpose: the success is recorded, and a body jq cannot parse must not turn it red. + rerun_failed_contract_only_siblings || true +fi diff --git a/.github/actions/ci-status/run.test.sh b/.github/actions/ci-status/run.test.sh index 39e7fa5..e729d0b 100755 --- a/.github/actions/ci-status/run.test.sh +++ b/.github/actions/ci-status/run.test.sh @@ -22,21 +22,44 @@ mkdir -p "$fixtures" "$calls" "$shim_directory" sha=deadbeefdeadbeefdeadbeefdeadbeefdeadbeef repository=melodic-software/ci-workflows -# A no-op `sleep` first on PATH. The carry-forward wait accounts for elapsed -# time arithmetically against its own 15-second interval, so shimming the sleep -# keeps the ceiling arithmetic real while the suite runs instantly. Unlike a -# test-only poll-interval knob, it adds nothing to the shipped runner. +# The carry-forward red, one remedy per state read; see `fail_carry_forward`. +fail_prefix="::error::no successful ci-lanes status on ${sha}; " +absent_remedy='if a full run on this SHA is in flight, its ci-status check supersedes this one; otherwise re-run the full workflow' +failure_remedy="the lanes verdict is failure; fix the failing lane, or re-run that run's failed jobs if the failure was transient" +writer_failure_remedy="the lanes verdict is failure (https://github.com/${repository}/actions/runs/4000); fix the failing lane, or re-run that run's failed jobs if the failure was transient" +pending_remedy='full run on this SHA is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled' +manual_closing="Once ci-lanes on ${sha} is success, re-run this run instead." +automatic_closing="A full run that records ci-lanes=success on ${sha} re-runs this run; re-run it yourself only if it stays red after that." + +# A `sleep` and a `date` first on PATH that share a fake clock: sleeping +# advances it, and `date +%s` reads it. The wait's ceiling is elapsed wall-clock +# time, so this keeps the ceiling arithmetic real while the suite runs +# instantly, and adds nothing test-only to the shipped runner. Every sleep is +# logged so a case can assert that none happened. +clock_file="$temporary_directory/clock" +sleep_log="$temporary_directory/sleep.log" cat >"$shim_directory/sleep" <<'SLEEP_SHIM' #!/usr/bin/env bash -exit 0 +printf '%s\n' "$1" >>"$SLEEP_LOG" +printf '%s\n' "$(($(cat "$FAKE_CLOCK") + $1))" >"$FAKE_CLOCK" SLEEP_SHIM chmod +x "$shim_directory/sleep" +cat >"$shim_directory/date" <<'DATE_SHIM' +#!/usr/bin/env bash +if [[ "$*" != '+%s' ]]; then + echo "date shim: only +%s is supported, got: $*" >&2 + exit 1 +fi +cat "$FAKE_CLOCK" +DATE_SHIM +chmod +x "$shim_directory/date" # A `gh` shim first on PATH: it serves fixture JSON keyed by method plus API # path, records every call so the harness can assert on writes that did and did # not happen, can fail a keyed call a fixed number of times before succeeding -# (the retry case), and can serve a DIFFERENT body on the Nth call to the same -# key (`..json`), which is how a status appearing mid-wait is fixtured. +# (the retry case), can serve a DIFFERENT body on the Nth call to the same key +# (`..json`), which is how a status appearing mid-wait is fixtured, and +# can make every call to a key take time on the fake clock (`.advance`). cat >"$shim_directory/gh" <<'SHIM' #!/usr/bin/env bash set -uo pipefail @@ -93,6 +116,10 @@ if [[ -f "$count_file" ]]; then fi printf '%s\n' "$call_number" >"$count_file" +if [[ -f "$GH_FIXTURES/${key}.advance" ]]; then + printf '%s\n' "$(($(cat "$FAKE_CLOCK") + $(cat "$GH_FIXTURES/${key}.advance")))" >"$FAKE_CLOCK" +fi + fail_times="$GH_FIXTURES/${key}.fail-times" if [[ -f "$fail_times" ]]; then remaining="$(cat "$fail_times")" @@ -140,6 +167,8 @@ run_case() { shift $(($# > 5 ? 5 : $#)) local actual_status : >"$gh_log" + : >"$sleep_log" + printf '0\n' >"$clock_file" rm -rf -- "$calls" mkdir -p "$calls" set +e @@ -148,6 +177,8 @@ run_case() { GH_LOG="$gh_log" \ GH_FIXTURES="$fixtures" \ GH_CALLS="$calls" \ + FAKE_CLOCK="$clock_file" \ + SLEEP_LOG="$sleep_log" \ GH_TOKEN=fixture-token \ RESULTS="$results" \ TREAT_SKIPPED_AS="$treat_skipped_as" \ @@ -155,11 +186,14 @@ run_case() { SAME_REPO="$same_repo" \ STATUS_CONTEXT=ci-lanes \ CARRY_FORWARD_WAIT_SECONDS=0 \ + RERUN_CONTRACT_ONLY_SIBLINGS=false \ + RECORD_PENDING=false \ REPOSITORY="$repository" \ SHA="$sha" \ STATUS_RETRY_BASE_DELAY=0 \ GITHUB_SERVER_URL=https://github.com \ GITHUB_RUN_ID=4242 \ + GITHUB_EVENT_NAME=pull_request \ "$@" \ bash "$script_directory/run.sh" >"$log_file" 2>&1 actual_status=$? @@ -223,6 +257,30 @@ expect_gh_call_before() { fi } +# expect_status_reads +# The shim counts calls per key, so this pins how many times the status list +# was read, independent of which fixture answered. +expect_status_reads() { + local expected="$1" actual=0 + local count_file="$calls/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.count" + if [[ -f "$count_file" ]]; then + actual="$(cat "$count_file")" + fi + if [[ "$actual" -ne "$expected" ]]; then + echo "FAIL: expected $expected status read(s), got $actual" + cat "$gh_log" + failures=$((failures + 1)) + fi +} + +expect_no_sleep() { + if [[ -s "$sleep_log" ]]; then + echo 'FAIL: expected no sleep, got:' + cat "$sleep_log" + failures=$((failures + 1)) + fi +} + # The shim logs `$*`, which never contains the literal `gh api` — it starts at # `api -X GET …` — so asserting on that string could never fire. Assert the log # is empty instead. @@ -352,7 +410,7 @@ echo 'case: full mode fails the run when the status write is refused after retri printf '%s\n' 99 >"$fixtures/POST_repos_melodic-software_ci-workflows_statuses_${sha}.fail-times" run_case 1 'success success' pass '' expect_log 'All lanes passed or were skipped.' -expect_log "::error::could not record ci-lanes on ${sha} (500); the ci-status job needs statuses: write" +expect_log "::error::could not record ci-lanes on ${sha} (500); the calling job needs statuses: write" rm -f -- "$fixtures/POST_repos_melodic-software_ci-workflows_statuses_${sha}.fail-times" # Without the retry loop the first 500 fails the run. @@ -460,19 +518,22 @@ expect_no_log 'All lanes passed' echo 'case: carry-forward fails on a recorded ci-lanes failure' status_list "[$(bot_status 100 failure)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}. ${manual_closing}" # Without the explicit `== success` test, a pending status would ride through. +# Without the pending branch of the message, the reader is told to re-run a +# full run that is still in flight. echo 'case: carry-forward fails on a pending ci-lanes status' status_list "[$(bot_status 100 pending)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${pending_remedy}. ${manual_closing}" +expect_no_log 're-run the full workflow' # Without the context filter, another context's success would satisfy the gate. echo 'case: carry-forward fails when no entry carries the ci-lanes context' status_list '[{"context":"other-lane","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the context filter, the FIRST entry (a failure under another context) # would decide. @@ -488,7 +549,7 @@ echo 'case: carry-forward fails closed when the status list cannot be read' rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.json" printf '%s\n' 'gh: Not Found (HTTP 404)' >"$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.err" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statuses.err" # --- carry-forward: forged statuses ---------------------------------------- @@ -501,7 +562,7 @@ rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_commits_${sha}_statu echo 'case: a forged success by a user account does not satisfy the carry-forward' status_list "[$(user_status 200 success),$(bot_status 100 failure)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" # Without newest-first selection an older user failure would shadow the bot's # real success. @@ -515,26 +576,35 @@ expect_log "Carried forward: ci-lanes is success on ${sha}" echo 'case: a bot failure newer than a bot success fails' status_list "[$(bot_status 200 failure),$(bot_status 100 success)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" + +# The marker a full run's first job writes under `record-pending`. Without +# newest-id selection the older success is carried forward while that full run +# is still in flight, which is the stale green the marker exists to stop. +echo 'case: a bot pending newer than a bot success fails' +status_list "[$(writer_status 200 pending 2026-09-05T12:00:20Z 4000),$(bot_status 100 success)]" +run_case 1 'skipped skipped' pass true +expect_log "${fail_prefix}full run https://github.com/${repository}/actions/runs/4000 is still in flight" +expect_no_log 'Carried forward' # Without the empty-list guard an absent status would read as an empty state. echo 'case: an empty status list fails the carry-forward' status_list '[]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the creator filter, a context that ONLY a user ever wrote satisfies # the gate — the plant-then-label attack in its simplest form. echo 'case: the ci-lanes context present only from a user account fails' status_list "[$(user_status 100 success)]" run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the Bot type check, an account merely NAMED like the bot passes. echo 'case: a user account impersonating the bot login fails' status_list '[{"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"User"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without `max_by(.id)` the selection depends on the array order the API # happens to return; an oldest-first list would then hand back the stale @@ -542,7 +612,7 @@ expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full wo echo 'case: an oldest-first status list still selects the newest bot entry' status_list '[{"id":10,"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}},{"id":20,"context":"ci-lanes","state":"failure","creator":{"login":"github-actions[bot]","type":"Bot"}}]' run_case 1 'skipped skipped' pass true -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" echo 'case: an oldest-first status list still carries a newer bot success forward' status_list '[{"id":10,"context":"ci-lanes","state":"failure","creator":{"login":"github-actions[bot]","type":"Bot"}},{"id":20,"context":"ci-lanes","state":"success","creator":{"login":"github-actions[bot]","type":"Bot"}}]' @@ -586,7 +656,25 @@ current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${pending_remedy} (waited 30s of 30s on in-flight run(s): 4000)" + +# The ceiling is elapsed wall-clock time, API calls included. Each listing here +# takes 10 seconds, so after one 15-second sleep 35 seconds have passed and the +# 30-second ceiling is reached on the second poll. Counting only the sleeps +# (15, then 30) would poll a third time and report 30 of 30, which is how a +# large ceiling used to outlast the job budget sized for it. +echo 'case: the ceiling counts elapsed wall-clock time, not summed sleeps' +clear_status_fixtures +clear_run_fixtures +status_list "[$(bot_status 100 pending)]" +current_run +workflow_runs "[${earlier_full_run}]" +printf '%s\n' 10 >"$fixtures/${workflow_runs_key}.advance" +run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 +expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 10s of 30s)." +expect_log "${fail_prefix}${pending_remedy} (waited 35s of 30s on in-flight run(s): 4000)" +expect_no_log 'waited 30s of 30s' +rm -f -- "$fixtures/${workflow_runs_key}.advance" # What decides the gate is the verdict the sibling left behind, not whatever was # on the SHA when this run started. Serving nothing on the first read and a @@ -601,7 +689,7 @@ workflow_runs_on_call 1 "[${earlier_full_run}]" workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log 'Waiting 15s for in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${failure_remedy} (waited 15s of 60s on in-flight run(s): 4000)" # A settled failure ends the wait even with a sibling incomplete, when that # sibling is not the writer re-running (this status names no writer at all). @@ -614,7 +702,7 @@ current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${failure_remedy}" # Without the self-exclusion term this run waits on itself forever; without the # status filter it waits on a run that already finished. @@ -628,7 +716,7 @@ run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_gh_call 'actions/workflows/777/runs' expect_no_log 'Waiting ' expect_no_log 'in-flight run(s):' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" # Without the 403 branch the missing scope reads as "no earlier run", which is # the same outcome but unattributable. The run degrades to @@ -641,7 +729,7 @@ current_run printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/${workflow_runs_key}.err" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::repos/${repository}/actions/workflows/777/runs returned HTTP 403; the ci-status job needs 'actions: read'" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' # Without discarding the earlier poll's read, a listing that fails PART WAY @@ -672,16 +760,27 @@ expect_log "::warning::repos/${repository}/actions/runs/4242 returned HTTP 403; expect_no_gh_call 'actions/workflows/' # Without the `> 0` guard, a consumer that disabled the wait still pays two -# Actions API calls and still needs the `actions: read` scope. -echo 'case: carry-forward-wait-seconds 0 disables the wait and makes no Actions API call' +# Actions API calls, still needs the `actions: read` scope, and holds its runner +# while a sibling is in flight. Wait 0 is one status read and nothing else. +echo 'case: carry-forward-wait-seconds 0 reads the status once, never waits, and makes no Actions API call' clear_status_fixtures clear_run_fixtures status_list '[]' current_run workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=0 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_gh_call 'actions/' +expect_status_reads 1 +expect_no_sleep + +echo 'case: carry-forward-wait-seconds 0 carries a recorded success after one read' +clear_status_fixtures +status_list "[$(bot_status 100 success)]" +run_case 0 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=0 +expect_log "Carried forward: ci-lanes is success on ${sha}" +expect_status_reads 1 +expect_no_sleep # Without the runner-side default, an action.yml regression that stopped passing # the input would silently disable the wait rather than fall back to 240. @@ -742,7 +841,7 @@ current_run status_list '[]' workflow_runs "[$(run_entry 3000 completed 2026-09-05T11:59:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' expect_no_log 'waited ' @@ -760,7 +859,7 @@ workflow_runs "[$(run_entry 4100 in_progress "$current_run_created_at"),$(run_en run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log "Waiting 15s for in-flight run(s) 4100 4300 on ${sha} to finish (waited 0s of 30s)." expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4100 4300' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4100 4300)" +expect_log "${fail_prefix}${absent_remedy} (waited 30s of 30s on in-flight run(s): 4100 4300)" # Without `error` in the settled set this waits the full 60s on a verdict that # can never change, then fails anyway. `error` is a terminal commit-status state @@ -773,7 +872,7 @@ status_list "[$(bot_status 100 error)]" workflow_runs "[${earlier_full_run}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}the lanes verdict is error; fix the failing lane" # --- carry-forward: a stale failure while its writer re-runs ---------------- # @@ -809,7 +908,7 @@ status_list "[${stale_failure}]" workflow_runs "[${writer_rerun}]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=30 expect_log '::warning::reached the 30s carry-forward-wait-seconds ceiling with in-flight run(s) 4000' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 30s of 30s on in-flight run(s): 4000). Once ci-lanes on ${sha} is success, re-run this run instead." +expect_log "${fail_prefix}${writer_failure_remedy} (waited 30s of 30s on in-flight run(s): 4000). ${manual_closing}" # Without the start-time term this waits on the writer's own first attempt, # which records the status a moment before it completes, and turns a prompt @@ -822,7 +921,7 @@ status_list "[${stale_failure}]" workflow_runs "[$(attempt_entry 4000 in_progress 2026-09-05T11:59:00Z 1 2026-09-05T11:59:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${writer_failure_remedy}" # Without narrowing the wait to the writer, a failed SHA brings back the mutual # wait: each contract-only run holds the failure open for the other until the @@ -836,7 +935,7 @@ status_list "[${stale_failure}]" workflow_runs "[$(attempt_entry 4100 in_progress "$current_run_created_at" 1 "$current_run_created_at"),$(attempt_entry 4300 in_progress 2026-09-05T12:00:00Z 2 2026-09-05T12:01:00Z)]" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_no_log 'Waiting ' -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow." +expect_log "${fail_prefix}${writer_failure_remedy}. ${manual_closing}" # Without the unsettled-state fail a re-run of the full run that ends without # recording anything would leave an absent status to pass. No status names no @@ -850,20 +949,36 @@ workflow_runs_on_call 1 "[${writer_rerun}]" workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 0s of 60s)." -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)." +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000)." expect_no_log 'Carried forward' -# Without the second remedy the message only ever says to re-run the full -# workflow, and a red that outlived its failure gets a new commit (every lane -# again) where re-running this run is enough. -echo 'case: the failure message names the re-run-this-run remedy for a later success' +# Without the closing remedy a red that outlived its failure gets a new commit +# (every lane again) where re-running this run is enough. Without naming the +# writer the reader has to hunt for the run whose failed jobs to re-run. +echo 'case: the failure message names the writer and the re-run-this-run remedy for a later success' clear_status_fixtures clear_run_fixtures current_run status_list "[${stale_failure}]" workflow_runs '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow. Once ci-lanes on ${sha} is success, re-run this run instead." +expect_log "${fail_prefix}${writer_failure_remedy}. ${manual_closing}" + +# The same step runs in both modes, so this run's rerun-contract-only-siblings +# is the full run's. Without branching the closing on it, the reader is told to +# re-run by hand a run the full run is about to re-run itself. +echo 'case: with rerun-contract-only-siblings the closing says the full run re-runs this run' +clear_status_fixtures +clear_run_fixtures +status_list "[$(writer_status 200 pending 2026-09-05T12:00:20Z 4000)]" +run_case 1 'skipped skipped' pass true true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "${fail_prefix}full run https://github.com/${repository}/actions/runs/4000 is still in flight, and its own ci-status check supersedes this one when it finishes. Re-run that run only if it was cancelled. ${automatic_closing}" +expect_no_log "$manual_closing" +expect_no_log 're-run the full workflow' +# A contract-only run never re-runs anything, even with the input on: the +# re-run is a full-mode step, which is what keeps it from looping. +expect_no_gh_call 'rerun-failed-jobs' +expect_no_gh_call 'actions/' # The documented trade, pinned deliberately from the passing side. A recorded # success ends the wait even with a sibling in flight, so a re-run of this SHA @@ -893,7 +1008,7 @@ workflow_runs "[${earlier_full_run}]" printf '%s\n' 'gh: Not Found (HTTP 404)' >"$fixtures/${statuses_key}.err" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::could not read repos/${repository}/commits/${sha}/statuses (HTTP 404)" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_log 'gh: Not Found (HTTP 404)' rm -f -- "$fixtures/${statuses_key}.err" @@ -929,7 +1044,7 @@ workflow_runs_on_call 1 "[${earlier_full_run}]" workflow_runs_on_call 2 "[${earlier_full_run},$(run_entry 4700 in_progress "$current_run_created_at")]" printf '%s\n' 2 >"$fixtures/${statuses_key}.fail-on-call" run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000 4700)" +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000 4700)" rm -f -- "$fixtures/${statuses_key}.fail-on-call" # A sibling can finish without ever recording a verdict: cancelled, or failed @@ -945,7 +1060,7 @@ workflow_runs_on_call 2 '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "Waiting 15s for in-flight run(s) 4000 on ${sha} to finish (waited 0s of 60s)." expect_log "No run on ${sha} is still in flight after 15s; using the ci-lanes status." -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow (waited 15s of 60s on in-flight run(s): 4000)" +expect_log "${fail_prefix}${absent_remedy} (waited 15s of 60s on in-flight run(s): 4000)" # A truncated body is what a cut-off response looks like: valid JSON up to the # point the connection dropped. `status_list` writes its argument verbatim, so @@ -961,11 +1076,182 @@ status_list "[$(bot_status 100 success)" workflow_runs '[]' run_case 1 'skipped skipped' pass true true CARRY_FORWARD_WAIT_SECONDS=60 expect_log "::warning::could not read repos/${repository}/commits/${sha}/statuses" -expect_log "::error::no successful ci-lanes status on ${sha}; re-run the full workflow" +expect_log "${fail_prefix}${absent_remedy}" expect_no_log 'Waiting ' clear_status_fixtures +# --- pending mode ---------------------------------------------------------- + +status_write_key="POST_repos_melodic-software_ci-workflows_statuses_${sha}" + +# Without pending mode, a contract-only run that reads the status while a full +# run is in flight carries an older run's success forward over lanes nobody has +# run yet. The marker names this run, which is how a contract-only red names +# the run to wait for. No aggregation: `results` is empty here and must not fail. +echo 'case: record-pending marks ci-lanes pending on a same-repository pull request and stops' +run_case 0 '' pass false true RECORD_PENDING=true +expect_status_payload '"state": "pending"' +expect_status_payload '"context": "ci-lanes"' +expect_status_payload "\"target_url\": \"https://github.com/${repository}/actions/runs/4242\"" +expect_log "Recorded ci-lanes=pending on ${sha}" +expect_no_log 'results is required' +expect_no_gh_call 'actions/' +expect_no_gh_call "commits/${sha}/statuses" + +echo 'case: record-pending writes on a pull_request_target event too' +run_case 0 '' pass false true RECORD_PENDING=true GITHUB_EVENT_NAME=pull_request_target +expect_status_payload '"state": "pending"' + +# Without the contract-only guard this run would mark the verdict pending, and +# it and every contract-only sibling would go red with no full run coming to +# overwrite the marker. +echo 'case: record-pending on a contract-only event writes nothing and passes' +run_case 0 '' pass true true RECORD_PENDING=true +expect_log '::notice::contract-only event: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the fork guard the read-only token's refused write fails the job. +echo 'case: record-pending on a fork pull request writes nothing and passes' +run_case 0 '' pass false false RECORD_PENDING=true +expect_log '::notice::fork pull request: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the event guard a push leaves a pending marker on a main commit that +# no full pull-request run will ever overwrite. +echo 'case: record-pending on a push writes nothing and passes' +run_case 0 '' pass '' true RECORD_PENDING=true GITHUB_EVENT_NAME=push +expect_log '::notice::push event: ci-lanes is not marked pending' +expect_no_gh_calls_at_all + +# Without the load-bearing write a refused marker passes silently, and the stale +# green it exists to stop comes back with nothing to say why. +echo 'case: record-pending fails the run when the write is refused after retries' +printf '%s\n' 99 >"$fixtures/${status_write_key}.fail-times" +run_case 1 '' pass false true RECORD_PENDING=true +expect_log "::error::could not record ci-lanes on ${sha} (500); the calling job needs statuses: write" +rm -f -- "$fixtures/${status_write_key}.fail-times" + +echo 'case: an unrecognized record-pending value is rejected' +run_case 1 '' pass false true RECORD_PENDING=yes +expect_log "::error::record-pending must be 'true' or 'false', got: yes" +expect_no_gh_calls_at_all + +# --- full mode: re-running failed contract-only siblings ------------------- + +# completed_run +completed_run() { + printf '{"id":%s,"status":"completed","conclusion":"%s"}' "$1" "$2" +} + +# jobs_for ... +# The latest attempt's jobs of run , one per conclusion. +jobs_for() { + local id="$1" jobs="" conclusion + shift + for conclusion in "$@"; do + jobs="${jobs}${jobs:+,}{\"name\":\"job\",\"conclusion\":\"${conclusion}\"}" + done + printf '{"jobs":[%s]}' "$jobs" >"$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_${id}_jobs.json" +} + +clear_rerun_fixtures() { + rm -f -- "$fixtures"/GET_repos_melodic-software_ci-workflows_actions_runs_*_jobs.* \ + "$fixtures"/POST_repos_melodic-software_ci-workflows_actions_runs_* +} + +# 5000 and 4242 (this run, completed only in this fixture) are contract-only +# shaped: one failed gate, every other job skipped. 5100 is a failed full run, +# 5500 a failed single-job run, and 5200, 5300 and 5400 did not fail. +clear_run_fixtures +clear_rerun_fixtures +current_run +workflow_runs "[$(completed_run 5000 failure),$(completed_run 5100 failure),$(completed_run 5200 success),$(run_entry 5300 in_progress 2026-09-05T12:00:40Z),$(completed_run 5400 cancelled),$(completed_run 5500 failure),$(completed_run 4242 failure)]" +jobs_for 5000 failure skipped skipped +jobs_for 5100 success failure failure +jobs_for 5500 failure +jobs_for 4242 failure skipped skipped + +# Without the re-run the contract-only red stays beside this run's green until +# someone re-runs it by hand. Without the shape test the failed full run 5100 is +# re-run too, every lane again for a verdict already recorded; without the +# single-job guard so is 5500; without the conclusion filter the green, +# in-flight and cancelled runs are fetched; without excluding this run's id it +# re-runs itself. +echo 'case: a full-mode success re-runs only the failed contract-only siblings' +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "Recorded ci-lanes=success on ${sha}." +expect_gh_call "POST repos/${repository}/actions/runs/5000/rerun-failed-jobs" +expect_log "Re-running failed contract-only run 5000 on ${sha}." +expect_log "Re-ran 1 failed contract-only run(s) on ${sha}." +for not_rerun in 5100 5500 4242; do + expect_no_gh_call "actions/runs/${not_rerun}/rerun-failed-jobs" +done +for not_fetched in 5200 5300 5400 4242; do + expect_no_gh_call "actions/runs/${not_fetched}/jobs" +done +# The re-run reads the status, so it must come after the write. +expect_gh_call_before "POST repos/${repository}/statuses/${sha}" 'rerun-failed-jobs' + +# Without the success gate a failing run re-runs siblings that can only read +# its failure and go red again. +echo 'case: a full-mode failure re-runs nothing' +run_case 1 'success failure' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_status_payload '"state": "failure"' +expect_no_gh_call 'actions/' + +# Without the opt-in every consumer would need `actions: write` on upgrade. +echo 'case: rerun-contract-only-siblings off by default makes no Actions call' +run_case 0 'success success' pass false true +expect_no_gh_call 'actions/' + +echo 'case: a fork run re-runs nothing' +run_case 0 'success success' pass false false RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_no_gh_call 'actions/' + +# The verdict is already recorded, so a refused re-run must not turn it red, +# and must not stop the next candidate. +echo 'case: a refused re-run warns naming actions: write and moves to the next sibling' +jobs_for 5600 failure skipped +workflow_runs "[$(completed_run 5000 failure),$(completed_run 5600 failure)]" +printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/POST_repos_melodic-software_ci-workflows_actions_runs_5000_rerun-failed-jobs.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::re-running run 5000 returned HTTP 403; rerun-contract-only-siblings needs 'actions: write' on the ci-status job." +expect_gh_call "POST repos/${repository}/actions/runs/5600/rerun-failed-jobs" +expect_log "Re-ran 1 failed contract-only run(s) on ${sha}." +rm -f -- "$fixtures/POST_repos_melodic-software_ci-workflows_actions_runs_5000_rerun-failed-jobs.err" + +echo 'case: a jobs read failure skips that sibling and still passes' +printf '%s\n' 'gh: Internal Server Error (HTTP 500)' >"$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_5000_jobs.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::could not read repos/${repository}/actions/runs/5000/jobs (HTTP 500). Not re-running run 5000." +expect_no_gh_call 'actions/runs/5000/rerun-failed-jobs' +expect_gh_call "POST repos/${repository}/actions/runs/5600/rerun-failed-jobs" +rm -f -- "$fixtures/GET_repos_melodic-software_ci-workflows_actions_runs_5000_jobs.err" + +echo 'case: a 403 listing the runs warns naming actions: write and still passes' +printf '%s\n' 'gh: Resource not accessible by integration (HTTP 403)' >"$fixtures/${workflow_runs_key}.err" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "::warning::repos/${repository}/actions/workflows/777/runs returned HTTP 403; the ci-status job needs 'actions: write'. Not re-running failed contract-only runs." +expect_no_gh_call 'rerun-failed-jobs' + +# Without errexit off inside the re-run, jq failing on a cut-off body exits the +# script nonzero after the success is recorded, turning a green gate red. +echo 'case: a malformed run list re-runs nothing and still passes' +rm -f -- "$fixtures/${workflow_runs_key}.err" +printf '%s' '{"workflow_runs":[' >"$fixtures/${workflow_runs_key}.json" +run_case 0 'success success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=true +expect_log "Recorded ci-lanes=success on ${sha}." +expect_no_gh_call 'rerun-failed-jobs' + +echo 'case: an unrecognized rerun-contract-only-siblings value is rejected' +run_case 1 'success' pass false true RERUN_CONTRACT_ONLY_SIBLINGS=on +expect_log "::error::rerun-contract-only-siblings must be 'true' or 'false', got: on" +expect_no_gh_calls_at_all + +clear_run_fixtures +clear_rerun_fixtures + # --- input validation ------------------------------------------------------ # # Every one of these values is interpolated into a `gh api` path. @@ -991,19 +1277,29 @@ expect_status_payload '"context": "CI Lanes"' echo 'case: action.yml still wires every input this harness exercises' action_metadata="$script_directory/action.yml" -for input_name in results treat-skipped-as contract-only same-repo status-context carry-forward-wait-seconds token repository sha; do +for input_name in results treat-skipped-as contract-only same-repo status-context carry-forward-wait-seconds rerun-contract-only-siblings record-pending token repository sha; do if ! grep -qE "^ ${input_name}:" "$action_metadata"; then echo "FAIL: action.yml declares no '${input_name}' input" failures=$((failures + 1)) fi done -for environment_name in RESULTS TREAT_SKIPPED_AS CONTRACT_ONLY SAME_REPO STATUS_CONTEXT CARRY_FORWARD_WAIT_SECONDS GH_TOKEN REPOSITORY SHA; do +for environment_name in RESULTS TREAT_SKIPPED_AS CONTRACT_ONLY SAME_REPO STATUS_CONTEXT CARRY_FORWARD_WAIT_SECONDS RERUN_CONTRACT_ONLY_SIBLINGS RECORD_PENDING GH_TOKEN REPOSITORY SHA; do if ! grep -qF " ${environment_name}: " "$action_metadata"; then echo "FAIL: action.yml does not pass '${environment_name}' to run.sh" failures=$((failures + 1)) fi done +# Both A+ inputs are opt-in: a consumer that passes nothing keeps today's +# behavior and needs no new permission. +for input_name in rerun-contract-only-siblings record-pending; do + input_default="$(awk -v key=" ${input_name}:" '$0 == key { found = 1; next } found && /^ default:/ { print; exit }' "$action_metadata")" + if [[ "$input_default" != " default: 'false'" ]]; then + echo "FAIL: action.yml does not default ${input_name} to 'false' (got: ${input_default})" + failures=$((failures + 1)) + fi +done + # The published default is the contract consumers inherit when they pass # nothing, and the design fixes it at 240 seconds (below cursor-plugins' # five-minute ci-status job budget). A drift here is invisible to every case diff --git a/README.md b/README.md index 67aa81a..51f539e 100644 --- a/README.md +++ b/README.md @@ -276,14 +276,25 @@ consumer to audit it. land between the two calls and report both no status and nothing in flight, failing a SHA that does carry a verdict. - **The wait never turns a verdict green.** Reaching the ceiling prints - `::error::no successful status on ; re-run the full workflow`, - extended with how long it waited and on which run ids, and exits 1. There is - no pass-on-timeout path. Every carry-forward red ends with `Once - on is success, re-run this run instead.`: the run cannot see a verdict - recorded after it, and once the lanes pass, re-running the red run replaces - its check run where a new commit would re-run every lane. - - **The calling job needs `actions: read`.** Under an explicit `permissions:` + `::error::no successful status on ; `, extended + with how long it waited and on which run ids, and exits 1. There is no + pass-on-timeout path. The ceiling counts elapsed wall-clock time from the + start of the wait, API calls included, so it overruns by at most the few + seconds one poll's calls take. + + Every carry-forward red names the remedy for the state it read. `pending` + names the full run in flight, whose own `ci-status` check run supersedes the + red one, so nothing needs re-running unless that run was cancelled. A + `failure` or `error` names the run that recorded it: fix the lane, or re-run + that run's failed jobs. An absent status says a full run in flight will + supersede the red one and to re-run the full workflow only when none is. The + closing sentence says who replaces the red once the lanes pass: with + `rerun-contract-only-siblings` on, the full run re-runs it (the same step runs + in both modes, so the contract-only run knows); otherwise `Once on + is success, re-run this run instead.`, which replaces its check run + where a new commit would re-run every lane. + + **A wait above `0` needs `actions: read`.** Under an explicit `permissions:` block the default for every scope is none, so both Actions calls 403 without it. A 403 prints a `::warning::` naming the missing scope and then degrades to a single status read, which is the pre-6b contract: loud, and never a pass on @@ -300,8 +311,73 @@ consumer to audit it. The 15-second poll interval is deliberately not a caller input: the only knob a consumer should have to reason about is the ceiling. - **Size the ceiling from the repository's own measured full run, then derive - `timeout-minutes` from it.** `carry-forward-wait-seconds` must cover the wall + **Recommended: no run waits for another.** Three opt-in settings together + remove the wait, and with it the dependency on how long the full run takes: + + - `carry-forward-wait-seconds: '0'` and `timeout-minutes: 3` on the + `ci-status` job. A contract-only run reads the status once, makes no + Actions call, and finishes in seconds: green on a recorded `success`, red + otherwise. + - `record-pending: 'true'`, called as the first step of the full run's first + job (for example `changes`). It marks `status-context` `pending` on the + head SHA, so a contract-only run on that SHA while the lanes run reads + `pending` and goes red instead of carrying an older `success` forward (a + draft run's, for example). The full run's gate overwrites the marker with + its verdict. It writes only on a same-repository pull request event that is + not contract-only, and passes with a notice otherwise, so it is safe in a + job that runs on every event. `results` is ignored. + - `rerun-contract-only-siblings: 'true'` on the `ci-status` step. After the + full run records `success`, it re-runs, through + `POST /repos/{owner}/{repo}/actions/runs/{run_id}/rerun-failed-jobs`, every + failed run of the same workflow on the SHA whose latest attempt skipped + every job but one failed gate. Full runs and the run itself are never + re-run. A re-run attempt keeps its original event, so it is contract-only + again, reads the new `success`, replaces the red check run, and never + re-runs anything itself. A refusal warns and does not change the verdict. + + A red contract-only check does not hold a merge in the meantime: GitHub's + docs do not say how several same-name check runs on one SHA are judged, and + on claude-code-plugins the merge gate followed the newest one in every case + measured, so the full run's later `ci-status` supersedes the red. The + remaining gap: a contract-only run still in flight when the full run lists + its siblings read the status before it was written, finishes red, and is not + re-run; its message ends by saying to re-run it if it stays red. + + Permissions: the first job needs `statuses: write`; the `ci-status` job needs + `statuses: write` and `actions: write` (which covers the `actions: read` a + wait above `0` needs). + + ```yaml + jobs: + changes: + permissions: + contents: read + statuses: write + steps: + - name: Mark the lanes verdict pending + uses: melodic-software/ci-workflows/.github/actions/ci-status@ # + with: + record-pending: 'true' + # ... change detection + ci-status: + if: always() + timeout-minutes: 3 + permissions: + contents: read + statuses: write + actions: write + steps: + - name: Aggregate lane results + uses: melodic-software/ci-workflows/.github/actions/ci-status@ # + with: + carry-forward-wait-seconds: '0' + rerun-contract-only-siblings: 'true' + results: ${{ needs.changes.result }} ... + ``` + + **To keep a wait instead, size the ceiling from the repository's own measured + full run, then derive `timeout-minutes` from it.** + `carry-forward-wait-seconds` must cover the wall time the contract-only run may have to wait out: the queue wait plus the p95 wall of the full `ci` run on this repository. Set `timeout-minutes` to at least that figure plus two minutes, then set `carry-forward-wait-seconds` to @@ -321,10 +397,11 @@ consumer to audit it. The 60-second margin survives as a **constraint, not a sizing rule**: a ceiling at or above the job budget lets the job timeout preempt the fail-closed error, which reports as a cancelled job rather than the - actionable "re-run the full workflow" message. + actionable message. Because the ceiling counts wall-clock time, the margin + holds at any ceiling size. The `240` default therefore suits only a repository whose full run finishes in - well under two minutes. The values in use across the fleet today: + well under two minutes. The waiting values in use across the fleet today: | repository | `timeout-minutes` | `carry-forward-wait-seconds` | | --- | --- | --- |