Skip to content

fix(ci-status): wait for the re-run of a failed full run before carrying its failure - #647

Draft
kyle-sexton wants to merge 1 commit into
mainfrom
fix/ci-status-stale-rerun-wait
Draft

kyle-sexton wants to merge 1 commit into
mainfrom
fix/ci-status-stale-rerun-wait

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

Closes #646

Summary

In carry-forward mode, ci-status ended its wait as soon as the newest ci-lanes status was success, failure or error. After a full run failed and was re-run, a contract-only event on the same SHA read the old failure at once and went red. It did not wait for the re-run. Reproduced in melodic-software/claude-code-plugins#4670: contract-only run 36675804726 went red at 06:00:01 on the 05:56:46 failure of run 36673988526, while that run's attempt 2 (started 05:57:10) was in flight. Attempt 2 passed at 06:12:49.

The failure message also only ever said re-run the full workflow. Once ci-lanes is success, re-running the red contract-only run is enough.

Fix

  • A settled failure or error is held open only while its writer re-runs. read_carried_state now also reads the chosen entry's created_at and the run id its target_url names. The status stays open, and the wait narrows to that one run, when that run is in flight with a run_started_at later than the status's created_at. A success still ends the wait at once.
  • No mutual wait between contract-only siblings. Only a full run writes the status, so a contract-only sibling is never the writer. Two contract-only runs on a failed SHA still stop at once. A run_attempt > 1 test was rejected because it also matches a re-run contract-only run. A jobs-API discriminator was already rejected in fix(ci-status): wait on any in-flight sibling before reading ci-lanes #562.
  • Start-time term. It keeps the writer's own first attempt out of the wait: that attempt writes the status a moment before it completes.
  • Known limit. A re-run of a different full run on the same SHA does not hold the failure open. That case fails at once, as before.
  • Fail-closed paths are unchanged. A writer re-run slower than the ceiling runs to it and fails. A timeout, a missing status or a skipped lane never passes.
  • Both remedies in one message. Every carry-forward red goes through fail_carry_forward. It keeps re-run the full workflow and appends Once ci-lanes on <sha> is success, re-run this run instead. The run cannot see a verdict recorded after it, so it names both.
  • The run.sh header, the carry-forward-wait-seconds input description and the README section are updated to match.

Verification

  • bash .github/actions/ci-status/run.test.sh: all cases pass on this branch.
  • The same test file run against main's run.sh fails 5 new cases:
    • Behavior: "a stale failure waits for the in-flight re-run of its writer" and "a writer re-run that outlasts the ceiling fails closed" fail on main.
    • Message only: "two contract-only siblings on a failed SHA do not wait on each other", "a missing status still fails after the re-run ends without recording one" and "the failure message names the re-run-this-run remedy" fail on main only on the new message text. main already behaves correctly in the first two. They are regression guards for this change.
  • Mutants of this branch's run.sh:
    • A failure held open by any in-flight sibling fails the "two contract-only siblings" case and three others.
    • run_attempt > 1 in place of the writer test fails the "two contract-only siblings" case.
    • Dropping the start-time term fails "the writer attempt that recorded the failure is not waited on".
  • The fixture timestamps come from the real runs. The failure status 55241522644 (created_at 05:56:46Z) names run 36673988526 in target_url. That run's attempt 1 run_started_at is 05:33:54Z and attempt 2's is 05:57:10Z. The consumer pin 4610c31e writes the same target_url shape.
  • shellcheck on run.sh and run.test.sh, shfmt -d .github/actions/ci-status/, actionlint, markdownlint-cli2 README.md, editorconfig-checker: clean. typos flags only a queueing on run.sh:20, which is already on main. node --test .github/scripts/ci-fanout-consolidation.test.cjs: pass.
  • Not verified here: a live run of the new composite. That needs a release and a consumer repin, and claude-code-plugins#4670 tracks it.

Related

🤖 Generated with Claude Code

…ing its failure

In carry-forward mode a settled failure or error ended the wait at once, so a
contract-only event during a re-run of the failed full run read the stale
failure and went red. The wait now holds a failure open while the run its
target_url names is in flight with an attempt started after the status was
created. Only a full run writes the status, so contract-only siblings on a
failed SHA still stop at once. A success still ends the wait immediately, and
every path stays fail-closed.

Every carry-forward red now also says to re-run this run once ci-lanes is
success, where a new commit would re-run every lane.

Closes #646

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit to melodic-software/claude-code-plugins that referenced this pull request Sep 30, 2026
…5592)

Refs #4670

## Summary

The Contract-only `ci-status` section of `docs/ci-runner-routing.md`
told operators to "try re-running" a red contract-only run and to push a
commit if it stayed red. A throwaway-PR reproduction on #4670 showed
what actually happens, so the section now says so.

## Fix

- Remedy: re-running a red contract-only run after `ci-lanes` is
`success` passes the composite and replaces the red check run. No new
commit is needed.
- Defect (a): a body edit while a failed full run is being re-run reads
the stale `ci-lanes` failure at once and goes red. The fix is upstream:
melodic-software/ci-workflows#646.
- The ruleset behavior for two same-name `ci-status` check runs stays
marked unverified. The probe could not observe it (the `do-not-merge`
label kept every run red), and no stale-check-clearing step is added.
- `ci.yml` is not repinned here.

## Verification

- `npx markdownlint-cli2 docs/ci-runner-routing.md` reports 0 issues.
- The diff touches only the Contract-only `ci-status` section.
- Run ids and timestamps for each stated behavior are in the results
comment on #4670:
#4670 (comment)

## Related

- #4670 stays open: AC2 to AC5 need the ci-workflows release and a
repin.
- melodic-software/ci-workflows#646 (issue),
melodic-software/ci-workflows#647 (draft fix).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci-status: a stale ci-lanes failure ends the carry-forward wait while its full run is re-running

1 participant