Skip to content
Merged
2 changes: 1 addition & 1 deletion .github/requirements-ci-animation.txt
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# CI-only Linux x64 cp314 wheel pins for the animation regression suites
# (plugins/animation/scripts/test_inkstats.py, test_woodcut_marks.py), so they
# RUN in test-bash instead of skipping green. The versions must equal
# RUN in test-python instead of skipping green. The versions must equal
# plugins/animation/requirements.in, so bump the two together.
# Kept apart from requirements-ci.txt so the jobs that read only that file do
# not install these wheels. Regenerate with:
Expand Down
711 changes: 316 additions & 395 deletions .github/workflows/ci.yml

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions .github/workflows/selection-audit.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,8 +57,8 @@ jobs:
with:
persist-credentials: false
# The suites read what they read only when their tools are present, so
# the toolchains match ci.yml's test-bash lane (Node, Python, ShellCheck,
# shfmt, the inventory parser packages, DuckDB); a suite that skips for a
# the toolchains match ci.yml's test-bash and test-python lanes (Node,
# Python, ShellCheck, shfmt, the inventory parser packages, DuckDB); a suite that skips for a
# missing tool reads less and can hide a gap, never invent one.
- name: Set up Node
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
Expand Down
172 changes: 119 additions & 53 deletions .github/workflows/test-windows.yml

Large diffs are not rendered by default.

44 changes: 30 additions & 14 deletions docs/ci-runner-routing.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,10 +40,14 @@ reviewed pull request.

The `ci-status` required check depends on every **required** workload lane
(`scope`, `lint-repo`, `lint-shell`, `check-plugins`, `check-skills`,
`test-bash`) and requires
`test-bash`, `test-python`, `test-node`) and requires
each result to be `success`, failing closed through execution
(`!cancelled()`, never a success-guard, so a skipped lane cannot report
success to branch protection). On a draft pull request every lane but
success to branch protection). The one exception is an ordered skip: a test
lane the change gives no work skips as a job, and `ci-status` counts that skip
as `success` only when `scope` succeeded and its row for that lane
(`run_bash`, `run_python`, `run_node`) is `false`; any other skip stays red.
On a draft pull request every lane but
`scope` carries a draft gate, so a draft run lints and tests nothing, and its
`ci-status` fails (`draft: lanes not run`) and records `ci-lanes=failure`
(`scripts/check-docs-only-gate.sh` pins both). A green draft would be the
Expand All @@ -57,7 +61,10 @@ longest job in the run: time-to-green is measured to the run's completion, so an
advisory lane was setting the number. Its `on.paths` filters repeat the `shell`,
`python` and `powershell` groups of `ci.yml`'s change-detection table, so a diff
that touches none of them starts no Windows run; the two are kept in step by
hand. The
hand. Inside a run, its `scope-windows` job runs the same planner over the same
diff, and a Windows step runs only when the change selects its suite
(`scripts/test-windows-plan.txt`); a schedule run (09:17 and 16:17 UTC), a
dispatch, and a change to that workflow run every step. The
metadata checks (Conventional Commits title,
`do-not-merge` label, issue linkage) run as the `pr-contract` composite step
inside the same `ci-status` job on the same hosted runner, so they no longer
Expand All @@ -71,13 +78,15 @@ hygiene, the CI configuration, the docs conventions); `lint-shell` runs
ShellCheck and the gates that read shell source; `check-plugins` holds the
plugin manifests, changelogs, hook declarations, `claude plugin validate`, the
counter ceilings and every shared-library sync; `check-skills` holds the skill
and eval contracts, check 25 included; `test-bash` runs the affected contract
suites, with the Node sub-projects and the disk-hygiene module as steps. A
domain is its own job only when that shortens the critical path by more than a
job's fixed cost (about 15 s bare, about 40 s with a toolchain), which is why
those last two ride `test-bash`. Every gate keeps its step name and its `id`,
which is its aggregator key, and each job carries its own
`aggregate-hygiene-results.sh` feed over exactly its own gate steps.
and eval contracts, check 25 included; `test-bash`, `test-python` and
`test-node` run the shell suites, the Python suites and the Node packages the
change selects, and each skips when it has none. A domain is its own job only
when that shortens the critical path by more than a job's fixed cost (about
15 s bare, about 40 s with a toolchain), or, for the test lanes, when it lets a
change that touches none of its ecosystem start no runner for it. Every gate
keeps its step name and its `id`, which is its aggregator key, and each job
carries its own `aggregate-hygiene-results.sh` feed over exactly its own gate
steps.

## What each event tests

Expand Down Expand Up @@ -111,10 +120,17 @@ to the whole tree when there is no diff base or `ci.yml` changed, and the
scheduled run scans everything. Replayed on 20 recent pull requests, every
skipped or narrowed scan landed on a whole-tree success.

`scope` also plans `test-bash`: one leg per 25 selected suites, one to
four (four on the whole tree or an UNMAPPED file), and, per leg, whether its
slice needs the animation wheels, the inventory's parser packages or the DuckDB
CLI. A leg installs only those; the shfmt and DuckDB downloads are cached.
`scope` also plans the test lanes, once, with `scripts/plan-test-lanes.sh`:
each selected suite goes to the lane of its ecosystem (a Node suite with a
sibling `.test.sh` runs through it in `test-bash`), `test-bash` gets one to six
legs of about 120 suite-seconds each and `test-python` one to four of about
180, packed longest first from the measured seconds in
`scripts/suite-seconds.txt`, and each leg installs only the optional toolchains
(the animation wheels, the inventory's parser packages, the DuckDB CLI) its
suites need. `test-node` runs the Node packages the change reaches. An
UNMAPPED file adds the whole corpus of its language. A Python pin runs every
Python suite, a Node pin every Node package, and a change to `ci.yml` or
`.github/actions/` every suite of every lane.

A suite that scans a directory never names the file that changed, so it
declares what it reads in a `# test-scope:` header, and the selector's rule R8
Expand Down
5 changes: 3 additions & 2 deletions scripts/affected-tests-no-suite.txt
Original file line number Diff line number Diff line change
Expand Up @@ -102,8 +102,9 @@ scripts/sync-plugin-options-docs.py
# consumer's plugin cache (which fires only on a plugin-root lockfile) never
# installs its dev toolchain. Its tests are vitest (`*.test.ts`), which is not
# one of the four suite conventions this tool selects, so the whole tree needs
# this entry. The lane covering it is the miro block of the `test` job in
# .github/workflows/ci.yml, gated on `run_node` (matcher `plugins/miro/**`):
# this entry. The lane covering it is the miro block of the `test-node` job in
# .github/workflows/ci.yml, which runs whenever a file under plugins/miro/server/
# changes (scripts/plan-test-lanes.sh):
# `npm ci`, `npm run typecheck`, `npm run lint`, `npm test`, and a stdio smoke
# test that starts src/launch.ts twice from an empty plugin data directory (the
# first launch installs from the lockfile, the second reuses it). Manifests and prose under the
Expand Down
21 changes: 11 additions & 10 deletions scripts/affected-tests.test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1413,11 +1413,11 @@ fi
rm -rf "$repo3"

# --- LIVE repo: ci.yml actually fans the selection out -----------------------
# `--shard` only buys anything if the workflow both PASSES it and creates more
# than one runner to pass it from. Either half alone is silently useless: a
# shard spec with no matrix runs leg 0 of 1 (everything, as before), and a
# matrix with no shard spec runs the whole selection on every leg. Pinned
# together, in the job that owns them.
# The plan's legs only buy anything if the workflow both sizes its matrix from
# them and has each leg run its own suites. Either half alone is silently
# useless: a plan read by no matrix runs one leg, and a matrix whose legs do
# not pick their own list runs nothing or everything. Pinned together, in the
# job that owns them.
live_ci="$REPO_ROOT/.github/workflows/ci.yml"
test_bash_block="$(awk '
/^ [A-Za-z_][A-Za-z0-9_-]*:[[:blank:]]*(#.*)?$/ {
Expand All @@ -1426,13 +1426,14 @@ test_bash_block="$(awk '
job == "test-bash"
' "$live_ci")"
# shellcheck disable=SC2016 # deliberate: these are workflow literals to match, not shell expansions.
if grep -q 'affected-tests\.sh --run --jobs 3 --shard "\$LEG/\$LEGS"' <<<"$test_bash_block" &&
grep -q '^ strategy:' <<<"$test_bash_block" &&
if grep -q 'run-plugin-tests\.sh --jobs 3 --suites-from "\$RUNNER_TEMP/leg-suites\.txt"' <<<"$test_bash_block" &&
grep -q "leg: \${{ fromJSON(needs\.scope\.outputs\.bash_legs || '\[0\]') }}" <<<"$test_bash_block" &&
grep -q 'LEG: \${{ strategy\.job-index }}' <<<"$test_bash_block" &&
grep -q 'LEGS: \${{ strategy\.job-total }}' <<<"$test_bash_block"; then
ok "ci.yml test-bash declares a matrix, runs three suites at a time, and passes the leg through to --shard"
grep -q 'PLAN: \${{ needs\.scope\.outputs\.bash_plan }}' <<<"$test_bash_block" &&
grep -q "jq -r --arg leg \"\$LEG\" '\.\[\$leg\] // \[\] | \.\[\]' <<<\"\$PLAN\"" <<<"$test_bash_block"; then
ok "ci.yml test-bash sizes its matrix from the plan and runs each leg's suites three at a time"
else
fail "ci.yml test-bash no longer fans the affected selection across a matrix at --jobs 3"
fail "ci.yml test-bash no longer fans the planned selection across a matrix at --jobs 3"
fi

# --- ci.yml never asks for more than the proven three ------------------------
Expand Down
Loading
Loading