Shared reusable GitHub Actions workflows for the narduk-enterprises estate (CI-5).
This repository is public. Its own CI and public callers run on GitHub-hosted capacity, with no org-variable or Blacksmith routing. Private callers retain reusable-workflow compatibility and may use manifest-routed self-hosted capacity only where their repository policy permits it. The workflows hold no secrets — every
workflow_callsecret is optional and skips cleanly, andapple.ymlandpython-data.ymldeclare none at all.reusable-browser-tests.ymldeclares one optional secret (NARDUK_PLATFORM_GH_PACKAGES_READ) that falls back to the ephemeralgithub.tokenwhen a caller passes nothing.Callers must pin the full 40-character commit SHA that
v2points to, with a# v2comment (new work;@v1is frozen, see v1 is frozen; adopt v2 when touched) — never a bare@v2tag and never@main— and a public (or possibly-future-public) caller must never pass a self-hosted runner label. That warning is still correct, but it rests on the caller, not on this repo:runneris a free-form caller-supplied value, a repo can go public later, and a fork PR on a public caller can run attacker-controlled code on estate infrastructure. A reusable job whose runner input is left empty resolves it per run from the caller's own visibility: a private caller gets the fleet manifest'slinux-ciorganization-group route, a public (or visibility-unknown) caller gets GitHub-hostedubuntu-latest— see Default route.apple.yml's Mac route is the one input with no default, deliberately.
Fix CI in one place, not 100. Application repos call these workflows via
workflow_call instead of blob-copying YAML. This repo replaces the broken
pattern where ~20 repos carried copies of weekly-drift-check.yml pointing at
narduk-enterprises/narduk-nuxt-template/.github/workflows/reusable-quality.yml@main
— a repo name that no longer exists (renamed to narduk-template), so every
scheduled run 404'd silently.
CI-5 phase 2 (company-hq strategy/workflow-consistency-proposal.md) adds the
three app-shaped CI gates below, each producing the identical
ci / Required check context so branch protection and org rulesets can
require the same string across every repo of that type. See
§2 "Proposed standard"
in that proposal for the full rationale.
apple.yml and python-data.yml were the two named gaps left after that pass,
and company-hq#173 (R-14) is the issue that closed them: "adopt the shared
workflow" had been the standing answer for CI duplication, and for the Apple
and Python-data families there was nothing to adopt — so their duplication was
structural, not neglectful, and naming a workflow that did not exist made the
gap look like an adoption backlog. Both now exist and both have a real adopter
(see "Adopters" below). Still not built: a nuxt-cloudflare-deploy.yml
sibling for push-to-main deploys with real Cloudflare secrets.
The standalone browser / Playwright gap named by company-hq#197 is now
reusable-browser-tests.yml. It consumes a same-run production build and fans
Chromium (and opt-in WebKit) into three shards on the manifest-routed isolated
pool. Each shard reports in its own job log; no report is uploaded or merged
(see "The CI artifact store (2026-09-26)" below). The older browser inputs in
nuxt-cloudflare.yml remain for backward compatibility; adopters that separate
browser CI use the dedicated callable instead of re-embedding the pool contract
in their app-class workflow.
| Workflow | Purpose |
|---|---|
apple.yml |
CI gate for Apple repos (Swift packages, iOS/macOS apps): SwiftLint (official Linux binary) and boundary/plist checks on a Linux runner, swift build / swift test / xcodebuild on the repo-scoped Mac. Two separately-routed runner inputs — that split is the point. Release/signing stays per-repo |
python-data.yml |
CI gate for Python / data-pipeline repos: explicit uv or Python provisioning, pytest, opt-in exact-version Ruff and Pyright (real static checking, not py_compile), plus an extra-checks hook |
reusable-browser-tests.yml |
Private-repo browser CI: validates the exact manifest browser-group object before any shard is scheduled, consumes a same-run production build, asserts the immutable Playwright package/browser image and real launch, runs three Chromium shards plus opt-in WebKit, each reporting in its own job log (no report upload or merge) |
docs-governance.yml |
Thin generic gate for docs/handbook-shaped repos: checkout, optionally provision Python/Node, run one repo-provided check command. Generalizes company-hq's handbook-spine-check.yml / untangle-project-sync.yml shape |
node-library.yml |
CI gate for library / cli project-lifecycle surfaces: script-probed lint/typecheck/test/build, with an optional per-package matrix generalizing narduk-libs' package-gates + verify pattern |
nuxt-cloudflare.yml |
CI gate for nuxt-web / cloudflare-worker surfaces: typecheck (worker + Nuxt split, matching hydrogen), optional unit tests, build, optional extra-scripts, optional web-foundation conformance check, optional Playwright e2e — optionally sharded onto a separately-routed browser pool, the prebuilt application reaching each shard through the R2 CI artifact store — optional wrangler deploy --dry-run validation. CI only — no deploy job (see below) |
reusable-node-ci.yml |
Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8): zero live callers estate-wide, re-verified via gh search code / gh api search/code across narduk-enterprises, narduk-incubator and narduk-enterprises-clients. node-library.yml was always the richer, preferred surface — see workflows#14/#16 |
code-review.yml |
Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8, extending the 2026-09-19 advisory retirement below): zero live pool adopters, file deleted. Current review is the CT650 PR review bot (runners#212) |
cursor-review.yml |
Retired 2026-09-28 (ci-reset W6): replaced by the CT650 PR review bot in narduk-enterprises/runners scripts/pr_review_bot/ (runners#212), which answers /review comments and the review-now / review-deep labels without Actions runner time; file, cursor-review-self.yml, scripts/cursor_review.py, its tests and brief deleted. The prompt moved verbatim to the bot |
closing-syntax-check.yml |
Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8): zero adopters, file deleted. narduk-enterprises/agent-infrastructure keeps its own local, canonical invocation of the same (commit-scanning) checker (agent-infrastructure#837, #1085) |
reusable-weekly-drift-check.yml |
Retired 2026-07-26: no live caller; see workflows#20 and the 2026-07-26 Actions-optimization audit |
red-main-listener.yml |
New 2026-09-24 (narduk-reboot P3-C2 / O-D8). Opens or refreshes ONE "main is red: <workflow-name>" issue, labelled red-main, per repo per listened workflow, the first time that workflow goes red on the default branch; closes it on the next green run. Prerequisite: the adopting repo must already have the red-main label (gh label create red-main --color B60205 --description "Main default-branch CI is failing") — opening the first issue passes labels[]=red-main and GitHub 422s a repo without it, so an adopter that skips this step fails closed on its first real red run (PR #142 review, second round). Optional fail-open portal-mirror POST hook. No automatic revert or rollback — that stays in the app's own deploy tooling |
design-ledger.yml |
New 2026-09-27 (agent-infrastructure#1804). Advisory, never required. Runs the design ledger checker (scripts/dc_ledger.py, a byte copy of agent-infrastructure's canonical file) against a product's design/<canvas>/ledger.json. mode: check goes red while a canvas screen and its code have drifted and writes the table to the job summary; mode: flag keeps one Design drift: <canvas> issue. See Design ledger |
flake-digest.yml |
New 2026-09-24 (narduk-reboot P3-C2). Weekly, advisory: scans a caller's completed runs of one workflow, counts failures in jobs matching quarantine-job-pattern (nuxt-cloudflare.yml's e2e-quarantine by default), and files or refreshes one digest issue per ISO week. Never gates main, never opens red-main |
All ten shipped callables are on: workflow_call only — none of them declare their own
triggers, and none declare concurrency: (see "How to consume" below for why).
A reusable workflow with no adopter is the same defect as an adopter with no
workflow, so this table is part of the catalog rather than a footnote. "Enforced"
means ci / Required is an actual required status check on the repo's default
branch, read back from the API — not that the caller parses.
| Workflow | First adopter | Enforced |
|---|---|---|
apple.yml |
narduk-enterprises/GeoGridKit |
not yet |
python-data.yml |
narduk-enterprises/narduk-data (earth-data-ci.yml) |
not yet |
reusable-browser-tests.yml |
narduk-enterprises/been-sober-for (first proof PR) |
not yet |
docs-governance.yml |
narduk-enterprises/company-hq |
yes — repo ruleset require-docs-governance |
node-library.yml |
narduk-enterprises/narduk-charts |
yes — repo ruleset require-ci-required |
nuxt-cloudflare.yml |
hydrogen |
no — hydrogen has no branch protection; it called @v1 unenforced for months, which is the failure mode this column exists to make visible |
reusable-node-ci.yml |
retired 2026-09-24 — zero live callers estate-wide, file deleted (narduk-reboot P3-C2 / O-D8) | — |
code-review.yml |
retired 2026-09-19, file deleted 2026-09-24 — no live pool adopters | n/a — historical advisory callable; current review is the CT650 PR review bot (runners#212) |
closing-syntax-check.yml |
retired 2026-09-24 — zero adopters, file deleted (narduk-reboot P3-C2 / O-D8); narduk-enterprises/agent-infrastructure stays on its own local, canonical invocation of the same (now commit-scanning) checker |
— |
reusable-weekly-drift-check.yml |
retired — zero live callers verified across narduk-enterprises and narduk-incubator |
— |
red-main-listener.yml |
narduk-enterprises/workflows itself (red-main-self.yml, listening to this repo's own CI) |
not yet |
flake-digest.yml |
none yet — no adopter here has a quarantine lane of its own; nuxt-cloudflare.yml adopters with e2e-quarantine-args set are the natural first callers |
not yet |
design-ledger.yml |
narduk-enterprises/mybo-at-v2 (planned) |
never — advisory by design (Logan, 2026-09-27: "Red check + drift issue (Recommended)") |
This is the actual point of the repo, not an implementation detail.
docs-governance.yml, node-library.yml, and nuxt-cloudflare.yml each end
with a job named exactly Required. (Since the CI reset, 2026-09-28,
docs-governance.yml's Required is its only job and runs the check itself:
with nothing to aggregate, a second job only cost a runner allocation per
call. lint_callables.py R5 exempts exactly that self-contained shape from
always().) An aggregating Required needs: every other job the
workflow defines, runs with if: always() (or, in nuxt-cloudflare.yml,
if: "!cancelled()", below), and explicitly checks each
needs.<job>.result — a job that's enabled must report success; it may
report skipped only when its controlling input is off. In particular,
run-e2e: true makes the plan and every E2E shard mandatory. A skipped
enabled job is failure, not an acceptable
substitute for a toolchain check that never ran. if: always() jobs succeed by
default if you don't check anything explicitly — these don't skip that check.
nuxt-cloudflare.yml's Required, Fast and escalated Fast use
!cancelled() rather than always() (CI reset, 2026-09-28). Both run after a
failed or skipped need; they differ only when the whole run is cancelled.
always() then still started the job, which queued for a runner only to fail,
and held the caller's concurrency group meanwhile: operator-portal run
36494070028 kept the next main run waiting about 14 minutes. With
!cancelled(), GitHub reports the job cancelled without a runner. That
conclusion never satisfies a required check (only success, neutral and
skipped do), so a cancelled run still cannot pass. In cancelled runs
acre-oracle 36490590502 and operator-portal 36494070028, every job whose if:
was false after the cancel reported cancelled, never skipped.
lint_callables.py R5 accepts exactly !cancelled() as the alternative.
This generalizes a pattern narduk-libs already proved in production: its ci.yml
runs a 12-lane package matrix, then a verify job that needs: package-gates
and fails unless needs.package-gates.result == 'success' — one stable
pass/fail signal regardless of how many matrix lanes ran underneath.
The subtlety worth restating: for a called reusable workflow, GitHub
composes the check-context string as <caller's job id> / <job's name>. So
even with an identical reusable workflow, a caller whose calling job is named
quality: or verify: instead of ci: produces a different context string.
The convention has to cover both ends:
- Every reusable workflow here ends with a job named
Required. - Every caller names its calling job id
ci.
That composes to one identical string on every adopting repo, regardless of app type or what runs underneath:
ci / Required
That string is what a branch-protection rule or org ruleset requires. This is
what reopens untangle/TRACKER.yaml D-PKG-3 Part B: once a publisher repo's
caller conforms, the org ruleset can require ci / Required without a
per-repo verification pass, because the string is enforced by convention, not
discovered by inspection.
Every Node workflow here (node-library.yml, nuxt-cloudflare.yml,
docs-governance.yml) used to hand cache: to
actions/setup-node and let it manage the whole round trip. That is the wrong
shape, and company-hq#269 measured why.
setup-node writes a fresh cache tarball whenever the primary key misses,
scoped to the ref the job ran on. A cache written on refs/pull/N/merge is
restorable only by that same pull request — no other PR and not the default
branch. So the sequence for any lockfile-changing PR was: miss → install →
pay a Post Run actions/setup-node tar for an entry nothing would ever read
→ branch dies → merge to main pays the identical tar a second time.
Measured on real runs before the change:
| Repo / job | restore | install | Post Run setup-node (save) |
job wall |
|---|---|---|---|---|
marketing-web / Build |
1s (miss) | 13s | 19s | 67s |
narduk-charts / package / default |
1s (miss) | 7s | 14s | 52s |
status-apps / browser shards (pool) |
5–7s | 19–20s | 32–54s | — |
earthdata-viewer / chromium shard |
— | 10.6s | 29s avg, 115s max | — |
nvault / Verify |
— | 9.7s | 31.5s avg, 315s max | — |
The debris was visible in the API as well as the clock: narduk-charts
carried 8 cache entries, four of them refs/pull/* duplicates of a
refs/heads/main entry; vtraceroute held a 317MB refs/pull/2/merge copy
of the 317MB entry main already had.
So these workflows now split the two halves explicitly:
actions/cache/restoreon every ref, with arestore-keysprefix so a pull request whose lockfile moved still falls back to the default-branch entry — which is the entry it actually wants.actions/cache/saveonly whengithub.ref_nameequals the caller's default branch, and only when the exact key missed.
github.* resolves against the caller's event inside a reusable workflow,
so ref_name is the caller's branch on a push and N/merge on a pull
request. If github.event.repository is ever absent the comparison is simply
false and the write is skipped — the safe direction.
Two deliberate exceptions:
python-data.ymlkeepssetup-uv'senable-cache: auto.automeans on for GitHub-hosted, off for self-hosted, which is already correct: a persistent linux-ci guest keeps~/.cache/uvbetween jobs, so uploading a tarball of an already-warm cache would be a regression for the only current adopter. Onlysave-cacheis gated, for a future hosted caller.reusable-weekly-drift-check.ymlkeepscache: pnpm. It is a weeklyscheduleonubuntu-latest, so it runs on the default branch almost every time — the write it makes is the one the next run reads.
This is not a caller-visible change: no inputs were added or removed, and a
caller pinned to @v1 picks it up when v1 moves.
docs-governance.yml, node-library.yml, nuxt-cloudflare.yml and
python-data.yml each accept a runner input: a JSON-encoded string,
decoded with fromJSON() at every job's runs-on:. apple.yml takes the
same encoding but splits it into two inputs — lint-runner and
apple-runner — because routing Apple CI per job rather than per repo is the
whole reason that file exists. It accepts three shapes:
runner: '"ubuntu-latest"' # plain string
runner: '["self-hosted","Linux","X64","proxmox","linux-ci"]' # JSON array
runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}' # JSON objectThe object form matches Config/github-runner-fleet.json's runsOn shape
exactly, so a private manifest-routed caller can paste that value verbatim.
The value must be valid JSON — a bare string still needs its own quotes,
which is why the explicit hosted value is the four-character JSON string
"ubuntu-latest", not the bare word.
The runner inputs of docs-governance.yml, node-library.yml,
python-data.yml, red-main-listener.yml and flake-digest.yml, and
apple.yml's lint-runner, default to the empty
string (row 18 Q3, Logan 2026-09-18, "Flip the default (Recommended)"). Every
runs-on: resolves the effective route as:
inputs.runner || github.event.repository.private == true && '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}' || '"ubuntu-latest"'
| Caller | Before | After |
|---|---|---|
| private, passes nothing | GitHub-hosted ubuntu-latest (policy drift, agent-infrastructure docs/standards/CI-RUNNER-POLICY.md §1) |
linux-ci organization group, group and labels (§4) |
public, or an event with no repository payload, passes nothing |
ubuntu-latest |
ubuntu-latest (§3) |
| any caller, explicit value | that value | that value — unchanged, proven by scripts/test_runner_default.py |
The comparison is == true, so an unknown visibility falls to hosted, the safe
direction. Because visibility is read at run time, a caller that goes public
lands on hosted on its next run with no edit. Blacksmith overflow and
CI_LIGHTWEIGHT_RUNNER apply to the effective route exactly as they applied to
inputs.runner before; an effective "ubuntu-latest" is never sent to
Blacksmith. reusable-browser-tests.yml's contract job and its Required
fallback use the same visibility gate with no caller input at all, so the
route contract is still validated on a runner this file chose.
The route literal is the fleet manifest's linux-ci class
(fleet/Config/github-runner-fleet.json, organization group linux-ci). It is
one string repeated at each site; test_runner_default.py fails if any copy
diverges, but nothing here re-reads the manifest, so a manifest label change
must be mirrored here.
The linux-ci group is selected-visibility. A private caller the
manifest does not list in that group gets a job that queues with no eligible
runner when it passes nothing. Such a repo must be added to the manifest group
(the policy fix) or pass '"ubuntu-latest"' explicitly, and a repo holding a
§2 hosted exception (for example package-delivery, exception 5) must pass
'"ubuntu-latest"' explicitly. See the versioning note on this change below.
A PUBLIC CALLER MUST NEVER PASS A SELF-HOSTED LABEL. A fork PR on a public
caller can run attacker-controlled code, so a self-hosted runner value hands
that PR estate infrastructure. A repo can also become public later, and runner is a
free-form string these workflows cannot police. Only private, manifest-routed
callers may pass a self-hosted value — resolve it first with
python3 scripts/github_runner_fleet.py route (or
scripts/onboard-proxmox-runner.sh --route-only) in the agent-infrastructure
checkout, then copy the returned runsOn object verbatim. Every workflow file
in this repo repeats this rule in a loud top-of-file comment; don't rely on
this README alone when adding the next one.
D-CI-CAP-1 (c) (2026-09-07, extends D-BLACKSMITH-2; Operator Portal decision log,
fleet#337) makes Blacksmith a kill-switched overflow for the linux-ci
class's ordinary private CI. Every runs-on: ${{ fromJSON(inputs.runner) }}
site in nuxt-cloudflare.yml, node-library.yml, docs-governance.yml, and
python-data.yml
(16 sites, none of them a deploy-credentialed job — nuxt-cloudflare.yml's
only deploy-shaped job, deploy-dry-run, runs wrangler deploy --dry-run
with zero Cloudflare secrets in scope) resolves as:
runs-on: ${{ fromJSON(vars.BLACKSMITH_RUNNERS_ENABLED == 'true' && inputs.runner != '"ubuntu-latest"' && format('"{0}"', vars.BLACKSMITH_LINUX_LABEL || 'blacksmith-2vcpu-ubuntu-2404') || inputs.runner) }}
In the callables with an empty runner default, each inputs.runner above is
the parenthesised effective route from Default route,
so a private caller that passes nothing is Blacksmith-eligible (it is on the
linux-ci class this overflow serves) and a public one never is.
- Switch: the org Actions variable
BLACKSMITH_RUNNERS_ENABLED(defaultfalse, visibility restricted to private repos). Only the exact stringtrueselects Blacksmith. A repo-level variable of the same name overrides the org default for exactly that repo (GitHub's normal repo-over-org precedence) — the mechanism for canarying or kill-switching one adopter without moving the org default. - Label, not a group id: the vendor's documented
blacksmith-2vcpu-ubuntu-2404string, overridable viavars.BLACKSMITH_LINUX_LABEL. No Blacksmith runner-group id is pinned anywhere — D-BLACKSMITH-4's live proof established that provider-created groups are not a stable routing contract; only the label is. - Public
node-library.ymlcallers are never affected: its routing expressions first requiregithub.event.repository.private == true, so they ignore both Blacksmith andCI_LIGHTWEIGHT_RUNNERfor public repositories. Their caller-supplied hostedrunnervalue remains the route. - Manual, not automatic fallback: GitHub does not move an already-queued
job to another
runs-on:target. If Blacksmith cannot schedule or its free allowance is exhausted, flip the variable back tofalse(or remove the repo-level override) and rerun — same operational shape as every other Blacksmith cohort (D-BLACKSMITH-3). - Cost boundary (D-BLACKSMITH-2/3, unchanged): free allowance only, no payment method, no paid overage; disable at 2,400 equivalent 2-vCPU minutes in a monthly cycle, any non-zero amount due, or any unexpected billing state, whichever comes first.
- Never routes here: production/deploy jobs (all app-owned and bespoke,
outside these six CI-only callables), the Playwright/browser class
(
reusable-browser-tests.yml's browser-runner job is untouched — its own trust boundary per agent-infrastructuredocs/standards/CI-RUNNER-POLICY.md§5), and Apple builds (apple.ymlis untouched — it has its own D-APPLE-CI-1 ladder). - Per-repo ordinary route (
CI_LINUX_RUNNER, CI reset 2026-09-28): innuxt-cloudflare.yml, the ordinary private route ofBuild,Checks,Extra gate,Deploy dry runand bothFastjobs reads a repo-level Actions variableCI_LINUX_RUNNER(a JSONrunsOnvalue). When set on a private caller it wins over the caller'srunnerinput (ci-reset W10, W8 item 8): every sampled nuxt caller passesrunner:with the linux-ci route verbatim, so the variable used to be unreachable exactly where the overflow was needed:(github.event.repository.private == true && vars.CI_LINUX_RUNNER || inputs.runner || github.event.repository.private == true && '<linux-ci route>' || '"ubuntu-latest"'). A controller sets or deletes it to move one burst repository's ordinary CI (for example to a Blacksmith label) without a pin bump. Setting it is a deliberate per-repo routing decision that also overrides an explicitrunner:(a CI-RUNNER-POLICY §2 hosted exception included), so only set it where that repo's policy allows the target. Unset, the route is byte-for-byte the previous one (scripts/test_runner_default.pyevaluates both). It never applies to a public caller and does not touch the lightweight, browser or preview routes. The Blacksmith switch above still applies to the resulting route unchanged. Configuration variables in a called workflow resolve from the CALLER's repository (GitHub docs, "Variables": "For reusable workflows, the variables from the caller workflow's repository are used"), so a repo-level value reaches this callable. - Adding a job is not additive here without checking this section again:
a new job that copies
${{ fromJSON(inputs.runner) }}verbatim does not get Blacksmith overflow automatically — use the expression above, or it silently stays off the overflow route.
The cursor-review.yml callable, this repository's cursor-review-self.yml,
scripts/cursor_review.py, scripts/test_cursor_review.py and
scripts/cursor_review_brief.md were deleted in ci-reset W6 (workflows#166),
and each repository's caller is removed by its own PR. A caller that still
exists pins a full SHA, so it keeps working until its removal merges. The
callable held a runner seat for the whole Cursor wait.
Reviews now come from the PR review bot on CT650:
- Code:
narduk-enterprises/runnersscripts/pr_review_bot/, operator guidedocs/pr-review-bot.md, runners#212. - Triggers: a PR comment whose first line is
/reviewor/review deep, or thereview-now/review-deeplabel, from an organization member with write access. - It keeps this callable's prompt (moved verbatim), models (Composer 2.5
standard, Grok 4.6 xhigh deep,
fast: false) and review format, and posts a real review fromnarduk-durable-agent-sessions[bot]. - Nothing to adopt: no caller file and no repository secret.
Git history holds the deleted callable and its design notes.
Every caller is a thin workflow with one job named ci (this is what
makes the ci / Required convention work — see above) and triggers on
pull_request, push to the default branch, and workflow_dispatch.
Never trigger PR-only — a PR-only trigger means the default branch's tip
carries no CI status after merge, which is exactly what happened to
narduk-skills and (pre-merge) narduk-eslint-config per the CI-5 phase 2
audit.
name: CI
on:
pull_request:
push:
branches: [main] # match the repo's actual default branch
workflow_dispatch:
concurrency:
group: ci-${{ github.repository }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
ci: # <-- must be named `ci`; GitHub composes "ci / Required" from this + the reusable workflow's job name
uses: narduk-enterprises/workflows/.github/workflows/<workflow>.yml@<sha-of-v2> # v2
with:
# ...per-workflow inputs, see below
secrets:
NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}Every caller must set its own workflow-level concurrency, as in the
template above. No workflow in this repo declares one, and none ever should.
This is a hard rule, not a gap waiting to be filled — the structural gate in
.github/workflows/ci.yml fails the build if a callable grows a
concurrency: block (rule R6).
The reason is stronger than "it does not propagate". It is that a group here is evaluated in the caller's context, which GitHub states plainly:
A called workflow uses the name of its caller workflow in
${{ github.workflow }}, so using this context as the value ofjobs.<job_id>.concurrency.groupin both caller and called workflows will cause the caller workflow to be cancelled when the called workflow runs.
So the obvious-looking group — ${{ github.workflow }}-${{ github.ref }},
which is what almost everyone writes — would collide with the caller's own
group and cancel the run that is calling us. Seven repos would start
cancelling their own CI the moment v1 moved, and the symptom (a run that
cancels itself for no visible reason) points at the adopter, not at here.
A callable-level group that was carefully uniquified to avoid the collision
would still buy nothing: cancel-in-progress on the caller's workflow-level
group already supersedes the entire previous run, jobs of this callable
included. A second gate underneath it can only add a way to be wrong.
concurrency is a permitted key on a job that calls a reusable workflow,
so a caller with an unusual need can scope it at jobs.ci.concurrency — but
per the same doc, do not reuse the callable's group value there either.
Do not put paths: or paths-ignore: on a caller's on: trigger. If the
workflow does not run, ci / Required is never reported, and GitHub shows a
required check that never arrives as permanently "Expected" — the pull request
becomes unmergeable and stays that way. That is company-hq#146, and it does not
fail loudly; it just quietly stops being mergeable.
The safe mechanism already exists and needs no change to these workflows:
the opt-in gates are ordinary boolean inputs, so a caller can compute them from
its own diff and pass the result. A gate turned off this way reports skipped,
which every Required job already accepts, and Required itself still runs and
still reports. The context never disappears.
jobs:
changes: # cheap; no checkout of the heavy tree needed
runs-on: ubuntu-latest
timeout-minutes: 5
permissions: { contents: read, pull-requests: read }
outputs:
code: ${{ steps.f.outputs.code }}
steps:
- uses: dorny/paths-filter@<full-sha> # vX.Y.Z
id: f
with:
filters: |
code:
- '!(**/*.md|docs/**)'
ci: # still named `ci`; still composes `ci / Required`
needs: changes
uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2
with:
run-e2e: ${{ needs.changes.outputs.code == 'true' }}
wrangler-dry-run: ${{ needs.changes.outputs.code == 'true' }}Which gates each callable exposes this way:
| Callable | Caller-gatable | Always runs |
|---|---|---|
nuxt-cloudflare.yml |
run-e2e (e2e, e2e-plan), wrangler-dry-run, run-tests |
build, checks (its steps run inside build with checks-in-build) |
apple.yml |
run-swiftlint / linux-checks (lint), run-build, run-tests |
xcode |
python-data.yml |
run-ruff (ruff steps in test), run-tests, run-pyright |
test |
reusable-browser-tests.yml |
run-webkit (webkit) |
validate, chromium |
node-library.yml |
run-lint, run-typecheck, run-tests, run-build |
package |
reusable-weekly-drift-check.yml |
all three jobs | — |
docs-governance.yml |
— | the single job |
The "always runs" column is deliberate and is not a gap to be closed. Those
jobs are the ones the required check actually certifies. Giving them a path
condition means ci / Required can report green on a change that was never
built — a false green, which is strictly worse than the wasted minutes it saves,
and which nobody discovers by looking at a passing pull request. If a caller
wants a docs-only change to cost less, it turns off the opt-in gates above and
still builds.
python-data.yml declares no secrets. apple.yml accepts an optional
DEPENDENCY_SSH_KEY for a private SwiftPM repository. Pass a read-only deploy
key scoped to that dependency; never a release, signing, or account key.
Build/test rewrite HTTPS package URLs to SSH only for the caller's GitHub
organization, using Git's process environment. A pinned GitHub Ed25519 host
key authenticates the server. The private key lives in a mode-0600 temporary
file removed on success or failure; Git config, Keychain and GITHUB_ENV stay
unchanged. Xcode callers should use -scmProvider system so resolution uses
this process configuration. Callers that omit the secret keep their existing
behavior. Do not pass credentials to untrusted code or fork pull requests.
secrets:
DEPENDENCY_SSH_KEY: ${{ secrets.PRIVATE_SWIFTPM_SSH_KEY }}reusable-browser-tests.yml is the one exception: it declares a single
optional NARDUK_PLATFORM_GH_PACKAGES_READ secret (required: false), mirroring
node-library.yml and nuxt-cloudflare.yml's existing CI-alias-for-the-PAT
convention, so that a browser CI caller installing cross-repo
@narduk-enterprises/* packages can pass a real read credential. A caller that
passes nothing keeps today's behavior unchanged: every consuming step falls
back to the ephemeral github.token.
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/apple.yml@<sha-of-v2> # v2
with:
# Repo-scoped Mac from Config/github-runner-fleet.json `appleRepositories`.
# REQUIRED — there is no default, on purpose.
apple-runner: '["self-hosted","macOS","apple-imac"]'
run-swiftlint: true
build-command: swift build
test-command: swift testTwo runner inputs, routed per job, because the estate has one always-on Mac
slot and AGENTS.md requires that only work needing Xcode/macOS occupy it:
apple-runner—xcodebuild,swift build/swift test, anything needing the Apple toolchain. Required, no default.lint-runner— SwiftLint pluslinux-checksfor shell/grep boundary gates and plist checks viapython3plistlib(not PlistBuddy/plutil, which are macOS-only). SwiftLint does not compile the project, but its official Linux binary still dynamically loadslibsourcekitdInProc.so; self-hostedlinux-cirunners provide the pinned Swift SourceKit runtime layer. Empty by default: a private caller gets thelinux-ciorganization-group route and a public one"ubuntu-latest"(see Default route). Passing the route explicitly still works and still wins:
lint-runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
linux-checks: |
python3 scripts/check_plists.py
! grep -rn "import UIKit" Sources/MyCorerun-swiftlint is opt-in and the SwiftLint archive is verified against a
pinned swiftlint-sha256 before it is unpacked. Enabling SwiftLint on a repo
with no .swiftlint.yml runs the full default rule set and is usually red on
first contact — GeoGridKit's first run produced 158 errors — so adopt a
repo-owned config in the same change. Prefer only_rules: over
disabled_rules: there: an allowlist cannot be broken by a future SwiftLint
release adding a rule.
The checksum verifies the downloaded SwiftLint archive; it does not provide the
Swift SourceKit runtime. The install step separately fails closed unless
/usr/lib/libsourcekitdInProc.so is readable, exports that exact path through
LINUX_SOURCEKIT_LIB_PATH, and still locates the unpacked swiftlint binary
with find because upstream archive layouts have changed.
Archive/sign/notarize/TestFlight/Sparkle are not here, for the same reason
node-library.yml omits publish: they need temporary keychains, the host-wide
Apple build lock and per-repo credentials, and folding them in would put
release credentials behind a PR-triggered gate. That stays the
apple-release-pipeline skill's per-repo workflow.
uv (the default), with a lockfile check:
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/python-data.yml@<sha-of-v2> # v2
with:
runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
working-directory: services/my-pipeline
venv-path: .venv-ci
uv-sync-args: "--locked --extra test"
test-command: python -m pytest tests -qpip/venv, for the repos that have not moved to uv:
with:
dependency-manager: pip
pip-install-args: '-e ".[test]"'Notes:
-
The environment's
bin/is prepended to$GITHUB_PATHafter install rather thansource-ing an activate script, because activation dies with the step's shell. That is what makesextra-checkswork from any directory. -
extra-checksis a multi-line shell hook that runs after the tests, in the same environment, and is covered byci / Required. It exists because a caller cannot add steps to a called workflow's job, so a repo-specific gate that needs the installed package would otherwise have to stay behind as a second job duplicating the entire install. -
run-ruffis opt-in andruff-versionis pinned exactly, so a ruff release cannot turn a green repo red. Pointing ruff at a repo that never had a static-analysis gate is usually red on first contact — narduk-data'searth-data-pipelinehas 36 violations today, including sixF821undefined-name — so the template does not decide for the caller when to take that on. Since ci-reset W10 (W8 item 10) ruff runs as the last steps of thetestjob, not in alintjob of its own: that job did about 0.1 min of work for a whole runner allocation. The steps run under!cancelled()after the tests, so a lint finding and a test failure both show, and ruff gets its own interpreter (called by path), so it never installs into the project environment. A caller's ruleset that required<job> / lintwould lose that context; none does (narduk-data'searth-data-ciis the only caller, and narduk-data'srequire-checksruleset names noci / lint). -
run-pyrightis opt-in forv1compatibility, but it is a real static gate: the workflow provisions Node 24 explicitly, installs exactpyright-versionunder$RUNNER_TEMP, and runs it in the installed Python environment. This is intentionally notpy_compile, which proves syntax and nothing about names or types. -
Several isolated pytest invocations (narduk-data's
ci.ymlneeds them, because two suites share a module basename with no__init__.py) go intest-commandas a multi-line string, or inextra-checks. -
extra-envvalues are expanded on the runner. Write$GITHUB_WORKSPACE, never${{ github.workspace }}(workflows#4):# RIGHT — expanded on the runner by the workflow itself extra-env: | PYTHONPATH=$GITHUB_WORKSPACE # WRONG — silently becomes `PYTHONPATH=` and breaks a later step extra-env: | PYTHONPATH=${{ github.workspace }}
with:inputs are evaluated in the caller, and ajobs.<id>.uses:job is never assigned a runner, sogithub.workspacethere is the empty string. Until this was fixed the workflow acceptedPYTHONPATH=as a well-formedKEY=VALUEand the failure surfaced four steps later asModuleNotFoundError: No module named 'pipelines', with nothing anywhere namingextra-env.$NAMEand${NAME}are expanded against the runner's environment. Nothing else is: no$(...), no backticks, no${NAME:-default}, no globbing, noeval. An empty value — or one whose every reference is unset — is now a hard error naming this trap, because a silently-unset variable is the worst outcome.scripts/test_extra_env.pylocks all of that down against the step text extracted from the YAML itself, so the tests cannot drift from the shipped script.
nuxt-cloudflare.lightweight-runner routes E2E plan and Required
independently of the build. Selection is explicit input, then
CI_LIGHTWEIGHT_RUNNER, then the existing build/Blacksmith route. Public
callers always use ubuntu-latest for these jobs. An unset override is
backward compatible; required check names and failure/skip semantics are unchanged.
CI_LIGHTWEIGHT_RUNNER is an organization Actions variable containing a JSON
runs-on value, initially "ubuntu-slim" for the authorized company gates.
Every shared Required job reads it except python-data.yml's, which runs on
the caller's own test route (CI reset 2026-09-28): its only caller runs
those jobs GitHub-hosted, so a lightweight Required on linux-ci was the one
self-hosted job in each run. Changing this one value changes routing
for subsequent jobs without changing callable code or repinning callers.
Package, browser, Apple, and deployment jobs retain their own routes.
node-library.required-runner remains an explicit per-caller override; avoid
setting it on ordinary callers that should follow the central route. Repository
variables take precedence over organization variables, so reserve repository
values for documented exceptions.
Inline planners or final assertions need a one-time adoption of the same expression, preserving their job names and dependency conditions:
runs-on: ${{ fromJSON(vars.CI_LIGHTWEIGHT_RUNNER || '"ubuntu-slim"') }}Use this only for bounded, secrets-free metadata and result checks with no private-network requirement. Their existing policy authorization still applies. It is a routing contract, not automatic classification by job duration. Already queued jobs keep their selected route. A composite action cannot select a runner because it starts after GitHub has assigned one.
Consumers pinned before this feature need one reviewed SHA update. That initial adoption is unavoidable; later capacity changes require only the organization variable. Keep immutable workflow pins. Unsetting the variable restores each callable's prior fallback, and malformed JSON fails visibly. Public callers must continue to use GitHub-hosted runners; never give the variable a self-hosted route in an organization that exposes it to public callers.
required-runner optionally separates the small Required aggregation job
from the package runner. It accepts the same JSON runner shape as runner;
empty uses the organization-level CI_LIGHTWEIGHT_RUNNER route for private
callers, falling back to existing routing (including Blacksmith) when that
variable is absent; public callers retain their supplied hosted route. An
explicit value takes precedence for Required only. The gate checks dependency results
without checking out source, installing packages, or receiving registry secrets.
For a secrets-free gate before a privileged self-hosted release, the caller may
pass required-runner: '"ubuntu-slim"' under company-hq's CI runner policy
exception 2. This keeps completion reporting independent of a saturated build
queue. It does not grant other callers a hosted-runner policy exception.
A monorepo may group packages into a bounded number of install lanes. Use
filter: "", a lane-level extra-scripts: "ci:batch", and set all four
run-* inputs to false. The root ci:batch script receives the complete
lane in PACKAGE_MATRIX_JSON; for example, a lane can carry
packages: ["@example/core", "@example/auth"]. That repository-owned script
must validate every selection, reject missing gates, retain every nonzero
exit and print per-package results. Run the same script locally. The callable
continues to own runner setup, one install per lane and auth cleanup; existing
single-package lanes are unchanged. Choose the batch count from measured
setup cost and capacity, and retain the caller's final integration aggregate.
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
with:
node-version: "22"
package-manager: pnpm
secrets:
NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}install-args and extra-scripts were added for narduk-charts' adoption
(company-hq#172) and are worth knowing about, because between them they are
the difference between retiring a local job and retiring half of one:
with:
package-manager: npm
install-args: "--legacy-peer-deps" # charts cannot `npm ci` without it
extra-scripts: size # runs after build, in the same laneextra-scripts runs each named package script after build, in the same job,
so a gate that needs build output (size-limit) does not require a second job
with a second full install and build. Each name is probed first and fails when
it is absent by default.
With a per-package matrix (generalizes narduk-libs' package-gates):
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
with:
package-matrix: |
[
{
"label": "narduk-core",
"filter": "@narduk-enterprises/narduk-core",
"extra-scripts": "check:dist"
},
{
"label": "narduk-auth",
"filter": "@narduk-enterprises/narduk-auth",
"extra-scripts": ""
}
]
secrets:
NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}Matrix lanes own their extras. The optional lane-level extra-scripts field
overrides the legacy shared input even when it is ""; that explicit empty
value means the lane has no extra gates. Omitting the field preserves the
shared-input behavior for existing callers, while declaring it on every lane
prevents one package's requirement from leaking into another.
The org Actions secret NARDUK_PLATFORM_GH_PACKAGES_READ maps into
GH_PACKAGES_READ on the auth and install steps. The legacy process alias is
also supplied for existing caller bootstrap scripts; it is an interface
compatibility detail, not another secret to create. There is no implicit
github.token fallback. Missing credentials fail before private-package
installs; public-only installs need no credential. A caller whose committed
project .npmrc routes @narduk-enterprises to the anonymous
https://npm.nard.uk mirror (company-hq D-PKG-6) counts as public for that
scope and can drop the secret, unless it also depends on @narduk-geo (not
mirrored) or its lockfile still names npm.pkg.github.com (workflows#106).
The web-foundation check (foundation-check: true) follows the same route
(workflows#109): a mirror-routed caller runs it with no credential at all,
and the pinned dlx download fetches from https://npm.nard.uk. That needs
narduk-app-tools 0.10.0 or later (the foundation-check-tool-version default),
because older releases hard-code GitHub Packages for their live registry
lookup. A caller that has adopted the tool as a dependency must depend on
0.10.0 or later too.
Without a caller bootstrap,
the callable writes a temporary user config containing a literal variable
reference, then removes it after the install.
A public monorepo with only workspace packages under an estate-looking scope can opt out of that automatic name-based detection without forwarding a token:
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
with:
runner: '"ubuntu-latest"'
required-runner: '"ubuntu-latest"'
package-registry-auth: disableddisabled asserts that every installed dependency is public. The default
auto remains fail-closed for private registry dependencies and preserves the
existing package-read-secret contract for private callers.
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2
# A called workflow's jobs may only request permissions the caller granted;
# asking for one it did not kills the whole run at startup (workflows#59).
# `pull-requests: write` is required by the `preview` lane's sticky comment
# -- see "Preview checks" below before pinning a ref that carries it.
permissions:
contents: read
packages: read
pull-requests: write
actions: read # E2E proof lookup (required by the new revision)
with:
node-version: "24"
package-manager: npm # hydrogen's current package manager; pnpm is the default
# Optional: delegate the complete install to a caller-owned wrapper.
# The callable passes the distinct service-token secret only as
# NVAULT_TOKEN and skips its legacy direct package-registry path.
install-script: ci:install
typecheck-worker-script: typecheck
typecheck-web-script: web:typecheck
run-e2e: true
wrangler-dry-run: true
# The estate web quality gates (foundation check, performance budget,
# live security headers), each with a reasoned opt-out. See
# "Quality level" below; new apps start here.
quality-level: standard
secrets:
# Foundation checks and direct registry installs use the canonical PAT.
NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}
# Caller-owned install scripts receive the service token separately.
NVAULT_TOKEN: ${{ secrets.NVAULT_TOKEN }}
# The CI artifact store (see "The CI artifact store (2026-09-26)").
# Optional: without them each E2E job builds its own application and no
# reuse proof is published or honoured.
CI_ARTIFACTS_R2_ACCOUNT_ID: ${{ secrets.CI_ARTIFACTS_R2_ACCOUNT_ID }}
CI_ARTIFACTS_R2_ACCESS_KEY_ID: ${{ secrets.CI_ARTIFACTS_R2_ACCESS_KEY_ID }}
CI_ARTIFACTS_R2_SECRET_ACCESS_KEY: ${{ secrets.CI_ARTIFACTS_R2_SECRET_ACCESS_KEY }}install-script is for repositories whose install wrapper exchanges an nVault
service token for the package credential, materializes any temporary registry
configuration itself, runs the package manager, and removes the configuration on exit.
The value must be one package.json script name using letters, digits, :, _,
or -; the callable rejects missing or unsafe names. Existing callers that
leave it empty keep the legacy install path unchanged.
NARDUK_PLATFORM_GH_PACKAGES_READ always carries the org package-read PAT;
foundation checks use it by default. NVAULT_TOKEN carries the caller's
service token when an explicitly selected install-script needs nVault; that
installer resolves GH_PACKAGES_READ from nVault before starting npm/pnpm.
Generic caller-owned install scripts may omit it and validate their own
credentials. When NVAULT_TOKEN is empty (Dependabot runs see only
Dependabot-store secrets, and NVAULT_TOKEN is Actions-only) and
foundation-check-auth is not nvault, the caller script instead receives
the org package-read secret as GH_PACKAGES_READ, so it can install without
an nVault exchange. It never receives both (workflows#98). Raw PAT consumers omit install-script and pass only
secrets.NARDUK_PLATFORM_GH_PACKAGES_READ.
foundation-check-auth: nvault remains a compatibility mode for callers that
previously mapped their service token into the package-read secret. New callers
use the default package-token mode and map both secrets as shown above.
Local workstations use gh-packages-run, and Workers Builds uses the protected
build secret GH_PACKAGES_READ. Neither uses the Actions input name as a vault
key. The credential route
provides the exact nVault selector and value-free diagnostics.
Deploying with real Cloudflare credentials on push-to-main is not this
workflow's job — that stays a separate nuxt-cloudflare-deploy.yml sibling
(deferred, not built in this pass), matching hydrogen's existing two-job
ci / deploy split rather than folding deploy secrets into the CI gate.
expected-candidate-sha is an optional full, lowercase 40-character commit SHA.
For an explicit validation request, create a fresh branch ref
narduk-validation/<full-sha>/<request-id> pointing at that already-pushed
candidate. The app-owned validation caller listens only to pushes under
narduk-validation/**, retains job ID ci, and passes github.sha as this
input. Ordinary development branches do not trigger it. This is an explicit
full-validation operation, never part of a local development deployment.
Every source checkout is pinned to the requested SHA, then verifies the push
event, ref-encoded candidate, event SHA and actual checkout before running
package code. A mismatched ref or checkout fails the real ci / Required gate.
Leaving the input empty preserves normal callers. The requesting CLI verifies
the selected source branch still points at the requested candidate before
creating the validation ref and records its reason and resulting run.
Do not use workflow_dispatch for release evidence: GitHub excludes those
job checks from pull-request required-check evaluation, even on the correct
head commit. See GitHub's event eligibility documentation.
The dedicated push ref lets the existing real required check count without
changing branch protections or introducing a synthetic success check.
The explicit caller must pass the application's full normal release checks, including its browser coverage. Exact-candidate mode disables browser path skipping and reuse; push semantics select the full shard/argument configuration. Do not reuse an off-main expression that disables E2E, add a success substitute for held workflows, or attach promotion/migration jobs. Required checks and review rules stay intact. Record the exact successful run and candidate; a later merged revision requires its own exact-revision evidence.
| Input | Type | Default | Purpose |
|---|---|---|---|
node-version |
string | "24" |
Node.js version passed straight to actions/setup-node |
node-version-file |
string | "" |
Optional path (relative to working-directory) to a file declaring the Node version, e.g. .node-version or .nvmrc |
Empty (the default) preserves prior behaviour exactly: every setup-node
step in this callable uses node-version. Setting node-version-file
lets a caller single-source its Node version from a file it already
maintains — .node-version, .nvmrc, package.json's engines via a
generated file, etc. — instead of ALSO pinning it as this input's own value,
which is how a caller's Node pin and its CI pin drift apart.
actions/setup-node rejects node-version and node-version-file together,
so every setup-node step here resolves them as two mutually exclusive
expressions rather than passing both: node-version-file is passed through
unchanged, and node-version resolves to an empty string whenever
node-version-file is set (an empty string is setup-node's own "not
provided" sentinel for either input). A caller that sets both gets
node-version-file; node-version is silently ignored in that case, exactly
as if the caller had left it unset.
Every nuxt-cloudflare.yml adopter gets Caller lint as part of Required
(Logan, 2026-09-17 askme round, "Caller lint + job timeouts in
workflows (Recommended)"; company-hq#745). Since the CI reset (2026-09-28) it
is a set of steps inside the Required job, not a Caller lint job of its
own: the separate job cost a whole runner allocation for about ten seconds of
lint on every call. The steps run after Required's gate steps under
!cancelled(), so a lane failure and a lint finding both show in one run, and
a finding fails Required directly. On a protected-path pull request the job
holding Fast (fast-escalated) runs the same anchored scripts, and when
Required is folded into Build (see checks-in-build below) Build runs them
at its own full checkout. Required's own checkout is sparse (ci-reset W10,
W8 item 9): only .github and action.yml/action.yaml at any depth, which is all
the lint reads, so it no longer fetches the whole tree for ten seconds of lint.
It checks out the CALLING
repository (not this one), runs pinned actionlint over the caller's own
.github/workflows/*.yml, and runs a small inline Python audit that fails
the job when a caller workflow:
- has no workflow-level
concurrency:block (skipped for a file whose only trigger isworkflow_call— a callable must NOT declare one; see "Concurrency is the caller's job" above); - has a job that does not
uses:a reusable workflow and has notimeout-minutes(a job that DOESuses:one is reported as an informational::notice::naming the called workflow instead — that job cannot declaretimeout-minutesat all, because the called workflow's own jobs own it, and flagging it as a finding was a false positive this repo used to ship inagent-infrastructure's ownaudit_workflows.py); - has any
uses:step or job not pinned to a full 40-character commit SHA; - is missing
permissions:at the workflow level or on any job.
This is a caller-side hygiene check, distinct from what actionlint alone
proves (schema/expression validity) and distinct from what lint_callables.py
proves about THIS repo's own callables — Caller lint proves the same class
of thing about the repository that adopted one.
| Input | Type | Default | Purpose |
|---|---|---|---|
dependency-audit |
boolean | true |
Fail the build on a high/critical advisory that has a published fix |
audit-ignore |
string | "" |
Comma-separated GHSA-xxxx-xxxx-xxxx=reason suppressions; the reason is required |
The estate security bar (company-hq#745; Logan, askme round 2026-09-17, "Fail on
fixable high/critical (Recommended)"; company-hq D-ORG-1 (g), 2026-09-02) is
alerts on, and no fixable high or critical advisory in the tree. That is
deliberately not what pnpm audit --audit-level=high reports on its own: its
exit status goes non-zero for any high/critical finding, including ones
upstream has published no patch for. A gate that cannot tell "you have not
upgraded" from "there is nothing to upgrade to" is a gate whose only
sustainable reaction is || true, and once that lands the fixable advisories
stop being caught too.
So the audit command's exit status is discarded on purpose and the JSON report is what decides:
| Finding | Result |
|---|---|
| high/critical with a published fix | ::error:: — the build job fails |
| high/critical with no published fix | ::warning:: — the build passes |
| moderate / low / info | counted in the job summary, never blocking |
| report missing, empty, unparseable, or in an unrecognised shape | ::error:: — hard failure |
That last row is the same fail-closed rule require-scripts exists for: "the
audit did not run" must never be indistinguishable from "the audit found
nothing".
A pull request only warns (Logan, 2026-09-29, "Main + nightly"). An
advisory is published on its own schedule, not the pull request's, so a fixable
high/critical finding used to block an unrelated PR the day it appeared. On
pull_request and pull_request_target every ::error:: row above is
re-labelled ::warning:: (the same advisory list, still in the job summary)
and the step passes; the closing line says the finding will fail the default
branch. On a push to the default branch, a schedule run and a
workflow_dispatch the step fails exactly as before, so a red audit is caught
by the merge that follows and by the nightly run (see
E2E off the pull request,
whose nightly caller is not the place for it: the audit runs in the ci
workflow, so add schedule: to that caller too if the audit should be checked
nightly).
Both report shapes are parsed, and the package-manager input selects the
command, never the parser — npm changed this format once already, and a
parser keyed on the input would silently read zero advisories the next time it
changes:
- pnpm / npm 6 — the
advisoriesmap. Fixable meanspatched_versionsis a real range rather than the"<0.0.0"no-patch sentinel. - npm 7+ (
auditReportVersion: 2) — thevulnerabilitiesmap. Fixable meansfixAvailableistrueor a{name, version, isSemVerMajor}object; a fix that needs a major bump still counts as a fix.
Estate contract pins are not advisories and do not count here — only what the package manager's own audit reports does.
audit-ignore: >-
GHSA-aaaa-bbbb-cccc=no upstream release yet, tracked in company-hq#812, review 2026-12-01,
GHSA-dddd-eeee-ffff=unreachable code path behind a disabled flag, review 2026-11-01Each entry is <id>=<reason>, entries separated by commas. The reason is
required: an entry with no =reason fails the gate rather than silently
muting an advisory, which mirrors the written-reason convention a Dependabot
ignore: block carries. Every suppression is echoed as a ::warning:: with its
reason attached, so a muted advisory cannot become invisible tribal knowledge,
and an entry that matches nothing in the current report is reported as a stale
suppression so it gets removed instead of accumulating. Advisories the report
carries no GHSA id for are matched by NPM-<numeric-id>.
The tree it audits is the one build just installed. A standalone lightweight
job on the lightweight runner class would cost a second checkout, a second
setup-node and a second full install — 60–120s and a second runner slot on
this repo's adopters — to re-derive state that already exists in build, for
the ~5–10s the audit command itself takes. It never touches the browser pool.
Required covers it through build, which it already demands success from;
there is no separate result to aggregate.
dependency-audit: false is a bounded remediation, not a setting — the same
status require-scripts: false has. Unlike run-tests and foundation-check,
this input defaults to true: those two run a caller-specific script that
may not exist, while this one reads the lockfile every adopter already has, and
it is a security bar rather than an optional lane. It is still a new gate that
can turn an existing adopter red, so the v1 tag must not move onto it until
the adopters have been checked — see Versioning policy.
Playwright is the slowest lane and it almost never catches a bug that the merge-time run would not (Logan, 2026-09-29: "skip on PRs; run after each merge, newest wins, plus nightly; one org switch"). Two pieces, both additive: with neither in play a caller behaves exactly as before.
1. The org switch: vars.CI_E2E_IN_CI. When the organization (or one repo,
which overrides the org value) sets CI_E2E_IN_CI to false (GitHub compares strings case-insensitively, so False and FALSE switch it off too), E2E plan, E2E
and E2E quarantine are skipped in an ordinary CI run on every event, and
Required reads that skip as success (a lane that runs anyway still fails
it). Build then skips packing and uploading the prebuilt application for the E2E
jobs, and a checks-in-build caller with no other lane folds Required into
Build (one job, no extra queue hop). A pull request that skipped E2E this way mints no Required proof (the proof key does not
include the variable, so a reusable proof would let a later push, after the variable is removed
or overridden, go green without E2E ever having run). Unset, empty or any other value changes
nothing. The switch deliberately does not override two promises to run
browsers: e2e-full-paths (a caller that sets it wants protected-path escalation,
which needs E2E plan to decide, so acre-oracle keeps browsers on auth/payment
changes) and expected-candidate-sha (explicit release validation cannot skip
configured browser coverage).
2. The post-merge and nightly run: mode: e2e. With mode: e2e the
callable runs only Build -> E2E plan -> E2E (and the quarantine lane) and a
Required verdict. Every other lane, the Caller lint and the dependency audit
are skipped, the CI_E2E_IN_CI variable is ignored, and the run never
path-skips (a newer push cancels this run, so no diff can prove a skip).
Required is red if the mode is not ci or e2e, if run-e2e is not true,
if journey-smoke-url is set, or if any E2E lane fails, is cancelled, or was
skipped without the plan's say-so. The mode has its own workflow because an
app's promote.yml fires on workflow_run completion of the whole CI
workflow: E2E left in ci.yml would delay every promotion.
Name the caller workflow E2E (the reaper and the red-main listener look for
that name) and stamp it exactly, with the same with: values as the app's ci
job for the E2E inputs and the same runner inputs:
name: E2E
# Post-merge and nightly Playwright. The newest merge wins: a push cancels the
# run still in flight, and every run tests the whole suite.
on:
push:
branches: [main]
schedule:
- cron: '17 8 * * *' # 08:17 UTC = 3:17 AM CT
workflow_dispatch:
concurrency:
group: e2e-${{ github.repository }}
cancel-in-progress: true
permissions:
contents: read
jobs:
ci:
permissions:
contents: read
packages: read
actions: read
pull-requests: write
uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha> # v2
secrets: inherit
with:
mode: e2e
run-e2e: true
# ...the app's ci.yml E2E and runner inputs, verbatim (e2e-runner,
# e2e-shards, e2e-args, e2e-build-artifact-path, e2e-quarantine-args,
# install-script, node-version, working-directory, ...)Format the stamped file with the app's own Prettier config: apps that run format:check over .github/workflows (cloudflarestat-us, single quotes) fail a double-quoted cron.
The calling job id stays ci so the composed context reads E2E / ci / Required.
concurrency sits in the caller, never in the callable (R6). Pass the same
permissions: block the app's ci job grants: a job in the callable may only
use what its caller granted (R12).
A red post-merge or nightly run reaches the same red-main issue flow as CI:
add the workflow to the app's red-main listener, and accept schedule.
on:
workflow_run:
workflows: ["CI", "E2E"]
types: [completed]
jobs:
red-main:
if: |
(github.event.workflow_run.event == 'push' || github.event.workflow_run.event == 'workflow_dispatch' || github.event.workflow_run.event == 'schedule') &&
github.event.workflow_run.head_branch == github.event.repository.default_branchEach workflow keeps its own "main is red: <name>" issue, opened by the first
red run and closed by the next green one. A run cancelled by a newer merge is
not a verdict and does nothing.
The estate has web quality tools that gated nothing on most apps, because each
one was an opt-in nobody opted into. quality-level: standard turns them on as
one set. It adds no job, no check name and no permission:
| Gate | What runs | Where | Blocks on |
|---|---|---|---|
foundation-check |
narduk-app foundation:check (the existing foundation-check steps) |
Build |
FAIL or UNKNOWN |
performance-budget |
narduk-app performance-budget --json [performance-budget-args] over .output/public |
Build, after build-script and extra-scripts |
any violation, a missing build output, an unreadable report |
security-headers |
narduk-app foundation:check:security-headers --base-url <origin> [--path …] |
Preview on pull requests (the PR's own Cloudflare preview); Journey smoke in post-deploy mode (security-headers-url, else journey-smoke-url) |
exit 1 (FAIL) or 2 (UNKNOWN) |
To adopt, an app adds one line to its ci: job's with: block, in its own
next change:
with:
quality-level: standardlegacy (the default) runs exactly the gates this workflow ran before the
input existed. An app never gets these gates from a pin bump alone.
Why an input, not a default flip. Callers pin a full SHA (25 of 31
nuxt-cloudflare.yml uses: lines across 18 repositories on 2026-09-27; the
other 6 are frozen @v1; none use @main). A pin looks like a natural
one-app-at-a-time gate, but Dependabot opens pin bumps for this callable on most
of those repositories, so a default flip on v2, or in a new major, would
arrive as a wave of simultaneously red pull requests. The explicit input is the
only switch that moves exactly one app, in a change that app chose. A later
major can flip the default once the fleet has adopted.
Opting out. Each standard gate turns off only through quality-opt-out,
with a written reason:
quality-opt-out: >-
performance-budget=hero video ships in the next release (tracked in the app's issue 42),
security-headers=no Workers Builds preview until the account cutoverEntries are check=reason, separated by commas, so a reason cannot contain a
comma. An entry with no reason, an unknown check name or a duplicate fails
Build. Every opt-out is echoed as a ::warning:: with its reason and a
job-summary line on every run, so an opted-out gate stays visible. An opt-out at
legacy warns that it does nothing. foundation-check: true together with a
foundation-check opt-out is a contradiction and fails.
Rules the resolver enforces up front (the Resolve quality gates step,
before any install, so misconfiguration fails the adoption pull request
itself):
security-headersneeds a pull-request preview. Withpreview-checks: nonethe app either enables a preview check or opts out with its reason.performance-budgetneeds a build. Withbuild-script: ""the app opts out.performance-budget-argscannot carry--jsonor--report-only: the workflow owns the report, and the gate never becomes a warning.security-headers-pathsentries start with/;security-headers-urlis onehttps://URL.
Where each gate reads from. performance-budget and both header probes run
the caller's own installed @narduk-enterprises/narduk-app-tools
(pnpm exec / npx --no-install), so neither needs a credential beyond the
install the job already did. The budget runs from working-directory; point it
at a workspace app with performance-budget-args: --app-dir apps/web. The
probes run from preview-working-directory (else working-directory). The
header probe needs a narduk-app-tools release that has
foundation:check:security-headers (narduk-libs#360, 2026-09-16). The tool's own
budget passes silently when .output/public is absent, so the step checks for
the directory itself.
The pull request probes its preview, never production. A change that fixes
the headers must be able to go green before it ships. Headers a Cloudflare zone
adds on the custom domain, rather than the app itself, are absent on the
workers.dev preview; the estate narduk-core security preset sets them in
the app. The post-deploy probe in Journey smoke reads production as just
deployed, runs whether or not the smoke passed, and never triggers the rollback
hook.
Not here: the axe accessibility ratchet. narduk-testkit's
expectAccessible is an assertion inside the app's own Playwright suite, which
run-e2e already runs. This workflow cannot tell a suite that calls it from one
that does not, short of grepping test sources, which would prove presence and
not coverage. The app generator and each app's AGENTS.md own that rule.
with:
run-tests: true # opt-in — see below for why the default is false
test-script: test # vitest, jest, whatever the repo already runs
extra-scripts: check:vendorUntil run-tests existed this workflow had no unit-test expression at all,
while node-library.yml had one. A Nuxt app with a vitest suite therefore
could not adopt its own class's callable without dropping its unit tests — the
same defect the browser-shard gap was, and the same consequence: the repo keeps
hand-rolling. earthdata-viewer (company-hq#278) is the adopter that surfaced
it; its unit job ran npm test and had nowhere to go.
run-tests defaults to false on purpose, even though the script runs
--if-present. test is a near-universal package.json script, so a default
of true would hand every existing adopter a brand-new gate the instant the
moving v1 tag advanced — and a suite that was never in a repo's CI turning
its default branch red is not a backward-compatible change, whatever the input
is labelled. Existing callers opt in when they mean to.
extra-scripts mirrors node-library.yml's input of the same name and runs
after build-script in the same lane, so a gate needing build output doesn't
pay for a second install. earthdata-viewer's vendored-package pin check is the
first case.
Everything in build delays the prebuilt E2E application, and so every E2E
shard, the preview and the deploy dry run. The typecheck and unit-test lanes
therefore run in their own parallel checks job, and e2e-plan no longer
waits for build. A caller whose runner pool, not its critical path, is the
constraint passes checks-in-build: true. The same three steps (YAML
aliases, not copies) then run in build before build-script, checks is
skipped, and Required demands exactly that. This saves one runner
allocation, checkout and install per run, and the E2E shards, preview and
deploy dry run start after the checks instead of beside them. An extra script that does not need build output belongs
in extra-gate-scripts, which also runs in parallel. On riverstatus, two
migration-proof scripts in extra-scripts held its E2E back by about 8 minutes.
Required folded into Build (ci-reset W10, W8 item 3). With
checks-in-build: true and no other lane in the run (no extra-gate-scripts,
fast-scripts, run-e2e, preview checks, wrangler-dry-run,
journey-smoke-url or required-reuse-pr-results), Build is the only lane,
so a separate Required job only waited a second runner queue hop to read one
result (operator-portal PRs: 1.7 min queued for 0.2 min of work). In exactly
that case the build job is named Required and runs Caller lint itself,
and the aggregating job skips; its check shows the raw name expression, never
Required. The check ci / Required keeps its name, so no ruleset changes,
and exactly one started job carries it in every run. Any of the lanes above
brings back the usual Build plus aggregating Required. lint_callables.py
R5 pins the shape (the same condition on both jobs, checks-in-build in it,
the aggregator's if: exactly !cancelled() && !(<condition>)).
concurrent-scripts (opt-in, ci-reset W10, W8 item 2a). Package scripts
listed here start in the background just before build-script and are awaited
after extra-scripts; each one's output is replayed in its own log group, and
a failed, missing (under require-scripts) or killed script fails Build like
any other gate. Use it for checks that need the installed tree but not the
build output (lint, a vendored-package pin check), to take them off the
critical path without a separate job. Two caveats:
- Memory. The scripts share the build's runner and memory. A Nuxt build
plus
vue-tscor ESLint can exceed a 2-vCPU guest's memory; an OOM kill of a background script fails Build ("exited without a status"), and an OOM kill of the build itself fails it too. Preferextra-gate-scripts(its own job) on small runners. - Build outputs. A concurrent script must not read or write what
build-scriptproduces (.nuxt,.output,dist). A script with apre<name>hook is refused, because aprelint: nuxt preparewould rewrite.nuxtunder the running build.
Every gate in this file used to run --if-present, which means a gate whose
script does not exist matches nothing, exits 0, and reports a green lane
that ran nothing. run-tests: true is a caller asserting tests exist; the
callable was taking that assertion on faith.
node-library.yml got the first fix in workflows#14. This file did not, and
the missing web:typecheck default became the dominant dead lane across its
adopters. typecheck-web-script now defaults to "": a second web typecheck
is opt-in, while a non-empty name remains an assertion that the script exists.
Every gate now probes for the script first, and reacts by require-scripts:
require-scripts |
script missing | effect |
|---|---|---|
false (explicit remediation opt-out) |
named but absent | ::warning:: + a job-summary line naming the script; lane still green |
true (default) |
named but absent | ::error:: and the job fails |
| either | empty script name | lane skipped silently — the caller declared it absent |
The default is now true. The migration first declared absent lanes and proved
every remaining name; callers retain false only as an explicit, temporary
remediation opt-out.
A caller declares absent lanes by passing an empty script name. This is what
makes require-scripts usable when some repos genuinely have no build or split
typecheck lane — "I have no web surface" remains distinct from "I named a
script that isn't there":
with:
typecheck-worker-script: typecheck
typecheck-web-script: "" # no web surface in this repo
build-script: "" # no build step; wrangler dry-run IS the build
run-tests: true
require-scripts: true # now every remaining lane is proven to runsoftware-delivery is the reference caller for that shape.
The gate text is not tested as a copy: scripts/test_script_gates.py extracts
each gate's run: block from the YAML and executes that exact text against real
package.json fixtures (54 cases across both Node callables), including
colon-bearing names like web:typecheck — the probe resolves scripts through
npm pkg get scripts.<name>, a dot-path, and a name that broke that lookup
would report every colon-bearing script as missing.
New browser adopters separate application CI from browser execution. The caller's own build job uploads one same-run build artifact; the standalone browser callable consumes it without rebuilding:
jobs:
build:
# ...checkout, install and build the application, then:
outputs:
build-artifact-id: ${{ steps.upload.outputs.artifact-id }}
steps:
- id: upload
uses: actions/upload-artifact@<sha> # v7
with:
name: e2e-build-${{ github.run_id }}-${{ github.run_attempt }}
path: apps/web/.output
include-hidden-files: true
retention-days: 1
browser:
needs: build
uses: narduk-enterprises/workflows/.github/workflows/reusable-browser-tests.yml@<sha-of-v2> # v2
with:
build-artifact-id: ${{ needs.build.outputs.build-artifact-id }}
linux-runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
browser-runner: '{"group":"playwright-isolated","labels":["self-hosted","Linux","X64","proxmox-playwright-x64"]}'
working-directory: apps/web
build-artifact-path: .output
build-artifact-marker: server/index.mjs
playwright-version: 1.61.1
e2e-script: test:e2e:ci
chromium-args: "--project=web --workers=1"
shards: 3That handoff is still a GitHub artifact, so it counts against the
organization's Actions artifact storage. nuxt-cloudflare.yml can no longer
produce it. Its prebuilt application now goes to the R2 CI artifact store for
its own E2E jobs. Before that, the per-call name suffix added in workflows#146
had already stopped matching this callable's default name. Moving this
handoff to the CI artifact store is workflows#150.
The two route inputs are not suggestions. Resolve both from the fleet manifest
and copy each runsOn object verbatim. A hosted contract job checks the exact
group and ordered labels before any caller-controlled self-hosted route is
scheduled. It also rejects absolute or escaping artifact paths, unsupported
package managers, non-exact Playwright versions, and invalid shard counts.
Required is hosted for the same reason: even a bad route still produces a
visible failing gate instead of scheduling the failure aggregator on that bad
route.
Chromium and opt-in WebKit are the only jobs on
proxmox-playwright-x64. Production-build validation runs on linux-ci.
The browser jobs declare no secrets and receive only GitHub's short-lived
token for checkout, same-run artifact transfer, and optional package reads. The caller must provide workflow-level concurrency because
overlapping runs share fixed runner paths; the callable deliberately declares
none.
The pool's current defects are treated as fail-closed constraints:
- agent-infrastructure#323: the root-owned
/opt/playwright-cibrowser tree is never written and noplaywright installfallback exists. Caller pin, installed packages, image package, manifests, executable ancestry, and a real headless launch must all agree. - agent-infrastructure#267: the workflow does not alter guest firewall or DNS state, add sleeps, or hide egress failure.
- agent-infrastructure#248/#237: workflow code never manipulates leases or
quarantine. An unavailable guest can leave work queued, but it cannot become
a skipped green; every enabled shard result is mandatory in
Required.
When the producer is an app-owned build job, export the upload step's
artifact-id as a job output and pass it as build-artifact-id. This keeps
failed-job reruns tied to the successful producer. Omitting it retains
the existing same-run, same-attempt artifact-name convention.
The older nuxt-cloudflare.yml e2e-* inputs remain supported for existing
callers and for repos without isolated-pool approval. New isolated-pool
adoptions use the standalone callable so the pool contract has one owner.
The Nuxt callable reuses build output by default when run-e2e is enabled.
It requires exactly one of .output/server/index.mjs or
apps/web/.output/server/index.mjs under working-directory. Other layouts
name their output explicitly:
run-e2e: true
e2e-build-artifact-path: apps/web/.outputThe build job publishes that relative path to the CI artifact store (see "The
CI artifact store (2026-09-26)") only after its gates pass. It uploads one gzip
tarball and hands its object key, SHA-256 and size to the E2E jobs as a job
output, with the signature of a presigned GET for that one object. A missing
build output fails Build. Hidden files are included because Nuxt's canonical
output directory is .output; the input is an explicit caller-selected path,
not a repository-wide sweep.
Every E2E shard fetches the tarball through that presigned GET. The shard
checks the SHA-256 and size against Build's output, unpacks it to the same path
and receives E2E_PREBUILT_ARTIFACT=1. The consumer's launcher must treat that
variable as an assertion: validate the expected entry points and fail when any
are absent, rather than silently rebuilding. A tarball that does not match
Build's digest fails the job and is never unpacked.
Any other miss makes the shard run build-script itself to produce the same
path. Misses include a run without the store's credentials (a fork, Dependabot,
a caller that does not pass them), an expired link and a storage outage. The
shard then fails if the script is missing or leaves no output there, so the
suite still tests a built application.
The object key includes the repository, github.run_id, github.run_attempt
and a per-call scope (artifact-scope). The reference travels as a job output,
so rerunning failed shards within the link's 24 hours reuses the original
build even when the attempt number advances. After that they rebuild in-job.
The default auto publishes nothing when E2E is disabled. Set an explicitly
empty path only for suites that do not test a built application. Launchers
must honor E2E_PREBUILT_ARTIFACT=1: the workflow cannot suppress a build
hardcoded inside an application script.
E2E now skips by default when every changed path is documentation or
repository metadata: root Markdown, Markdown under docs/, app/package-root
READMEs, agent guidance files, licenses, issue/PR templates, or CODEOWNERS. Runtime
Markdown under content/, executable files under docs/, app code, CSS,
dependencies, configuration, workflows, tests and unknown paths still run.
The same rule applies to pull requests and forward pushes, including the merge
push. Manual, scheduled and merge-group runs always run E2E. A skipped plan
emits no shards; Required checks that E2E and its report actually skipped.
e2e-skip-paths replaces the default with repository-root-relative globs
(* does not cross directories; ** does). An explicitly empty string disables
path skipping. Narrow or disable the list if your app renders those documents;
never add CSS or executable documentation to the list. e2e-full-paths takes
precedence over a matching skip pattern.
The plan reads GitHub's compare endpoint with contents: read, including both
names of renamed files. Empty or failed comparisons, new/deleted branch SHAs,
force-push divergence, unsupported filenames and the API's 300-file cap all
run E2E. See GitHub's compare API.
Default-branch pushes first look for a successful PR run of the same Git tree, caller workflow path and callable inputs. This handles squash and merge commits whose SHA differs but whose tested contents are identical. Other CI checks and production delivery still run normally.
Required publishes a proof to the CI artifact store only after every required
gate passed and the actual E2E arguments and shard count matched the full
suite. A PR subset, docs-only skip, failed shard or fork cannot publish proof.
The proof is a small pointer, proof/<owner>/<repo>/<key>.json. It names the
PR run and attempt that wrote it. The key digests the tested tree, the caller
workflow path and the callable inputs.
The push does not trust the pointer on its own. It reuses the proof only after GitHub's API confirms three things:
- the run is a completed, successful
pull_requestrun of the same caller workflow; - the run's head repository is this repository;
- the named attempt has a successful
Requiredjob whosePublish full E2E proof <key>step, carrying this exact key, succeeded.
An absent, expired (30-day lifecycle), changed or unreadable proof, or one
GitHub does not confirm, runs E2E, as does a run without the store's
credentials. Scheduled and manual runs always run. Set
e2e-reuse-pr-results: false for tests with intentionally different PR/push
behavior or external state that must be rechecked after merge.
required-reuse-pr-results reuses the whole gate the same way, from the
Publish Required proof <key> step.
Permission migration: adopting this revision requires actions: read on
all Nuxt callable ci jobs, even if E2E is disabled. GitHub validates the
permission ceiling before evaluating job conditions. The proof lookup needs
only read access to Actions results to confirm the run a proof names; it gains
no write permission. The proof itself lives in the CI artifact store, so pass
the three CI_ARTIFACTS_R2_* secrets as well (see the example above). Add this
alongside the existing contents: read, packages: read, and
pull-requests: write grants before or with the SHA bump. Existing pinned
callers do not change. This is a breaking permission change: do not advance
v2 over it; publish a new major only after an adopter canary is green.
permissions:
contents: read
packages: read
pull-requests: write
actions: read # read the prior PR's E2E proof
with:
run-e2e: true
# Defaults: conservative path skipping and full-PR proof reuse.(workflows#83) e2e-skip-paths above is all-or-nothing: it either runs the
whole matrix or none of it. These two inputs are the middle setting — a
caller-defined subset on pull requests, the full suite on the default
branch.
run-e2e: true
e2e-shards: 3 # push / default branch: unchanged
e2e-args: ""
e2e-pr-shards: 1 # pull request: one lane
e2e-pr-args: "--project=smoke" # pull request: the caller's own projectWhy this exists: the browser class now has three active 8 GiB on-prem primary
guests (343, 345 and 346), three 4 GiB fsn1-pve01 fallback guests (340–342),
and CT 344 is a configured dormant guest rather than an active slot. The old
three-effective-slots/seven-declared queue measurements are historical; heavy
jobs request memory-8g, while not every browser guest is 8 GiB. Capacity and
tiering belong to narduk-enterprises/fleet, not to this callable. See the
canonical host inventory for names and placement.
The primary/fallback flags do not establish GitHub scheduling priority: normal
browser jobs can land on all six active guests; memory-8g matches the three
active on-prem guests.
Fewer lanes is not by itself faster — it is fewer slots. Measured on gonogo
with e2e-pr-shards: 1 and no e2e-pr-args: the same suite ran 527 / 407 /
292 s in three lanes plus a 32 s report (a 559 s critical path), and 913 s in
one. Aggregate pool occupancy dropped 27% (1258 runner-seconds over four jobs
to 913 over one) and the run held one isolated slot instead of three, which
is the part that shortens every other repository's queue — but the pull
request's own wall clock got longer, because the tests were redistributed
rather than reduced. e2e-pr-shards alone is a courtesy to the pool. To make
your own pull request faster, pair it with an e2e-pr-args subset that runs
genuinely fewer tests.
The rules, all of which fail toward running more:
- Unset is today's behaviour.
e2e-pr-shards: 0ande2e-pr-args: ""are the defaults and mean "no override". Every current adopter passes neither, so they see no change whatsoever. - Pull-request events only.
push,schedule,workflow_dispatchand anything unrecognised rune2e-shards/e2e-argseven when an override is configured — the same confinemente2e-skip-pathshas, and for the same reason: the default branch is the canonical validation, and neither input may weaken it. e2e-pr-argsreplaces, it does not append. A pull-request subset cannot silently inherit a conflicting--projectfrome2e-args.- Empty means inherit. A caller that wants arguments on a push and none
on a pull request cannot express that here; put a
github.event_nameexpression in the caller's ownwith:block instead. A sentinel for "explicitly empty" would be a second, weaker way to say the same thing. - The subset is the caller's to define. This workflow never guesses what
"smoke" means — it passes your arguments to your
e2e-script. Tag the subset in your ownplaywright.config(a project) or with--grep.
Required is unaffected as a gate: it still demands E2E succeed, and it
checks the same single resolved shard count E2E sharded on.
A subset is a smaller gate, not a weaker one. Whatever a pull request
stops running, the default-branch push still runs — but it runs it after the
merge. Choose the subset so a failure it cannot catch is one you are willing
to find on main.
e2e-full-paths is the escape hatch from the pull-request subset: a
space-separated list of globs (the same syntax as e2e-skip-paths). When any
file a pull request changes matches one, that pull request runs the full
e2e-shards/e2e-args instead of e2e-pr-shards/e2e-pr-args.
with:
e2e-pr-shards: 1
e2e-pr-args: "--project=smoke"
e2e-full-paths: "server/database/** drizzle/** playwright.config.ts"The use is a fast pull-request smoke by default, with the whole suite reserved for the paths whose breakage the smoke tier cannot see — schema, auth, routing, the Playwright config itself. The rules:
- Unset is today's behaviour. An empty
e2e-full-paths(the default) never forces the full suite. - It beats
e2e-skip-paths. A file matching both lists runs the full suite; it is never skipped. - Unknown means full, for an opted-in caller. When the changed files
cannot be listed (no base/head SHA, no
gh, a compare API error, an empty list, or the 300-file compare cap), a caller that sete2e-full-pathsgets the full suite. A caller that did not keeps its pull-request tier. - Full-tier selection is for pull requests. Pushes use the full tier when neither path skipping nor equivalent PR proof applies.
- The
E2E planjob summary lists which changed files forced the full run. e2e-full-pathsmatches case-insensitively (ASCII only).*auth*matchesuseAuth.ts,AuthPanel.vueandOAuthCallback.ts;*session*matchesuseSession.ts.*still does not cross/. An emptye2e-full-pathsstill never escalates.e2e-skip-pathsstays case-sensitive: each list folds case only in the direction that adds proof, so the default skip list skips exactly what it did before and an oddly casedREADME.MDruns E2E.
The check name stays Fast in both cases, so the required context
ci / Fast does not change. What differs is the sibling check in the same
workflow run:
| Run | Check named Fast |
Other Fast checks |
|---|---|---|
| Plain | the fast job (lint and unit scripts); on a pull request of a caller with e2e-full-paths (or an exact candidate), fast-escalable |
none (fast-escalated is skipped) |
| Escalated | the fast-escalated job (the full Required gate) |
Fast lanes (escalated) (fast-escalable: the lint and unit scripts) |
fast and fast-escalable run one step list (a YAML alias) on disjoint
domains (ci-reset W10, W8 item 4). Only a pull request of a caller that set
e2e-full-paths, or an exact candidate, can escalate, so only there does the
lane wait for E2E plan; everywhere else fast needs only Reuse plan and
starts with the run (riverstatus main 36497919511: Fast queued 12.0 min behind
a 0.1 min plan). On a push, plain fast no longer fails closed on a failed
E2E plan (its full output never applies to Fast off a pull request);
Required still demands the plan succeed.
So a readiness check that needs to know whether a ci / Fast run over 180
seconds was the full gate looks for a ci / Fast lanes (escalated) check run
in the same check suite. Reading check-run names needs no extra permission,
so callers grant nothing new for this. For people reading the run, the job
holding Fast also writes a ### Fast escalation step-summary section with
the line escalated: true or escalated: false; that step never fails the
job. A Fast job that does not apply is skipped and takes no runner: both jobs
when fast-scripts is empty, in reused runs and in journey-smoke mode, and
fast-escalated on every run that does not escalate (CI reset, 2026-09-28;
until then both started as no-op jobs, which cost two runner allocations per
call). GitHub never evaluates a skipped job's name, so such a check shows the
raw name expression (for example
(inputs.fast-scripts != '' && ...) && 'Fast' || 'Fast (escalation not needed)'),
never Fast, and cannot satisfy a required ci / Fast check. A readiness
check should match ci / Fast lanes (escalated) exactly, not as a substring:
a skipped fast job's raw expression contains that text. Required demands
each Fast job succeed when it applies and be skipped when it does not. The
raw name cannot be made readable: a skipped job's name is its literal template
whatever contexts it reads (inputs, github and vars included, probed
2026-09-29), and a static Fast would be a passing ci / Fast.
The Fast lane (fast and fast-escalable) runs any eslint CLI process its
fast-scripts start with --cache --cache-strategy content, so unchanged files
are not linted again. No caller change is needed: a preload
(NODE_OPTIONS=--require, set for the Run fast scripts step only) appends the
flags to the process whose entry script is eslint's own bin/eslint.js, however
it was reached (pnpm --filter web run lint, turbo run lint, a wrapper
script), and leaves everything else alone: other programs, eslint --print-config, --version, and an eslint that already passes --cache*
flags. (Appending -- --cache to the script does not work: the scripts are
wrappers, and turbo run lint --cache would be turbo's own flag.) narduk-lint
(narduk-libs' lint-budget wrapper) is not covered: it takes no
--cache-strategy.
The cache is ../.ci-cache/eslint/<lockfile hash> beside the checkout. A
persistent self-hosted runner keeps it between runs. GitHub-hosted runners
restore it everywhere and save it only on the default branch (the dependency
cache rule). A changed lockfile starts cold.
Only Fast gets it. Build, Checks, Extra gate and the escalated Fast
lint cold: ESLint's cache is keyed on a file's own content and config, so a type
change in another file does not re-lint a cached file, and a type-aware rule
(@typescript-eslint/no-floating-promises) can pass an unchanged file that a
changed file just broke. The cold full gate is the net for that. Set
eslint-cache: false to run Fast lint cold too. scripts/test_eslint_cache.py
executes the shipped step and preload.
A flaky test is taken out of the gate by tagging it (for example
@quarantine) and excluding that tag from the caller's gating Playwright
projects. It still needs somewhere to run, or it can never show it is fixed.
e2e-quarantine-args is that place:
with:
e2e-quarantine-args: "--project=quarantine --retries=0"When set, an extra E2E (quarantine) job runs e2e-script with exactly
these arguments, on every event E2E runs on:
- It cannot fail the gate. The job is
continue-on-error: true,Requireddoes not list it, and no jobneeds:it. A red quarantine run shows as a warning and a job-summary line. It never turnsci / Requiredor the caller's run red.lint_callables.pyenforces this through itsNON_GATING_JOBSexemption, which is the only job allowed outside R5. - It is the gate's setup. It runs
E2E's own steps through a YAML alias: the same prebuilt application, runner route, toolchain checks and auth cleanup. Only the arguments differ. It is unsharded. - Its evidence is separate. Its report is in its own job log, apart from the gate's shards.
- It skips with E2E. A docs-only PR skipped by
e2e-skip-pathsruns neither.
The job's history on the default branch is the "N consecutive green runs"
record a test needs to leave quarantine. Pass --retries=0 so a retry can't
hide a flake. Empty (the default) adds no job.
| Input | Type | Default | Purpose |
|---|---|---|---|
preview-checks |
string | og |
none, og, e2e-subset, or og,e2e-subset |
preview-url-source |
string | pr-comment |
pr-comment or url-template |
preview-url-template |
string | "" |
URL template for url-template; {branch}, {branch-alias}, {sha} |
preview-timeout-minutes |
number | 20 |
Whole-job budget; the bounded wait is this minus five minutes |
preview-working-directory |
string | "" |
Directory the preview checks run from; empty uses working-directory |
This is a BREAKING interface change. The
previewjob requestspull-requests: write, and a called workflow asking for a permission its caller did not grant kills the caller's entire run at startup —startup_failure, zero jobs, no logs, no annotation (workflows#59, which took down every@v1adopter on 2026-09-04). Every adopter must addpull-requests: writeto itsci:job before thev1tag moves onto this commit, and an adopter whose repository has no Workers Builds preview must also passpreview-checks: nonein the same change:jobs: ci: uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2 permissions: contents: read packages: read pull-requests: write # sticky preview comment
Workers Builds' GitHub App creates check runs, commit statuses and a
pull request comment, and "a preview URL will be provided for any builds
which perform wrangler versions upload"
(Cloudflare: Workers Builds GitHub integration).
It does not create a GitHub Deployment — that is the Pages integration's
shape — so there is no environment_url to read, and this workflow does not
offer a source that pretends there is. Adding one would also have cost every
adopter a deployments: read grant to carry dead code.
That leaves two honest sources:
pr-comment(default) — read the pull request's comments and take the first*.workers.devURL posted by an author whose login containscloudflare. Aworkers.devlink posted by anyone else is ignored.url-template— derive the URL locally, with no GitHub read at all. Cloudflare's aliased preview URLs are<ALIAS>-<WORKER_NAME>.<SUBDOMAIN>.workers.dev, so a typical template ishttps://{branch-alias}-myworker.myaccount.workers.dev.{branch-alias}is the head ref lowercased with every character outside[a-z0-9]replaced by-. That reproduces the transform this estate has observed (buoys' branchcodex/buoys-complete→codex-buoys-complete-buoys.narduk-enterprises.workers.dev); Cloudflare does not publish the truncation rule it applies to long branch names, so a repository with long branch names should stay onpr-comment, which reads the URL Cloudflare actually minted rather than predicting it.
A preview link in a comment says a build was attempted. The wait is not
satisfied until the URL itself answers and its x-build-version header is a
prefix of the pull request's head SHA. That header is the estate's existing
live-proof convention (buoys docs/workers-builds.md records
x-build-version: a84fa2163903 for commit a84fa216390321…), and it is
compared by prefix because Cloudflare emits a 12-character short SHA while
git rev-parse --short defaults to 7. A fixed-width comparison would be a gate
that never passes, and a gate that never passes is a gate somebody deletes.
Each round of the bounded wait is one comment read and one HTTP read — no sleep
loop that can outlive the job, and the wait is preview-timeout-minutes minus
five so the checks and the comment still have room after it.
A preview that never becomes ready inside the bound is a FAILURE, not a
skip, and so is one that answers 4xx/5xx, carries no x-build-version, or
serves a different commit. "The preview never showed up" is the single most
common way a Workers Builds connection silently breaks, and it is
indistinguishable from "this repository does not use previews" only if you
refuse to make the caller say which it is. That is what preview-checks: none
is for.
The lane runs on pull_request / pull_request_target events only. A push to
the default branch has no pull-request preview to check, and Required expects
the job skipped there — a preview job that runs on a push is a failure
too.
-
og—narduk-app og:check --live --base-url <preview>, run from the caller's own installed@narduk-enterprises/narduk-app-tools. Unlikefoundation-check, there is no pinneddlxfallback here:preview-checks: ogis a caller asserting it has the tool, and a second resolution path would mean a second credential path and a second version to keep in step. -
e2e-subset— runse2e-scriptwithe2e-pr-args(the same pull-request argument sete2e-pr-shards/e2e-pr-argsintroduced in workflows#83) against the preview withPLAYWRIGHT_BASE_URLset. An emptye2e-pr-argsis a hard failure rather than a silent full-suite run against a shared preview.The caller's Playwright config must honour
PLAYWRIGHT_BASE_URLand must not start its ownwebServer, or the suite will test localhost and report green. buoys cannot use this yet for exactly that reason — itsplaywright.config.tshas only dev-server and prebuilt-Worker modes and "no fixture accepts an override base URL" (docs/e2e-testing.md).
Lightweight (the same route as Required and E2E plan) unless
e2e-subset is selected, which needs a browser guest and therefore the
e2e-runner route. On that route the lane runs the same YAML nodes as the
e2e job's isolated-route guard, image-equality assertion and browser
installer — shared by anchor, not copied, because an image-equality gate that
exists twice is an image-equality gate that drifts. lint_callables.py R11 now
holds every job routed to e2e-runner to that preflight, not just e2e.
Note that with e2e-subset selected the bounded wait holds a
playwright-isolated slot (three effective slots across roughly ten repos —
see "A smaller suite on pull requests"). That is why preview-timeout-minutes
is a caller input and why og is the default.
The lane posts a single comment carrying the URL, the build version, the head
SHA and each check's result, marked with <!-- narduk-ci:preview -->. Later
runs find it by marker and edit it; they never append. The same table also
goes to the job summary, so the URL survives even when the comment API call
does not — and a failure to post warns rather than turning an otherwise passing
lane red.
playwright-isolated is download-free and image-owned. The current pool pins
Playwright 1.61.1: Chromium headless shell revision 1228 (Chromium
149.0.7827.55) and WebKit revision 2311 (WebKit 26.5). The image keeps
the package, browsers.json, and browser payloads root-owned/read-only under
/opt/playwright-ci; the ephemeral guest exports a job-visible symlink tree
through PLAYWRIGHT_BROWSERS_PATH.
Before a shard starts its suite, the callable now fails unless all of these are true:
- the caller directly pins
@playwright/testto an exact version (no caret, tilde, tag, or range); - that pin, the installed
@playwright/test, installedplaywright-core, and image Playwright version are identical; - the caller and image
browsers.jsonSHA-256 values match, and each requested engine's revision/upstream version matches; - the selected executable exists, resolves into the immutable image browser tree, has root-owned non-writable ancestry, and passes a real headless launch canary;
PLAYWRIGHT_BROWSERS_PATHis the absolute/optpath exported by the guest, never a workspace or$RUNNER_TEMPcache.
The isolated-route guard runs before dependency installation. It rejects
e2e-install-browsers: true and every e2e-browsers-path override, while the
installer step independently excludes both the playwright-isolated group and
proxmox-playwright-x64 label. A mismatch names consumer/image versions,
manifest digests, browser revisions, and the upgrade choice, then exits
non-zero without invoking an installer. Hosted browser jobs may still opt into
the explicit installer because they own their toolchain rather than consuming
this pool.
Path-filtering is the caller's job — this workflow doesn't know the caller's repo layout:
name: handbook-spine-check
on:
pull_request:
paths: [docs/**, scripts/check-handbook-spine.py]
push:
branches: [main]
paths: [docs/**, scripts/check-handbook-spine.py]
workflow_dispatch:
concurrency:
group: handbook-spine-check-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
ci:
uses: narduk-enterprises/workflows/.github/workflows/docs-governance.yml@<sha-of-v2> # v2
with:
command: python3 scripts/check-handbook-spine.py
fetch-depth: 1command can be multi-line to install its own light dependencies first — for
example, a PyYAML-consuming script (like company-hq's
untangle-project-sync.yml, which is not itself a docs-governance.yml
caller today, but shares its shape):
with:
command: |
python3 -m pip install --disable-pip-version-check --no-cache-dir "pyyaml==6.*"
python3 untangle/sync-to-project.py --project-number 1 --dry-runA product keeps one design ledger beside each Claude Design canvas
(design/<canvas>/ledger.json, agent-infrastructure#1804). It maps each canvas
screen to the code it specifies and records when the two last matched. This
callable reads drift in both directions: not-built (the canvas moved),
design-stale (the code moved) and diverged (both moved).
mode: checkrunsdc_ledger.py status <ledger> --check --github-summary. It exits 1 while any entry is flagged, prints one fix line per entry, and appends the table to the job summary.mode: flagrunsdc_ledger.py flag <ledger>with the caller'sgithub.token. It opens, edits, reopens or closes oneDesign drift: <canvas>issue on the ledger'sissue.repo, whose labels must exist.issue.repomust be the calling repository: the token can write no other, so the checker errors when it differs fromGITHUB_REPOSITORY, and whenissue.repois missing.flagnever runs on a pull request. It runs only onpush,scheduleorworkflow_dispatch(v2.1.1). On any other event theflagjob is skipped and thecheckjob fails the run with an error, so a pull request's code never writes issues.- Commit the boards. The callable never runs the ledger's
buildcommand, so the canvas boards the ledger'sprojectpoints at must be in git. A gitignored, generatedprojectreads every entry as "board or screen missing", which stays red once anything is built.statusprints a hint whenprojectis absent from the checkout. - A malformed ledger is an error (exit 2, naming the entry):
codemust be a list of paths,gate.stateexactlyopenorcleared,builtnull or all three hashes. Ledger text is treated as untrusted: it cannot start a workflow command in the log, break a summary or issue table, or mention anyone. - Advisory. There is no
Requiredjob, so never add<job> / checkto branch protection, and do not name the calling jobci. - Grant
contents: readandissues: writein both modes. GitHub checks the skippedflagjob's permissions too, so a check-only caller that grants onlycontents: readends instartup_failure(proven on workflows#155). Thecheckjob itself runs withcontents: read. - The checker comes from this repo at
job.workflow_sha, the commit the caller pinned, via a sparse checkout without credentials. A pull request cannot change the code that judges it, and the ledger'sbuildcommand is never run. scripts/dc_ledger.pyis a byte copy. The canonical file is agent-infrastructure'sskills/claude-design-ops/scripts/dc_ledger.py, and itsscripts/check-dc-ledger-parityfails when the two differ. To change it: open the PR here with the new copy; land the canonical change in agent-infrastructure first, its pin at this PR's head commit (the Cursor review here reads agent-infrastructure main as the canonical); merge this PR and tag it; then move agent-infrastructure's pin to the merge commit.
name: design-ledger
on:
pull_request:
paths: [design/<canvas>/**, apps/web/app/**] # the ledger's mapped paths
push:
branches: [main]
permissions:
contents: read
jobs:
design-ledger:
uses: narduk-enterprises/workflows/.github/workflows/design-ledger.yml@<sha-of-v2.1.1> # v2.1.1
permissions:
contents: read
issues: write
with:
ledger: design/<canvas>/ledger.json
mode: ${{ github.event_name == 'push' && 'flag' || 'check' }}runs-on follows the Default route: leave it
empty and a private caller lands on linux-ci.
reusable-weekly-drift-check.yml was removed after a fresh organization-wide
caller search found no live consumer (workflows#20, Actions-optimization
audit). narduk-template-smoke-app is disabled and documentation references
were not runtime callers.
reusable-node-ci.yml was removed in narduk-reboot P3-C2 / O-D8: a fresh
organization-wide code search across narduk-enterprises,
narduk-incubator, and narduk-enterprises-clients (2026-09-24, gh search code / gh api search/code, restricted to path:.github/workflows) found
zero live callers — the same result the 2026-07-27 search recorded when
node-library.yml shipped as its richer, preferred replacement (workflows#14,
workflows#16). node-library.yml is a strict superset for library/cli
surfaces: the Required aggregator job and the optional package-matrix
input (see "The ci / Required convention" above) plus a runner input that
accepts the JSON array/object form, not only a plain string. New callers use
node-library.yml; there is no migration to perform because there was no
live caller to migrate.
code-review.yml asked the estate's now-retired ephemeral agent pool for one
read-only review of a pull request head, via a repository_dispatch at
agent-infrastructure; it was never a CI gate and had no Required job. The
pool was retired 2026-09-19 with an empty allowlist, and adoption moved to
cursor-review.yml (Refs narduk-enterprises/agent-infrastructure#333,
D-AGENT-POOL-1), itself retired 2026-09-28 for the CT650 PR review bot (see
"Cursor review (retired 2026-09-28)"). A 2026-09-24 re-verify (gh search code / gh api search/code across narduk-enterprises, narduk-incubator and
narduk-enterprises-clients) found zero live pool adopters, matching the state
this section already described, so the file was deleted in narduk-reboot
P3-C2 / O-D8 rather than kept as a dead compatibility surface.
- Callers pin full commit SHAs, never
@main.@mainis how the last outage happened; it is not a supported reference. Caller lint (inRequired) and narduk-app-tools foundation item 5.1 both reject a bare@v2tag, so a caller pins the SHA the tag points to and names the tag in a comment:uses: …/<workflow>.yml@<40-char sha> # v2. Existing@v1pins still pass whilev1is frozen. - Maintainers cut tags.
vNmajor tags (v1,v2, …) move forward only for backward-compatible changes within that major; breaking changes (renamed or newly-required inputs, removed jobs, changed secret names) get a new major. - To adopt a fix, callers bump their pinned tag/SHA deliberately — nothing
changes under them silently unless they chose a moving
vNmajor tag. v1is a moving major tag and is advanced by hand after a merge, never as a side effect of merging. The CI-5 phase 2 files shipped ahead of the tag andv1was advanced to cover them afterwards;apple.yml,python-data.ymlandnode-library.yml's two new optional inputs are backward-compatible additions to the same major, sov1moves again rather than av2being cut. An adopter merged before the tag moves must pin the exact commit SHA and switch to@v1once the tag covers it.v1is frozen; adoptv2when touched. See the section below.- A fan-out canary precedes an advance: before moving the tag, trigger or find at least one adopter run on the new commit and confirm it resolves and succeeds, per Logan, 2026-09-04: "Fan-out canary before the tag advances (Recommended)" (Refs company-hq#536).
- The empty-
runnerdefault flip (row 18 Q3) must not reachv1until its unlisted private callers are fixed. It changes no input name, job, check context or permission, so it stays withinv1, but it moves every private caller that passes nothing onto theselected-visibilitylinux-cigroup. At the flip, the@v1callers passing nothing werecoding-standards(docs-governance.yml) andx-event-recap(node-library.yml) — both outside that group — andpackage-deliveryandsoftware-deliveryonnuxt-cloudflare.yml(whose flip ships separately). Beforev1advances over this change: add each such repo to the fleet manifest'slinux-cigroup, or have it pass'"ubuntu-latest"'explicitly (mandatory for a §2 hosted exception such aspackage-delivery); then run the fan-out canary. SHA-pinned callers pick the change up only when they bump their pin. - Adding a workflow, or adding an optional input with a default, is within-major. Renaming or newly requiring an input, removing a job, renaming a job (which renames the composed check context and silently orphans every branch-protection rule that required it), or changing a secret name is a new major.
- A new job-level
permissions:scope is a BREAKING change, not an addition. A job in a called workflow may only request permissions the caller granted on itsuses:job; ask for one it did not and GitHub fails the caller's entire run at startup —startup_failure,jobs: [], no logs, no check-run annotation — before any job is created. Nothing in the adopter's repository changed, so it reads as an infrastructure outage rather than an interface break, and it lands on every@v1adopter simultaneously the moment the tag moves. On 2026-09-04pull-requests: readon nuxt-cloudflare.yml'sE2E planjob did exactly that to harvest-tracker, marketing-web, vtraceroute and hydrogen, while the two repos pinned to older commit SHAs kept running (workflows#59). Each callable's permitted set is declared inscripts/lint_callables.py'sCALLER_GRANTSand enforced as R12; widening one means updating every adopter'spermissions:block first, then the map, and only then moving the tag. nuxt-cloudflare.yml's browser-shard inputs (e2e-runner,e2e-shards,e2e-args,e2e-install-browsers,e2e-browsers-path) and its two new jobs are within-major on the same rule: five optional inputs whose defaults reproduce the previous behaviour, plus added jobs. Adding a job is not a breaking change here specifically because the composed context comes from the caller's job id and this workflow'sRequiredjob — neither of which moved.v1moved again rather than av2being cut.nuxt-cloudflare.yml's pull-request subset inputs (e2e-pr-shards,e2e-pr-args, workflows#83) are within-major for the same reason: both are optional, both default to an UNSET sentinel (0/""), and with them unset every event resolves the shard count and argument list exactly as before. No job, check name, orRequiredexpectation changed — the effective shard count simply moved frominputs.e2e-shardsto anE2E planoutput that equals it whenever no override is supplied.e2e-full-pathsis within-major on the same rule: optional, default"", and with it unset every event plans exactly as before.e2e-quarantine-argsis too: optional, default"", and when unset its job is skipped. The job it adds is neverneeds:-ed and never gates, so no check contextRequiredreads changes.reusable-browser-tests.ymlis a new callable, so its required route, artifact, and exact-version inputs do not break an existing caller.python-data.yml'srun-pyright, exactpyright-version, andpyright-argsinputs are optional additions; existing Python callers keep their prior behavior until they opt into static analysis.reusable-browser-tests.yml's later addition of the optionalNARDUK_PLATFORM_GH_PACKAGES_READsecret (required: false) is within-major on the same rule as a new optional input: a caller that passes nothing gets byte-identical behavior to before the secret existed (workflows#50).nuxt-cloudflare.yml'srun-tests/test-script/extra-scriptsare within-major on the same rule — three optional inputs, no new job, one conditional step each.run-testsdefaults tofalseprecisely so that it is within-major:testis a near-universal package script, so defaulting it on would have made a movingv1tag introduce a gate to callers who never asked for one, which is a breaking change dressed as an additive input. Getting the default wrong is how an "additive" change breaks people.nuxt-cloudflare.yml'smodeinput andvars.CI_E2E_IN_CIswitch (E2E off the pull request) are within-major on the same rule: one optional input defaulting toci, no new job, no new permission, and an unset variable reproduces every adopter's behaviour exactly. TheDependency auditstep's pull-request downgrade (warn onpull_request, block on push, schedule and dispatch) only loosens a gate on one event, so no adopter turns red because of it.nuxt-cloudflare.yml'sfoundation-check/foundation-check-tool-version(agent-infrastructure docs/standards/WEB-FOUNDATION-CHECK.md, D-WEBFOUND-2 Q5/Q9 (a), D-WEBFOUND-3) are within-major on the same rule — two optional inputs, no new job, three added steps inside the existingbuildjob,Required'sneeds:graph unchanged.foundation-checkdefaults tofalsefor the same reasonrun-testsdoes: the web-foundation program is a multi-wave fleet migration (D-WEBFOUND-2 Q4/Q10), most fleet apps do not conform to the seven-item contract yet, and a movingv1tag must not hand every existing adopter a brand-new red gate the day the tag advances.nuxt-cloudflare.yml'squality-level,quality-opt-out,performance-budget-args,security-headers-pathsandsecurity-headers-urlare within-major on the same rule: five optional inputs, no new job, no new permission,Required'sneeds:graph unchanged.quality-leveldefaults tolegacy, which resolves every new gate off, so movingv2over it changes nothing for an app that did not writequality-level: standard. The gates are default-on only insidestandard. See Quality level for why this is an input rather than a default flip (Dependabot pin bumps would otherwise turn the fleet red at once).foundation-check-authdefaults topackage-token.nvaultremains only for callers whose legacyNARDUK_PLATFORM_GH_PACKAGES_READmapping carries an nVault service token; new callers use the canonical package PAT there and map the service token separately asNVAULT_TOKENforinstall-script. The callable resolves the existing package-read grant for that legacy one step and suppliesNODE_AUTH_TOKENto both the installed checker and the optional pinned download. The checker needs this credential for its live N-1 registry lookup even when dependencies are already installed, unless the caller's project.npmrcroutes@narduk-enterprisestonpm.nard.uk(workflows#109): then no credential is resolved or exported. Neither the service token nor the resolved package token is exported to later steps; temporary npm configuration contains only an environment-variable reference. Missing or rejected credentials leave a blocking UNKNOWN artifact rather than falling back togithub.token.- The D-CI-CAP-1 (c) Blacksmith-overflow change (see "Blacksmith overflow"
above) is within-major on the same rule: no new input, no new job, no new
job-level
permissions:, and the default (vars.BLACKSMITH_RUNNERS_ENABLEDundeclared) reproduces every existing adopter's behavior byte-for-byte.v1moves again rather than av2being cut, behind the usual fan-out canary.
Logan, 2026-09-18 (askme, 13:41 CT): "Moving v2 tag; repos move when touched (Recommended)".
-
v1stays at f8e3cc6 and does not move again. Everything after it is on thev2line.v2.0.0(a29cd16, the pull-request preview gate) is a breaking change fornuxt-cloudflare.yml: itspreviewjob requestspull-requests: write, and a caller that does not grant it hitsstartup_failureon its whole run (workflows#59). Movingv1over it would have broken every@v1nuxt-cloudflare caller at once. On 2026-09-18 none of the seven (hydrogen, marketing-web, my-farm, narduk-nvr, package-delivery, software-delivery, vtraceroute) granted it. -
v2is the moving major tag, advanced by hand behind the fan-out canary, the same wayv1was. -
A repo moves to
v2in the next PR that works on it, not in a sweep. Make these changes in that same PR:- every
uses: …/<workflow>.yml@v1becomes the full SHAv2points to, with a# v2comment (git ls-remote https://github.com/narduk-enterprises/workflows refs/tags/v2;3cc8c736d07aa821daeb417e7f6682ecd3935aa8on 2026-09-18). A bare@v2fails caller-lint ("is not pinned to a full 40-character commit SHA") and foundation item 5.1; - a
nuxt-cloudflare.ymlcaller addspull-requests: writeto itsci:job'spermissions:, and passespreview-checks: noneif the repo has no Workers Builds preview (see the preview gate); - a private caller that passes no
runnernow lands on thelinux-ciorganization group (Default route). Confirm the repo is in that group. If it is not, it queues forever rather than failing. Otherwise passrunner:explicitly (for example'"ubuntu-latest"'for a named CI-RUNNER-POLICY §2 hosted exception); nuxt-cloudflare.yml'sdependency-auditdefaults totrueonv2and fails on fixable high/critical advisories. Fix them, or passdependency-audit: falsewith its written reason.
The PR's own CI run on the
v2SHA is the proof it moved cleanly. - every
.github/workflows/ci.ymlgates this repo (~7s).actionlint+scripts/lint_callables.py+ behavior tests includingscripts/test_playwright_toolchain.py. The structural gate enforces every convention in this list, so none of them can regress silently: see the rule table (R1–R11) at the top ofscripts/lint_callables.py. Run it locally before pushing:python3 scripts/lint_callables.py.- Third-party and first-party actions are pinned to full commit SHAs with a
version comment, targeting the current Actions Node runtime (enforced: R4).
All nine pins currently resolve to their claimed tags and every one runs on
node24.astral-sh/setup-uvis deliberately held atv8.3.2rather thanv9.0.0: v9's sole breaking change flipsprune-cachetofalse, which would grow cache usage for every adopter, and v8.3.2 is already on the current runtime — so the bump would be a behaviour change with no hardening benefit. - Jobs declare minimal
permissionsand explicittimeout-minutes(enforced: R1/R2/R3).
Every job has had a finite timeout-minutes since the day it was written — the
guard is not new. What was missing was a basis. Measured 2026-07-25 from the
jobs endpoint, counting only jobs the callable actually composed (a
ci / … or drift-check / … name), execution time only with queue excluded,
never extrapolated from a run count.
That filter is the whole measurement, not a detail. A first pass that matched on
the bare job name mixed each adopter's pre-adoption local job into the same
bucket and reported python-data / test at p95 444s / max 722s. Split
correctly, the callable's ci / test is p95 90s / max 91s and the 722s belongs
to the local test job narduk-data ran before it adopted — a 5× error, in the
direction that would have made a fine timeout look nearly breached.
| Callable | Job | Timeout | p50 | p95 | max | n | repos | ×p95 |
|---|---|---|---|---|---|---|---|---|
apple.yml |
xcode |
45 (xcode-timeout-minutes) |
116s | 228s | 231s | 8 | 4 | 11.8× |
apple.yml |
lint |
15 (lint-timeout-minutes) |
8s | 12s | 12s | 4 | 2 | 75.9× |
docs-governance.yml |
check |
15 | 10s | 18s | 22s | 39 | 1 | 49.7× |
node-library.yml |
package / <label> |
30 | 78s | 289s | 298s | 17 | 4 | 6.2× |
nuxt-cloudflare.yml |
Build |
30 | 72s | 131s | 145s | 16 | 4 | 13.8× |
nuxt-cloudflare.yml |
Deploy dry run |
15 | 18s | 20s | 20s | 4 | 1 | 45.7× |
python-data.yml |
test |
30 (test-timeout-minutes) |
75s | 90s | 91s | 5 | 1 | 20.0× |
reusable-weekly-drift-check.yml |
Typecheck |
15 | 69s | 69s | 69s | 1 | 1 | 13.0× |
reusable-weekly-drift-check.yml |
Unit Tests |
15 | 50s | 50s | 50s | 1 | 1 | 18.0× |
reusable-weekly-drift-check.yml |
Template Drift Check |
10 | 45s | 45s | 45s | 1 | 1 | 13.3× |
| (all five) | Required |
5 | 4–5s | 5–7s | 7s | 83 | 11 | 43–65× |
Jobs with no data at all — their timeouts are declared, finite, and unmeasured. Do not read the numbers above onto them:
| Callable | Job | Timeout | Why nothing ran |
|---|---|---|---|
nuxt-cloudflare.yml |
E2E, E2E plan |
30 / 5 | never executed on any adopter. hydrogen and software-delivery set run-e2e: false; marketing-web and vtraceroute leave it at the default. 35 skipped instances, 0 runs |
python-data.yml |
lint |
10 | narduk-data leaves run-ruff false — skipped in every run |
Nothing was changed as a result. Every measured timeout sits between 6.2×
and 76× its observed p95, so none is close to producing a false red. The one job
near the 4–6× target band is node-library / package (6.2×), which is correct
as-is. The rest are looser than a band would suggest, deliberately:
- The sample is tiny and 24 hours old. Composed jobs first appear
2026-07-24T21:54Z and most adopters landed the next day. Only
docs-governance / check(n=39) is a distribution; below n≈10 a percentile is one observation wearing a hat. - For a short job the floor is not p95. It is "long enough that a cold cache or a slow guest is not a false red", which is minutes regardless of a 12s p95.
apple / xcodeat 11.8× is the loosest that matters, because it holds the estate's single Mac slot. It stays: n=8 across four small Swift repos is far too thin to justify tightening a real iOS archive toward ~20 minutes, and it is already a caller-tunable input. An adopter that knows its build should set it.
The regression guard is rule R2 in scripts/lint_callables.py, which makes it
impossible to add a job without a finite timeout — including via an input whose
numeric default was removed.
timeout-minutes measures execution, never queue — and here that gap is
enormous. docs-governance / check executes in 10s and has waited 1973s
(33 min) for a linux-ci runner; its Required job has waited 1044s. Anyone
sizing a timeout from a run's wall-clock duration would set it wildly wrong.
When adding a job, size its timeout from the same place — the jobs endpoint, per job, never extrapolated from a run count.
- This is a public repository. Its own gate and public callers use GitHub-hosted runners; private callers retain reusable-workflow compatibility and may use manifest-routed self-hosted runners only where policy permits.
- New reusable workflows follow both estate-wide conventions added by CI-5
phase 2: the workflow's last job is named exactly
Requiredandneeds:everything else (see above), andruns-on:decodes the effective route (inputs.runner, else the visibility-gated default) from a JSON-encodedrunnerinput defaulting to''(see "Default route" above; add the new callable toscripts/test_runner_default.py). - Estate conventions live in the
ci-workflow-authorskill (agent-infrastructure repo); consult it before adding workflows here.
nuxt-cloudflare.yml neither uploads nor downloads a GitHub artifact, and
reusable-browser-tests.yml uploads none. On 2026-09-26 the org's Actions
artifact storage passed its included allowance with a $0 budget, GitHub
refused every artifact upload estate-wide, and CI went red. The owner first
dropped the evidence uploads, then chose to move the handoffs CI actually
needs to a Cloudflare R2 bucket the estate owns.
Dropped (evidence only; nothing gated on it):
-
nuxt-cloudflare.ymlno longer has:- the Playwright evidence upload;
- the
E2E reportmerge job; - the journey-smoke evidence upload;
- the foundation-check artifact.
Evaluate web-foundation conformance checknow printsfoundation-check.jsonin its own log. -
reusable-browser-tests.ymlno longer has:- the Chromium and WebKit blob-report uploads;
- the
reportmerge job, with its HTML report artifact and itsreport-artifactoutput.
report-timeout-minutesstays as a deprecated, unused input so callers that pass it keep validating. Moving these to R2 would have put a bucket write credential on the isolated Playwright pool for evidence nobody gates on, so they were dropped, not moved. -
Every shard reports with the caller's own Playwright reporter in its job log. There is no
--reporter=blob, which only existed to be merged.
Moved to R2:
- the prebuilt E2E application Build hands to every E2E job;
- the
Requiredproof (required-reuse-pr-results); - the full E2E proof (
e2e-reuse-pr-results).
Not moved yet: reusable-browser-tests.yml still takes its build as a
same-run GitHub artifact the caller uploads (workflows#150).
| Bucket | narduk-ci-artifacts, in the narduk-enterprises Cloudflare account (location WNAM) |
| Prebuilt application | prebuilt/<owner>/<repo>/<run>-<attempt>-<scope>.tar.gz, expires after 7 days |
| Reuse proof | proof/<owner>/<repo>/<proof key>.json, expires after 30 days |
| Credential probe | probe/<owner>/<repo>/<name>, expires after 1 day |
| Incomplete multipart upload | aborted after 7 days |
Every key is scoped by github.repository and shape-checked before use.
The workflow uses three org Actions secrets with selected visibility: only
the repositories on each secret's repository list receive them. All three are
declared required: false on the callable:
| Secret | Value |
|---|---|
CI_ARTIFACTS_R2_ACCOUNT_ID |
the Cloudflare account ID |
CI_ARTIFACTS_R2_ACCESS_KEY_ID |
the R2 token's ID |
CI_ARTIFACTS_R2_SECRET_ACCESS_KEY |
the SHA-256 of the token value |
The token has Object Read and Write on this one bucket and nothing else. It
expires on 2027-09-26. Its source is nvault
cloudflare/prd/narduk-enterprises-ci-artifacts (CI_ARTIFACTS_R2_TOKEN,
CI_ARTIFACTS_R2_ACCOUNT_ID), persona
cloudflare-narduk-enterprises-ci-artifacts.
A reusable workflow sees only the secrets its caller passes, so a caller adds
the three lines shown in the main example above, or uses secrets: inherit.
A caller that repins onto a commit with the store must also be added to the repository list of all three secrets. Until it is, GitHub hands it empty values, and its runs fall back as described under "Without credentials, and on failure". An org owner adds a repository without touching any value:
repo_id=$(gh api repos/narduk-enterprises/<repo> --jq .id)
for name in CI_ARTIFACTS_R2_ACCOUNT_ID CI_ARTIFACTS_R2_ACCESS_KEY_ID CI_ARTIFACTS_R2_SECRET_ACCESS_KEY; do
gh api -X PUT "orgs/narduk-enterprises/actions/secrets/$name/repositories/$repo_id"
done
# Read the list back:
gh api orgs/narduk-enterprises/actions/secrets/CI_ARTIFACTS_R2_SECRET_ACCESS_KEY/repositories \
--jq '.repositories[].full_name'Who holds the secret access key:
- Build, to publish the prebuilt application;
Required, on a pull request, to publish proofs;Reuse planandE2E plan, which read proofs on a default-branch push.
The E2E jobs, including those on the isolated browser pool, get only the two IDs. They read the prebuilt application through a presigned GET that Build signed for that one object, valid for 24 hours.
That signature is not a secret, and it appears in the logs. It travels as a job output, and GitHub drops any job output that carries a secret, so each E2E job prints it in the fetch step's environment. GitHub masks the account ID and the access key ID, but they are identifiers, not keys. So anyone who can read that run's logs can fetch that one tarball until the link expires. That is the same audience, for the same day, that could download the GitHub artifact it replaces.
One token serves the whole estate, so nothing is trusted just because it is in the bucket.
Prebuilt application. E2E unpacks the tarball only when its SHA-256 and size equal what its own run's Build reported through job outputs. A mismatch fails the job before anything is unpacked.
Proof. A proof object is only a pointer to a run. It is honoured only when all of the following hold:
- GitHub's API reports that run as a completed, successful
pull_requestrun of the same caller workflow; - the run's repository and head repository are both this repository;
- the run attempt is the one that wrote the pointer;
- that attempt's
Requiredjob succeeded; - that job has a successful step named
Publish Required proof <key>orPublish full E2E proof <key>, with this exact key.
The key is unchanged from the artifact era: the digest of the tested Git tree, the caller workflow and every callable input.
The store can cost time but never coverage, and it never passes a gate falsely. A run has no store credentials when it is a fork pull request, a Dependabot run (Dependabot sees only Dependabot secrets), a public caller, a caller that does not pass the secrets, or a caller missing from their repository list. Then:
- Build publishes nothing, with a
::warning::naming the three secrets (a notice until ci-reset W10, W8 item 1, and nobody read it). - Each E2E job warns that it rebuilds the application, then runs
build-scriptitself and fails if it cannot produce the output. - Proofs. No proof is published or honoured, so the default-branch push runs the full gate.
A storage error, an expired link or a stalled transfer does the same, with a warning. A tarball transfer gives up after 3 attempts of 120 s. That is about 6 minutes in the worst case, well inside the 30-minute job limit, so Build still finishes and each E2E job still has time to build its own application.
Only two things turn a job red:
- a prebuilt application that fails its digest check;
- an in-job fallback build that fails.
scripts/test_ci_artifact_store.py runs the client exactly as the workflow
installs it. It checks:
- AWS's published Signature V4 examples;
- a round trip and tampering against a local S3 double that authenticates every request with an independent signer;
- the fallbacks;
- proof publish and lookup;
- the rule that no GitHub artifact action or artifact API appears in
nuxt-cloudflare.yml.
test_e2e_build_artifact.py, test_e2e_reuse.py and test_fast_path.py
cover the wiring and the proof lookups job by job.