Skip to content

Latest commit

 

History

129 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

workflows

Shared reusable GitHub Actions workflows for the narduk-enterprises estate (CI-5).

This repository is public. Its own CI and public callers run on GitHub-hosted capacity, with no org-variable or Blacksmith routing. Private callers retain reusable-workflow compatibility and may use manifest-routed self-hosted capacity only where their repository policy permits it. The workflows hold no secrets — every workflow_call secret is optional and skips cleanly, and apple.yml and python-data.yml declare none at all. reusable-browser-tests.yml declares one optional secret (NARDUK_PLATFORM_GH_PACKAGES_READ) that falls back to the ephemeral github.token when a caller passes nothing.

Callers must pin the full 40-character commit SHA that v2 points to, with a # v2 comment (new work; @v1 is frozen, see v1 is frozen; adopt v2 when touched) — never a bare @v2 tag and never @main — and a public (or possibly-future-public) caller must never pass a self-hosted runner label. That warning is still correct, but it rests on the caller, not on this repo: runner is a free-form caller-supplied value, a repo can go public later, and a fork PR on a public caller can run attacker-controlled code on estate infrastructure. A reusable job whose runner input is left empty resolves it per run from the caller's own visibility: a private caller gets the fleet manifest's linux-ci organization-group route, a public (or visibility-unknown) caller gets GitHub-hosted ubuntu-latest — see Default route. apple.yml's Mac route is the one input with no default, deliberately.

Fix CI in one place, not 100. Application repos call these workflows via workflow_call instead of blob-copying YAML. This repo replaces the broken pattern where ~20 repos carried copies of weekly-drift-check.yml pointing at narduk-enterprises/narduk-nuxt-template/.github/workflows/reusable-quality.yml@main — a repo name that no longer exists (renamed to narduk-template), so every scheduled run 404'd silently.

CI-5 phase 2 (company-hq strategy/workflow-consistency-proposal.md) adds the three app-shaped CI gates below, each producing the identical ci / Required check context so branch protection and org rulesets can require the same string across every repo of that type. See §2 "Proposed standard" in that proposal for the full rationale.

apple.yml and python-data.yml were the two named gaps left after that pass, and company-hq#173 (R-14) is the issue that closed them: "adopt the shared workflow" had been the standing answer for CI duplication, and for the Apple and Python-data families there was nothing to adopt — so their duplication was structural, not neglectful, and naming a workflow that did not exist made the gap look like an adoption backlog. Both now exist and both have a real adopter (see "Adopters" below). Still not built: a nuxt-cloudflare-deploy.yml sibling for push-to-main deploys with real Cloudflare secrets.

The standalone browser / Playwright gap named by company-hq#197 is now reusable-browser-tests.yml. It consumes a same-run production build and fans Chromium (and opt-in WebKit) into three shards on the manifest-routed isolated pool. Each shard reports in its own job log; no report is uploaded or merged (see "The CI artifact store (2026-09-26)" below). The older browser inputs in nuxt-cloudflare.yml remain for backward compatibility; adopters that separate browser CI use the dedicated callable instead of re-embedding the pool contract in their app-class workflow.

Catalog

Workflow Purpose
apple.yml CI gate for Apple repos (Swift packages, iOS/macOS apps): SwiftLint (official Linux binary) and boundary/plist checks on a Linux runner, swift build / swift test / xcodebuild on the repo-scoped Mac. Two separately-routed runner inputs — that split is the point. Release/signing stays per-repo
python-data.yml CI gate for Python / data-pipeline repos: explicit uv or Python provisioning, pytest, opt-in exact-version Ruff and Pyright (real static checking, not py_compile), plus an extra-checks hook
reusable-browser-tests.yml Private-repo browser CI: validates the exact manifest browser-group object before any shard is scheduled, consumes a same-run production build, asserts the immutable Playwright package/browser image and real launch, runs three Chromium shards plus opt-in WebKit, each reporting in its own job log (no report upload or merge)
docs-governance.yml Thin generic gate for docs/handbook-shaped repos: checkout, optionally provision Python/Node, run one repo-provided check command. Generalizes company-hq's handbook-spine-check.yml / untangle-project-sync.yml shape
node-library.yml CI gate for library / cli project-lifecycle surfaces: script-probed lint/typecheck/test/build, with an optional per-package matrix generalizing narduk-libs' package-gates + verify pattern
nuxt-cloudflare.yml CI gate for nuxt-web / cloudflare-worker surfaces: typecheck (worker + Nuxt split, matching hydrogen), optional unit tests, build, optional extra-scripts, optional web-foundation conformance check, optional Playwright e2e — optionally sharded onto a separately-routed browser pool, the prebuilt application reaching each shard through the R2 CI artifact store — optional wrangler deploy --dry-run validation. CI only — no deploy job (see below)
reusable-node-ci.yml Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8): zero live callers estate-wide, re-verified via gh search code / gh api search/code across narduk-enterprises, narduk-incubator and narduk-enterprises-clients. node-library.yml was always the richer, preferred surface — see workflows#14/#16
code-review.yml Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8, extending the 2026-09-19 advisory retirement below): zero live pool adopters, file deleted. Current review is the CT650 PR review bot (runners#212)
cursor-review.yml Retired 2026-09-28 (ci-reset W6): replaced by the CT650 PR review bot in narduk-enterprises/runners scripts/pr_review_bot/ (runners#212), which answers /review comments and the review-now / review-deep labels without Actions runner time; file, cursor-review-self.yml, scripts/cursor_review.py, its tests and brief deleted. The prompt moved verbatim to the bot
closing-syntax-check.yml Retired 2026-09-24 (narduk-reboot P3-C2 / O-D8): zero adopters, file deleted. narduk-enterprises/agent-infrastructure keeps its own local, canonical invocation of the same (commit-scanning) checker (agent-infrastructure#837, #1085)
reusable-weekly-drift-check.yml Retired 2026-07-26: no live caller; see workflows#20 and the 2026-07-26 Actions-optimization audit
red-main-listener.yml New 2026-09-24 (narduk-reboot P3-C2 / O-D8). Opens or refreshes ONE "main is red: <workflow-name>" issue, labelled red-main, per repo per listened workflow, the first time that workflow goes red on the default branch; closes it on the next green run. Prerequisite: the adopting repo must already have the red-main label (gh label create red-main --color B60205 --description "Main default-branch CI is failing") — opening the first issue passes labels[]=red-main and GitHub 422s a repo without it, so an adopter that skips this step fails closed on its first real red run (PR #142 review, second round). Optional fail-open portal-mirror POST hook. No automatic revert or rollback — that stays in the app's own deploy tooling
design-ledger.yml New 2026-09-27 (agent-infrastructure#1804). Advisory, never required. Runs the design ledger checker (scripts/dc_ledger.py, a byte copy of agent-infrastructure's canonical file) against a product's design/<canvas>/ledger.json. mode: check goes red while a canvas screen and its code have drifted and writes the table to the job summary; mode: flag keeps one Design drift: <canvas> issue. See Design ledger
flake-digest.yml New 2026-09-24 (narduk-reboot P3-C2). Weekly, advisory: scans a caller's completed runs of one workflow, counts failures in jobs matching quarantine-job-pattern (nuxt-cloudflare.yml's e2e-quarantine by default), and files or refreshes one digest issue per ISO week. Never gates main, never opens red-main

All ten shipped callables are on: workflow_call only — none of them declare their own triggers, and none declare concurrency: (see "How to consume" below for why).

Adopters

A reusable workflow with no adopter is the same defect as an adopter with no workflow, so this table is part of the catalog rather than a footnote. "Enforced" means ci / Required is an actual required status check on the repo's default branch, read back from the API — not that the caller parses.

Workflow First adopter Enforced
apple.yml narduk-enterprises/GeoGridKit not yet
python-data.yml narduk-enterprises/narduk-data (earth-data-ci.yml) not yet
reusable-browser-tests.yml narduk-enterprises/been-sober-for (first proof PR) not yet
docs-governance.yml narduk-enterprises/company-hq yes — repo ruleset require-docs-governance
node-library.yml narduk-enterprises/narduk-charts yes — repo ruleset require-ci-required
nuxt-cloudflare.yml hydrogen no — hydrogen has no branch protection; it called @v1 unenforced for months, which is the failure mode this column exists to make visible
reusable-node-ci.yml retired 2026-09-24 — zero live callers estate-wide, file deleted (narduk-reboot P3-C2 / O-D8) —
code-review.yml retired 2026-09-19, file deleted 2026-09-24 — no live pool adopters n/a — historical advisory callable; current review is the CT650 PR review bot (runners#212)
closing-syntax-check.yml retired 2026-09-24 — zero adopters, file deleted (narduk-reboot P3-C2 / O-D8); narduk-enterprises/agent-infrastructure stays on its own local, canonical invocation of the same (now commit-scanning) checker —
reusable-weekly-drift-check.yml retired — zero live callers verified across narduk-enterprises and narduk-incubator —
red-main-listener.yml narduk-enterprises/workflows itself (red-main-self.yml, listening to this repo's own CI) not yet
flake-digest.yml none yet — no adopter here has a quarantine lane of its own; nuxt-cloudflare.yml adopters with e2e-quarantine-args set are the natural first callers not yet
design-ledger.yml narduk-enterprises/mybo-at-v2 (planned) never — advisory by design (Logan, 2026-09-27: "Red check + drift issue (Recommended)")

The ci / Required convention

This is the actual point of the repo, not an implementation detail.

docs-governance.yml, node-library.yml, and nuxt-cloudflare.yml each end with a job named exactly Required. (Since the CI reset, 2026-09-28, docs-governance.yml's Required is its only job and runs the check itself: with nothing to aggregate, a second job only cost a runner allocation per call. lint_callables.py R5 exempts exactly that self-contained shape from always().) An aggregating Required needs: every other job the workflow defines, runs with if: always() (or, in nuxt-cloudflare.yml, if: "!cancelled()", below), and explicitly checks each needs.<job>.result — a job that's enabled must report success; it may report skipped only when its controlling input is off. In particular, run-e2e: true makes the plan and every E2E shard mandatory. A skipped enabled job is failure, not an acceptable substitute for a toolchain check that never ran. if: always() jobs succeed by default if you don't check anything explicitly — these don't skip that check.

nuxt-cloudflare.yml's Required, Fast and escalated Fast use !cancelled() rather than always() (CI reset, 2026-09-28). Both run after a failed or skipped need; they differ only when the whole run is cancelled. always() then still started the job, which queued for a runner only to fail, and held the caller's concurrency group meanwhile: operator-portal run 36494070028 kept the next main run waiting about 14 minutes. With !cancelled(), GitHub reports the job cancelled without a runner. That conclusion never satisfies a required check (only success, neutral and skipped do), so a cancelled run still cannot pass. In cancelled runs acre-oracle 36490590502 and operator-portal 36494070028, every job whose if: was false after the cancel reported cancelled, never skipped. lint_callables.py R5 accepts exactly !cancelled() as the alternative.

This generalizes a pattern narduk-libs already proved in production: its ci.yml runs a 12-lane package matrix, then a verify job that needs: package-gates and fails unless needs.package-gates.result == 'success' — one stable pass/fail signal regardless of how many matrix lanes ran underneath.

The subtlety worth restating: for a called reusable workflow, GitHub composes the check-context string as <caller's job id> / <job's name>. So even with an identical reusable workflow, a caller whose calling job is named quality: or verify: instead of ci: produces a different context string. The convention has to cover both ends:

  • Every reusable workflow here ends with a job named Required.
  • Every caller names its calling job id ci.

That composes to one identical string on every adopting repo, regardless of app type or what runs underneath:

ci / Required

That string is what a branch-protection rule or org ruleset requires. This is what reopens untangle/TRACKER.yaml D-PKG-3 Part B: once a publisher repo's caller conforms, the org ruleset can require ci / Required without a per-repo verification pass, because the string is enforced by convention, not discovered by inspection.

Dependency caching: restore on every ref, write only on the default branch

Every Node workflow here (node-library.yml, nuxt-cloudflare.yml, docs-governance.yml) used to hand cache: to actions/setup-node and let it manage the whole round trip. That is the wrong shape, and company-hq#269 measured why.

setup-node writes a fresh cache tarball whenever the primary key misses, scoped to the ref the job ran on. A cache written on refs/pull/N/merge is restorable only by that same pull request — no other PR and not the default branch. So the sequence for any lockfile-changing PR was: miss → install → pay a Post Run actions/setup-node tar for an entry nothing would ever read → branch dies → merge to main pays the identical tar a second time.

Measured on real runs before the change:

Repo / job restore install Post Run setup-node (save) job wall
marketing-web / Build 1s (miss) 13s 19s 67s
narduk-charts / package / default 1s (miss) 7s 14s 52s
status-apps / browser shards (pool) 5–7s 19–20s 32–54s —
earthdata-viewer / chromium shard — 10.6s 29s avg, 115s max —
nvault / Verify — 9.7s 31.5s avg, 315s max —

The debris was visible in the API as well as the clock: narduk-charts carried 8 cache entries, four of them refs/pull/* duplicates of a refs/heads/main entry; vtraceroute held a 317MB refs/pull/2/merge copy of the 317MB entry main already had.

So these workflows now split the two halves explicitly:

  • actions/cache/restore on every ref, with a restore-keys prefix so a pull request whose lockfile moved still falls back to the default-branch entry — which is the entry it actually wants.
  • actions/cache/save only when github.ref_name equals the caller's default branch, and only when the exact key missed.

github.* resolves against the caller's event inside a reusable workflow, so ref_name is the caller's branch on a push and N/merge on a pull request. If github.event.repository is ever absent the comparison is simply false and the write is skipped — the safe direction.

Two deliberate exceptions:

  • python-data.yml keeps setup-uv's enable-cache: auto. auto means on for GitHub-hosted, off for self-hosted, which is already correct: a persistent linux-ci guest keeps ~/.cache/uv between jobs, so uploading a tarball of an already-warm cache would be a regression for the only current adopter. Only save-cache is gated, for a future hosted caller.
  • reusable-weekly-drift-check.yml keeps cache: pnpm. It is a weekly schedule on ubuntu-latest, so it runs on the default branch almost every time — the write it makes is the one the next run reads.

This is not a caller-visible change: no inputs were added or removed, and a caller pinned to @v1 picks it up when v1 moves.

Runner routing (runner input)

docs-governance.yml, node-library.yml, nuxt-cloudflare.yml and python-data.yml each accept a runner input: a JSON-encoded string, decoded with fromJSON() at every job's runs-on:. apple.yml takes the same encoding but splits it into two inputs — lint-runner and apple-runner — because routing Apple CI per job rather than per repo is the whole reason that file exists. It accepts three shapes:

runner: '"ubuntu-latest"'                                              # plain string
runner: '["self-hosted","Linux","X64","proxmox","linux-ci"]'           # JSON array
runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'  # JSON object

The object form matches Config/github-runner-fleet.json's runsOn shape exactly, so a private manifest-routed caller can paste that value verbatim. The value must be valid JSON — a bare string still needs its own quotes, which is why the explicit hosted value is the four-character JSON string "ubuntu-latest", not the bare word.

Default route (empty runner)

The runner inputs of docs-governance.yml, node-library.yml, python-data.yml, red-main-listener.yml and flake-digest.yml, and apple.yml's lint-runner, default to the empty string (row 18 Q3, Logan 2026-09-18, "Flip the default (Recommended)"). Every runs-on: resolves the effective route as:

inputs.runner || github.event.repository.private == true && '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}' || '"ubuntu-latest"'
Caller Before After
private, passes nothing GitHub-hosted ubuntu-latest (policy drift, agent-infrastructure docs/standards/CI-RUNNER-POLICY.md §1) linux-ci organization group, group and labels (§4)
public, or an event with no repository payload, passes nothing ubuntu-latest ubuntu-latest (§3)
any caller, explicit value that value that value — unchanged, proven by scripts/test_runner_default.py

The comparison is == true, so an unknown visibility falls to hosted, the safe direction. Because visibility is read at run time, a caller that goes public lands on hosted on its next run with no edit. Blacksmith overflow and CI_LIGHTWEIGHT_RUNNER apply to the effective route exactly as they applied to inputs.runner before; an effective "ubuntu-latest" is never sent to Blacksmith. reusable-browser-tests.yml's contract job and its Required fallback use the same visibility gate with no caller input at all, so the route contract is still validated on a runner this file chose.

The route literal is the fleet manifest's linux-ci class (fleet/Config/github-runner-fleet.json, organization group linux-ci). It is one string repeated at each site; test_runner_default.py fails if any copy diverges, but nothing here re-reads the manifest, so a manifest label change must be mirrored here.

The linux-ci group is selected-visibility. A private caller the manifest does not list in that group gets a job that queues with no eligible runner when it passes nothing. Such a repo must be added to the manifest group (the policy fix) or pass '"ubuntu-latest"' explicitly, and a repo holding a §2 hosted exception (for example package-delivery, exception 5) must pass '"ubuntu-latest"' explicitly. See the versioning note on this change below.

A PUBLIC CALLER MUST NEVER PASS A SELF-HOSTED LABEL. A fork PR on a public caller can run attacker-controlled code, so a self-hosted runner value hands that PR estate infrastructure. A repo can also become public later, and runner is a free-form string these workflows cannot police. Only private, manifest-routed callers may pass a self-hosted value — resolve it first with python3 scripts/github_runner_fleet.py route (or scripts/onboard-proxmox-runner.sh --route-only) in the agent-infrastructure checkout, then copy the returned runsOn object verbatim. Every workflow file in this repo repeats this rule in a loud top-of-file comment; don't rely on this README alone when adding the next one.

Blacksmith overflow (BLACKSMITH_RUNNERS_ENABLED)

D-CI-CAP-1 (c) (2026-09-07, extends D-BLACKSMITH-2; Operator Portal decision log, fleet#337) makes Blacksmith a kill-switched overflow for the linux-ci class's ordinary private CI. Every runs-on: ${{ fromJSON(inputs.runner) }} site in nuxt-cloudflare.yml, node-library.yml, docs-governance.yml, and python-data.yml (16 sites, none of them a deploy-credentialed job — nuxt-cloudflare.yml's only deploy-shaped job, deploy-dry-run, runs wrangler deploy --dry-run with zero Cloudflare secrets in scope) resolves as:

runs-on: ${{ fromJSON(vars.BLACKSMITH_RUNNERS_ENABLED == 'true' && inputs.runner != '"ubuntu-latest"' && format('"{0}"', vars.BLACKSMITH_LINUX_LABEL || 'blacksmith-2vcpu-ubuntu-2404') || inputs.runner) }}

In the callables with an empty runner default, each inputs.runner above is the parenthesised effective route from Default route, so a private caller that passes nothing is Blacksmith-eligible (it is on the linux-ci class this overflow serves) and a public one never is.

  • Switch: the org Actions variable BLACKSMITH_RUNNERS_ENABLED (default false, visibility restricted to private repos). Only the exact string true selects Blacksmith. A repo-level variable of the same name overrides the org default for exactly that repo (GitHub's normal repo-over-org precedence) — the mechanism for canarying or kill-switching one adopter without moving the org default.
  • Label, not a group id: the vendor's documented blacksmith-2vcpu-ubuntu-2404 string, overridable via vars.BLACKSMITH_LINUX_LABEL. No Blacksmith runner-group id is pinned anywhere — D-BLACKSMITH-4's live proof established that provider-created groups are not a stable routing contract; only the label is.
  • Public node-library.yml callers are never affected: its routing expressions first require github.event.repository.private == true, so they ignore both Blacksmith and CI_LIGHTWEIGHT_RUNNER for public repositories. Their caller-supplied hosted runner value remains the route.
  • Manual, not automatic fallback: GitHub does not move an already-queued job to another runs-on: target. If Blacksmith cannot schedule or its free allowance is exhausted, flip the variable back to false (or remove the repo-level override) and rerun — same operational shape as every other Blacksmith cohort (D-BLACKSMITH-3).
  • Cost boundary (D-BLACKSMITH-2/3, unchanged): free allowance only, no payment method, no paid overage; disable at 2,400 equivalent 2-vCPU minutes in a monthly cycle, any non-zero amount due, or any unexpected billing state, whichever comes first.
  • Never routes here: production/deploy jobs (all app-owned and bespoke, outside these six CI-only callables), the Playwright/browser class (reusable-browser-tests.yml's browser-runner job is untouched — its own trust boundary per agent-infrastructure docs/standards/CI-RUNNER-POLICY.md §5), and Apple builds (apple.yml is untouched — it has its own D-APPLE-CI-1 ladder).
  • Per-repo ordinary route (CI_LINUX_RUNNER, CI reset 2026-09-28): in nuxt-cloudflare.yml, the ordinary private route of Build, Checks, Extra gate, Deploy dry run and both Fast jobs reads a repo-level Actions variable CI_LINUX_RUNNER (a JSON runsOn value). When set on a private caller it wins over the caller's runner input (ci-reset W10, W8 item 8): every sampled nuxt caller passes runner: with the linux-ci route verbatim, so the variable used to be unreachable exactly where the overflow was needed: (github.event.repository.private == true && vars.CI_LINUX_RUNNER || inputs.runner || github.event.repository.private == true && '<linux-ci route>' || '"ubuntu-latest"'). A controller sets or deletes it to move one burst repository's ordinary CI (for example to a Blacksmith label) without a pin bump. Setting it is a deliberate per-repo routing decision that also overrides an explicit runner: (a CI-RUNNER-POLICY §2 hosted exception included), so only set it where that repo's policy allows the target. Unset, the route is byte-for-byte the previous one (scripts/test_runner_default.py evaluates both). It never applies to a public caller and does not touch the lightweight, browser or preview routes. The Blacksmith switch above still applies to the resulting route unchanged. Configuration variables in a called workflow resolve from the CALLER's repository (GitHub docs, "Variables": "For reusable workflows, the variables from the caller workflow's repository are used"), so a repo-level value reaches this callable.
  • Adding a job is not additive here without checking this section again: a new job that copies ${{ fromJSON(inputs.runner) }} verbatim does not get Blacksmith overflow automatically — use the expression above, or it silently stays off the overflow route.

Cursor review (retired 2026-09-28)

The cursor-review.yml callable, this repository's cursor-review-self.yml, scripts/cursor_review.py, scripts/test_cursor_review.py and scripts/cursor_review_brief.md were deleted in ci-reset W6 (workflows#166), and each repository's caller is removed by its own PR. A caller that still exists pins a full SHA, so it keeps working until its removal merges. The callable held a runner seat for the whole Cursor wait.

Reviews now come from the PR review bot on CT650:

  • Code: narduk-enterprises/runners scripts/pr_review_bot/, operator guide docs/pr-review-bot.md, runners#212.
  • Triggers: a PR comment whose first line is /review or /review deep, or the review-now / review-deep label, from an organization member with write access.
  • It keeps this callable's prompt (moved verbatim), models (Composer 2.5 standard, Grok 4.6 xhigh deep, fast: false) and review format, and posts a real review from narduk-durable-agent-sessions[bot].
  • Nothing to adopt: no caller file and no repository secret.

Git history holds the deleted callable and its design notes.

How to consume

Caller template

Every caller is a thin workflow with one job named ci (this is what makes the ci / Required convention work — see above) and triggers on pull_request, push to the default branch, and workflow_dispatch. Never trigger PR-only — a PR-only trigger means the default branch's tip carries no CI status after merge, which is exactly what happened to narduk-skills and (pre-merge) narduk-eslint-config per the CI-5 phase 2 audit.

name: CI

on:
  pull_request:
  push:
    branches: [main] # match the repo's actual default branch
  workflow_dispatch:

concurrency:
  group: ci-${{ github.repository }}-${{ github.event.pull_request.number || github.ref }}
  cancel-in-progress: true

jobs:
  ci: # <-- must be named `ci`; GitHub composes "ci / Required" from this + the reusable workflow's job name
    uses: narduk-enterprises/workflows/.github/workflows/<workflow>.yml@<sha-of-v2> # v2
    with:
      # ...per-workflow inputs, see below
    secrets:
      NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}

Concurrency is the caller's job, and putting it here would be actively dangerous

Every caller must set its own workflow-level concurrency, as in the template above. No workflow in this repo declares one, and none ever should. This is a hard rule, not a gap waiting to be filled — the structural gate in .github/workflows/ci.yml fails the build if a callable grows a concurrency: block (rule R6).

The reason is stronger than "it does not propagate". It is that a group here is evaluated in the caller's context, which GitHub states plainly:

A called workflow uses the name of its caller workflow in ${{ github.workflow }}, so using this context as the value of jobs.<job_id>.concurrency.group in both caller and called workflows will cause the caller workflow to be cancelled when the called workflow runs.

— Reusing workflow configurations

So the obvious-looking group — ${{ github.workflow }}-${{ github.ref }}, which is what almost everyone writes — would collide with the caller's own group and cancel the run that is calling us. Seven repos would start cancelling their own CI the moment v1 moved, and the symptom (a run that cancels itself for no visible reason) points at the adopter, not at here.

A callable-level group that was carefully uniquified to avoid the collision would still buy nothing: cancel-in-progress on the caller's workflow-level group already supersedes the entire previous run, jobs of this callable included. A second gate underneath it can only add a way to be wrong.

concurrency is a permitted key on a job that calls a reusable workflow, so a caller with an unusual need can scope it at jobs.ci.concurrency — but per the same doc, do not reuse the callable's group value there either.

Path filtering: use the boolean inputs, never paths / paths-ignore

Do not put paths: or paths-ignore: on a caller's on: trigger. If the workflow does not run, ci / Required is never reported, and GitHub shows a required check that never arrives as permanently "Expected" — the pull request becomes unmergeable and stays that way. That is company-hq#146, and it does not fail loudly; it just quietly stops being mergeable.

The safe mechanism already exists and needs no change to these workflows: the opt-in gates are ordinary boolean inputs, so a caller can compute them from its own diff and pass the result. A gate turned off this way reports skipped, which every Required job already accepts, and Required itself still runs and still reports. The context never disappears.

jobs:
  changes: # cheap; no checkout of the heavy tree needed
    runs-on: ubuntu-latest
    timeout-minutes: 5
    permissions: { contents: read, pull-requests: read }
    outputs:
      code: ${{ steps.f.outputs.code }}
    steps:
      - uses: dorny/paths-filter@<full-sha> # vX.Y.Z
        id: f
        with:
          filters: |
            code:
              - '!(**/*.md|docs/**)'

  ci: # still named `ci`; still composes `ci / Required`
    needs: changes
    uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2
    with:
      run-e2e: ${{ needs.changes.outputs.code == 'true' }}
      wrangler-dry-run: ${{ needs.changes.outputs.code == 'true' }}

Which gates each callable exposes this way:

Callable Caller-gatable Always runs
nuxt-cloudflare.yml run-e2e (e2e, e2e-plan), wrangler-dry-run, run-tests build, checks (its steps run inside build with checks-in-build)
apple.yml run-swiftlint / linux-checks (lint), run-build, run-tests xcode
python-data.yml run-ruff (ruff steps in test), run-tests, run-pyright test
reusable-browser-tests.yml run-webkit (webkit) validate, chromium
node-library.yml run-lint, run-typecheck, run-tests, run-build package
reusable-weekly-drift-check.yml all three jobs —
docs-governance.yml — the single job

The "always runs" column is deliberate and is not a gap to be closed. Those jobs are the ones the required check actually certifies. Giving them a path condition means ci / Required can report green on a change that was never built — a false green, which is strictly worse than the wasted minutes it saves, and which nobody discovers by looking at a passing pull request. If a caller wants a docs-only change to cost less, it turns off the opt-in gates above and still builds.

python-data.yml declares no secrets. apple.yml accepts an optional DEPENDENCY_SSH_KEY for a private SwiftPM repository. Pass a read-only deploy key scoped to that dependency; never a release, signing, or account key. Build/test rewrite HTTPS package URLs to SSH only for the caller's GitHub organization, using Git's process environment. A pinned GitHub Ed25519 host key authenticates the server. The private key lives in a mode-0600 temporary file removed on success or failure; Git config, Keychain and GITHUB_ENV stay unchanged. Xcode callers should use -scmProvider system so resolution uses this process configuration. Callers that omit the secret keep their existing behavior. Do not pass credentials to untrusted code or fork pull requests.

    secrets:
      DEPENDENCY_SSH_KEY: ${{ secrets.PRIVATE_SWIFTPM_SSH_KEY }}

reusable-browser-tests.yml is the one exception: it declares a single optional NARDUK_PLATFORM_GH_PACKAGES_READ secret (required: false), mirroring node-library.yml and nuxt-cloudflare.yml's existing CI-alias-for-the-PAT convention, so that a browser CI caller installing cross-repo @narduk-enterprises/* packages can pass a real read credential. A caller that passes nothing keeps today's behavior unchanged: every consuming step falls back to the ephemeral github.token.

apple.yml

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/apple.yml@<sha-of-v2> # v2
    with:
      # Repo-scoped Mac from Config/github-runner-fleet.json `appleRepositories`.
      # REQUIRED — there is no default, on purpose.
      apple-runner: '["self-hosted","macOS","apple-imac"]'
      run-swiftlint: true
      build-command: swift build
      test-command: swift test

Two runner inputs, routed per job, because the estate has one always-on Mac slot and AGENTS.md requires that only work needing Xcode/macOS occupy it:

  • apple-runner — xcodebuild, swift build/swift test, anything needing the Apple toolchain. Required, no default.
  • lint-runner — SwiftLint plus linux-checks for shell/grep boundary gates and plist checks via python3 plistlib (not PlistBuddy/plutil, which are macOS-only). SwiftLint does not compile the project, but its official Linux binary still dynamically loads libsourcekitdInProc.so; self-hosted linux-ci runners provide the pinned Swift SourceKit runtime layer. Empty by default: a private caller gets the linux-ci organization-group route and a public one "ubuntu-latest" (see Default route). Passing the route explicitly still works and still wins:
      lint-runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
      linux-checks: |
        python3 scripts/check_plists.py
        ! grep -rn "import UIKit" Sources/MyCore

run-swiftlint is opt-in and the SwiftLint archive is verified against a pinned swiftlint-sha256 before it is unpacked. Enabling SwiftLint on a repo with no .swiftlint.yml runs the full default rule set and is usually red on first contact — GeoGridKit's first run produced 158 errors — so adopt a repo-owned config in the same change. Prefer only_rules: over disabled_rules: there: an allowlist cannot be broken by a future SwiftLint release adding a rule.

The checksum verifies the downloaded SwiftLint archive; it does not provide the Swift SourceKit runtime. The install step separately fails closed unless /usr/lib/libsourcekitdInProc.so is readable, exports that exact path through LINUX_SOURCEKIT_LIB_PATH, and still locates the unpacked swiftlint binary with find because upstream archive layouts have changed.

Archive/sign/notarize/TestFlight/Sparkle are not here, for the same reason node-library.yml omits publish: they need temporary keychains, the host-wide Apple build lock and per-repo credentials, and folding them in would put release credentials behind a PR-triggered gate. That stays the apple-release-pipeline skill's per-repo workflow.

python-data.yml

uv (the default), with a lockfile check:

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/python-data.yml@<sha-of-v2> # v2
    with:
      runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
      working-directory: services/my-pipeline
      venv-path: .venv-ci
      uv-sync-args: "--locked --extra test"
      test-command: python -m pytest tests -q

pip/venv, for the repos that have not moved to uv:

    with:
      dependency-manager: pip
      pip-install-args: '-e ".[test]"'

Notes:

  • The environment's bin/ is prepended to $GITHUB_PATH after install rather than source-ing an activate script, because activation dies with the step's shell. That is what makes extra-checks work from any directory.

  • extra-checks is a multi-line shell hook that runs after the tests, in the same environment, and is covered by ci / Required. It exists because a caller cannot add steps to a called workflow's job, so a repo-specific gate that needs the installed package would otherwise have to stay behind as a second job duplicating the entire install.

  • run-ruff is opt-in and ruff-version is pinned exactly, so a ruff release cannot turn a green repo red. Pointing ruff at a repo that never had a static-analysis gate is usually red on first contact — narduk-data's earth-data-pipeline has 36 violations today, including six F821 undefined-name — so the template does not decide for the caller when to take that on. Since ci-reset W10 (W8 item 10) ruff runs as the last steps of the test job, not in a lint job of its own: that job did about 0.1 min of work for a whole runner allocation. The steps run under !cancelled() after the tests, so a lint finding and a test failure both show, and ruff gets its own interpreter (called by path), so it never installs into the project environment. A caller's ruleset that required <job> / lint would lose that context; none does (narduk-data's earth-data-ci is the only caller, and narduk-data's require-checks ruleset names no ci / lint).

  • run-pyright is opt-in for v1 compatibility, but it is a real static gate: the workflow provisions Node 24 explicitly, installs exact pyright-version under $RUNNER_TEMP, and runs it in the installed Python environment. This is intentionally not py_compile, which proves syntax and nothing about names or types.

  • Several isolated pytest invocations (narduk-data's ci.yml needs them, because two suites share a module basename with no __init__.py) go in test-command as a multi-line string, or in extra-checks.

  • extra-env values are expanded on the runner. Write $GITHUB_WORKSPACE, never ${{ github.workspace }} (workflows#4):

        # RIGHT — expanded on the runner by the workflow itself
        extra-env: |
          PYTHONPATH=$GITHUB_WORKSPACE
    
        # WRONG — silently becomes `PYTHONPATH=` and breaks a later step
        extra-env: |
          PYTHONPATH=${{ github.workspace }}

    with: inputs are evaluated in the caller, and a jobs.<id>.uses: job is never assigned a runner, so github.workspace there is the empty string. Until this was fixed the workflow accepted PYTHONPATH= as a well-formed KEY=VALUE and the failure surfaced four steps later as ModuleNotFoundError: No module named 'pipelines', with nothing anywhere naming extra-env.

    $NAME and ${NAME} are expanded against the runner's environment. Nothing else is: no $(...), no backticks, no ${NAME:-default}, no globbing, no eval. An empty value — or one whose every reference is unset — is now a hard error naming this trap, because a silently-unset variable is the worst outcome. scripts/test_extra_env.py locks all of that down against the step text extracted from the YAML itself, so the tests cannot drift from the shipped script.

Central lightweight-job routing

nuxt-cloudflare.lightweight-runner routes E2E plan and Required independently of the build. Selection is explicit input, then CI_LIGHTWEIGHT_RUNNER, then the existing build/Blacksmith route. Public callers always use ubuntu-latest for these jobs. An unset override is backward compatible; required check names and failure/skip semantics are unchanged.

CI_LIGHTWEIGHT_RUNNER is an organization Actions variable containing a JSON runs-on value, initially "ubuntu-slim" for the authorized company gates. Every shared Required job reads it except python-data.yml's, which runs on the caller's own test route (CI reset 2026-09-28): its only caller runs those jobs GitHub-hosted, so a lightweight Required on linux-ci was the one self-hosted job in each run. Changing this one value changes routing for subsequent jobs without changing callable code or repinning callers. Package, browser, Apple, and deployment jobs retain their own routes. node-library.required-runner remains an explicit per-caller override; avoid setting it on ordinary callers that should follow the central route. Repository variables take precedence over organization variables, so reserve repository values for documented exceptions.

Inline planners or final assertions need a one-time adoption of the same expression, preserving their job names and dependency conditions:

runs-on: ${{ fromJSON(vars.CI_LIGHTWEIGHT_RUNNER || '"ubuntu-slim"') }}

Use this only for bounded, secrets-free metadata and result checks with no private-network requirement. Their existing policy authorization still applies. It is a routing contract, not automatic classification by job duration. Already queued jobs keep their selected route. A composite action cannot select a runner because it starts after GitHub has assigned one.

Consumers pinned before this feature need one reviewed SHA update. That initial adoption is unavoidable; later capacity changes require only the organization variable. Keep immutable workflow pins. Unsetting the variable restores each callable's prior fallback, and malformed JSON fails visibly. Public callers must continue to use GitHub-hosted runners; never give the variable a self-hosted route in an organization that exposes it to public callers.

node-library.yml

required-runner optionally separates the small Required aggregation job from the package runner. It accepts the same JSON runner shape as runner; empty uses the organization-level CI_LIGHTWEIGHT_RUNNER route for private callers, falling back to existing routing (including Blacksmith) when that variable is absent; public callers retain their supplied hosted route. An explicit value takes precedence for Required only. The gate checks dependency results without checking out source, installing packages, or receiving registry secrets. For a secrets-free gate before a privileged self-hosted release, the caller may pass required-runner: '"ubuntu-slim"' under company-hq's CI runner policy exception 2. This keeps completion reporting independent of a saturated build queue. It does not grant other callers a hosted-runner policy exception.

A monorepo may group packages into a bounded number of install lanes. Use filter: "", a lane-level extra-scripts: "ci:batch", and set all four run-* inputs to false. The root ci:batch script receives the complete lane in PACKAGE_MATRIX_JSON; for example, a lane can carry packages: ["@example/core", "@example/auth"]. That repository-owned script must validate every selection, reject missing gates, retain every nonzero exit and print per-package results. Run the same script locally. The callable continues to own runner setup, one install per lane and auth cleanup; existing single-package lanes are unchanged. Choose the batch count from measured setup cost and capacity, and retain the caller's final integration aggregate.

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
    with:
      node-version: "22"
      package-manager: pnpm
    secrets:
      NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}

install-args and extra-scripts were added for narduk-charts' adoption (company-hq#172) and are worth knowing about, because between them they are the difference between retiring a local job and retiring half of one:

    with:
      package-manager: npm
      install-args: "--legacy-peer-deps" # charts cannot `npm ci` without it
      extra-scripts: size                # runs after build, in the same lane

extra-scripts runs each named package script after build, in the same job, so a gate that needs build output (size-limit) does not require a second job with a second full install and build. Each name is probed first and fails when it is absent by default.

With a per-package matrix (generalizes narduk-libs' package-gates):

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
    with:
      package-matrix: |
        [
          {
            "label": "narduk-core",
            "filter": "@narduk-enterprises/narduk-core",
            "extra-scripts": "check:dist"
          },
          {
            "label": "narduk-auth",
            "filter": "@narduk-enterprises/narduk-auth",
            "extra-scripts": ""
          }
        ]
    secrets:
      NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}

Matrix lanes own their extras. The optional lane-level extra-scripts field overrides the legacy shared input even when it is ""; that explicit empty value means the lane has no extra gates. Omitting the field preserves the shared-input behavior for existing callers, while declaring it on every lane prevents one package's requirement from leaking into another.

The org Actions secret NARDUK_PLATFORM_GH_PACKAGES_READ maps into GH_PACKAGES_READ on the auth and install steps. The legacy process alias is also supplied for existing caller bootstrap scripts; it is an interface compatibility detail, not another secret to create. There is no implicit github.token fallback. Missing credentials fail before private-package installs; public-only installs need no credential. A caller whose committed project .npmrc routes @narduk-enterprises to the anonymous https://npm.nard.uk mirror (company-hq D-PKG-6) counts as public for that scope and can drop the secret, unless it also depends on @narduk-geo (not mirrored) or its lockfile still names npm.pkg.github.com (workflows#106). The web-foundation check (foundation-check: true) follows the same route (workflows#109): a mirror-routed caller runs it with no credential at all, and the pinned dlx download fetches from https://npm.nard.uk. That needs narduk-app-tools 0.10.0 or later (the foundation-check-tool-version default), because older releases hard-code GitHub Packages for their live registry lookup. A caller that has adopted the tool as a dependency must depend on 0.10.0 or later too. Without a caller bootstrap, the callable writes a temporary user config containing a literal variable reference, then removes it after the install.

A public monorepo with only workspace packages under an estate-looking scope can opt out of that automatic name-based detection without forwarding a token:

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/node-library.yml@<sha-of-v2> # v2
    with:
      runner: '"ubuntu-latest"'
      required-runner: '"ubuntu-latest"'
      package-registry-auth: disabled

disabled asserts that every installed dependency is public. The default auto remains fail-closed for private registry dependencies and preserves the existing package-read-secret contract for private callers.

nuxt-cloudflare.yml

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2
    # A called workflow's jobs may only request permissions the caller granted;
    # asking for one it did not kills the whole run at startup (workflows#59).
    # `pull-requests: write` is required by the `preview` lane's sticky comment
    # -- see "Preview checks" below before pinning a ref that carries it.
    permissions:
      contents: read
      packages: read
      pull-requests: write
      actions: read # E2E proof lookup (required by the new revision)
    with:
      node-version: "24"
      package-manager: npm # hydrogen's current package manager; pnpm is the default
      # Optional: delegate the complete install to a caller-owned wrapper.
      # The callable passes the distinct service-token secret only as
      # NVAULT_TOKEN and skips its legacy direct package-registry path.
      install-script: ci:install
      typecheck-worker-script: typecheck
      typecheck-web-script: web:typecheck
      run-e2e: true
      wrangler-dry-run: true
      # The estate web quality gates (foundation check, performance budget,
      # live security headers), each with a reasoned opt-out. See
      # "Quality level" below; new apps start here.
      quality-level: standard
    secrets:
      # Foundation checks and direct registry installs use the canonical PAT.
      NARDUK_PLATFORM_GH_PACKAGES_READ: ${{ secrets.NARDUK_PLATFORM_GH_PACKAGES_READ }}
      # Caller-owned install scripts receive the service token separately.
      NVAULT_TOKEN: ${{ secrets.NVAULT_TOKEN }}
      # The CI artifact store (see "The CI artifact store (2026-09-26)").
      # Optional: without them each E2E job builds its own application and no
      # reuse proof is published or honoured.
      CI_ARTIFACTS_R2_ACCOUNT_ID: ${{ secrets.CI_ARTIFACTS_R2_ACCOUNT_ID }}
      CI_ARTIFACTS_R2_ACCESS_KEY_ID: ${{ secrets.CI_ARTIFACTS_R2_ACCESS_KEY_ID }}
      CI_ARTIFACTS_R2_SECRET_ACCESS_KEY: ${{ secrets.CI_ARTIFACTS_R2_SECRET_ACCESS_KEY }}

install-script is for repositories whose install wrapper exchanges an nVault service token for the package credential, materializes any temporary registry configuration itself, runs the package manager, and removes the configuration on exit. The value must be one package.json script name using letters, digits, :, _, or -; the callable rejects missing or unsafe names. Existing callers that leave it empty keep the legacy install path unchanged.

NARDUK_PLATFORM_GH_PACKAGES_READ always carries the org package-read PAT; foundation checks use it by default. NVAULT_TOKEN carries the caller's service token when an explicitly selected install-script needs nVault; that installer resolves GH_PACKAGES_READ from nVault before starting npm/pnpm. Generic caller-owned install scripts may omit it and validate their own credentials. When NVAULT_TOKEN is empty (Dependabot runs see only Dependabot-store secrets, and NVAULT_TOKEN is Actions-only) and foundation-check-auth is not nvault, the caller script instead receives the org package-read secret as GH_PACKAGES_READ, so it can install without an nVault exchange. It never receives both (workflows#98). Raw PAT consumers omit install-script and pass only secrets.NARDUK_PLATFORM_GH_PACKAGES_READ.

foundation-check-auth: nvault remains a compatibility mode for callers that previously mapped their service token into the package-read secret. New callers use the default package-token mode and map both secrets as shown above.

Local workstations use gh-packages-run, and Workers Builds uses the protected build secret GH_PACKAGES_READ. Neither uses the Actions input name as a vault key. The credential route provides the exact nVault selector and value-free diagnostics.

Deploying with real Cloudflare credentials on push-to-main is not this workflow's job — that stays a separate nuxt-cloudflare-deploy.yml sibling (deferred, not built in this pass), matching hydrogen's existing two-job ci / deploy split rather than folding deploy secrets into the CI gate.

Exact candidate validation

expected-candidate-sha is an optional full, lowercase 40-character commit SHA. For an explicit validation request, create a fresh branch ref narduk-validation/<full-sha>/<request-id> pointing at that already-pushed candidate. The app-owned validation caller listens only to pushes under narduk-validation/**, retains job ID ci, and passes github.sha as this input. Ordinary development branches do not trigger it. This is an explicit full-validation operation, never part of a local development deployment.

Every source checkout is pinned to the requested SHA, then verifies the push event, ref-encoded candidate, event SHA and actual checkout before running package code. A mismatched ref or checkout fails the real ci / Required gate. Leaving the input empty preserves normal callers. The requesting CLI verifies the selected source branch still points at the requested candidate before creating the validation ref and records its reason and resulting run.

Do not use workflow_dispatch for release evidence: GitHub excludes those job checks from pull-request required-check evaluation, even on the correct head commit. See GitHub's event eligibility documentation. The dedicated push ref lets the existing real required check count without changing branch protections or introducing a synthetic success check.

The explicit caller must pass the application's full normal release checks, including its browser coverage. Exact-candidate mode disables browser path skipping and reuse; push semantics select the full shard/argument configuration. Do not reuse an off-main expression that disables E2E, add a success substitute for held workflows, or attach promotion/migration jobs. Required checks and review rules stay intact. Record the exact successful run and candidate; a later merged revision requires its own exact-revision evidence.

node-version-file: single-sourcing the Node version

Input Type Default Purpose
node-version string "24" Node.js version passed straight to actions/setup-node
node-version-file string "" Optional path (relative to working-directory) to a file declaring the Node version, e.g. .node-version or .nvmrc

Empty (the default) preserves prior behaviour exactly: every setup-node step in this callable uses node-version. Setting node-version-file lets a caller single-source its Node version from a file it already maintains — .node-version, .nvmrc, package.json's engines via a generated file, etc. — instead of ALSO pinning it as this input's own value, which is how a caller's Node pin and its CI pin drift apart.

actions/setup-node rejects node-version and node-version-file together, so every setup-node step here resolves them as two mutually exclusive expressions rather than passing both: node-version-file is passed through unchanged, and node-version resolves to an empty string whenever node-version-file is set (an empty string is setup-node's own "not provided" sentinel for either input). A caller that sets both gets node-version-file; node-version is silently ignored in that case, exactly as if the caller had left it unset.

Caller lint: hygiene gate over the caller's OWN workflows

Every nuxt-cloudflare.yml adopter gets Caller lint as part of Required (Logan, 2026-09-17 askme round, "Caller lint + job timeouts in workflows (Recommended)"; company-hq#745). Since the CI reset (2026-09-28) it is a set of steps inside the Required job, not a Caller lint job of its own: the separate job cost a whole runner allocation for about ten seconds of lint on every call. The steps run after Required's gate steps under !cancelled(), so a lane failure and a lint finding both show in one run, and a finding fails Required directly. On a protected-path pull request the job holding Fast (fast-escalated) runs the same anchored scripts, and when Required is folded into Build (see checks-in-build below) Build runs them at its own full checkout. Required's own checkout is sparse (ci-reset W10, W8 item 9): only .github and action.yml/action.yaml at any depth, which is all the lint reads, so it no longer fetches the whole tree for ten seconds of lint. It checks out the CALLING repository (not this one), runs pinned actionlint over the caller's own .github/workflows/*.yml, and runs a small inline Python audit that fails the job when a caller workflow:

  • has no workflow-level concurrency: block (skipped for a file whose only trigger is workflow_call — a callable must NOT declare one; see "Concurrency is the caller's job" above);
  • has a job that does not uses: a reusable workflow and has no timeout-minutes (a job that DOES uses: one is reported as an informational ::notice:: naming the called workflow instead — that job cannot declare timeout-minutes at all, because the called workflow's own jobs own it, and flagging it as a finding was a false positive this repo used to ship in agent-infrastructure's own audit_workflows.py);
  • has any uses: step or job not pinned to a full 40-character commit SHA;
  • is missing permissions: at the workflow level or on any job.

This is a caller-side hygiene check, distinct from what actionlint alone proves (schema/expression validity) and distinct from what lint_callables.py proves about THIS repo's own callables — Caller lint proves the same class of thing about the repository that adopted one.

Dependency audit (dependency-audit, audit-ignore)

Input Type Default Purpose
dependency-audit boolean true Fail the build on a high/critical advisory that has a published fix
audit-ignore string "" Comma-separated GHSA-xxxx-xxxx-xxxx=reason suppressions; the reason is required

The estate security bar (company-hq#745; Logan, askme round 2026-09-17, "Fail on fixable high/critical (Recommended)"; company-hq D-ORG-1 (g), 2026-09-02) is alerts on, and no fixable high or critical advisory in the tree. That is deliberately not what pnpm audit --audit-level=high reports on its own: its exit status goes non-zero for any high/critical finding, including ones upstream has published no patch for. A gate that cannot tell "you have not upgraded" from "there is nothing to upgrade to" is a gate whose only sustainable reaction is || true, and once that lands the fixable advisories stop being caught too.

So the audit command's exit status is discarded on purpose and the JSON report is what decides:

Finding Result
high/critical with a published fix ::error:: — the build job fails
high/critical with no published fix ::warning:: — the build passes
moderate / low / info counted in the job summary, never blocking
report missing, empty, unparseable, or in an unrecognised shape ::error:: — hard failure

That last row is the same fail-closed rule require-scripts exists for: "the audit did not run" must never be indistinguishable from "the audit found nothing".

A pull request only warns (Logan, 2026-09-29, "Main + nightly"). An advisory is published on its own schedule, not the pull request's, so a fixable high/critical finding used to block an unrelated PR the day it appeared. On pull_request and pull_request_target every ::error:: row above is re-labelled ::warning:: (the same advisory list, still in the job summary) and the step passes; the closing line says the finding will fail the default branch. On a push to the default branch, a schedule run and a workflow_dispatch the step fails exactly as before, so a red audit is caught by the merge that follows and by the nightly run (see E2E off the pull request, whose nightly caller is not the place for it: the audit runs in the ci workflow, so add schedule: to that caller too if the audit should be checked nightly).

Both report shapes are parsed, and the package-manager input selects the command, never the parser — npm changed this format once already, and a parser keyed on the input would silently read zero advisories the next time it changes:

  • pnpm / npm 6 — the advisories map. Fixable means patched_versions is a real range rather than the "<0.0.0" no-patch sentinel.
  • npm 7+ (auditReportVersion: 2) — the vulnerabilities map. Fixable means fixAvailable is true or a {name, version, isSemVerMajor} object; a fix that needs a major bump still counts as a fix.

Estate contract pins are not advisories and do not count here — only what the package manager's own audit reports does.

audit-ignore: a suppression must carry its reason
audit-ignore: >-
  GHSA-aaaa-bbbb-cccc=no upstream release yet, tracked in company-hq#812, review 2026-12-01,
  GHSA-dddd-eeee-ffff=unreachable code path behind a disabled flag, review 2026-11-01

Each entry is <id>=<reason>, entries separated by commas. The reason is required: an entry with no =reason fails the gate rather than silently muting an advisory, which mirrors the written-reason convention a Dependabot ignore: block carries. Every suppression is echoed as a ::warning:: with its reason attached, so a muted advisory cannot become invisible tribal knowledge, and an entry that matches nothing in the current report is reported as a stale suppression so it gets removed instead of accumulating. Advisories the report carries no GHSA id for are matched by NPM-<numeric-id>.

Why a step in build and not its own job

The tree it audits is the one build just installed. A standalone lightweight job on the lightweight runner class would cost a second checkout, a second setup-node and a second full install — 60–120s and a second runner slot on this repo's adopters — to re-derive state that already exists in build, for the ~5–10s the audit command itself takes. It never touches the browser pool. Required covers it through build, which it already demands success from; there is no separate result to aggregate.

Turning it off

dependency-audit: false is a bounded remediation, not a setting — the same status require-scripts: false has. Unlike run-tests and foundation-check, this input defaults to true: those two run a caller-specific script that may not exist, while this one reads the lockfile every adopter already has, and it is a security bar rather than an optional lane. It is still a new gate that can turn an existing adopter red, so the v1 tag must not move onto it until the adopters have been checked — see Versioning policy.

E2E off the pull request (mode: e2e and CI_E2E_IN_CI)

Playwright is the slowest lane and it almost never catches a bug that the merge-time run would not (Logan, 2026-09-29: "skip on PRs; run after each merge, newest wins, plus nightly; one org switch"). Two pieces, both additive: with neither in play a caller behaves exactly as before.

1. The org switch: vars.CI_E2E_IN_CI. When the organization (or one repo, which overrides the org value) sets CI_E2E_IN_CI to false (GitHub compares strings case-insensitively, so False and FALSE switch it off too), E2E plan, E2E and E2E quarantine are skipped in an ordinary CI run on every event, and Required reads that skip as success (a lane that runs anyway still fails it). Build then skips packing and uploading the prebuilt application for the E2E jobs, and a checks-in-build caller with no other lane folds Required into Build (one job, no extra queue hop). A pull request that skipped E2E this way mints no Required proof (the proof key does not include the variable, so a reusable proof would let a later push, after the variable is removed or overridden, go green without E2E ever having run). Unset, empty or any other value changes nothing. The switch deliberately does not override two promises to run browsers: e2e-full-paths (a caller that sets it wants protected-path escalation, which needs E2E plan to decide, so acre-oracle keeps browsers on auth/payment changes) and expected-candidate-sha (explicit release validation cannot skip configured browser coverage).

2. The post-merge and nightly run: mode: e2e. With mode: e2e the callable runs only Build -> E2E plan -> E2E (and the quarantine lane) and a Required verdict. Every other lane, the Caller lint and the dependency audit are skipped, the CI_E2E_IN_CI variable is ignored, and the run never path-skips (a newer push cancels this run, so no diff can prove a skip). Required is red if the mode is not ci or e2e, if run-e2e is not true, if journey-smoke-url is set, or if any E2E lane fails, is cancelled, or was skipped without the plan's say-so. The mode has its own workflow because an app's promote.yml fires on workflow_run completion of the whole CI workflow: E2E left in ci.yml would delay every promotion.

Name the caller workflow E2E (the reaper and the red-main listener look for that name) and stamp it exactly, with the same with: values as the app's ci job for the E2E inputs and the same runner inputs:

name: E2E

# Post-merge and nightly Playwright. The newest merge wins: a push cancels the
# run still in flight, and every run tests the whole suite.
on:
  push:
    branches: [main]
  schedule:
    - cron: '17 8 * * *' # 08:17 UTC = 3:17 AM CT
  workflow_dispatch:

concurrency:
  group: e2e-${{ github.repository }}
  cancel-in-progress: true

permissions:
  contents: read

jobs:
  ci:
    permissions:
      contents: read
      packages: read
      actions: read
      pull-requests: write
    uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha> # v2
    secrets: inherit
    with:
      mode: e2e
      run-e2e: true
      # ...the app's ci.yml E2E and runner inputs, verbatim (e2e-runner,
      # e2e-shards, e2e-args, e2e-build-artifact-path, e2e-quarantine-args,
      # install-script, node-version, working-directory, ...)

Format the stamped file with the app's own Prettier config: apps that run format:check over .github/workflows (cloudflarestat-us, single quotes) fail a double-quoted cron.

The calling job id stays ci so the composed context reads E2E / ci / Required. concurrency sits in the caller, never in the callable (R6). Pass the same permissions: block the app's ci job grants: a job in the callable may only use what its caller granted (R12).

A red post-merge or nightly run reaches the same red-main issue flow as CI: add the workflow to the app's red-main listener, and accept schedule.

on:
  workflow_run:
    workflows: ["CI", "E2E"]
    types: [completed]
jobs:
  red-main:
    if: |
      (github.event.workflow_run.event == 'push' || github.event.workflow_run.event == 'workflow_dispatch' || github.event.workflow_run.event == 'schedule') &&
      github.event.workflow_run.head_branch == github.event.repository.default_branch

Each workflow keeps its own "main is red: <name>" issue, opened by the first red run and closed by the next green one. A run cancelled by a newer merge is not a verdict and does nothing.

Quality level (quality-level, quality-opt-out)

The estate has web quality tools that gated nothing on most apps, because each one was an opt-in nobody opted into. quality-level: standard turns them on as one set. It adds no job, no check name and no permission:

Gate What runs Where Blocks on
foundation-check narduk-app foundation:check (the existing foundation-check steps) Build FAIL or UNKNOWN
performance-budget narduk-app performance-budget --json [performance-budget-args] over .output/public Build, after build-script and extra-scripts any violation, a missing build output, an unreadable report
security-headers narduk-app foundation:check:security-headers --base-url <origin> [--path …] Preview on pull requests (the PR's own Cloudflare preview); Journey smoke in post-deploy mode (security-headers-url, else journey-smoke-url) exit 1 (FAIL) or 2 (UNKNOWN)

To adopt, an app adds one line to its ci: job's with: block, in its own next change:

    with:
      quality-level: standard

legacy (the default) runs exactly the gates this workflow ran before the input existed. An app never gets these gates from a pin bump alone.

Why an input, not a default flip. Callers pin a full SHA (25 of 31 nuxt-cloudflare.yml uses: lines across 18 repositories on 2026-09-27; the other 6 are frozen @v1; none use @main). A pin looks like a natural one-app-at-a-time gate, but Dependabot opens pin bumps for this callable on most of those repositories, so a default flip on v2, or in a new major, would arrive as a wave of simultaneously red pull requests. The explicit input is the only switch that moves exactly one app, in a change that app chose. A later major can flip the default once the fleet has adopted.

Opting out. Each standard gate turns off only through quality-opt-out, with a written reason:

      quality-opt-out: >-
        performance-budget=hero video ships in the next release (tracked in the app's issue 42),
        security-headers=no Workers Builds preview until the account cutover

Entries are check=reason, separated by commas, so a reason cannot contain a comma. An entry with no reason, an unknown check name or a duplicate fails Build. Every opt-out is echoed as a ::warning:: with its reason and a job-summary line on every run, so an opted-out gate stays visible. An opt-out at legacy warns that it does nothing. foundation-check: true together with a foundation-check opt-out is a contradiction and fails.

Rules the resolver enforces up front (the Resolve quality gates step, before any install, so misconfiguration fails the adoption pull request itself):

  • security-headers needs a pull-request preview. With preview-checks: none the app either enables a preview check or opts out with its reason.
  • performance-budget needs a build. With build-script: "" the app opts out.
  • performance-budget-args cannot carry --json or --report-only: the workflow owns the report, and the gate never becomes a warning.
  • security-headers-paths entries start with /; security-headers-url is one https:// URL.

Where each gate reads from. performance-budget and both header probes run the caller's own installed @narduk-enterprises/narduk-app-tools (pnpm exec / npx --no-install), so neither needs a credential beyond the install the job already did. The budget runs from working-directory; point it at a workspace app with performance-budget-args: --app-dir apps/web. The probes run from preview-working-directory (else working-directory). The header probe needs a narduk-app-tools release that has foundation:check:security-headers (narduk-libs#360, 2026-09-16). The tool's own budget passes silently when .output/public is absent, so the step checks for the directory itself.

The pull request probes its preview, never production. A change that fixes the headers must be able to go green before it ships. Headers a Cloudflare zone adds on the custom domain, rather than the app itself, are absent on the workers.dev preview; the estate narduk-core security preset sets them in the app. The post-deploy probe in Journey smoke reads production as just deployed, runs whether or not the smoke passed, and never triggers the rollback hook.

Not here: the axe accessibility ratchet. narduk-testkit's expectAccessible is an assertion inside the app's own Playwright suite, which run-e2e already runs. This workflow cannot tell a suite that calls it from one that does not, short of grepping test sources, which would prove presence and not coverage. The app generator and each app's AGENTS.md own that rule.

The unit-test lane and extra-scripts

    with:
      run-tests: true       # opt-in — see below for why the default is false
      test-script: test     # vitest, jest, whatever the repo already runs
      extra-scripts: check:vendor

Until run-tests existed this workflow had no unit-test expression at all, while node-library.yml had one. A Nuxt app with a vitest suite therefore could not adopt its own class's callable without dropping its unit tests — the same defect the browser-shard gap was, and the same consequence: the repo keeps hand-rolling. earthdata-viewer (company-hq#278) is the adopter that surfaced it; its unit job ran npm test and had nowhere to go.

run-tests defaults to false on purpose, even though the script runs --if-present. test is a near-universal package.json script, so a default of true would hand every existing adopter a brand-new gate the instant the moving v1 tag advanced — and a suite that was never in a repo's CI turning its default branch red is not a backward-compatible change, whatever the input is labelled. Existing callers opt in when they mean to.

extra-scripts mirrors node-library.yml's input of the same name and runs after build-script in the same lane, so a gate needing build output doesn't pay for a second install. earthdata-viewer's vendored-package pin check is the first case.

Everything in build delays the prebuilt E2E application, and so every E2E shard, the preview and the deploy dry run. The typecheck and unit-test lanes therefore run in their own parallel checks job, and e2e-plan no longer waits for build. A caller whose runner pool, not its critical path, is the constraint passes checks-in-build: true. The same three steps (YAML aliases, not copies) then run in build before build-script, checks is skipped, and Required demands exactly that. This saves one runner allocation, checkout and install per run, and the E2E shards, preview and deploy dry run start after the checks instead of beside them. An extra script that does not need build output belongs in extra-gate-scripts, which also runs in parallel. On riverstatus, two migration-proof scripts in extra-scripts held its E2E back by about 8 minutes.

Required folded into Build (ci-reset W10, W8 item 3). With checks-in-build: true and no other lane in the run (no extra-gate-scripts, fast-scripts, run-e2e, preview checks, wrangler-dry-run, journey-smoke-url or required-reuse-pr-results), Build is the only lane, so a separate Required job only waited a second runner queue hop to read one result (operator-portal PRs: 1.7 min queued for 0.2 min of work). In exactly that case the build job is named Required and runs Caller lint itself, and the aggregating job skips; its check shows the raw name expression, never Required. The check ci / Required keeps its name, so no ruleset changes, and exactly one started job carries it in every run. Any of the lanes above brings back the usual Build plus aggregating Required. lint_callables.py R5 pins the shape (the same condition on both jobs, checks-in-build in it, the aggregator's if: exactly !cancelled() && !(<condition>)).

concurrent-scripts (opt-in, ci-reset W10, W8 item 2a). Package scripts listed here start in the background just before build-script and are awaited after extra-scripts; each one's output is replayed in its own log group, and a failed, missing (under require-scripts) or killed script fails Build like any other gate. Use it for checks that need the installed tree but not the build output (lint, a vendored-package pin check), to take them off the critical path without a separate job. Two caveats:

  • Memory. The scripts share the build's runner and memory. A Nuxt build plus vue-tsc or ESLint can exceed a 2-vCPU guest's memory; an OOM kill of a background script fails Build ("exited without a status"), and an OOM kill of the build itself fails it too. Prefer extra-gate-scripts (its own job) on small runners.
  • Build outputs. A concurrent script must not read or write what build-script produces (.nuxt, .output, dist). A script with a pre<name> hook is refused, because a prelint: nuxt prepare would rewrite .nuxt under the running build.

require-scripts: a lane that matched no script is not a passing lane

Every gate in this file used to run --if-present, which means a gate whose script does not exist matches nothing, exits 0, and reports a green lane that ran nothing. run-tests: true is a caller asserting tests exist; the callable was taking that assertion on faith.

node-library.yml got the first fix in workflows#14. This file did not, and the missing web:typecheck default became the dominant dead lane across its adopters. typecheck-web-script now defaults to "": a second web typecheck is opt-in, while a non-empty name remains an assertion that the script exists.

Every gate now probes for the script first, and reacts by require-scripts:

require-scripts script missing effect
false (explicit remediation opt-out) named but absent ::warning:: + a job-summary line naming the script; lane still green
true (default) named but absent ::error:: and the job fails
either empty script name lane skipped silently — the caller declared it absent

The default is now true. The migration first declared absent lanes and proved every remaining name; callers retain false only as an explicit, temporary remediation opt-out.

A caller declares absent lanes by passing an empty script name. This is what makes require-scripts usable when some repos genuinely have no build or split typecheck lane — "I have no web surface" remains distinct from "I named a script that isn't there":

    with:
      typecheck-worker-script: typecheck
      typecheck-web-script: ""   # no web surface in this repo
      build-script: ""           # no build step; wrangler dry-run IS the build
      run-tests: true
      require-scripts: true      # now every remaining lane is proven to run

software-delivery is the reference caller for that shape.

The gate text is not tested as a copy: scripts/test_script_gates.py extracts each gate's run: block from the YAML and executes that exact text against real package.json fixtures (54 cases across both Node callables), including colon-bearing names like web:typecheck — the probe resolves scripts through npm pkg get scripts.<name>, a dot-path, and a name that broke that lookup would report every colon-bearing script as missing.

Browser shards and the isolated pool

New browser adopters separate application CI from browser execution. The caller's own build job uploads one same-run build artifact; the standalone browser callable consumes it without rebuilding:

jobs:
  build:
    # ...checkout, install and build the application, then:
    outputs:
      build-artifact-id: ${{ steps.upload.outputs.artifact-id }}
    steps:
      - id: upload
        uses: actions/upload-artifact@<sha> # v7
        with:
          name: e2e-build-${{ github.run_id }}-${{ github.run_attempt }}
          path: apps/web/.output
          include-hidden-files: true
          retention-days: 1

  browser:
    needs: build
    uses: narduk-enterprises/workflows/.github/workflows/reusable-browser-tests.yml@<sha-of-v2> # v2
    with:
      build-artifact-id: ${{ needs.build.outputs.build-artifact-id }}
      linux-runner: '{"group":"linux-ci","labels":["self-hosted","Linux","X64","proxmox","linux-ci"]}'
      browser-runner: '{"group":"playwright-isolated","labels":["self-hosted","Linux","X64","proxmox-playwright-x64"]}'
      working-directory: apps/web
      build-artifact-path: .output
      build-artifact-marker: server/index.mjs
      playwright-version: 1.61.1
      e2e-script: test:e2e:ci
      chromium-args: "--project=web --workers=1"
      shards: 3

That handoff is still a GitHub artifact, so it counts against the organization's Actions artifact storage. nuxt-cloudflare.yml can no longer produce it. Its prebuilt application now goes to the R2 CI artifact store for its own E2E jobs. Before that, the per-call name suffix added in workflows#146 had already stopped matching this callable's default name. Moving this handoff to the CI artifact store is workflows#150.

The two route inputs are not suggestions. Resolve both from the fleet manifest and copy each runsOn object verbatim. A hosted contract job checks the exact group and ordered labels before any caller-controlled self-hosted route is scheduled. It also rejects absolute or escaping artifact paths, unsupported package managers, non-exact Playwright versions, and invalid shard counts. Required is hosted for the same reason: even a bad route still produces a visible failing gate instead of scheduling the failure aggregator on that bad route.

Chromium and opt-in WebKit are the only jobs on proxmox-playwright-x64. Production-build validation runs on linux-ci. The browser jobs declare no secrets and receive only GitHub's short-lived token for checkout, same-run artifact transfer, and optional package reads. The caller must provide workflow-level concurrency because overlapping runs share fixed runner paths; the callable deliberately declares none.

The pool's current defects are treated as fail-closed constraints:

  • agent-infrastructure#323: the root-owned /opt/playwright-ci browser tree is never written and no playwright install fallback exists. Caller pin, installed packages, image package, manifests, executable ancestry, and a real headless launch must all agree.
  • agent-infrastructure#267: the workflow does not alter guest firewall or DNS state, add sleeps, or hide egress failure.
  • agent-infrastructure#248/#237: workflow code never manipulates leases or quarantine. An unavailable guest can leave work queued, but it cannot become a skipped green; every enabled shard result is mandatory in Required.

When the producer is an app-owned build job, export the upload step's artifact-id as a job output and pass it as build-artifact-id. This keeps failed-job reruns tied to the successful producer. Omitting it retains the existing same-run, same-attempt artifact-name convention.

The older nuxt-cloudflare.yml e2e-* inputs remain supported for existing callers and for repos without isolated-pool approval. New isolated-pool adoptions use the standalone callable so the pool contract has one owner.

Reusing build output in E2E

The Nuxt callable reuses build output by default when run-e2e is enabled. It requires exactly one of .output/server/index.mjs or apps/web/.output/server/index.mjs under working-directory. Other layouts name their output explicitly:

      run-e2e: true
      e2e-build-artifact-path: apps/web/.output

The build job publishes that relative path to the CI artifact store (see "The CI artifact store (2026-09-26)") only after its gates pass. It uploads one gzip tarball and hands its object key, SHA-256 and size to the E2E jobs as a job output, with the signature of a presigned GET for that one object. A missing build output fails Build. Hidden files are included because Nuxt's canonical output directory is .output; the input is an explicit caller-selected path, not a repository-wide sweep.

Every E2E shard fetches the tarball through that presigned GET. The shard checks the SHA-256 and size against Build's output, unpacks it to the same path and receives E2E_PREBUILT_ARTIFACT=1. The consumer's launcher must treat that variable as an assertion: validate the expected entry points and fail when any are absent, rather than silently rebuilding. A tarball that does not match Build's digest fails the job and is never unpacked.

Any other miss makes the shard run build-script itself to produce the same path. Misses include a run without the store's credentials (a fork, Dependabot, a caller that does not pass them), an expired link and a storage outage. The shard then fails if the script is missing or leaves no output there, so the suite still tests a built application.

The object key includes the repository, github.run_id, github.run_attempt and a per-call scope (artifact-scope). The reference travels as a job output, so rerunning failed shards within the link's 24 hours reuses the original build even when the attempt number advances. After that they rebuild in-job. The default auto publishes nothing when E2E is disabled. Set an explicitly empty path only for suites that do not test a built application. Launchers must honor E2E_PREBUILT_ARTIFACT=1: the workflow cannot suppress a build hardcoded inside an application script.

Skipping irrelevant changes (e2e-skip-paths)

E2E now skips by default when every changed path is documentation or repository metadata: root Markdown, Markdown under docs/, app/package-root READMEs, agent guidance files, licenses, issue/PR templates, or CODEOWNERS. Runtime Markdown under content/, executable files under docs/, app code, CSS, dependencies, configuration, workflows, tests and unknown paths still run.

The same rule applies to pull requests and forward pushes, including the merge push. Manual, scheduled and merge-group runs always run E2E. A skipped plan emits no shards; Required checks that E2E and its report actually skipped.

e2e-skip-paths replaces the default with repository-root-relative globs (* does not cross directories; ** does). An explicitly empty string disables path skipping. Narrow or disable the list if your app renders those documents; never add CSS or executable documentation to the list. e2e-full-paths takes precedence over a matching skip pattern.

The plan reads GitHub's compare endpoint with contents: read, including both names of renamed files. Empty or failed comparisons, new/deleted branch SHAs, force-push divergence, unsupported filenames and the API's 300-file cap all run E2E. See GitHub's compare API.

Reusing full PR E2E after merge (e2e-reuse-pr-results)

Default-branch pushes first look for a successful PR run of the same Git tree, caller workflow path and callable inputs. This handles squash and merge commits whose SHA differs but whose tested contents are identical. Other CI checks and production delivery still run normally.

Required publishes a proof to the CI artifact store only after every required gate passed and the actual E2E arguments and shard count matched the full suite. A PR subset, docs-only skip, failed shard or fork cannot publish proof. The proof is a small pointer, proof/<owner>/<repo>/<key>.json. It names the PR run and attempt that wrote it. The key digests the tested tree, the caller workflow path and the callable inputs.

The push does not trust the pointer on its own. It reuses the proof only after GitHub's API confirms three things:

  • the run is a completed, successful pull_request run of the same caller workflow;
  • the run's head repository is this repository;
  • the named attempt has a successful Required job whose Publish full E2E proof <key> step, carrying this exact key, succeeded.

An absent, expired (30-day lifecycle), changed or unreadable proof, or one GitHub does not confirm, runs E2E, as does a run without the store's credentials. Scheduled and manual runs always run. Set e2e-reuse-pr-results: false for tests with intentionally different PR/push behavior or external state that must be rechecked after merge. required-reuse-pr-results reuses the whole gate the same way, from the Publish Required proof <key> step.

Permission migration: adopting this revision requires actions: read on all Nuxt callable ci jobs, even if E2E is disabled. GitHub validates the permission ceiling before evaluating job conditions. The proof lookup needs only read access to Actions results to confirm the run a proof names; it gains no write permission. The proof itself lives in the CI artifact store, so pass the three CI_ARTIFACTS_R2_* secrets as well (see the example above). Add this alongside the existing contents: read, packages: read, and pull-requests: write grants before or with the SHA bump. Existing pinned callers do not change. This is a breaking permission change: do not advance v2 over it; publish a new major only after an adopter canary is green.

    permissions:
      contents: read
      packages: read
      pull-requests: write
      actions: read # read the prior PR's E2E proof
    with:
      run-e2e: true
      # Defaults: conservative path skipping and full-PR proof reuse.

A smaller suite on pull requests (e2e-pr-shards, e2e-pr-args)

(workflows#83) e2e-skip-paths above is all-or-nothing: it either runs the whole matrix or none of it. These two inputs are the middle setting — a caller-defined subset on pull requests, the full suite on the default branch.

      run-e2e: true
      e2e-shards: 3                  # push / default branch: unchanged
      e2e-args: ""
      e2e-pr-shards: 1               # pull request: one lane
      e2e-pr-args: "--project=smoke" # pull request: the caller's own project

Why this exists: the browser class now has three active 8 GiB on-prem primary guests (343, 345 and 346), three 4 GiB fsn1-pve01 fallback guests (340–342), and CT 344 is a configured dormant guest rather than an active slot. The old three-effective-slots/seven-declared queue measurements are historical; heavy jobs request memory-8g, while not every browser guest is 8 GiB. Capacity and tiering belong to narduk-enterprises/fleet, not to this callable. See the canonical host inventory for names and placement. The primary/fallback flags do not establish GitHub scheduling priority: normal browser jobs can land on all six active guests; memory-8g matches the three active on-prem guests.

Fewer lanes is not by itself faster — it is fewer slots. Measured on gonogo with e2e-pr-shards: 1 and no e2e-pr-args: the same suite ran 527 / 407 / 292 s in three lanes plus a 32 s report (a 559 s critical path), and 913 s in one. Aggregate pool occupancy dropped 27% (1258 runner-seconds over four jobs to 913 over one) and the run held one isolated slot instead of three, which is the part that shortens every other repository's queue — but the pull request's own wall clock got longer, because the tests were redistributed rather than reduced. e2e-pr-shards alone is a courtesy to the pool. To make your own pull request faster, pair it with an e2e-pr-args subset that runs genuinely fewer tests.

The rules, all of which fail toward running more:

  • Unset is today's behaviour. e2e-pr-shards: 0 and e2e-pr-args: "" are the defaults and mean "no override". Every current adopter passes neither, so they see no change whatsoever.
  • Pull-request events only. push, schedule, workflow_dispatch and anything unrecognised run e2e-shards/e2e-args even when an override is configured — the same confinement e2e-skip-paths has, and for the same reason: the default branch is the canonical validation, and neither input may weaken it.
  • e2e-pr-args replaces, it does not append. A pull-request subset cannot silently inherit a conflicting --project from e2e-args.
  • Empty means inherit. A caller that wants arguments on a push and none on a pull request cannot express that here; put a github.event_name expression in the caller's own with: block instead. A sentinel for "explicitly empty" would be a second, weaker way to say the same thing.
  • The subset is the caller's to define. This workflow never guesses what "smoke" means — it passes your arguments to your e2e-script. Tag the subset in your own playwright.config (a project) or with --grep.

Required is unaffected as a gate: it still demands E2E succeed, and it checks the same single resolved shard count E2E sharded on.

A subset is a smaller gate, not a weaker one. Whatever a pull request stops running, the default-branch push still runs — but it runs it after the merge. Choose the subset so a failure it cannot catch is one you are willing to find on main.

The full suite for risky paths (e2e-full-paths)

e2e-full-paths is the escape hatch from the pull-request subset: a space-separated list of globs (the same syntax as e2e-skip-paths). When any file a pull request changes matches one, that pull request runs the full e2e-shards/e2e-args instead of e2e-pr-shards/e2e-pr-args.

    with:
      e2e-pr-shards: 1
      e2e-pr-args: "--project=smoke"
      e2e-full-paths: "server/database/** drizzle/** playwright.config.ts"

The use is a fast pull-request smoke by default, with the whole suite reserved for the paths whose breakage the smoke tier cannot see — schema, auth, routing, the Playwright config itself. The rules:

  • Unset is today's behaviour. An empty e2e-full-paths (the default) never forces the full suite.
  • It beats e2e-skip-paths. A file matching both lists runs the full suite; it is never skipped.
  • Unknown means full, for an opted-in caller. When the changed files cannot be listed (no base/head SHA, no gh, a compare API error, an empty list, or the 300-file compare cap), a caller that set e2e-full-paths gets the full suite. A caller that did not keeps its pull-request tier.
  • Full-tier selection is for pull requests. Pushes use the full tier when neither path skipping nor equivalent PR proof applies.
  • The E2E plan job summary lists which changed files forced the full run.
  • e2e-full-paths matches case-insensitively (ASCII only). *auth* matches useAuth.ts, AuthPanel.vue and OAuthCallback.ts; *session* matches useSession.ts. * still does not cross /. An empty e2e-full-paths still never escalates. e2e-skip-paths stays case-sensitive: each list folds case only in the direction that adds proof, so the default skip list skips exactly what it did before and an oddly cased README.MD runs E2E.

What ci / Fast reports when a protected path matches

The check name stays Fast in both cases, so the required context ci / Fast does not change. What differs is the sibling check in the same workflow run:

Run Check named Fast Other Fast checks
Plain the fast job (lint and unit scripts); on a pull request of a caller with e2e-full-paths (or an exact candidate), fast-escalable none (fast-escalated is skipped)
Escalated the fast-escalated job (the full Required gate) Fast lanes (escalated) (fast-escalable: the lint and unit scripts)

fast and fast-escalable run one step list (a YAML alias) on disjoint domains (ci-reset W10, W8 item 4). Only a pull request of a caller that set e2e-full-paths, or an exact candidate, can escalate, so only there does the lane wait for E2E plan; everywhere else fast needs only Reuse plan and starts with the run (riverstatus main 36497919511: Fast queued 12.0 min behind a 0.1 min plan). On a push, plain fast no longer fails closed on a failed E2E plan (its full output never applies to Fast off a pull request); Required still demands the plan succeed.

So a readiness check that needs to know whether a ci / Fast run over 180 seconds was the full gate looks for a ci / Fast lanes (escalated) check run in the same check suite. Reading check-run names needs no extra permission, so callers grant nothing new for this. For people reading the run, the job holding Fast also writes a ### Fast escalation step-summary section with the line escalated: true or escalated: false; that step never fails the job. A Fast job that does not apply is skipped and takes no runner: both jobs when fast-scripts is empty, in reused runs and in journey-smoke mode, and fast-escalated on every run that does not escalate (CI reset, 2026-09-28; until then both started as no-op jobs, which cost two runner allocations per call). GitHub never evaluates a skipped job's name, so such a check shows the raw name expression (for example (inputs.fast-scripts != '' && ...) && 'Fast' || 'Fast (escalation not needed)'), never Fast, and cannot satisfy a required ci / Fast check. A readiness check should match ci / Fast lanes (escalated) exactly, not as a substring: a skipped fast job's raw expression contains that text. Required demands each Fast job succeed when it applies and be skipped when it does not. The raw name cannot be made readable: a skipped job's name is its literal template whatever contexts it reads (inputs, github and vars included, probed 2026-09-29), and a static Fast would be a passing ci / Fast.

ESLint cache in the Fast lane (eslint-cache, default true)

The Fast lane (fast and fast-escalable) runs any eslint CLI process its fast-scripts start with --cache --cache-strategy content, so unchanged files are not linted again. No caller change is needed: a preload (NODE_OPTIONS=--require, set for the Run fast scripts step only) appends the flags to the process whose entry script is eslint's own bin/eslint.js, however it was reached (pnpm --filter web run lint, turbo run lint, a wrapper script), and leaves everything else alone: other programs, eslint --print-config, --version, and an eslint that already passes --cache* flags. (Appending -- --cache to the script does not work: the scripts are wrappers, and turbo run lint --cache would be turbo's own flag.) narduk-lint (narduk-libs' lint-budget wrapper) is not covered: it takes no --cache-strategy.

The cache is ../.ci-cache/eslint/<lockfile hash> beside the checkout. A persistent self-hosted runner keeps it between runs. GitHub-hosted runners restore it everywhere and save it only on the default branch (the dependency cache rule). A changed lockfile starts cold.

Only Fast gets it. Build, Checks, Extra gate and the escalated Fast lint cold: ESLint's cache is keyed on a file's own content and config, so a type change in another file does not re-lint a cached file, and a type-aware rule (@typescript-eslint/no-floating-promises) can pass an unchanged file that a changed file just broke. The cold full gate is the net for that. Set eslint-cache: false to run Fast lint cold too. scripts/test_eslint_cache.py executes the shipped step and preload.

A non-blocking quarantine lane (e2e-quarantine-args)

A flaky test is taken out of the gate by tagging it (for example @quarantine) and excluding that tag from the caller's gating Playwright projects. It still needs somewhere to run, or it can never show it is fixed. e2e-quarantine-args is that place:

    with:
      e2e-quarantine-args: "--project=quarantine --retries=0"

When set, an extra E2E (quarantine) job runs e2e-script with exactly these arguments, on every event E2E runs on:

  • It cannot fail the gate. The job is continue-on-error: true, Required does not list it, and no job needs: it. A red quarantine run shows as a warning and a job-summary line. It never turns ci / Required or the caller's run red. lint_callables.py enforces this through its NON_GATING_JOBS exemption, which is the only job allowed outside R5.
  • It is the gate's setup. It runs E2E's own steps through a YAML alias: the same prebuilt application, runner route, toolchain checks and auth cleanup. Only the arguments differ. It is unsharded.
  • Its evidence is separate. Its report is in its own job log, apart from the gate's shards.
  • It skips with E2E. A docs-only PR skipped by e2e-skip-paths runs neither.

The job's history on the default branch is the "N consecutive green runs" record a test needs to leave quarantine. Pass --retries=0 so a retry can't hide a flake. Empty (the default) adds no job.

Preview checks (preview-checks, preview-url-source)

Input Type Default Purpose
preview-checks string og none, og, e2e-subset, or og,e2e-subset
preview-url-source string pr-comment pr-comment or url-template
preview-url-template string "" URL template for url-template; {branch}, {branch-alias}, {sha}
preview-timeout-minutes number 20 Whole-job budget; the bounded wait is this minus five minutes
preview-working-directory string "" Directory the preview checks run from; empty uses working-directory

This is a BREAKING interface change. The preview job requests pull-requests: write, and a called workflow asking for a permission its caller did not grant kills the caller's entire run at startup — startup_failure, zero jobs, no logs, no annotation (workflows#59, which took down every @v1 adopter on 2026-09-04). Every adopter must add pull-requests: write to its ci: job before the v1 tag moves onto this commit, and an adopter whose repository has no Workers Builds preview must also pass preview-checks: none in the same change:

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/nuxt-cloudflare.yml@<sha-of-v2> # v2
    permissions:
      contents: read
      packages: read
      pull-requests: write # sticky preview comment
What Cloudflare actually exposes, and what this workflow therefore reads

Workers Builds' GitHub App creates check runs, commit statuses and a pull request comment, and "a preview URL will be provided for any builds which perform wrangler versions upload" (Cloudflare: Workers Builds GitHub integration). It does not create a GitHub Deployment — that is the Pages integration's shape — so there is no environment_url to read, and this workflow does not offer a source that pretends there is. Adding one would also have cost every adopter a deployments: read grant to carry dead code.

That leaves two honest sources:

  • pr-comment (default) — read the pull request's comments and take the first *.workers.dev URL posted by an author whose login contains cloudflare. A workers.dev link posted by anyone else is ignored.
  • url-template — derive the URL locally, with no GitHub read at all. Cloudflare's aliased preview URLs are <ALIAS>-<WORKER_NAME>.<SUBDOMAIN>.workers.dev, so a typical template is https://{branch-alias}-myworker.myaccount.workers.dev. {branch-alias} is the head ref lowercased with every character outside [a-z0-9] replaced by -. That reproduces the transform this estate has observed (buoys' branch codex/buoys-complete → codex-buoys-complete-buoys.narduk-enterprises.workers.dev); Cloudflare does not publish the truncation rule it applies to long branch names, so a repository with long branch names should stay on pr-comment, which reads the URL Cloudflare actually minted rather than predicting it.
The URL is not the proof — x-build-version is

A preview link in a comment says a build was attempted. The wait is not satisfied until the URL itself answers and its x-build-version header is a prefix of the pull request's head SHA. That header is the estate's existing live-proof convention (buoys docs/workers-builds.md records x-build-version: a84fa2163903 for commit a84fa216390321…), and it is compared by prefix because Cloudflare emits a 12-character short SHA while git rev-parse --short defaults to 7. A fixed-width comparison would be a gate that never passes, and a gate that never passes is a gate somebody deletes.

Each round of the bounded wait is one comment read and one HTTP read — no sleep loop that can outlive the job, and the wait is preview-timeout-minutes minus five so the checks and the comment still have room after it.

Fail closed

A preview that never becomes ready inside the bound is a FAILURE, not a skip, and so is one that answers 4xx/5xx, carries no x-build-version, or serves a different commit. "The preview never showed up" is the single most common way a Workers Builds connection silently breaks, and it is indistinguishable from "this repository does not use previews" only if you refuse to make the caller say which it is. That is what preview-checks: none is for.

The lane runs on pull_request / pull_request_target events only. A push to the default branch has no pull-request preview to check, and Required expects the job skipped there — a preview job that runs on a push is a failure too.

The checks
  • og — narduk-app og:check --live --base-url <preview>, run from the caller's own installed @narduk-enterprises/narduk-app-tools. Unlike foundation-check, there is no pinned dlx fallback here: preview-checks: og is a caller asserting it has the tool, and a second resolution path would mean a second credential path and a second version to keep in step.

  • e2e-subset — runs e2e-script with e2e-pr-args (the same pull-request argument set e2e-pr-shards/e2e-pr-args introduced in workflows#83) against the preview with PLAYWRIGHT_BASE_URL set. An empty e2e-pr-args is a hard failure rather than a silent full-suite run against a shared preview.

    The caller's Playwright config must honour PLAYWRIGHT_BASE_URL and must not start its own webServer, or the suite will test localhost and report green. buoys cannot use this yet for exactly that reason — its playwright.config.ts has only dev-server and prebuilt-Worker modes and "no fixture accepts an override base URL" (docs/e2e-testing.md).

Runner class

Lightweight (the same route as Required and E2E plan) unless e2e-subset is selected, which needs a browser guest and therefore the e2e-runner route. On that route the lane runs the same YAML nodes as the e2e job's isolated-route guard, image-equality assertion and browser installer — shared by anchor, not copied, because an image-equality gate that exists twice is an image-equality gate that drifts. lint_callables.py R11 now holds every job routed to e2e-runner to that preflight, not just e2e.

Note that with e2e-subset selected the bounded wait holds a playwright-isolated slot (three effective slots across roughly ten repos — see "A smaller suite on pull requests"). That is why preview-timeout-minutes is a caller input and why og is the default.

One sticky comment

The lane posts a single comment carrying the URL, the build version, the head SHA and each check's result, marked with <!-- narduk-ci:preview -->. Later runs find it by marker and edit it; they never append. The same table also goes to the job summary, so the URL survives even when the comment API call does not — and a failure to post warns rather than turning an otherwise passing lane red.

The isolated Playwright toolchain gate

playwright-isolated is download-free and image-owned. The current pool pins Playwright 1.61.1: Chromium headless shell revision 1228 (Chromium 149.0.7827.55) and WebKit revision 2311 (WebKit 26.5). The image keeps the package, browsers.json, and browser payloads root-owned/read-only under /opt/playwright-ci; the ephemeral guest exports a job-visible symlink tree through PLAYWRIGHT_BROWSERS_PATH.

Before a shard starts its suite, the callable now fails unless all of these are true:

  • the caller directly pins @playwright/test to an exact version (no caret, tilde, tag, or range);
  • that pin, the installed @playwright/test, installed playwright-core, and image Playwright version are identical;
  • the caller and image browsers.json SHA-256 values match, and each requested engine's revision/upstream version matches;
  • the selected executable exists, resolves into the immutable image browser tree, has root-owned non-writable ancestry, and passes a real headless launch canary;
  • PLAYWRIGHT_BROWSERS_PATH is the absolute /opt path exported by the guest, never a workspace or $RUNNER_TEMP cache.

The isolated-route guard runs before dependency installation. It rejects e2e-install-browsers: true and every e2e-browsers-path override, while the installer step independently excludes both the playwright-isolated group and proxmox-playwright-x64 label. A mismatch names consumer/image versions, manifest digests, browser revisions, and the upgrade choice, then exits non-zero without invoking an installer. Hosted browser jobs may still opt into the explicit installer because they own their toolchain rather than consuming this pool.

docs-governance.yml

Path-filtering is the caller's job — this workflow doesn't know the caller's repo layout:

name: handbook-spine-check

on:
  pull_request:
    paths: [docs/**, scripts/check-handbook-spine.py]
  push:
    branches: [main]
    paths: [docs/**, scripts/check-handbook-spine.py]
  workflow_dispatch:

concurrency:
  group: handbook-spine-check-${{ github.event.pull_request.number || github.ref }}
  cancel-in-progress: true

jobs:
  ci:
    uses: narduk-enterprises/workflows/.github/workflows/docs-governance.yml@<sha-of-v2> # v2
    with:
      command: python3 scripts/check-handbook-spine.py
      fetch-depth: 1

command can be multi-line to install its own light dependencies first — for example, a PyYAML-consuming script (like company-hq's untangle-project-sync.yml, which is not itself a docs-governance.yml caller today, but shares its shape):

    with:
      command: |
        python3 -m pip install --disable-pip-version-check --no-cache-dir "pyyaml==6.*"
        python3 untangle/sync-to-project.py --project-number 1 --dry-run

design-ledger.yml

A product keeps one design ledger beside each Claude Design canvas (design/<canvas>/ledger.json, agent-infrastructure#1804). It maps each canvas screen to the code it specifies and records when the two last matched. This callable reads drift in both directions: not-built (the canvas moved), design-stale (the code moved) and diverged (both moved).

  • mode: check runs dc_ledger.py status <ledger> --check --github-summary. It exits 1 while any entry is flagged, prints one fix line per entry, and appends the table to the job summary.
  • mode: flag runs dc_ledger.py flag <ledger> with the caller's github.token. It opens, edits, reopens or closes one Design drift: <canvas> issue on the ledger's issue.repo, whose labels must exist. issue.repo must be the calling repository: the token can write no other, so the checker errors when it differs from GITHUB_REPOSITORY, and when issue.repo is missing.
  • flag never runs on a pull request. It runs only on push, schedule or workflow_dispatch (v2.1.1). On any other event the flag job is skipped and the check job fails the run with an error, so a pull request's code never writes issues.
  • Commit the boards. The callable never runs the ledger's build command, so the canvas boards the ledger's project points at must be in git. A gitignored, generated project reads every entry as "board or screen missing", which stays red once anything is built. status prints a hint when project is absent from the checkout.
  • A malformed ledger is an error (exit 2, naming the entry): code must be a list of paths, gate.state exactly open or cleared, built null or all three hashes. Ledger text is treated as untrusted: it cannot start a workflow command in the log, break a summary or issue table, or mention anyone.
  • Advisory. There is no Required job, so never add <job> / check to branch protection, and do not name the calling job ci.
  • Grant contents: read and issues: write in both modes. GitHub checks the skipped flag job's permissions too, so a check-only caller that grants only contents: read ends in startup_failure (proven on workflows#155). The check job itself runs with contents: read.
  • The checker comes from this repo at job.workflow_sha, the commit the caller pinned, via a sparse checkout without credentials. A pull request cannot change the code that judges it, and the ledger's build command is never run.
  • scripts/dc_ledger.py is a byte copy. The canonical file is agent-infrastructure's skills/claude-design-ops/scripts/dc_ledger.py, and its scripts/check-dc-ledger-parity fails when the two differ. To change it: open the PR here with the new copy; land the canonical change in agent-infrastructure first, its pin at this PR's head commit (the Cursor review here reads agent-infrastructure main as the canonical); merge this PR and tag it; then move agent-infrastructure's pin to the merge commit.
name: design-ledger

on:
  pull_request:
    paths: [design/<canvas>/**, apps/web/app/**]   # the ledger's mapped paths
  push:
    branches: [main]

permissions:
  contents: read

jobs:
  design-ledger:
    uses: narduk-enterprises/workflows/.github/workflows/design-ledger.yml@<sha-of-v2.1.1> # v2.1.1
    permissions:
      contents: read
      issues: write
    with:
      ledger: design/<canvas>/ledger.json
      mode: ${{ github.event_name == 'push' && 'flag' || 'check' }}

runs-on follows the Default route: leave it empty and a private caller lands on linux-ci.

Weekly drift check (retired 2026-07-26)

reusable-weekly-drift-check.yml was removed after a fresh organization-wide caller search found no live consumer (workflows#20, Actions-optimization audit). narduk-template-smoke-app is disabled and documentation references were not runtime callers.

Generic Node CI (retired 2026-09-24)

reusable-node-ci.yml was removed in narduk-reboot P3-C2 / O-D8: a fresh organization-wide code search across narduk-enterprises, narduk-incubator, and narduk-enterprises-clients (2026-09-24, gh search code / gh api search/code, restricted to path:.github/workflows) found zero live callers — the same result the 2026-07-27 search recorded when node-library.yml shipped as its richer, preferred replacement (workflows#14, workflows#16). node-library.yml is a strict superset for library/cli surfaces: the Required aggregator job and the optional package-matrix input (see "The ci / Required convention" above) plus a runner input that accepts the JSON array/object form, not only a plain string. New callers use node-library.yml; there is no migration to perform because there was no live caller to migrate.

Historical advisory code review (retired 2026-09-19, file deleted 2026-09-24)

code-review.yml asked the estate's now-retired ephemeral agent pool for one read-only review of a pull request head, via a repository_dispatch at agent-infrastructure; it was never a CI gate and had no Required job. The pool was retired 2026-09-19 with an empty allowlist, and adoption moved to cursor-review.yml (Refs narduk-enterprises/agent-infrastructure#333, D-AGENT-POOL-1), itself retired 2026-09-28 for the CT650 PR review bot (see "Cursor review (retired 2026-09-28)"). A 2026-09-24 re-verify (gh search code / gh api search/code across narduk-enterprises, narduk-incubator and narduk-enterprises-clients) found zero live pool adopters, matching the state this section already described, so the file was deleted in narduk-reboot P3-C2 / O-D8 rather than kept as a dead compatibility surface.

Versioning policy

  • Callers pin full commit SHAs, never @main. @main is how the last outage happened; it is not a supported reference. Caller lint (in Required) and narduk-app-tools foundation item 5.1 both reject a bare @v2 tag, so a caller pins the SHA the tag points to and names the tag in a comment: uses: …/<workflow>.yml@<40-char sha> # v2. Existing @v1 pins still pass while v1 is frozen.
  • Maintainers cut tags. vN major tags (v1, v2, …) move forward only for backward-compatible changes within that major; breaking changes (renamed or newly-required inputs, removed jobs, changed secret names) get a new major.
  • To adopt a fix, callers bump their pinned tag/SHA deliberately — nothing changes under them silently unless they chose a moving vN major tag.
  • v1 is a moving major tag and is advanced by hand after a merge, never as a side effect of merging. The CI-5 phase 2 files shipped ahead of the tag and v1 was advanced to cover them afterwards; apple.yml, python-data.yml and node-library.yml's two new optional inputs are backward-compatible additions to the same major, so v1 moves again rather than a v2 being cut. An adopter merged before the tag moves must pin the exact commit SHA and switch to @v1 once the tag covers it.
  • v1 is frozen; adopt v2 when touched. See the section below.
  • A fan-out canary precedes an advance: before moving the tag, trigger or find at least one adopter run on the new commit and confirm it resolves and succeeds, per Logan, 2026-09-04: "Fan-out canary before the tag advances (Recommended)" (Refs company-hq#536).
  • The empty-runner default flip (row 18 Q3) must not reach v1 until its unlisted private callers are fixed. It changes no input name, job, check context or permission, so it stays within v1, but it moves every private caller that passes nothing onto the selected-visibility linux-ci group. At the flip, the @v1 callers passing nothing were coding-standards (docs-governance.yml) and x-event-recap (node-library.yml) — both outside that group — and package-delivery and software-delivery on nuxt-cloudflare.yml (whose flip ships separately). Before v1 advances over this change: add each such repo to the fleet manifest's linux-ci group, or have it pass '"ubuntu-latest"' explicitly (mandatory for a §2 hosted exception such as package-delivery); then run the fan-out canary. SHA-pinned callers pick the change up only when they bump their pin.
  • Adding a workflow, or adding an optional input with a default, is within-major. Renaming or newly requiring an input, removing a job, renaming a job (which renames the composed check context and silently orphans every branch-protection rule that required it), or changing a secret name is a new major.
  • A new job-level permissions: scope is a BREAKING change, not an addition. A job in a called workflow may only request permissions the caller granted on its uses: job; ask for one it did not and GitHub fails the caller's entire run at startup — startup_failure, jobs: [], no logs, no check-run annotation — before any job is created. Nothing in the adopter's repository changed, so it reads as an infrastructure outage rather than an interface break, and it lands on every @v1 adopter simultaneously the moment the tag moves. On 2026-09-04 pull-requests: read on nuxt-cloudflare.yml's E2E plan job did exactly that to harvest-tracker, marketing-web, vtraceroute and hydrogen, while the two repos pinned to older commit SHAs kept running (workflows#59). Each callable's permitted set is declared in scripts/lint_callables.py's CALLER_GRANTS and enforced as R12; widening one means updating every adopter's permissions: block first, then the map, and only then moving the tag.
  • nuxt-cloudflare.yml's browser-shard inputs (e2e-runner, e2e-shards, e2e-args, e2e-install-browsers, e2e-browsers-path) and its two new jobs are within-major on the same rule: five optional inputs whose defaults reproduce the previous behaviour, plus added jobs. Adding a job is not a breaking change here specifically because the composed context comes from the caller's job id and this workflow's Required job — neither of which moved. v1 moved again rather than a v2 being cut.
  • nuxt-cloudflare.yml's pull-request subset inputs (e2e-pr-shards, e2e-pr-args, workflows#83) are within-major for the same reason: both are optional, both default to an UNSET sentinel (0 / ""), and with them unset every event resolves the shard count and argument list exactly as before. No job, check name, or Required expectation changed — the effective shard count simply moved from inputs.e2e-shards to an E2E plan output that equals it whenever no override is supplied. e2e-full-paths is within-major on the same rule: optional, default "", and with it unset every event plans exactly as before. e2e-quarantine-args is too: optional, default "", and when unset its job is skipped. The job it adds is never needs:-ed and never gates, so no check context Required reads changes.
  • reusable-browser-tests.yml is a new callable, so its required route, artifact, and exact-version inputs do not break an existing caller. python-data.yml's run-pyright, exact pyright-version, and pyright-args inputs are optional additions; existing Python callers keep their prior behavior until they opt into static analysis. reusable-browser-tests.yml's later addition of the optional NARDUK_PLATFORM_GH_PACKAGES_READ secret (required: false) is within-major on the same rule as a new optional input: a caller that passes nothing gets byte-identical behavior to before the secret existed (workflows#50).
  • nuxt-cloudflare.yml's run-tests / test-script / extra-scripts are within-major on the same rule — three optional inputs, no new job, one conditional step each. run-tests defaults to false precisely so that it is within-major: test is a near-universal package script, so defaulting it on would have made a moving v1 tag introduce a gate to callers who never asked for one, which is a breaking change dressed as an additive input. Getting the default wrong is how an "additive" change breaks people.
  • nuxt-cloudflare.yml's mode input and vars.CI_E2E_IN_CI switch (E2E off the pull request) are within-major on the same rule: one optional input defaulting to ci, no new job, no new permission, and an unset variable reproduces every adopter's behaviour exactly. The Dependency audit step's pull-request downgrade (warn on pull_request, block on push, schedule and dispatch) only loosens a gate on one event, so no adopter turns red because of it.
  • nuxt-cloudflare.yml's foundation-check / foundation-check-tool-version (agent-infrastructure docs/standards/WEB-FOUNDATION-CHECK.md, D-WEBFOUND-2 Q5/Q9 (a), D-WEBFOUND-3) are within-major on the same rule — two optional inputs, no new job, three added steps inside the existing build job, Required's needs: graph unchanged. foundation-check defaults to false for the same reason run-tests does: the web-foundation program is a multi-wave fleet migration (D-WEBFOUND-2 Q4/Q10), most fleet apps do not conform to the seven-item contract yet, and a moving v1 tag must not hand every existing adopter a brand-new red gate the day the tag advances.
  • nuxt-cloudflare.yml's quality-level, quality-opt-out, performance-budget-args, security-headers-paths and security-headers-url are within-major on the same rule: five optional inputs, no new job, no new permission, Required's needs: graph unchanged. quality-level defaults to legacy, which resolves every new gate off, so moving v2 over it changes nothing for an app that did not write quality-level: standard. The gates are default-on only inside standard. See Quality level for why this is an input rather than a default flip (Dependabot pin bumps would otherwise turn the fleet red at once).
  • foundation-check-auth defaults to package-token. nvault remains only for callers whose legacy NARDUK_PLATFORM_GH_PACKAGES_READ mapping carries an nVault service token; new callers use the canonical package PAT there and map the service token separately as NVAULT_TOKEN for install-script. The callable resolves the existing package-read grant for that legacy one step and supplies NODE_AUTH_TOKEN to both the installed checker and the optional pinned download. The checker needs this credential for its live N-1 registry lookup even when dependencies are already installed, unless the caller's project .npmrc routes @narduk-enterprises to npm.nard.uk (workflows#109): then no credential is resolved or exported. Neither the service token nor the resolved package token is exported to later steps; temporary npm configuration contains only an environment-variable reference. Missing or rejected credentials leave a blocking UNKNOWN artifact rather than falling back to github.token.
  • The D-CI-CAP-1 (c) Blacksmith-overflow change (see "Blacksmith overflow" above) is within-major on the same rule: no new input, no new job, no new job-level permissions:, and the default (vars.BLACKSMITH_RUNNERS_ENABLED undeclared) reproduces every existing adopter's behavior byte-for-byte. v1 moves again rather than a v2 being cut, behind the usual fan-out canary.

v1 is frozen; adopt v2 when touched

Logan, 2026-09-18 (askme, 13:41 CT): "Moving v2 tag; repos move when touched (Recommended)".

  • v1 stays at f8e3cc6 and does not move again. Everything after it is on the v2 line. v2.0.0 (a29cd16, the pull-request preview gate) is a breaking change for nuxt-cloudflare.yml: its preview job requests pull-requests: write, and a caller that does not grant it hits startup_failure on its whole run (workflows#59). Moving v1 over it would have broken every @v1 nuxt-cloudflare caller at once. On 2026-09-18 none of the seven (hydrogen, marketing-web, my-farm, narduk-nvr, package-delivery, software-delivery, vtraceroute) granted it.

  • v2 is the moving major tag, advanced by hand behind the fan-out canary, the same way v1 was.

  • A repo moves to v2 in the next PR that works on it, not in a sweep. Make these changes in that same PR:

    1. every uses: …/<workflow>.yml@v1 becomes the full SHA v2 points to, with a # v2 comment (git ls-remote https://github.com/narduk-enterprises/workflows refs/tags/v2; 3cc8c736d07aa821daeb417e7f6682ecd3935aa8 on 2026-09-18). A bare @v2 fails caller-lint ("is not pinned to a full 40-character commit SHA") and foundation item 5.1;
    2. a nuxt-cloudflare.yml caller adds pull-requests: write to its ci: job's permissions:, and passes preview-checks: none if the repo has no Workers Builds preview (see the preview gate);
    3. a private caller that passes no runner now lands on the linux-ci organization group (Default route). Confirm the repo is in that group. If it is not, it queues forever rather than failing. Otherwise pass runner: explicitly (for example '"ubuntu-latest"' for a named CI-RUNNER-POLICY §2 hosted exception);
    4. nuxt-cloudflare.yml's dependency-audit defaults to true on v2 and fails on fixable high/critical advisories. Fix them, or pass dependency-audit: false with its written reason.

    The PR's own CI run on the v2 SHA is the proof it moved cleanly.

Maintainer conventions

  • .github/workflows/ci.yml gates this repo (~7s). actionlint + scripts/lint_callables.py + behavior tests including scripts/test_playwright_toolchain.py. The structural gate enforces every convention in this list, so none of them can regress silently: see the rule table (R1–R11) at the top of scripts/lint_callables.py. Run it locally before pushing: python3 scripts/lint_callables.py.
  • Third-party and first-party actions are pinned to full commit SHAs with a version comment, targeting the current Actions Node runtime (enforced: R4). All nine pins currently resolve to their claimed tags and every one runs on node24. astral-sh/setup-uv is deliberately held at v8.3.2 rather than v9.0.0: v9's sole breaking change flips prune-cache to false, which would grow cache usage for every adopter, and v8.3.2 is already on the current runtime — so the bump would be a behaviour change with no hardening benefit.
  • Jobs declare minimal permissions and explicit timeout-minutes (enforced: R1/R2/R3).

Timeout basis

Every job has had a finite timeout-minutes since the day it was written — the guard is not new. What was missing was a basis. Measured 2026-07-25 from the jobs endpoint, counting only jobs the callable actually composed (a ci / … or drift-check / … name), execution time only with queue excluded, never extrapolated from a run count.

That filter is the whole measurement, not a detail. A first pass that matched on the bare job name mixed each adopter's pre-adoption local job into the same bucket and reported python-data / test at p95 444s / max 722s. Split correctly, the callable's ci / test is p95 90s / max 91s and the 722s belongs to the local test job narduk-data ran before it adopted — a 5× error, in the direction that would have made a fine timeout look nearly breached.

Callable Job Timeout p50 p95 max n repos ×p95
apple.yml xcode 45 (xcode-timeout-minutes) 116s 228s 231s 8 4 11.8×
apple.yml lint 15 (lint-timeout-minutes) 8s 12s 12s 4 2 75.9×
docs-governance.yml check 15 10s 18s 22s 39 1 49.7×
node-library.yml package / <label> 30 78s 289s 298s 17 4 6.2×
nuxt-cloudflare.yml Build 30 72s 131s 145s 16 4 13.8×
nuxt-cloudflare.yml Deploy dry run 15 18s 20s 20s 4 1 45.7×
python-data.yml test 30 (test-timeout-minutes) 75s 90s 91s 5 1 20.0×
reusable-weekly-drift-check.yml Typecheck 15 69s 69s 69s 1 1 13.0×
reusable-weekly-drift-check.yml Unit Tests 15 50s 50s 50s 1 1 18.0×
reusable-weekly-drift-check.yml Template Drift Check 10 45s 45s 45s 1 1 13.3×
(all five) Required 5 4–5s 5–7s 7s 83 11 43–65×

Jobs with no data at all — their timeouts are declared, finite, and unmeasured. Do not read the numbers above onto them:

Callable Job Timeout Why nothing ran
nuxt-cloudflare.yml E2E, E2E plan 30 / 5 never executed on any adopter. hydrogen and software-delivery set run-e2e: false; marketing-web and vtraceroute leave it at the default. 35 skipped instances, 0 runs
python-data.yml lint 10 narduk-data leaves run-ruff false — skipped in every run

Nothing was changed as a result. Every measured timeout sits between 6.2× and 76× its observed p95, so none is close to producing a false red. The one job near the 4–6× target band is node-library / package (6.2×), which is correct as-is. The rest are looser than a band would suggest, deliberately:

  • The sample is tiny and 24 hours old. Composed jobs first appear 2026-07-24T21:54Z and most adopters landed the next day. Only docs-governance / check (n=39) is a distribution; below n≈10 a percentile is one observation wearing a hat.
  • For a short job the floor is not p95. It is "long enough that a cold cache or a slow guest is not a false red", which is minutes regardless of a 12s p95.
  • apple / xcode at 11.8× is the loosest that matters, because it holds the estate's single Mac slot. It stays: n=8 across four small Swift repos is far too thin to justify tightening a real iOS archive toward ~20 minutes, and it is already a caller-tunable input. An adopter that knows its build should set it.

The regression guard is rule R2 in scripts/lint_callables.py, which makes it impossible to add a job without a finite timeout — including via an input whose numeric default was removed.

timeout-minutes measures execution, never queue — and here that gap is enormous. docs-governance / check executes in 10s and has waited 1973s (33 min) for a linux-ci runner; its Required job has waited 1044s. Anyone sizing a timeout from a run's wall-clock duration would set it wildly wrong.

When adding a job, size its timeout from the same place — the jobs endpoint, per job, never extrapolated from a run count.

  • This is a public repository. Its own gate and public callers use GitHub-hosted runners; private callers retain reusable-workflow compatibility and may use manifest-routed self-hosted runners only where policy permits.
  • New reusable workflows follow both estate-wide conventions added by CI-5 phase 2: the workflow's last job is named exactly Required and needs: everything else (see above), and runs-on: decodes the effective route (inputs.runner, else the visibility-gated default) from a JSON-encoded runner input defaulting to '' (see "Default route" above; add the new callable to scripts/test_runner_default.py).
  • Estate conventions live in the ci-workflow-author skill (agent-infrastructure repo); consult it before adding workflows here.

The CI artifact store (2026-09-26)

nuxt-cloudflare.yml neither uploads nor downloads a GitHub artifact, and reusable-browser-tests.yml uploads none. On 2026-09-26 the org's Actions artifact storage passed its included allowance with a $0 budget, GitHub refused every artifact upload estate-wide, and CI went red. The owner first dropped the evidence uploads, then chose to move the handoffs CI actually needs to a Cloudflare R2 bucket the estate owns.

Dropped (evidence only; nothing gated on it):

  • nuxt-cloudflare.yml no longer has:

    • the Playwright evidence upload;
    • the E2E report merge job;
    • the journey-smoke evidence upload;
    • the foundation-check artifact.

    Evaluate web-foundation conformance check now prints foundation-check.json in its own log.

  • reusable-browser-tests.yml no longer has:

    • the Chromium and WebKit blob-report uploads;
    • the report merge job, with its HTML report artifact and its report-artifact output.

    report-timeout-minutes stays as a deprecated, unused input so callers that pass it keep validating. Moving these to R2 would have put a bucket write credential on the isolated Playwright pool for evidence nobody gates on, so they were dropped, not moved.

  • Every shard reports with the caller's own Playwright reporter in its job log. There is no --reporter=blob, which only existed to be merged.

Moved to R2:

  • the prebuilt E2E application Build hands to every E2E job;
  • the Required proof (required-reuse-pr-results);
  • the full E2E proof (e2e-reuse-pr-results).

Not moved yet: reusable-browser-tests.yml still takes its build as a same-run GitHub artifact the caller uploads (workflows#150).

Bucket, keys and lifecycle

Bucket narduk-ci-artifacts, in the narduk-enterprises Cloudflare account (location WNAM)
Prebuilt application prebuilt/<owner>/<repo>/<run>-<attempt>-<scope>.tar.gz, expires after 7 days
Reuse proof proof/<owner>/<repo>/<proof key>.json, expires after 30 days
Credential probe probe/<owner>/<repo>/<name>, expires after 1 day
Incomplete multipart upload aborted after 7 days

Every key is scoped by github.repository and shape-checked before use.

Credentials

The workflow uses three org Actions secrets with selected visibility: only the repositories on each secret's repository list receive them. All three are declared required: false on the callable:

Secret Value
CI_ARTIFACTS_R2_ACCOUNT_ID the Cloudflare account ID
CI_ARTIFACTS_R2_ACCESS_KEY_ID the R2 token's ID
CI_ARTIFACTS_R2_SECRET_ACCESS_KEY the SHA-256 of the token value

The token has Object Read and Write on this one bucket and nothing else. It expires on 2027-09-26. Its source is nvault cloudflare/prd/narduk-enterprises-ci-artifacts (CI_ARTIFACTS_R2_TOKEN, CI_ARTIFACTS_R2_ACCOUNT_ID), persona cloudflare-narduk-enterprises-ci-artifacts.

A reusable workflow sees only the secrets its caller passes, so a caller adds the three lines shown in the main example above, or uses secrets: inherit.

A caller that repins onto a commit with the store must also be added to the repository list of all three secrets. Until it is, GitHub hands it empty values, and its runs fall back as described under "Without credentials, and on failure". An org owner adds a repository without touching any value:

repo_id=$(gh api repos/narduk-enterprises/<repo> --jq .id)
for name in CI_ARTIFACTS_R2_ACCOUNT_ID CI_ARTIFACTS_R2_ACCESS_KEY_ID CI_ARTIFACTS_R2_SECRET_ACCESS_KEY; do
  gh api -X PUT "orgs/narduk-enterprises/actions/secrets/$name/repositories/$repo_id"
done
# Read the list back:
gh api orgs/narduk-enterprises/actions/secrets/CI_ARTIFACTS_R2_SECRET_ACCESS_KEY/repositories \
  --jq '.repositories[].full_name'

Who holds the secret access key:

  • Build, to publish the prebuilt application;
  • Required, on a pull request, to publish proofs;
  • Reuse plan and E2E plan, which read proofs on a default-branch push.

The E2E jobs, including those on the isolated browser pool, get only the two IDs. They read the prebuilt application through a presigned GET that Build signed for that one object, valid for 24 hours.

That signature is not a secret, and it appears in the logs. It travels as a job output, and GitHub drops any job output that carries a secret, so each E2E job prints it in the fetch step's environment. GitHub masks the account ID and the access key ID, but they are identifiers, not keys. So anyone who can read that run's logs can fetch that one tarball until the link expires. That is the same audience, for the same day, that could download the GitHub artifact it replaces.

Trust

One token serves the whole estate, so nothing is trusted just because it is in the bucket.

Prebuilt application. E2E unpacks the tarball only when its SHA-256 and size equal what its own run's Build reported through job outputs. A mismatch fails the job before anything is unpacked.

Proof. A proof object is only a pointer to a run. It is honoured only when all of the following hold:

  • GitHub's API reports that run as a completed, successful pull_request run of the same caller workflow;
  • the run's repository and head repository are both this repository;
  • the run attempt is the one that wrote the pointer;
  • that attempt's Required job succeeded;
  • that job has a successful step named Publish Required proof <key> or Publish full E2E proof <key>, with this exact key.

The key is unchanged from the artifact era: the digest of the tested Git tree, the caller workflow and every callable input.

Without credentials, and on failure

The store can cost time but never coverage, and it never passes a gate falsely. A run has no store credentials when it is a fork pull request, a Dependabot run (Dependabot sees only Dependabot secrets), a public caller, a caller that does not pass the secrets, or a caller missing from their repository list. Then:

  • Build publishes nothing, with a ::warning:: naming the three secrets (a notice until ci-reset W10, W8 item 1, and nobody read it).
  • Each E2E job warns that it rebuilds the application, then runs build-script itself and fails if it cannot produce the output.
  • Proofs. No proof is published or honoured, so the default-branch push runs the full gate.

A storage error, an expired link or a stalled transfer does the same, with a warning. A tarball transfer gives up after 3 attempts of 120 s. That is about 6 minutes in the worst case, well inside the 30-minute job limit, so Build still finishes and each E2E job still has time to build its own application.

Only two things turn a job red:

  • a prebuilt application that fails its digest check;
  • an in-job fallback build that fails.

Tests

scripts/test_ci_artifact_store.py runs the client exactly as the workflow installs it. It checks:

  • AWS's published Signature V4 examples;
  • a round trip and tampering against a local S3 double that authenticates every request with an independent signer;
  • the fallbacks;
  • proof publish and lookup;
  • the rule that no GitHub artifact action or artifact API appears in nuxt-cloudflare.yml.

test_e2e_build_artifact.py, test_e2e_reuse.py and test_fast_path.py cover the wiring and the proof lookups job by job.

About

Shared reusable GitHub Actions workflows for the narduk-enterprises estate (CI-5). Callers pin @vn tags or SHAs, never @main.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages