-
Notifications
You must be signed in to change notification settings - Fork 10
302 lines (290 loc) · 16 KB
/
Copy pathboard-snapshot.yml
File metadata and controls
302 lines (290 loc) · 16 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
name: Board Snapshot
# The standing caller for `scripts/pm/board-snapshot.mjs` (#17390).
#
# ## Why a scheduled backup, and why it lands on a branch
#
# Three fleet accounts were suspended in two months and the appeals are
# unanswered. #17374 F3 measured what a suspension destroys: every issue, pull
# request and comment the account authored — while every branch and commit
# survives, because those belong to the repository rather than to a user. The
# board IS the fleet's state by rule (pm-dispatch SKILL.md, 全体座位的不变量:
# 「GitHub 之外永不维护任何跟踪状态」), so today one suspension erases state that
# nothing else holds; six cards, one of them a p1 security decision, are gone.
#
# The maintainer's answer, verbatim and untranslated, in chat on 2026-09-10:
# 「或者用 ci 写个定时任务备份?」 — this file is that scheduled task. It runs as
# `github-actions[bot]`, which is not a fleet account and cannot be suspended
# with one, and it writes to an ORPHAN BRANCH of this same repository: no second
# repo, no second credential, and a branch is precisely the thing a suspension
# leaves standing.
#
# ## An archive is not a tracker — the one-board rule still holds
#
# ⛔ Nothing reads `board-archive` for state. Not a seat, not a patrol, not a
# gate, not a query; every reading of the board still goes to GitHub. The
# archive answers exactly one question, after a loss: what did the record say?
# The script it runs has NO write path to GitHub in any mode — two of its
# self-test cases read its own source and hold that structurally — and
# `--restore` PRINTS a recreate payload for a seat to post, or for nobody to.
#
# ## The budget this run is capped against, and who else spends it
#
# `GITHUB_TOKEN` is limited to 1,000 requests per hour PER REPOSITORY. That is
# ONE pool for every workflow in this repo that calls the API — not a per-workflow
# allowance — and this job is far from its only caller:
#
# this archiver 800 per run (the `--max-requests` cap below), 4/day
# half-state patrol a board sweep, 4/day at :37 on hours 1,7,13,19
# required-set patrol 2/day at :23 on hours 4,16
# release-coverage patrol daily at 04:19; platform-checklist watchdog at 02:51
# merged-branch reaper Mondays 04:37, an API sweep over branches and PRs
# every CI, lint and smoke run
# a handful of requests each, hourly and per push
#
# 800 is chosen against that pool rather than against this job's appetite: it is
# the largest cap that still leaves a fifth of the hour to whatever else lands in
# the same window, and the schedule below keeps the two heavy board readers out
# of each other's hours entirely. ⛔ Raising the cap is not a tuning decision,
# ⛔ there is no retry loop, and ⛔ there is no second token — a run that wants
# more requests waits for the next slot, which is what the resume cursor is for.
#
# ## Why a run may stop before it is finished — and what it reads FIRST
#
# A first full snapshot of this board — several thousand numbers, each with a
# comment thread — does not fit one run and must not try, so the script stops at
# its own `--max-requests` ceiling, writes its phase cursor into the manifest and
# exits 0. The next scheduled run continues from that cursor. Four runs a day
# walk the backlog in a few days without ever exceeding the budget, and a
# steady-state incremental run costs a few hundred requests.
#
# WHICH records it reads first is the part that matters, and it is measured
# rather than assumed. The first scheduled run of this workflow walked
# `state=all` oldest-first and spent all 800 requests on 137 closed issues and
# 268 closed pull requests from February — ZERO open cards, with days of runs
# still to go. The open board is exactly what a suspension destroys (#17374 F3),
# so it was the half being archived last. A first walk now runs the OPEN phase to
# completion — issues and pull requests, one listing — before the closed history
# starts, and a run that runs out of budget inside the open set resumes inside
# it. The manifest's `walk_phase` says which phase a run is in, so "open set
# complete, history resuming" is never read as "still inside the open set".
#
# ⛔ On a real rate-limit refusal the script does NOT retry: it stops, writes the
# cursor and exits non-zero with the reset time. A loop against a spent budget
# starves every other automated caller in the repository for the rest of the hour.
#
# ## The `pull_request` trigger, and what it is allowed to do
#
# Every patrol in this tree exercises itself on the pull requests that change
# it, and that is not a style preference: `check:pm-dispatch-gates` pins that NO
# discovered gate family is reachable only through a schedule, so a workflow
# whose script CI never runs at PR time reddens that gate by existing. A first
# draft of this file omitted the trigger to keep PR-time cost at zero and was
# refused by exactly that pin.
#
# So a pull_request run DOES execute the snapshot on a real runner — the
# transport, the flags, the page walk and the rendering, proven where they will
# actually run — and it writes NOTHING: no archive branch is checked out, the
# run is `--dry-run` into the runner's temp dir, and the commit step is skipped.
# It is also capped hard (`--limit`, `--max-requests` below), because the
# repository's request budget is shared with the half-state patrol and the
# closed-card sweep and must not be spent per push. ⛔ A pull request must never
# write to the board OR to the archive.
#
# The offline half is held for free beside it: `check:pm-board-snapshot` runs the
# script's 70-case `--self-test` inside the required `Lint & Repo Gates` job,
# which has no paths filter (`check:self-test-wired` is what requires that step
# to exist). This job runs the same self-test again before it touches the
# archive, so a broken tool fails loudly instead of committing a broken backup.
#
# ## Adopting this in a sibling repo
#
# Copy FOUR files, unchanged: `scripts/pm/board-snapshot.mjs`,
# `scripts/pm/check-half-states.mjs` (the archiver imports its proxy-re-exec plan
# and its repo resolver), `scripts/invoked-as.mjs` (imported by both) and this
# file. A repo that has already adopted the half-state patrol has the middle two.
# ⛔ Do not shorten this list from memory — the imports decide it, not this
# comment, and the patrol's own adopt list was wrong for months in exactly that
# way, which installed a dead sweeper that failed before its first predicate.
#
# There is nothing to configure. The board archived is `github.repository`,
# passed explicitly below, and the script REFUSES its resolver's default rather
# than falling back to one: a copy of this file quietly archiving the repo it was
# copied FROM would produce a complete, well-formed, green archive of the wrong
# board.
on:
schedule:
# Four times a day, six hours apart, at :07 past the hour. The HOURS are
# chosen against the other standing spenders listed above, not for tidiness.
#
# This job first shipped at `37 1,7,13,19` — the half-state patrol's cron to
# the minute. Both are scheduled board readers, both authenticate as this
# repository's `GITHUB_TOKEN`, and that budget is per repository per hour, so
# a snapshot run spending its full 800 leaves the patrol the remainder of the
# same window. Measured on the day this moved: the 13:37Z patrol run and the
# 13:41Z snapshot run were BOTH green, so nothing had been starved yet — the
# offset is PREVENTION, taken while the archive was still small enough that
# no run had yet spent its whole cap against the patrol's window.
#
# Of the five hour sets that share no hour with the patrol's 1,7,13,19, this
# is the least contended on this repository's cron inventory: 08, 14 and 20
# UTC carry no other scheduled workflow at all, and 02 UTC carries only the
# platform-checklist watchdog at :51 — 44 minutes after this run starts —
# plus CodeQL on Mondays. `4,10,16,22` was the obvious alternative and is
# worse on the same measurement: it drops an 800-request run into the busiest
# scheduled hour on this board (04:00 rerun-safety, 04:19 release-coverage
# patrol, 04:23 required-set patrol, 04:37 branch reaper on Mondays, 04:41
# create-smoke) and shares hour 16 with the required-set patrol.
#
# The minute is off the top of the hour on purpose, as in every patrol here:
# scheduled workflows queue behind everyone else's :00 cron.
#
# Six-hourly is the loss window this accepts: a card created and destroyed
# inside one interval was never archived, and nothing cheaper than a webhook
# closes that, which is a different card.
- cron: '7 2,8,14,20 * * *'
workflow_dispatch: {}
# Changes to the archiver itself get exercised before they merge. The paths
# name every file the run actually loads — the archiver, the module it imports
# its proxy plan and repo resolver from, the entry-guard helper both import,
# and this file. An undeclared dependency is one a change can break without
# this trigger firing, which is the same defect the adopt list above warns
# about, one layer along.
pull_request:
paths:
- 'scripts/pm/board-snapshot.mjs'
- 'scripts/pm/check-half-states.mjs'
- 'scripts/invoked-as.mjs'
- '.github/workflows/board-snapshot.yml'
# Least privilege. `contents: write` is for the archive BRANCH and nothing else;
# the two read scopes are what the snapshot reads. ⛔ This job never writes to
# `main` and never writes to any issue, pull request, label or comment — there is
# no scope here that would let it, which is the mechanical half of the one-board
# rule stated above.
permissions:
contents: write
issues: read
pull-requests: read
# One archiver at a time. Two runs racing the same branch would have the loser
# push a snapshot built from a stale checkout, silently discarding the winner's.
concurrency:
group: board-snapshot
cancel-in-progress: false
jobs:
archive:
name: Snapshot the board to the archive branch
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@v7
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: '22'
# No `pnpm install`: the archiver imports only `node:` builtins, global
# `fetch`, and two repo-local modules — no npm dependency, so installing
# the workspace here would buy nothing and would give a scheduled backup a
# lockfile it could fail on.
- name: The tool's own cases, before it touches the archive
run: node scripts/pm/board-snapshot.mjs --self-test
# The archive branch is an ORPHAN: it shares no history with `main`, so a
# checkout of it carries the archive and nothing else. First run creates it
# locally; the push below is what publishes it.
- name: Check out or create the archive branch
if: github.event_name != 'pull_request'
run: |
set -euo pipefail
git fetch --depth=1 origin '+refs/heads/board-archive:refs/remotes/origin/board-archive' || true
if git rev-parse --verify --quiet refs/remotes/origin/board-archive > /dev/null; then
git worktree add archive -b board-archive refs/remotes/origin/board-archive
echo "archive branch exists; checked out at $(git -C archive rev-parse --short HEAD)"
else
git worktree add --detach archive HEAD
git -C archive checkout --orphan board-archive
git -C archive rm -rf --quiet .
# An orphan that still carries main's tree would commit the whole
# repository into the archive branch on its first run, so this is
# asserted rather than assumed.
tracked=$(git -C archive ls-files | wc -l)
if [ "$tracked" != "0" ]; then
echo "::error::the new orphan branch still tracks $tracked file(s) from main; refusing to seed the archive with them."
exit 1
fi
echo "archive branch created as an orphan (first run)"
fi
- name: Run the snapshot
id: snapshot
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
# WHICH board this run archives: the repo this workflow is installed
# in, always. A copy of this file cannot end up archiving the repo it
# was copied FROM, and the script refuses to guess when this is unset.
PM_SWEEP_REPO: ${{ github.repository }}
# A pull_request run PROVES the tool on a runner and writes nothing:
# the temp dir has no manifest, so it is a full read, capped to a
# handful of numbers and a budget the shared quota will not notice.
# ⛔ The two spellings differ in WHERE they write and in nothing else,
# so what a PR exercises is the path a schedule takes.
SNAPSHOT_ARGS: ${{ github.event_name == 'pull_request' && '--out=board --dry-run --limit=5 --max-requests=25' || '--out=archive/board' }}
run: |
set +e
node scripts/pm/board-snapshot.mjs $SNAPSHOT_ARGS \
> "$RUNNER_TEMP/snapshot.md" 2> "$RUNNER_TEMP/snapshot.err"
code=$?
set -e
# Captured with NO pipe in between. `cmd | tail` reports the PIPE's
# status — `tail` essentially never fails, so a clean snapshot and a
# failed count check would both read as 0, and this script's exit
# register (0 / 2 / 3 / 4) is the whole point of having one.
echo "exit_code=$code" >> "$GITHUB_OUTPUT"
echo "board-snapshot exited $code"
cat "$RUNNER_TEMP/snapshot.md" || true
cat "$RUNNER_TEMP/snapshot.err" >&2 || true
# Land the truth, THEN raise the alarm: whatever the snapshot managed to
# read is committed before the run is allowed to go red, so a failing count
# check never costs the records the same run archived.
- name: Commit and push the archive
id: commit
if: github.event_name != 'pull_request'
run: |
set -euo pipefail
if [ -z "$(git -C archive status --porcelain)" ]; then
echo "committed=none" >> "$GITHUB_OUTPUT"
echo "no change on the board since the last run — nothing committed."
exit 0
fi
git -C archive add -A
files=$(git -C archive diff --cached --name-only | wc -l)
git -C archive \
-c user.name='github-actions[bot]' \
-c user.email='41898282+github-actions[bot]@users.noreply.github.com' \
commit -q -m "board snapshot $(date -u +%Y-%m-%dT%H:%M:%SZ)" \
-m "run ${{ github.run_id }} · ${files} file(s) changed"
git -C archive push origin board-archive
echo "committed=${files}" >> "$GITHUB_OUTPUT"
echo "pushed ${files} changed file(s) to board-archive"
- name: Publish the run to the summary
if: always()
run: |
{
echo "### Board snapshot — exit ${{ steps.snapshot.outputs.exit_code }}"
echo
echo "Committed: ${{ steps.commit.outputs.committed || 'nothing (the run stopped before the commit step)' }}"
echo
echo '```'
cat "$RUNNER_TEMP/snapshot.md" 2>/dev/null || echo '(no report produced)'
echo '```'
echo
echo '<details><summary>stderr</summary>'
echo
echo '```'
cat "$RUNNER_TEMP/snapshot.err" 2>/dev/null || true
echo '```'
echo
echo '</details>'
} >> "$GITHUB_STEP_SUMMARY"
- name: Fail the run if the snapshot did not read the board
if: steps.snapshot.outputs.exit_code != '0'
run: |
echo "::error::board-snapshot exited ${{ steps.snapshot.outputs.exit_code }} — 2 is a failed count check (the archive and the board disagree about how many open issues exist), 3 is a prerequisite the runner did not meet, 4 is a rate-limit stop with a resume cursor written. Whatever was read HAS been committed; see this run's summary."
exit 1