Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 44 additions & 9 deletions .claude/skills/babysit-pipeline/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,9 +102,13 @@ with evidence — never let the user ask "still running?".
- **Defer, don't halt**: one stuck library must not stop a multi-spec
queue — log it, cap ~90 min/spec, move on, and sweep leftovers at
the end with targeted regens.
- **Halt the whole queue** only on a failure CLUSTER (≥5 failed
`Generate:` runs in minutes = model daily quota exhausted; a
fallback model via `-f model=` can finish leftovers).
- **Halt the whole queue** only on a failure CLUSTER: **≥3 distinct
(spec, library) pairs** failing within minutes of each other = model
daily quota exhausted; a fallback model via `-f model=` can finish
leftovers. The threshold is pairs, full stop — the raw run count
says nothing, because one impossible pair spends three runs on its
own auto-retries, so two bad pairs already produce six failures with
the pipeline entirely healthy (observed 2026-08-24).
- Closing/merging pipeline PRs or issues yourself is out of bounds
(CLAUDE.md external-write rule + mandatory workflow).

Expand Down Expand Up @@ -160,6 +164,24 @@ ggplot2/network-force-directed, all three of which had been read as
capability gaps. Two failures on the same pair → `deferred.log` with
the reason, move on, sweep at the end.

Check the run list before deferring, though: since #10627 the
workflow spends its own three attempts per campaign, so a pair that
comes back missing may already have failed three times, and the
manual retry adds nothing. There is no per-spec filter on the API, so
filter the output by run title:

```bash
gh run list --workflow=impl-generate.yml --limit 40 \
--json conclusion,createdAt,displayTitle \
--jq '.[] | select(.displayTitle | test("<spec>")) | "\(.createdAt) \(.conclusion) \(.displayTitle)"'
```

Three `Generate: <lib> for <spec>` failures minutes apart is a measured
gap; one failure is a flake worth the retry. The same check covers
`RESULT=TIMEOUT` with `recent generate failures: 0`, which only means
the failures fell outside the driver's 25-minute window, not that
nothing failed.

## Gotchas

- **`Marking <lib> as failed: N generation attempts` counts more than
Expand All @@ -185,12 +207,25 @@ the reason, move on, sweep at the end.
failed 17:55, retried and succeeded 18:02, no human involved), so a
single occurrence is noise — only a repeat on the same pair means
anything.
- **Static library + interactive/3D spec is the one gap that is
usually real.** plotnine, pygal, seaborn, matplotlib and ggplot2
against `*-interactive`, `*-drilldown`, `*-realtime`, `*-streaming`,
`slider-*`, `*-3d` specs make up most genuine dead ends; 18 of the
45 real gaps behind those labels were exactly this shape. Still
give them the one retry, then defer without further ceremony.
- **Do not predict capability gaps — measure them.** The 2026-08-24
backfill guessed four times which pairs were genuinely impossible
(chartjs on treemap/sankey, plotnine on 3D, ggplot2 on wireframe,
"static library vs. interactive spec") and every category-level
guess was wrong: 17 of 20 parked pairs generated fine, pygal and
chartjs each succeeded on the very plot types they were written off
for, and `bar-3d-categorical` succeeded in plotnine while
`scatter-3d` did not. Only three pairs failed with a full budget —
plotnine on `scatter-3d`, `contour-3d`, `line-3d-trajectory`, all
needing a spatial projection plotnine does not have, while the
"3D" spec that is representable in 2D went through. A gap is real
when three attempts under a fresh campaign budget say so, never
because the pairing sounds implausible.
- **A spec id from an issue title may not exist.** `impl:*:failed`
labels outlive their specs: of 26 spec IDs harvested that way, 14
had no `plots/<spec>/` on main at all, and a dispatch for one dies
at `Validate specification exists` seconds in. Intersect any
label-derived list with `git ls-tree -d origin/main plots/` before
queueing it.
- **The scripts' library list is a copy** of `core/constants.py`'s
registry (15 libs). When a library is added/removed, update
`monitor_spec.sh`'s `ALL_LIBS` and `run_spec.sh`'s `lang_of` in the
Expand Down
21 changes: 19 additions & 2 deletions .claude/skills/babysit-pipeline/run_spec.sh
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,24 @@ recent_generate_failures() {
|| echo "?"
}

git -C "$REPO" fetch origin main --quiet
# Two drivers polling the same checkout collide on `.git/refs/remotes/origin/main`
# ("cannot lock ref ... is at X but expected Y"). A lost fetch leaves origin/main
# stale, which makes meta_present understate progress — so retry, and say so in
# the log when all three attempts lose, rather than silently polling a stale ref.
fetch_main() {
Comment on lines +80 to +84

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, and my fault: the script commit came after I wrote the description. Title and body updated — the scope line now reads "prose plus one executable change" and there is a dedicated section for the run_spec.sh behaviour change with its own verification.

On the log detail: fetch_main now captures git stderr and prints it with the warning. A ref-lock collision and an auth or network failure call for different responses, so discarding the message threw away the one bit that distinguishes them. Exercised against a bad ref → last error: fatal: couldn't find remote ref ….

local i err
for i in 1 2 3; do
# Keep stderr: a ref-lock collision and, say, an auth or network failure
# need different responses, and discarding the message hides which it was.
err=$(git -C "$REPO" fetch origin main --quiet 2>&1) && return 0
sleep $(( i * 3 ))
done
echo "[warn] git fetch origin main failed 3x — origin/main may be stale this round" | tee -a "$LOG"
echo "[warn] last error: ${err}" | tee -a "$LOG"
return 1
}

fetch_main
TODO=()
for lib in "${LIBS[@]}"; do
if meta_present "$lib"; then
Expand Down Expand Up @@ -106,7 +123,7 @@ iters=$(( MAXMIN * 60 / INTERVAL ))
idle=0; last_done=-1
for ((i=1; i<=iters; i++)); do
sleep "$INTERVAL"
git -C "$REPO" fetch origin main --quiet
fetch_main
missing=(); done_n=0
for lib in "${TODO[@]}"; do
if meta_present "$lib"; then done_n=$((done_n+1)); else missing+=("$lib"); fi
Expand Down
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,6 +180,21 @@ aggregate instead: an italic *Catalog* line at the end of the version section an

### Changed

- **`babysit-pipeline`'s gap-backfill guidance corrected against a completed run** — the
section shipped in #10628 was written mid-backfill and three of its claims did not
survive the rest of it. The halt-on-cluster threshold counted raw failed runs, but one
impossible pair now burns three runs on its own retries, so two bad pairs tripped it
while the pipeline was healthy; it counts distinct `(spec, library)` pairs now. The
"static library against an interactive or 3D spec is the gap that is usually real"
gotcha is replaced by its opposite: every category-level prediction made during the
backfill was wrong — 17 of 20 parked pairs generated fine, and `bar-3d-categorical`
succeeded in plotnine while `scatter-3d` did not — so a gap counts as real only after
three attempts under a fresh budget. Added: spec IDs harvested from issue titles must be
intersected with the actual `plots/` directories (14 of 26 pointed at specs that no
longer exist), and a `TIMEOUT` with zero recent failures means the failures aged out of
the driver's window, not that nothing failed. `run_spec.sh` also retries its
`git fetch`: two drivers polling one checkout collide on the `origin/main` ref lock,
and a lost fetch made the poller read stale metadata and understate progress (#10660).
- **`babysit-pipeline` learned gap backfill** — closing coverage gaps across the catalogue is
a different job from watching one fresh spec, and the skill only described the latter. It
now documents reading the missing set from `plots/*/metadata/` instead of from labels
Expand Down
Loading