From 9aa55bb8fd32a0c93d4f00e40f54e43731cd7121 Mon Sep 17 00:00:00 2001 From: Claude Date: Fri, 25 Sep 2026 20:49:48 +0000 Subject: [PATCH 1/8] Add an sshfs backend to eval-under sshfs is the filesystem people actually reach for when they mount a remote over ssh, and it breaks git-annex in a way no local filesystem does. This adds it as a backend so a report against it can be reproduced directly. Two modes: - Loopback (default): a throwaway sshd on the first free port at or above 2222, with its own host key, authorized_keys and pid file inside the run's scratch directory, and a fresh backing directory sshfs-mounted back over it. The system sshd is not used and ~/.ssh/authorized_keys is never written to. - `--host` mounts a real remote using the caller's ssh config, so a reporter's own server and mount options can be used verbatim. Knobs for the options reports actually turn on: `--no-cache`, `--workaround`, `--opt`, `--port`, `--user`, `--remote-dir`, plus the common `--mount-point` / `--set-home` / `--keep`. Notes on the two non-obvious pieces: - Dropping privileges deliberately calls `sudo -u` literally rather than going through the `${SUDO[@]}` array the other backends use: that array is empty when we are already root, which is exactly the case that needs the drop. - sshd runs with `UsePAM yes`. With it off, sshd refuses any account whose shadow entry is locked (`!`), which is the normal state for service and CI accounts, and the mount fails for a reason that looks nothing like its cause. Its log is dumped on mount failure for the same reason. `bin/ci/run-under.sh` refuses sshfs together with a root-requiring target (pjdfstest): a FUSE mount belongs to whoever mounted it, there is no --no-root-squash equivalent, and the run would measure privilege rather than the filesystem. No matrix cell and no badge -- the backend exists for on-demand reproduction, and GOTCHAS.md records what it does to hardlink identity, mtime granularity and fifos before anyone reads a red result as a filesystem bug. Split out of the filesystem-probing branch, which had grown too large to review as one change. Verified: shellcheck clean, all 28 bats tests pass (the four "every installed backend" cases now cover this backend), the mount works as root and via the privilege-drop path, and the refusal above exits 2. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- GOTCHAS.md | 111 +++++++++ README.md | 9 + bin/ci/install-backend.sh | 20 +- bin/ci/run-under.sh | 18 +- bin/eval-under-sshfs | 506 ++++++++++++++++++++++++++++++++++++++ 5 files changed, 660 insertions(+), 4 deletions(-) create mode 100755 bin/eval-under-sshfs diff --git a/GOTCHAS.md b/GOTCHAS.md index 06bba2b..477e98a 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -75,6 +75,117 @@ So `--no-root-squash` (env: `EVAL_UNDER_NFS_NO_ROOT_SQUASH`) exports with user would. The loop and BeeGFS backends already run the wrapped command as root, so the flag is a no-op there. +### sshfs (`bin/eval-under-sshfs`) + +A throwaway sshd is started on the first free port at or above 2222 -- +its own host key, its own `authorized_keys`, its own pid file, all inside +the run's scratch directory -- and a fresh backing directory is +sshfs-mounted back over it. The system sshd is not used and +`~/.ssh/authorized_keys` is never written to. With `--host` it mounts a +real remote instead, using the caller's ssh config. + +| Knob | Value | Why | +| --- | --- | --- | +| Port | first free `>= 2222` | Two runs at once (a reproduction while a suite is going, two CI cells on one runner) must not collide. `--port` pins it. | +| `-o reconnect,ServerAliveInterval=15,ServerAliveCountMax=3` | always | A dropped connection should fail the command, not wedge it forever. | +| `-o cache=no` | only with `--no-cache` | sshfs caches attributes by default, which hides stale-stat behaviour. This turns off sshfs's own cache, not all caching -- the kernel's 1-second attribute timeout still applies, and the stale-size effect above survives it. On is what users actually have. | +| `-o workaround=rename` | only with `--workaround rename` | Makes sshfs emulate rename-over-existing by unlinking first -- non-atomic, which is what SFTP servers without the POSIX-rename extension force. | +| Mount/command user | the invoking user | A FUSE mount belongs to whoever ran `sshfs`; root cannot read it without `allow_other`. Same treatment `eval-under-nfs` gives `root_squash`. | + +**Hardlinks exist but are not observable, and that is the whole story +for git-annex.** `ln a b` succeeds over SFTP, and then `a` and `b` report +*different* inode numbers and `nlink=1` each: + +``` +ext4 name=a ino=1884275 nlink=2 name=b ino=1884275 nlink=2 +sshfs name=a ino=3 nlink=1 name=b ino=4 nlink=1 +``` + +This is not a misconfiguration, and the link is not fake: in the backing +directory on the server both names really do share one inode with +`nlink=2`. What SFTP cannot carry is *identity* -- its attribute record +has neither an inode number nor a link count -- so sshfs synthesises an +`st_ino` per path and reports `nlink=1` for everything. There is no knob +for it: `use_ino` (removed in libfuse 3) only ever passed through inode +numbers that a filesystem supplies, and sshfs has none to supply, so it +would not have helped under FUSE2 either; sshfs 3.7 offers only +`disable_hardlink`, which makes `link()` fail outright. + +Worse than invisible, and worth knowing when a report mentions truncated +files: immediately after writing one name, the *other* name still reads +back with size 0 through the mount -- with `-o cache=no` as well. The +capability probe reports the inode half as +`hardlink-same-inode=no` / `hardlink-nlink=no` while `hardlink=yes` -- +which is exactly why those two checks exist. `hardlink` alone called +sshfs healthy. + +What it costs, measured with git-annex 10.20240129: + +- `git annex add` on a **locked** branch: fine. The file becomes a + symlink into `.git/annex/objects`, and symlinks work. +- **Any unlocked `git annex add`: fails** -- `foo failed to link to + annex`. add hardlinks the content into the annex and then verifies the + link, and the verification cannot succeed on a filesystem where the + link is invisible. This is not limited to an adjusted branch: a plain + v10 repo fails identically with `annex.addunlocked=true` in git + config, set through `git annex config`, or passed as `-c`. +- `git add` through git's own filter (with `annex.largefiles` matching): + **fine** -- the pointer is committed and `git annex fsck` is clean. + So is `git annex unlock` of an already-committed file. +- `git clone` of a local repo: **fails** -- + `fatal: hardlink different from source at '...'`. git's local-clone + path hardlinks objects and runs the same check. + +So the boundary is not locked-vs-unlocked, and not adjusted-vs-plain: +it is **who does the ingest**. git-annex hardlinking content into the +annex fails; git's filter writing a pointer does not. + +That distinction is easy to get backwards from the test suite alone, +because `git annex test` shows `Repo Tests v10 unlocked` **green** and +only `v10 adjusted unlocked branch` red (11 of 12 in its Init Tests +group, once `add` fails). The green group is not evidence that unlocked +repos are fine here -- the suite's unlocked mode ingests with `git add`. +An earlier revision of this file drew exactly that inference and was +wrong. + +It matters for DataLad, which calls `git annex add`: plain v10 unlocked +repos break for it too, not just adjusted ones. + +**`git annex test` wedges partway through, reproducibly.** Both +full-suite runs stopped at the same place -- `Remote Tests / unavailable +remote / removeKey` -- and sat there until killed (15+ minutes on the +second). On ext4 that same test takes **0.02s and passes**, and the whole +suite finishes in 1m21s. + +It is not an I/O hang. While wedged: + +- the mount stays responsive (`ls` returns immediately), +- git-annex holds **no open files on the mount and no sockets**, +- its threads sit in `futex_do_wait` / `ep_poll`, with no child + processes outstanding. + +That is a process waiting on something internal, not one blocked on the +filesystem. Note also that git-annex sets `annex.sshcaching = false` here +on its own, because ssh control sockets need unix sockets and this mount +has none -- so the run is already on a different code path from a normal +one. + +Two caveats before anyone reports this upstream: the same test passes in +seconds when selected on its own with `-p '/unavailable remote/'`, so it +needs the full-suite context; and this was git-annex 10.20240129 from +Ubuntu 24.04, not a daily build. Re-run it through the reproduce +workflow, which installs the daily build from con/git-annex, before +filing anything. + +**Timestamps are quantised to the second.** Five files created back to +back get one or two distinct mtimes, where every other filesystem +measured gives five. Nothing else in the matrix has a clock this coarse, +which makes sshfs the row that exercises git's racy-timestamp handling. + +**No fifos, no unix sockets.** git-annex says so itself at init +("Detected a filesystem without fifo support") and adapts. Worth knowing +before reading it as a failure. + ### BeeGFS (`bin/eval-under-beegfs`) A containerised cluster (`fixtures/beegfs/docker-compose-v{7,8}.yml`) plus diff --git a/README.md b/README.md index 23f7932..e687c9e 100644 --- a/README.md +++ b/README.md @@ -161,6 +161,14 @@ sudo bin/eval-under loop --fs xfs --size 200 --set-home -- \ # the fsync-heavy slow path) sudo bin/eval-under nfs --set-home -- bash -c 'cd "$HOME" && git annex test' +# Under sshfs, with the attribute cache off +sudo bin/eval-under sshfs --no-cache --set-home -- \ + bash -c 'cd "$HOME" && git annex test' + +# ...or against the reporter's own server, with their mount options +sudo bin/eval-under sshfs --host store.example.org --remote-dir /data/scratch \ + --workaround rename --set-home -- git annex fsck + # Skip teardown to poke around after a failure sudo bin/eval-under beegfs --set-home --keep -- some-failing-command @@ -204,6 +212,7 @@ into `bin/eval-under`, bumped with each release tag. | `bin/eval-under` | Dispatcher: routes to `bin/eval-under-` | | `bin/eval-under-beegfs` | BeeGFS backend (containerised cluster + kernel client mount) | | `bin/eval-under-nfs` | NFS backend (localhost loopback export) | +| `bin/eval-under-sshfs` | sshfs backend (throwaway loopback sshd, or a remote you name) | | `bin/eval-under-loop` | Loop-device backend (dd + losetup + mkfs. + mount) | | `fixtures/beegfs/docker-compose-v7.yml` | BeeGFS v7 test cluster (mgmtd + meta + storage), `network_mode: host` | | `fixtures/beegfs/docker-compose-v8.yml` | Same, for BeeGFS v8.x (different mgmtd command style / gRPC control plane) | diff --git a/bin/ci/install-backend.sh b/bin/ci/install-backend.sh index 6c5ee54..4d76d7a 100755 --- a/bin/ci/install-backend.sh +++ b/bin/ci/install-backend.sh @@ -9,17 +9,18 @@ # usage: # bin/ci/install-backend.sh # -# backend = beegfs | nfs | loop +# backend = beegfs | nfs | loop | sshfs # version = for beegfs: point release (e.g. 7.4.6, 8.1.0) # for loop: filesystem type (e.g. vfat, ext4, xfs, btrfs) # for nfs: literal "n/a" +# for sshfs: literal "n/a" # # Idempotent enough for CI re-runs; not a full package manager. set -euo pipefail export DEBIAN_FRONTEND=noninteractive -BACKEND="${1:?backend required (beegfs|nfs|loop)}" +BACKEND="${1:?backend required (beegfs|nfs|loop|sshfs)}" VERSION="${2:?version required (BeeGFS version | loop fs name | 'n/a' for nfs)}" # Give unattended-upgrades a moment on ubuntu-22.04 runners rather than @@ -72,6 +73,18 @@ install_nfs() { command -v exportfs } +install_sshfs() { + apt_update + # openssh-sftp-server is what actually serves the mount; on Ubuntu it + # is pulled in by openssh-server, but name it so a slimmer image + # cannot leave us without an sftp-server binary. + apt_install sshfs openssh-server openssh-sftp-server + command -v sshfs + # The backend starts its own sshd, so the system one need not run -- + # but its privilege-separation directory must exist. + sudo mkdir -p /run/sshd +} + install_loop() { local pkg case "$VERSION" in @@ -90,5 +103,6 @@ case "$BACKEND" in beegfs) install_beegfs ;; nfs) install_nfs ;; loop) install_loop ;; - *) echo "unknown backend: $BACKEND (expected beegfs|nfs|loop)" >&2; exit 1 ;; + sshfs) install_sshfs ;; + *) echo "unknown backend: $BACKEND (expected beegfs|nfs|loop|sshfs)" >&2; exit 1 ;; esac diff --git a/bin/ci/run-under.sh b/bin/ci/run-under.sh index ff1a2cb..e35b2f8 100755 --- a/bin/ci/run-under.sh +++ b/bin/ci/run-under.sh @@ -13,10 +13,11 @@ # usage: # bin/ci/run-under.sh [target] # -# backend = beegfs | nfs | loop +# backend = beegfs | nfs | loop | sshfs # version = for beegfs: point release (e.g. 7.4.6, 8.1.0) # for loop: filesystem type (e.g. vfat, ext4) # for nfs: literal "n/a" +# for sshfs: literal "n/a" # target = git-annex (default) | git | stress-ng | pjdfstest # # env overrides: @@ -46,6 +47,20 @@ target_known "$TARGET" || { exit 1 } +# A root-requiring suite under a backend that can only run as the +# invoking user would measure nothing: every privileged syscall it exists +# to test fails for want of privilege, not because of the filesystem. The +# NFS backend has --no-root-squash for this; sshfs has no equivalent, +# because a FUSE mount belongs to whoever mounted it. Refuse the pair up +# front rather than produce a meaningless red result. +if [ "$BACKEND" = sshfs ] && target_needs_root "$TARGET"; then + echo "$TARGET needs root, and the sshfs backend always runs the wrapped" >&2 + echo "command as the invoking user (a FUSE mount belongs to its mounter)." >&2 + echo "There is no --no-root-squash equivalent here, so this combination" >&2 + echo "cannot measure what the target is for. Try the nfs or loop backend." >&2 + exit 2 +fi + # The target scripts re-derive their own defaults from matrix.sh, but an # override handed to us must survive into the wrapped child. export EVAL_UNDER_SRC_DIR @@ -62,6 +77,7 @@ case "$BACKEND" in # need an export that does not squash root, and need to keep # their privileges rather than being dropped to the invoker. target_needs_root "$TARGET" && opts=(--no-root-squash) ;; + sshfs) opts=() ;; *) echo "unknown backend: $BACKEND" >&2; exit 1 ;; esac diff --git a/bin/eval-under-sshfs b/bin/eval-under-sshfs new file mode 100755 index 0000000..3c952fa --- /dev/null +++ b/bin/eval-under-sshfs @@ -0,0 +1,506 @@ +#!/bin/bash +# SPDX-FileCopyrightText: 2026 Yaroslav Halchenko +# SPDX-License-Identifier: MIT +# +# Generated with Claude Code +# +# eval-under-sshfs: run a command with TMPDIR / DATALAD_TESTS_TEMP_DIR (and +# optionally HOME) pointing at an sshfs (FUSE-over-SFTP) mount. +# +# Backend of the eval-under framework. +# +# Two modes: +# +# loopback (default) -- brings up a throwaway sshd on a high port, with +# its own host key and its own authorized_keys, and sshfs-mounts a +# fresh backing directory back over it. Nothing outside the scratch +# dir is touched: the user's ~/.ssh is never written to. +# +# remote (--host) -- sshfs-mounts a directory on a host you name, using +# your existing ssh config/agent. For reproducing a reporter's +# actual setup, where the interesting variable is often their +# server's SFTP implementation rather than sshfs itself. + +set -eu + +# Empty MNT means "derive a per-run mountpoint next to the scratch dir" +# (see the MNT_BASE block below); only an explicit --mount-point / +# EVAL_UNDER_MOUNT pins the mount to a fixed path. +MNT="${EVAL_UNDER_MOUNT:-}" +KEEP="${EVAL_UNDER_KEEP:-0}" +SET_HOME="${EVAL_UNDER_HOME_ON_MOUNT:-0}" + +# Loopback-mode knobs. An empty PORT means "pick a free one": two +# eval-under runs on the same machine (or two CI cells on one runner) +# must not fight over a fixed port. +PORT="${EVAL_UNDER_SSHFS_PORT:-}" +# Remote mode: unset means loopback. +HOST="${EVAL_UNDER_SSHFS_HOST:-}" +REMOTE_USER="${EVAL_UNDER_SSHFS_USER:-}" +REMOTE_DIR="${EVAL_UNDER_SSHFS_DIR:-}" + +# sshfs -o options. EVAL_UNDER_SSHFS_OPTS is a comma-separated string +# appended after ours, so a caller can override any default we set. +EXTRA_OPTS="${EVAL_UNDER_SSHFS_OPTS:-}" +NO_CACHE="${EVAL_UNDER_SSHFS_NO_CACHE:-0}" +WORKAROUND="${EVAL_UNDER_SSHFS_WORKAROUND:-}" + +usage() { + cat <<'EOF' +Usage: eval-under-sshfs [OPTIONS] -- CMD [ARGS...] + +Run CMD with TMPDIR / DATALAD_TESTS_TEMP_DIR (and optionally HOME) +pointing at an sshfs mount. + +By default this is a loopback mount: a throwaway sshd is started on a +high port with its own host key and its own authorized_keys file, and a +fresh backing directory is sshfs-mounted back over it. Your ~/.ssh is +not read or written. With --host, mounts a remote directory instead, +using your normal ssh configuration. + +Requires the `sshfs` package (and, for loopback mode, `openssh-server`). +Install with: apt install sshfs openssh-server + +Options (flag / env var / default / purpose): + + --mount-point PATH EVAL_UNDER_MOUNT (next to scratch dir) + Where to mount sshfs on the host. Default: a fresh per-run + directory next to this run's scratch dir, both under $TMPDIR + (/tmp if unset) -- e.g. with TMPDIR=~/.tmp: + + ~/.tmp/eval-under-sshfs-B3UXj.scratch keys, sshd conf, backing + ~/.tmp/eval-under-sshfs-B3UXj.sshfs where this run mounts it + + A fixed default would collide between concurrent runs and hide + which run owned the mount. Pass this option when a stable path is + wanted. + + --set-home EVAL_UNDER_HOME_ON_MOUNT (unset) + Also set HOME=/home for the wrapped command. + + --keep EVAL_UNDER_KEEP (unset) + Skip teardown; leave the mount (and the throwaway sshd) up for + debugging. + + --port N EVAL_UNDER_SSHFS_PORT (first free >=2222) + Port for the throwaway sshd (loopback mode only). Left unset, the + first free port at or above 2222 is used, and a port lost to a + concurrent run between the scan and the bind is retried on the + next one up, so concurrent runs do not collide. + + --host HOST EVAL_UNDER_SSHFS_HOST (unset) + Mount from a remote host instead of starting a local sshd. Uses + your ssh config, keys and agent as-is. Implies that --port refers + to that host's sshd. + + --user USER EVAL_UNDER_SSHFS_USER (invoking user) + Remote username. Requires --host: the throwaway sshd generated for + loopback mode only admits the invoking user, so any other name + there could only ever be refused. + + --remote-dir PATH EVAL_UNDER_SSHFS_DIR (remote: required) + Directory on the remote host to mount. In loopback mode this + defaults to a fresh scratch directory and should be left alone. + + --no-cache EVAL_UNDER_SSHFS_NO_CACHE (unset) + Pass `-o cache=no`. sshfs caches attributes and directory listings + by default, which papers over exactly the stale-stat behaviour a + reporter may be hitting. Turn the cache off to see the raw + protocol semantics. + + --workaround LIST EVAL_UNDER_SSHFS_WORKAROUND (unset) + Pass `-o workaround=LIST` (sshfs's own compatibility switches, + e.g. `rename`, `truncate`, `buflimit`). `rename` is the + interesting one here: it makes sshfs emulate rename-over-existing + by unlinking the destination first, which is what it has to do + against SFTP servers lacking the POSIX-rename extension -- a + non-atomic rename, and precisely the semantics git and git-annex + assume they can rely on. + + --opt OPTS EVAL_UNDER_SSHFS_OPTS (unset) + Extra comma-separated `-o` options, appended last so they win. + May be given more than once. + + -h, --help + Print this help. + +Environment variables set FOR the wrapped command (all backends): + + TMPDIR = + DATALAD_TESTS_TEMP_DIR = + HOME = /home (only if --set-home) + +Exit status: the wrapped command's exit status. Teardown runs on any +exit (unless --keep). + +Examples: + + # git-annex test over loopback sshfs, cache off + sudo bin/eval-under sshfs --no-cache --set-home -- \ + bash -c 'cd "$HOME" && git init t && cd t && git annex init && git annex test' + + # Reproduce non-atomic rename-over-existing + sudo bin/eval-under sshfs --workaround rename --set-home -- some-command + + # Against a reporter's own server + sudo bin/eval-under sshfs --host store.example.org --remote-dir /data/scratch \ + --set-home -- git annex fsck +EOF +} + +while [ $# -gt 0 ]; do + case "$1" in + --keep) KEEP=1; shift ;; + --mount-point) MNT="$2"; shift 2 ;; + --set-home) SET_HOME=1; shift ;; + --port) PORT="$2"; shift 2 ;; + --host) HOST="$2"; shift 2 ;; + --user) REMOTE_USER="$2"; shift 2 ;; + --remote-dir) REMOTE_DIR="$2"; shift 2 ;; + --no-cache) NO_CACHE=1; shift ;; + --workaround) WORKAROUND="$2"; shift 2 ;; + --opt) EXTRA_OPTS="${EXTRA_OPTS:+$EXTRA_OPTS,}$2"; shift 2 ;; + -h|--help) usage; exit 0 ;; + --) shift; break ;; + *) echo "unknown arg: $1" >&2; usage >&2; exit 2 ;; + esac +done + +[ $# -gt 0 ] || { echo "no command given" >&2; usage >&2; exit 2; } + + +# Did the caller pin the port? Decided before start_local_sshd starts +# reassigning PORT while hunting for a free one. +PORT_PINNED="${PORT:+1}" + +if [ "$(id -u)" -ne 0 ]; then + command -v sudo >/dev/null || { + echo "must run as root (or have sudo available)" >&2; exit 2; } + SUDO=(sudo) +else + SUDO=() +fi + +# The FUSE mount belongs to whoever runs sshfs, and the wrapped command +# has to be able to read it. Both therefore run as the invoking user -- +# the same treatment bin/eval-under-nfs gives root_squash. +if [ -n "${SUDO_UID:-}" ]; then + INVOKER_UID="$SUDO_UID"; INVOKER_GID="${SUDO_GID:-$SUDO_UID}" + INVOKER_USER="${SUDO_USER:-$INVOKER_UID}" +else + INVOKER_UID="$(id -u)"; INVOKER_GID="$(id -g)" + INVOKER_USER="$(id -un)" +fi + +# In loopback mode the generated sshd config has AllowUsers , so +# a different --user could only ever be refused -- with nothing but +# "Connection reset by peer" to explain it. +if [ -n "$REMOTE_USER" ] && [ -z "$HOST" ]; then + echo "--user needs --host: the throwaway sshd only accepts $INVOKER_USER" >&2 + exit 2 +fi + +# Dropping privileges needs `sudo` itself, so this deliberately does NOT +# go through ${SUDO[@]}: that array is EMPTY when we are already root, +# which is exactly the case that needs the drop. Under +# `sudo eval-under sshfs ...` -- the documented invocation, and the one +# bin/ci/run-under.sh uses -- id -u is 0 while INVOKER_UID is the real +# user, so "${SUDO[@]}" -u would expand to a bare `-u`. bin/eval-under-nfs +# spells out `sudo -u` for the same reason. +as_invoker() { + if [ "$INVOKER_UID" = "$(id -u)" ]; then + "$@" + else + sudo -u "$INVOKER_USER" "$@" + fi +} + +# Dropping privileges is only possible with sudo present, whether or not +# we needed it to become root in the first place. +if [ "$INVOKER_UID" != "$(id -u)" ]; then + command -v sudo >/dev/null || { + echo "ERROR: need sudo to run as the invoking user ($INVOKER_USER)" >&2 + exit 2 + } +fi + +log() { printf '\nI: %s\n' "$*"; } + +command -v sshfs >/dev/null 2>&1 || { + echo "ERROR: sshfs not installed. Install with: apt install sshfs" >&2 + exit 3 +} + +# The scratch dir (throwaway host key, sshd config, authorized_keys and +# the backing directory that gets re-exported) and the mountpoint derive +# from one mktemp -u base, so a run leaves a self-describing pair of +# siblings under $TMPDIR -- the convention every backend here follows +# (see "Adding a new backend" in the README). +MNT_BASE="$(mktemp -u "${TMPDIR:-/tmp}/eval-under-sshfs-XXXXX")" +SCRATCH="$MNT_BASE.scratch" +MNT="${MNT:-$MNT_BASE.sshfs}" +mkdir -p "$SCRATCH" +# 755, not mktemp's 700: sshd reads its config and authorized_keys from +# here as root while the mount and the wrapped command run as the +# invoking user. The mode is not enough on its own -- root created the +# directory, so the invoker also has to own it or every as_invoker write +# into it (the keys, authorized_keys, the backing dir) is EACCES. +chmod 755 "$SCRATCH" +chown "$INVOKER_UID:$INVOKER_GID" "$SCRATCH" +SSHD_PID="" + +# Whether teardown created the mountpoint and may therefore remove it. +MNT_CREATED=0 + +teardown() { + local rc=$? + set +e + if [ "$KEEP" = 1 ]; then + echo "I: --keep set; leaving mount up (exit=$rc)" + echo "I: mount: $MNT" + echo "I: scratch: $SCRATCH" + [ -n "$SSHD_PID" ] && echo "I: sshd pid: $SSHD_PID (kill it yourself)" + return "$rc" + fi + log "teardown" + # Kill the wrapped command first. On a signal it is still running, and + # it must not survive the filesystem it is working on. + pkill -P $$ 2>/dev/null + if command -v fuser >/dev/null 2>&1; then + "${SUDO[@]}" fuser -k -M "$MNT" 2>/dev/null + fi + # fusermount as the mount's owner; fall back to a lazy root umount. + as_invoker fusermount3 -u "$MNT" 2>/dev/null \ + || as_invoker fusermount -u "$MNT" 2>/dev/null \ + || "${SUDO[@]}" umount -l "$MNT" 2>/dev/null + # Kill the whole process group: the daemon forks a child per + # connection, and killing only the listener leaves those behind + # holding the port. + [ -n "$SSHD_PID" ] && { "${SUDO[@]}" kill -- "-$SSHD_PID" 2>/dev/null \ + || "${SUDO[@]}" kill "$SSHD_PID" 2>/dev/null; } + # Only remove the mountpoint if this run made it: an explicit + # --mount-point may well be a directory that predates us. + if [ "$MNT_CREATED" = 1 ]; then + "${SUDO[@]}" rm -rf "$MNT" 2>/dev/null + fi + "${SUDO[@]}" rm -rf "$SCRATCH" 2>/dev/null + return "$rc" +} +# Teardown on a signal too, not just a clean exit: otherwise a SIGTERM +# unmounts and rm -rf's the filesystem out from under a wrapped command +# that is still running, and the orphan keeps writing into a deleted +# inode. on_signal drops the EXIT trap so teardown runs exactly once. +on_signal() { + trap - EXIT + teardown + exit 130 +} +trap teardown EXIT +trap on_signal INT TERM HUP + +# First free TCP port at or above $1 on loopback. `ss` if we have it, +# otherwise a bash /dev/tcp connect probe: a refused connection means +# nothing is listening. +find_free_port() { + local p="$1" limit=$(( $1 + 200 )) + while [ "$p" -lt "$limit" ]; do + if command -v ss >/dev/null 2>&1; then + ss -ltnH "sport = :$p" 2>/dev/null | grep -q . || { echo "$p"; return 0; } + else + if ! (exec 3<>"/dev/tcp/127.0.0.1/$p") 2>/dev/null; then + echo "$p"; return 0 + fi + fi + p=$(( p + 1 )) + done + return 1 +} + +find_sftp_server() { + local p + for p in /usr/lib/openssh/sftp-server \ + /usr/libexec/openssh/sftp-server \ + /usr/lib/ssh/sftp-server \ + /usr/libexec/sftp-server; do + [ -x "$p" ] && { echo "$p"; return 0; } + done + return 1 +} + +# Bring up an sshd that exists only for this run: own port, own host key, +# own authorized_keys, own pid file. Deliberately NOT the system sshd -- +# we would otherwise have to append to the user's authorized_keys and +# hope teardown removes it again. +start_local_sshd() { + local sftp_server sshd_bin + sshd_bin="$(command -v sshd || echo /usr/sbin/sshd)" + [ -x "$sshd_bin" ] || { + echo "ERROR: sshd not found. Install with: apt install openssh-server" >&2 + exit 3 + } + sftp_server="$(find_sftp_server)" || { + echo "ERROR: no sftp-server binary found (openssh-sftp-server?)" >&2 + exit 3 + } + + as_invoker ssh-keygen -t ed25519 -N '' -q -f "$SCRATCH/id" -C eval-under-sshfs + as_invoker ssh-keygen -t ed25519 -N '' -q -f "$SCRATCH/hostkey" -C eval-under-sshfs-host + as_invoker cp "$SCRATCH/id.pub" "$SCRATCH/authorized_keys" + chmod 600 "$SCRATCH/authorized_keys" + + # sshd refuses to start without its privilege-separation directory, + # which is absent in minimal containers (it is created at boot by the + # packaging, not by the package install). + "${SUDO[@]}" mkdir -p /run/sshd + + # Finding a free port and binding it are two steps, and a concurrent + # run can take the port in between -- with four runs on one machine, + # three used to die on "sshd did not start". So treat a failed bind as + # a lost race and try the next free port up, unless the caller pinned + # one with --port. + local tries=0 from=2222 + while [ "$tries" -lt 5 ]; do + tries=$((tries + 1)) + [ -n "$PORT_PINNED" ] || PORT="$(find_free_port "$from")" || { + echo "ERROR: no free port at or above $from" >&2; exit 3; } + write_sshd_config "$sftp_server" + log "starting throwaway sshd on 127.0.0.1:$PORT" + # Run it as root so any AllowUsers target can log in; it is bound to + # loopback and accepts a single ephemeral key. -E keeps its log, which + # is the only place an auth refusal is ever explained. + if "${SUDO[@]}" "$sshd_bin" -f "$SCRATCH/sshd_config" \ + -E "$SCRATCH/sshd.log" && wait_for_sshd_pid; then + log "sshd pid $SSHD_PID (port $PORT)" + return 0 + fi + # Reap anything that did start before we move the port: otherwise it + # survives teardown, because SSHD_PID is still empty. + "${SUDO[@]}" pkill -f "sshd -f $SCRATCH/sshd_config" 2>/dev/null || true + [ -z "$PORT_PINNED" ] || break + from=$((PORT + 1)) + done + + echo "ERROR: sshd did not start" >&2 + dump_sshd_log + exit 3 +} + +# The config is rewritten per attempt because the Port line changes. +write_sshd_config() { + cat > "$SCRATCH/sshd_config" <&2 + sed 's/^/E: /' "$SCRATCH/sshd.log" >&2 +} + +# Assemble the -o list. Ours first, caller's last so theirs wins. +build_opts() { + local opts="reconnect,ServerAliveInterval=15,ServerAliveCountMax=3" + if [ -z "$HOST" ]; then + # Loopback: a throwaway host key we just generated is not worth + # recording, and would collide on the next run. + opts="$opts,IdentityFile=$SCRATCH/id,IdentitiesOnly=yes" + opts="$opts,StrictHostKeyChecking=no,UserKnownHostsFile=/dev/null" + fi + [ "$NO_CACHE" != 0 ] && opts="$opts,cache=no" + [ -n "$WORKAROUND" ] && opts="$opts,workaround=$WORKAROUND" + [ -n "$EXTRA_OPTS" ] && opts="$opts,$EXTRA_OPTS" + echo "$opts" +} + +mount_sshfs() { + local spec opts + opts="$(build_opts)" + + if [ -z "$HOST" ]; then + start_local_sshd + REMOTE_DIR="${REMOTE_DIR:-$SCRATCH/backing}" + as_invoker mkdir -p "$REMOTE_DIR" + spec="${REMOTE_USER:-$INVOKER_USER}@127.0.0.1:$REMOTE_DIR" + else + [ -n "$REMOTE_DIR" ] || { + echo "ERROR: --host requires --remote-dir" >&2; exit 2; } + spec="${REMOTE_USER:-$INVOKER_USER}@$HOST:$REMOTE_DIR" + fi + + [ -d "$MNT" ] || MNT_CREATED=1 + "${SUDO[@]}" mkdir -p "$MNT" + # Only re-own a directory we made. An explicit --mount-point may be + # someone's existing directory, and teardown has no way to put its + # ownership back. + if [ "$MNT_CREATED" = 1 ]; then + "${SUDO[@]}" chown "$INVOKER_UID:$INVOKER_GID" "$MNT" + fi + + # Remote mode never called start_local_sshd, so PORT may still be + # empty: that is plain ssh, which means 22. + log "sshfs -p ${PORT:=22} -o $opts $spec $MNT" + if ! as_invoker sshfs -p "$PORT" -o "$opts" "$spec" "$MNT"; then + echo "ERROR: sshfs mount failed" >&2 + dump_sshd_log + exit 3 + fi + + # An sshfs mount can appear in the mount table before the first + # request round-trips. Touch it once so a failure surfaces here rather + # than inside the wrapped command. + as_invoker test -d "$MNT" || { echo "ERROR: mount not usable" >&2; exit 3; } + "${SUDO[@]}" mount | grep -F "$MNT" | sed 's/^/I: /' +} + +mount_sshfs + +RUN_HOME="$HOME" +if [ "$SET_HOME" = 1 ]; then + RUN_HOME="$MNT/home" + as_invoker mkdir -p "$RUN_HOME" +fi + +if [ "$INVOKER_UID" != "$(id -u)" ]; then + log "running (as uid=$INVOKER_UID): $*" + # Literal sudo, not "${SUDO[@]}" -- same reason as in as_invoker: the + # array is empty precisely when we are root and need the drop. + sudo -u "$INVOKER_USER" -E \ + env HOME="$RUN_HOME" TMPDIR="$MNT" DATALAD_TESTS_TEMP_DIR="$MNT" \ + "$@" +else + log "running: $*" + HOME="$RUN_HOME" TMPDIR="$MNT" DATALAD_TESTS_TEMP_DIR="$MNT" "$@" +fi From 0808a73c6480cf997f197f433be2f4df1c2cc0c6 Mon Sep 17 00:00:00 2001 From: Claude Date: Sat, 26 Sep 2026 00:46:24 +0000 Subject: [PATCH 2/8] Put sshfs in the CI matrix, with two cells rather than four The backend was on-demand only, so nothing watched it. Add it as a matrix row: `sshfs (loopback) / git-annex test` and `sshfs (loopback) / git testsuite`, taking the matrix from 20 cells to 22. Not four cells. stress-ng and pjdfstest are needs-root, and a FUSE mount belongs to whoever mounted it -- there is no --no-root-squash equivalent the way there is for NFS -- so those two could only ever report on privilege rather than on the filesystem. run-under.sh already refuses that pair with exit 2; the matrix now expresses the same rule as data. Rather than a hand-kept exclusion list, the backend row carries `no-root: true` and `cell_enabled()` in matrix.sh derives the gap from it plus the target's existing `needs-root`. So the reason lives in one place, a future unprivileged backend gets the behaviour for free, and the grid cannot silently grow a cell that measures privilege. Every consumer asks `cell_enabled` instead of assuming backends x targets: - matrix-json.sh omits the pair, so the workflow never schedules it; - gen-readme-matrix.sh prints `n/a` instead of a badge (README regenerated); - update-status.py keeps it out of status.json, so it publishes no permanently-unknown badge; - render-report.py renders an explicit gap rather than an "unknown" badge, which would read as "not measured yet". `sshfs-git-annex` is expected red and is annotated as such on the report page and in GOTCHAS.md: invisible hardlinks (link() succeeds, st_ino differs, nlink=1) break `git annex add`'s post-link verification. `sshfs-git` is the control -- same mount, a suite that never hardlinks into an object store. Its result is not predicted here; the first run measures it, and GOTCHAS.md gets the answer either way. Verified: shellcheck clean and 28/28 bats; matrix-json.sh emits 22 cells with exactly the two sshfs ones; update-status.py's cell set agrees and omits sshfs-stress-ng / sshfs-pjdfstest; and render-report.py, run over a fabricated 22-cell status.json, writes 22 cell badges plus overall, shows two n/a gaps in the sshfs row, and carries the expected-red note on sshfs-git-annex. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- README.md | 1 + bin/ci/gen-readme-matrix.sh | 15 ++++++++++++++- bin/ci/matrix-json.sh | 3 +++ bin/ci/matrix.sh | 28 ++++++++++++++++++++++++++++ evals/matrix.yaml | 16 +++++++++++++++- 5 files changed, 61 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index e687c9e..de42ee7 100644 --- a/README.md +++ b/README.md @@ -28,6 +28,7 @@ itself is both backend- and suite-agnostic: new filesystems drop in as | BeeGFS 7.4.6 | [![BeeGFS 7.4.6 / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-7.4.6-git-annex.svg)](https://con.github.io/eval-under/#beegfs-7.4.6-git-annex) | [![BeeGFS 7.4.6 / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-7.4.6-git.svg)](https://con.github.io/eval-under/#beegfs-7.4.6-git) | [![BeeGFS 7.4.6 / stress-ng](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-7.4.6-stress-ng.svg)](https://con.github.io/eval-under/#beegfs-7.4.6-stress-ng) | [![BeeGFS 7.4.6 / pjdfstest](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-7.4.6-pjdfstest.svg)](https://con.github.io/eval-under/#beegfs-7.4.6-pjdfstest) | | BeeGFS 8.1.0 | [![BeeGFS 8.1.0 / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-8.1.0-git-annex.svg)](https://con.github.io/eval-under/#beegfs-8.1.0-git-annex) | [![BeeGFS 8.1.0 / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-8.1.0-git.svg)](https://con.github.io/eval-under/#beegfs-8.1.0-git) | [![BeeGFS 8.1.0 / stress-ng](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-8.1.0-stress-ng.svg)](https://con.github.io/eval-under/#beegfs-8.1.0-stress-ng) | [![BeeGFS 8.1.0 / pjdfstest](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/beegfs-8.1.0-pjdfstest.svg)](https://con.github.io/eval-under/#beegfs-8.1.0-pjdfstest) | | NFS (localhost) | [![NFS (localhost) / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/nfs-git-annex.svg)](https://con.github.io/eval-under/#nfs-git-annex) | [![NFS (localhost) / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/nfs-git.svg)](https://con.github.io/eval-under/#nfs-git) | [![NFS (localhost) / stress-ng](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/nfs-stress-ng.svg)](https://con.github.io/eval-under/#nfs-stress-ng) | [![NFS (localhost) / pjdfstest](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/nfs-pjdfstest.svg)](https://con.github.io/eval-under/#nfs-pjdfstest) | +| sshfs (loopback) | [![sshfs (loopback) / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/sshfs-git-annex.svg)](https://con.github.io/eval-under/#sshfs-git-annex) | [![sshfs (loopback) / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/sshfs-git.svg)](https://con.github.io/eval-under/#sshfs-git) | n/a | n/a | | Loop vfat | [![Loop vfat / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-vfat-git-annex.svg)](https://con.github.io/eval-under/#loop-vfat-git-annex) | [![Loop vfat / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-vfat-git.svg)](https://con.github.io/eval-under/#loop-vfat-git) | [![Loop vfat / stress-ng](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-vfat-stress-ng.svg)](https://con.github.io/eval-under/#loop-vfat-stress-ng) | [![Loop vfat / pjdfstest](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-vfat-pjdfstest.svg)](https://con.github.io/eval-under/#loop-vfat-pjdfstest) | | Loop ext4 | [![Loop ext4 / git-annex test](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-ext4-git-annex.svg)](https://con.github.io/eval-under/#loop-ext4-git-annex) | [![Loop ext4 / git testsuite](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-ext4-git.svg)](https://con.github.io/eval-under/#loop-ext4-git) | [![Loop ext4 / stress-ng](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-ext4-stress-ng.svg)](https://con.github.io/eval-under/#loop-ext4-stress-ng) | [![Loop ext4 / pjdfstest](https://raw.githubusercontent.com/con/eval-under/gh-pages/badges/loop-ext4-pjdfstest.svg)](https://con.github.io/eval-under/#loop-ext4-pjdfstest) | diff --git a/bin/ci/gen-readme-matrix.sh b/bin/ci/gen-readme-matrix.sh index ebc4713..3832dfe 100755 --- a/bin/ci/gen-readme-matrix.sh +++ b/bin/ci/gen-readme-matrix.sh @@ -75,6 +75,12 @@ table_md() { IFS='|' read -r backend version label <<< "$cell" printf '| %s |' "$label" for target in "${EVAL_UNDER_TARGETS[@]}"; do + # A pair that is not a cell gets no badge: there is nothing to + # report, and a badge would imply a result we never measure. + if ! cell_enabled "$backend" "$version" "$target"; then + printf ' n/a |' + continue + fi slug="$(cell_slug "$backend" "$version" "$target")" printf ' [![%s / %s](%s)](%s) |' \ "$label" "$(target_label "$target")" \ @@ -103,4 +109,11 @@ if [ "$CHECK" = 1 ]; then fi cp "$new_readme" "$README" -echo "refreshed the README CI matrix ($((${#EVAL_UNDER_BACKENDS[@]} * ${#EVAL_UNDER_TARGETS[@]})) cells)" +cells=0 +for cell in "${EVAL_UNDER_BACKENDS[@]}"; do + IFS='|' read -r backend version _ <<< "$cell" + for target in "${EVAL_UNDER_TARGETS[@]}"; do + cell_enabled "$backend" "$version" "$target" && cells=$((cells + 1)) + done +done +echo "refreshed the README CI matrix ($cells cells)" diff --git a/bin/ci/matrix-json.sh b/bin/ci/matrix-json.sh index ebf2948..b4b68b6 100755 --- a/bin/ci/matrix-json.sh +++ b/bin/ci/matrix-json.sh @@ -36,6 +36,9 @@ entries=() for cell in "${EVAL_UNDER_BACKENDS[@]}"; do IFS='|' read -r backend version label <<< "$cell" for target in "${EVAL_UNDER_TARGETS[@]}"; do + # Not every backend x target pair is a cell -- see cell_enabled() + # in matrix.sh. + cell_enabled "$backend" "$version" "$target" || continue entries+=("$backend|$version|$label|$target|$(target_label "$target")|$(cell_slug "$backend" "$version" "$target")") done done diff --git a/bin/ci/matrix.sh b/bin/ci/matrix.sh index 3700166..437391c 100755 --- a/bin/ci/matrix.sh +++ b/bin/ci/matrix.sh @@ -65,6 +65,14 @@ out.append("declare -A _EU_NEEDS_ROOT=(%s)" % " ".join( out.append("declare -A _EU_NEEDS_GA=(%s)" % " ".join( "[%s]=%s" % (q(t["name"]), q(int(bool(t["needs-git-annex"])))) for t in targets)) +def bslug(b): + return b["backend"] if b["version"] == "n/a" else "%s-%s" % (b["backend"], b["version"]) + +# Optional per-backend flag, defaulting to false, so existing rows need +# no edit. +out.append("declare -A _EU_BACKEND_NO_ROOT=(%s)" % " ".join( + "[%s]=%s" % (q(bslug(b)), q(int(bool(b.get("no-root", False))))) for b in backends)) + # Env overrides win, so these are defaults only. out.append(": \"${EVAL_UNDER_REPO_SLUG:=%s}\"" % q(d["repo-slug"])) out.append(": \"${EVAL_UNDER_SRC_DIR:=%s}\"" % q(d["src-dir"])) @@ -138,3 +146,23 @@ target_needs_git_annex() { [ "${_EU_NEEDS_GA[$1]:-0}" = 1 ]; } cell_output_dir() { echo "${EVAL_UNDER_OUTPUT_DIR:-/tmp/eval-under-output/$(cell_slug "$1" "$2" "$3")}" } + +# Can this backend hand the wrapped suite privilege at all? Takes a +# backend *slug* (as backend_slug prints it), not a bare backend name. +backend_no_root() { [ "${_EU_BACKEND_NO_ROOT[$1]:-0}" = 1 ]; } + +# Is x a cell the matrix actually defines? +# The grid is deliberately not fully populated: a backend that cannot run +# as root has no cell for a target that needs root, because such a cell +# would measure privilege rather than the filesystem. Every consumer of +# the matrix asks this rather than assuming backends x targets, so the +# workflow, the README grid, the badges and the report page agree on +# which cells exist. +cell_enabled() { + local backend="$1" version="$2" target="$3" + if backend_no_root "$(backend_slug "$backend" "$version")" \ + && target_needs_root "$target"; then + return 1 + fi + return 0 +} diff --git a/evals/matrix.yaml b/evals/matrix.yaml index 350a674..aad8a17 100644 --- a/evals/matrix.yaml +++ b/evals/matrix.yaml @@ -51,7 +51,8 @@ refs: # than a bare GitHub cancellation. # loop-size-mb backing-image size for `eval-under loop --size`. # needs-root the suite is meaningless without privilege, so the -# NFS backend must export no_root_squash for it. +# NFS backend must export no_root_squash for it, and a +# backend marked `no-root` gets no cell for it at all. # pjdfstest is half privileged-vs-unprivileged # assertions and refuses to run otherwise; stress-ng's # chown/mknod stressors need CAP_CHOWN / CAP_MKNOD. @@ -100,6 +101,19 @@ backends: - backend: nfs version: n/a label: NFS (localhost) + + # `no-root` says this backend cannot hand the wrapped suite privilege: + # a FUSE mount belongs to whoever mounted it, and there is no + # --no-root-squash equivalent the way there is for NFS. So this row has + # no cell for a needs-root target (stress-ng, pjdfstest) -- such a cell + # would report on privilege rather than on the filesystem, which is the + # same reason bin/ci/run-under.sh refuses the pair outright. The two + # cells that remain are the ones that found something: git-annex is red + # here by design (invisible hardlinks), git is the control. + - backend: sshfs + version: n/a + label: sshfs (loopback) + no-root: true - backend: loop version: vfat label: Loop vfat From 925e507848f5881d1d462d0da8d86deddf110314 Mon Sep 17 00:00:00 2001 From: Claude Date: Sat, 26 Sep 2026 01:23:17 +0000 Subject: [PATCH 3/8] Characterise the sshfs git cell, and stop counting TODOs as failures The `sshfs / git testsuite` cell came back red, and the previous commit had called it "the control: a suite that never hardlinks into an object store". That was wrong, and measuring it turned up a bug in our own reporting. Root cause of the cell: `git clone ` hardlinks each object and then sanity-checks the result, comparing st_mode/st_ino/st_dev/ st_size/st_uid/st_gid against the source (builtin/clone.c). sshfs synthesises st_ino per path, so the check fails and the clone dies with "hardlink different from source". Plain `git clone` of a local path does not work on sshfs at all -- a broader statement than the git-annex one, and the same mechanism. It accounts for t1507-rev-parse-upstream (20), t1013-read-tree-submodule (58) and t0035-safe-bare-repository (2), each of which clones or adds a submodule in setup. A second, independent cause is the absence of unix sockets: t0301-credential-cache (37, "unable to bind ... Operation not permitted") and t0052-simple-ipc (9/9). `-o disable_hardlink` fixes both this and the git-annex failure, which is the counter-intuitive part worth telling a reporter: sshfs then fails link() with EPERM instead of pretending, and every caller here has a copy fallback that only an honest failure reaches. Measured, with an ext4 control: default disable_hardlink ext4 git annex add (locked) ok ok ok git annex add, annex.addunlocked=true FAIL ok ok git clone FAIL ok ok The reporting bug: `not ok N ... # TODO known breakage` is git's test_expect_failure -- a TAP TODO directive that prove counts as an expected result, not a failure. dump-failure-logs.sh counted those, so this cell reported 351 failed assertions where prove saw 166, and the two worst-looking scripts (t1517-outside-repo at 104, t0450-txt-doc-vs-help at 54) are ones prove reports as *ok*. That sends triage after failures that do not exist; it inflated every git cell, vfat included, not just this one. Both the count and the detail listing now exclude the directive, matching prove. GOTCHAS.md gets the mechanisms, the workaround table, a single-script reproduction recipe, and -- deliberately -- the list of small failures that are *not* explained (t0003-attributes 6, t1091 8, t0610 3, t0021 3, ten scripts with one each), so nobody assumes they share a cause. Method: built the pinned git v2.55.0 locally and ran the whole t0*/t1* range under the sshfs backend (174 scripts, 10366 assertions) plus an ext4 control, which passes the five scripts under suspicion. Reproduced the clone failure minimally, read git's check in clone.c, and confirmed link() returns EPERM under disable_hardlink. One hypothesis was tested and rejected rather than left implied: synthesised inode numbers do not collide at 4000 entries, so that is not behind the whole-suite failures. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- bin/ci/dump-failure-logs.sh | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/bin/ci/dump-failure-logs.sh b/bin/ci/dump-failure-logs.sh index 6253b18..23b1215 100755 --- a/bin/ci/dump-failure-logs.sh +++ b/bin/ci/dump-failure-logs.sh @@ -74,10 +74,19 @@ if [ "$TARGET" = "git" ]; then echo "=== git testsuite failures ($results) ===" if [ -d "$results" ]; then # Pass 1: which scripts failed, and how badly. + # + # `not ok N ... # TODO known breakage` is git's test_expect_failure: + # a TAP TODO directive, which prove counts as an expected result and + # not as a failure -- a script whose only "not ok" lines are TODOs is + # reported ok by the harness. Counting them here inflated every git + # cell (the sshfs cell read 351 failed assertions where prove saw + # 166, and two scripts prove called ok appeared as the worst + # offenders at 104 and 54), which sends triage after failures that do + # not exist. Exclude the directive, and match prove. names=() counts=() for out in "$results"/*.out; do [ -e "$out" ] || continue - n="$(grep -c '^not ok ' "$out" 2>/dev/null || true)" + n="$(grep '^not ok ' "$out" 2>/dev/null | grep -vc '# TODO' || true)" [ "${n:-0}" -gt 0 ] || continue names+=("$(basename "${out%.out}")") counts+=("$n") @@ -94,7 +103,7 @@ if [ "$TARGET" = "git" ]; then shown=$((shown + 1)) out="$results/${names[$i]}.out" echo "--- ${names[$i]}: ${counts[$i]} failed ---" - grep '^not ok ' "$out" | head -40 || true + grep '^not ok ' "$out" | grep -v '# TODO' | head -40 || true echo " ... last $GIT_DUMP_TAIL_LINES lines of ${names[$i]}.out:" tail -"$GIT_DUMP_TAIL_LINES" "$out" | sed 's/^/ | /' || true echo From 80dd4f405dc54e820d3e128047fd32d57a7744eb Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 19:58:36 +0000 Subject: [PATCH 4/8] Teach the shared matrix model which pairs are not cells The rebase dropped my copies of this rule from update-status.py and render-report.py: master now enumerates cells once, in matrix_cells() in bin/ci/evals.py, which is a better home for it than the two duplicates the pre-rebase branch had. cell_enabled() states the rule -- a `no-root` backend has no cell for a `needs-root` target, because such a cell could only report on privilege rather than on the filesystem -- and matrix_cells() skips those pairs, so status.json, the badges and the report page agree with the workflow instead of publishing a permanently-unknown badge. render-report.py draws an explicit `n/a` gap for them; an "unknown" badge would read as "not measured yet". The shell side already has the same rule in matrix.sh, and run-under.sh refuses the pair outright. Verified: both sides report 22 cells with exactly sshfs-git-annex and sshfs-git, and neither reports sshfs-stress-ng or sshfs-pjdfstest; the README grid shows two n/a entries; run-checks.sh passes in full. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- bin/ci/evals.py | 23 ++++++++++++++++++++++- bin/ci/render-report.py | 18 ++++++++++++------ 2 files changed, 34 insertions(+), 7 deletions(-) diff --git a/bin/ci/evals.py b/bin/ci/evals.py index 38121fd..6e0e88f 100644 --- a/bin/ci/evals.py +++ b/bin/ci/evals.py @@ -31,12 +31,33 @@ def load_matrix(path: Path = MATRIX_FILE) -> dict: return yaml.safe_load(fh) +def cell_enabled(b: dict, t: dict) -> bool: + """Is this backend x target pair a cell the matrix actually defines? + + The grid is deliberately not fully populated. A backend marked + `no-root` cannot hand the wrapped suite privilege -- a FUSE mount + belongs to whoever mounted it, and there is no --no-root-squash + equivalent the way there is for NFS -- so it has no cell for a + `needs-root` target: the cell could only report on privilege rather + than on the filesystem. bin/ci/run-under.sh refuses the same pair + outright. Keep in sync with cell_enabled() in bin/ci/matrix.sh. + """ + return not (b.get("no-root") and t.get("needs-root")) + + def matrix_cells(m: dict) -> dict[str, dict]: - """slug -> cell metadata, in matrix (row, column) order.""" + """slug -> cell metadata, in matrix (row, column) order. + + Skips the pairs cell_enabled() rules out, so every consumer -- the + status file, the badges, the report page -- agrees on which cells + exist instead of publishing a permanently-unknown one. + """ cells = {} for b in m["backends"]: bslug = backend_slug(b["backend"], b["version"]) for t in m["targets"]: + if not cell_enabled(b, t): + continue cells[f"{bslug}-{t['name']}"] = { "backend": b["backend"], "version": b["version"], diff --git a/bin/ci/render-report.py b/bin/ci/render-report.py index 405e74d..88dca65 100755 --- a/bin/ci/render-report.py +++ b/bin/ci/render-report.py @@ -28,7 +28,7 @@ from pathlib import Path import known_issues -from evals import backend_slug, load_matrix +from evals import backend_slug, cell_enabled, load_matrix HERE = Path(__file__).resolve().parent @@ -167,8 +167,8 @@ def main() -> int: # Matrix order, so the page reads like the README grid. m = load_matrix() - backends = [(backend_slug(b["backend"], b["version"]), b["label"]) for b in m["backends"]] - targets = [(t["name"], t["label"]) for t in m["targets"]] + backends = [(backend_slug(b["backend"], b["version"]), b["label"], b) for b in m["backends"]] + targets = [(t["name"], t["label"], t) for t in m["targets"]] _, issues = known_issues.load_valid() def state(c: dict) -> str: @@ -186,9 +186,15 @@ def state(c: dict) -> str: f"{npass}/{total} passing" + (f", {nnew} unexpected" if nnew else "")) rows = [] - for bslug, blabel in backends: + for bslug, blabel, bdef in backends: tds = [f"{html.escape(blabel)}"] - for tname, tlabel in targets: + for tname, tlabel, tdef in targets: + # A pair the matrix does not define (see cell_enabled) is an + # explicit gap, not an "unknown" badge: nothing was measured + # here and nothing ever will be. + if not cell_enabled(bdef, tdef): + tds.append('n/a') + continue slug = f"{bslug}-{tname}" c = cells.get(slug, {"conclusion": "unknown"}) st = state(c) @@ -214,7 +220,7 @@ def state(c: dict) -> str: tds.append(f'{body}{meta}{cell_notes(c, st, fixed)}') rows.append("" + "".join(tds) + "") - head = "".join(f"{html.escape(l)}" for _, l in targets) + head = "".join(f"{html.escape(l)}" for _, l, _ in targets) repo = m.get("repo-slug", "con/eval-under") doc = f""" From c0ccb1476f2bd52180bf21f21b261a3e9ff88ba2 Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 20:01:47 +0000 Subject: [PATCH 5/8] Record the sshfs git failures as known issues The rebase put this branch on master's known-issues machinery, which judges a cell per test rather than by the suite's exit code. The two new sshfs cells had no entries, so they would have reported `failing-new` even though their failures are the documented findings this backend exists to surface. The prose that used to say so by hand in GOTCHAS.md belongs in evals/known-issues.yaml now, where the verdict, the status page and GOTCHAS all read it from one place. Measured, not asserted: ran the cell here (`run-under.sh sshfs n/a git` then `check-cell.sh`), which reported 185 failures across 30 scripts, and attributed the scripts by the message in their own logs rather than by guessing: - `sshfs-git-local-clone-hardlink` (114 failures, 16 scripts, fs-divergence) -- `git clone ` hardlinks each object then compares st_ino/st_dev/... against the source (builtin/clone.c), and sshfs synthesises st_ino per path, so the clone dies with `fatal: hardlink different from source`. Every script here clones or adds a submodule in setup. Plain `git clone` of a local path does not work on sshfs at all; --no-hardlinks, file:// and `-o disable_hardlink` do. - `sshfs-git-no-unix-sockets` (46, fs-limitation) -- `bind()` on the mount fails with EPERM, so the credential-cache daemon never starts and simple-ipc finds no server. Mirrors the existing vfat entry; `unix-socket=no` in fs-capabilities.sh predicts it. - `sshfs-git-untriaged` (25, needs-triage) -- kept deliberately separate: neither signature appears in these scripts' logs, so folding them into the mechanisms above would be a guess. Candidates noted (1-second mtime granularity, no xattr/fifo). With these, the cell judges `failing-known` -> conclusion=success: 9734 pass, 185 fail all known, 0 new. GOTCHAS.md's generated section is regenerated from them, and run-checks.sh passes in full. Still outstanding: `sshfs (loopback) / git-annex test` has no entry yet, so it will report `failing-new` on its first run here. Its test ids are best taken from CI's own verdict -- CI runs a git-annex daily build, this container has 10.20240129, and ids from the wrong build would mask real failures rather than document them. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- GOTCHAS.md | 55 +++++++++++++++++++++++++++ evals/known-issues.yaml | 83 +++++++++++++++++++++++++++++++++++++++++ 2 files changed, 138 insertions(+) diff --git a/GOTCHAS.md b/GOTCHAS.md index 477e98a..1435a1d 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -316,6 +316,61 @@ See: , [BeeGFS See: [Loop git-annex cells: annex.diskreserve](#loop-git-annex-cells-annexdiskreserve) + +### `sshfs-git-local-clone-hardlink`: local `git clone` verifies its hardlinks, and sshfs synthesises st_ino + +**Cells:** `sshfs-git` \ +**Tags:** `fs-divergence` \ +**Tests:** `t0001-init.sh#28,37`, `t0003-attributes.sh#24-25,29,32-34,41`, `t0021-conversion.sh#28-30`, `t0033-safe-directory.sh#16`, `t0035-safe-bare-repository.sh#1,13`, `t0410-partial-clone.sh#34,38`, `t0610-reftable-basics.sh#26-28,48`, `t1013-read-tree-submodule.sh#1-8,10-15,18-28,30-48,51-60,65-68`, `t1060-object-corruption.sh#12`, `t1091-sparse-checkout-builtin.sh#31-32,36-38,47,49,77`, `t1350-config-hooks-path.sh#4`, `t1423-ref-backend.sh#36`, `t1460-refs-migrate.sh#9,24`, `t1500-rev-parse.sh#77`, `t1507-rev-parse-upstream.sh#1-7,9-11,13-14,17-18,21,23-27`, `t1600-index.sh#6` + +SFTP's `ATTRS` carries no inode number, so sshfs synthesises `st_ino` +per path. `git clone ` hardlinks each object and then +compares `st_mode`/`st_ino`/`st_dev`/`st_size`/`st_uid`/`st_gid` +against the source (`builtin/clone.c`), so the check fails and the +clone dies with `fatal: hardlink different from source`. **Plain +`git clone` of a local path does not work on sshfs at all.** + +Each script here clones, or adds a submodule (which clones), in its +setup, so one failed setup cascades through the script; the scripts +were attributed by that message appearing in their own logs. +`git clone --no-hardlinks`, `git clone file://...` and mounting with +`-o disable_hardlink` all work -- with the option, sshfs fails +`link()` with `EPERM` instead of pretending, and git falls back to +copying. + +See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) + + +### `sshfs-git-no-unix-sockets`: git's IPC and credential-cache tests need Unix sockets on the work tree + +**Cells:** `sshfs-git` \ +**Tags:** `fs-limitation` \ +**Tests:** `t0052-simple-ipc.sh#1-9`, `t0301-credential-cache.sh#2-3,7-8,10-11,13-23,25-26,28,30,32-33,37-38,40-41,43-52` + +sshfs has no Unix sockets: `bind()` on the mount fails with +`Operation not permitted`, so the credential-cache daemon never +starts (`unable to bind to .../credential/socket`) and simple-ipc +finds `no server listening`. `unix-socket=no` in +`bin/ci/fs-capabilities.sh` predicts both. The same limitation is +recorded for vfat as `vfat-git-no-unix-sockets`. + +See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) + + +### `sshfs-git-untriaged`: remaining git failures on sshfs, not yet attributed + +**Cells:** `sshfs-git` \ +**Tags:** `needs-triage` \ +**Tests:** `t0017-env-helper.sh#4`, `t0027-auto-crlf.sh#2076,2078,2081,2096,2102-2104`, `t0450-txt-doc-vs-help.sh#167,347`, `t1002-read-tree-m-u-2way.sh#17,22`, `t1004-read-tree-m-u-wf.sh#9-10,12,15-17`, `t1092-sparse-checkout-compatibility.sh#55`, `t1300-config.sh#435`, `t1301-shared-repo.sh#17`, `t1410-reflog.sh#21`, `t1430-bad-ref-name.sh#17`, `t1461-refs-list.sh#398`, `t1700-split-index.sh#9` + +Acknowledged, root cause not run down. Deliberately kept apart from +the two mechanisms above rather than folded into them: neither the +hardlink message nor a socket refusal appears in these scripts' +logs. Candidates worth checking first are sshfs's 1-second mtime +granularity and its lack of `xattr`/`fifo` support. + +See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) + ## Root-cause notes diff --git a/evals/known-issues.yaml b/evals/known-issues.yaml index ef52a2e..4d295da 100644 --- a/evals/known-issues.yaml +++ b/evals/known-issues.yaml @@ -166,3 +166,86 @@ issues: tests: ["*"] tags: [harness, needs-triage] links: ["GOTCHAS.md#loop-git-annex-cells-annexdiskreserve"] + + # --------------------------------------------------------------- sshfs + - id: sshfs-git-local-clone-hardlink + title: local `git clone` verifies its hardlinks, and sshfs synthesises st_ino + backends: [sshfs] + targets: [git] + tests: + - "t0001-init.sh#28,37" + - "t0003-attributes.sh#24-25,29,32-34,41" + - "t0021-conversion.sh#28-30" + - "t0033-safe-directory.sh#16" + - "t0035-safe-bare-repository.sh#1,13" + - "t0410-partial-clone.sh#34,38" + - "t0610-reftable-basics.sh#26-28,48" + - "t1013-read-tree-submodule.sh#1-8,10-15,18-28,30-48,51-60,65-68" + - "t1060-object-corruption.sh#12" + - "t1091-sparse-checkout-builtin.sh#31-32,36-38,47,49,77" + - "t1350-config-hooks-path.sh#4" + - "t1423-ref-backend.sh#36" + - "t1460-refs-migrate.sh#9,24" + - "t1500-rev-parse.sh#77" + - "t1507-rev-parse-upstream.sh#1-7,9-11,13-14,17-18,21,23-27" + - "t1600-index.sh#6" + tags: [fs-divergence] + links: ["GOTCHAS.md#sshfs-bineval-under-sshfs"] + notes: | + SFTP's `ATTRS` carries no inode number, so sshfs synthesises `st_ino` + per path. `git clone ` hardlinks each object and then + compares `st_mode`/`st_ino`/`st_dev`/`st_size`/`st_uid`/`st_gid` + against the source (`builtin/clone.c`), so the check fails and the + clone dies with `fatal: hardlink different from source`. **Plain + `git clone` of a local path does not work on sshfs at all.** + + Each script here clones, or adds a submodule (which clones), in its + setup, so one failed setup cascades through the script; the scripts + were attributed by that message appearing in their own logs. + `git clone --no-hardlinks`, `git clone file://...` and mounting with + `-o disable_hardlink` all work -- with the option, sshfs fails + `link()` with `EPERM` instead of pretending, and git falls back to + copying. + + - id: sshfs-git-no-unix-sockets + title: git's IPC and credential-cache tests need Unix sockets on the work tree + backends: [sshfs] + targets: [git] + tests: + - "t0052-simple-ipc.sh#1-9" + - "t0301-credential-cache.sh#2-3,7-8,10-11,13-23,25-26,28,30,32-33,37-38,40-41,43-52" + tags: [fs-limitation] + links: ["GOTCHAS.md#sshfs-bineval-under-sshfs"] + notes: | + sshfs has no Unix sockets: `bind()` on the mount fails with + `Operation not permitted`, so the credential-cache daemon never + starts (`unable to bind to .../credential/socket`) and simple-ipc + finds `no server listening`. `unix-socket=no` in + `bin/ci/fs-capabilities.sh` predicts both. The same limitation is + recorded for vfat as `vfat-git-no-unix-sockets`. + + - id: sshfs-git-untriaged + title: remaining git failures on sshfs, not yet attributed + backends: [sshfs] + targets: [git] + tests: + - "t0017-env-helper.sh#4" + - "t0027-auto-crlf.sh#2076,2078,2081,2096,2102-2104" + - "t0450-txt-doc-vs-help.sh#167,347" + - "t1002-read-tree-m-u-2way.sh#17,22" + - "t1004-read-tree-m-u-wf.sh#9-10,12,15-17" + - "t1092-sparse-checkout-compatibility.sh#55" + - "t1300-config.sh#435" + - "t1301-shared-repo.sh#17" + - "t1410-reflog.sh#21" + - "t1430-bad-ref-name.sh#17" + - "t1461-refs-list.sh#398" + - "t1700-split-index.sh#9" + tags: [needs-triage] + links: ["GOTCHAS.md#sshfs-bineval-under-sshfs"] + notes: | + Acknowledged, root cause not run down. Deliberately kept apart from + the two mechanisms above rather than folded into them: neither the + hardlink message nor a socket refusal appears in these scripts' + logs. Candidates worth checking first are sshfs's 1-second mtime + granularity and its lack of `xattr`/`fifo` support. From 7adc041dc0275f54846ce8b2f5dba9ebb3207bcd Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 21:20:22 +0000 Subject: [PATCH 6/8] Seed the sshfs git known issues from CI's verdict, not a local run The first draft of these three entries was derived from a run in this development container, and CI disagreed with it: `sshfs (loopback) / git testsuite` reported `failing-new` with 11 failures no entry covered, while 28 of the ids I had listed passed there (4 under `sshfs-git-local-clone-hardlink`, 24 of the 25 under `sshfs-git-untriaged`). The mechanisms were right; the assertion numbers were not mine to guess. The container runs as root. That both flips permission-dependent assertions and renumbers the scripts that define tests conditionally, so a locally derived `script#N` can name a different assertion than the same `script#N` on a runner -- which is exactly what happened: my `t0003-attributes.sh#41` passes in CI while `#48` fails, and `t0450`'s `#167,347` became `#797`. This is the same reason the git-annex entry is still outstanding rather than invented here, applied to the cell where I had not applied it. Reseeded from run 36476334300: - `sshfs-git-local-clone-hardlink`: 114 -> 110 ids, dropping t0001#28, t0003#41, t0610#48 and t1091#47. - `sshfs-git-no-unix-sockets`: unchanged. It reproduced exactly, 46/46, which is what a mechanism that owns two whole scripts should do. - `sshfs-git-untriaged`: 25 -> 12 ids. Only t1092#55 survived; the 11 new failures join it. Three are now named from their own assertions, since CI's log quotes them: t0003#48 "builtin object mode attributes work (dir and regular paths)", t0061#6 "run_command can run a script without a #! line", t0061#18 "run_command is asked to abort gracefully" -- so the mode bits execve() and builtin_objectmode read join mtime granularity and xattr/fifo as the candidates to check first. Also recorded: t0003#48 is in `vfat-git-untriaged` too, and the t0061 and t1091/t1300 assertions here sit beside the ones listed there, so some of this is likely one non-POSIX cause shared with vfat rather than anything sshfs invented. Worth noting these are sshfs's own: the same run's `NFS (localhost) / git testsuite` is green with no nfs git entries at all, so none of the 11 is an environment or build artefact. Verified by replaying CI's exact failure set (all 168 ids, reconstructed and cross-checked per script against the job's own summary table -- zero mismatches) through `known_issues.py check`: `failing-known`, known 168, new 0, no now-passing notices, exit 0, where before it was `failing-new` with 11. `run-checks.sh` passes in full, and GOTCHAS.md is regenerated. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- GOTCHAS.md | 26 +++++++++++++++++++--- evals/known-issues.yaml | 49 +++++++++++++++++++++++++++-------------- 2 files changed, 56 insertions(+), 19 deletions(-) diff --git a/GOTCHAS.md b/GOTCHAS.md index 1435a1d..a6182fd 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -321,7 +321,7 @@ See: [Loop git-annex cells: annex.diskreserve](#loop-git-annex-cells-annexdiskre **Cells:** `sshfs-git` \ **Tags:** `fs-divergence` \ -**Tests:** `t0001-init.sh#28,37`, `t0003-attributes.sh#24-25,29,32-34,41`, `t0021-conversion.sh#28-30`, `t0033-safe-directory.sh#16`, `t0035-safe-bare-repository.sh#1,13`, `t0410-partial-clone.sh#34,38`, `t0610-reftable-basics.sh#26-28,48`, `t1013-read-tree-submodule.sh#1-8,10-15,18-28,30-48,51-60,65-68`, `t1060-object-corruption.sh#12`, `t1091-sparse-checkout-builtin.sh#31-32,36-38,47,49,77`, `t1350-config-hooks-path.sh#4`, `t1423-ref-backend.sh#36`, `t1460-refs-migrate.sh#9,24`, `t1500-rev-parse.sh#77`, `t1507-rev-parse-upstream.sh#1-7,9-11,13-14,17-18,21,23-27`, `t1600-index.sh#6` +**Tests:** `t0001-init.sh#37`, `t0003-attributes.sh#24-25,29,32-34`, `t0021-conversion.sh#28-30`, `t0033-safe-directory.sh#16`, `t0035-safe-bare-repository.sh#1,13`, `t0410-partial-clone.sh#34,38`, `t0610-reftable-basics.sh#26-28`, `t1013-read-tree-submodule.sh#1-8,10-15,18-28,30-48,51-60,65-68`, `t1060-object-corruption.sh#12`, `t1091-sparse-checkout-builtin.sh#31-32,36-38,49,77`, `t1350-config-hooks-path.sh#4`, `t1423-ref-backend.sh#36`, `t1460-refs-migrate.sh#9,24`, `t1500-rev-parse.sh#77`, `t1507-rev-parse-upstream.sh#1-7,9-11,13-14,17-18,21,23-27`, `t1600-index.sh#6` SFTP's `ATTRS` carries no inode number, so sshfs synthesises `st_ino` per path. `git clone ` hardlinks each object and then @@ -338,6 +338,9 @@ were attributed by that message appearing in their own logs. `link()` with `EPERM` instead of pretending, and git falls back to copying. +Test ids seeded from run 36476334300 (git v2.55.0), 110 assertions +across 16 scripts. + See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) @@ -361,13 +364,30 @@ See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) **Cells:** `sshfs-git` \ **Tags:** `needs-triage` \ -**Tests:** `t0017-env-helper.sh#4`, `t0027-auto-crlf.sh#2076,2078,2081,2096,2102-2104`, `t0450-txt-doc-vs-help.sh#167,347`, `t1002-read-tree-m-u-2way.sh#17,22`, `t1004-read-tree-m-u-wf.sh#9-10,12,15-17`, `t1092-sparse-checkout-compatibility.sh#55`, `t1300-config.sh#435`, `t1301-shared-repo.sh#17`, `t1410-reflog.sh#21`, `t1430-bad-ref-name.sh#17`, `t1461-refs-list.sh#398`, `t1700-split-index.sh#9` +**Tests:** `t0003-attributes.sh#48`, `t0061-run-command.sh#6,18`, `t0302-credential-store.sh#57`, `t0450-txt-doc-vs-help.sh#797`, `t1091-sparse-checkout-builtin.sh#48`, `t1092-sparse-checkout-compatibility.sh#55`, `t1300-config.sh#194,285,494`, `t1403-show-ref.sh#9`, `t1450-fsck.sh#36` Acknowledged, root cause not run down. Deliberately kept apart from the two mechanisms above rather than folded into them: neither the hardlink message nor a socket refusal appears in these scripts' logs. Candidates worth checking first are sshfs's 1-second mtime -granularity and its lack of `xattr`/`fifo` support. +granularity, its lack of `xattr`/`fifo` support, and how it reports +the mode bits `execve()` and `builtin_objectmode` read -- three of +these are named from their own assertions: +`t0003#48` "builtin object mode attributes work (dir and regular +paths)", `t0061#6` "run_command can run a script without a #! line" +and `t0061#18` "run_command is asked to abort gracefully". + +`t0003-attributes.sh#48` is also in `vfat-git-untriaged`, and this +list's `t0061` and `t1091`/`t1300` assertions sit next to the ones +recorded there, so some of these are likely one non-POSIX cause +shared with vfat rather than anything sshfs invented. + +Test ids seeded from run 36476334300 (git v2.55.0): 12 assertions, +of which 11 were not in this file's first draft. That draft was +derived from a local run in a container that runs as **root**, +which both flips permission-dependent assertions and renumbers the +scripts that define tests conditionally -- so its ids named partly +different assertions. CI's verdict is the one to seed from. See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) diff --git a/evals/known-issues.yaml b/evals/known-issues.yaml index 4d295da..223b04d 100644 --- a/evals/known-issues.yaml +++ b/evals/known-issues.yaml @@ -173,16 +173,16 @@ issues: backends: [sshfs] targets: [git] tests: - - "t0001-init.sh#28,37" - - "t0003-attributes.sh#24-25,29,32-34,41" + - "t0001-init.sh#37" + - "t0003-attributes.sh#24-25,29,32-34" - "t0021-conversion.sh#28-30" - "t0033-safe-directory.sh#16" - "t0035-safe-bare-repository.sh#1,13" - "t0410-partial-clone.sh#34,38" - - "t0610-reftable-basics.sh#26-28,48" + - "t0610-reftable-basics.sh#26-28" - "t1013-read-tree-submodule.sh#1-8,10-15,18-28,30-48,51-60,65-68" - "t1060-object-corruption.sh#12" - - "t1091-sparse-checkout-builtin.sh#31-32,36-38,47,49,77" + - "t1091-sparse-checkout-builtin.sh#31-32,36-38,49,77" - "t1350-config-hooks-path.sh#4" - "t1423-ref-backend.sh#36" - "t1460-refs-migrate.sh#9,24" @@ -207,6 +207,9 @@ issues: `link()` with `EPERM` instead of pretending, and git falls back to copying. + Test ids seeded from run 36476334300 (git v2.55.0), 110 assertions + across 16 scripts. + - id: sshfs-git-no-unix-sockets title: git's IPC and credential-cache tests need Unix sockets on the work tree backends: [sshfs] @@ -229,18 +232,15 @@ issues: backends: [sshfs] targets: [git] tests: - - "t0017-env-helper.sh#4" - - "t0027-auto-crlf.sh#2076,2078,2081,2096,2102-2104" - - "t0450-txt-doc-vs-help.sh#167,347" - - "t1002-read-tree-m-u-2way.sh#17,22" - - "t1004-read-tree-m-u-wf.sh#9-10,12,15-17" + - "t0003-attributes.sh#48" + - "t0061-run-command.sh#6,18" + - "t0302-credential-store.sh#57" + - "t0450-txt-doc-vs-help.sh#797" + - "t1091-sparse-checkout-builtin.sh#48" - "t1092-sparse-checkout-compatibility.sh#55" - - "t1300-config.sh#435" - - "t1301-shared-repo.sh#17" - - "t1410-reflog.sh#21" - - "t1430-bad-ref-name.sh#17" - - "t1461-refs-list.sh#398" - - "t1700-split-index.sh#9" + - "t1300-config.sh#194,285,494" + - "t1403-show-ref.sh#9" + - "t1450-fsck.sh#36" tags: [needs-triage] links: ["GOTCHAS.md#sshfs-bineval-under-sshfs"] notes: | @@ -248,4 +248,21 @@ issues: the two mechanisms above rather than folded into them: neither the hardlink message nor a socket refusal appears in these scripts' logs. Candidates worth checking first are sshfs's 1-second mtime - granularity and its lack of `xattr`/`fifo` support. + granularity, its lack of `xattr`/`fifo` support, and how it reports + the mode bits `execve()` and `builtin_objectmode` read -- three of + these are named from their own assertions: + `t0003#48` "builtin object mode attributes work (dir and regular + paths)", `t0061#6` "run_command can run a script without a #! line" + and `t0061#18` "run_command is asked to abort gracefully". + + `t0003-attributes.sh#48` is also in `vfat-git-untriaged`, and this + list's `t0061` and `t1091`/`t1300` assertions sit next to the ones + recorded there, so some of these are likely one non-POSIX cause + shared with vfat rather than anything sshfs invented. + + Test ids seeded from run 36476334300 (git v2.55.0): 12 assertions, + of which 11 were not in this file's first draft. That draft was + derived from a local run in a container that runs as **root**, + which both flips permission-dependent assertions and renumbers the + scripts that define tests conditionally -- so its ids named partly + different assertions. CI's verdict is the one to seed from. From a4f4da73b0ee813ea8f7332b5570ba58049c6355 Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 21:54:39 +0000 Subject: [PATCH 7/8] sshfs git residue: it is flaky, so stop pinning ids and say so My previous commit reseeded these entries from CI run 36476334300 and predicted the cell would go green. It did not: run 36485450259, on near-identical code, reported 18 residual failures with **no overlap** with the 12 that run reported. Every id I had just added passed, and ids I had just removed as "now passing" -- t0003#41, t0017#4, t1430#26, t1700#9 -> #10-15 -- were back. The residue is not a fixed set of divergences, so no `script#N` list can cover it, and the explanation I committed last time (root renumbering the scripts) was wrong: locally t0003 yields 24,25,29,32,33,34, exactly CI's ids. What is actually true, measured: - The two mechanism entries are deterministic. `local-clone-hardlink` reproduced 110/110 and `no-unix-sockets` 46/46 on both CI runs, and t0003's six hardlink assertions failed in all six local runs. - The residue is flaky at a low per-test rate. Six local runs of `t0003 t0017 t0040 t1700`: t0017#4 failed in one, t0040#37 and t1700#10-15 and t0003#41/#48 in none, though CI has flagged each. - It needs concurrent load. `prove --jobs 1` is no cleaner than `--jobs 4`, and 14 runs of t0017 on its own were all clean. - Some of it cannot be filesystem semantics at all: t0040#37 "OPT_CALLBACK() and OPT_BIT() work" and t0017#4 "test-tool env-helper --type=ulong" parse arguments and environment variables, and t0450 compares documentation against `-h` output. What they share is capturing output into `>out`/`2>err` in the trash directory on the mount and grepping it, which points at the mount losing or delaying writes under load. So this commit stops pretending the list gates anything. The entry now carries the union observed across both CI runs, labelled as a sample, plus the measurements above and the untried candidate that would actually resolve it: mounting with `-o attr_timeout=0 -o entry_timeout=0` (possibly with `-o max_conns=N` or `-o sync_read`) and re-measuring, which if it works makes the cell deterministic and properly xfail-able. This does NOT make the cell green -- it will still report `failing-new` whenever a run draws a flake outside the sample, and that is now the honest state rather than a hidden one. run-checks.sh passes and GOTCHAS.md is regenerated. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- GOTCHAS.md | 59 ++++++++++++++++++++--------------- evals/known-issues.yaml | 68 ++++++++++++++++++++++++++--------------- 2 files changed, 79 insertions(+), 48 deletions(-) diff --git a/GOTCHAS.md b/GOTCHAS.md index a6182fd..cb39c7e 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -364,30 +364,41 @@ See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) **Cells:** `sshfs-git` \ **Tags:** `needs-triage` \ -**Tests:** `t0003-attributes.sh#48`, `t0061-run-command.sh#6,18`, `t0302-credential-store.sh#57`, `t0450-txt-doc-vs-help.sh#797`, `t1091-sparse-checkout-builtin.sh#48`, `t1092-sparse-checkout-compatibility.sh#55`, `t1300-config.sh#194,285,494`, `t1403-show-ref.sh#9`, `t1450-fsck.sh#36` - -Acknowledged, root cause not run down. Deliberately kept apart from -the two mechanisms above rather than folded into them: neither the -hardlink message nor a socket refusal appears in these scripts' -logs. Candidates worth checking first are sshfs's 1-second mtime -granularity, its lack of `xattr`/`fifo` support, and how it reports -the mode bits `execve()` and `builtin_objectmode` read -- three of -these are named from their own assertions: -`t0003#48` "builtin object mode attributes work (dir and regular -paths)", `t0061#6` "run_command can run a script without a #! line" -and `t0061#18` "run_command is asked to abort gracefully". - -`t0003-attributes.sh#48` is also in `vfat-git-untriaged`, and this -list's `t0061` and `t1091`/`t1300` assertions sit next to the ones -recorded there, so some of these are likely one non-POSIX cause -shared with vfat rather than anything sshfs invented. - -Test ids seeded from run 36476334300 (git v2.55.0): 12 assertions, -of which 11 were not in this file's first draft. That draft was -derived from a local run in a container that runs as **root**, -which both flips permission-dependent assertions and renumbers the -scripts that define tests conditionally -- so its ids named partly -different assertions. CI's verdict is the one to seed from. +**Tests:** `t0003-attributes.sh#41,48`, `t0017-env-helper.sh#4`, `t0040-parse-options.sh#37`, `t0061-run-command.sh#6,18`, `t0302-credential-store.sh#57`, `t0450-txt-doc-vs-help.sh#131,647,797`, `t0610-reftable-basics.sh#61`, `t1091-sparse-checkout-builtin.sh#21,43,48`, `t1092-sparse-checkout-compatibility.sh#55`, `t1300-config.sh#194,197,237,285,494`, `t1403-show-ref.sh#9`, `t1430-bad-ref-name.sh#26`, `t1450-fsck.sh#36`, `t1461-refs-list.sh#415`, `t1503-rev-parse-verify.sh#4`, `t1700-split-index.sh#10-12,14-15` + +**These are flaky, not fixed divergences, and this entry cannot +gate the cell.** Two mechanisms above are deterministic; this +residue is not. Measured three ways: + +- The same cell on two CI runs of near-identical code + (36476334300, then 36485450259) reported 12 and 18 residual + failures with **no overlap**: every id the first run flagged + passed in the second, and vice versa. The two mechanism entries + meanwhile reproduced exactly both times, 110 and 46. +- Locally, six runs of `t0003 t0017 t0040 t1700` under sshfs: + `t0003`'s six hardlink assertions failed in all six runs, while + `t0017#4` failed in one and `t0040#37`, `t1700#10-15` and + `t0003#41`/`#48` in none -- although CI has flagged each of them. +- `prove --jobs 1` is no cleaner than `--jobs 4`, and 14 runs of + `t0017` alone were all clean, so it takes the fuller suite's + concurrent load to show up at all. + +Some of these tests touch no filesystem semantics whatsoever -- +`t0040#37` "OPT_CALLBACK() and OPT_BIT() work" and `t0017#4` +"test-tool env-helper --type=ulong" parse arguments and +environment variables, and `t0450` compares documentation against +`-h` output. What they do share is capturing output into `>out` / +`2>err` inside the trash directory on the mount and then grepping +it, which points at the mount losing or delaying writes under +concurrent load rather than at any semantic divergence. + +So a `script#N` list is the wrong instrument here: each run draws a +different sample, and pinning one run's sample is what made this +cell report `failing-new` twice. Untried candidates, in order: +mounting with `-o attr_timeout=0 -o entry_timeout=0` (and possibly +`-o max_conns=N` or `-o sync_read`) to see whether the residue +disappears, which would make the cell deterministic and properly +xfail-able. See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) diff --git a/evals/known-issues.yaml b/evals/known-issues.yaml index 223b04d..6e29003 100644 --- a/evals/known-issues.yaml +++ b/evals/known-issues.yaml @@ -232,37 +232,57 @@ issues: backends: [sshfs] targets: [git] tests: - - "t0003-attributes.sh#48" + # Observed across CI runs 36476334300 and 36485450259. This list + # DOCUMENTS a sample; it does not gate -- see the notes. + - "t0003-attributes.sh#41,48" + - "t0017-env-helper.sh#4" + - "t0040-parse-options.sh#37" - "t0061-run-command.sh#6,18" - "t0302-credential-store.sh#57" - - "t0450-txt-doc-vs-help.sh#797" - - "t1091-sparse-checkout-builtin.sh#48" + - "t0450-txt-doc-vs-help.sh#131,647,797" + - "t0610-reftable-basics.sh#61" + - "t1091-sparse-checkout-builtin.sh#21,43,48" - "t1092-sparse-checkout-compatibility.sh#55" - - "t1300-config.sh#194,285,494" + - "t1300-config.sh#194,197,237,285,494" - "t1403-show-ref.sh#9" + - "t1430-bad-ref-name.sh#26" - "t1450-fsck.sh#36" + - "t1461-refs-list.sh#415" + - "t1503-rev-parse-verify.sh#4" + - "t1700-split-index.sh#10-12,14-15" tags: [needs-triage] links: ["GOTCHAS.md#sshfs-bineval-under-sshfs"] notes: | - Acknowledged, root cause not run down. Deliberately kept apart from - the two mechanisms above rather than folded into them: neither the - hardlink message nor a socket refusal appears in these scripts' - logs. Candidates worth checking first are sshfs's 1-second mtime - granularity, its lack of `xattr`/`fifo` support, and how it reports - the mode bits `execve()` and `builtin_objectmode` read -- three of - these are named from their own assertions: - `t0003#48` "builtin object mode attributes work (dir and regular - paths)", `t0061#6` "run_command can run a script without a #! line" - and `t0061#18` "run_command is asked to abort gracefully". + **These are flaky, not fixed divergences, and this entry cannot + gate the cell.** Two mechanisms above are deterministic; this + residue is not. Measured three ways: - `t0003-attributes.sh#48` is also in `vfat-git-untriaged`, and this - list's `t0061` and `t1091`/`t1300` assertions sit next to the ones - recorded there, so some of these are likely one non-POSIX cause - shared with vfat rather than anything sshfs invented. + - The same cell on two CI runs of near-identical code + (36476334300, then 36485450259) reported 12 and 18 residual + failures with **no overlap**: every id the first run flagged + passed in the second, and vice versa. The two mechanism entries + meanwhile reproduced exactly both times, 110 and 46. + - Locally, six runs of `t0003 t0017 t0040 t1700` under sshfs: + `t0003`'s six hardlink assertions failed in all six runs, while + `t0017#4` failed in one and `t0040#37`, `t1700#10-15` and + `t0003#41`/`#48` in none -- although CI has flagged each of them. + - `prove --jobs 1` is no cleaner than `--jobs 4`, and 14 runs of + `t0017` alone were all clean, so it takes the fuller suite's + concurrent load to show up at all. - Test ids seeded from run 36476334300 (git v2.55.0): 12 assertions, - of which 11 were not in this file's first draft. That draft was - derived from a local run in a container that runs as **root**, - which both flips permission-dependent assertions and renumbers the - scripts that define tests conditionally -- so its ids named partly - different assertions. CI's verdict is the one to seed from. + Some of these tests touch no filesystem semantics whatsoever -- + `t0040#37` "OPT_CALLBACK() and OPT_BIT() work" and `t0017#4` + "test-tool env-helper --type=ulong" parse arguments and + environment variables, and `t0450` compares documentation against + `-h` output. What they do share is capturing output into `>out` / + `2>err` inside the trash directory on the mount and then grepping + it, which points at the mount losing or delaying writes under + concurrent load rather than at any semantic divergence. + + So a `script#N` list is the wrong instrument here: each run draws a + different sample, and pinning one run's sample is what made this + cell report `failing-new` twice. Untried candidates, in order: + mounting with `-o attr_timeout=0 -o entry_timeout=0` (and possibly + `-o max_conns=N` or `-o sync_read`) to see whether the residue + disappears, which would make the cell deterministic and properly + xfail-able. From f6cc4fb71ad7422633ebfba8ca56dea7d9643fd2 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 29 Sep 2026 12:18:49 +0000 Subject: [PATCH 8/8] dump-failure-logs.sh: an sshfs case, and what cache=no actually buys Two things this branch owed. `dump-failure-logs.sh` had no `sshfs)` arm, so every failing sshfs cell printed `unknown backend: sshfs` and no diagnostics at all -- visible in the git-annex cell's log, which is the one place they would have helped. The new arm prints any surviving `/tmp/eval-under-sshfs-*.scratch/sshd.log` (the backend removes its scratch dir on teardown, so the log is only there when teardown was skipped -- a crash, or --keep), any leftover `mount -t fuse.sshfs`, which is itself the finding when a suite timed out, and fuse-tagged dmesg. And the `cache=no` candidate the last commit left untried is now measured: five runs of `t0*.sh` at `--jobs 4`, counting failures beyond the 64 assertions this subset's two mechanisms own. Baseline drew 4, 8 and 8; `cache=no` drew 1 and 2, at about 7% more wall clock. So sshfs's caching is implicated, and `cache=no` is worth having as a mitigation -- but it does not make the cell green, because known_issues.py sets `failing-new` on the first uncovered failure and has no notion of a flake budget. One flake per run still fails the cell most runs. That turns the remaining question into a decision rather than a patch, and the note in evals/known-issues.yaml now says so: cover the whole cell with `tests: ["*"]` the way `vfat-pjdfstest` does (green, but it masks the 110 and 46 findings along with any future regression), make the cell non-gating, or give the verdict machinery a flake budget -- which belongs in the machinery, not in this backend's PR. Not choosing one here on my own. shellcheck clean over 26 scripts; run-checks.sh passes; GOTCHAS.md regenerated. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- GOTCHAS.md | 22 +++++++++++++++++----- bin/ci/dump-failure-logs.sh | 18 ++++++++++++++++++ evals/known-issues.yaml | 22 +++++++++++++++++----- 3 files changed, 52 insertions(+), 10 deletions(-) diff --git a/GOTCHAS.md b/GOTCHAS.md index cb39c7e..fd8ccaf 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -394,11 +394,23 @@ concurrent load rather than at any semantic divergence. So a `script#N` list is the wrong instrument here: each run draws a different sample, and pinning one run's sample is what made this -cell report `failing-new` twice. Untried candidates, in order: -mounting with `-o attr_timeout=0 -o entry_timeout=0` (and possibly -`-o max_conns=N` or `-o sync_read`) to see whether the residue -disappears, which would make the cell deterministic and properly -xfail-able. +cell report `failing-new` twice. + +Mounting with `-o cache=no` (`--no-cache`) was measured, not +guessed: five runs of `t0*.sh` at `--jobs 4`, counting failures +beyond the 64 this subset's two mechanisms own. Baseline drew 4, 8 +and 8; `cache=no` drew 1 and 2, for about 7% more wall clock (107s +-> 114s). So the cache is implicated and `cache=no` is worth having +as a mitigation -- but it does not fix the gate, because +`known_issues.py` sets `failing-new` on the *first* uncovered +failure and has no tolerance for a flake budget. A rate of one per +run still fails the cell most runs. + +Which leaves a decision rather than a patch: cover the whole cell +with `tests: ["*"]` as `vfat-pjdfstest` does (green, but it masks +the 110 and 46 findings and any future regression), make the cell +non-gating, or give the verdict machinery a flake budget. That last +one belongs in the machinery, not in this backend's PR. See: [sshfs (`bin/eval-under-sshfs`)](#sshfs-bineval-under-sshfs) diff --git a/bin/ci/dump-failure-logs.sh b/bin/ci/dump-failure-logs.sh index 23b1215..bf69778 100755 --- a/bin/ci/dump-failure-logs.sh +++ b/bin/ci/dump-failure-logs.sh @@ -50,6 +50,24 @@ case "$BACKEND" in | grep -iE "loop|nfs|${VERSION:-nomatch}" \ | tail -30 || true ;; + sshfs) + # bin/eval-under-sshfs removes its scratch dir on teardown, so + # sshd.log is only still here when teardown was skipped (a crash, + # or --keep). Print it when it is: a mount that died mid-suite + # says so there and nowhere else. + for log in /tmp/eval-under-sshfs-*.scratch/sshd.log; do + [ -f "$log" ] || continue + echo "=== $log (last 50) ===" + sudo tail -50 "$log" || true + done + # A leaked mount means teardown did not finish, which is itself + # the finding when the suite timed out. + mounts="$(mount -t fuse.sshfs 2>/dev/null)" + echo "=== fuse.sshfs mounts still present ===" + echo "${mounts:-(none)}" + echo "=== dmesg (fuse-tagged, last 30) ===" + sudo dmesg 2>/dev/null | grep -iE 'fuse|sshfs' | tail -30 || true + ;; *) echo "unknown backend: $BACKEND" >&2 ;; diff --git a/evals/known-issues.yaml b/evals/known-issues.yaml index 6e29003..af41cbb 100644 --- a/evals/known-issues.yaml +++ b/evals/known-issues.yaml @@ -281,8 +281,20 @@ issues: So a `script#N` list is the wrong instrument here: each run draws a different sample, and pinning one run's sample is what made this - cell report `failing-new` twice. Untried candidates, in order: - mounting with `-o attr_timeout=0 -o entry_timeout=0` (and possibly - `-o max_conns=N` or `-o sync_read`) to see whether the residue - disappears, which would make the cell deterministic and properly - xfail-able. + cell report `failing-new` twice. + + Mounting with `-o cache=no` (`--no-cache`) was measured, not + guessed: five runs of `t0*.sh` at `--jobs 4`, counting failures + beyond the 64 this subset's two mechanisms own. Baseline drew 4, 8 + and 8; `cache=no` drew 1 and 2, for about 7% more wall clock (107s + -> 114s). So the cache is implicated and `cache=no` is worth having + as a mitigation -- but it does not fix the gate, because + `known_issues.py` sets `failing-new` on the *first* uncovered + failure and has no tolerance for a flake budget. A rate of one per + run still fails the cell most runs. + + Which leaves a decision rather than a patch: cover the whole cell + with `tests: ["*"]` as `vfat-pjdfstest` does (green, but it masks + the 110 and 46 findings and any future regression), make the cell + non-gating, or give the verdict machinery a flake budget. That last + one belongs in the machinery, not in this backend's PR.