Skip to content

ENH: NFS mount/export option pass-through + two stat-stability targets - #11

Open
yarikoptic-gitmate wants to merge 6 commits into
masterfrom
enh-mtime-stability-targets
Open

yarikoptic-gitmate wants to merge 6 commits into
masterfrom
enh-mtime-stability-targets

Conversation

@yarikoptic-gitmate

@yarikoptic-gitmate yarikoptic-gitmate commented Sep 24, 2026 •

Copy link
Copy Markdown

Proposal, from chasing con/git-annex#293: git-annex fails intermittently on NFS with failed to link to annex / unlock failed, because it records (inode, size, high-resolution mtime) for a file, copies it, stats it again and compares the two exactly — and an NFS client can report two different mtimes for a file nothing wrote to. In con/git-annex CI the mtime of an unmodified file moved by 12 ms, 279 ms and 41 s across a copy, and mounting with actimeo=0 turned 3/3 failing runs into 3/3 passing ones.

Three things that investigation needed and this framework could not express. All small, and useful beyond that one bug.

1. --mount-opts / --export-opts on the NFS backend

NFS_OPTS was hardcoded to rw,async / rw,sync, so the client-side knobs that decide this class of question — actimeo=0, noac, lookupcache=none, nocto, vers=, rsize=/wsize= — were unreachable without editing the script. Now:

sudo bin/eval-under nfs --set-home --mount-opts actimeo=0 -- \
  bash -c 'cd "$HOME" && git annex test'

Mount and export options stay in separate variables (EVAL_UNDER_NFS_MOUNT_OPTS / EVAL_UNDER_NFS_EXPORT_OPTS), since an option valid in one namespace is rejected by the other — same reasoning as the existing sync/no_root_squash split. This is the highest-value part of the PR: it makes "the filesystem lied" vs "the code is wrong" a one-flag experiment.

2. New target: mtime-stability

Write a file, stat it, copy it the way git-annex copies (cp --reflink=auto -a --no-preserve=xattr), stat it again, compare exactly, report a rate. No git-annex involved, so a red cell says the filesystem, not the application. -j for parallel load (the race is a timing one), --delay to age the attribute cache past acregmin/acregmax, --no-copy as a control separating "the copy revalidates" from "the cache expired on its own". Output is in git-annex's own inode size secs nsecs format so it lines up with what git-annex prints.

This is the property Annex/Content.hs:linkAnnex and prepSendAnnex assume, so the cell answers the question for BeeGFS and vfat too — relevant to con/git-annex#288.

3. New target: git-annex-linkannex

The same question at the git-annex level: loop git annex unlock (linkFromAnnex') and unlocked git annex add (linkToAnnex), report a failure rate in minutes instead of a pass/fail of the ~20-minute suite. Useful for verifying a fix, and for watching the rate move as mount options change.

Both targets emit TAP and are scored per row by bin/ci/collect-results.py → known_issues.py, like stress-ng: mtime-stability reports one point per stat field (inode-stable, size-stable, mtime-stable, with the rate and worst mtime delta in the description), git-annex-linkannex one per direction (unlock, add-unlocked). Per-field/per-mode ids rather than per-round ones, so a known issue can name "this filesystem moves the mtime" instead of "round 417".

What CI found

All 30 cells green on run 36790703938 (f358985), in 18 minutes. The git-annex-linkannex measurements, since they are the point of the PR:

cell rate wall
NFS (localhost) 0/200 unlock, 0/200 add-unlocked 105 s
BeeGFS 7.4.6 0/200, 0/200 196 s
BeeGFS 8.1.0 0/200, 0/200 220 s (83 s + 64 s in the loop)
Loop ext4 0/200, 0/200 39 s
Loop vfat 0/200, 0/200 28 s

mtime-stability passes on all five backends: no filesystem here moves an unmodified file's (inode, size, mtime) across a copy.

The NFS cells being clean is itself the finding. CI installs the current daily build (10.20260901+git74-g3cdda5ef2a) — the same git-annex whose nfs-home cell fails in con/git-annex — so this is the exact failing operations on the exact failing version, with zero failures. The variable is the kernel: this workflow pins ubuntu-22.04 (kernel 6.8) because BeeGFS's DKMS modules cannot build against 6.17, while con/git-annex's test-annex moved to ubuntu-24.04, whose image now ships 6.17. See con/git-annex#294, where 22.04 passed 4/4 and 24.04 failed 4/4 on the same commit.

So this framework cannot currently reproduce that bug, by construction. A 24.04 row would have to skip the BeeGFS backends; I have not added one, as that is a matrix-shape decision rather than part of this PR.

Three things CI taught us about the harness along the way, each fixed here:

  • annex.diskreserve vs small loop images. On a fresh ext4 image with 83 MB free, git annex unlock of a 14-byte file refuses with "not enough free space, need 13.73 MB more". Every round failed before linkAnnex was reached. The loop now sets annex.diskreserve=0 in its probe repos — it measures the inode-cache comparison, not git-annex's disk-space policy. This is the same root cause as the existing loop-annex-diskreserve known issue, which covers the git-annex target; that entry is untouched and still applies there.
  • A full disk looked exactly like the bug. ENOSPC makes git-annex print the same failed to link to annex / unlock failed lines as an inode-cache mismatch, and the loop was counting them — reporting 800/800 "failures" on ext4, the negative control. It now recognises those messages, refuses to report a rate, and exits 4; the target turns that into an incomplete cell rather than a verdict the run cannot support.
  • Sizing. At 200×4 rounds, BeeGFS spent 967 s on one of two modes against a 1200 s budget, so both BeeGFS cells timed out with only mode 1 reported. Rounds per worker are now 50 (200 per mode). The trade-off is detection power — 200 rounds per mode will not reliably show a rate below ~1% — so bin/ci/linkannex-loop.sh -n is the knob for hunting something rare, not this default.

Notes for review

  • Both targets run on every backend, so the README grid goes 20 → 30 cells (gen-readme-matrix.sh re-run). If that is more than you want, dropping the git-annex-linkannex entry from evals/matrix.yaml is a one-line change — the script stays usable by hand.
  • matrix-json.sh now carries needs-git-annex per cell and the workflow gates the daily-build fetch on it, instead of on matrix.target == 'git-annex'. Without that, a second git-annex-using target runs with no git-annex installed — which is exactly what happened. target_needs_git_annex() in matrix.sh had been dead code; it now has a caller, and tests/test_matrix_json.py holds both halves of the contract.
  • mtime-stability needs nothing installed; git-annex-linkannex reuses install_git_annex.
  • No known-issues.yaml entries added: every cell passes, so there is nothing to record.
  • Observed but not addressed here: at the old round count the BeeGFS cluster degraded under the load (Receive failed from node_meta_1, then metadata and storage nodes to offline, with teardown wedged in FhgfsOps_flush). That also inflated the 967 s above — at 200 rounds per mode the same cell does both modes in 147 s, far better than linear scaling would predict, and the cluster stays healthy. Worth knowing if any future target leans harder on BeeGFS metadata, but it needs no change here.
  • Not included, but proposed: a --tmpdir-local flag so TMPDIR can stay off the mount. con/git-annex's nfs-home flavor puts only HOME on NFS, and that difference is part of why the first local reproduction attempt diverged from CI. It belongs in all three backends rather than just this one, so it seemed better as its own PR.

Testing

  • bin/ci/run-checks.sh fully green: shellcheck, pyflakes, bats (30 tests, two new for the NFS flags), known-issues, 33 unit tests (five new for the matrix contract).
  • Both targets run end-to-end into a cell dir and collect with # complete: yes.
  • The ENOSPC abort path exercised on a deliberately pre-filled image: real ENOSPC, exit 4, rate withheld, cell incomplete rather than failing.
  • mtime-stability's failure path exercised with a stubbed cp that bumps the source mtime, which reproduces the NFS signature exactly (inode and size stable, mtime-stable not ok).
  • The cleanup trap verified as an unprivileged user against real annex objects (dr-xr-xr-x): plain rm -rf leaves them behind and floods the log, chmod -R u+w first leaves nothing and says nothing. On the NFS cell that cut the uploaded log bundle from 78,288 to 1,573 bytes.
  • vfat could not be exercised locally (no mkfs.vfat, and this container's apt is blocked), so its results come from CI only.

🤖 Generated with Claude Code

https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf


Generated by Claude Code

Yaroslav Halchenko and others added 2 commits September 30, 2026 19:16
Chasing con/git-annex#293 (git-annex intermittently failing on NFS with
"failed to link to annex" / "unlock failed") needed three things this
framework could not express. All three are small:

* `eval-under nfs --mount-opts OPTS` / `--export-opts OPTS`. NFS_OPTS
  was hardcoded to rw,async|rw,sync, so the client-side knobs that
  matter for this class of bug -- actimeo=0, noac, lookupcache=none,
  nocto, vers= -- were unreachable. Mount and export options stay in
  separate variables because an option valid in one is rejected by the
  other.

* target `mtime-stability`: write, stat, copy the way git-annex copies,
  stat again, compare (inode, size, high-res mtime) exactly, report a
  rate. No git-annex involved, so a red cell says "the filesystem",
  not "the application". This is the property git-annex assumes in
  Annex/Content.hs:linkAnnex and prepSendAnnex, and that NFS attribute
  caching breaks: in con/git-annex CI the mtime of an unmodified file
  moved by 12 ms, 279 ms and 41 s across a copy, and mounting with
  actimeo=0 made 3/3 failing runs pass.

* target `git-annex-linkannex`: the same question at the git-annex
  level -- loop `git annex unlock` (linkFromAnnex') and unlocked
  `git annex add` (linkToAnnex), report a failure rate in minutes
  rather than a pass/fail of the ~20 minute suite.

Both targets run on every backend, so BeeGFS and vfat get answered for
free; that grows the README grid from 20 cells to 30.

Tested: shellcheck clean; both targets run end-to-end on ext4 (0%, the
negative control); option strings verified for each flag combination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
The rebase onto master brought in per-test cell judging (#14): each
target's output goes through bin/ci/collect-results.py into
results.tsv, and known_issues.py judges that instead of the suite's
exit code. collect-results.py dispatches on the target name and raises
`no results adapter for target ...` otherwise, so as first written
these two targets would have produced incomplete cells.

Both now print TAP and are scored like stress-ng:

* mtime-stability: one point per stat field -- inode-stable,
  size-stable, mtime-stable -- with the rate and the worst mtime delta
  in the description, and the first few mismatches as TAP comments.
  Per-field ids rather than per-round ones because a known issue has to
  be able to name "this filesystem moves the mtime", not "round 417".
* git-annex-linkannex: one point per direction, `unlock` and
  `add-unlocked`. linkannex-loop.sh gained --report FILE so the wrapper
  reads a count instead of parsing prose, and a loop that dies before
  reporting becomes a failing point rather than a silent pass.
* collect_tap_target() serves both, taking the row id from the
  description's first word, as collect_stress_ng does.

No known-issues entries are added: whether NFS actually fails
mtime-stable on this CI's ubuntu-22.04 runners (kernel 6.8) is an open
question -- con/git-annex only sees the git-annex failures on 24.04
(kernel 6.17), see con/git-annex#294 -- so let the first run answer it
rather than pre-declaring a verdict.

Tested: bin/ci/run-checks.sh fully green (shellcheck 29 scripts,
pyflakes, bats 30 tests, known-issues, 28 unit tests). Both targets run
end-to-end into a cell dir and collect cleanly on ext4; the TAP failure
path was exercised with a stubbed cp that bumps the source mtime, which
yields exactly the NFS signature (inode and size stable, mtime-stable
not ok). Two bats tests cover the new flags.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
@yarikoptic-gitmate
yarikoptic-gitmate force-pushed the enh-mtime-stability-targets branch from 7a139fc to b8aa45a Compare September 30, 2026 19:21

Copy link
Copy Markdown
Author

Rebased onto current master (60 commits, up to e8b1291 / #16). One textual conflict, in the matrix file — git followed its move to evals/matrix.yaml, and the collision was only that master reworded the pjdfstest comment where my two entries were inserted above it; master's wording kept, my entries kept. Everything else auto-merged, and the README grid regenerated to the same 30 cells.

The rebase also surfaced something a textual merge cannot: #14 changed how cells are judged, from the suite's exit code to per-test rows via bin/ci/collect-results.py → known_issues.py. That collector dispatches on target name and raises no results adapter for target … for anything it does not know, so as originally written both new targets would have produced incomplete cells. Fixed in the second commit — both now print TAP and are scored like stress-ng:

target TAP points (row ids)
mtime-stability inode-stable, size-stable, mtime-stable — rate and worst mtime delta in the description, first few mismatches as TAP comments
git-annex-linkannex unlock, add-unlocked — one per direction

Per-field and per-mode ids rather than per-round ones, so a known issue can name "this filesystem moves the mtime" instead of "round 417". linkannex-loop.sh gained --report FILE so the wrapper reads a count instead of parsing prose, and a loop that dies before reporting becomes a failing point rather than a silent pass. collect_tap_target() serves both and takes the row id from the description's first word, as collect_stress_ng does.

No known-issues.yaml entries added on purpose: whether NFS fails mtime-stable on this CI's ubuntu-22.04 runners is genuinely open. The failure that started all of this only reproduces on kernel 6.17 (ubuntu-24.04); on 6.8 — what ubuntu-22.04 runs today, and what this workflow uses — the same git-annex build passes, 4/4 vs 4/4 in a one-variable run (con/git-annex#294). So this PR's own first run is the interesting data point, and pre-declaring a verdict would only hide it. Worth considering a 24.04 row for the NFS backend separately — on 22.04 alone this grid cannot see that class of bug at all.

Verification, all local:

  • bin/ci/run-checks.sh fully green — shellcheck (29 scripts), pyflakes, bats (30 tests, two new ones covering --mount-opts/--export-opts), known-issues valid, 28 unit tests.
  • Both targets run end-to-end into a cell dir and collect cleanly on ext4: inode-stable/size-stable/mtime-stable and unlock/add-unlocked, # complete: yes.
  • The TAP failure path was exercised with a stubbed cp that bumps the source mtime — it yields exactly the NFS signature, not ok mtime-stable with inode and size stable.

Generated by Claude Code

The five git-annex-linkannex cells of run 36765084880 all died the same
way, on every backend including ext4:

    I: running: timeout 1200 .../bin/ci/target-git-annex-linkannex.sh
    git: 'annex' is not a git command. See 'git --help'.

git-annex was never installed. evals/matrix.yaml declares
`needs-git-annex: true` for the target and install-target.sh routes it to
the git-annex arm, but the workflow step that actually fetches the daily
build was gated on `matrix.target == 'git-annex'` -- the target's *name*,
not the flag -- so it never ran. ext4 failing was the tell: it is the
negative control, and a filesystem-independent failure is the harness.

matrix-json.sh now carries `needs-git-annex` per cell, so the workflow
gates on the data ("each entry carries everything the job body needs",
per its own header) and target_needs_git_annex() in matrix.sh gets its
first caller -- it was dead code, which is why the duplication went
unnoticed. Adding a git-annex-using target stays a data edit.

tests/test_matrix_json.py holds both halves of that contract: every
cell's flag matches the data file, the flag is a JSON boolean (the string
"0" is truthy in a GitHub expression), and no step is gated on a target
name. Checked against master's workflow, where the last of those fails.

Also drops the stale target list from install-target.sh's missing-arg
message, which had gone stale exactly this way; target_known already
prints the live list from evals/matrix.yaml.

Tested: the failure reproduced locally by running the target with a PATH
holding everything but git-annex -- same `git: 'annex' is not a git
command`, same "no TAP plan from target-git-annex-linkannex.sh (suite
died?)" from the collector -- and the same target emitting `1..2` with
two passing points once git-annex is present. Old vs new gate over the
rendered matrix: 5 cells vs 10, the extra 5 being exactly the cells that
failed. bin/ci/run-checks.sh fully green (shellcheck, pyflakes, bats 30,
known-issues, 33 unit tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Run 36774070493 got git-annex installed (the previous commit's fix) and
promptly reported the opposite nonsense: 800/800 rounds "failed" in both
modes, 100.00%, on ext4 -- the negative control, where the true rate is
0. The log's reason:

    not enough free space, need 21.73 MB more (use --force to override
      this check or adjust annex.diskreserve)
    f106 failed to link to annex

Two separate defects, and the first is the dangerous one.

A full filesystem makes git-annex print exactly the same "failed to link
to annex" / "unlock failed" lines as the inode-cache mismatch this loop
exists to measure, and the loop tallied them. So a cell that ran out of
room reported a 100% linkAnnex failure rate -- indistinguishable, in the
results, from the NFS bug this target was written to quantify. Left in,
it would have made the first red NFS cell unreadable. The loop now
recognises the space messages, aborts instead of counting, and exits 4
(documented alongside the other statuses); the target turns that into an
incomplete cell rather than a filesystem verdict the run cannot support.

Second, the reserve. Measured on a fresh ext4 image with 83MB free,
`git annex unlock` of a 14-byte file refuses with "need 13.73 MB more" --
it wants ~97MB free to rewrite 14 bytes, which a 100MB image cannot give,
so every round failed from round 0 without linkAnnex being reached at
all. The loop sets annex.diskreserve=0 in each probe repo: it is
measuring the inode-cache comparison, not git-annex's disk-space policy,
and real ENOSPC is still caught by the check above.

loop-size-mb 100 -> 256 as well. With the reserve gone the cell passes at
100MB, but it consumed ~52MB of the ~83MB usable, and vfat keeps a
worktree copy of every file because it has no symlinks, so it needs more
than ext4 did. The backing image is sparse, so the headroom is free.

Tested: the ext4 cell reproduced the 800/800 exactly at the old settings,
and now reports `ok 1 - unlock 0/800` / `ok 2 - add-unlocked 0/800`,
finishing with 153.8M of 256M free (was 30.9M of 100M). The new abort
path was exercised on a deliberately pre-filled 60MB image: genuine
ENOSPC, exit 4, the rate withheld, no report file, and through the target
a cell that is "incomplete", not "failing". vfat could not be measured
here -- no mkfs.vfat and apt is blocked -- so CI is its first run.
bin/ci/run-checks.sh fully green (shellcheck, pyflakes, bats 30,
known-issues, 33 unit tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
…only

The NFS cell of run 36774070493 passed (0/800 both modes) but printed
dozens of these first:

    rm: cannot remove '.../annex/objects/x1/V2/SHA256E-s14--.../SHA256E-s14--...':
      Permission denied

git-annex sets each object's directory to dr-xr-xr-x, and nothing can
unlink through a directory it cannot write. Running as root hides this,
which is why it never showed locally; under the root-squashed export this
target uses (needs-root: false) it does not, so the EXIT trap's `rm -rf`
failed on every object. Two costs: a screenful of noise ahead of the TAP
output, in a target whose only job is to be read, and the probe repos
left on the mount, where the next mode's 800 rounds still have to fit --
which is also why the ext4 headroom measured in the previous commit was
pessimistic. chmod -R u+w first, as git-annex's own test suite does.

Tested: reproduced as an unprivileged user on a repo with real annex
objects (dr-xr-xr-x, owner nobody) -- plain `rm -rf` gives exactly the
message above and leaves the object behind, chmod-then-rm leaves nothing
and says nothing. The ext4 cell still reports `ok 1 - unlock 0/800` /
`ok 2 - add-unlocked 0/800`, now with zero "cannot remove" lines and no
leftover probe dirs. bin/ci/run-checks.sh fully green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Both BeeGFS cells of run 36776837242 came back incomplete (exit 124).
Not a filesystem verdict and not the cleanup: BeeGFS 8.1.0 finished mode
`unlock` cleanly at 0/800 (0.00%) -- and took 967s of the target's 1200s
budget to do it, so mode `add-unlocked` was killed four minutes in. NFS,
ext4 and vfat each finish a mode in 100-200s; BeeGFS is an order of
magnitude slower, exactly as evals/matrix.yaml already says of it.

Rounds per worker 200 -> 50, so 200 rounds per mode instead of 800. That
brings BeeGFS to roughly 240s per mode, fits the existing timeout with
room to spare, and restores what this target is for: a rate in minutes
rather than a pass/fail after twenty. Raising the timeout instead would
have cost up to ~50min per BeeGFS cell and risked the 60-minute job cap
-- teardown alone hung 15 minutes in that run -- turning a timeout into a
bare GitHub cancellation with no logs.

The cost is detection power: 200 rounds per mode will not reliably show a
rate below ~1%. Both the target's header and this message say so, and
point at bin/ci/linkannex-loop.sh -n for hunting something rare by hand,
because that is the knob to reach for rather than this default.

Separately, and not this PR's to fix: the BeeGFS cluster degraded during
that run -- "Receive failed from node_meta_1" from t=927s, then metadata
and storage nodes to `offline`, and teardown wedged in FhgfsOps_flush.
A metadata-heavy workload appears to be enough to upset it.

loop-size-mb stays 256: its comment now says the 52MB measurement was at
800 rounds per mode, and that the sparse image leaves room to raise the
rounds by hand without editing the matrix too.

Tested: the ext4 cell runs in 38s wall (was ~4min), reporting
`ok 1 - unlock 0/200` / `ok 2 - add-unlocked 0/200`, 19s and 13s per
mode, and collects with `# complete: yes`. bin/ci/run-checks.sh fully
green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants