DO NOT MERGE — CI diagnostic harness for the intermittent NFS LinkAnnexFailed - #293
yarikoptic-gitmate wants to merge 3 commits into
Conversation
The nightly `test-annex (nfs-home)` job fails in a different test each
night, always at `git annex add` of an unlocked file ("<file> failed to
link to annex") or at `git annex unlock` ("unlock failed"). Both are
linkAnnex returning LinkAnnexFailed, which is reachable by two paths
that print nothing at all:
* linkOrCopy returned Nothing - a swallowed IOException, or cp /
preserveGitMode returning False;
* the source file's inode cache (inode, size, high resolution mtime)
differed before and after the copy, so checksrcunchanged deleted
the destination.
The transcripts in the CI logs cannot tell those apart, so add a warning
to each, including the before/after inode caches. Diagnostic only, on
paths that already fail; drop it once the cause is known.
Not compile-tested locally (no GHC available where this was written);
a PR touching patches/*.patch triggers the builds, which will verify it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
This branch exists only to exercise the LinkAnnexFailed diagnostic patch against NFS; it is not meant to be merged. Point its Ubuntu workflow at that single question: * test-annex runs only the nfs-home flavor, three times per invocation (one run reproduces the failure in roughly two nights out of three). * test-annex-more and test-datalad are disabled; they say nothing about this failure and cost ~40 minutes of runners per invocation. * The failure e-mails are disabled: failures here are the expected outcome, and an "Ubuntu build failed" mail per night would read as if master were broken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Answered on the first run: it is
|
| inode | size | mtime before | mtime after | delta | |
|---|---|---|---|---|---|
import1/f |
8952864 (same) | 16 (same) | 13:38:49.078602968 | 13:38:49.090671216 | +12.07 ms |
| annex object | 9455246 (same) | 20 (same) | 13:45:52.023778495 | 13:46:46.769318488 | +54.75 s |
Same inode, same size, only the mtime moves — and in the second case the source is a frozen annex object that nothing in the test writes to. compareStrong requires exact equality of the high-resolution mtime (Utility/InodeCache.hs), so linkAnnex concludes the source changed mid-copy, deletes the destination and returns LinkAnnexFailed. That surfaces as failed to link to annex on the To side and unlock failed on the From side, i.e. the two messages the nightlies have been alternating between.
Two flavours of the same root cause, both NFS attribute-cache artefacts rather than real modifications:
- the 12 ms one is a sub-second mtime that differs between two stats within the same second — the client's locally-set value being replaced by the server's;
- the 54.7 s one is a plainly stale cached attribute (
acregmaxis 60 s by default) that only gets revalidated during the copy.
Note also the knock-on damage: when the deleted destination is the annex object, everything downstream in that repo fails. Job rep 2 goes on to produce a run of content not available to send, openBinaryFile: does not exist and cp: cannot stat .git/annex/objects/ZP/kg/SHA256E-s2097152--… in the testremote type directory group — a missing object, not an independent bug.
Next steps
- Worth confirming by mounting the export with
-o noac(oractimeo=0) in this branch's workflow: if the failures disappear, attribute caching is proven and NFS users have a workaround. - Worth reporting upstream: the strict
compareStrongmtime equality inlinkAnnex'schecksrcunchangedis not a safe assumption on NFS, where two stats of an unmodified file can disagree in the sub-second field, or return a value up toacregmaxstale.
The harness has answered what it was built for; it can keep running for more samples, but the patch should not be merged.
Generated by Claude Code
The first harness run showed both failures are linkAnnex's checksrcunchanged, with inode and size unchanged and only the mtime moving (+12 ms in one case, +54.7 s in the other) — consistent with NFS client attribute caching rather than a real modification. Test that directly: run the nfs-home flavor three ways, three times each, in one invocation so the arms see the same runner conditions. * default — what the nightlies mount, the control; * actimeo0 — `-o actimeo=0`, attribute caching off and nothing else, the clean probe of the hypothesis; * noac — `-o noac`, that plus synchronous writes, the workaround usually recommended to NFS users. The effective mount options are printed per job so the log records what was actually negotiated. The suite timeout goes 3600 -> 7200s because noac writes synchronously and is much slower; a job dying at the timeout would make the arm unreadable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Attribute caching confirmed:
|
| arm | mount | result | wall clock |
|---|---|---|---|
default |
as the nightlies mount | 3/3 failed | 8m17s, 8m54s, 10m27s |
actimeo0 |
-o actimeo=0 |
3/3 passed | 32m01s, 32m24s, 50m47s |
noac |
-o noac |
3/3 passed | 29m17s, 34m30s, 34m54s |
The mount options each job recorded, confirming the arms really differed:
default: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,…
actimeo0: rw,relatime,vers=4.2,…,namlen=255,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,…
noac: rw,relatime,sync,vers=4.2,…,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,noac,…
actimeo=0 is sufficient. It disables attribute caching and changes nothing else; noac adds synchronous writes (note the extra sync in its options) and does no better. So the failure is the client's attribute cache, not the write path — which is what the mtime evidence implied and now has a controlled result behind it.
Cost of the workaround
Roughly 4×: a passing default job in the previous run (rep 3, same 26 groups) took 8m20s, against 32m01s for actimeo0 here. The control arm's 8–10 min above is not a fair comparison on its own, since a group that fails ends early.
New: a second check site with the same root cause
default,2 produced the familiar diagnostics — note both deltas are again mtime-only, inode and size unchanged:
add ../dir2/foo
DIAGNOSTIC: linkAnnex: inode cache of ../dir2/foo changed while copying it; before: 9211545 20 1790178434 30902619; after: 9211545 20 1790178434 309979362
../dir2/foo failed to link to annex → +0.279 s, same second
unlock foo
DIAGNOSTIC: linkAnnex: inode cache of .git/annex/objects/Kj/0x/SHA256E-s20--e394…/SHA256E-s20--e394… changed while copying it; before: 9464846 20 1790178735 562779131; after: 9464846 20 1790178776 778504714
git-annex: unlock failed → +41.216 s (stale cached attribute)
But default,1 and default,3 failed with no DIAGNOSTIC line at all. Their logs (complete, 1764 and 1810 lines) contain no failed to link to annex and no unlock failed; they died in testremote type directory with:
storeKey: FAIL
Exception: content changed while it was being sent
present True: FAIL
That message is Annex/Content.hs:730, the sameInodeCache check in prepSendAnnex — a different call site, same assumption: an object's inode cache must compare equal before and after the transfer. compareInodeCaches only relaxes to compareWeak when inodesChanged (the sentinal file) says inodes moved, and here they never do, so it stays exact and the unmodified object looks modified.
So the strict high-resolution mtime comparison bites in at least two places, and a fix confined to linkAnnex would leave transfers still failing on NFS.
Where this leaves it
- Root cause: established. An NFS client can return two different high-resolution mtimes for a file nothing wrote to — sub-second within the same second (12 ms, 279 ms seen) or plainly stale (41 s, 55 s seen, against a 60 s default
acregmax). git-annex'scompareStrongtreats that as "the file changed". - Workaround for NFS users:
actimeo=0, at roughly 4× the wall clock.noacis not needed. - Upstream report: worth filing with these numbers — the comparison is exact-equality on a timestamp the filesystem does not promise to be stable, at both
linkAnnex'schecksrcunchangedandprepSendAnnex'ssameInodeCache. - The harness has now answered both questions it was built for. The patch still must not be merged; the branch can go once a report exists.
Generated by Claude Code
|
Closing: the harness answered both questionsThis branch was never meant to merge — it existed to make an intermittent NFS failure say why it was failing. It did, twice over, so it is being closed rather than left open to rot. Summary and pointers, for whoever finds this from a future NFS failure. What it found
Two call sites, not one — a fix confined to the first would leave transfers broken:
Controlled run 35882579073, three arms × three reps on one build: Not a regression in the failing code: nothing has touched Where it lives now
What was in here, for reference
Unrelated, but recorded because it showed up here: the Windows The branch can be deleted. Generated by Claude Code |
Chasing con/git-annex#293 (git-annex intermittently failing on NFS with "failed to link to annex" / "unlock failed") needed three things this framework could not express. All three are small: * `eval-under nfs --mount-opts OPTS` / `--export-opts OPTS`. NFS_OPTS was hardcoded to rw,async|rw,sync, so the client-side knobs that matter for this class of bug -- actimeo=0, noac, lookupcache=none, nocto, vers= -- were unreachable. Mount and export options stay in separate variables because an option valid in one is rejected by the other. * target `mtime-stability`: write, stat, copy the way git-annex copies, stat again, compare (inode, size, high-res mtime) exactly, report a rate. No git-annex involved, so a red cell says "the filesystem", not "the application". This is the property git-annex assumes in Annex/Content.hs:linkAnnex and prepSendAnnex, and that NFS attribute caching breaks: in con/git-annex CI the mtime of an unmodified file moved by 12 ms, 279 ms and 41 s across a copy, and mounting with actimeo=0 made 3/3 failing runs pass. * target `git-annex-linkannex`: the same question at the git-annex level -- loop `git annex unlock` (linkFromAnnex') and unlocked `git annex add` (linkToAnnex), report a failure rate in minutes rather than a pass/fail of the ~20 minute suite. Both targets run on every backend, so BeeGFS and vfat get answered for free; that grows the README grid from 20 cells to 30. Tested: shellcheck clean; both targets run end-to-end on ext4 (0%, the negative control); option strings verified for each flag combination. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
This PR is a diagnostic harness and must not be merged. It exists so that the nightly
test-annex (nfs-home)failure can be run repeatedly with instrumentation, against a build of currentupstream/master, without putting a diagnostic patch into production builds and releases.The failure being chased
test-annex (nfs-home, ubuntu-24.04)fails on roughly two nightlies out of three, in a different test each time, and it is the only failing job in those runs:conflictor failed to link to annex(atgit annex add)git-annex: unlock failed(atgit annex unlock)Both messages are the same underlying failure.
Annex/Ingest.hs:213(failed to link to annex) andCommand/Unlock.hs:61(unlock failed) are two callers oflinkAnnex, andlinkAnnexreachesLinkAnnexFailedthrough exactly two exits, neither of which reports anything:Annex/Content.hs:633—linkOrCopyreturnedNothing: anIOExceptionswallowed bycatchDefaultIO Nothing(Annex/Content/LowLevel.hs:50), orcp/preserveGitModereturningFalseviacatchBoolIO(:78).Annex/Content.hs:646—checksrcunchanged: the source file'sInodeCache(inode, size, high-resolution mtime, compared withcompareStrong, i.e. exact equality) differed before and after the copy, so the destination was deleted.The second is the interesting one: on the
unlockpath the source is a frozen annex object that nothing is modifying, so it would mean twostats of an unchanged file disagreeing — plausible with NFS attribute caching, not with a local filesystem. But the CI transcripts cannot tell the two exits apart, because both are silent.What is in this PR
patches/20260922-f68b252dbd-diag-linkannex-failure.patch— adds a warning to each of those exits, including the before/after inode caches:DIAGNOSTIC: linkAnnex: inode cache of <src> changed while copying it; before: <ino size secs nsecs>; after: …— says which field moved;DIAGNOSTIC: linkOrCopy threw: <exception>,DIAGNOSTIC: checkedCopyFile: copyFileExternal failed/preserveGitMode failed/exception: ….Diagnostic only, on paths that already fail. All callers of the patched functions guard on the object being present, so these should not appear in normal operation. DEP-3 + SPDX header; verified to
git applycleanly againstf68b252dbdand not to reverse-apply.Branch-local tuning of
.github/workflows/build-ubuntu.yaml, so repeated runs cost little and stay quiet:test-annexruns only thenfs-homeflavor, three times per invocation;test-annex-moreandtest-dataladare disabled;These workflow changes are for this branch only, which is one more reason the PR should not be merged.
How it is run
Opening the PR builds the patch on Ubuntu, macOS, macOS ARM64 and Windows (they all trigger on
patches/*.patch), which also serves as the compile check — the patch was written without a Haskell toolchain available. After that, only the Ubuntu workflow needs repeating: a nightlyworkflow_dispatchofbuild-ubuntu.yamlwithref=claude/elegant-meitner-y7sdk7builds currentupstream/masterwith the patch applied (BUILD_COMMITdefaults toorigin/upstream/master;patches/comes from the dispatched ref) and runs threenfs-homejobs.Findings get reported back on this PR. Once a failing run names the cause, the patch has done its job and the branch can be deleted — nothing here is meant to land on
master.🤖 Generated with Claude Code
https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Generated by Claude Code