Skip to content

DO NOT MERGE — CI diagnostic harness for the intermittent NFS LinkAnnexFailed - #293

Closed
yarikoptic-gitmate wants to merge 3 commits into
masterfrom
claude/elegant-meitner-y7sdk7
Closed

yarikoptic-gitmate wants to merge 3 commits into
masterfrom
claude/elegant-meitner-y7sdk7

Conversation

@yarikoptic-gitmate

Copy link
Copy Markdown
Collaborator

This PR is a diagnostic harness and must not be merged. It exists so that the nightly test-annex (nfs-home) failure can be run repeatedly with instrumentation, against a build of current upstream/master, without putting a diagnostic patch into production builds and releases.

The failure being chased

test-annex (nfs-home, ubuntu-24.04) fails on roughly two nightlies out of three, in a different test each time, and it is the only failing job in those runs:

run date test(s) message
35701645924 2026-09-22 adjusted branch merge regression conflictor failed to link to annex (at git annex add)
35497963785 2026-09-20 conflict resolution movein regression, conversion annexed to git git-annex: unlock failed (at git annex unlock)
35429352903, 35319817132, 35196743368, 34944231019, 34574161202 09-19 … 09-11 (various) —

Both messages are the same underlying failure. Annex/Ingest.hs:213 (failed to link to annex) and Command/Unlock.hs:61 (unlock failed) are two callers of linkAnnex, and linkAnnex reaches LinkAnnexFailed through exactly two exits, neither of which reports anything:

  • Annex/Content.hs:633 — linkOrCopy returned Nothing: an IOException swallowed by catchDefaultIO Nothing (Annex/Content/LowLevel.hs:50), or cp / preserveGitMode returning False via catchBoolIO (:78).
  • Annex/Content.hs:646 — checksrcunchanged: the source file's InodeCache (inode, size, high-resolution mtime, compared with compareStrong, i.e. exact equality) differed before and after the copy, so the destination was deleted.

The second is the interesting one: on the unlock path the source is a frozen annex object that nothing is modifying, so it would mean two stats of an unchanged file disagreeing — plausible with NFS attribute caching, not with a local filesystem. But the CI transcripts cannot tell the two exits apart, because both are silent.

What is in this PR

  1. patches/20260922-f68b252dbd-diag-linkannex-failure.patch — adds a warning to each of those exits, including the before/after inode caches:

    • DIAGNOSTIC: linkAnnex: inode cache of <src> changed while copying it; before: <ino size secs nsecs>; after: … — says which field moved;
    • DIAGNOSTIC: linkOrCopy threw: <exception>, DIAGNOSTIC: checkedCopyFile: copyFileExternal failed / preserveGitMode failed / exception: ….

    Diagnostic only, on paths that already fail. All callers of the patched functions guard on the object being present, so these should not appear in normal operation. DEP-3 + SPDX header; verified to git apply cleanly against f68b252dbd and not to reverse-apply.

  2. Branch-local tuning of .github/workflows/build-ubuntu.yaml, so repeated runs cost little and stay quiet:

    • test-annex runs only the nfs-home flavor, three times per invocation;
    • test-annex-more and test-datalad are disabled;
    • the failure e-mails are disabled — failures here are the expected outcome and should not read as master being broken.

    These workflow changes are for this branch only, which is one more reason the PR should not be merged.

How it is run

Opening the PR builds the patch on Ubuntu, macOS, macOS ARM64 and Windows (they all trigger on patches/*.patch), which also serves as the compile check — the patch was written without a Haskell toolchain available. After that, only the Ubuntu workflow needs repeating: a nightly workflow_dispatch of build-ubuntu.yaml with ref=claude/elegant-meitner-y7sdk7 builds current upstream/master with the patch applied (BUILD_COMMIT defaults to origin/upstream/master; patches/ comes from the dispatched ref) and runs three nfs-home jobs.

Findings get reported back on this PR. Once a failing run names the cause, the patch has done its job and the branch can be deleted — nothing here is meant to land on master.

🤖 Generated with Claude Code

https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf


Generated by Claude Code

Yaroslav Halchenko and others added 2 commits September 22, 2026 13:18
The nightly `test-annex (nfs-home)` job fails in a different test each
night, always at `git annex add` of an unlocked file ("<file> failed to
link to annex") or at `git annex unlock` ("unlock failed").  Both are
linkAnnex returning LinkAnnexFailed, which is reachable by two paths
that print nothing at all:

  * linkOrCopy returned Nothing - a swallowed IOException, or cp /
    preserveGitMode returning False;
  * the source file's inode cache (inode, size, high resolution mtime)
    differed before and after the copy, so checksrcunchanged deleted
    the destination.

The transcripts in the CI logs cannot tell those apart, so add a warning
to each, including the before/after inode caches.  Diagnostic only, on
paths that already fail; drop it once the cause is known.

Not compile-tested locally (no GHC available where this was written);
a PR touching patches/*.patch triggers the builds, which will verify it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
This branch exists only to exercise the LinkAnnexFailed diagnostic patch
against NFS; it is not meant to be merged. Point its Ubuntu workflow at
that single question:

* test-annex runs only the nfs-home flavor, three times per invocation
  (one run reproduces the failure in roughly two nights out of three).
* test-annex-more and test-datalad are disabled; they say nothing about
  this failure and cost ~40 minutes of runners per invocation.
* The failure e-mails are disabled: failures here are the expected
  outcome, and an "Ubuntu build failed" mail per night would read as if
  master were broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf

Copy link
Copy Markdown
Collaborator Author

Answered on the first run: it is checksrcunchanged, and the mtime is what moves

Run 35733660775: the patch compiled on all four platforms, and two of the three nfs-home jobs failed. Both printed a diagnostic, and both are the same exit — Annex/Content.hs:646, the inode-cache comparison after the copy. Neither hit the linkOrCopy/swallowed-exception path.

Job rep 2, test import (To direction — worktree file into the annex):

      import import1/f 
        DIAGNOSTIC: linkAnnex: inode cache of import1/f changed while copying it; before: 8952864 16 1790084329 78602968; after: 8952864 16 1790084329 90671216

        import1/f failed to link to annex

Job rep 1, test edit (no pre-commit) (From direction — annex object out to the worktree):

        DIAGNOSTIC: linkAnnex: inode cache of .git/annex/objects/Kj/0x/SHA256E-s20--e394a389d787383843decc5d3d99b6d184ffa5fddeec23b911f9ee7fc8b9ea77/SHA256E-s20--e394a389d787383843decc5d3d99b6d184ffa5fddeec23b911f9ee7fc8b9ea77 changed while copying it; before: 9455246 20 1790084752 23778495; after: 9455246 20 1790084806 769318488

showInodeCache prints inode size seconds nanoseconds, so decoded:

inode size mtime before mtime after delta
import1/f 8952864 (same) 16 (same) 13:38:49.078602968 13:38:49.090671216 +12.07 ms
annex object 9455246 (same) 20 (same) 13:45:52.023778495 13:46:46.769318488 +54.75 s

Same inode, same size, only the mtime moves — and in the second case the source is a frozen annex object that nothing in the test writes to. compareStrong requires exact equality of the high-resolution mtime (Utility/InodeCache.hs), so linkAnnex concludes the source changed mid-copy, deletes the destination and returns LinkAnnexFailed. That surfaces as failed to link to annex on the To side and unlock failed on the From side, i.e. the two messages the nightlies have been alternating between.

Two flavours of the same root cause, both NFS attribute-cache artefacts rather than real modifications:

  • the 12 ms one is a sub-second mtime that differs between two stats within the same second — the client's locally-set value being replaced by the server's;
  • the 54.7 s one is a plainly stale cached attribute (acregmax is 60 s by default) that only gets revalidated during the copy.

Note also the knock-on damage: when the deleted destination is the annex object, everything downstream in that repo fails. Job rep 2 goes on to produce a run of content not available to send, openBinaryFile: does not exist and cp: cannot stat .git/annex/objects/ZP/kg/SHA256E-s2097152--… in the testremote type directory group — a missing object, not an independent bug.

Next steps

  1. Worth confirming by mounting the export with -o noac (or actimeo=0) in this branch's workflow: if the failures disappear, attribute caching is proven and NFS users have a workaround.
  2. Worth reporting upstream: the strict compareStrong mtime equality in linkAnnex's checksrcunchanged is not a safe assumption on NFS, where two stats of an unmodified file can disagree in the sub-second field, or return a value up to acregmax stale.

The harness has answered what it was built for; it can keep running for more samples, but the patch should not be merged.


Generated by Claude Code

The first harness run showed both failures are linkAnnex's
checksrcunchanged, with inode and size unchanged and only the mtime
moving (+12 ms in one case, +54.7 s in the other) — consistent with NFS
client attribute caching rather than a real modification.

Test that directly: run the nfs-home flavor three ways, three times
each, in one invocation so the arms see the same runner conditions.

* default  — what the nightlies mount, the control;
* actimeo0 — `-o actimeo=0`, attribute caching off and nothing else,
  the clean probe of the hypothesis;
* noac     — `-o noac`, that plus synchronous writes, the workaround
  usually recommended to NFS users.

The effective mount options are printed per job so the log records what
was actually negotiated. The suite timeout goes 3600 -> 7200s because
noac writes synchronously and is much slower; a job dying at the timeout
would make the arm unreadable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf

Copy link
Copy Markdown
Collaborator Author

Attribute caching confirmed: actimeo=0 alone makes it go away, 9/9 clean split

Run 35882579073 (head 756a55be) ran the nfs-home flavor three ways, three times each, on the same build and the same runner batch:

arm mount result wall clock
default as the nightlies mount 3/3 failed 8m17s, 8m54s, 10m27s
actimeo0 -o actimeo=0 3/3 passed 32m01s, 32m24s, 50m47s
noac -o noac 3/3 passed 29m17s, 34m30s, 34m54s

The mount options each job recorded, confirming the arms really differed:

default:  rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,…
actimeo0: rw,relatime,vers=4.2,…,namlen=255,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,…
noac:     rw,relatime,sync,vers=4.2,…,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,noac,…

actimeo=0 is sufficient. It disables attribute caching and changes nothing else; noac adds synchronous writes (note the extra sync in its options) and does no better. So the failure is the client's attribute cache, not the write path — which is what the mtime evidence implied and now has a controlled result behind it.

Cost of the workaround

Roughly 4×: a passing default job in the previous run (rep 3, same 26 groups) took 8m20s, against 32m01s for actimeo0 here. The control arm's 8–10 min above is not a fair comparison on its own, since a group that fails ends early.

New: a second check site with the same root cause

default,2 produced the familiar diagnostics — note both deltas are again mtime-only, inode and size unchanged:

add ../dir2/foo
  DIAGNOSTIC: linkAnnex: inode cache of ../dir2/foo changed while copying it; before: 9211545 20 1790178434 30902619; after: 9211545 20 1790178434 309979362
  ../dir2/foo failed to link to annex          →  +0.279 s, same second

unlock foo
  DIAGNOSTIC: linkAnnex: inode cache of .git/annex/objects/Kj/0x/SHA256E-s20--e394…/SHA256E-s20--e394… changed while copying it; before: 9464846 20 1790178735 562779131; after: 9464846 20 1790178776 778504714
  git-annex: unlock failed                     →  +41.216 s (stale cached attribute)

But default,1 and default,3 failed with no DIAGNOSTIC line at all. Their logs (complete, 1764 and 1810 lines) contain no failed to link to annex and no unlock failed; they died in testremote type directory with:

storeKey:   FAIL
  Exception: content changed while it was being sent
present True: FAIL

That message is Annex/Content.hs:730, the sameInodeCache check in prepSendAnnex — a different call site, same assumption: an object's inode cache must compare equal before and after the transfer. compareInodeCaches only relaxes to compareWeak when inodesChanged (the sentinal file) says inodes moved, and here they never do, so it stays exact and the unmodified object looks modified.

So the strict high-resolution mtime comparison bites in at least two places, and a fix confined to linkAnnex would leave transfers still failing on NFS.

Where this leaves it

  1. Root cause: established. An NFS client can return two different high-resolution mtimes for a file nothing wrote to — sub-second within the same second (12 ms, 279 ms seen) or plainly stale (41 s, 55 s seen, against a 60 s default acregmax). git-annex's compareStrong treats that as "the file changed".
  2. Workaround for NFS users: actimeo=0, at roughly 4× the wall clock. noac is not needed.
  3. Upstream report: worth filing with these numbers — the comparison is exact-equality on a timestamp the filesystem does not promise to be stable, at both linkAnnex's checksrcunchanged and prepSendAnnex's sameInodeCache.
  4. The harness has now answered both questions it was built for. The patch still must not be merged; the branch can go once a report exists.

Generated by Claude Code

Copy link
Copy Markdown
Collaborator Author

test-datalad (release) on Windows: pre-existing on master, not this PR

test-datalad (release) failed on this PR's head (756a55be). It is not this PR's failure:

  • it fails identically on master's scheduled Windows run from this morning, 35838651247 — where it is likewise the only failing job — and on every scheduled build-windows run since 09-19 (35706078128, 35581602966, 35500294962, 35431263046), all on base 2c4a6c2e;

  • both jobs end on the same single assertion:

    FAILED datalad/core/local/tests/test_run.py::test_run_merge_sub_under_plain_dir
      - AssertionError: no run-info merge commit found
    

    with the same environment (Python 3.10.11, pytest 9.1.1, datalad 1.6.3), and the two logs are the same length to the character.

It is also not plausibly reachable from this PR's diff: the only non-Ubuntu-workflow change here is the diagnostic patch, which adds warning calls on two paths that already fail (linkAnnex's LinkAnnexFailed exits), and this datalad test fails on a missing run-info merge commit, with no DIAGNOSTIC: line anywhere in either log.

No fix to port: the failure lives in datalad's own test against Windows, and there is no PR in this repository carrying one. Someone may want to track it separately — it has been red on master for five nights.

I did not spend a re-run on it. The re-run exists to confirm a failure that only looks unrelated; here the base branch already demonstrates it independently, and a Windows test-datalad job costs ~40 minutes to tell us what master's nightly says for free. Nothing about it blocks this branch, which is a diagnostic harness and never merges.


Generated by Claude Code

Copy link
Copy Markdown
Collaborator Author

Closing: the harness answered both questions

This branch was never meant to merge — it existed to make an intermittent NFS failure say why it was failing. It did, twice over, so it is being closed rather than left open to rot. Summary and pointers, for whoever finds this from a future NFS failure.

What it found

test-annex (nfs-home) fails because git-annex records a file's (inode, size, high-resolution mtime), copies the file, stats it again and compares the two exactly — and an NFS client can return two different mtimes for a file nothing wrote to. Four instrumented catches, all with inode and size unchanged and only the mtime moving: +12 ms, +279 ms (sub-second, same second) and +54.7 s, +41.2 s (a cached attribute gone stale, against the 60 s default acregmax).

Two call sites, not one — a fix confined to the first would leave transfers broken:

  • Annex/Content.hs:646 — checksrcunchanged in linkAnnex → failed to link to annex, unlock failed
  • Annex/Content.hs:730 — sameInodeCache in prepSendAnnex → content changed while it was being sent

Controlled run 35882579073, three arms × three reps on one build: default 3/3 failed, actimeo=0 3/3 passed, noac 3/3 passed. Since actimeo=0 disables attribute caching and nothing else, it is the attribute cache rather than the write path. Workaround for NFS users: mount with actimeo=0, at roughly 4× the wall clock (8m20s → 32m01s for the same 26 groups).

Not a regression in the failing code: nothing has touched compareStrong's use since at least 2025-01-01, and the strict comparison predates 2022-08. The nightly passed on 2026-03-13 and was failing by 2026-08-08, so what changed is the environment or the timing, not git-annex. Details in the issue below.

Where it lives now

  • #294 — the tracking issue: full evidence, the regression analysis, the workaround, and the upstream report still to be filed at git-annex.branchable.com.
  • con/eval-under#11 — reproducers, so nobody has to rebuild this harness: a mtime-stability target that shows the failure with no git-annex involved, a git-annex-linkannex target that reports a failure rate in minutes instead of a 20-minute pass/fail, and eval-under nfs --mount-opts, without which actimeo=0 could not be tried locally at all.

What was in here, for reference

  1. patches/20260922-f68b252dbd-diag-linkannex-failure.patch — warnings on the two silent LinkAnnexFailed exits, printing the before/after inode caches. It compiled on all four platforms and did its job on the first run. It is a debugging aid, not a fix, and should not be carried in patches/.
  2. Branch-local tuning of build-ubuntu.yaml — nfs-home only, repeated, other jobs and the failure mails off, later extended with the default/actimeo0/noac arms. Useful as a template if a similar question ever needs a harness; not useful on master.

Unrelated, but recorded because it showed up here: the Windows test-datalad (release) failure on this PR is not this branch's — it is datalad 1.6.3, where the is_managed_branch() xfail in test_run_merge_sub_under_plain_dir is registered after the assertions it is meant to cover, so it never applies. Fixed in datalad 1.6.4, released 2026-09-24, so the nightly should clear itself once CI picks that up.

The branch can be deleted.


Generated by Claude Code

yarikoptic-gitmate pushed a commit to con/eval-under that referenced this pull request Sep 30, 2026
Chasing con/git-annex#293 (git-annex intermittently failing on NFS with
"failed to link to annex" / "unlock failed") needed three things this
framework could not express. All three are small:

* `eval-under nfs --mount-opts OPTS` / `--export-opts OPTS`. NFS_OPTS
  was hardcoded to rw,async|rw,sync, so the client-side knobs that
  matter for this class of bug -- actimeo=0, noac, lookupcache=none,
  nocto, vers= -- were unreachable. Mount and export options stay in
  separate variables because an option valid in one is rejected by the
  other.

* target `mtime-stability`: write, stat, copy the way git-annex copies,
  stat again, compare (inode, size, high-res mtime) exactly, report a
  rate. No git-annex involved, so a red cell says "the filesystem",
  not "the application". This is the property git-annex assumes in
  Annex/Content.hs:linkAnnex and prepSendAnnex, and that NFS attribute
  caching breaks: in con/git-annex CI the mtime of an unmodified file
  moved by 12 ms, 279 ms and 41 s across a copy, and mounting with
  actimeo=0 made 3/3 failing runs pass.

* target `git-annex-linkannex`: the same question at the git-annex
  level -- loop `git annex unlock` (linkFromAnnex') and unlocked
  `git annex add` (linkToAnnex), report a failure rate in minutes
  rather than a pass/fail of the ~20 minute suite.

Both targets run on every backend, so BeeGFS and vfat get answered for
free; that grows the README grid from 20 cells to 30.

Tested: shellcheck clean; both targets run end-to-end on ext4 (0%, the
negative control); option strings verified for each flag combination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B89nUooZLfThcTA4fSMPGf
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant