Reproducible test harness for running arbitrary commands under a
temporarily-mounted filesystem. Purpose-built for catching
filesystem-specific behaviour bugs (rename semantics, close-on-exec,
timestamp granularity, locking, permissions, adjusted-branch fallbacks,
root_squash interactions) in tools like git-annex, DataLad, and
rsync -- classes of bugs that don't show up in plain-ext4 CI.
git-annex is the immediate demo target: this repo runs git annex test
-- plus git's own testsuite, pjdfstest, and a curated set of stress-ng
filesystem stressors -- against every backend on every push. The harness
itself is both backend- and suite-agnostic: new filesystems drop in as
bin/eval-under-<name> scripts, new suites as bin/ci/target-<name>.sh
(see below).
Read GOTCHAS.md before drawing conclusions from a red cell. It records the exact mkfs / mount / export settings each backend uses -- results only mean something relative to those -- and the root cause of every failure already run down, so nobody re-investigates a known one.
| Backend | git-annex test | git testsuite | stress-ng | pjdfstest |
|---|---|---|---|---|
| BeeGFS 7.4.6 | ||||
| BeeGFS 8.1.0 | ||||
| NFS (localhost) | ||||
| Loop vfat | ||||
| Loop ext4 |
Rows are backends (which filesystem the work happens on), columns
are targets (which suite is run on it). All 20 cells are one job
matrix in .github/workflows/test.yaml,
fanned out from evals/matrix.yaml -- adding a
filesystem or a suite is a data edit, not a code edit.
The badges are ours, not GitHub's: GitHub publishes one badge per
workflow file, so a per-cell grid is not something it can serve for a
single matrix workflow. Each run renders an SVG per cell onto the
gh-pages branch, and each badge links to that
cell's own job log via the status page --
GitHub has no stable URL for "the latest job of this matrix cell", so
the page's #<cell> anchor supplies the indirection. The page also
carries what a badge cannot: which run produced the result, how long
ago, which known issues a red cell's failures fall under, and which
failures are new.
State lives in
status.json on that branch and is
merged across runs, so a partial run -- "Re-run failed jobs" being the
common one -- updates only its own cells instead of blanking the rest.
Run bin/ci/gen-readme-matrix.sh after editing the matrix to refresh
the table.
A red cell is not automatically a bug: vfat has no symlinks, ownership, or xattrs, and NFS has its own locking and close-to-open rules. The matrix exists to make which filesystem breaks which layer visible at a glance.
Failures already understood are listed per test in
evals/known-issues.yaml, and each cell is
judged against it:
| Cell outcome | CI job | Badge |
|---|---|---|
| all tests pass | green | passing |
| every failure covered by a known issue | green | failing (known), still red |
| any failure no issue covers | red | N new failing |
| suite timed out, died, or its totals disagree with the parse | red | incomplete |
Separately, a known issue whose tests all pass is flagged (+N fixed?
on the badge, a notice on the run) without affecting the job, and one
that matched no test that ran is warned about as stale.
The "Known issues" section of GOTCHAS.md is generated from
the same file (bin/ci/known_issues.py gotchas). For a new failure,
bin/ci/known_issues.py draft <verdict.json> prints an issue stub from
the cell's logs-* artifact.
Budgets below are what the suite itself costs on the loop-ext4 cell; a
sync-heavy backend (NFS, BeeGFS) runs the same work considerably slower,
which is why each target carries its own timeout in evals/matrix.yaml.
| Target | What it runs | Prep | Budget |
|---|---|---|---|
git-annex |
The full git annex test suite |
Daily build from con/git-annex | tens of minutes |
git |
174 scripts of git's own testsuite (t0*.sh t1*.sh, ~10k assertions) |
Source build, pinned tag | ~4 min build + ~2 min on ext4 |
stress-ng |
20 curated filesystem stressors, one at a time, --verify |
apt install stress-ng |
~2 min |
pjdfstest |
POSIX conformance: 238 scripts, ~8800 assertions | Source build, pinned commit | ~3.5 min on ext4 |
Why these four, in order of how deep they sit:
- stress-ng is the cheapest and the bluntest: it hammers
rename,link/symlink,locka/lockf,xattr,chmod/chown,utimewith verification on, so a filesystem that returns success while doing the wrong thing is caught before anything is built on top of it. Each stressor runs separately, so "vfat has no xattrs" reports as a skip rather than poisoning the run. - pjdfstest is the specification check: it tells you which syscall
returned which errno where POSIX says otherwise. Where
git annex testsays "something is wrong", this saysrenamereturnsEEXISTinstead ofENOTEMPTY. - git is the integration check for the layer git-annex actually
stands on -- index, refs, object store, attributes. Git's suite is the
canonical filesystem-picky testbed and is maintained far more
rigorously than anything hand-written here. Only the trash directories
(where the tests' file operations happen) live on the mount; the build
itself stays on the runner's disk. It is driven through
prove, git's TAP harness, rather than the defaultmake testtarget -- undermakethe first failing script aborts the run before the totals are ever printed, whereasprovefinishes every script and ends with a summary naming each failure. - git-annex is the original motivation and the top of the stack.
Pinned upstream refs live in evals/matrix.yaml (with why each is
pinned where it is) and are bumped deliberately: with a moving
testsuite, a newly-red cell is ambiguous.
A 2024 report of 35+ git-annex test failures on
BeeGFS 7.4.6
tracked back to a rename-semantics quirk that BeeGFS exposes but
ext4/tmpfs usually mask. That's not a new class of story: DataLad's CI
has for years included NFS and vfat-loop flavours precisely because
those filesystems break git-annex assumptions in ways plain ext4
doesn't (eval_under_nfs,
eval_under_testloopfs).
This repo consolidates and generalises that pattern: one dispatcher
(bin/eval-under), one CI workflow driven by an external matrix
definition, a small backend script per filesystem, and a small target
script per test suite. Both axes slot in uniformly; the
git-annex-under-BeeGFS coverage that motivated the repo is now one cell
of that matrix.
# git-annex smoke under BeeGFS 7.4.6, HOME on the mount
sudo bin/eval-under beegfs --set-home -- bash -c '
cd "$HOME" && git init t && cd t && git annex init && git annex test
'
# Same, under a 200 MB xfs loop
sudo bin/eval-under loop --fs xfs --size 200 --set-home -- \
bash -c 'cd "$HOME" && git annex test'
# Under a localhost NFS export (async by default; --sync to reproduce
# the fsync-heavy slow path)
sudo bin/eval-under nfs --set-home -- bash -c 'cd "$HOME" && git annex test'
# Skip teardown to poke around after a failure
sudo bin/eval-under beegfs --set-home --keep -- some-failing-command
# Discover backends / read backend help
bin/eval-under --list
bin/eval-under nfs --help
# Version
bin/eval-under --version # -> eval-under 0.1.0-4-g1a2b3c4All backends accept --mount-point, --set-home, --keep, and their
own backend-specific options. See bin/eval-under BACKEND --help for
the full flag / env-var / default table per backend.
Without --mount-point (or EVAL_UNDER_MOUNT), a backend mounts on a
fresh per-run directory under $TMPDIR (/tmp if unset), next to the
backing state it creates there -- not on a fixed /mnt/<backend>:
~/.tmp/eval-under-nfs-B3UXj.orig exported backing dir
~/.tmp/eval-under-nfs-B3UXj.nfs the mount TMPDIR points at
So concurrent runs don't collide, nothing has to pre-exist under /mnt,
and a leftover mount from --keep says which run left it. An explicit
--mount-point still pins the mount wherever you want it; a directory
that was already there is only unmounted on teardown, never removed.
--version (-V) reports git describe --tags --always --dirty when
run from a checkout -- so 0.1.0 on a release tag, 0.1.0-4-g1a2b3c4
four commits past it, and a -dirty suffix for uncommitted changes. An
installed copy outside a checkout reports the VERSION_FALLBACK baked
into bin/eval-under, bumped with each release tag.
| Path | Purpose |
|---|---|
Vagrantfile + provision/ |
Ubuntu 24.04 libvirt VM with docker + BeeGFS + NFS + loop deps + git-annex |
bin/eval-under |
Dispatcher: routes to bin/eval-under-<backend> |
bin/eval-under-beegfs |
BeeGFS backend (containerised cluster + kernel client mount) |
bin/eval-under-nfs |
NFS backend (localhost loopback export) |
bin/eval-under-loop |
Loop-device backend (dd + losetup + mkfs. + mount) |
fixtures/beegfs/docker-compose-v7.yml |
BeeGFS v7 test cluster (mgmtd + meta + storage), network_mode: host |
fixtures/beegfs/docker-compose-v8.yml |
Same, for BeeGFS v8.x (different mgmtd command style / gRPC control plane) |
fixtures/beegfs/beegfs-*.conf.template |
Minimal client + helperd confs for the throwaway cluster |
evals/matrix.yaml |
Single source of truth: backends x targets, pinned upstream refs, per-target knobs |
bin/ci/matrix.sh |
Shell accessors over evals/matrix.yaml, sourced by every other bin/ci script |
bin/ci/matrix-json.sh |
Renders that file as the workflow's matrix: value (via fromJson) |
bin/ci/install-target.sh |
Runner-side prep for a target (apt package, or source build at a pinned tag) |
bin/ci/target-<target>.sh |
The suite itself, run inside the mount by bin/ci/run-under.sh |
evals/known-issues.yaml |
Known failures per cell and test |
bin/ci/collect-results.py |
Turns a suite's output into per-test results.tsv |
bin/ci/evals.py |
The matrix model and results.tsv format shared by the bin/ci Python scripts |
bin/ci/known_issues.py |
Validates the issues, judges a cell against them, regenerates GOTCHAS.md's list |
bin/ci/gen-readme-matrix.sh |
Regenerates the README badge grid from evals/matrix.yaml |
bin/ci/render-badge.sh |
Renders one status badge as a self-contained SVG |
bin/ci/update-status.py |
Merges a run's per-cell results into the persistent status.json |
bin/ci/render-report.py |
Renders status.json into the badge set + the report page |
bin/ci/publish-status.sh |
Ties those together and pushes the site to gh-pages |
bin/ci/run-checks.sh |
The repo's own checks (-h lists them); what checks.yaml runs |
tests/eval-under.bats |
CLI entry point: options, backend discovery, dispatch, --version |
tests/test_*.py, tests/data/ |
Unit tests of the results parsers and the known-issues classifier |
.github/workflows/test.yaml |
The whole matrix: one matrix job, 20 test cells, one publish job |
.github/workflows/checks.yaml |
run-checks.sh on every push and PR; minutes, no root, no mount |
drafts/git-annex-test-beegfs.yaml |
Copy-target workflow for con/git-annex (external PR target) |
The dev container this repo is usually edited in lacks CAP_SYS_MODULE
(no BeeGFS kmod) and doesn't run its own NFS server. Use the Vagrant VM:
vagrant up # first time: ~10 min
vagrant ssh
cd /vagrant
# Iterate on any backend:
sudo bin/eval-under beegfs --set-home -- bash -c '
cd "$HOME" && git init t && cd t && git annex init && git annex fsck
'
sudo bin/eval-under nfs --set-home -- git annex test
sudo bin/eval-under loop --fs vfat --set-home -- git annex test
# Or run a whole CI cell exactly as the runner would. install-target.sh
# is the one-off prep (source builds land in $EVAL_UNDER_SRC_DIR, not on
# the mount); run-under.sh then wraps the suite in the backend.
bin/ci/install-target.sh pjdfstest
sudo -E bin/ci/run-under.sh loop ext4 pjdfstest
sudo -E bin/ci/run-under.sh nfs n/a stress-ngOptional: install act in the VM to replay the GitHub workflow locally.
VAGRANT_INSTALL_ACT=1 vagrant provision
vagrant ssh -c 'cd /vagrant && act -j test'act runs the workflow in a container, so it can validate the YAML
flow but cannot exercise the BeeGFS kernel module or the host's NFS
server -- useful for shaking out workflow bugs, not for actual
filesystem testing.
Three layers, deliberately separate:
tests/*.bats-- theeval-underCLI itself: option handling, backend discovery, dispatch,--version. Unprivileged, mounts nothing, runs in about a second. The dispatcher is exercised against throwaway trees of stub backends, so adding a real backend does not mean rewriting the suite.tests/test_*.py-- the results parsers and the known-issues classifier, on trimmed real suite output undertests/data/.- the CI matrix (
.github/workflows/test.yaml) -- the backends' actual mount and teardown logic, by running upstream suites under them. Needs root, a kernel module and a live cluster; that is what the badge grid at the top reports.
# All checks, exactly what .github/workflows/checks.yaml runs:
bin/ci/run-checks.sh
# Or one at a time (-h lists them):
bin/ci/run-checks.sh shellcheck
bin/ci/run-checks.sh unit
# Or bats directly, when you want its own flags:
bats tests/
bats --filter version tests/bin/ci/install-check-deps.sh installs what they need; the Vagrant VM
runs the same script.
A --version caveat worth knowing when a check fails only in CI:
git describe needs tags, and actions/checkout fetches none by
default, which is why the checks workflow asks for fetch-depth: 0.
- Drop
bin/eval-under-<newbackend>next to the existing backends. It's picked up automatically bybin/eval-under --list. - Follow the pattern:
set -eu,${SUDO[@]}arrays for root,trap teardown EXIT, a here-doc'dusage(),EVAL_UNDER_<BACKEND>_*env vars for backend-specific options, and the common--mount-point/--set-home/--keepflags on top. Default the mountpoint to$MNT_BASE.<backend>, derived from onemktemp -u "${TMPDIR:-/tmp}/eval-under-<backend>-XXXXX"base shared with the backend's backing state -- never a fixed/mnt/<backend>. - At the end, run the wrapped command with
TMPDIR,DATALAD_TESTS_TEMP_DIR, and (if--set-home)HOMEpointing at the mount. - Add a row to
backends:inevals/matrix.yaml, then runbin/ci/gen-readme-matrix.shto refresh the README grid above. The workflow picks the new cells up on its own. Commit the result. - Teach
bin/ci/install-backend.shhow to install its client packages. - Update
provision/setup.shif the backend needs new host packages.
- Write
bin/ci/target-<name>.sh. It runs inside the mount, withTMPDIR(andHOME, with--set-home) already pointing at the filesystem under test. Exit non-zero on failure; skip -- loudly -- rather than fail on operations the filesystem genuinely cannot do. - Teach
bin/ci/install-target.shhow to prepare it on the runner. Build source trees into$EVAL_UNDER_SRC_DIR(the runner's own disk), never onto the mount: only the suite's I/O should exercise the filesystem under test. Pin any upstream checkout to a tag. - Add an entry to
targets:inevals/matrix.yamlwith itslabel,timeout,loop-size-mb,needs-root, andneeds-git-annex. - Teach
bin/ci/collect-results.pyto turn its output into per-test results. Prefer a suite that speaks TAP, and cross-check the parse against the suite's own totals; without an adapter every cell of the new column reportsincomplete. - Run
bin/ci/gen-readme-matrix.shand commit the new README column. bin/ci/shellcheck.shbefore committing (CI runs the same check).
bin/eval-under-beegfs+fixtures/beegfs/*+ a copy ofdrafts/git-annex-test-beegfs.yaml-> PR to con/git-annex once validated, either as a new workflow or as a matrix flavour oftest-annexinbuild-ubuntu.yaml.- Optionally, a slimmed-down smoke workflow -> PR to ThinkParQ/beegfs-containers addressing their issue #21.
- libvirt +
vagrant-libvirtplugin (Debian/Ubuntu:apt install vagrant libvirt-daemon-systemthenvagrant plugin install vagrant-libvirt) - ~30 GB free disk, ~6 GB free RAM for the VM
Default box is cloud-image/ubuntu-24.04 (Canonical's official image,
has a libvirt provider). To use virtualbox instead:
VAGRANT_DEFAULT_PROVIDER=virtualbox vagrant up -- that switches the
box to bento/ubuntu-24.04 automatically.
Machine-readable per the REUSE specification:
- License texts live in
LICENSES/(MIT for project-native files, Apache-2.0 for the ThinkParQ-derived BeeGFS Docker Compose fixtures). - Path-to-license mappings live in
REUSE.toml. - Absorbed files preserve upstream attribution:
bin/eval-under-nfsandbin/eval-under-loopcredit the DataLad developers (MIT);fixtures/beegfs/docker-compose-v{7,8}.ymlcredit ThinkParQ GmbH (Apache-2.0). - Validate with
uvx --from reuse reuse lint(orpip install reuse && reuse lint). This project ships 100% REUSE-compliant.
New files contributed to this repo don't need per-file SPDX headers;
they are covered by the REUSE.toml catchall block. Add per-file
headers (or an additional [[annotations]] block) only if the file
carries a different license or additional attribution.
- Bug that motivated the BeeGFS coverage: https://git-annex.branchable.com/bugs/35_failed_tests_on_beegfs/
- Datalad's original per-filesystem wrappers:
tools/eval_under_nfs,tools/eval_under_testloopfs - BeeGFS containerisation guidance: https://doc.beegfs.io/latest/advanced_topics/containers.html
- BeeGFS CSI driver e2e workflow (client install patterns): https://github.com/ThinkParQ/beegfs-csi-driver/blob/main/.github/workflows/build-test-publish.yaml