Skip to content

Redfs ubuntu noble 6.8.0 58.60 updates - #207

Open
achhenderson wants to merge 10 commits into
DDNStorage:redfs-ubuntu-noble-6.8.0-58.60from
achhenderson:redfs-ubuntu-noble-6.8.0-58.60_updates
Open

Redfs ubuntu noble 6.8.0 58.60 updates#207
achhenderson wants to merge 10 commits into
DDNStorage:redfs-ubuntu-noble-6.8.0-58.60from
achhenderson:redfs-ubuntu-noble-6.8.0-58.60_updates

Conversation

@achhenderson

Copy link
Copy Markdown

No description provided.

hbirth and others added 10 commits August 24, 2026 20:41
A FUSE_NOTIFY_INVAL_INODE data invalidation means another (remote) entity
is modifying the file.

Rather than react to a single notify, keep a per-inode moving average of
how fast data invalidations arrive for the whole file: an EWMA of the
inter-arrival interval, updated under fi->lock on every notify
(fuse_notify_inval_hot()). The inode is latched only once the average
spacing drops below an internal threshold (FUSE_NOTIFY_DIO_INTERVAL) while
a local writer is open; a lone or occasional notify keeps the average high
and does not trip the switch.  The heuristic has no external knob -- its
parameters (EWMA weight, threshold, seed) are source-level constants.

Introduce the forced-direct-IO latch (FUSE_I_FORCE_DIO):

  - fuse_reverse_inval_inode() folds each data invalidation into the moving
    average and sets the latch when it trips with a local writer present;
  - fuse_file_{read,write}_iter() and fuse_cache_write_iter() route to the
    direct path while latched; fuse_dio_{wr_exclusive_lock,lock,unlock}()
    use the shared parallel-dio path and bypass the cached/uncached
    accounting;
  - fuse_file_io_open() opens new files uncached so they do not re-enter
    caching mode;
  - fuse_prepare_release() clears the latch once the last writer is gone
    and fuse_file_release() drops any clean folios a racing read
    repopulated; fuse_file_mmap() reverts to caching mode (a mapping needs
    the page cache).

Latching to direct IO is only coherent if no buffered write can deposit
dirty folios into the page cache after it has been dropped. Add a
per-inode rw_semaphore, wb_inval_rwsem, to serialise the buffered-write
page-cache dirtying against the latch transition. The writeback path
holds it for read around the dirtying and re-checks the latch under it;
fuse_reverse_inval_inode() holds it for write around its invalidate +
latch set. The notification may be delivered by the same server thread
that still owes a reply to an in-flight write holding the inode lock, so
it takes the rwsem with a trylock and never blocks: if the writer has gone
it skips the latch and only invalidates the notified range. The writer's
read-side section stays free of server round-trips because under fc->dlm
the partial-write RMW read is skipped.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit eeec69b)
[ahenderson: 6.8 has no iomap buffered-write path, so the writeback
  selector and the fuse_writeback_write_iter() context are dropped.
  generic_file_write_iter() is open-coded in fuse_cache_write_iter() so the
  coherency gate is taken inside i_rwsem, preserving the i_rwsem -> gate order
  the later gating commits depend on.  fuse_inode_uncached_io_start() takes one
  argument here; there is no FOPEN_PASSTHROUGH arm in fuse_file_io_open(),
  fuse_file_mmap() or fuse_file_{read,write}_iter(), and no fuse_inode_backing(),
  so those clauses are dropped]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
fuse_dio_lock() takes an uncached_io reference (via
fuse_inode_uncached_io_start()) only when FUSE_I_FORCE_DIO is clear,
while fuse_dio_unlock() decided whether to drop it by re-reading
FUSE_I_FORCE_DIO. On this tree the latch is toggled asynchronously by
the inode-invalidation notify-storm path, so the bit can differ between
the lock and the unlock of a single direct write:

 - clear at lock (reference taken), set before unlock: the reference is
   never dropped, leaving fi->iocachectr permanently negative and
   hanging the next caching-mode open;

 - set at lock (no reference), cleared before unlock:
   fuse_inode_uncached_io_end() is called without a matching start,
   tripping WARN_ON(fi->iocachectr >= 0) and corrupting the counter.

Record in fuse_dio_lock() whether a reference was actually taken and
have fuse_dio_unlock() drop it based on that captured decision instead
of re-testing the racy bit, so the accounting stays balanced regardless
of any FORCE_DIO transition mid-write.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit 6bf7ec0)
[ahenderson: 6.8 keeps the past-EOF exclusive-lock trigger, as the
  parallel-direct-write rework is not ported, so the accounting decision becomes a
  nested if / else-if rather than upstream's flat form;
  fuse_inode_uncached_io_start() takes one argument]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
Acquire the dlm lock from fuse server for the normal buffer read path
to ensure the distributed page cache across different nodes can be
co-existing and consistency. More importantly, this change will correct
the DLM lock and folio locks ordering for the buffer read path, thus
can avoid the potential deadlock between the buffer read and page cache
invalidation processes.

Signed-off-by Hai Zhong Zhou <hazhou@ddn.com>
(cherry picked from commit cf349bc)
[ahenderson: context only - 6.8 retains the older fc->dlm block in
  fuse_cache_write_iter() and its 'goto writethrough' structure, so the write-side
  hunk is re-indented.  The change itself is unmodified]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
A FUSE_NOTIFY_INVAL_INODE is a coherency event: once the server signals a
remote modify, no local read may return a page it has superseded.  The
invalidate ran unserialized against cache-serving reads, so a buffered read
could hand back a stale folio it still held a reference to.

Convert the per-inode wb_inval_rwsem to a percpu_rw_semaphore and take its
read side around the cache-serving read as well as the existing buffered
write.  The read side is per-CPU, so it scales on a shared file; the NOTIFY
takes the write side blocking, giving the invalidate priority -- it parks
new readers, drains in-flight ones, then drops the cache.  Every gated
invalidate now runs under the write side, not just the storm-latching one.

The gate is allocated only for writeback+dlm regular files and is NULL
elsewhere (best-effort invalidate, as before).  The blocking write side may
run on the notify-delivering server thread, so it is safe only under a
server that services request replies on other threads; redfs' dlm server
provides that contract.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit d6180a2)
[ahenderson: fuse_cache_wr_unlock() comes from the unported
  shared-i_rwsem commit, so the gate's early exits use plain inode_unlock();
  fuse_init_file_inode() gains the struct fuse_conn local that 6.8 lacks; no
  'writeback' selector, as 6.8 has no iomap branch]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
The buffered read path now acquires a DLM read lock via
fuse_get_dlm_lock(..., FUSE_PAGE_LOCK_READ). Before sending the
request to the server, fuse_get_dlm_lock() calls
fuse_dlm_range_is_locked() to skip regions we already hold.

That coverage check compared the held lock mode for exact equality
(range->mode != lock_mode), so a range we already hold with an
exclusive WRITE lock was reported as not-locked for a READ request.
Because fuse_dlm_lock_range() intentionally does not downgrade a
WRITE lock on a read, the region stays WRITE-locked and every
subsequent read re-requests a DLM read lock from the server. This
made read-after-write and re-read workloads flood the server with
redundant FUSE_DLM_WB_LOCK requests, never converging.

A held WRITE lock (exclusive) subsumes a READ lock. Treat a range as
uncovered only when the held mode is strictly weaker than the
requested mode (range->mode < lock_mode). READ requests are now
satisfied by either a READ or a WRITE lock, while WRITE requests
still require an existing WRITE lock (upgrade otherwise), matching
the compatibility rules already documented in fuse_dlm_lock_range().

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit 4e3bef0)
[ahenderson: applies unchanged]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
fuse_dlm_try_merge() locates the first merge candidate by walking
from rb_first_cached() until it reaches the region just granted.
The walk runs under the write-held cache rwsem on every
fuse_dlm_lock_range() call, and the tree it walks holds every cached
grant of the inode.  Strided writers (IOR hard-write) accumulate
grants that cannot merge with each other, so the tree keeps growing
and every new grant pays a scan of all grants below it -- quadratic
over the run, with fuse_dlm_range_is_locked() readers blocked behind
each scan.

Seed the merge with fuse_page_it_iter_first() on the region widened
by one unit to each side instead; finding the lowest overlapping
range is what the interval tree is there for.

This also repairs two edge cases of the linear scan: a region
starting at offset 0 made 'start - 1' wrap so the scan degenerated
and merging was silently skipped, and a region ending at U64_MAX
overflowed 'end + 1' in the loop bound, ending the merge after the
first range.  Both bounds now saturate.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit 67b2479)
[ahenderson: applies unchanged]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
The NOTIFY_INVAL_INODE revoke computed

    fuse_dlm_unlock_range(fi, offset, pg_end == -1 ? 0 : offset + len - 1)

which is wrong at both degenerate ends: a to-EOF invalidate (len <= 0,
e.g. a remote truncate) with offset > 0 becomes the inverted range
[offset, 0] and removes nothing, so the revoked grant stays visible to
the re-validating IO paths forever -- cached writes with no
server-side lock, zero-filled RMW reads; and an invalidate of byte 0
(offset 0, len 1) becomes [0, 0], the "destroy everything" sentinel,
wiping every grant of the inode.

Map the range in one helper shared by the gated and the ungated
branch: to-EOF revokes through U64_MAX, and the bounds widen to page
boundaries to match how grants are recorded -- revoking too much only
costs a re-request, too little leaves a stale grant.

Drop the in-band (0, 0) sentinel: whole-file invalidates walk the
normal removal path, release-all is fuse_dlm_cache_release_locks(),
and an inverted range is rejected with -EINVAL instead of silently
ignored.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit e6f09f5)
[ahenderson: applied ahead of the grant re-validation commit in this
  series, so this also moves the revoke under the gate write side - work upstream
  did in that later commit.  'backing' dropped from the ungated-branch comment, as
  6.8 has no backing-file support]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
Both cached IO paths request their DLM lock first and then go to sleep
on things a NOTIFY invalidate can be holding: the read path blocks on
the coherency gate (writer priority), the write path additionally
sleeps on a contended i_rwsem.  A NOTIFY invalidate running in that
window revokes exactly the lock just granted (fuse_dlm_unlock_range()),
so the task wakes up and populates or dirties the page cache with no
DLM coverage.

Close the window without ever sending a FUSE_DLM_WB_LOCK request while
holding the gate (a grant that had to wait on an invalidate delivered
to this same client would deadlock against our own gate hold):

 - Drop the lock record under the gate write side in
   fuse_reverse_inval_inode(), so revocation and page drop are one
   atomic step with respect to the gate.

 - After entering the gate read side, re-check the grant against the
   live lock tree; if it was revoked while we waited, drop the gate,
   re-request, re-enter and check again.  With the revoke now gated,
   passing the check means the lock cannot go away for the whole gate
   hold: a revoke arriving mid-operation parks until the IO is done.

 - Keep the write path's lock request ahead of the inode lock: the
   round trip must not capture the writer-priority i_rwsem for
   unbounded cluster-grant latency, and the in-gate re-validation
   already closes the grant-to-use window.  Only O_APPEND moves below
   the lock, because its range is the current EOF -- stable only under
   the exclusive inode lock.  This also fixes the append range itself:
   generic_write_checks() rewrites ki_pos to i_size for IOCB_APPEND,
   so the old 'i_size + ki_pos' double-counted (ki_pos is absolute,
   not relative) and locked a range disjoint from where the data
   lands.

fuse_get_dlm_lock() now reports whether the grant is recorded, and the
re-validation never re-requests a grant that failed, so it cannot spin
(the read path seeds this from its pre-gate request instead of
discarding that result).  A grant the server issued but that could not
be recorded (small-allocation -ENOMEM) reports
FUSE_DLM_GRANT_UNRECORDED: coverage exists cluster-wide, so failing
the IO would be wrong -- it proceeds, it just cannot re-validate.
Empty ranges are trivially held, so a zero-length IO neither sends a
doomed request nor spins in the retry loops.  The write path returns a
real failure to the caller instead of dirtying the cache without DLM
coverage; only -ENOSYS still degrades to a plain cached write, since
it means the server has no DLM at all and clears fc->dlm.  The read
path keeps falling through unlocked and additionally bounds its retry:
a reader-only inode has no force-DIO latch to end a revoke storm, so
after a few re-requests the read is served unlocked rather than
looping in the kernel for the duration of the storm.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit 259ade5)
[ahenderson: fuse_cache_wr_exclusive_lock() and fuse_cache_wr_unlock()
  come from the unported shared-i_rwsem commit; the cached write keeps the
  exclusive inode lock, so only fuse_cache_wr_dlm_lock() is taken from that hunk.
  The inode.c changes are already present, having come with the revoke-range
  commit applied earlier here.  The in-gate re-validation is omitted from the
  writethrough path, which never requests a DLM lock on 6.8]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
A FUSE_DLM_WB_LOCK reply and a NOTIFY invalidate are serviced on
different threads, so a revoke aimed at the grant a reply carries can
be processed before fuse_get_dlm_lock() records it: the revoke finds
nothing to remove, and the requester then records an already-dead
grant that no later NOTIFY will target -- a permanent false positive
for the re-validating IO paths.

Add a revocation generation to the lock cache, bumped under the cache
lock by every revoke path -- unconditionally, because the racing
revoke sees an empty overlap precisely when the grant is in flight.
fuse_get_dlm_lock() samples it before sending and records through
fuse_dlm_lock_range_gen(), which refuses with -EAGAIN once the
generation has moved; the grant is then re-requested instead of
recorded, bounded so a revoke storm cannot pin the IO here (past the
bound the failure reports like any request failure).

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit 8983871)
[ahenderson: applies unchanged]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
…y gate

Two revocation paths bypassed the revoke-under-gate invariant the IO
paths re-validate against:

 - fuse_reverse_inval_inode() skipped the gate once mapping_mapped()
   turned true, revoking concurrently with gate holders -- but
   fuse_cache_read_iter()/fuse_cache_write_iter() enter the gate
   unconditionally, so a single mmap() reopened the race.  Keep the
   gate for mmapped inodes; only the force-DIO latch stays disabled
   for them (a mapping needs the page cache, and fuse_file_mmap()
   reverts any latch it races with).

 - The local truncates in fuse_do_setattr() -- the atomic-O_TRUNC open
   shortcut and the after-setattr trim -- revoked and dropped the
   cache with no gate at all, so an already re-validated reader could
   repopulate the truncated range.  Take the gate write side around
   revoke + drop.  This cannot deadlock: both run under exclusive
   i_rwsem, which no gate holder waits on (the write path takes
   i_rwsem before the gate, the read path never takes it).

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
(cherry picked from commit be0662c)
[ahenderson: !fuse_inode_backing() dropped from the gate condition and
  its comment, as 6.8 has no backing-file support; the atomic-O_TRUNC hunk lands
  without the fi->lock / server_size lines, which come from the unported
  zero-fill-expanding-writes commit]
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
@hbirth

hbirth commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator

Do we have xfstest runs for this?
I like the effort to make sure 6.8 gets what is possible with the limited features it provides.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants