Skip to content

Redfs rhel10 0 - #250

Merged
hbirth merged 5 commits into
DDNStorage:redfs-rhel10_0from
hbirth:redfs-rhel10_0
Oct 11, 2026
Merged

hbirth merged 5 commits into
DDNStorage:redfs-rhel10_0from
hbirth:redfs-rhel10_0

Conversation

@hbirth

@hbirth hbirth commented Oct 11, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

hbirth and others added 5 commits October 11, 2026 09:37
Convert space indentation to tabs, put block comment delimiters on
their own lines, align continuation lines to the open parenthesis and
drop the double dash as comment punctuation.

Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
With the writeback cache, foreground requests can wait in userspace on
DLM revokes. A revoke waits on page invalidation, which waits on
writeback. If foreground requests held every entry of an io-uring
queue, background writeback could not be sent and none of them would
complete.

When the connection has the writeback cache, each queue now keeps a
share of its entries for background requests. Background requests may
use a quarter of the entries, but at least 2. Foreground requests may
use the rest, but at least 8. A foreground request over the limit waits
on the new fuse_req_fg_queue and only moves to fuse_req_queue when a
slot frees up. Background requests keep their own path to free entries
and are held to their cap, except that each queue may always have one
active background request. If a queue has fewer than 10 entries both
minimums cannot hold, and the foreground minimum wins. The foreground
count is released on completion, on commit errors and on entry
teardown.

Requests marked no_fg_limit skip the foreground limit. These are the
page cache reads and writes (redfs_do_readfolio, redfs_send_readpages
and redfs_send_write_pages) and the flush from write_inode. A page
fault holds the inode range lock while its read is outstanding, and the
revoke that is waiting for that range lock would never finish if the
read were queued behind other foreground requests. The flush is sent by
writeback itself. Without the writeback cache, dispatch is unchanged.

Signed-off-by: Hai Zhong Zhou <hazhou@ddn.com>
…ache

With the writeback cache, a request can wait in userspace on a DLM
revoke. The revoke waits on page invalidation on this node, which waits
on writeback and on locked pages being read. Those page reads and
writes are themselves fuse requests. If requests that may block in
userspace held every entry of an io-uring queue, the page I/O could
never be sent and none of the requests would complete.

Add a uring_critical flag to fuse_args for requests on the path that
invalidation waits for. These are the page reads in fuse_do_readfolio()
and fuse_send_readpages() (both sync and async readahead), the page
writes in fuse_send_write_pages() and fuse_send_writepage(), and
fuse_flush_times(), which writeback calls. The daemon must never block
a critical request on a DLM revoke that waits on page invalidation on
this node.

When the connection has the writeback cache, a quarter of the entries
of each queue, but at least 2, are critical entries. Non-critical
requests, foreground and background together, may hold only the other
entries, counted in active_noncritical. A queue always keeps at least
one entry for non-critical requests, and at least one critical entry
when it has 2 or more entries. Critical requests are not counted in
that limit and may use any entry.

A non-critical foreground request over the limit waits on the new
fuse_req_fg_queue. A non-critical background request waits on
fuse_req_bg_queue, now limited by both the background limits and the
non-critical limit. Critical background requests wait on the new
fuse_req_bg_crit_queue, so they never sit behind non-critical ones, and
are admitted first. They are still held to max_background, except that
each queue may always have one critical background request active.
Critical foreground requests go straight to fuse_req_queue.

When a non-critical slot is freed, it is offered first to the same
kind of request, foreground or background, that freed it, so neither
kind can starve the other. Registering a new entry raises the limit and
admits waiting requests of both kinds. The counts are released through
one helper, fuse_uring_end_active(), on completion, on commit errors
and on request removal. Entry teardown and queue abort also drop the
non-critical count. Without the writeback cache, dispatch is unchanged.

Signed-off-by: Hai Zhong Zhou <hazhou@ddn.com>
fuse_dlm_lock_range() flipped the mode of every READ range overlapping
a newly recorded WRITE grant in its entirety, not just the part the
grant covers.  A whole-file READ grant -- read(2) leaves one behind via
readahead -- overlapped by any WRITE grant was thus recorded as a
whole-file WRITE grant, and fuse_dlm_lock_is_held() then reported WRITE
coverage the server never gave.  Every write-side check trusts that
record: the fault fast path, fuse_get_dlm_lock()'s early exit, and the
write path's re-validation all see the range as covered, no lock
request reaches the server, and writeback sends WRITEs the server
observes under a read lock.

Split the overlapping READ range at the grant boundaries and upgrade
only the intersection.  The splits preserve their mode, so an
allocation failure partway through still leaves the tree describing
exactly the grants held.

Found probing the fault-around fix in the next commit: the probe's
first-touch write fault should have requested a WRITE grant for its
batch window, but the poisoned cache reported the page already covered
and no request reached the server.

Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
With the writeback cache on, page_mkwrite skipped the DLM lock request
entirely, on the assumption that fuse_filemap_fault() had already taken
a write grant for any page a shared-writable mapping could dirty.  That
assumption does not hold: fault-around (filemap_map_pages) PTE-maps any
resident folio in its window without calling ->fault for it, so a folio
read in under a READ grant -- by read(2) or readahead -- can take its
first write fault with no WRITE grant ever requested.  The page is
dirtied anyway and writeback later sends a WRITE the server sees under
a read lock (DFS cases 040, 080, 081, 090 and 200).

page_mkwrite cannot take the grant itself: it is called with mmap_lock
(or the per-VMA lock) held and has no VM_FAULT_RETRY protocol to drop
it, while the buffered write path holds the IO range lock LOCKED while
faulting in its source buffer, which takes mmap_lock.  With a pending
mmap_lock writer queued between the two, blocking in page_mkwrite on
the range lock or the DLM round trip recreates the three-party deadlock
that 94effa2 ("fuse: do not block the fault path under mmap_lock")
removed from the fault path.

Instead, give shared-writable VMAs under the writeback cache with DLM a
vm_operations_struct without ->map_pages, so the first touch of every
page goes through fuse_filemap_fault(), which records a WRITE grant for
its batch window before any PTE is installed and already drops
mmap_lock to block.  page_mkwrite keeps a non-blocking check of the
recorded grant -- the same LOCKED fence the fault fast path uses -- and
returns VM_FAULT_NOPAGE on a miss.  A verified miss can only mean a
revoke in flight, whose invalidation clears the PTE, so the re-fault
resolves through fuse_filemap_fault(); a trylock-contended miss
re-faults with mmap_lock released in between, so the range lock holder
can make progress.

This also makes fuse_filemap_fault()'s FAULT_FLAG_TRIED pass take the
drop-and-retry slow path instead of blocking in place: with every first
touch of a shared-writable page now funnelled through ->fault, blocking
there on the range lock with mmap_lock held steps into page_mkwrite's
old role in the same three-party cycle, and did so instantly under a
buffered write sourced from an mmap of the file racing munmap and a
write fault.  Retries may repeat (4064b98 "mm: allow VM_FAULT_RETRY for
multiple times"), so waiting for a contended range must happen with
mmap_lock dropped, every time; blocking in place remains only for
callers that allow no retries at all.

The non-writeback-cache arm keeps the existing single-page
fuse_get_page_mkwrite_lock() behaviour.

Reported-by: Hai Zhong Zhou
Suggested-by: Yong Ze Chen
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
@hbirth
hbirth requested review from hazhou-ddn and yongzech October 11, 2026 07:48

@hazhou-ddn hazhou-ddn left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@hbirth
hbirth merged commit 52b0cfb into DDNStorage:redfs-rhel10_0 Oct 11, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants