Repository navigation
Redfs rhel10 0 - #250
Merged
Merged
Redfs rhel10 0#250
Conversation
Convert space indentation to tabs, put block comment delimiters on their own lines, align continuation lines to the open parenthesis and drop the double dash as comment punctuation. Signed-off-by: Horst Birthelmer <hbirthelmer@ddn.com>
With the writeback cache, foreground requests can wait in userspace on DLM revokes. A revoke waits on page invalidation, which waits on writeback. If foreground requests held every entry of an io-uring queue, background writeback could not be sent and none of them would complete. When the connection has the writeback cache, each queue now keeps a share of its entries for background requests. Background requests may use a quarter of the entries, but at least 2. Foreground requests may use the rest, but at least 8. A foreground request over the limit waits on the new fuse_req_fg_queue and only moves to fuse_req_queue when a slot frees up. Background requests keep their own path to free entries and are held to their cap, except that each queue may always have one active background request. If a queue has fewer than 10 entries both minimums cannot hold, and the foreground minimum wins. The foreground count is released on completion, on commit errors and on entry teardown. Requests marked no_fg_limit skip the foreground limit. These are the page cache reads and writes (redfs_do_readfolio, redfs_send_readpages and redfs_send_write_pages) and the flush from write_inode. A page fault holds the inode range lock while its read is outstanding, and the revoke that is waiting for that range lock would never finish if the read were queued behind other foreground requests. The flush is sent by writeback itself. Without the writeback cache, dispatch is unchanged. Signed-off-by: Hai Zhong Zhou <hazhou@ddn.com>
…ache With the writeback cache, a request can wait in userspace on a DLM revoke. The revoke waits on page invalidation on this node, which waits on writeback and on locked pages being read. Those page reads and writes are themselves fuse requests. If requests that may block in userspace held every entry of an io-uring queue, the page I/O could never be sent and none of the requests would complete. Add a uring_critical flag to fuse_args for requests on the path that invalidation waits for. These are the page reads in fuse_do_readfolio() and fuse_send_readpages() (both sync and async readahead), the page writes in fuse_send_write_pages() and fuse_send_writepage(), and fuse_flush_times(), which writeback calls. The daemon must never block a critical request on a DLM revoke that waits on page invalidation on this node. When the connection has the writeback cache, a quarter of the entries of each queue, but at least 2, are critical entries. Non-critical requests, foreground and background together, may hold only the other entries, counted in active_noncritical. A queue always keeps at least one entry for non-critical requests, and at least one critical entry when it has 2 or more entries. Critical requests are not counted in that limit and may use any entry. A non-critical foreground request over the limit waits on the new fuse_req_fg_queue. A non-critical background request waits on fuse_req_bg_queue, now limited by both the background limits and the non-critical limit. Critical background requests wait on the new fuse_req_bg_crit_queue, so they never sit behind non-critical ones, and are admitted first. They are still held to max_background, except that each queue may always have one critical background request active. Critical foreground requests go straight to fuse_req_queue. When a non-critical slot is freed, it is offered first to the same kind of request, foreground or background, that freed it, so neither kind can starve the other. Registering a new entry raises the limit and admits waiting requests of both kinds. The counts are released through one helper, fuse_uring_end_active(), on completion, on commit errors and on request removal. Entry teardown and queue abort also drop the non-critical count. Without the writeback cache, dispatch is unchanged. Signed-off-by: Hai Zhong Zhou <hazhou@ddn.com>
fuse_dlm_lock_range() flipped the mode of every READ range overlapping a newly recorded WRITE grant in its entirety, not just the part the grant covers. A whole-file READ grant -- read(2) leaves one behind via readahead -- overlapped by any WRITE grant was thus recorded as a whole-file WRITE grant, and fuse_dlm_lock_is_held() then reported WRITE coverage the server never gave. Every write-side check trusts that record: the fault fast path, fuse_get_dlm_lock()'s early exit, and the write path's re-validation all see the range as covered, no lock request reaches the server, and writeback sends WRITEs the server observes under a read lock. Split the overlapping READ range at the grant boundaries and upgrade only the intersection. The splits preserve their mode, so an allocation failure partway through still leaves the tree describing exactly the grants held. Found probing the fault-around fix in the next commit: the probe's first-touch write fault should have requested a WRITE grant for its batch window, but the poisoned cache reported the page already covered and no request reached the server. Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
With the writeback cache on, page_mkwrite skipped the DLM lock request
entirely, on the assumption that fuse_filemap_fault() had already taken
a write grant for any page a shared-writable mapping could dirty. That
assumption does not hold: fault-around (filemap_map_pages) PTE-maps any
resident folio in its window without calling ->fault for it, so a folio
read in under a READ grant -- by read(2) or readahead -- can take its
first write fault with no WRITE grant ever requested. The page is
dirtied anyway and writeback later sends a WRITE the server sees under
a read lock (DFS cases 040, 080, 081, 090 and 200).
page_mkwrite cannot take the grant itself: it is called with mmap_lock
(or the per-VMA lock) held and has no VM_FAULT_RETRY protocol to drop
it, while the buffered write path holds the IO range lock LOCKED while
faulting in its source buffer, which takes mmap_lock. With a pending
mmap_lock writer queued between the two, blocking in page_mkwrite on
the range lock or the DLM round trip recreates the three-party deadlock
that 94effa2 ("fuse: do not block the fault path under mmap_lock")
removed from the fault path.
Instead, give shared-writable VMAs under the writeback cache with DLM a
vm_operations_struct without ->map_pages, so the first touch of every
page goes through fuse_filemap_fault(), which records a WRITE grant for
its batch window before any PTE is installed and already drops
mmap_lock to block. page_mkwrite keeps a non-blocking check of the
recorded grant -- the same LOCKED fence the fault fast path uses -- and
returns VM_FAULT_NOPAGE on a miss. A verified miss can only mean a
revoke in flight, whose invalidation clears the PTE, so the re-fault
resolves through fuse_filemap_fault(); a trylock-contended miss
re-faults with mmap_lock released in between, so the range lock holder
can make progress.
This also makes fuse_filemap_fault()'s FAULT_FLAG_TRIED pass take the
drop-and-retry slow path instead of blocking in place: with every first
touch of a shared-writable page now funnelled through ->fault, blocking
there on the range lock with mmap_lock held steps into page_mkwrite's
old role in the same three-party cycle, and did so instantly under a
buffered write sourced from an mmap of the file racing munmap and a
write fault. Retries may repeat (4064b98 "mm: allow VM_FAULT_RETRY for
multiple times"), so waiting for a contended range must happen with
mmap_lock dropped, every time; blocking in place remains only for
callers that allow no retries at all.
The non-writeback-cache arm keeps the existing single-page
fuse_get_page_mkwrite_lock() behaviour.
Reported-by: Hai Zhong Zhou
Suggested-by: Yong Ze Chen
Signed-off-by: Allison Henderson <allison.henderson@ddn.com>
yongzech
approved these changes
Oct 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.