Skip to content

Avoid blocking folder navigation during directory size calculation - #118

Merged
thisisgm merged 1 commit into
thisisgm:mainfrom
nerdislb:fix/nonblocking-directory-sizes
Sep 22, 2026
Merged

thisisgm merged 1 commit into
thisisgm:mainfrom
nerdislb:fix/nonblocking-directory-sizes

Conversation

@nerdislb

@nerdislb nerdislb commented Sep 11, 2026 •

Copy link
Copy Markdown

Opening a small directory can wait behind a recursive size calculation for a directory in the previous view. dirsize currently walks inline in the backend event loop, so the 150 ms loading indicator can appear even when listing the destination itself takes less than a millisecond.

This moves size calculation to one persistent worker. The event loop continues handling navigation, and sizes arrive through the existing event channel. Viewport-only requests, the completed-result cache, the 2 s walk deadline, and the wire format are preserved.

  • List/sort invalidation and dirsizecancel invalidate unfinished work with a generation counter. Late replies cannot populate a different listing or a reused row index.
  • Cancellation is checked during traversal. At most one size job runs; a cancelled job keeps its slot until it returns, so a blocked filesystem does not cause an accumulating set of replacement threads.
  • Shutdown cancels size work without joining a potentially blocked filesystem call. Queued rows are discarded; already completed sizes survive an explicit dirsizecancel.

Cancellation is cooperative: a blocked syscall can still delay subsequent size results until it returns. It no longer blocks the navigation event loop. Other synchronous operations, including scanning the destination itself, are outside this change.

Validation

  • cargo build --locked and cargo build --release --locked passed.
  • cargo test --locked and cargo test --release --locked: 606 passed, 0 failed in each profile on this isolated branch.
  • BIN=./target/debug/flea bash tests/protocol.sh passed.
  • git diff --check passed; the four existing compiler warnings are unchanged.

This is targeted validation of the Rust/backend change, not a claim that the entire upstream GUI/test runner is green.

Four new unit tests cover cancellation before/during traversal, a deterministically paused worker whose stale result is rejected after row-index reuse, and a real 1,900-level directory tree. The worker reserves a 16 MiB stack for the existing recursive walk.

The protocol suite also uses a test-only LD_PRELOAD/opendir barrier to pin a running size walk and verify cancel/re-request, list, sort, and quit. All four scenarios pass against debug and release; unchanged upstream fails the responsiveness check as expected. This helper requires cc and Python 3. The production binary has no test hooks or new dependencies.

Local latency check

Five requests per release build on the same local Btrfs filesystem:

Build Request-to-rows samples (ms) Median (ms)
Upstream b992e764 1150.49, 237.51, 238.94, 237.87, 236.12 237.87
This patch 0.25, 0.17, 0.19, 0.19, 0.23 0.19

These are backend request-to-first-rows times, not end-to-end GUI frame times. For each request pair, list the parent of a large local source tree, request that tree's dirsize, wait 10 ms, then send a list for a small directory. Measure until the corresponding rows reply. Both arms use release builds and the same paths; no cache dropping or filesystem modifications are involved.

Summary by CodeRabbit

  • New Features

    • Directory size calculations now run in the background, keeping navigation and other commands responsive.
    • Directory size work can be cancelled when listing, sorting, or cancelling size calculations.
    • Completed results are checked for freshness, preventing outdated sizes from appearing after directory changes.
  • Bug Fixes

    • Improved shutdown behavior so the application can quit without waiting for an in-progress directory scan.
  • Documentation

    • Updated protocol documentation to describe background processing, cancellation, and result freshness.

@coderabbitai

coderabbitai Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 5ea275f3-340a-48c9-9af8-9cc428e028b3

📥 Commits

Reviewing files that changed from the base of the PR and between b992e76 and 21dace5.

📒 Files selected for processing (12)
  • AGENTS.md
  • docs/protocol.md
  • src/backend/dirsize.rs
  • src/backend/dirsizereq.rs
  • src/backend/dirsizeworker.rs
  • src/backend/events.rs
  • src/backend/mod.rs
  • src/backend/run.rs
  • src/backend/state.rs
  • tests/dirsize-async.py
  • tests/dirsize-block.c
  • tests/protocol.sh

Included review availability: Your plan provides up to 2 included reviews per hour; 1 remains after this review.


📝 Walkthrough

Walkthrough

Directory-size traversal now runs in a cancellable background worker. The event loop receives generation-checked results while remaining available for navigation. New unit, integration, protocol, and documentation updates cover cancellation and stale-result handling.

Changes

Asynchronous directory-size processing

Layer / File(s) Summary
Cancellable directory traversal
src/backend/dirsize.rs
The walker checks cancellation during recursive traversal and returns partial results when cancellation occurs.
Background worker and completion events
src/backend/dirsizeworker.rs, src/backend/events.rs, src/backend/mod.rs
A single-threaded worker runs size jobs with generation tracking and sends completed results through Event::DirSize.
Queue and event-loop integration
src/backend/dirsizereq.rs, src/backend/run.rs, src/backend/state.rs
The backend starts queued jobs, accepts valid completions, and cancels active work during cancellation, row removal, and shutdown.
Protocol documentation and regression coverage
docs/protocol.md, tests/dirsize-async.py, tests/dirsize-block.c, tests/protocol.sh, AGENTS.md
Documentation and tests cover asynchronous responses, cancellation, stale results, refreshed sizes, and quit behavior during a blocked walk.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Bug fix

Sequence Diagram(s)

sequenceDiagram
  participant Backend loop
  participant dirsizereq
  participant Worker
  participant Filesystem
  Backend loop->>dirsizereq: start_next(State)
  dirsizereq->>Worker: start(row, path)
  Worker->>Filesystem: walk_cancellable(path)
  Filesystem-->>Worker: DirSize result
  Worker-->>Backend loop: Event::DirSize(Done)
  Backend loop->>dirsizereq: report_done(Done)
Loading

Suggested reviewers: thisisgm

Merge Risk: ⚪ Minimal · up to 21dac

Directory-size calculation moves off the event loop while preserving cancellation and result freshness behavior. No current merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 20.59% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 34 functions across 10 files. (2 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: moving directory size calculation away from the event loop to keep folder navigation responsive.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 20.59% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 34 functions across 10 files. (2 skipped: 2 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@thisisgm

Copy link
Copy Markdown
Owner

YOU COOKED, will merge on next!

@viktorvillalobos

Copy link
Copy Markdown

Adding a data point in case it helps prioritise this one. On 0.2.1-3 (the Omarchy package) the inline walk is what makes flea feel slow to start for me.

My home has two folders full of JS monorepo worktrees, about 8.2M entries between them, mostly node_modules. With Start in set to Home:

  • the window maps at ~1.1 s and the listing is drawn at ~1.3 s
  • then the backend's main thread sits at 100% for ~3.6 s (warm page cache) walking those two folders for the Size column. Nothing else is answered until it finishes: entering a folder, the context menu, Open with
  • on a cold cache each big folder runs to the 2 s deadline off disk, so the first launch after boot is noticeably worse, and the same stall comes back whenever I pass through a folder whose visible subfolders are that large

Sampling /proc/<backend pid>/fd during the burst shows only directory fds under those trees, and the busy thread is the main one (openat / statx / getdents64), which matches the description here.

To confirm, I made DirSizes.plan() return [] locally. Backend CPU in the first 8 s after launch goes from 3.77 s to 0.00 s with the same 35 rows drawn, and dirSizeRequests goes from 1 to 0. That is what I'm running for now, at the cost of "-" in the Size column.

This PR would fix it properly and keep the sizes. It has shown as conflicting with main since 0.3.0 landed. I can test a rebased build against this tree and report numbers if that is useful.

@viktorvillalobos

Copy link
Copy Markdown

Follow-up with numbers: I built this branch as it stands (21dace5, on top of v0.2.0) and ran it next to the installed 0.2.1-3 on the same tree as above (~8.2M entries under two folders in ~).

Through the GUI: launch at ~, wait for the listing, give it 250 ms so the size walk is under way, then send Enter to that window to open a folder. Three alternating rounds each, warm page cache. Timings read through the flea IPC target (path, listInFlight, visibleRows), CPU from /proc/<backend>/task/*/stat.

0.2.1-3 installed this branch
window mapped 1166 / 1121 / 1142 ms 1135 / 1119 / 1139 ms
home listing drawn 1341 / 1211 / 1233 ms 1315 / 1309 / 1224 ms
Enter → folder opened 1599 / 1625 / 1630 ms 55 / 54 / 63 ms
backend main thread CPU, first 10 s 3.51 / 3.52 / 3.53 s 0.00 / 0.00 / 0.00 s
backend worker CPU, first 10 s 0 1.97 / 1.95 / 2.00 s

Startup itself is the same on both. The difference is what happens right after: on 0.2.1 the first action waits about 1.6 s behind the walk, on this branch it is about 57 ms.

Driving the backend directly over the protocol says the same thing. list ~, dirsize for the 15 folder rows, then list of a small folder straight away: the second rows reply comes back after 1845 ms on 0.2.1 and after 0.3 ms here.

Sizes still arrive. Staying in ~, all 15 dirsized replies came back within 4.1 s, two of them partial at the 2 s deadline, which is what I'd expect for those two trees. The walk also stops early once I navigate away, so the branch spends less CPU overall (about 2.0 s against 3.5 s).

Caveats: warm cache only, and I tested the branch as written, not a rebase onto current main. The conflicts are in src/backend/dirsize.rs and src/backend/dirsizereq.rs.

@thisisgm
thisisgm merged commit 02e865b into thisisgm:main Sep 22, 2026
1 check passed
@thisisgm

Copy link
Copy Markdown
Owner

Shipped in 02e865b2 (v0.3.2): your asynchronous folder sizes are in, and e5f8d4ee then made a size sort order folders by those walked bytes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants