Skip to content

Latest commit

 

History

History
163 lines (134 loc) · 8.42 KB

File metadata and controls

163 lines (134 loc) · 8.42 KB

Architecture

Diskern is a Cargo workspace with one engine and two frontends, plus a marketing site.

┌─────────────┐   ┌──────────────────┐
│ diskern-cli │   │ app/ (Tauri v2)  │   two thin frontends
└──────┬──────┘   └────────┬─────────┘
       │                   │
       └───────┬───────────┘
               ▼
      ┌─────────────────┐
      │  diskern-core   │   the engine: all logic lives here
      └─────────────────┘

Where things live

Path What it is
crates/diskern-core Engine: scanner, dedup, rules, risk, graph, quarantine
crates/diskern-cli diskern binary — engine from the terminal
app/ Tauri v2 desktop app (React frontend, Rust backend)
site/ Landing page, deployed to GitHub Pages
docs/ Project docs — running locally, rules, releasing
.github/workflows/ CI, release builds, Pages deploy

A scan, end to end

One function orchestrates almost all of it: report::build_with. Reading it alongside this section is the fastest way into the engine.

scanner ──► graph ──► rules + risk ──► dedup ──► report

1. Walk. scanner::scan walks the roots in parallel with jwalk, skipping excluded directories, and returns a FileEntry per file: path, size, modified and accessed times, whether it is a symlink. Metadata only — nothing is read or hashed here, and nothing is written ever.

2. Graph. graph::ImpactGraph::from_entries makes one pass looking for two things: directories holding a project marker (Cargo.toml, package.json, pyproject.toml), and directories that are dependency stores (target, node_modules, a virtualenv). It links each project to the store it owns, so the engine can later answer "how many live projects reference this?".

3. Classify. rules::RulesDb::classify matches the normalized path against the rule globs — first match wins, which is why protected rules are listed first — yielding a Category and a base Verdict. risk::downgrade then applies the graph's answer. Evidence can only make a verdict more cautious, never less, and Protected is final.

4. Dedup. dedup::find_duplicates_filtered buckets by size, BLAKE3-hashes only the files whose sizes collide (which skips most of a real disk), then buckets by hash. It runs after classification so that entries nothing will act on can sit it out — they have no place in an offer to keep one copy and drop the rest, and hashing them is the most expensive way to produce an unusable number.

5. Report. risk::assess adds an informational score and per-file evidence, and each entry becomes a Finding carrying its category, verdict, reclaimable bytes and the reasons that justify them. The headline total counts findings in full and adds only the duplicate copies nothing has counted yet.

Acting on a finding is a separate call: actions::quarantine_finding is the report-bound safety entry point. It looks up an exact path in the completed report, combines the report's graph-aware verdict with a fresh static-rule verdict by taking the stricter result, checks the report's size, modification time and symlink state without using access time, and then delegates to actions::quarantine. Missing or changed findings fail closed. The desktop backend invalidates the report when a newer scan starts, and a generation lease prevents a superseded scan from publishing or continuing an action. A completed report is a user-review snapshot: filesystem graph changes made without a new scan are not silently treated as a new report, so the user must scan again before relying on them. The final path-based move still has an ordinary OS-level TOCTOU window; the metadata check is a bounded stale-report defense, not a universal filesystem identity proof.

Adding a feature

Work out which layer owns it before writing anything — the answer is usually further down than it first looks.

Does a rule cover it? Teaching Diskern that some directory is a cache is data, not code: add it to rules/base.json with a test. No Rust, no new code paths, and it reaches both frontends at once.

Does it change a verdict? Then it belongs in rules, risk or graph, and the constraint in design decisions applies: evidence may only make a verdict more cautious. Add the evidence as a reason too — a verdict the user can't see the basis for is the thing Diskern exists not to ship.

Does it change what a scan finds or reports? scanner, dedup and report. Watch the cancellation flag: anything that loops over every entry has to check it, or a Cancel arriving during that stage does nothing.

Only then, the frontends. Add a Tauri command in commands.rs and register it in lib.rs, or a flag in main.rs. Both should be thin enough that the interesting part of your change is already tested in the engine before either sees it.

Design decisions

Frontends are thin. The CLI and the app both call the same engine functions; neither contains scanning or classification logic. Any new capability goes into diskern-core first.

Safety is layered, not sprinkled. Read-only scanning, quarantine instead of deletion, and deterministic rule-based verdicts are enforced in the engine (see the core README), so no frontend can accidentally weaken them.

AI is narration-only. The optional ai feature explains findings in plain language; it can never change a verdict. This keeps the engine fully auditable and offline-capable.

Quarantine destination safety

Within one process, quarantine holds the manifest mutex from destination selection through the move and manifest append (including rollback on append failure). Restore-from-manifest and purge share that mutex. This prevents two files whose names flatten alike from selecting the same destination, and keeps manifest rewrites from dropping a concurrent append. Copying a large file holds the lock for the duration of that operation. Report metadata and the caller's scan authority are rechecked under the lock immediately before the move, so requests waiting behind another action cannot use an already stale approval.

Moves reserve destinations without replacing existing entries: a hard link followed by source removal on supported filesystems, or a copy opened with create_new(true) otherwise. Broken symlinks also count as occupied names. The copy fallback preserves file permissions and only accepts regular files; it refuses symlinks rather than copying their targets. If source removal fails, the move attempts to remove its new destination while leaving the source intact. If the source can no longer be verified, it preserves the destination; cleanup failures and retained destination paths are reported in the error. These steps are not an atomic transaction or a power-loss recovery mechanism.

The mutex is process-local. Do not run independent processes against the same quarantine directory: payload creation refuses to overwrite another entry, but manifest updates are not protected by an OS-level lock. Cross-process locking and crash journaling belong to the recovery proposal below.

Proposed quarantine recovery

The quarantine crash recovery proposal documents durable intent records, startup reconciliation, legacy compatibility, and a staged implementation/test plan for issue #196. It is a design proposal; the current engine does not yet implement this recovery protocol.