Diskern is a Cargo workspace with one engine and two frontends, plus a marketing site.
┌─────────────┐ ┌──────────────────┐
│ diskern-cli │ │ app/ (Tauri v2) │ two thin frontends
└──────┬──────┘ └────────┬─────────┘
│ │
└───────┬───────────┘
▼
┌─────────────────┐
│ diskern-core │ the engine: all logic lives here
└─────────────────┘
| Path | What it is |
|---|---|
crates/diskern-core |
Engine: scanner, dedup, rules, risk, graph, quarantine |
crates/diskern-cli |
diskern binary — engine from the terminal |
app/ |
Tauri v2 desktop app (React frontend, Rust backend) |
site/ |
Landing page, deployed to GitHub Pages |
docs/ |
Project docs — running locally, rules, releasing |
.github/workflows/ |
CI, release builds, Pages deploy |
One function orchestrates almost all of it:
report::build_with. Reading it
alongside this section is the fastest way into the engine.
scanner ──► graph ──► rules + risk ──► dedup ──► report
1. Walk. scanner::scan
walks the roots in parallel with jwalk, skipping excluded directories,
and returns a FileEntry per file: path, size, modified and accessed
times, whether it is a symlink. Metadata only — nothing is read or
hashed here, and nothing is written ever.
2. Graph.
graph::ImpactGraph::from_entries
makes one pass looking for two things: directories holding a project
marker (Cargo.toml, package.json, pyproject.toml), and directories
that are dependency stores (target, node_modules, a virtualenv). It
links each project to the store it owns, so the engine can later answer
"how many live projects reference this?".
3. Classify. rules::RulesDb::classify
matches the normalized path against the rule globs — first match wins,
which is why protected rules are listed first — yielding a Category
and a base Verdict. risk::downgrade
then applies the graph's answer. Evidence can only make a verdict more
cautious, never less, and Protected is final.
4. Dedup. dedup::find_duplicates_filtered
buckets by size, BLAKE3-hashes only the files whose sizes collide (which
skips most of a real disk), then buckets by hash. It runs after
classification so that entries nothing will act on can sit it out — they
have no place in an offer to keep one copy and drop the rest, and
hashing them is the most expensive way to produce an unusable number.
5. Report. risk::assess adds
an informational score and per-file evidence, and each entry becomes a
Finding carrying its category, verdict, reclaimable bytes and the
reasons that justify them. The headline total counts findings in full
and adds only the duplicate copies nothing has counted yet.
Acting on a finding is a separate call:
actions::quarantine_finding is the
report-bound safety entry point. It looks up an exact path in the completed
report, combines the report's graph-aware verdict with a fresh static-rule
verdict by taking the stricter result, checks the report's size, modification
time and symlink state without using access time, and then delegates to
actions::quarantine. Missing or
changed findings fail closed. The desktop backend invalidates the report when
a newer scan starts, and a generation lease prevents a superseded scan from
publishing or continuing an action. A completed report is a user-review
snapshot: filesystem graph changes made without a new scan are not silently
treated as a new report, so the user must scan again before relying on them.
The final path-based move still has an ordinary OS-level TOCTOU window; the
metadata check is a bounded stale-report defense, not a universal filesystem
identity proof.
Work out which layer owns it before writing anything — the answer is usually further down than it first looks.
Does a rule cover it? Teaching Diskern that some directory is a cache
is data, not code: add it to
rules/base.json with a test.
No Rust, no new code paths, and it reaches both frontends at once.
Does it change a verdict? Then it belongs in rules, risk or
graph, and the constraint in design decisions
applies: evidence may only make a verdict more cautious. Add the evidence
as a reason too — a verdict the user can't see the basis for is the
thing Diskern exists not to ship.
Does it change what a scan finds or reports? scanner, dedup and
report. Watch the cancellation flag: anything that loops over every
entry has to check it, or a Cancel arriving during that stage does
nothing.
Only then, the frontends. Add a Tauri command in
commands.rs and register it in
lib.rs, or a flag in
main.rs. Both should be thin
enough that the interesting part of your change is already tested in the
engine before either sees it.
Frontends are thin. The CLI and the app both call the same engine
functions; neither contains scanning or classification logic. Any new
capability goes into diskern-core first.
Safety is layered, not sprinkled. Read-only scanning, quarantine instead of deletion, and deterministic rule-based verdicts are enforced in the engine (see the core README), so no frontend can accidentally weaken them.
AI is narration-only. The optional ai feature explains findings in
plain language; it can never change a verdict. This keeps the engine fully
auditable and offline-capable.
Within one process, quarantine holds the manifest mutex from destination selection through the move and manifest append (including rollback on append failure). Restore-from-manifest and purge share that mutex. This prevents two files whose names flatten alike from selecting the same destination, and keeps manifest rewrites from dropping a concurrent append. Copying a large file holds the lock for the duration of that operation. Report metadata and the caller's scan authority are rechecked under the lock immediately before the move, so requests waiting behind another action cannot use an already stale approval.
Moves reserve destinations without replacing existing entries: a hard link
followed by source removal on supported filesystems, or a copy opened with
create_new(true) otherwise. Broken symlinks also count as occupied names.
The copy fallback preserves file permissions and only accepts regular files;
it refuses symlinks rather than copying their targets. If source removal fails,
the move attempts to remove its new destination while leaving the source intact.
If the source can no longer be verified, it preserves the destination; cleanup
failures and retained destination paths are reported in the error. These steps
are not an atomic transaction or a power-loss recovery mechanism.
The mutex is process-local. Do not run independent processes against the same quarantine directory: payload creation refuses to overwrite another entry, but manifest updates are not protected by an OS-level lock. Cross-process locking and crash journaling belong to the recovery proposal below.
The quarantine crash recovery proposal documents durable intent records, startup reconciliation, legacy compatibility, and a staged implementation/test plan for issue #196. It is a design proposal; the current engine does not yet implement this recovery protocol.