Skip to content

feat(client): offline DB compaction for nethermind and besu, geth journal check, db compact CLI - #351

Draft
qu0b wants to merge 5 commits into
masterfrom
qu0b/db-compaction-rocksdb
Draft

qu0b wants to merge 5 commits into
masterfrom
qu0b/db-compaction-rocksdb

Conversation

@qu0b

@qu0b qu0b commented Oct 8, 2026

Copy link
Copy Markdown
Member

Offline database compaction for nethermind and besu, a journal check for geth, and a CLI that runs it all outside a run. The motivation is the post-Gloas shadowfork images: benchmarks found the contracts deployed after the snapshot in higher DB tiers, and for geth still in the journal, so they read faster than real mainnet state would. Every image is cleaned with exactly the commands db_compaction runs.

1. nethermind and besu: RocksDB ldb

Neither client ships an offline compactor (besu storage has none as of 26.6.1; trie-log prune deletes rather than rewrites). Both now return DBMaintenanceCommands, so db_compaction (before_pre_runs / before_benchmarks) works for them:

  • Compaction runs RocksDB 11.8.1 ldb (the version nethermind binds; besu's rocksdbjni 10.6.2 databases open with it) from a separate image, ghcr.io/ethpandaops/benchmarkoor-rocksdb-ldb:11.8.1, built DEBUG_LEVEL=0 PORTABLE=1 from the new Dockerfile.rocksdb-ldb and published by deploy-docker.rocksdb-ldb.yaml.
  • Every directory under the datadir with an OPTIONS-* file is a database (nethermind: one per store, ~13 on a fresh mainnet datadir; besu: one, with a column family per segment). Each column family is compacted with --try_load_options, i.e. the options the client wrote it with.
  • list_column_families prints a header line, then {a, b, c} with names raw. Besu names its families by single bytes, and two of them are 0x09 and 0x0a, so the parser splits on ", " only and never trims.
  • DBMaintenanceCommands gains CompactImage: Compact runs in that image with Compact[0] as entrypoint, while inspect and prepare keep the client image. db_compaction.image overrides it, and an image other than the instance's is pulled if-not-present (that was missing for db_compaction.image too).

Two ldb behaviours would otherwise turn a wrong compaction into a green one:

  • ldb opens a family without a merge operator with a string-append one. A database whose OPTIONS names a merge operator ldb lacks (nethermind's log index, off by default, has one) is refused rather than compacted with the wrong merge.
  • ldb compact --column_family=<unknown> compacts the default family and exits 0. After each family the script requires rocksdb.num-files-at-level0 to be 0 through get_property, which does fail on an unknown name.

2. geth: the journal check

A clean stop writes the 128 diff layers and the write buffer below them (up to 256 MiB) to geth/triedb/merkle.journal, and the next start loads both into memory. The 2026-10-08 image journaled 1,354 blocks this way. After the compaction the runner now parses the journal, following triedb/pathdb/journal.go at v1.17.7 (version 3, disk layer then diff layers, with the optional node origins told apart by rlp.Kind as geth does). It fails if the buffer is non-empty or there are more than 128 diff layers. Any other journal version is an error, not a guess. db_compaction.verify: false skips it; a volume datadir is not checked (the runner can't read it).

3. benchmarkoor db compact

benchmarkoor db compact --client <geth|nethermind|besu|erigon|reth|ethrex> --datadir <dir> [--prepare <step>]...

Runs the phase's own runDBCompactionContainers (inspect, prepare, compact, verify, inspect) against a stopped client's host datadir. It exits non-zero on any failure, the journal check included. reth: no-op with a log line (MDBX, persisted every block). ethrex: error. Optional --image, --compaction-image, --extra-arg, --timeout, --no-inspect mirror the config. No marker is written.

Verified

  • Real nethermind 2.1.0 and besu 26.9.0 datadirs (mainnet genesis, no peers): compacted by the CLI (13 nethermind databases / 25 families; besu's 17 families including 0x09, 0x0a), then each client reopened them and served the genesis hash 0xd4e5…8fa3 and a genesis balance over RPC.
  • Real geth v1.17.7 journals (--dev): geth logged layers=25 → check passes with 25 diff layers; layers=150 → fails, 128 diff layers (blocks 23–150) over a non-empty buffer; after --cache.gc=0 (buffer=0.00B) layers=128 → passes, 128 diff layers (213–340).
  • TestCompactDatadir_RealLdb (runs when BENCHMARKOOR_TEST_LDB_IMAGE is set): two real RocksDB databases with besu-style byte names and nethermind-style names, several level-0 files each. Every family ends with level 0 empty and identical scan output, and a second pass succeeds.
  • Unit tests: command construction; the script against a stub ldb for raw-byte family parsing, every database, extra args, unreadable database, non-empty level 0, foreign merge operator, ldb's own merge operator; the journal parser on journals encoded with geth's types (128/129 layers, dirty buffer, no origins, wrong version, truncated); runner tool image, entrypoint and pull; verify failure, verify: false, volume.
  • go test ./... passes; golangci-lint v2.3.0 (go 1.24) reports nothing new. Two issues it reports are already on master: pkg/docker, pkg/podman.

Not verified

  • Not run against a synced nethermind or besu datadir, or one with real chain history: only genesis state. Per-family timings on a mainnet-sized database are unmeasured.
  • The ghcr.io/…-rocksdb-ldb:11.8.1 image doesn't exist until the new workflow runs on master. Until then, set db_compaction.image / --compaction-image to a local build of Dockerfile.rocksdb-ldb.

Behaviour change

verify defaults to true. A geth db_compaction whose datadir was stopped with a non-empty write buffer, which is the default geth cache config after any pre-run, now fails where it used to pass. That is what the check is for, but existing geth configs need --cache.gc=0 on the run before the stop, or verify: false.

qu0b added 4 commits October 8, 2026 16:35
Neither client ships an offline compaction command, so db_compaction now
runs RocksDB's ldb (11.8.1, Dockerfile.rocksdb-ldb) against the stopped
datadir: every directory with an OPTIONS-* file is a database, and each of
its column families is compacted with --try_load_options.

- DBMaintenanceCommands.CompactImage runs Compact in a tool image, with
  Compact[0] as the entrypoint; inspection and prepare steps keep the
  client image. db_compaction.image overrides either, and an image other
  than the instance's is pulled if not present.
- list_column_families prints raw names after a header line, and besu's
  are single bytes including 0x09 and 0x0a, so the script splits on ", "
  only and never trims.
- A database whose OPTIONS names a merge operator ldb lacks is refused,
  since ldb would substitute a string-append operator; and because
  `ldb compact` silently falls back to the default family for an unknown
  name, each family must show an empty level 0 afterwards.
…8 blocks

geth's path scheme writes its 128 diff layers and the unflushed write
buffer below them to geth/triedb/merkle.journal on a clean stop, and loads
them back into memory on the next start, so that state is never read from
the database a compaction rewrote.

After the compaction the runner now parses the journal (layout version 3,
triedb/pathdb/journal.go in geth v1.17.7) and fails the phase when the
write buffer is not empty or there are more than 128 diff layers. Another
journal version is an error rather than a guess. DBMaintenanceCommands
gains Verify for this; db_compaction.verify: false skips it, and a volume
datadir, which the runner cannot read, is not checked.

Checked against three journals written by geth v1.17.7 --dev: layers=25
and layers=128 (after --cache.gc=0) pass with 25 and 128 diff layers;
layers=150 fails with 128 diff layers over a non-empty buffer.
… a run

    benchmarkoor db compact --client <geth|nethermind|besu|erigon|reth|ethrex> \
      --datadir <dir> [--prepare <step>]...

runs the same containers and checks as a db_compaction phase
(runner.CompactDatadir drives runDBCompactionContainers) against a stopped
client's host datadir, and exits non-zero on any failure, the geth journal
check included. reth is a no-op with a log line (MDBX, persisted every
block); a client without offline compaction is an error. --image,
--compaction-image, --extra-arg, --timeout and --no-inspect mirror the
db_compaction settings. No marker is written: that belongs to a run.

TestCompactDatadir_RealLdb runs it through the docker daemon against real
RocksDB databases when BENCHMARKOOR_TEST_LDB_IMAGE names an ldb image.
…ult in db compact

A node stopped with a non-empty write buffer (any default-cache run) fails the
check, so defaulting it on in db_compaction would break existing configs.
benchmarkoor db compact keeps it on (--no-verify skips it).
…d after compaction

A 52-minute compaction of besu's 1.1 TB jochemnet database logged only the
family names and a total duration: nothing in the output showed whether the
database changed. The level 0 check now reads the same levelstats.

@redpandabot redpandabot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Adds offline RocksDB ldb compaction for nethermind/besu, a geth journal check, and a benchmarkoor db compact CLI, wired through a new CompactImage/Verify on DBMaintenanceCommands and a verify config key. The geth journal parser matches geth v1.17.7's journal layout and ldb's argv/column-family handling lines up with RocksDB 11.8.1, but the script's post-compaction level-0 guard is defeated by shell pipeline error masking, so one of the two documented safety checks cannot fail.

Issues

  • 🟡 pkg/client/rocksdb_compact.sh:74 — the level-0 guard can't fire: levels() swallows ldb errors — see the thread on that line

Reviewed @ cc9861ef
"Testing shows the presence, not the absence of bugs." — Edsger Dijkstra

# `compact` falls back to the default family for an unknown name and
# still exits 0; get_property does not, and an empty level 0 is what a
# finished full compaction leaves.
after=$(levels "$db" "$cf")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 the level-0 guard can't fire: levels() swallows ldb errors

levels() (lines 30-33) pipes ldb ... get_property into awk/sed, and POSIX gives a pipeline the status of its last command, so a failing get_property produces an empty string and a zero status — set -e never sees it. At line 74 after is then empty, which matches neither L0= nor the failure case, so the script prints compacted ... duration=Ns and exits 0. That is exactly the case the guard was added for: ldb compact --column_family=<unknown> silently compacts the default family and exits 0, then get_property fails for the unknown name, so a wrong compaction is reported green. I ran the actual script under dash with an ldb stub whose get_property fails for one listed family: it logged compacted ... column_family=bogus and exited 0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant