Skip to content

feat(snapshot): benchmarkoor snapshot export - #350

Draft
qu0b wants to merge 11 commits into
masterfrom
qu0b/snapshot-export
Draft

qu0b wants to merge 11 commits into
masterfrom
qu0b/snapshot-export

Conversation

@qu0b

@qu0b qu0b commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Adds benchmarkoor snapshot export, which publishes a stopped client datadir in the snapshots.ethpandaops.io layout so that ethereum-package's network_sync_base_url and the snapshot_fetcher role can read it unchanged. It replaces the publish.sh / verify.sh / blockhash.py / metadata.sh scripts the shadowfork image builds have used so far.

benchmarkoor snapshot export --client <c> --datadir <dir> --network <name> --block <n> --bucket <bucket> \
  [--prefix ""] --head-block-file <eth_getBlockByNumber full-tx JSON> [--metadata-file <json>] \
  [--write-latest] [--verify-public-base <url>] [--zstd-level 6] [--verify-blocks 3] [--archive-name <name>]

Credentials come from S3_ENDPOINT_URL, AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY. Full docs are in docs/snapshot-export.md.

1. Streamed, never a local archive

The command walks the datadir, drops the excludes and pipes the sorted member list through tar --null --no-recursion -T - | zstd -<level> -T0 straight into a multipart upload. Members are in tar --sort=name order, so the same datadir always gives the same archive.

  • Part size is derived from the datadir. It starts at 64 MiB and grows until the archive fits in 9,000 parts. The old fixed 64 MiB parts capped an object at 640 GiB; a 1.5 TB datadir now gets ~160 MiB parts. This is a new upload.UploadStream / upload.StreamPartSize in pkg/upload, built on the existing S3 client and manager rather than a second one.
  • A failed tar or zstd aborts the upload. This includes tar's "file changed as we read it" when the datadir is still running. The failure surfaces as a read error, so the multipart upload is aborted instead of completed with a truncated archive. Ctrl-C does the same.

2. Hashed and verified

  • The sha256 of every 1 GiB of the uploaded stream is recorded on the way through.
  • After the upload, the object's size must match the bytes sent, and N random 1 GiB ranges are re-read and compared against those hashes.
    • With --verify-public-base, the ranges are read through the public URL as consumers read them, and each must be a 206.
    • Without it, they are read through S3.
  • _snapshot_eth_getBlockByNumber.json (published byte-for-byte, so Amsterdam header fields survive; its result.number must equal --block) and _snapshot_metadata.json are written only after verification passes. latest is written only with --write-latest, and last.
  • The metadata is {"data_size_bytes": …} with --metadata-file merged over it, nested objects key by key, so shadowfork.amsterdam_time and docker_image pass through.

3. Per-client packing

client root excluded beyond nodekey, LOCK, nodes/, logs/, _snapshot_* archive
geth contents of <datadir>/geth, no prefix (snapshots.ethpandaops.io's layout; consumers extract into <datadir>/geth) snapshot.tar.zst
reth ./ discovery-secret, known-peers.json snapshot-v2.tar.zst
erigon ./ snapshot-pruned.tar.zst
besu ./ key snapshot.tar.zst
ethrex ./ node.key, node_config.json snapshot.tar.zst
nethermind contents of <datadir>/nethermind_db, no prefix (mainnet/ at the root, as jochemnet's tarball; consumers extract into <datadir>/nethermind_db) */peers, */discoveryNodes snapshot.tar.zst
  • The archive names are the ones the jochemnet and msf-2 inventories fetch. --archive-name overrides them, e.g. snapshot.tar.zst for reth to match snapshots.ethpandaops.io.
  • Excludes match from the archive root, so a database's own LOCK (e.g. geth/chaindata/LOCK) is kept.

The image gains tar and zstd.

Verified

  • Unit tests (make test-core):
    • the key layout and the archive names;
    • the excludes for every client;
    • packing order, symlinks and determinism;
    • a tar failure on an unreadable file;
    • block hashing across odd write sizes;
    • verification: size mismatch, a corrupted block named by index and offset, and an HTTP server that ignores Range (refused on its 200);
    • part sizing up to 40 TiB;
    • the head-block check and the metadata merge.
  • Integration test against a throwaway MinIO (make test-integration-core, added to CI). It uses cgr.dev/chainguard/minio, because minio/minio is no longer on Docker Hub. It checks that:
    • a 12 MiB geth datadir uploads as a three-part object (ETag -3);
    • the object downloads and extracts to the datadir minus the excludes, symlink included;
    • the head block is byte-identical and the metadata is merged;
    • latest is absent unless requested;
    • a byte flipped in the second part is caught by the range verification (block 6);
    • a tar failure mid-multipart leaves no object, no metadata, no latest and no open multipart upload.
  • The built binary against MinIO: a full export with env credentials and --write-latest; verification through an anonymous-read bucket via --verify-public-base; and fail-fast on a wrong --block, missing credentials and a public base that 403s.
  • golangci-lint v2.3.0 (go 1.24): 0 issues in the changed packages, integration tag included.
  • Nothing was uploaded to ethpandaops-shadowfork-images.

qu0b added 4 commits October 8, 2026 16:26
StreamPartSize grows the 64 MiB part size until the expected stream fits in
9,000 parts, so a >1.5 TB archive no longer hits the 640 GiB ceiling.
NewS3Client exposes the package's client so other packages reuse it.
Publishes a stopped datadir in the snapshots.ethpandaops.io layout: tar | zstd
streamed into a multipart upload, hashed every 1 GiB, verified by size and
random ranges (through a public base or S3), then the head block, metadata and,
with --write-latest, latest last. A failed tar or zstd aborts the upload.

@redpandabot redpandabot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

The PR adds benchmarkoor snapshot export, streaming a stopped client datadir through tar | zstd into an S3 multipart upload, hashing 1 GiB blocks, verifying random ranges, then writing the head-block/metadata/latest objects. The layout, packing, hashing and abort-on-truncation logic are well structured and tested. Two issues remain: Ctrl-C does not actually abort the multipart upload, and a negative --verify-blocks panics.

Issues

  • 🟡 pkg/upload/stream.go:57 — Ctrl-C leaves the multipart upload un-aborted — see the thread on that line
  • 🟢 pkg/snapshot/verify.go:91 — Negative --verify-blocks panics — see the thread on that line

Reviewed @ ad96ab38
"If debugging is the process of removing software bugs, then programming must be the process of putting them in." — Edsger Dijkstra

Comment thread pkg/upload/stream.go
u.Concurrency = concurrency
})

_, err := u.Upload(ctx, &s3.PutObjectInput{ //nolint:staticcheck // SA1019: successor is pre-v1

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Ctrl-C leaves the multipart upload un-aborted

UploadStream passes the CLI's cancellable context straight into manager.Uploader.Upload. On SIGINT/SIGTERM cmd/benchmarkoor/snapshot.go cancels that same context, and the SDK's failure path aborts via AbortMultipartUpload(u.ctx, ...) with the already-canceled context, so the abort request never reaches the store and the incomplete multipart upload is orphaned — the opposite of the documented "Ctrl-C does the same". Abort with a detached context (e.g. context.WithoutCancel, or the UploadID from the returned error) so cleanup survives the interrupt.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 26000da (with a test that fails without the fix).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was my finding, and 26000da addresses it. UploadStream now re-aborts on a context that outlives the canceled one (pkg/upload/stream.go:65-80):

if err != nil {
    var mu manager.MultiUploadFailure
    if errors.As(err, &mu) && mu.UploadID() != "" {
        abortCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), time.Minute)
        defer cancel()
        if _, aerr := client.AbortMultipartUpload(abortCtx, &s3.AbortMultipartUploadInput{...}); aerr != nil {
            err = errors.Join(err, fmt.Errorf("aborting upload %s: %w", mu.UploadID(), aerr))
        }
    }
    ...
}

I checked the pinned SDK (feature/s3/manager v1.22.5): its fail() aborts with u.ctx, i.e. the canceled context, and Upload returns an error implementing manager.MultiUploadFailure that carries the upload ID, so errors.As + mu.UploadID() is the correct recovery. The signal context from cmd/benchmarkoor/snapshot.go:106 does reach pkg/snapshot/export.go:120, so this is exactly the path that was broken — the finding was right and is now fixed.

The new TestUploadStream_CanceledAbortsUpload (pkg/upload/stream_test.go:38) cancels on the first part and asserts the abort DELETE for up-1. I couldn't execute it here (no Go toolchain in the checkout), but the control flow is direct and it targets the right request.

One small nit, not a blocker: the SDK's own fail() already attempts the abort first, so on a non-cancel failure where that abort succeeds, this second call can return NoSuchUpload and get joined into the returned error. Cleanup is unaffected; only the error text changes.


@ 26000daa

Comment thread pkg/snapshot/verify.go
return fmt.Errorf("object is %d bytes, uploaded %d", got, size)
}

for _, i := range rand.Perm(len(sums))[:min(n, len(sums))] {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Negative --verify-blocks panics

rand.Perm(len(sums))[:min(n, len(sums))] with a negative n (e.g. --verify-blocks -1) indexes the slice with a negative high bound and panics; VerifyBlocks is never validated in ExportOptions.Validate or the CLI. Clamp n to >= 0 or reject the flag value.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 26000da (with a test that fails without the fix).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 26000da. ExportOptions.Validate now rejects the value (pkg/snapshot/export.go:54-56):

if o.VerifyBlocks < 0 {
    return fmt.Errorf("verify blocks must be >= 0, got %d", o.VerifyBlocks)
}

Export calls it first thing (pkg/snapshot/export.go:73), before it reaches Verify at line 138, and the CLI path goes through Export (cmd/benchmarkoor/snapshot.go:109), so --verify-blocks -1 is now rejected rather than panicking. TestExportOptionsValidate (pkg/snapshot/layout_test.go:47-55) pins the negative case. Confirmed addressed.


@ 26000daa

…-verify-blocks

The SDK aborts on the upload's own context, which a Ctrl-C has already
canceled, so the abort never reached the store. Abort again on a context that
outlives it. A negative --verify-blocks panicked in Verify.

@redpandabot redpandabot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adds benchmarkoor snapshot export: it walks a stopped datadir, streams a deterministic tar|zstd straight into a multipart S3 upload (part size derived from the datadir), hashes 1 GiB blocks, verifies size plus random ranges through S3 or a public URL, then writes the head-block/metadata and optional latest. The follow-up commit validates a negative --verify-blocks and retries a Ctrl-C-canceled multipart abort on a detached context; both previous findings are fixed and I found nothing new.


Reviewed 17 changed file(s) @ 26000daa — no blocking issues found.
"If debugging is the process of removing software bugs, then programming must be the process of putting them in." — Edsger Dijkstra

qu0b added 2 commits October 8, 2026 17:47
snapshots.ethpandaops.io publishes geth (mainnet and jochemnet alike) as the
contents of <datadir>/geth, and consumers extract into <datadir>/geth. A
geth/-prefixed archive extracted that way nests a second geth/, and extracted
flat geth opens chaindata/ but misses triedb/merkle.journal and rewinds its
head (found by the e2e prototype on the jochemnet geth snapshot).
jochemnet's nethermind tarball has mainnet/ at its root and every shadowfork
inventory extracts it into <datadir>/nethermind_db; the CI hosts hold the
same tree under <dir>/nethermind_db/mainnet, so packing the datadir would
have nested a second nethermind_db/ (found by the e2e prototype).
redpandabot[bot]

This comment was marked as outdated.

qu0b added 2 commits October 8, 2026 21:14
A 1.2 TB export logged one line at its start and nothing for the next hours
(found by the e2e prototype: geth's export, ~120 MB/s compressed, ~2.5 h).
Also gofmt pack.go.
The published geth image carried .snapshot_fetcher_started and
download_snapshot.sh from the source tarball (found by the e2e prototype).

@redpandabot redpandabot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

This PR adds benchmarkoor snapshot export, a streamed tar|zstd-to-multipart-S3 publisher with random-range verification, plus the upload/verification helpers and tests behind it. The two previously-flagged issues (Ctrl-C leaving the multipart upload open, and a negative --verify-blocks panicking) are now fixed, but one open finding remains: the docs example still passes the geth instance dir as --datadir. I also found one new silent-failure edge case when the archive root is a symlink.

Issues

  • 🟡 pkg/snapshot/pack.go:121 — a symlinked archive root yields a silently empty snapshot — see the thread on that line
  • 🟡 docs/snapshot-export.md:18 — example passes the geth instance dir as --datadir — see the thread on that line
  • 🟢 pkg/upload/stream.go:69 — re-abort on a path the SDK already aborted — The aws-sdk-go-v2 manager already calls AbortMultipartUpload before returning a MultiUploadFailure when the upload context is not canceled, so on every non-Ctrl-C failure (e.g. a tar read error) this second abort gets NoSuchUpload and joins it into the error. Harmless to the outcome, but it makes real failures noisier.

Reviewed @ c6540790
"If debugging is the process of removing software bugs, then programming must be the process of putting them in." — Edsger Dijkstra

Comment thread pkg/snapshot/pack.go

m := &Manifest{Dir: root}

err = filepath.WalkDir(root, func(p string, d fs.DirEntry, err error) error {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 a symlinked archive root yields a silently empty snapshot

os.Stat(root) follows a symlink and accepts a symlinked directory, but filepath.WalkDir uses os.Lstat on the root: a symlinked root is a non-directory DirEntry, so nothing is descended and m.Members ends up as just .. tar --no-recursion then emits only the ./ entry (I reproduced this: piping a lone . member produces an archive containing only ./), so an empty archive is uploaded, it passes size/range verification against itself, and latest is written. Resolve the root with filepath.EvalSymlinks (or walk with os.ReadDir) before building the manifest. Only triggers when --datadir (or <datadir>/geth, <datadir>/nethermind_db) is a symlink to the real directory.

qu0b added 2 commits October 11, 2026 12:29
ethrex keeps its chain-id-1 db at <datadir>/chain-1 (sst, blob, metadata.json
at its root). jochemnet's archive is that directory's contents and msf-2
extracts it into /data/ethrex/chain-1; packing the datadir would have nested a
second chain-1/ (the same class as geth's and nethermind's roots, found on the
2026-10-11 build).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant